Fault Diagnosis Method Based on Multi-Class Semi-Supervised Support Matrix Machine

Through the multi-classification semi-supervised support matrix machine model, the adaptive low-rank approximation algorithm and multi-class hinge loss terms are used to directly process matrix data, which solves the shortcomings of existing fault diagnosis methods in fault detection and cause identification, and achieves high-precision and robust fault diagnosis, reducing dependence on marking samples.

CN118711612BActive Publication Date: 2025-05-27TONGJI UNIV +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410741926.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-07
Publication Date
2025-05-27
Estimated Expiration
2044-06-07

AI Technical Summary

Technical Problem

The existing fault diagnosis methods have shortcomings in fault detection and cause identification, especially the topological damage caused by the processing of two-dimensional signals, and the application of semi-supervised learning in fault diagnosis is insufficient.

Method used

The multi-classified semi-supervised support matrix machine model is adopted, and the matrix data is directly processed through the adaptive low-rank approximation algorithm and multi-class hinge loss terms, low-rank information is extracted, and the spatial structure of the signal is maintained. The semi-supervised learning strategy is used to gradually improve the model performance.

Benefits of technology

It improves the accuracy and robustness of fault diagnosis, reduces the error caused by subjective judgment, maintains the internal structure of the data, enhances the understanding and discrimination ability of fault characteristics, and reduces the dependence on a large number of labeled samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118711612B_ABST
    Figure CN118711612B_ABST
Patent Text Reader

Abstract

The present invention relates to the field of fault diagnosis technology, and specifically to a fault diagnosis method based on a multi-classification semi-supervised support matrix machine. In the present invention, original sound signals are collected from a public database to generate sound waveform images, and each sound waveform image is converted into a feature matrix of a fixed size through downsampling and grayscale technology to generate samples, and the samples include labeled samples and unlabeled samples; an objective function including multi-class hinge loss terms and combined regularization terms is designed, and a multi-classification semi-supervised support matrix machine model is trained based on the samples; the multi-classification semi-supervised support matrix machine model is trained according to labeled samples and unlabeled samples, and the model is trained through labeled samples, and the model is used to predict unlabeled samples to obtain the probability that each sample belongs to each label, and the unlabeled samples with a probability higher than the threshold are screened out through a predefined confidence threshold θ, and the unlabeled samples are used as pseudo-labeled samples, and the model is iteratively repeated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of fault diagnosis, and particularly to a fault diagnosis method based on a multi-class semi-supervised support matrix machine. Background Art

[0002] Currently, many methods have been developed to diagnose faults, which can be mainly summarized as threshold-based strategies, expert system strategies, and signal processing-driven strategies. A method based on the switch motor current proposes to detect faults by setting empirical thresholds. However, since the thresholds are set empirically, they will be affected by subjective judgments. In addition, the threshold-based method can only achieve fault detection and cannot achieve the purpose of fault cause identification. The fault diagnosis method based on an expert system needs to rely on a large amount of expert experience, which may have a negative impact on the diagnosis result. With the development of signal processing technology, the signal processing-driven strategy provides the possibility to solve the above problems.

[0003] However, the existing methods still have some deficiencies: relying on the SL paradigm design, getting rid of this dependence requires a large number of labeled images, yet the fault diagnosis method based on the SSL concept is still lacking; the collected signals can naturally exhibit two-dimensional (2D) fault features, and the spatial relationship between adjacent pixels can enhance the fault diagnosis effect. However, the above methods must vectorize each feature (including manual features and deep features) extracted from the two-dimensional image, thus destroying the topological structure embedded in each image during the fault recognition process.

[0004] The features in matrix format can more effectively preserve the structural information of the signal in time and space. On the contrary, vectorization may lead to the collapse of the topological structure and the loss of structural information. The Support Matrix Machine (SMM) has been identified as a powerful matrix learning algorithm that can directly perform the classification step on two-dimensional matrix features, eliminating the vectorization step and maintaining the structured information of the data. Although it seems natural and feasible to split the multi-class classification task into a series of binary classification tasks through the one-versus-rest (OvR) or one-versus-one (OvO) strategy, there are also some obvious shortcomings in this solution. First, in the prediction stage, if the confidence scales of the binary classifiers are different from each other, it may lead to bias. Second, when using the OvR strategy, the large excess of negative samples may already cause the input sample distribution to be unbalanced; the process of training multiple binary classifiers will be very time-consuming. Third, using the nuclear norm to approximate the rank is suboptimal. The nuclear norm is the sum of the singular values, and the size of the singular value reflects its importance. Using the nuclear norm to approximate the rank of a matrix treats all singular values equally instead of adaptively selecting them, which greatly reduces its adaptability. In addition, in the field of matrix data classification, the idea of semi-supervised learning (SSL) has not received sufficient research attention. This strategy usually requires a large number of labeled samples to learn the model parameters. However, performing label annotation requires rich professional experience and domain knowledge, which is obviously time-consuming and laborious in practical engineering applications. Summary of the Invention

[0005] The purpose of the present invention is to provide a fault diagnosis method based on a multi-class semi-supervised support matrix machine to solve the problems proposed in the above background technology.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A fault diagnosis method based on a multi-class semi-supervised support matrix machine, and its method steps include:

[0007] S1. Collect the original sound signals from the public database to generate sound waveform images, and convert each sound waveform image into a feature matrix of a fixed size through downsampling and grayscale technology to generate samples, where the samples include labeled samples and unlabeled samples, and the labeled samples represent the data containing the true label results;

[0008] S2. Design an objective function including multi-class hinge loss terms and a combined regularization term, and train a multi-class semi-supervised support matrix machine model based on the samples, where an adaptive low-rank approximation algorithm is used to extract strongly correlated low-rank information from the samples;

[0009] S3. The multi-class semi-supervised support matrix machine model trains the model based on the labeled samples and unlabeled samples. The model is trained with the labeled samples, and then the trained model is used to predict the unlabeled samples to obtain the probability of each sample belonging to each label. The unlabeled samples with probabilities higher than the predefined confidence threshold θ are screened out and used as pseudo-labeled samples, which are added to the labeled samples as new labeled samples. Then the model is retrained based on the new labeled samples. By continuously iterating and repeating the above process until the prediction probabilities of the unlabeled samples are all lower than the predefined confidence threshold θ.

[0010] As a further improvement of this technical solution, the objective function is expressed as follows:

[0011] In the formula, is the regression parameter in tensor form, and the Frobenius norm is a regularization term; is an adaptive low-rank regularization term, and an adaptive low-rank approximation algorithm is used to process the samples. is the multi-class hinge loss term, where τ of the adaptive low-rank regularization term and C in the multi-class hinge loss term are positive scalars that constrain the adaptive low-rank regularization term and the multi-class hinge loss term respectively; ξ represents the sequence of slack variables of the hinge loss. represents any label and x i , the difference in feature mapping between the true label y i is defined as follows: This feature mapping is a sparse tensor.

[0012] As a further improvement of this technical solution, the adaptive low-rank approximation algorithm is used to solve the adaptive low-rank regularization problem, and the adaptive low-rank approximation problem is constructed as follows:

[0013] where σ yi is the singular value of the i-th y in the formula, σ i is the singular value of the regression parameter W, ε is a sufficiently small positive number to ensure that |σ i | + ε is not zero; the solution to this adaptive low-rank approximation problem is:

[0014] where When the singular value σ yi is larger, the proximal operator prox τ (σ yi ) is closer to σ yi , when the singular value σyi The smaller it is, the closer the proximal operator prox τ (σ yi ) is to zero; the adaptive low-rank regularization term retains the large singular values of the regression parameter W and discards the small singular values as irrelevant outliers..

[0015] As a further improvement of this technical solution, the objective function introduces the tensor form of the regression parameter and the adaptive low-rank regularization term to process the feature matrix for classification, so as to avoid feature vectorization.

[0016] As a further improvement of this technical solution, the multi-class hinge loss term is represented by expanding the margin rescaling loss, which is used to support samples in matrix form.

[0017] As a further improvement of this technical solution, the combined regularization term is the square of the Frobenius norm of the model parameters and the adaptive low-rank regularization term of the matrix-type hyperplane extracted from the model parameters.

[0018] As a further improvement of this technical solution, the predefined confidence threshold θ is set to be above 0.5.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0020] 1. The fault diagnosis method based on the multi-class semi-supervised support matrix machine proposes a new multi-class semi-supervised support matrix machine model for multi-class fault diagnosis in matrix form. By directly using the spatial information of the matrix, through the use of the adaptive low-rank approximation algorithm and the multi-class hinge loss term, while retaining the spatial structure characteristics of the original sound signal, it effectively extracts the key low-rank information, thereby enhancing the accuracy and robustness of fault diagnosis. This technology avoids the limitations of simply judging the signal by threshold or relying too much on expert experience in traditional methods, thus reducing the bias caused by subjective judgment. At the same time, by directly processing matrix-type data without vectorization, this method better preserves the internal structure of the data, enabling the model to better capture the complex interrelationships between fault features, and thus showing higher efficiency and accuracy in diagnosing different types of rail transit faults.

[0021] 2. The fault diagnosis method based on multi-class semi-supervised support matrix machine makes full use of a large number of unlabeled sample data through the strategy of semi-supervised learning, and gradually improves the model performance through continuous iteration. In each iteration, the model screens out high-confidence unlabeled samples as pseudo-labeled samples through a predefined confidence threshold, continuously enriching the labeled sample set. This strategy significantly reduces the dependence on a large number of labeled samples, reduces the time and economic costs of the labeling work, and makes the model more scalable and practical in actual fault diagnosis applications. This model not only improves the diagnostic accuracy, but also gradually improves the understanding and discrimination ability of complex fault states through self-learning. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 It is a schematic diagram of the method steps of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0023] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0024] Please refer to Figure 1 , the present invention provides a technical solution: a fault diagnosis method based on multi-class semi-supervised support matrix machine, including the following method steps:

[0025] S1. Collect the original sound signals from a public database to generate sound waveform images, and convert each sound waveform image into a feature matrix of a fixed size through downsampling and grayscale technology to generate samples, where the samples include labeled samples and unlabeled samples;

[0026] S2. Design an objective function including multi-class hinge loss terms and combined regularization terms, and train a multi-class semi-supervised support matrix machine model based on the samples, where an adaptive low-rank approximation algorithm is used to extract strongly correlated low-rank information from the samples;

[0027] S3. The multi-class semi-supervised support matrix machine model trains the model according to the labeled samples and unlabeled samples. The model is trained with the labeled samples, and the model is used to predict the unlabeled samples to obtain the probability of each sample belonging to each label. The unlabeled samples with probabilities higher than this threshold are screened out through a predefined confidence threshold θ and used as pseudo-labeled samples, which are added to the labeled samples as new labeled samples, and the model is retrained according to the new labeled samples. The above process is continuously iterated until the prediction probabilities of the unlabeled samples are all lower than the predefined confidence threshold θ.

[0028] Specifically, it includes:

[0029] S1. Collect the original sound signals from a public database, generate sound waveform images from the original sound signals, and convert each sound waveform image into a feature matrix of a fixed size through downsampling and grayscale techniques. Specifically, reduce the sampling rate of the signal by reducing the number of sampling points in the signal, convert each sampling point into a grayscale value, map the intensity or amplitude of the sound waveform to different grayscale levels, and finally adjust it to a matrix of a fixed size, such as a feature matrix of 45×60; this step does not require complex signal processing calculations and does not damage the spatial structure of the current curve; use the collected feature matrix as a sample, and the sample includes labeled samples and unlabeled samples, and the labeled samples represent the data containing the true label results.

[0030] The representation form of the sample is as follows:

[0031] A matrix-type sample of k classes (k≥2): Where is the i-th feature matrix in the sample, and y i ∈{1,2,3,...,k} is the i-th true label in the sample.

[0032] S2. Design an objective function that includes multi-class hinge loss terms and a combined regularization term, and train a multi-class semi-supervised support matrix machine model based on the sample. The representation form of the objective function is as follows:

[0033] In the formula, is the regression parameter in the form of a tensor, and the Frobenius norm is a regularization term used to avoid overfitting and make the model follow the principle of minimizing the structural risk; is an adaptive low-rank regularization term, which uses an adaptive low-rank approximation algorithm to process the sample and is used to extract strongly correlated low-rank information. The adaptive low-rank approximation problem can be constructed as: where σ yi is the singular value of the i-th y in the formula, σ i is the singular value of the regression parameter W, and ε is a sufficiently small positive number to ensure that |σ i |+ε is not zero; the solution to this adaptive low-rank approximation problem is:

[0034] where When the singular value σ yi is larger, the proximal operator prox τ (σ yi ) is closer to σyi , when the singular value σ yi is smaller, the proximal operator prox τ (σ yi ) is closer to zero. Since larger singular values are related to the strongly correlated information of the matrix, they are retained, while smaller singular values are related to irrelevant information and are discarded. In summary, the adaptive low-rank regularization term retains the large singular values of the regression parameter W and discards the small singular values as irrelevant outliers. Therefore, a more accurate weight matrix is obtained;

[0035] In addition, the multi-class hinge loss term is where τ of the adaptive low-rank regularization term and C in the multi-class hinge loss term are positive scalars that constrain the adaptive low-rank regularization term and the multi-class hinge loss term respectively; ξ represents the sequence of slack variables of the hinge loss; represents any label and the true label y i of x i The difference in feature mapping between them is defined as follows: This feature mapping is a sparse tensor.

[0036] In summary, in S2, in order to directly utilize the feature matrix information of the samples, the objective function consists of two parts: a multi-class hinge loss term and a combined regularization term that considers the matrix form data structure information; the multi-class hinge loss term is represented by extending the margin rescaled loss to support samples in matrix form; the combined regularization term is the square of the Frobenius norm of the model parameters and the adaptive low-rank regularization term of the matrix-type hyperplane extracted from the model parameters; the square Frobenius norm controls the complexity of the model and prevents overfitting problems during the training phase; in order to fully utilize the structural information of the matrix (i.e., the correlation between rows and columns), adaptive low-rank regularization is used to capture the global structure in the matrix data; the singular values of all hyperplanes are penalized, retaining the large singular values of the regression parameter W and discarding the small singular values as irrelevant outliers. Based on the basic assumption that the true model parameters are sparse in rank, such a combined regularization term obtains the optimal solution and fully encodes the structural information of the matrix, thereby obtaining better classification performance. Note that τ is a key parameter of the objective function, which determines how much penalty is added to the adaptive regularization, and thus implicitly reflects how much structural information of the matrix should be included in the classification. This model is constructed based on the regularized risk minimization framework.

[0037] In addition, by introducing these two parts, the tensor form of the regression parameter and the adaptive low-rank regularization term, the model is allowed to directly manipulate and learn the features of two-dimensional matrices (or higher dimensions) without flattening these features into one-dimensional feature vectors.

[0038] S3. The output form of the multi-class semi-supervised support matrix machine model is as follows:

[0039] The probability output corresponding to each sample is obtained through this formula. The probability output represents the confidence level of the matrix-type input belonging to its true label. A predefined confidence threshold θ is used to determine which predictions of unlabeled samples are reliable enough to be converted into pseudo-labeled samples. Usually, there is no standard fixed value for the setting of the predefined confidence threshold θ, which largely depends on the specific application scenario, dataset characteristics, and the initialization performance of the model. It is usually above 0.5. Since an output probability greater than 0.5 means that the model believes that the sample is more likely to belong to this class than other classes, 0.5 is a natural starting point. However, in actual applications, in order to ensure the high quality of the introduced pseudo-labeled samples, the threshold is usually set higher, such as 0.7, 0.8 or higher.

[0040] Use the labeled samples and unlabeled samples to fully train the multi-class semi-supervised support matrix machine model; the predicted label distribution of the unlabeled samples is consistent with that of the labeled samples. Therefore, select the unlabeled samples with the maximum probability output greater than the predefined confidence threshold θ to enrich the labeled samples, specifically including:

[0041] Train the multi-class semi-supervised support matrix machine model based on the labeled samples. Use the trained model to predict the unlabeled samples to obtain the probability distribution of each sample belonging to each label. This probability distribution reflects the confidence level of the model for each sample belonging to each label;

[0042] Filter the prediction results of the unlabeled samples according to the predefined confidence threshold θ. If the probability that a sample is predicted by the model to belong to a certain label is higher than the predefined confidence threshold θ, the sample is considered credible. Regard this predicted label as a pseudo-label and add the sample to the labeled sample set. By adding the samples with pseudo-labels obtained through confidence screening to the original labeled samples as new labeled samples, the amount of labeled samples is amplified;

[0043] The new labeled samples include the original labeled samples and the pseudo-labeled samples; retrain the model according to the new labeled samples; after training is completed, use this model again to perform probability output on the unlabeled samples (the remaining ones and those with confidence levels lower than the threshold), and repeat the steps of confidence screening and sample update;

[0044] This process iterates. As new unlabeled samples are added to the labeled dataset, the model is updated and improved accordingly; this loop continues until there are no more unlabeled samples that can be added (i.e., the prediction probabilities of the unlabeled samples are all lower than the threshold), or the set maximum number of iterations is reached.

[0045] In S3, a semi-supervised strategy is used. The semi-supervised strategy utilizes relatively limited labeled samples and a large number of unlabeled samples to ensure the fault diagnosis performance of the model, reduces the economic cost and time cost generated by the sample annotation work, and realizes fault diagnosis under the condition of a small number of sample labels.

[0046] The above shows and describes the basic principles, main features and advantages of the present invention. Those skilled in the art should understand that the present invention is not limited by the above embodiments. The above embodiments and the descriptions in the specification are only preferred examples of the present invention and are not used to limit the present invention. Without departing from the spirit and scope of the present invention, the present invention will have various changes and improvements, and these changes and improvements all fall within the scope of the present invention claimed. The scope of protection claimed by the present invention is defined by the appended claims and their equivalents.

Claims

1. A fault diagnosis method based on a multi-classification semi-supervised support matrix machine, characterized in that: The method steps are as follows: S1. Collect original sound signals from a public database to generate sound waveform images, and convert each sound waveform image into a feature matrix of a fixed size through downsampling and grayscale technology to generate samples. The samples include labeled samples and unlabeled samples, where labeled samples represent data containing true label results; S2. Design an objective function including a multi-class hinge loss term and a combined regularization term, and train a multi-classification semi-supervised support matrix machine model based on the sample, wherein an adaptive low-rank approximation algorithm is used to extract strongly correlated low-rank information from the sample, and regularization constraints are performed on parameters in the multi-classification semi-supervised support matrix machine model based on the low-rank information; S3. The multi-classification semi-supervised support matrix machine model is trained based on labeled samples and unlabeled samples. The model is trained through labeled samples, and the model is used to predict unlabeled samples to obtain the probability of each sample belonging to each label. The unlabeled samples with a probability higher than the predefined confidence threshold θ are screened out and used as pseudo-labeled samples. They are added to the labeled samples as new labeled samples, and the model is retrained based on the new labeled samples. The above process is repeated through continuous iteration until the predicted probability of the unlabeled samples is lower than the predefined confidence threshold θ.

2. The fault diagnosis method based on multi-classification semi-supervised support matrix machine according to claim 1 is characterized in that: The objective function is expressed as follows: In the formula, is the regression parameter in tensor form, the Frobenius norm is a regularization term; is an adaptive low-rank regularization term, using an adaptive low-rank approximation algorithm to process the sample, is a multi-class hinge loss term, where τ of the adaptive low-rank regularization term and C in the multi-class hinge loss term are positive scalars constraining the adaptive low-rank regularization term and the multi-class hinge loss term, respectively; ξ represents the slack variable sequence of the hinge loss; Represents any label With x i, The true value label y i The difference between the feature maps, defined as: The feature map is a sparse tensor.

3. The fault diagnosis method based on multi-classification semi-supervised support matrix machine according to claim 1 is characterized in that: The adaptive low-rank approximation problem approximated by the adaptive low-rank approximation algorithm is constructed as follows: where σ yi is the singular value of the ith y in the formula, σ i is the singular value of the regression parameter W, and ε is a small enough positive number to ensure that |σ i |+ε is not zero; the solution to this adaptive low-rank approximation problem is: in When the singular value σ yi The larger the value, the proximal operator prox τ (σ yi ) is closer to σ yi , when the singular value σ yi The smaller the value, the proximal operator prox τ (σ yi ) is closer to zero; Adaptive low-rank regularization term Large singular values ​​of the regression parameter W are retained, and small singular values ​​are discarded as irrelevant outliers.

4. The fault diagnosis method based on multi-classification semi-supervised support matrix machine according to claim 2 is characterized in that: The objective function introduces the tensor form of regression parameters and an adaptive low-rank regularization term to process the feature matrix for classification, so as to avoid feature vectorization.

5. The fault diagnosis method based on multi-classification semi-supervised support matrix machine according to claim 2 is characterized in that: The multi-class hinge loss term is formulated by extending the margin rescaling loss for samples in support matrix form.

6. The fault diagnosis method based on multi-classification semi-supervised support matrix machine according to claim 2 is characterized in that: The combined regularization term is the square of the Frobenius norm of the model parameters and an adaptive low-rank regularization term of a matrix-type hyperplane extracted from the model parameters.

7. The fault diagnosis method based on multi-classification semi-supervised support matrix machine according to claim 1 is characterized in that: The predefined confidence threshold θ is set to be greater than 0.5.

Citation Information

Patent Citations

  • Semi-supervised learning hydraulic reversing valve fault diagnosis method based on multi-sensor information

    CN115659256A

  • Semi-supervised gearbox fault diagnosis method based on infrared thermogram

    CN116612316A