Industrial fault detection method for adaptively estimating number of unknown fault categories

By constructing teacher models and student models, combining pruning and knowledge distillation technology, the problem of adaptive estimating the number of unknown fault categories is solved, efficient fault detection and clustering is achieved, and the detection accuracy of unknown fault categories is improved.

CN120372428APending Publication Date: 2025-07-25NANJING UNIV +1
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510469649.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-14
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The existing industrial fault detection methods are difficult to adaptively estimate the number of unknown fault categories, and cannot effectively utilize the clustering structure of unlabeled industrial data, resulting in inefficient detection and classification of unknown fault categories.

Method used

The teacher model is constructed and the number of unknown fault categories is estimated through pruning and knowledge distillation techniques, and the student model is constructed to classify and cluster fault categories using labeled and unlabeled industrial data training.

Benefits of technology

While saving manual labeling costs, accurately estimate the number of unknown fault categories, improving the detection and clustering accuracy of known and unknown fault categories.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120372428A_ABST
    Figure CN120372428A_ABST
Patent Text Reader

Abstract

The invention discloses an industrial fault detection method for adaptively estimating the number of unknown fault categories, and the method comprises the steps: obtaining to-be-processed industrial data which comprises a marked industrial data set and an unmarked industrial data set; and constructing and training a teacher model, and predicting whether the input industrial data is an unknown fault category by using the teacher model. And pruning unknown fault category classification heads which are not activated in the industrial data, calculating to obtain an estimated unknown fault category number, and constructing and training a student model based on the estimated unknown fault category number until convergence. During detection, to-be-detected industrial data is input into the student model, and the fault type of the to-be-detected industrial data can be predicted and output. According to the method, based on the parameterized classifier and through knowledge distillation, marked noise filtering and other technologies, known fault category classification, unknown fault category detection and clustering are effectively carried out, and meanwhile, the number of unknown fault categories can be more accurately estimated according to the activation degree of classification heads and the closeness degree of clustering.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an industrial fault detection method for adaptively estimating the number of unknown fault categories, and belongs to the field of power information technology. Background Art

[0002] The normal operation of industrial equipment such as power systems and mechanical devices is the basis of engineering production. Therefore, it is extremely important to efficiently and intelligently detect the faults of industrial equipment in real time to ensure the stability and efficiency of engineering production. At present, industrial fault detection methods based on deep learning and computer vision have received attention and gradually replaced traditional manual detection methods, becoming an indispensable part of building intelligent manufacturing enterprises. It automatically obtains real-time data of the operation of industrial production equipment through industrial vision equipment, or performs real-time detection by means of sensing signals of automatic detection equipment. The forms of industrial data used include forms such as industrial equipment images and industrial equipment sensing signals.

[0003] Industrial fault detection methods need to use a large amount of labeled industrial fault data as the basis for training fault detection models. The processing of unlabeled fault data faces challenges such as high labeling costs, difficult identification of unknown defects, and large data distribution differences. And labeling a large amount of industrial fault data requires high labor costs. Therefore, semi-supervised learning methods that can utilize both labeled and unlabeled industrial data are needed. Semi-supervised learning shows unique advantages in resource-constrained scenarios by balancing the utilization of labeled and unlabeled data. However, the current mainstream semi-supervised industrial fault detection methods can only detect known fault categories and unknown fault categories in unlabeled industrial data, but ignore the clustering structure of unlabeled industrial data. Then, transferring the features learned by the detection model from known fault categories to unknown fault categories of unlabeled industrial data can not only classify industrial data of known fault categories, but also cluster industrial data of unknown fault categories, which is more conducive to the analysis and repair of unknown fault categories by industrial production engineers.

[0004] In addition, traditional unknown fault detection methods are divided into parametric methods and non-parametric methods. Non-parametric methods can automatically estimate the number of unknown fault categories, but their classification performance is not as good as that of methods based on parametric classifiers. And the methods based on parametric classifiers rely on the prior knowledge of the number of unknown fault categories, which is difficult to obtain in real scenarios. Therefore, how to design a method that can adaptively estimate the number of unknown fault categories and at the same time has excellent fault detection performance is an urgent problem to be solved. Summary of the Invention

[0005] The content part of this application is used to briefly introduce concepts, which will be described in detail in the specific implementation part later. The content part of this application is not intended to identify the key features or essential features of the claimed technical solution, nor is it intended to limit the scope of the claimed technical solution.

[0006] Aiming at the problems and deficiencies in the prior art, the purpose of the present invention is to provide an industrial fault detection method for adaptively estimating the number of unknown fault categories. A teacher model is constructed and trained using labeled industrial data and unlabeled industrial data until convergence. Whether the industrial data is an unknown fault category is predicted by the teacher model, and pruning operations are performed on the results predicted by the teacher model. A student model is constructed and trained using the pruned results until convergence. Finally, the fault category of the industrial data is output by the student model. The present invention can accurately and effectively classify industrial fault data of known fault categories and detect and cluster industrial fault data of unknown fault categories. It is used to solve the problems raised in the above background technology.

[0007] To achieve the above purpose, the present invention provides the following technical solutions:

[0008] The present invention discloses an industrial fault detection method for adaptively estimating the number of unknown fault categories, including the following steps:

[0009] Step 1, in response to obtaining industrial data to be processed, including a labeled industrial data set and an unlabeled industrial data set;

[0010] Step 2, construct a teacher model and train it to convergence according to the labeled industrial data set and the unlabeled industrial data set;

[0011] Step 3, use the teacher model to predict whether the industrial data is an unknown fault category, and prune the classification head of the unknown fault category to obtain an estimated number of unknown fault categories;

[0012] Step 4, construct a student model and train it to convergence according to the estimated number of unknown fault categories;

[0013] Step 5, input the industrial data to be tested into the student model, and predict and output the fault category of the industrial data to be tested.

[0014] Preferably, the construction and training of the teacher model in step 2 include the following steps:

[0015] The construction and training of the teacher model in step 2 include the following steps:

[0016] Step 2.1, initialize the parameterized classifier, feature projector and parameterized classifier of the teacher model;

[0017] Step 2.2, augment the industrial data using an industrial data augmentation method to obtain weakly augmented industrial data;

[0018] Step 2.3, input the weakly augmented industrial data into a feature extractor to obtain a vector embedding;

[0019] Step 2.4, input the vector embedding into a feature projector to obtain a feature vector;

[0020] Step 2.5, use the feature vector as a parameter to input into a parameterized classifier, and train the teacher model in combination with a loss function until convergence.

[0021] Preferably, the rule for training the teacher model using the loss function in Step 2.5 is:

[0022] Use cross-entropy loss for the labeled industrial data set and supervised contrastive loss for training

[0023] Use self-distillation cross-prediction loss for the unlabeled industrial data set and unsupervised contrastive loss for training

[0024] Use average entropy maximization regularization loss for the industrial data in the same training batch for training.

[0025] Preferably, Step 3 further includes the following steps:

[0026] Step 3.1, for the industrial data that has not been predicted by the teacher model as a certain unknown fault category, prune the unactivated unknown fault category classification heads;

[0027] Step 3.2, extract the remaining unknown fault category prototype vectors in the teacher model after pruning, and calculate the cosine distance between them and other unknown fault category prototype vectors;

[0028] Step 3.3, initialize a Gaussian mixture model, and fit the posterior probability of the Gaussian components through a maximization algorithm;

[0029] Step 3.4, according to the posterior probability, prune the unknown fault category classification heads again to obtain an estimated number of unknown fault categories.

[0030] Preferably, the construction and training of the student model in Step 4 include the following steps:

[0031] Step 4.1, initialize a feature extractor and a feature projector;

[0032] Step 4.2, initialize a parameterized classifier with the estimated number of unknown fault categories and the number of known fault categories;

[0033] Step 4.3, use the RandAugment data augmentation method to augment the industrial data to obtain strongly augmented industrial data;

[0034] Step 4.4, in each training round, input the strongly augmented and weakly augmented industrial data, calculate the clean threshold of the unlabeled industrial data, and screen out the clean industrial data set and the noisy industrial data set;

[0035] Step 4.5, train the student model on the parameterized classifier using the loss function until convergence.

[0036] Preferably, the rule for training the student model using the loss function in Step 4.5 is:

[0037] Use cross-entropy loss and supervised contrastive loss to train on the labeled industrial data set;

[0038] Use knowledge distillation loss and unsupervised contrastive loss to train on the unlabeled industrial data set;

[0039] Use mean entropy maximization regularization loss to train on the industrial data of the same training batch.

[0040] Preferably, the total loss of the teacher model is expressed as

[0041]

[0042] where λ represents the weight of the loss of the labeled industrial data in the teacher model, ω represents the weight of the regularization loss, represents the average of the predicted probability vectors p i of the industrial data in the k-th training batch and the predicted probability vectors of its weakly augmented industrial data, and k represents the training batch.

[0043] Preferably, the total loss of the student model is expressed as

[0044]

[0045] where τ n and τ m both represent temperature coefficients and τ n >τ m , λ s represents the weight of the loss of the labeled industrial data in the student model, is the predicted probability vector of the weakly augmented industrial data in the teacher model, represents the set of clean industrial data, represents the set of noisy industrial data.

[0046] As a second aspect of the present application, the present invention also discloses an electronic device, including:

[0047] at least one processor, and a memory communicatively connected to the at least one processor;

[0048] The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the steps of the above-mentioned processing method for quickly browsing common items and drawing files and merging and displaying multiple windows.

[0049] As a third aspect of the present application, the present invention also discloses a computer storage medium, on which a computer program is stored, characterized in that when the computer program is executed by a processor, it realizes the steps of the above-mentioned processing method for quickly browsing common items and drawing files and merging and displaying multiple windows.

[0050] Compared with the prior art, the beneficial effects of the present invention are:

[0051] The present invention provides an industrial fault detection method for adaptively estimating the number of unknown fault categories. In response to obtaining industrial data to be processed, it includes a labeled industrial data set and an unlabeled industrial data set. A teacher model is constructed and trained, and the teacher model is used to predict whether the input industrial data is an unknown fault category. Prune the unactivated unknown fault category classification heads in the industrial data, calculate the estimated number of unknown fault categories, and construct and train a student model based on the estimated number of unknown fault categories until convergence. When detecting, input the industrial data to be measured into the student model, and the fault category of the industrial data to be measured can be predicted and output. The method of the present invention is based on a parametric classifier and through technologies such as knowledge distillation and labeled noise filtering, effectively classifies known fault categories, detects unknown fault categories and clusters, and at the same time can more accurately estimate the number of unknown fault categories according to the activation degree of the classification head and the tightness of clustering. While saving a large amount of manual annotation costs, it adaptively estimates the number of unknown fault categories. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] The drawings constituting a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and descriptions of the drawings of this application are used to explain this application and do not constitute an improper limitation of this application.

[0053] In the drawings:

[0054] Figure 1 This is the main step connection diagram of the industrial fault detection method in the embodiment of the present invention;

[0055] Figure 2 This is the overall process connection diagram of the industrial fault detection method in the embodiment of the present invention;

[0056] Figure 3 This is the process connection diagram for the construction and training of the teacher model in the embodiment of the present invention;

[0057] Figure 4 This is the process connection diagram for the construction and training of the student model in the embodiment of the present invention;

[0058] Figure 5 This is the structural schematic diagram of the electronic device in the embodiment of the present invention. Detailed implementation manners

[0059] Hereinafter, embodiments of the present disclosure will be described in more detail with reference to the accompanying drawings. Although some embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. On the contrary, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present disclosure are only for exemplary purposes and are not used to limit the protection scope of the present disclosure.

[0060] In addition, it should be noted that for ease of description, only parts related to the relevant invention are shown in the drawings. Without conflict, the embodiments in the present disclosure and the features in the embodiments can be combined with each other.

[0061] The present invention discloses an industrial fault detection method for adaptively estimating the number of unknown fault categories. Hereinafter, the present disclosure will be described in detail with reference to the accompanying drawings and in combination with embodiments. Referring to Figure 1 and Figure 2 as shown, it mainly includes the following steps:

[0062] Step 1, in response to obtaining industrial data to be trained, including a labeled industrial data set and an unlabeled industrial data set;

[0063] Step 2, construct a teacher model and train it to convergence according to the labeled industrial data set and the unlabeled industrial data set;

[0064] Step 3, use the teacher model to predict whether the industrial data is an unknown fault category, and prune the classification head of the unknown fault category to obtain an estimated number of unknown fault categories;

[0065] Step 4, construct a student model and train it to convergence according to the estimated number of unknown fault categories;

[0066] Step 5: Input the industrial data to be measured into the student model, and predict and output the fault category of the industrial data to be measured.

[0067] First, we need to obtain the industrial data x to be processed i , and use the model pre-trained from the industrial dataset as a feature extractor to extract industrial data, including the labeled industrial dataset and the unlabeled industrial dataset. The categories in the labeled industrial dataset are called known fault categories, and the unlabeled industrial dataset contains categories not seen in the labeled industrial dataset, which are called unknown fault categories. Specifically, the industrial data x i takes forms such as industrial equipment image data, industrial equipment sensing signals, etc., and the feature extractor is denoted as f(·|θ f ). The labeled industrial dataset is denoted as where there are N l labeled industrial data, and the category space is The unlabeled industrial dataset is denoted as where there are N u unlabeled industrial data, and the category space is The category space of the known fault categories is Then the category space of the unknown fault categories is

[0068] Next, construct and train the teacher model. The teacher model includes a feature extractor, a feature projector, and a parameterized classifier, as Figure 3 shown. For the construction and training of the teacher model, it specifically includes the following steps:

[0069] Step 2.1: Initialize the parameterized classifier, feature projector, and parameterized classifier of the teacher model;

[0070] Step 2.2: Use the industrial data augmentation method to augment the industrial data to obtain weakly augmented industrial data;

[0071] Step 2.3: Input the weakly augmented industrial data into the feature extractor to obtain vector embeddings;

[0072] Step 2.4: Input the vector embeddings into the feature projector to obtain feature vectors, which are used as the parameters of the parameterized classifier;

[0073] Step 2.5: Input the feature vectors into the parameterized classifier, and train the teacher model in combination with the loss function until convergence. Generally speaking, the feature extractor is responsible for extracting high-dimensional features from the original data, using the feature projector to map the high-dimensional features to a low-dimensional space to enhance the feature separability, and then using the parameterized classifier to make classification decisions based on the low-dimensional features. In the present invention, the number of heads of the parameterized classifier is set to the sum of the number of known fault categories and the number of unknown fault categories, and a relatively large number K of unknown fault categories needs to be initializedinit , then the number of heads of the parametric classifier is represented as K = K old + K init . Initialize the class prototypes The class prototypes of the parametric classifier usually manifest as an explicit expression of the model parameters for the class features. Among them, the set of known fault class prototypes is represented as The set of unknown fault class prototypes is represented as The parameters of the parametric classifier are the feature vectors within each set of fault class prototypes. Use the industrial data augmentation method to augment the industrial data x i to obtain weakly augmented industrial data The multi-layer perceptron for initializing the feature projector is represented as Input the industrial data x i into the feature extractor to obtain the vector embedding h of the industrial data i , and then project the input multi-layer perceptron of the vector embedding feature projector to obtain the feature vector z i , and this feature vector z i is used as the parameter input of the parametric classifier.

[0074] Furthermore, the loss function is used to train the labeled industrial dataset and the unlabeled industrial dataset respectively. That is, the cross-entropy loss and the supervised contrastive loss are used to train the model on the labeled industrial dataset, and the self-distillation cross-prediction loss and the unsupervised contrastive loss are used to train the model on the unlabeled industrial dataset. The average entropy maximization regularization loss is used to train the model for the industrial data in the same training batch. Specifically, it is necessary to calculate the predicted fault class probability of the teacher model for the industrial data. Taking the industrial data x i as an example, its predicted probability on the k-th fault class is expressed as:

[0075]

[0076] Among them, τ t is the temperature coefficient, p i represents the predicted probability vector, p i,k represents the predicted probability of the k-th fault class, z i represents the projected feature vector, k' represents the k'-th fault class, c k and c k' represent the prototype vectors of the k-th and k'-th fault classes respectively. For the labeled industrial dataset, calculate the cross-entropy loss on the parametric classifier and is expressed as:

[0077]

[0078] Among them, represents the training batch, represents the set of labeled industrial data in the training batch, represents the predicted probability vector after weak augmentation. And the supervised contrast loss is expressed as:

[0079]

[0080] Among them, represents the set of positive sample feature vectors of industrial data in the training batch, represents the set of positive samples of industrial data in the training batch, z j represents the feature vector of the j-th sample in the set of positive samples. The positive sample is defined as the industrial data with the same label as the industrial data x i having the same label.

[0081] And for the unlabeled industrial data set, the self-distillation cross-prediction loss calculated on the parameterized classifier is expressed as:

[0082]

[0083] Among them, represents the set of unlabeled industrial data in the training batch, p i represents the predicted probability vector, represents the predicted probability vector after weak augmentation. And the unsupervised contrast loss is expressed as:

[0084]

[0085] The average entropy maximization regularization loss is expressed as:

[0086]

[0087] Here, the predicted probability vector p of industrial data in the k-th training batch i and the average of its predicted probability vector of weakly augmented industrial data and, represents the training batch. It is used to calculate the average entropy of the training batch samples. Therefore, the total loss function for training the teacher model is expressed as:

[0088]

[0089] Here Where λ is the weight of the labeled industrial data loss, and ω is the weight of the regularization loss. Repeat the above steps for training until the teacher model converges, so that the model has certain classification and clustering capabilities, and a trained teacher model including a feature extractor, a feature projector, and a parameterized classifier can be obtained.

[0090] The teacher model has certain fault detection performance. However, the teacher model cannot adaptively estimate the number of fault types. In the following steps, the performance of the teacher model will be further improved by pruning the classification head and denoising knowledge distillation. As described in step 3, use the teacher model to predict whether the industrial data is an unknown fault category, and prune the classification head of the unknown fault category to obtain an estimate of the number of unknown fault categories. The following steps are also included:

[0091] Step 3.1, for the industrial data that the teacher model does not predict to belong to an unknown fault category, prune the unactivated classification head of the unknown fault category;

[0092] Step 3.2, extract the prototype vectors of the remaining unknown fault categories after pruning in the teacher model, and calculate the cosine distance between them and the prototype vectors of other unknown fault categories;

[0093] Step 3.3, initialize the Gaussian mixture model, and fit the posterior probability of the Gaussian components through the maximization algorithm;

[0094] Step 3.4, prune the classification head of the unknown fault category according to the posterior probability to obtain an estimate of the number of unknown fault categories. Specifically, first use the teacher model to predict whether the industrial data is an unknown fault category, and calculate the number of industrial data predicted to be each unknown fault category. If no industrial data is predicted to be an unknown fault category, prune the unactivated classification head of the unknown fault category. The set of pruned fault category prototypes is expressed as Where, represents the prototype vector of the k1-th unknown fault category, represents the data set predicted by the teacher model to be the k1-th unknown fault category, represents the predicted unknown fault category of the i-th sample by the teacher model. Extract the set of prototype vectors of the remaining unknown fault categories after pruning in the trained teacher model, which is expressed as Calculate the cosine distance between each unknown fault category type vector and the nearest prototype vector of other unknown fault categories. Taking the k-th unknown prototype as an example, it is expressed as

[0095] Initialize a Gaussian mixture model with two components. The probability density function of the Gaussian mixture model can be expressed as:

[0096]

[0097] Among them, is the probability density function of the Gaussian distribution, and π1 and π2 represent the weights of two Gaussian classifications. By fitting the distance calculated by the expectation-maximization algorithm, the posterior probability p(μ1|d k ) of each unknown fault category belonging to the Gaussian component with a smaller mean can be obtained. When the posterior probability is greater than the set threshold u, the classification heads of the unknown fault categories that meet the conditions are further pruned twice to obtain denoted as And after pruning, the final estimated number K of unknown fault categories can be obtained pruned , denoted as K init is the number of unknown fault categories. Among them, represents the prototype vector of the k2-th unknown fault category, represents the posterior probability that the k2-th unknown fault category belongs to the Gaussian component with a smaller mean, and represent the number of classification heads of unknown fault categories for the first pruning and the second pruning, respectively.

[0098] Next, we construct and train the student model. The student model also includes a feature extractor, a feature projector, and a parameterized classifier, as Figure 4 shown. Its training and construction specifically include the following steps:

[0099] Step 4.1, initialize the feature extractor and the feature projector;

[0100] Step 4.2, initialize the parameterized classifier with the estimated number of unknown fault categories and the number of known fault categories;

[0101] Step 4.3, use the RandAugment data augmentation method to augment the industrial data to obtain strongly augmented industrial data;

[0102] Step 4.4, input the strongly augmented and weakly augmented industrial data in each training round, calculate the clean threshold of the unlabeled industrial data, and screen out the clean industrial data set and the noisy industrial data set;

[0103] Step 4.5, train the student model on the parameterized classifier using the loss function until convergence.

[0104] Re-initialize the parameterized classifier with the estimated number of unknown fault categories and the number of known fault categories obtained above, denoted as where K' = K pruned + K old is the estimated total number of categories. In addition, the feature extractor and the feature projector At this time, the model including the feature extractor, feature projector, and parameterized classifier is called the student model. The RandAugment data augmentation method is used to augment the industrial data to obtain a strongly augmented version of the industrial data. The strongly and weakly augmented versions of the unlabeled industrial data are input into the student model, and the clean threshold u i of each industrial data x i (l) is calculated, and finally the set of clean industrial data The clean threshold u i is specifically expressed as:

[0105] u i (l) = (1 - β)u i (l - 1) + βp i (l - 1), u i (0) = 0;

[0106] where l is the training epoch, the maximum prediction probability of the industrial data, and β is the weight of the moving average. And the set of clean industrial data selected in one training batch is expressed as:

[0107]

[0108] where, and are the clean thresholds of the weakly and strongly augmented versions respectively, and are the maximum prediction probabilities of the weakly and strongly augmented versions respectively. The set of noisy industrial data selected is expressed as

[0109] Furthermore, the cross-entropy loss and the supervised contrastive loss are used to train the student model on the labeled industrial dataset, and the knowledge distillation loss and the unsupervised contrastive loss are used to train the student model on the unlabeled industrial data, while the average entropy maximization regularization loss is used to train the student model on the industrial data in one training batch. Specifically, the target of the knowledge distillation loss is the prediction probability of the teacher model, and different temperature coefficients are given to the knowledge distillation losses of the clean industrial data set and the noisy industrial data set, which is specifically expressed as:

[0110]

[0111] where τ n and τ mare all expressed as temperature coefficients and τ n > τ m , λ s represents the weight of the labeled industrial data loss in the student model, is the predicted probability vector of the weakly augmented industrial data in the teacher model, represents the set of clean industrial data, represents the set of noisy industrial data. The calculation of the supervised contrast loss is expressed as The loss on the parameterized classifier is The average entropy maximization regularization loss is Then the total loss calculated on the student model is expressed as:

[0112]

[0113] Repeat the number of training epochs to obtain a student model with training convergence. Compared with the teacher model, the student model estimates the number of unknown fault categories more accurately, and also has higher classification and clustering accuracy for both known and unknown fault categories. In the test phase, the industrial data to be measured is input into the student model, and the fault category of the industrial data to be measured is determined according to the prediction result of its parameterized classifier.

[0114] To implement the above embodiments, the present application also discloses an electronic device. Referring to Figure 5 As shown, the electronic device 500 may include a processing device (such as a central processing unit, a graphics processing unit, etc.) 501, which may perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 602 or the program loaded from the storage device 508 into the random access memory (RAM) 503. In the RAM 503, various programs and data required for the operation of the electronic device 500 are also stored. The processing device 501, the ROM 502, and the RAM 503 are connected to each other through a bus 504. The input / output (I / O) interface 505 is also connected to the bus 504.

[0115] Generally, the following devices may be connected to the I / O interface 505: an input device 506 including, for example, a touch screen, a touch pad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 507 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a storage device 508 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 509. The communication device 509 may allow the electronic device 500 to communicate with other devices wirelessly or wiredly to exchange data. Although Figure 5 shows the electronic device 500 with various devices, it should be understood that it is not required to implement or have all the shown devices. Instead, more or fewer devices may be implemented or had. Figure 5Each block shown in the figure may represent a device, or multiple devices as required.

[0116] In particular, according to some embodiments of the present disclosure, the processes described above with reference to the flowcharts may be implemented as computer software programs. For example, some embodiments of the present disclosure include a computer program product, which includes a computer program carried on a computer storage medium, and the computer program includes program code for performing the methods shown in the flowcharts. In such some embodiments, the computer program may be downloaded and installed from a network through a communication device 509, or installed from a storage device 508, or installed from a ROM 502. When the computer program is executed by a processing device 501, the above-mentioned functions defined in the methods of some embodiments of the present disclosure are executed.

[0117] It should be noted that the computer storage medium in some embodiments of the present disclosure may be a computer-readable signal medium or a computer-readable storage medium or any combination of the two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of the computer-readable storage medium may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.

[0118] In some embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program, and the program can be used by or in combination with an instruction execution system, apparatus, or device. In some embodiments of the present disclosure, the computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries computer-readable program code. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. The computer-readable signal medium may also be any computer storage medium other than the computer-readable storage medium, and the computer-readable signal medium can send, propagate, or transmit a program for use by or in combination with an instruction execution system, apparatus, or device. The program code contained on the computer storage medium can be transmitted by any appropriate medium, including but not limited to: wires, optical cables, RF (radio frequency), etc., or any suitable combination of the above.

[0119] In some embodiments, the client and the server can communicate using any currently known or future-developed network protocol such as HTTP (Hyper Text Transfer Protocol), and can be interconnected with digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include local area networks ("LANs"), wide area networks ("WANs"), the Internet (e.g., the Internet), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.

[0120] The above computer storage medium can be included in the above electronic device; it can also exist separately without being assembled into the electronic device. The above computer storage medium carries one or more programs, and when the above one or more programs are executed by the electronic device, the electronic device can implement an industrial fault detection method for adaptively estimating the number of unknown fault categories.

[0121] Computer program code for performing the operations of some embodiments of the present disclosure can be written in one or more programming languages or combinations thereof. The above programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (e.g., by using an Internet service provider to connect through the Internet).

[0122] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in the flowchart or block diagram can represent a module, a program segment, or a part of code that contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the blocks may occur in a different order than that marked in the accompanying drawings.

[0123] For example, two consecutively represented boxes can actually be executed substantially in parallel, and they can sometimes also be executed in the reverse order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, as well as the combination of boxes in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions. The units described in some embodiments of the present disclosure can be implemented in software or in hardware. The described units can also be provided in a processor, and the names of these units do not, in some cases, constitute a limitation on the unit itself.

[0124] The functions described above can be performed at least in part by one or more hardware logic components. For example, without limitation, exemplary types of hardware logic components that can be used include: field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), and so on.

[0125] The above description is only some preferred embodiments of the present disclosure and an explanation of the technical principles applied. Those skilled in the art should understand that the scope of the invention involved in the embodiments of the present disclosure is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features having similar functions disclosed in the embodiments of the present disclosure.

Claims

1. An industrial fault detection method for adaptively estimating the number of unknown fault categories, characterized in that, It includes the following steps: Step 1, in response to obtaining industrial data to be processed, including a labeled industrial data set and an unlabeled industrial data set; Step 2, construct a teacher model and train it to convergence based on the labeled industrial data set and the unlabeled industrial data set; Step 3, use the teacher model to predict whether the industrial data is an unknown fault category, and prune the classification head of the unknown fault category to obtain an estimated number of unknown fault categories; Step 4, construct a student model and train it to convergence according to the estimated number of unknown fault categories; Step 5, input the industrial data to be measured into the student model, and predict and output the fault category of the industrial data to be measured.

2. An industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 1, characterized in that The construction and training of the teacher model in Step 2 include the following steps: Step 2.1, initialize the parameterized classifier, feature projector and parameterized classifier of the teacher model; Step 2.2, use the industrial data augmentation method to augment the industrial data to obtain weakly augmented industrial data; Step 2.3, input the weakly augmented industrial data into the feature extractor to obtain vector embeddings; Step 2.4, input the vector embeddings into the feature projector to obtain feature vectors; Step 2.5, use the feature vectors as parameters to input into the parameterized classifier, and train the teacher model to convergence in combination with the loss function.

3. An industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 2, characterized in that, The rule for training the teacher model using the loss function in Step 2.5 is: Using cross - entropy loss for the labeled industrial dataset and training with supervised contrastive loss Using self - distillation cross - prediction loss for the unlabeled industrial dataset and training with unsupervised contrastive loss Use the average entropy maximization regularization loss for the industrial data of the same training batch Train.

4. An industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 2, characterized in that, Step 3 further includes the following steps: Step 3.1, for the industrial data that is not predicted by the teacher model as a certain unknown fault category, prune the unactivated classification head of the unknown fault category; Step 3.2, extract the remaining prototype vectors of the unknown fault categories after pruning in the teacher model, and calculate the cosine distance between them and the prototype vectors of other unknown fault categories; Step 3.3, initialize the Gaussian mixture model, and fit and calculate the posterior probability of the Gaussian components through the maximization algorithm; Step 3.4, secondarily prune the classification head of the unknown fault category according to the posterior probability to obtain an estimated number of unknown fault categories.

5. An industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 4, characterized in that, The construction and training of the student model in Step 4 include the following steps: Step 4.1, initialize the feature extractor and the feature projector; Step 4.2, initialize the parameterized classifier with the estimated number of unknown fault categories and the number of known fault categories; Step 4.3, use the RandAugment data augmentation method to augment the industrial data to obtain strongly augmented industrial data; Step 4.4, in each training round, input the strongly augmented and weakly augmented industrial data, calculate the clean threshold of the unlabeled industrial data, and screen out the clean industrial data set and the noisy industrial data set; Step 4.5, train the student model to convergence using the loss function on the parameterized classifier.

6. The industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 5, characterized in that The rule for training the student model using the loss function in Step 4.5 is: Using cross-entropy loss for the labeled industrial dataset and supervised contrastive loss for training; Using knowledge distillation loss on unlabeled industrial datasets and unsupervised contrastive loss for training; Use the average entropy maximization regularization loss for the industrial data of the same training batch for training.

7. An industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 3, characterized in that: The total loss of the teacher model is expressed as Among them, λ represents the weight of the marked industrial data loss in the teacher model, and ω represents the weight of the regularization loss. represents the industrial data prediction probability vector p in the k-th training batch i and its weakly augmented industrial data prediction probability vector The average value of the sum, represents the training batch.

8. An industrial fault detection method for adaptively estimating the number of unknown fault categories according to claim 6, characterized in that: The total loss of the student model is expressed as Among them, τ n and τ m both represent the temperature coefficient and τ n > τ m , λ s represents the weight of the labeled industrial data loss in the student model, is the predicted probability vector of the weakly augmented industrial data in the teacher model, represents the set of clean industrial data, represents the set of noisy industrial data.

9. An electronic device, characterized in that, It includes: At least one processor, and a memory communicatively connected to the at least one processor; Instructions executable by the at least one processor are stored on the memory, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the steps of the method according to any one of claims 1 to 8.

10. A computer storage medium, on which a computer program is stored, characterized in that: When the computer program is executed by a processor, the steps described in any one of claims 1 to 8 are implemented.

Citation Information

Cited By

  • Equipment fault prediction method and device, equipment and medium

    CN121479522A