Cross-domain generalized zero sample industrial fault diagnosis method based on two-step discrimination strategy

Through a two-step discriminant strategy, the cross-domain generalized zero-sample industrial fault diagnosis method combined with deep feature extraction and unsupervised clustering, the sample scarcity and inter-domain differences in fault diagnosis in cross-domain scenarios are solved, and accurate identification of unseen faults and efficient application of models are achieved.

CN119989154APending Publication Date: 2025-05-13CHONGQING UNIV

Patent Information

Application Number
CN202510164526.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-14
Publication Date
2025-05-13

AI Technical Summary

Technical Problem

The existing cross-domain generalized zero-sample industrial fault diagnosis methods have problems in actual industrial applications such as scarcity of visible fault samples, neglecting inter-domain differences and lacking generalized zero-sample diagnostic capabilities, resulting in poor diagnostic performance of the model in cross-domain scenarios.

Method used

A cross-domain generalized zero-sample industrial fault diagnosis method adopts a two-step discriminant strategy, and a cross-domain alignment and accurate identification of no-see faults is achieved by building a deep feature extractor based on CNN and a fault detector based on Kullback-Leibler divergence, combining unsupervised clustering method and cross-entropy loss function.

Benefits of technology

It effectively overcomes the scarcity of unseen samples, alleviates the misclassification of unseen faults in cross-domain scenarios, and improves the application effect and generalization ability of the model in actual industrial scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119989154A_ABST
    Figure CN119989154A_ABST
Patent Text Reader

Abstract

The invention relates to a cross-domain generalized zero sample industrial fault diagnosis method based on a two-step discrimination strategy, and belongs to the technical field of industrial fault diagnosis, and the method comprises the following steps: S1, constructing a CNN-based depth feature extractor for feature extraction, and employing a feature alignment method based on a feature-level maximum mean difference strategy for cross-domain alignment; s2, a fault detector based on Kullback-Leibler divergence KLD is constructed to measure the difference between the two pieces of probability distribution; s3, further identifying the result of the step S2 by using a semantic similarity-based detector; and S4, the marked visible fault classification is evaluated by using the cross entropy loss Lc.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of industrial fault diagnosis, and relates to a cross-domain generalized zero-sample industrial fault diagnosis method with a two-step discrimination strategy. Background Art

[0002] Fault diagnosis is a key means to ensure safe and stable operation in modern industrial systems. In recent years, with the advancement of industrial technology, fault diagnosis technology based on machine learning and pattern recognition has gradually become mainstream. However, such methods usually rely on a large number of fault samples, which are identified and labeled by domain experts or technicians to train models. Since faults occur sporadically in industrial processes, relevant sample data is usually scarce, which not only increases the cost and complexity of data collection, but also leads to serious data shortage problems in the model training process, limiting the practical application effect of fault diagnosis models.

[0003] In this context, Generalized Zero-sample Fault Diagnosis (GZSFD) technology has become an effective way to solve the problem of sample scarcity by transferring knowledge or models obtained from visible faults to diagnose unseen faults. Traditional zero-shot methods mostly focus only on the classification of unseen faults, without considering the coexistence of known and unknown faults in actual scenarios, which leads to a significant decrease in the diagnostic performance of the model in generalized zero-shot scenarios. At the same time, the differences in data distribution between different industrial equipment and working conditions further increase the difficulty of realizing cross-domain generalized zero-shot fault diagnosis. Therefore, how to realize the generalized zero-shot task of industrial fault diagnosis in cross-domain scenarios and improve the generalization ability of the model has become the focus of technological development.

[0004] In actual industrial applications, operating conditions and equipment wear can cause distribution differences between the source domain and the target domain. To solve this problem, methods such as fault prototype adaptive networks, dual adversarial learning strategies, and unsupervised deep domain adaptation have been introduced into cross-domain fault diagnosis. Patent CN 113609569 B discloses a discriminative generalized zero-sample learning fault diagnosis method. By discriminating fault samples, the generalized zero-sample diagnosis task is decomposed into supervised learning and zero-sample learning tasks. Patent CN 114254677B discloses a cross-domain fault diagnosis method based on multi-task adversarial discriminant domain adaptation, which realizes domain-level distribution alignment through feature generators and domain discriminators. Patent CN 117708656 B discloses a rolling bearing cross-domain fault diagnosis method for a single source domain. The method generates multiple pseudo-domain samples by training a domain generation module with source domain samples, and improves the distribution difference between pseudo-domain samples and source domain samples and the distribution difference between pseudo-domain samples, simulates unknown target domains, and obtains unknown target domains from single domain generalization. Patent CN 115129029 B discloses an industrial system fault diagnosis method based on sub-domain adaptation dictionary learning. It uses labeled source domain operating condition data to train a fault classifier and performs transfer dictionary learning. The sub-domain differences between the source domain and the target domain are measured by LMMD distance.

[0005] However, most existing methods rely on a large number of labeled visible target samples to train the model, while in actual industrial applications, the number of visible target samples under new working conditions may be very limited. In addition, the distribution difference between the source domain and the target domain is often ignored in the context of GZSFD, which further limits the application potential of cross-domain fault diagnosis technology in real industrial scenarios.

[0006] The existing technologies have the following deficiencies in fault diagnosis tasks in cross-domain scenarios:

[0007] 1. Scarcity of visible fault samples: The training process usually requires a large number of visible fault samples. However, in actual industrial applications, visible target fault samples are often scarce, resulting in poor performance of the model in actual classification.

[0008] 2. Ignoring inter-domain differences: Existing methods usually ignore the distribution differences between the source domain and the target domain, which limits the generalization ability and practical application of fault diagnosis models.

[0009] 3. Lack of generalized zero-shot diagnosis capability: Most methods do not consider the problem of generalized zero-shot fault diagnosis in cross-domain scenarios, which limits their applicability in complex industrial fault diagnosis applications. Summary of the invention

[0010] In view of this, the purpose of the present invention is to provide a cross-domain generalized zero-shot industrial fault diagnosis method with a two-step discrimination strategy. Based on the conversion framework of the two-step identification strategy, the generalized zero-shot fault diagnosis GZSFD task is converted into a supervised domain adaptive fault diagnosis (SDAFD) task and a ZSFD task. Visible and unseen faults are distinguished from all test samples, and neighbor searches are performed in different spaces. The scarcity of unseen samples is effectively overcome, and the misclassification of unseen faults in cross-domain scenarios is reduced, thereby ensuring the application of the model in actual industrial scenarios.

[0011] In order to achieve the above object, the present invention provides the following technical solutions:

[0012] A cross-domain generalized zero-shot industrial fault diagnosis method with a two-step discrimination strategy includes the following steps:

[0013] S1: Construct a CNN-based deep feature extractor for feature extraction, and use a feature alignment method based on the feature-level maximum mean difference strategy for cross-domain alignment;

[0014] S2: Construct a fault detector based on Kullback-Leibler divergence KLD to measure the difference between two probability distributions;

[0015] S3: Use a semantic similarity-based detector to further identify the result of step S2;

[0016] S4: Using cross entropy loss L c Evaluate the classification of the marked visible faults.

[0017] Furthermore, in step S1, by superimposing these convolutional layers and maximum pooling layers in the CNN network model, a deep feature extractor f(·) is obtained to embed the original monitoring data into a feature vector, thereby extracting data features, and the fault prototype of the i-th visible fault for:

[0018]

[0019] Where m i is the number of samples of the i-th visible fault.

[0020] Furthermore, in step S1, the mean difference between the two distributions in the kernel space is quantified based on the feature-level maximum mean difference strategy, so that the knowledge obtained from the source domain can be transferred to the target domain. The feature-level maximum mean difference is defined as:

[0021]

[0022] where h is located in the reproducing kernel Hilbert space (RKHS) In RKHS, the function h is represented by the inner product of any point h(·); the mathematical expectation is calculated by the corresponding mean embedding function μ; the Gaussian kernel function is used to calculate the expected value. Instead of inner product

[0023] Furthermore, in step S2, the distribution difference is measured using a KLD-based fault detector:

[0024]

[0025] in

[0026] According to P in the kth visible fault category i and P j The value between the two distributions KLD k , set the threshold η of the kth visible fault category k It is expressed as:

[0027]

[0028] The function sort(·) is a descending sorting function;

[0029] For any given sample x i , according to its corresponding distribution q i The difference between the distribution P of the kth visible fault category and the distribution metric is expressed as:

[0030]

[0031] When η i >η k When the sample x i It is considered as no fault observed.

[0032] Furthermore, in step S3, the visible or invisible fault results detected in step S2 are further identified based on an unsupervised clustering method, and the cluster center is regarded as the fault prototype:

[0033]

[0034] Where m t is the number of samples in the tth cluster;

[0035] Based on the learned mapping, the predicted attribute matrix of the first detected fault is represented as:

[0036]

[0037] For the jth unseen fault, classification is performed based on the similarity between the predicted representation and the true semantics:

[0038]

[0039] in is the unknown attribute detected first, and d(·) is the Euclidean distance function.

[0040] Further, in step S4, for the visible fault category finally detected, the cross entropy loss L is used c The marked visible fault classification is evaluated and expressed as:

[0041]

[0042] The predicted attribute matrix of unseen faults is expressed as:

[0043]

[0044] Based on mapping learning, the feature weight W is expressed as;

[0045] A=[f(X s ),f(X du )]W

[0046] The final predicted fault label is expressed as:

[0047]

[0048] The beneficial effects of the present invention are:

[0049] Excellence: The industrial fault diagnosis method in a cross-domain scenario proposed in the present invention can perform domain alignment and out-of-distribution detection on data with distribution differences, effectively identify visible / unseen faults under different working conditions, and solve the problem of fault diagnosis under different working conditions.

[0050] Applicability: The method proposed in the present invention adopts measurement based on unsupervised clustering, which effectively avoids the requirement for labeled data. At the same time, in actual industrial production, fault samples under different working conditions may change, and there may be few fault samples under new working conditions, which makes it difficult to train a powerful model. The present invention can use data from other working conditions for model training, making the present invention more applicable to actual scenarios.

[0051] Other advantages, objectives and features of the present invention will be described in the following description to some extent, and to some extent, will be obvious to those skilled in the art based on the following examination and study, or can be taught from the practice of the present invention. The objectives and other advantages of the present invention can be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0052] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below in conjunction with the accompanying drawings, wherein:

[0053] Figure 1 Flow chart of the cross-domain generalized zero-shot industrial fault diagnosis method with a two-step discrimination strategy;

[0054] Figure 2 It is the fault semantic attribute matrix of mechanical rolling bearing. DETAILED DESCRIPTION

[0055] The following describes the embodiments of the present invention by specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention in a schematic manner, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0056] It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and thus the drawings only show components related to the present invention rather than being drawn according to the number, shape and size of components in actual implementation. In actual implementation, the type, quantity and proportion of each component may be changed arbitrarily, and the component layout may also be more complicated.

[0057] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0058] like Figure 1 As shown, the present invention provides a cross-domain generalized zero-sample industrial fault diagnosis method with a two-step discrimination strategy, and the specific process is as follows:

[0059] Step 1: Feature extraction and feature alignment

[0060] A deep feature extractor based on CNN is constructed to embed the original data into feature vectors to extract data features. A feature alignment method based on the feature-level maximum mean difference strategy is adopted to minimize the distance between the feature distributions of the source domain and the target domain, thereby aligning the representations of the source domain and the target domain.

[0061] In this embodiment, by superimposing these convolutional layers and maximum pooling layers in the CNN network model, a deep feature extractor f(·) is obtained, thereby embedding the original monitoring data into a feature vector. Since the output of the fully connected layer of CNN is the extracted feature, the fault prototype of the i-th visible fault for:

[0062]

[0063] Where m i is the number of samples of the i-th visible fault.

[0064] In the cross-domain scenario of GZSFD, there are distribution differences in the data collected under different working conditions or on different devices. Therefore, after feature extraction, a projection method based on the maximum mean difference at the feature level is used to quantify the mean difference between the two distributions in the kernel space, so that the knowledge obtained from the source domain can be transferred to the target domain. The maximum mean difference at the feature level can be defined as:

[0065]

[0066] where h is located in the reproducing kernel Hilbert space (RKHS) In RKHS, the function h can be expressed as the inner product of any point h(·). Therefore, the mathematical expectation can be calculated using the corresponding mean embedding function μ. At the same time, since the function h is essentially infinite-dimensional, the feature space of the kernel function has infinite dimensions. Therefore, the Gaussian kernel function can be used. Instead of inner product

[0067] Step 2: Unseen Fault Detection of GZSFD

[0068] Construct a visible / unseen fault detector. In the first step, the difference between the two probability distributions is measured by a KLD-based fault detector to identify visible / unseen faults. Based on the results of the first step, the visible / unseen faults are further identified by an unsupervised clustering method. In this embodiment, step 2 specifically includes the following steps:

[0069] 1) KLD-based fault detector measures the distribution difference:

[0070]

[0071] in

[0072] According to P in the kth visible fault category i and P j The value between the two distributions KLD k , the threshold η of the kth visible fault category can bek Expressed as:

[0073]

[0074] The function sort(·) is a descending sorting function.

[0075] For any given sample x i , according to its corresponding distribution q i The difference between the distribution P of the kth visible fault category and the distribution metric can be expressed as:

[0076]

[0077] When η i >η k When the sample x i It is considered as no fault observed.

[0078] 2) Based on the unsupervised clustering method, the visible / unseen fault results detected in the previous step are further identified. The cluster center can be regarded as the fault prototype:

[0079]

[0080] Where m t is the number of samples in the tth cluster.

[0081] The attribute matrix is Figure 2 As shown, based on the learning mapping, the predicted attribute matrix of the first detected fault is represented as:

[0082]

[0083] For the jth unseen fault, classification is performed based on the similarity between the predicted representation and the true semantics:

[0084]

[0085] in is the first detected unknown attribute, and d(·) is the Euclidean distance function.

[0086] Step 3: The GZSFD task in the cross-domain scenario is converted into the SDAFD task and the ZSFD task through a two-step detection strategy. For the visible fault category finally detected, the cross entropy loss L is used c The marked visible fault classification is evaluated and expressed as:

[0087]

[0088] The predicted attribute matrix of unseen faults is expressed as:

[0089]

[0090] Based on mapping learning, the feature weight W can be expressed as

[0091] A=[f(X s ),f(X du )]W

[0092] The final predicted fault label can be expressed as:

[0093]

[0094] The evaluation index is the average accuracy of visible faults and unseen faults Acc s and Acc u , the harmonic mean H 1 , and the overall accuracy H 2 :

[0095]

[0096] This embodiment experiments on mechanical rolling bearing fault data from the Case Western Reserve University (CWRU) Bearing Data Center, which is a representative benchmark dataset. It records acceleration vibration data under different motor loads from the drive-end bearing, resulting in different working conditions. In the CWRU dataset, there are three different motor loads: 0, 1, and 2 horsepower (hp). Correspondingly, the motor speeds are 1797, 1772, and 1750 revolutions per minute, respectively. For each operating condition, there are three fault conditions: inner race fault (IR), rolling ball fault (BF), and outer race fault (OR). For each fault, the fault diameters are 7, 14, and 21 mils, respectively. In this case, the fault samples from different operating conditions consist of a source domain and a target domain, and the GZSFD task in a cross-domain scenario is performed. The training samples only include 9 visible faults from 1 labeled source domain and 1 unlabeled target domain, as shown in Table 1.

[0097] Table 1

[0098]

[0099] The information of the three target domains of mechanical rolling bearings is shown in Table 2.

[0100] Table 2

[0101]

[0102] The comparison results in the mechanical rolling bearing dataset (CWRU) are shown in Table 3.

[0103] Table 3

[0104]

[0105]

[0106] This method is verified in the above-mentioned actual mechanical rolling bearing industrial case. The experimental results show that the generalized zero-sample industrial fault diagnosis method proposed in this invention can well handle the fault diagnosis task in cross-domain scenarios. In general, this method can meet the actual industrial fault diagnosis task.

[0107] In the above embodiments, the description's reference to "this embodiment" indicates that a particular feature, structure, or characteristic described in conjunction with the embodiment is included in at least some embodiments, but not necessarily all embodiments. Multiple occurrences of "this embodiment" do not necessarily all refer to the same embodiment.

[0108] In the above-described embodiments, although the invention has been described in conjunction with specific embodiments of the invention, many substitutions, modifications, and variations of these embodiments will be apparent to those of ordinary skill in the art based on the foregoing description. For example, other storage structures (e.g., dynamic RAM (DRAM)) may use the embodiments discussed. Embodiments of the invention are intended to encompass all such substitutions, modifications, and variations that fall within the broad scope of the appended claims.

[0109] This embodiment further provides a computer-readable storage medium on which a computer program is stored. When the program is executed by a processor, any one of the methods in this embodiment is implemented.

[0110] This embodiment also provides an electronic terminal, including: a processor and a memory;

[0111] The memory is used to store a computer program, and the processor is used to execute the computer program stored in the memory, so that the terminal executes any one of the methods in this embodiment.

[0112] The computer-readable storage medium in this embodiment can be understood by ordinary technicians in this field: all or part of the steps of implementing the above-mentioned method embodiments can be completed by hardware related to the computer program. The aforementioned computer program can be stored in a computer-readable storage medium. When the program is executed, the steps of the above-mentioned method embodiments are executed; and the aforementioned storage medium includes: ROM, RAM, magnetic disk or optical disk and other media that can store program codes.

[0113] The electronic terminal provided in this embodiment includes a processor, a memory, a transceiver and a communication interface. The memory and the communication interface are connected to the processor and the transceiver and complete communication with each other. The memory is used to store computer programs, the communication interface is used to communicate, and the processor and the transceiver are used to run computer programs so that the electronic terminal executes each step of the above method.

[0114] In this embodiment, the memory may include a random access memory (RAM), and may also include a non-volatile memory (non-volatile memory), such as at least one disk memory.

[0115] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0116] The present invention can be used in many general or special computing system environments or configurations, such as personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronic devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like.

[0117] The present invention may be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention may also be practiced in distributed computing environments where tasks are performed by remote processing devices connected through a communication network. In a distributed computing environment, program modules may be located in local and remote computer storage media, including storage devices.

[0118] Finally, it should be noted that the above embodiments are only used to illustrate the technical solution of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solution of the present invention can be modified or replaced by equivalents without departing from the purpose and scope of the technical solution, which should be included in the scope of the claims of the present invention.

Claims

1. A cross-domain generalized zero-sample industrial fault diagnosis method with a two-step discrimination strategy, characterized by: The following steps are involved: S1: Construct a CNN-based deep feature extractor for feature extraction, and use a feature alignment method based on the feature-level maximum mean difference strategy for cross-domain alignment; S2: Construct a fault detector based on Kullback-Leibler divergence KLD to measure the difference between two probability distributions; S3: Use a semantic similarity-based detector to further identify the result of step S2; S4: Using cross entropy loss L c Evaluate the classification of the marked visible faults.

2. The cross-domain generalized zero-sample industrial fault diagnosis method of the two-step discrimination strategy according to claim 1 is characterized by: In step S1, by superimposing these convolutional layers and maximum pooling layers in the CNN network model, a deep feature extractor f(·) is obtained to embed the original monitoring data into a feature vector, thereby extracting data features, and the fault prototype of the i-th visible fault for: Where m i is the number of samples of the i-th visible fault.

3. The cross-domain generalized zero-sample industrial fault diagnosis method of the two-step discrimination strategy according to claim 1 is characterized by: In step S1, the mean difference between two distributions in the kernel space is quantified based on the feature-level maximum mean difference strategy, so that the knowledge obtained from the source domain can be transferred to the target domain. The feature-level maximum mean difference is defined as: where h is located in the reproducing kernel Hilbert space (RKHS) In RKHS, the function h is represented by the inner product of any point h(·); the mathematical expectation is calculated by the corresponding mean embedding function μ; the Gaussian kernel function is used to calculate the expected value. Instead of inner product 4. The cross-domain generalized zero-sample industrial fault diagnosis method of the two-step discrimination strategy according to claim 1 is characterized by: In step S2, the distribution difference is measured using a KLD-based fault detector: in According to P in the kth visible fault category i and P j The value between the two distributions KLD k , set the threshold η of the kth visible fault category k It is expressed as: The function sort(·) is a descending sorting function; For any given sample x i , according to its corresponding distribution q i The difference between the distribution P of the kth visible fault category and the distribution metric is expressed as: When η i >η k When the sample x i It is considered as no fault observed.

5. The cross-domain generalized zero-sample industrial fault diagnosis method of the two-step discrimination strategy according to claim 1 is characterized by: In step S3, the visible or invisible fault results detected in step S2 are further identified based on an unsupervised clustering method, and the cluster center is regarded as the fault prototype: Where m t is the number of samples in the tth cluster; Based on the learned mapping, the predicted attribute matrix of the first detected fault is represented as: For the jth unseen fault, classification is performed based on the similarity between the predicted representation and the true semantics: in is the unknown attribute detected first, and d(·) is the Euclidean distance function.

6. The cross-domain generalized zero-sample industrial fault diagnosis method of the two-step discrimination strategy according to claim 1 is characterized by: In step S4, for the visible fault category finally detected, the cross entropy loss L is used c The marked visible fault classification is evaluated and expressed as: The predicted attribute matrix of unseen faults is expressed as: Based on mapping learning, the feature weight W is expressed as; A=[f(X s ),f(X du )]W The final predicted fault label is expressed as:

Citation Information

Patent Citations

  • A discriminative generalized zero-shot learning fault diagnosis method

    CN113609569B

  • Cross-domain fault diagnosis system and method based on multi-task adversarial discrimination domain adaptation

    CN114254677B

  • Industrial system fault diagnosis method and system based on sub-domain adaptation dictionary learning

    CN115129029B

  • A cross-domain fault diagnosis method for rolling bearings based on a single source domain

    CN117708656B

Cited By

  • Industrial safety cross-domain collaborative analysis system based on time series data and visual data

    CN121327446A

  • Industrial safety cross-domain collaborative analysis system based on time series data and visual data

    CN121327446B