Coupling fault diagnosis method and system

By constructing a model using a graph convolutional learning network and optimizing the loss function based on the similarity between the support set and the query set, the problem of coupled fault diagnosis with small samples is solved, and efficient fault identification under limited data is achieved.

CN120876931APending Publication Date: 2025-10-31NAT UNIV OF DEFENSE TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510861409.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-25
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing technologies struggle to achieve effective fault diagnosis under limited coupled fault marker signal data conditions, especially in non-widespread industries such as aerospace. How to perform coupled fault diagnosis under small sample conditions is a challenge that urgently needs to be addressed.

Method used

A graph convolutional learning network is used to construct a graph convolutional learning model. By constructing a graph representing the similarity relationship between known fault samples, the model is trained using a support set and a query set. Three similarity relationship metric structures and loss functions are designed to optimize the model to identify coupled fault signals.

Benefits of technology

With a small number of training samples, it can effectively identify the various labels of coupled faults, improving the model's generalization ability and diagnostic accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876931A_ABST
    Figure CN120876931A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a coupling fault diagnosis method and system, and the method comprises the steps: constructing a graph which represents the similarity relation between known fault samples for all known fault samples in a support set and a label corresponding to each known fault sample; constructing a graph convolutional learning model based on a graph convolutional learning network by adopting the graph; three similar relations are adopted as a measurement structure of a known fault sample, and a corresponding loss function is determined according to the measurement structure; training the graph convolutional learning network by using the support set to obtain an intermediate graph convolutional learning network model; training the intermediate graph convolutional learning network model by using a support set and an inquiry set to obtain a trained graph convolutional learning network model; and for any to-be-identified coupling fault signal, the trained graph convolution learning network model is adopted to output a predicted value of each type of labels, and the predicted value of each type of labels is used for representing whether the fault exists or not. And when only a small number of training samples exist, each label of the coupling fault can be identified.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of fault diagnosis, and more specifically to a coupled fault diagnosis method and system. Background Technology

[0002] With the increasing automation of industry, fault diagnosis technology based on the Predictive Health Management (PHM) framework is playing an increasingly important role. Due to the high integration of system components, coupled fault phenomena, which may involve multiple fault types, are becoming increasingly common in collected fault signals. Data-driven diagnostic methods have achieved good performance for this coupled fault mode and have become a mainstream research direction. These methods are typically based on machine learning theory and rely on a large amount of labeled fault signal data. However, obtaining sufficient labeled data is a costly task in practical applications. This is mainly due to two factors: first, when a system fault occurs, self-protection mechanisms quickly lead to component damage, thus limiting the number of fault signals that can be collected; second, the process of labeling data requires deep expertise, which is particularly expensive in non-traditional industries such as aerospace. Therefore, how to achieve effective fault diagnosis under limited coupled fault labeled signal data conditions—the so-called "small-sample coupled fault diagnosis" problem—has become a critical challenge that urgently needs to be addressed.

[0003] While numerous studies have been conducted on few-sample single-fault diagnosis in recent years, few researchers have focused on the few-sample problem of coupled faults. Representative few-sample single-fault diagnosis methods can be categorized into three types: the first type is data augmentation methods, which do not require additional datasets; the second type is transfer learning-based methods, which transfer models or data from cross-domain applications to few-sample tasks to achieve better results; and the third type is meta-learning-based methods, which construct different few-sample learning tasks from the source domain and attempt to learn fundamental differences. Despite the significant success of these methods, they all focus on the single-fault diagnosis problem.

[0004] Furthermore, since coupled fault data can be considered multi-label data, machine learning methods for multi-label small-sample learning primarily focus on image datasets and cannot be directly applied to fault diagnosis problems. Therefore, further in-depth research is needed on this challenging small-sample coupled fault diagnosis problem. Among the three types of small-sample single-fault diagnosis research methods, the last two require additional source data information, limiting their use in some specific scenarios. Therefore, we consider using the first, more general method to address this new problem.

[0005] In developing this invention, the applicant discovered four main challenges that need to be overcome for the problem of small-sample coupled fault diagnosis:

[0006] In small-sample coupled fault diagnosis, features may contain more than one fault type. Compared to single faults, the label space in coupled faults is much more complex. For example, if the label dimension of a single fault is c, then the number of possible fault combinations for each sample feature in coupled faults is 2^c. c According to machine learning theory, the larger the hypothesis space of a model, the more prone it is to overfitting. Therefore, it is more susceptible to overfitting than traditional single-fault few-shot learning problems, leading to lower test accuracy.

[0007] Without the help of source data, a model can only be learned using a limited number of training data features and labels. One feasible option is to utilize information from the query set, employing a semi-supervised approach to improve the model's generalization performance. However, effectively utilizing query set information is a challenge.

[0008] In few-shot learning, it is difficult to achieve optimal hyperparameters in the loss function. Choosing the optimal hyperparameters is another challenge.

[0009] In single-fault data, there is a one-to-one correspondence between features and labels. However, the correlation between features and labels in coupled faults is more complex; coupled fault data can be viewed as multi-label data. Therefore, how to utilize this one-to-many relationship is a challenging problem. Summary of the Invention

[0010] This invention provides a method for diagnosing coupled faults, which can solve the problem of diagnosing small-sample coupled faults in existing methods.

[0011] To achieve the above objectives, in a first aspect, embodiments of the present invention provide a method for diagnosing coupled faults, comprising:

[0012] Step 1: For all known fault samples in the support set and the label corresponding to each known fault sample, construct a graph representing the similarity relationship between the known fault samples. This graph is used as input to the graph convolutional learning network.

[0013] Step 2: Construct a graph convolutional learning model using the graph-based graph convolutional learning network;

[0014] Step 3: For the graph corresponding to the support set, three similarity relationships are used as the metric structure for known fault samples, and the corresponding loss function is determined based on the metric structure;

[0015] Step 4: In the first stage, the graph convolutional learning network is trained using the support set to obtain an intermediate graph convolutional learning network model; in the second stage, the intermediate graph convolutional learning network model is trained using the support set and the query set to obtain the trained graph convolutional learning network model.

[0016] Step 5: For any coupled fault signal to be identified, use the trained graph convolutional learning network model to output the predicted value of each label. The predicted value of each label is used to indicate whether the fault exists.

[0017] Secondly, embodiments of the present invention provide a coupled fault diagnosis system, comprising:

[0018] The graph construction unit is used to construct a graph representing the similarity relationship between known fault samples for all known fault samples in the support set and the label corresponding to each known fault sample. The graph is used as input to the graph convolutional learning network.

[0019] Model construction unit, used to construct a graph convolutional learning model using the graph-based graph convolutional learning network;

[0020] The metric structure construction unit is used to use three similarity relationships as the metric structure for known fault samples for the graph corresponding to the support set, and to determine the corresponding loss function based on the metric structure.

[0021] The training unit is used in the first stage to train the graph convolutional learning network using the support set to obtain an intermediate graph convolutional learning network model; in the second stage, the intermediate graph convolutional learning network model is trained using the support set and the query set to obtain the trained graph convolutional learning network model.

[0022] The fault identification unit, for any coupled fault signal to be identified, uses a trained graph convolutional learning network model to output a predicted value for each label. The predicted value for each label is used to indicate whether the fault of that type exists.

[0023] The above technical solution has the following beneficial effects: For coupled fault diagnosis, when there are only a few training samples, a graph convolutional learning network model can be constructed from a few training samples, thereby enabling the identification of the fault signal to be identified and identifying the labels of the coupled fault. Attached Figure Description

[0024] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0025] Figure 1 This is a flowchart of a coupled fault diagnosis method according to an embodiment of the present invention;

[0026] Figure 2 This is a structural diagram of a coupled fault diagnosis system according to an embodiment of the present invention;

[0027] Figure 3 This is a framework diagram of the GCMS method according to an embodiment of the present invention. Detailed Implementation

[0028] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0029] like Figure 1 As shown, in conjunction with embodiments of the present invention, a coupled fault diagnosis method is provided, comprising:

[0030] Step 1: For all known fault samples in the support set and the label corresponding to each known fault sample, construct a graph representing the similarity relationship between the known fault samples. This graph is used as input to the graph convolutional learning network.

[0031] Step 2: Construct a graph convolutional learning model using the graph-based graph convolutional learning network;

[0032] Step 3: For the graph corresponding to the support set, three similarity relationships are used as the metric structure for known fault samples, and the corresponding loss function is determined based on the metric structure;

[0033] Step 4: In the first stage, the graph convolutional learning network is trained using the support set to obtain an intermediate graph convolutional learning network model; in the second stage, the intermediate graph convolutional learning network model is trained using the support set and the query set to obtain the trained graph convolutional learning network model.

[0034] Step 5: For any coupled fault signal to be identified, use the trained graph convolutional learning network model to output the predicted value of each label. The predicted value of each label is used to indicate whether the fault exists.

[0035] Preferably, the coupled fault diagnosis method further includes:

[0036] Step 6: For each known fault sample, use Fast Fourier Transform to transform the features of the fault sample from time domain features to frequency domain features to obtain the support set. Here, the known fault samples are multi-label faults; the frequency domain features refer to a known fault sample being composed of a vector of length d, where each specific value in the vector represents a feature of the known fault sample.

[0037] Step six is ​​executed before step one.

[0038] Preferably, step one specifically includes:

[0039] Let the support set be {X} s ,Y s}, X s Let X represent the set of known fault samples. s ∈R n×d Y s X represents s The corresponding label matrix, where labels represent fault types; X s Each row represents the Y corresponding to a known fault sample. s The same row represents the label of the known fault sample; one known fault sample corresponds to at least one label; Y s It is the first vector of 1*c, and the value of the first vector consists of 0 and 1: Y s ∈{0,1} n×c n is the number of known fault samples, and c is the number of label categories; where X s The number of known fault samples within the training sample is less than the number of samples in regular training.

[0040] For the support set, construct a graph for the graph convolutional learning network: {X s ,Y s The corresponding diagram, {X s ,Y s The corresponding graph represents the similarity relationship between known fault samples, {X} s ,Y s The corresponding graph contains nodes and edges, with each node corresponding to X. s A known fault sample; edges are values ​​in the first adjacency matrix A, representing the connection between nodes. The rows and columns of the first adjacency matrix A are the number of known fault samples, n. The value in the i-th row and j-th column of the first adjacency matrix A represents the similarity between the i-th and j-th known fault samples, with similarity values ​​taking either 0 or 1. A third adjacency matrix is ​​used. This represents a self-loop graph, where each known faulty sample is connected to itself via a self-loop. It is represented as: Where I is the first identity matrix;

[0041] Construct a degree matrix D, which is a diagonal matrix. The degree matrix D represents the number of connection points for each known fault sample in D, i.e., the number of known fault samples connected to each known fault sample. Each diagonal element of the degree matrix D is the sum of the corresponding rows of the adjacency matrix A, expressed as D = diag(A1), where 1 is a column vector with all elements being 1. to indicate The degree.

[0042] Preferably, step two includes:

[0043] The graph convolutional learning network is configured to consist of three base layers with the same structure. The l-th neural network layer f in the GCL network... l The structure is as follows:

[0044]

[0045] Among them, X l W is the input to the l-th neural network layer. l Let be the weight matrix of the l-th neural network layer, and δ represent the activation function in the neural network layer;

[0046] The calculations performed by each neural network layer are represented by formula (1), X l+1 This represents the output of the l-th neural network layer, while X... l+1 It also represents the input of the (l+1)th neural network layer;

[0047] X l+1 The number of feature dimensions is determined by W l The number of columns determines W. l The column represents the number of neurons in the l-th layer of the neural network;

[0048] X l+1 and The relationship between rank and rank can be expressed as an inequality:

[0049] To scale the features, a normalization layer F was added after the first neural network layer. bn f 1 ,f 2 The activation function in all layers is ReLU, and the last neural network layer f 3 The activation function in the graph convolutional learning network model is the sigmoid function.

[0050] f gcl =sequence[f 1 ,F bn ,f 2 ,f 3 (2)

[0051] The loss function of the graph convolutional learning model can be expressed as:

[0052]

[0053] Among them, L ce For cross-entropy loss, It is constructed from the support set labels.

[0054] Preferably, step one further includes:

[0055] For known fault samples within the same label class, each known fault sample is configured to connect to only one neighbor. Labels are used to distinguish the adjacency information between known fault samples and their connected neighbors: if the label vector values ​​of a known fault sample and its connected neighbor are exactly the same, then the known fault sample and its connected neighbor belong to the same class; otherwise, the known fault sample and its connected neighbor belong to different classes. The third adjacency matrix is ​​then used to... In this code, the diagonal and opposite elements are set to 1, the bottom right element is set to 1, and all other elements are set to 0. This sets the elements corresponding to the c-th category tag to 1. Represented as

[0056]

[0057] Where the superscript 'c' represents the c-th class, The rank of is nc.

[0058] Preferably, step three specifically includes:

[0059] For the support set, the similarity relationships of all known fault samples are divided into three categories: the most similar sample pair X. ms Some similar samples for X ps and dissimilar sample pairs X ds ;

[0060] Construct an objective function based on the feature distance and label consistency of known fault samples. The objective function satisfies the following condition: if the feature distance between two samples is close, then their labels are also similar. The loss function of the objective function is expressed as:

[0061] dist(f(X ms ))<dist(f(X ps ))<dist(f(X ds (11)

[0062] Where f() represents the feature mapping function, corresponding to [f 1 ,f bn ,f 2 ], dist indicates the distance calculation method;

[0063] Transform inequality (11) into a solvable form:

[0064]

[0065] Equations (12)-(14) serve as a metric structure, using three ratios: and To ensure consistency between the feature metric distance and the label of known faulty samples, where, X represents the most similar sample with label a. ms The maximum distance and partially similar samples X ps The ratio of the average distance between them The smaller the value, the more similar the sample X is. ms The distance between them is greater than that between some similar samples X ps The smaller the distance between them; the subscript 'a' indicates and It is calculated based on the a-th label of the fault sample, indicating that the a-th label is the baseline label; X represents partially similar samples labeled 'a'. ps The average distance between dissimilar sample pairs X ds The ratio of the minimum distances between two samples X is used to make some similar samples X... ps The distance between them is less than that between dissimilar samples X ds The distance between them; X represents the most similar sample with label a. ms The maximum distance between dissimilar sample pairs X ds The ratio of the minimum distances between the two samples is used to find the most similar sample X. ms The distance between them is less than that between dissimilar samples X ds The distance between them; for class c labels, each class label is used as the base label a, and the formulas (12)-(14) are used as the measurement structure;

[0066] The neural network layer f 1 ,f bn ,f 2 Considered as a feature extraction module, the last neural network layer f 3 Treating it as a classification module, the metric structure constraints (corresponding to formulas (12)-(14)) are applied to the feature extraction module: f 1 ,f bn ,f 2 To strengthen the classification module f 3 Classification ability.

[0067] Preferably, step three specifically includes:

[0068] Calculate all baseline labels The corresponding average value is used as the corresponding loss function, expressed as:

[0069]

[0070] Where {Y} represents the relationship between Y and Y. s The tag set after removing duplicate values, |{Y}| represents the number of tag categories within {Y}. So, after removing duplicates, {Y} has c types of tags, with each category containing 1 tag.

[0071] Preferably, step four specifically includes:

[0072] In the first stage, a support set is used to train the graph convolutional learning network. Combining equations (6), (15), and (17), the first empirical loss function for training the graph convolutional learning model is set as follows:

[0073] L stage1 =L gcl (X s )+αL mp (X s )+αL pd (X s (18)

[0074] Where α is the first equilibrium parameter;

[0075] When the first empirical loss function converges to the preset value, the training of the graph convolutional learning network is completed, and the intermediate graph convolutional learning network model is obtained.

[0076] Let the query set be {X} q ,Y q}, X q Let X represent the set of unknown fault samples. q ∈R N×d Y q X represents q The corresponding label matrix; X q Each row represents an unknown fault sample, Y q The corresponding row represents the label of the unknown fault sample;

[0077] X q The frequency domain features corresponding to each unknown fault sample are input into the intermediate graph convolutional learning network model, which outputs a probability prediction vector. For each unknown fault sample, there are c predicted values. When the predicted value is greater than the threshold 1 / c, it indicates that the unknown fault sample has the label corresponding to the predicted value. Through all X... q The label corresponding to each unknown fault sample within Y constitutes Y q Y q It is the second vector of 1*c, and the value of the second vector consists of 0 and 1.

[0078] Preferably, step four specifically includes:

[0079] In the second stage, the intermediate graph convolutional learning network model is trained using a support set and a query set; combining equations (6), (15), (16), and (17), a second empirical loss function for training the graph convolutional learning model is set:

[0080] Lstage2 =L gcl (X s )+βL mp (X q )+βL pd (X q )+γL md (X q (19)

[0081] Where β and γ are the second and third balance parameters, respectively, and the second balance parameter β is less than the third balance parameter γ. The first term on the right side of equation (19) is to use the support set to maintain the accuracy of the classifier. The second, third and fourth terms are used to learn the metric structure of the query set.

[0082] When the second empirical loss function converges to the preset value, the training of the intermediate graph convolutional learning network model is completed, and the trained graph convolutional learning network model is obtained.

[0083] like Figure 2 As shown, in conjunction with embodiments of the present invention, a coupled fault diagnosis system is also included, comprising:

[0084] Graph construction unit 21 is used to construct a graph representing the similarity relationship between known fault samples for all known fault samples in the support set and the label corresponding to each known fault sample, and the graph is used as input to the graph convolutional learning network.

[0085] Model construction unit 22 is used to construct a graph convolutional learning model using the graph-based graph convolutional learning network;

[0086] The metric structure construction unit 23 is used to use three similarity relationships as the metric structure for known fault samples for the graph corresponding to the support set, and to determine the corresponding loss function based on the metric structure.

[0087] Training unit 24 is used to train the graph convolutional learning network in the first stage using the support set to obtain an intermediate graph convolutional learning network model; and in the second stage, to train the intermediate graph convolutional learning network model using the support set and the query set to obtain the trained graph convolutional learning network model.

[0088] The fault identification unit 25, for any coupled fault signal to be identified, uses the trained graph convolutional learning network model to output the predicted value of each label, which is used to indicate whether the fault of that type exists.

[0089] Preferably, the coupled fault diagnosis system further includes:

[0090] The feature transformation unit is used to transform the features of each known fault sample from time-domain features to frequency-domain features using the Fast Fourier Transform to obtain the support set. Here, the known fault samples are multi-label faults. The frequency-domain features refer to the fact that a known fault sample is composed of a vector of length d, and each specific value in the vector represents the feature of the known fault sample.

[0091] Preferably, the graph construction unit 21 is used for:

[0092] Let the support set be {X} s ,Y s}, X s Let X represent the set of known fault samples. s ∈R n×d Y s X represents s The corresponding label matrix, where labels represent fault types; X s Each row represents the Y corresponding to a known fault sample. s The same row represents the label of the known fault sample; one known fault sample corresponds to at least one label; Y s It is the first vector of 1*c, and the value of the first vector consists of 0 and 1: Y s ∈{0,1} n×c n is the number of known fault samples, and c is the number of label categories; where X s The number of known fault samples within the training sample is less than the number of samples in regular training.

[0093] For the support set, construct a graph for the graph convolutional learning network: {X s ,Y s The corresponding diagram, {X s ,Y s The corresponding graph represents the similarity relationship between known fault samples, {X} s ,Y s The corresponding graph contains nodes and edges, with each node corresponding to X. s A known fault sample; edges are values ​​in the first adjacency matrix A, representing the connection between nodes. The rows and columns of the first adjacency matrix A are the number of known fault samples, n. The value in the i-th row and j-th column of the first adjacency matrix A represents the similarity between the i-th and j-th known fault samples, with similarity values ​​taking either 0 or 1. A third adjacency matrix is ​​used. This represents a self-loop graph, where each known faulty sample is connected to itself via a self-loop. It is represented as: Where I is the first identity matrix;

[0094] Construct a degree matrix D, which is a diagonal matrix. The degree matrix D represents the number of connection points for each known fault sample in D, i.e., the number of known fault samples connected to each known fault sample. Each diagonal element of the degree matrix D is the sum of the corresponding rows of the adjacency matrix A, expressed as D = diag(A1), where 1 is a column vector with all elements being 1. to indicate The degree.

[0095] Preferably, the module building unit 22 is used for:

[0096] The graph convolutional learning network is configured to consist of three base layers with the same structure. The l-th neural network layer f in the GCL network... l The structure is as follows:

[0097]

[0098] Among them, X l W is the input to the l-th neural network layer. l Let be the weight matrix of the l-th neural network layer, and δ represent the activation function in the neural network layer;

[0099] The calculations performed by each neural network layer are represented by formula (1), X l+1 This represents the output of the l-th neural network layer, while X... l+1 It also indicates the lth +1 The input to each neural network layer;

[0100] X l+1 The number of feature dimensions is determined by W l The number of columns determines W. l The column represents the number of neurons in the l-th layer of the neural network;

[0101] X l+1 and The relationship between rank and rank can be expressed as an inequality:

[0102] To scale the features, a normalization layer F was added after the first neural network layer. bn f 1 ,f 2 The activation function in all layers is ReLU, and the last neural network layer f 3 The activation function in the graph convolutional learning network model is the sigmoid function.

[0103] f gcl =sequence[f 1 ,F bn ,f 2 ,f 3 (2)

[0104] The loss function of the graph convolutional learning model can be expressed as:

[0105]

[0106] Among them, L ce For cross-entropy loss, It is constructed from the support set labels.

[0107] Preferably, the graph construction unit 21 is used for:

[0108] For known fault samples within the same label class, each known fault sample is configured to connect to only one neighbor. Labels are used to distinguish the adjacency information between known fault samples and their connected neighbors: if the label vector values ​​of a known fault sample and its connected neighbor are exactly the same, then the known fault sample and its connected neighbor belong to the same class; otherwise, the known fault sample and its connected neighbor belong to different classes. The third adjacency matrix is ​​then used to... In this code, the diagonal and opposite elements are set to 1, the bottom right element is set to 1, and all other elements are set to 0. This sets the elements corresponding to the c-th category tag to 1. Represented as

[0109]

[0110] Where the superscript 'c' represents the c-th class, The rank of is nc.

[0111] Preferably, the metric structure building unit 23 is specifically used for:

[0112] For the support set, the similarity relationships of all known fault samples are divided into three categories: the most similar sample pair X. ms Some similar samples for X ps and dissimilar sample pairs X ds ;

[0113] Construct an objective function based on the feature distance and label consistency of known fault samples. The objective function satisfies the following condition: if the feature distance between two samples is close, then their labels are also similar. The loss function of the objective function is expressed as:

[0114] dist(f(X ms ))<dist(f(X ps ))<dist(f(X ds (11)

[0115] Where f() represents the feature mapping function, corresponding to [f 1 ,f bn ,f 2], dist indicates the distance calculation method;

[0116] Transform inequality (11) into a solvable form:

[0117]

[0118] Equations (12)-(14) serve as a metric structure, using three ratios: and To ensure consistency between the feature metric distance and the label of known faulty samples, where, X represents the most similar sample with label a. ms The maximum distance and partially similar samples X ps The ratio of the average distance between them The smaller the value, the more similar the sample X is. ms The distance between them is greater than that between some similar samples X ps The smaller the distance between them; the subscript 'a' indicates and It is calculated based on the a-th label of the fault sample, indicating that the a-th label is the baseline label; X represents partially similar samples labeled 'a'. ps The average distance between dissimilar sample pairs X ds The ratio of the minimum distances between two samples X is used to make some similar samples X... ps The distance between them is less than that between dissimilar samples X ds The distance between them; X represents the most similar sample with label a. ms The maximum distance between dissimilar sample pairs X ds The ratio of the minimum distances between the two samples is used to find the most similar sample X. ms The distance between them is less than that between dissimilar samples X ds The distance between them; for class c labels, each class label is used as the base label a, and the formulas (12)-(14) are used as the measurement structure;

[0119] The neural network layer f 1 ,f bn ,f 2 Considered as a feature extraction module, the last neural network layer f 3 Treating it as a classification module, the metric structure constraints (corresponding to formulas (12)-(14)) are applied to the feature extraction module: f 1 ,f bn ,f 2 To strengthen the classification module f 3 Classification ability.

[0120] Preferably, the metric structure building unit 23 is specifically used for:

[0121] Calculate all baseline labels The corresponding average value is used as the corresponding loss function, expressed as:

[0122]

[0123] Where {Y} represents the relationship between Y and Y. s The tag set after removing duplicate values, where |{Y}| represents the number of tag categories within {Y}, then the deduplicated {Y} has... c There are 1 type of tag, and the number of tags in each type is 1.

[0124] Preferably, the training unit 24 is specifically used for:

[0125] In the first stage, a support set is used to train the graph convolutional learning network. Combining equations (6), (15), and (17), the first empirical loss function for training the graph convolutional learning model is set as follows:

[0126] L stage1 =L gcl (X s )+αL mp (X s )+αL pd (X s ) (18)

[0127] Where α is the first equilibrium parameter;

[0128] When the first empirical loss function converges to the preset value, the training of the graph convolutional learning network is completed, and the intermediate graph convolutional learning network model is obtained.

[0129] Let the query set be {X} q ,Y q}, X q Let X represent the set of unknown fault samples. q ∈R N×d Y q X represents q The corresponding label matrix; X q Each row represents an unknown fault sample, Y q The corresponding row represents the label of the unknown fault sample;

[0130] X q The frequency domain features corresponding to each unknown fault sample are input into the intermediate graph convolutional learning network model, which outputs a probability prediction vector. For each unknown fault sample, there are c predicted values. When the predicted value is greater than the threshold 1 / c, it indicates that the unknown fault sample has the label corresponding to the predicted value. Through all X... qThe label corresponding to each unknown fault sample within Y constitutes Y q Y q It is the second vector of 1*c, and the value of the second vector consists of 0 and 1.

[0131] Preferably, the training unit 24 is specifically used for:

[0132] In the second stage, the intermediate graph convolutional learning network model is trained using a support set and a query set; combining equations (6), (15), (16), and (17), a second empirical loss function for training the graph convolutional learning model is set:

[0133] L stage2 =L gcl (X s )+βL mp (X q )+βL pd (X q )+γL md (X q (19)

[0134] Where β and γ are the second and third balance parameters, respectively, and the second balance parameter β is less than the third balance parameter γ. The first term on the right side of equation (19) is to use the support set to maintain the accuracy of the classifier. The second, third and fourth terms are used to learn the metric structure of the query set.

[0135] When the second empirical loss function converges to the preset value, the training of the intermediate graph convolutional learning network model is completed, and the trained graph convolutional learning network model is obtained.

[0136] The technical solutions of the present invention will be described in detail below with reference to specific application examples. For technical details not described in the implementation process, please refer to the relevant descriptions above.

[0137] A few-sample coupled fault diagnosis method that integrates multi-label graph convolutional learning with metric structure, such as Figure 3 As shown, firstly, the frequency domain features of the fault signal are obtained using Fast Fourier Transform (FFT) for feature enhancement. Secondly, at the data augmentation level, a graph convolutional network is designed, which consists of a self-looping graph (second adjacency matrix). This involves integrating a high-order nearest neighbor adjacency graph (i.e., a graph) to form a composite graph structure. Networks built using this graph can enhance samples and aggregate features, aiming to simultaneously expand the scale of training samples and integrate feature information, thereby reducing the negative impact of limited training data. Furthermore, at the metric structure learning level, multi-layer similarity relationships in multi-label data are mined, and three similarity ranking loss functions are designed and embedded in the network. These functions effectively construct the structure of the feature space and constrain the distance ranking of different similarities, ensuring that samples with different similarities maintain a reasonable distance between the support set and the query set. This structural constraint provides additional information support to the model, thereby improving its performance. This further narrows the scope of the hypothesis space and improves the model's generalization ability. Specifically, to increase the model's supervision information, the distance ranking relationships of most similar, partially similar, and dissimilar metrics are used, and three metric loss functions are designed to construct the metric structure. This metric structure can be seen as a structural constraint on the model, reducing the model's hypothesis space. For the query set, a label discrimination threshold is added to obtain pseudo-labels with higher confidence. Finally, pseudo-labels are used to optimize the three loss metrics to utilize the feature information of the query set and obtain the final predicted labels.

[0138] Furthermore, the proposed model is adaptive, and its hyperparameters can be fixed across different few-sample tasks. The effectiveness of the proposed method is demonstrated on various experimental data.

[0139] An adaptive optimization framework was designed to control the values ​​of each term in the loss function and adaptively adjust the training time on the support and query sets respectively. This framework can be used for different tasks with fixed hyperparameters.

[0140] First, problem definition and preprocessing.

[0141] This section introduces the basic concepts of few-shot learning. Given Corresponding tags in,

[0142] {X s ,Y s} is called the support set, X s Let X represent the set of known fault samples. s ∈R n×d Y s X represents s The corresponding label matrix, where labels represent fault types. X s Each row represents the Y corresponding to a known fault sample. s The same line represents the label of the known fault sample. A known fault sample may correspond to multiple labels (at least one label), indicating that a known fault sample may have multiple faults. sIt is a 1*c vector, and the values ​​of the vector consist of 0s and 1s. That is, Y. s ∈{0,1} n×c n is the number of known fault samples, c is the number of label categories, and k is the number of known fault samples corresponding to each label category.

[0143] d represents the number of feature dimensions, which is an inherent attribute of a sample. A sample is composed of a vector of length d, where each specific value in the vector is a feature of the sample. If we consider the features of this sample as a d-dimensional space, then the sample is a specific coordinate point in that d-dimensional space. The feature dimension is the dimension of the d-dimensional space, and it is also equal to the length of the vector.

[0144] {X q ,Y q} is called the query set, X q Let X represent the set of unknown fault samples. q ∈R N×d Y q X represents q The corresponding label matrix. X q Each row represents an unknown fault sample, Y q The corresponding row represents the label of this unknown fault sample. Y q The value inside is a 1*c vector, and the value of the vector consists of 0 and 1.

[0145] In a few-shot learning task, the number of known faulty samples for each label class is s, and X s The number of known fault samples s in the sample is finite, and the small sample task is based on X. s ,Y s Predict the set of unknown fault samples X q The corresponding label matrix Y q When this is the case, this task is called a few-shot learning task.

[0146] Learning unknown fault samples X using graph convolutional learning networks q The corresponding label Y q During the prediction process, Y q It is invisible. Because the set X of known fault samples is... s The number of corresponding label categories is c, and the number of known fault samples corresponding to each label category is k. This is called a c-way k-shot task.

[0147] The few-shot learning task is based on a combination of a limited number of known fault samples X. s Learn a suitable graph convolutional learning network model and predict query set X qThe fault information is further used to train the graph convolutional learning network model, and finally a graph convolutional learning network model is obtained to determine the label of the coupling fault to be identified.

[0148] Due to the specific nature of fault diagnosis problems, the research object is signal data, and the input data is time-domain data. Therefore, a common practice is to preprocess the data before training. The most common method is to use Fast Fourier Transform (FFT) to convert the original data (known fault sample signals and unknown fault sample signals) into frequency-domain data to obtain more significant features. Since the processed features are symmetric, the valuable feature dimensions are reduced to half of the original. After preprocessing, the number of feature dimensions in the support set and query set becomes... Both the training and prediction processes are based on preprocessed data.

[0149] Second, graph convolutional learning networks

[0150] In few-shot learning tasks, the most challenging problem is the limited number of known fault samples used for training. From the perspective of basic machine learning theory, this leads to a large generalization error. Therefore, how to maximize the use of feature information from a limited number of known fault samples is an important research direction. Graph Convolutional Learning (GCL) networks can effectively aggregate valuable features and increase the number of samples, making them an effective few-shot learning method. Therefore, a convolutional learning model (GCL model) for coupled fault diagnosis is obtained by training a graph convolutional learning network.

[0151] Before training, a key point in graph convolutional learning networks is how to construct the graph, through pre-constructing {X}. s Y s The graph is constructed and then applied to a graph convolutional network for training the graph convolutional learning network. s and Y s The main function of the corresponding graph is to represent the similarity relationships between known fault samples. {X s Y s The corresponding graph contains nodes and edges, with each node corresponding to X. s A known fault sample; edges are values ​​in the first adjacency matrix A, representing the connection between nodes. The rows and columns of the first adjacency matrix A are the number of known fault samples, n. The value in the i-th row and j-th column of the first adjacency matrix A represents the similarity between the i-th and j-th known fault samples. The similarity value can be 0 or 1. For example, if the i-th and j-th known fault samples belong to the same class and have the same label, the value is 1; otherwise, it is 0. A {X s Y s} corresponds to a graph. If {X s Y s} changes, {X s Y s The first adjacency matrix A of the corresponding graph will generally change accordingly. Each value in the first adjacency matrix A represents the similarity between two known fault samples. If it is 1, it means that the two known fault samples have the same label and are known fault samples of the same category; this relationship is called a connected sample pair. If it is 0 or -1, it means that the two known fault samples have different labels and are known fault samples of different categories; this relationship is called a disconnected sample pair.

[0152] The first adjacency matrix A contains all connected and disconnected sample pairs, allowing us to learn the distance between each known faulty sample and its known other faulty samples. In the GCL network, the key is to combine the aggregated features (weighted average features) of similar samples to achieve better performance. Simultaneously, a third adjacency matrix... To represent a self-loop graph, which shows the self-loops connecting each known faulty sample to itself, we can use the following representation: Where I is the identity matrix.

[0153] We also need to construct a degree matrix D, a diagonal matrix that reflects the number of connections for each known fault sample, i.e., the number of known fault samples connected to each known fault sample. Each diagonal element of the degree matrix D is the sum of the corresponding rows of the adjacency matrix A, expressed as D = diag(A1), where 1 represents a column vector with all elements being 1. When using a GCL network to learn a graph, the degree matrix D is usually used as a normalization matrix; To represent the third adjacency matrix The degree.

[0154] For example, if the i-th row of an adjacency matrix is ​​[0,0,1,1], then this row means that the i-th known fault sample, the 3rd known fault sample, and the 4th known fault sample are of the same type. In other words, the i-th known fault sample, the 3rd known fault sample, and the 4th known fault sample are connected to each other. The number of known fault samples connected to the i-th known fault sample is the sum of the values ​​in this row, which is 2. A GCL network consists of several base layers (neural network layers) with the same structure. The l-th neural network layer f in a GCL network... l The structure is as follows:

[0155]

[0156] Among them, X l W is the input data for the l-th neural network layer. lLet be the weight matrix of the l-th neural network layer, and δ represent the activation function in the neural network layer.

[0157] Formula (1) represents the computation performed by each layer of the neural network, X l+1 This indicates that X is the output of the l-th neural network layer. l+1 This also represents the input of the (l+1)th layer of the neural network. Therefore, X... l+1 The number of feature dimensions is determined by W l The number of columns determines W. l The column represents the number of neurons in the l-th layer of the neural network. Furthermore, X... l+1 and The rank inequality holds.

[0158] Therefore, when constructing a GCL model using a GCL network, different layer parameters need to be defined: the number of neurons in each neural network layer and δ, with the number of neurons in the last neural network layer set to the dimension of the label c. In this embodiment of the invention, three neural network layers are used to construct the GCL model. The number of neurons (i.e., the number of outputs) in each neural network layer are 160, 80, and c, respectively. Furthermore, to scale the features, a normalization layer F is added after the first neural network layer. bn f 1 ,f 2 The activation function in all layers is ReLU, and the last neural network layer f 3 The activation function in the model is the sigmoid function. The GCL model can be written as:

[0159] f gcl =sequence[f 1 ,F bn ,f 2 ,f 3 (2)

[0160] While current GCL models have achieved good results, two problems remain when applied to few-shot learning. The first problem is the balance between feature aggregation capability and the number of available known fault samples. If we need to enhance the feature aggregation capability of known fault samples, we need to increase the number of nearest neighbors for each known fault sample to enrich the aggregated samples; however, this leads to a reduction in the actual usable samples. For example, if each known fault sample is adjacent to every other known fault sample, i.e. If all the elements in the formula are 1, then the formula (1) contains... If each row in the algorithm has the same element (value), then only one known fault sample can actually be used for training in the next neural network layer, greatly reducing the number of trainable sample features. Therefore, it is necessary to balance feature aggregation with the number of available known fault samples.

[0161] Regarding the first question, for known fault samples within the same label class, each known fault sample is configured to connect only to one of its neighbors (known fault samples). Labels will be used to distinguish adjacent information (i.e., if the values ​​of two label vectors are exactly the same, then the sample pair is of the same class and adjacent; otherwise, they are different). In the code, the diagonal and opposite elements are set to 1, the bottom-right element is set to 1, and all other elements are set to 0. Specifically, the elements corresponding to the c-th category tag are set to 1. Represented as The specific form is as follows:

[0162]

[0163] Here, the superscript 'c' represents the c-th class. Thus, The rank is nc. In fact, since the entire GCL model is composed of multiple neural network layers, the adjacency matrix... Regardless of the individual neural network layers, as the number of neural network layers increases, more known fault samples can be aggregated.

[0164] In the first neural network layer of the GCL network, assuming that the known faulty samples x1, x2, x3, and x4 are samples with the same label, then:

[0165]

[0166] Among them, W 1 Let σ be the weight matrix of the first layer of the neural network, and let σ() represent the activation function in the neural network layer. and These represent the inputs to the first neural network layer, and the superscript 1 represents the output of the first neural network layer. and Let represent the output of the first neural network layer, which is the input of the second neural network layer; when data enters the second neural network layer, then:

[0167]

[0168] in, and These represent the outputs of the second neural network layer, using f l l = 1, 2, 3 represents the calculations performed by each neural network layer, and the superscript indicates the l-th neural network layer;

[0169] From formulas (4) and (5), it can be seen that although the number of connection points for each known fault sample is 1, as the number of neural network layers increases, the output of each neural network layer can still aggregate the features of different samples.

[0170] The second question is how to construct a new adjacency matrix (the second adjacency matrix). Increase the number of training samples. As we know from the computational method of graph convolutional neural networks, different graph adjacency matrices can construct different inputs for the neural network layers. Therefore, adding a new graph adjacency matrix can achieve the goal of increasing the number of training samples. Additionally, when using the current GCL model to test new data, it is still necessary to construct a new adjacency matrix (the second adjacency matrix) for the data. However, due to the high dimensionality of the sample features, the commonly used method of constructing an adjacency matrix based on the distances of the original known fault samples may not be reliable. Furthermore, it is desired that the final GCL model can predict the corresponding label even when the fault signal to be identified lacks neighbor information. Therefore, a second adjacency matrix is ​​constructed using [a different approach]. Used to construct a self-loop graph, where the second adjacency matrix The identity matrix. During training, the first identity matrix I and the first adjacency matrix are used simultaneously based on the label information of the support set. Consider the first identity matrix I and the first adjacency matrix The classification loss also considers the second adjacency matrix. Based on the classification loss, two loss functions are derived (the two terms on the right side of equation (6)). In this way, the GCL model can aggregate features during training and make predictions on the test dataset without constructing an adjacency matrix. On the other hand, this operation can also be seen as a sample augmentation technique to further increase the number of training samples.

[0171] In summary, the loss function of the GCL model can be expressed as:

[0172]

[0173] Among them, L ce For cross-entropy loss, This indicates that the support set labels are constructed.

[0174] Third, learning metric structures

[0175] In learning with small samples, insufficient training data presents a challenge. Machine learning theory includes the following: First, there are empirical loss and expected loss. The model loss obtained during training is called empirical loss, expressed as:

[0176]

[0177] Among them, Risk empirical (f) is the empirical loss function, representing the loss function of the model learned from the training data, used to measure the difference between the predicted labels and the true labels during training. In addition, there exists an ideal loss for the model, called the expected loss, expressed as:

[0178]

[0179] Here, Risk(f) represents the expected loss function, which is the true predicted loss on a population of data following a certain distribution. This distribution is usually unknown, and therefore the expected loss function is also unknown. Ideally, the expected loss should approach zero; however, this can only be approximated by empirical loss. Generally, the relationship between the empirical loss function and the expected loss function is as follows:

[0180]

[0181] Clearly, when the number of training samples approaches infinity, the empirical loss on the training data will converge to the expected loss. However, in actual training, the amount of data cannot be infinite. Therefore, in actual training, there will always be an error between the empirical risk and the expected risk, i.e.:

[0182] GeneralizationError=Risk(f)-Risk empirical (f) (10)

[0183] Equation (10) is called the generalization error. If both equation (7) and equation (10) can be made to approach 0, then equation (8) can be considered to approach 0 as well. Therefore, in the learning process, it is necessary not only to minimize the empirical loss function of the model, but also to minimize the generalization error (which refers to the difference between empirical risk and expected risk, i.e., equation (10)). In few-shot learning, the most critical issue is how to minimize the generalization error. An effective method is to use the prior information of the data (in statistics, prior information is information about parameters or models known before statistical inference. In this embodiment of the invention, prior information generally refers to generally accepted and reasonable assumptions) to constrain the direction of model optimization so that it follows the objective laws of the data. Mining the relationship between data similarity and labels can effectively utilize the prior information of the data, thus strengthening the constraint of the generalization error from this direction.

[0184] In fact, coupled fault data contains multiple fault signals, and therefore can also be viewed as multi-label data. The biggest difference between coupled fault data and single-fault diagnosis problems lies in the diversity of label information composition, leading to multi-layered similarity among sample pairs with different fault signal components. For example, suppose there are four signal samples in a gearbox, denoted as x1, x2, x3, and x4. A four-dimensional label vector is set to represent "stripping," "broken tooth," "shaft bending," and "shaft imbalance" faults, respectively. Using "0-1" encoding to encode the labels of the known fault samples, assuming the labels for x1, x2, x3, and x4 are [1,0,1,0], [1,0,1,0], [1,0,0,1], and [0,1,0,1], respectively, then x1 and x2 have completely identical labels and are the most similar sample pair. x1 and x4 have no shared labels and are dissimilar sample pairs. The labels for x1 and x3 are not completely identical; from a traditional single-label perspective, x1 and x3 should be considered dissimilar. However, both x1 and x3 contain "stripping" fault signals, and intuitively, there should be a weak similarity between x1 and x3. This can be considered a special characteristic of multi-label learning, that is, it contains multiple layers of similarity, including similarity, dissimilarity, and partial similarity.

[0185] This invention constructs an objective function based on the feature distance and label consistency of samples, which can better utilize the multi-layer similarity of coupled faults. This objective function is based on a fundamental assumption: if the feature distance between two known fault samples is close, then the labels of the two known fault samples are also similar. The specific settings are as follows: first, all known fault samples are divided into three categories of sample pairs: the most similar sample pair X... ms Some similar samples for X ps and dissimilar sample pairs X ds The loss function of the objective function can be written as:

[0186] dist(f(X ms ))<dist(f(X ps ))<dist(f(X ds (11)

[0187] Here, f() represents a feature mapping function concept, corresponding to [f 1 ,f bn ,f 2 ], dist represents the distance calculation method. In this embodiment, Euclidean distance is used.

[0188] However, inequality (11) cannot be solved. Therefore, inequality (11) needs to be transformed into a form that can be solved by optimization, as follows:

[0189]

[0190] Where, dist represents the distance calculation method, which in this embodiment uses Euclidean distance; mean, max, and min represent the average, maximum, and minimum distance values, respectively; three different criteria are used to calculate the distance between three sample pairs: the most similar sample pair X ms Calculate the maximum value of all distances, and find the partially similar sample pairs X. ps Calculate the mean distance for all distances, and then calculate the mean distance for dissimilar sample pairs X. ds Calculate the minimum of all distances, min. This is because the most similar sample pair X... ms The maximum value of the distance between dissimilar samples X is the upper bound of the similarity distance. ds The minimum distance is the lower bound of the distance. By minimizing the ratio of this upper bound to the lower bound, the absolute distance ranking relationship between these two sample pairs can be constructed. For partially similar sample pairs X... ps Since neither its upper nor lower bounds can reflect the overall situation, the average distance is used. Furthermore, to avoid a denominator of 0, a 1 is added to the denominator to control the size of the fraction. Thus, based on these three ratios: and This ensures consistency between the feature metric distance and the label of known faulty samples.

[0191] X represents the most similar sample with label a. ms The maximum distance and partially similar samples X ps The ratio of the average distance between them The smaller the value, the more similar the sample X is. ms The distance between them is greater than that between some similar samples X ps The smaller the distance between them; the subscript 'a' indicates and It is calculated based on the a-th label of the fault sample, indicating that the a-th label is the baseline label; X represents partially similar samples labeled 'a'. ps The average distance between dissimilar sample pairs X ds The ratio of the minimum distances between two samples X is used to make some similar samples X... ps The distance between them is less than that between dissimilar samples X ds The distance between them; X represents the most similar sample with label a. ms The maximum distance between dissimilar sample pairs X ds The ratio of the minimum distances between the two samples is used to find the most similar sample X. ms The distance between them is less than that between dissimilar samples X dsThe distance between them.

[0192] For the c-type label, each type of label is used as the base label a, and the formulas (12)-(14) are used as the measurement structure.

[0193] In graph convolutional learning networks, the neural network layer f 1 ,f bn ,f 2 Considered as a feature extraction module, the last neural network layer f 3 Since it is considered a classification module, f in equations (12)-(14) is the output of the feature extraction module, which is the second neural network layer f. 2 Layer output: f 3 Metric structure constraints (corresponding to formulas (12)-(14)) are applied to the feature extraction module to enhance classification capabilities.

[0194] Another challenge is how to partition the sample pairs. Since the known fault samples in the support set are labeled, the known fault sample pairs in the support set are partitioned by these labels. Because both similar and dissimilar sample pairs are based on a baseline label, a baseline label is set and different labels are iterated over to obtain three different types of sample pairs. Samples whose labels are exactly the same as the baseline label from other known fault samples are the most similar sample pairs; samples whose labels have no common labels with the baseline label from other known fault samples are the dissimilar sample pairs; and samples whose labels have some common labels with the baseline label from other known fault samples are the partially similar sample pairs. Finally, three types of losses are calculated in the iteration over different baseline labels. Their respective average values, for example, X s If the number of label categories is c, then for training sample X... s Here, the number of iterations is c. The c classes of labels on the support set are iterated sequentially, with each class used as the baseline label in turn, to calculate the three losses. This is the final loss, and the average values ​​of these three losses are as follows:

[0195]

[0196]

[0197] Where {Y} represents the relationship between Y and Y. s The set of labels after removing duplicate values, where |{Y}| represents the number of label categories within {Y}. For c - Way k-shot task: If the number of known fault samples corresponding to each label is k, then after deduplication, {Y} will only have c labels, with 1 label in each category. |{Y}| represents the number of elements in the set. This symbol is a fixed representation in mathematics, and its value here is c.

[0198] Fourth, phased learning strategies

[0199] In few-shot learning, to make deeper use of metric structure information, the query set needs to be incorporated into the training process. Therefore, the training process will be divided into two phases. The first phase uses the support set for training, and the second phase uses both the support set and the query set for training.

[0200] In the first stage, combining formulas (6), (15), and (17), the empirical loss function for training the GCL model is set as follows:

[0201] L stage1 =L gcl (X s )+αL mp (X s )+αL pd (X s (18)

[0202] Here, α is the first balancing parameter. After the GCL model is trained, its output will be a probability prediction vector between 0 and 1. To derive the labels, the probability prediction vector is converted into a 0-1 label vector. Simultaneously considering the labels of the fault signals (fault samples), the threshold for each value decision within the label vector is set to 1 / c.

[0203] In the second stage, formulas (6), (15), (16), and (17) are combined to train both the support set and the query set simultaneously. The empirical loss function for training the GCL model is expressed as:

[0204] L stage2 =L gcl (X s )+βL mp (X q )+βL pd (X q )+γL md (X q (19)

[0205] Where β and γ are the second and third balance parameters, respectively. The first term on the right side of equation (19) is used to maintain the accuracy of the classifier by using the support set. The second, third, and fourth terms are used to learn the metric structure of the query set. Since the true label of the query set is unknown, pseudo-labels are needed to construct three sample pairs. Pseudo-labels refer to the labels obtained after the first stage of training, by using the current GCL model to predict the unknown fault samples in the query set. Furthermore, for the query set, the label discrimination threshold 1 / c is used to determine the final label of the unknown fault sample, which can obtain pseudo-labels with higher confidence.

[0206] The biggest difference between equations (18) and (19) lies in the components that measure the structural loss. Equation (11) reflects the comparison relationship between the distances of the three types of samples. If the sample distances hold true, then dist(f(X) ms ))<dist(f(X ps )),dist(f(X ps ))<dist(f(X ds )), then dist(f(X) ms ))<dist(f(X ds This also holds true. However, this conclusion requires that the sample pairs strictly conform to the correspondence between labels and distances. In equation (18), since the support set has labels, there is no need to add an additional metric structure L. md (loss function L) md The query set requires pseudo-labels to determine similarity. Since the GCL model during training is not optimal, the pseudo-labels obtained at this time have certain errors, which in turn lead to errors in the segmentation of the three types of sample pairs in the query set. In addition, intuitively speaking, judging similarity and dissimilarity is more reliable than judging partial similarity. Therefore, it is necessary to add the metric structure of all three types of sample pairs (the second, third, and fourth terms of formula (19)) to the query set for training, and the second balance parameter β should be less than the third balance parameter γ.

[0207] Fifth, Experiment

[0208] The performance of the proposed diagnostic method will be evaluated through the following experiments. Before proceeding with the specific experiments, the experimental setup, comparison methods, and experimental dataset information will be introduced.

[0209] Six state-of-the-art few-shot diagnostic learning methods were selected for comparison, including two meta-learning methods, one transfer-based method, and two data augmentation methods. The data augmentation methods are Generalized Graph Contrastive Learning (GCN) and GCL, proposed in 2024. The difference between GCN and GCL is that GCN includes a pre-training process, while GCL does not. The transfer-based and meta-learning-based methods were proposed in 2020 and named FTN and MRN, respectively. Finally, to further validate the effectiveness of the proposed multi-label few-shot learning methods, another state-of-the-art multi-label few-shot learning method, proposed in 2024 and named BCR, was added. The last method is the baseline method, a ResNet model named FsRes. Since GCL, GCN, and FTN are designed for single faults, when applied to coupled faults, the softmax activation function of the last layer of their network's classification layer is converted to a sigmoid function. For MRN, to conform to the original paper, coupled faults are treated as a single fault type, and multi-labels are converted to multi-class labels for coupled fault diagnosis.

[0210] Two datasets were used in the experiment. The first dataset came from the 2009 PHM Data Challenge and consisted of fault data collected from spur gears. This dataset included various fault types, including chipped, eccentric, and broken gear faults; inner, outer, and ball bearing faults; and imbalance faults in shafts. It included single-fault data and coupled fault data involving multiple faults. The sampling frequency of this dataset was 66.67 kHz. Based on this dataset, dataset A was constructed using data under different operating conditions. To construct dataset A, data under five health states were selected from four different operating conditions: 30Hz low voltage (A1), 30Hz high voltage (A2), 40Hz low voltage (A3), and 40Hz high voltage (A4). The specific fault types corresponding to each health state in dataset A are detailed in Table 1. Each dataset contains 1200 sample points, with 200 samples for each health state, for a total of 1000 samples per dataset.

[0211] The second dataset, designated Dataset B, is a fault dataset collected from parallel gearboxes in a laboratory environment. This dataset includes chipped and broken gear faults, inner and ball bearing faults, and bent and imbalanced shaft faults. It includes data for normal operation and six coupled fault conditions, with the specific fault types for each health condition outlined in Table 2. The sampling frequency for this dataset is 51.2 kHz. All fault modes were collected at two speeds (1000 rpm and 2000 rpm) and two loads (0.02 A and 0.07 A), resulting in four datasets under four operating conditions, labeled B1, B2, B3, and B4. Each dataset contains 1200 samples, with 200 samples for each health condition, for a total of 1400 samples per dataset. Label tables are also used to illustrate the relationships between coupled faults and multi-labeled issues, as shown in Tables 3 and 4. Tables 1, 2, 3, and 4 show that all fault data under different states in the dataset are coupled from two or more fault types, making it reasonable to consider coupled fault data as multi-label data.

[0212] Each dataset is divided into two parts, each containing 100 samples from each class. The first part serves as the support set, and the second part as the query set. In each experimental task, different datasets are repeatedly selected as new support sets, while the query set remains constant, and the experiment is repeated 10 times. This means that during the experiment, to ensure the reliability and stability of the results, different subsets of the dataset are randomly selected multiple times to construct the support set, while the query set remains unchanged. In this way, the performance of the model under different support set conditions can be evaluated.

[0213] Table 1 shows the specific fault types for each health state in dataset A.

[0214] health status gear bearings axis 1 healthy healthy healthy 2 Peeling & Pitting healthy healthy 3 Pitting & Tooth Breakage Sphere failure healthy 4 Broken tooth Inner ring & outer ring & sphere malfunction unbalanced 5 healthy Outer ring & sphere malfunction unbalanced

[0215] Table 2 shows the specific fault types for each health state in dataset B.

[0216] health status gear bearings axis 1 healthy healthy healthy 2 Broken tooth Inner ring fault Shaft bending 3 healthy Sphere failure unbalanced 4 Broken tooth Inner ring fault healthy 5 Peeling Sphere failure unbalanced 6 Broken tooth Sphere failure healthy 7 Peeling Inner ring fault Shaft bending

[0217] Table 3 shows the 0-1 codes for each health state in dataset A.

[0218] health status healthy Peeling pitting Broken tooth Inner circle Outer ring sphere unbalanced 1 1 0 0 0 0 0 0 0 2 0 1 1 0 0 0 0 0 3 0 0 1 1 0 0 1 0 4 0 0 0 1 1 1 1 1 5 0 0 0 0 0 1 1 1

[0219] Table 4 shows the 0-1 codes for each health state in dataset B.

[0220] health status healthy Peeling Broken tooth Inner circle sphere Shaft bending unbalanced 1 1 0 0 0 0 0 0 2 0 0 1 1 0 1 0 3 0 0 0 0 1 0 1 4 0 0 1 1 0 0 0 5 0 1 0 0 1 0 1 6 0 0 1 0 1 0 0 7 0 1 0 1 0 1 0

[0221] Experiments were conducted on eight different datasets (A1, A2, A3, A4, B1, B2, B3, and B4) with three different sample sizes: 3, 5, and 10 samples, respectively. The accuracy of different comparison methods—FsRes, FTN, MRN, BCR, GCL, GCN, and the method GCMS from this embodiment—was recorded. The hyperparameters of GCMS were fixed, with α, β, and γ set to 0.01, 0.005, and 0.02, respectively, and the learning rate set to 2e^(- ... -4 These values ​​are based on experimental experience. Meta-learning-based MRN, BCR methods, and transfer learning-based FTN methods all require source datasets. Therefore, an additional dataset is provided, which only changes the working conditions; for example, when learning task A1, the A2 data will be used as the source data for these methods.

[0222] The evaluation criterion is multi-label accuracy, expressed as follows:

[0223]

[0224] Where m is the number of predicted samples, y i , These are the true labels and the predicted labels, respectively. This accuracy is a combination of precision and recall. If the precision is high, the numerator of equation (20) will be large. Conversely, if the recall is high, the denominator of equation (20) will be small. This calculation method is more rigorous than others, but it also takes into account the special characteristics of multi-label problems. The final results are shown in Table 5.

[0225] Table 5 shows that the proposed GCMS method achieves higher accuracy than other methods in most cases. From a dataset perspective, despite varying tasks, the method in this embodiment maintains high accuracy across different datasets. From a shot number perspective, the method performs best in scenarios with 3 shots. These two perspectives sufficiently demonstrate the effectiveness of the proposed method. Several additional conclusions can be drawn compared to other methods. The first conclusion is that learning through support sets is a valuable technical roadmap. For example, GCN and GCL have achieved good performance in half of the tasks. Pre-training with three different graphs is a feasible idea, but its effectiveness is not yet stable. The second conclusion is that transfer-based methods are greatly affected by differences in data distribution between the source and query sets. For example, in task A2, where the source data is A3, FTN's accuracy is worse than the baseline FsRes, while in other tasks, it outperforms the baseline method. This is attributed to the significantly different data distributions caused by the completely different working conditions of A2 and A3. The third conclusion is that multi-label, small-shot methods designed for machine learning may not be suitable for fault diagnosis problems. This is due to the data type. In BCR, it works on image datasets, which differ from signal data. A final conclusion is that the number of samples on the support set has a significantly smaller impact on meta-learning-based methods compared to other methods. Despite its poor performance, BCR maintains stability across different sample counts, which is attributed to its precise policy. In meta-learning, most methods predict labels based on distance from the support set, similar to nearest neighbor methods. Since the data is not used in training, the number of samples on the support set has less impact than other methods.

[0226] Table 5 shows the accuracy of all methods on different datasets.

[0227]

[0228]

[0229] The beneficial technical effects achieved by the embodiments of the present invention are as follows:

[0230] Current research on the problem of few-shot coupled fault diagnosis is limited. To address this challenging problem, we treat it as a special case of multi-label few-shot learning and propose a novel graph convolutional method with a metric structure. This method solves the problem through two levels. The first level is a data augmentation level, which controls the rank of the adjacency matrix to obtain more useful information while maintaining the advantages of feature integration learning. Furthermore, a self-loop graph is added to further expand the scale of the training samples. The second level is a feature transformation level. At this level, we leverage the multi-level similarity features in the multi-label problem and design a metric structure for the data in the support and query sets. This structure consists of three constraints that guide the network to learn valuable features, following the assumption that feature distance is consistent with label similarity. Experiments have significantly demonstrated the effectiveness of the method.

[0231] The specific order or hierarchy of steps in the disclosed process is an example of an exemplary method. Based on design preferences, the specific order or hierarchy of steps in the process may be rearranged without departing from the scope of this disclosure. The appended method claims provide elements of various steps in an exemplary order and are not intended to limit us to the specific order or hierarchy stated therein. In the detailed description above, various features are combined together in a single embodiment to simplify this disclosure. This method of disclosure should not be construed as reflecting an intention that embodiments of the claimed subject matter require more features than are explicitly stated in each claim. Rather, as reflected in the appended claims, the invention is in a state with fewer features than all the features of a single disclosed embodiment. Therefore, the appended claims are hereby explicitly incorporated into the detailed description, wherein each claim stands alone as a separate preferred embodiment of the invention.

Claims

1. A coupled fault diagnosis method, characterized in that, include: Step 1: For all known fault samples in the support set and the label corresponding to each known fault sample, construct a graph representing the similarity relationship between the known fault samples. This graph is used as input to the graph convolutional learning network. Step 2: Construct a graph convolutional learning model using the graph-based graph convolutional learning network; Step 3: For the graph corresponding to the support set, three similarity relationships are used as the metric structure for known fault samples, and the corresponding loss function is determined based on the metric structure; Step 4: In the first stage, the graph convolutional learning network is trained using the support set to obtain an intermediate graph convolutional learning network model; in the second stage, the intermediate graph convolutional learning network model is trained using the support set and the query set to obtain the trained graph convolutional learning network model. Step 5: For any coupled fault signal to be identified, use the trained graph convolutional learning network model to output the predicted value of each label. The predicted value of each label is used to indicate whether the fault exists.

2. The coupling fault diagnosis method according to claim 1 further includes: Step 6: For each known fault sample, use Fast Fourier Transform to transform the features of the fault sample from time domain features to frequency domain features to obtain the support set. Here, the known fault samples are multi-label faults; the frequency domain features refer to a known fault sample being composed of a vector of length d, where each specific value in the vector represents a feature of the known fault sample. Step six is ​​executed before step one.

3. The coupling fault diagnosis method according to claim 2, characterized in that, Step one specifically includes: Let the support set be {X} s ,Y s }, X s Let X represent the set of known fault samples. s ∈R n×d Y s X represents s The corresponding label matrix, where labels represent fault types; X s Each row represents the Y corresponding to a known fault sample. s The same row represents the label of the known fault sample; one known fault sample corresponds to at least one label; Y s It is the first vector of 1*c, and the value of the first vector consists of 0 and 1: Y s ∈{0,1} n×c n is the number of known fault samples, and c is the number of label categories; where X s The number of known fault samples within the training sample is less than the number of samples in regular training. For the support set, construct a graph for the graph convolutional learning network: {X s ,Y s The corresponding diagram, {X s ,Y s The corresponding graph represents the similarity relationship between known fault samples, {X} s ,Y s The corresponding graph contains nodes and edges, with each node corresponding to X. s A known fault sample; edges are values ​​in the first adjacency matrix A, representing the connection between nodes. The rows and columns of the first adjacency matrix A are the number of known fault samples, n. The value in the i-th row and j-th column of the first adjacency matrix A represents the similarity between the i-th and j-th known fault samples, with similarity values ​​taking either 0 or 1. A third adjacency matrix is ​​used. This represents a self-loop graph, where each known faulty sample is connected to itself via a self-loop. It is represented as: Where I is the first identity matrix; Construct a degree matrix D, which is a diagonal matrix. The degree matrix D represents the number of connection points for each known fault sample in D, i.e., the number of known fault samples connected to each known fault sample. Each diagonal element of the degree matrix D is the sum of the corresponding rows of the adjacency matrix A, expressed as D = diag(A1), where 1 is a column vector with all elements being 1. to indicate The degree.

4. The coupling fault diagnosis method according to claim 3, characterized in that, Step two includes: The graph convolutional learning network is configured to consist of three base layers with the same structure. The l-th neural network layer f in the GCL network... l The structure is as follows: Among them, X l W is the input to the l-th neural network layer. l Let be the weight matrix of the l-th neural network layer, and δ represent the activation function in the neural network layer; The calculations performed by each neural network layer are represented by formula (1), X l+1 This represents the output of the l-th neural network layer, while X... l+1 It also represents the input of the (l+1)th neural network layer; X l+1 The number of feature dimensions is determined by W l The number of columns determines W. l The column represents the number of neurons in the l-th layer of the neural network; X l+1 and The relationship between rank and rank can be expressed as an inequality: To scale the features, a normalization layer F was added after the first neural network layer. bn f 1 ,f 2 The activation function in all layers is ReLU, and the last neural network layer f 3 The activation function in the graph convolutional learning network model is the sigmoid function. f gcl =sequence[f 1 ,F bn ,f 2 ,f 3 ] (2) The loss function of the graph convolutional learning model can be expressed as: Among them, L ce For cross-entropy loss, It is constructed from the support set labels.

5. The coupling fault diagnosis method according to claim 3, characterized in that, Step one also includes: For known fault samples within the same label class, each known fault sample is configured to connect to only one neighbor. Labels are used to distinguish the adjacency information between known fault samples and their connected neighbors: if the label vector values ​​of a known fault sample and its connected neighbor are exactly the same, then the known fault sample and its connected neighbor belong to the same class; otherwise, the known fault sample and its connected neighbor belong to different classes. The third adjacency matrix is ​​then used to... In this code, the diagonal and opposite elements are set to 1, the bottom right element is set to 1, and all other elements are set to 0. This sets the elements corresponding to the c-th category tag to 1. Represented as Where the superscript 'c' represents the c-th class, The rank of is nc.

6. The coupling fault diagnosis method according to claim 4, characterized in that, Step three specifically includes: For the support set, the similarity relationships of all known fault samples are divided into three categories: the most similar sample pair X. ms Some similar samples for X ps and dissimilar sample pairs X ds ; Construct an objective function based on the feature distance and label consistency of known fault samples. The objective function satisfies the following condition: if the feature distance between two samples is close, then their labels are also similar. The loss function of the objective function is expressed as: dist(f(X ms ))<dist(f(X ps ))<dist(f(X ds )) (11) Where f() represents the feature mapping function, corresponding to [f 1 ,f bn ,f 2 ], dist indicates the distance calculation method; Transform inequality (11) into a solvable form: Equations (12)-(14) serve as a metric structure, using three ratios: and To ensure consistency between the feature metric distance and the label of known faulty samples, where, X represents the most similar sample with label a. ms The maximum distance and partially similar samples X ps The ratio of the average distance between them The smaller the value, the more similar the sample X is. ms The distance between them is greater than that between some similar samples X ps The smaller the distance between them; the subscript 'a' indicates and It is calculated based on the a-th label of the fault sample, indicating that the a-th label is the baseline label; X represents partially similar samples labeled 'a'. ps The average distance between dissimilar sample pairs X ds The ratio of the minimum distances between two samples X is used to make some similar samples X... ps The distance between them is less than that between dissimilar samples X ds The distance between them; X represents the most similar sample with label a. ms The maximum distance between dissimilar sample pairs X ds The ratio of the minimum distances between the two samples is used to find the most similar sample X. ms The distance between them is less than that between dissimilar samples X ds The distance between them; for class c labels, each class label is used as the base label a, and the formulas (12)-(14) are used as the measurement structure; The neural network layer f 1 ,f bn ,f 2 Considered as a feature extraction module, the last neural network layer f 3 Treating it as a classification module, the metric structure constraints (corresponding to formulas (12)-(14)) are applied to the feature extraction module: f 1 ,f bn ,f 2 To strengthen the classification module f 3 Classification ability.

7. The coupling fault diagnosis method according to claim 6, characterized in that, Step three specifically includes: Calculate all baseline labels The corresponding average value is used as the corresponding loss function, expressed as: Where {Y} represents the relationship between Y and Y. s The tag set after removing duplicate values, |{Y}| represents the number of tag categories within {Y}. So, after removing duplicates, {Y} has c types of tags, with each category containing 1 tag.

8. The coupling fault diagnosis method according to claim 7, characterized in that, Step four specifically includes: In the first stage, a support set is used to train the graph convolutional learning network. Combining equations (6), (15), and (17), the first empirical loss function for training the graph convolutional learning model is set as follows: L stage1 =L gcl (X s )+αL mp (X s )+αL pd (X s ) (18) Where α is the first equilibrium parameter; When the first empirical loss function converges to the preset value, the training of the graph convolutional learning network is completed, and the intermediate graph convolutional learning network model is obtained. Let the query set be {X} q ,Y q }, X q Let X represent the set of unknown fault samples. q ∈R N×d Y q X represents q The corresponding label matrix; X q Each row represents an unknown fault sample, Y q The corresponding row represents the label of the unknown fault sample; X q The frequency domain features corresponding to each unknown fault sample are input into the intermediate graph convolutional learning network model, which outputs a probability prediction vector. For each unknown fault sample, there are c predicted values. When the predicted value is greater than the threshold 1 / c, it indicates that the unknown fault sample has the label corresponding to the predicted value. Through all X... q The label corresponding to each unknown fault sample within Y constitutes Y q Y q It is the second vector of 1*c, and the value of the second vector consists of 0 and 1.

9. The coupling fault diagnosis method according to claim 8, characterized in that, Step four specifically includes: In the second stage, the intermediate graph convolutional learning network model is trained using a support set and a query set; combining equations (6), (15), (16), and (17), a second empirical loss function for training the graph convolutional learning model is set: L stage2 =L gcl (X s )+βL mp (X q )+βL pd (X q )+γL md (X q ) (19) Where β and γ are the second and third balance parameters, respectively, and the second balance parameter β is less than the third balance parameter γ. The first term on the right side of equation (19) is to use the support set to maintain the accuracy of the classifier. The second, third and fourth terms are used to learn the metric structure of the query set. When the second empirical loss function converges to the preset value, the training of the intermediate graph convolutional learning network model is completed, and the trained graph convolutional learning network model is obtained.

10. A coupled fault diagnosis system, characterized in that, include: The graph construction unit is used to construct a graph representing the similarity relationship between known fault samples for all known fault samples in the support set and the label corresponding to each known fault sample. The graph is used as input to the graph convolutional learning network. Model construction unit, used to construct a graph convolutional learning model using the graph-based graph convolutional learning network; The metric structure construction unit is used to use three similarity relationships as the metric structure for known fault samples for the graph corresponding to the support set, and to determine the corresponding loss function based on the metric structure. The training unit is used in the first stage to train the graph convolutional learning network using the support set to obtain an intermediate graph convolutional learning network model. In the second stage, the intermediate graph convolutional learning network model is trained using the support set and the query set to obtain the trained graph convolutional learning network model. The fault identification unit, for any coupled fault signal to be identified, uses a trained graph convolutional learning network model to output a predicted value for each label. The predicted value for each label is used to indicate whether the fault of that type exists.