A Few-Shot Fault Incremental Learning Method Based on Multi-View Shapelet Prototype Network

Through the incremental learning method of small sample failure based on multi-view shapelet prototype network, the problem of adapting to a small number of new faults in industrial fault diagnosis is solved, and better generalization ability and computing efficiency are achieved.

CN118586520BActive Publication Date: 2025-05-30ZHENGZHOU UNIVERSITY OF LIGHT INDUSTRY
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202410720728.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-06-05
Publication Date
2025-05-30
Estimated Expiration
2044-06-05

AI Technical Summary

Technical Problem

The prior art is difficult to effectively adapt to a small number of new failures in industrial fault diagnosis, resulting in catastrophic forgetting and overfitting problems, and computational complexity and inefficiency during incremental learning.

Method used

The multi-view shapelet prototype network is adopted to build a generalized feature space through the multi-view element learning framework, and the multi-view shapelet prototype classifier is used to enhance the generalization ability of shapelets, and the multi-view information is fused through the multi-view element correction module to reduce overfitting and forgetting.

Benefits of technology

The ability to adapt to a small number of new failures without forgetting previous failures is achieved, reducing the risk of catastrophic forgetting and overfitting, and improving computational efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118586520B_ABST
    Figure CN118586520B_ABST
Patent Text Reader

Abstract

The present invention discloses a few-shot fault incremental learning method based on a multi-view shapelet prototype network, which includes the following steps: Step 1, obtain an initial training sample D0, where the initial training sample D0 includes a training set and a validation set. Step 2, initialize the training set, and then input it into the multi-view shapelet prototype network for pre-training to obtain a pre-trained multi-view shapelet prototype network; the output of the multi-view shapelet prototype network includes a multi-view shapelet set. Step 3, perform meta-training on the pre-trained multi-view shapelet prototype network to obtain a final multi-view shapelet prototype network and a meta-correction module. The present invention horizontally links and fuses multi-perspective information, so the semantic gap between the old-class prototypes and the new-class prototypes is bridged. In short, the extracted multi-perspective small shapes contain some key information for discriminating waveforms and can identify new faults of any duration.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fault diagnosis, and particularly to a few-shot fault incremental learning method based on a multi-view shapelet prototype network. Background Technique

[0002] Fault diagnosis can detect abnormal states in the production process and find the root causes of faults. Timely diagnosis of industrial faults is the key to maintaining the safe, efficient, and economic operation of complex industrial systems. In recent years, data-driven methods, especially deep learning methods, have played an important role in various fault diagnosis models. However, for deep learning methods to obtain reliable modeling performance, there are two basic prerequisites: 1) A sufficient number of training data. The lack of sufficient data will lead to poor generalization and overfitting. 2) All fault categories are known a priori. When encountering new faults, the trained model will undergo a complete retraining process, which will result in high computational costs. Unfortunately, due to the dynamics of the environment and operations, new fault types will continuously appear in the actual continuous production process. In addition, due to the high cost of manual annotation, the number of new fault samples is small. Therefore, it is of great practical significance to fine-tune the basic fault diagnosis model based on a small number of new fault samples to diagnose all the faults encountered.

[0003] Currently, many studies have been conducted to solve the problem of fault diagnosis under industrial few-shot samples. For example, Wang et al. proposed a synthetic minority over-sampling method to deal with the imbalanced classification problem in the field of mechanical fault diagnosis. Xu et al. proposed a bearing dynamic model to simulate vibration signals and enhance few-shot samples. Hu et al. proposed a convolutional conditional variational auto-encoding method to generate synthetic samples to supplement scarce driving data. However, all these methods can only solve the overfitting and data imbalance problems related to few-shot samples. They ignore the high time complexity and catastrophic forgetting problems brought about by retraining the model when the model incrementally diagnoses new faults.

[0004] Incremental learning is an adaptive algorithm that avoids retraining from scratch when new data arrives and enables positive transfer of knowledge. Some studies have incrementally learned new faults by expanding the network or aligning features. However, when these models fit too closely to new faults with few samples, overfitting occurs, resulting in the loss of the ability to generalize to a large number of unseen test samples. Gu et al. proposed the Imbalance-corrected Convolutional Neural Network (II-CNN) to simultaneously address overfitting and catastrophic forgetting problems. II-CNN uses branch expansion to construct new networks for new fault patterns and adopts an imbalance correction method to solve the sample imbalance problem. However, as new faults continue to emerge, II-CNN requires additional memory to store the expanded network modules, leading to an increase in computational complexity and a decrease in computational efficiency. Ren et al. proposed a Heterogeneous Sample Enhancement Network with lifelong learning capabilities. The heterogeneous DA subnet in HSELL-Net is used to adjust the feature distribution, while the lifelong subnet is used to enhance small samples. However, HSELL-Net requires additional heterogeneous samples and cannot obtain learning capabilities from existing homogeneous samples. In industrial production processes, there are some common faults with sufficient samples. Therefore, it is natural to establish a basic model based on common faults. For existing fault diagnosis methods, how to make the basic model continuously adapt to new faults with few samples without catastrophic forgetting and overfitting is an unsolved problem. We name this problem Few-shot Fault Incremental Learning (FSFIL).

[0005] Recently, Tao et al. first proposed the Few-shot Class Incremental Learning (FSCIL) method, which can learn new tasks from only a few labeled samples without forgetting previously learned tasks. Meta-learning can acquire prior knowledge from experience and enable the model to achieve good generalization on limited labeled data. Zhou et al. proposed the Meta-learning based FSCIL Multi-stage Incremental Task (LIMIT). LIMIT learns transferable features in the base task with rich samples and uses these features to classify new classes with few samples. The concept of FSCIL provides a possible solution to the FSFIL problem. However, most existing FSCIL tasks focus on animal image classification. The visual features of animals include ears and noses, which are stable and easy to transfer. Different from these visual features, the industrial fault features reflected by industrial signals are uncertain, highly affected by noise, and unexplainable. At the same time, due to the emergence of new faults caused by changes in the environment and operations, the durations of different working conditions also vary. Therefore, it is a challenge to use existing FSCIL methods to diagnose new faults by transferring the features of seen faults.

[0006] Explanation of terms:

[0007] k-means algorithm: The k-means clustering algorithm is an iterative clustering analysis algorithm.

[0008] z-score method: Also known as the standard score method, it is a statistical concept used to describe the position of a value relative to the mean of the entire data set.

[0009] Soft minimum distance: That is, the soft minimum function, a distance proposed in the paper "Composition of Control Barrier Functions With Differing Relative Degrees for Safety Under Input Constraints". Summary of the Invention

[0010] The purpose of the present invention is to provide a few-shot fault incremental learning method based on a multi-view shapelet prototype network to solve the problems proposed in the above background technology.

[0011] To achieve the above purpose, the present invention provides the following technical solutions:

[0012] A few-shot fault incremental learning method based on a multi-view shapelet prototype network includes the following steps:

[0013] Step 1, obtain the initial training sample D 0 , the initial training sample D 0 includes the training set and the validation set

[0014] Step 2, initialize the multi-view shapelet of the training set, and then input it into the multi-view shapelet prototype network for pre-training to obtain a pre-trained multi-view shapelet prototype network; the output of the multi-view shapelet prototype network includes the multi-view shapelet set

[0015] Step 3, perform meta-training on the pre-trained multi-view shapelet prototype network to obtain the final multi-view shapelet prototype network and the meta-correction module.

[0016] For further improvement, in the second step, the method for initializing the multi-view shapelet of the training set is as follows: Use the k-means algorithm for the training set of the initial training sample D 0 of the L vCluster all subsequences of the length, and obtain the cluster centers, which are used to initialize the shapelet subset of the v-th view, and finally obtain the shapelets of multiple views after initialization; L v = L×g×v, where L represents the length of the time series shapelet, and g represents the length change ratio.

[0017] For further improvement, in the second step, the steps of pre-training are as follows:

[0018] 2.1) Input the training set into the multi-view shapelet prototype network, and map the m-th sample X in the initialized training set t,m to a latent space based on the multi-view shapelet set to obtain the representation of the corresponding multi-view shapelet set where, where t represents the t-th task stage, and in the pre-training stage t = 0; is the set of representations of the multi-view shapelet sets for all task stages and represents the set of representations of the multi-view shapelet set in the t-th task stage; is the set of representations;

[0019]

[0020] where,

[0021] represents the φ-th subsequence in the m-th sample X in the training set t,m in the t-th task stage; there are N - L + 1 subsequences in the m-th sample X t,m in the t-th task stage, where N represents the length of the input time series, and L represents the length of the time series shapelet; D() represents the calculation of the soft minimum distance between S k,v and X t,m ; represents the soft minimum distance between S k,v of k dimensions and X t,m , and S k,v represents the k-th shapelet of the v-th view; represents the l-th data point in the k-th shapelet of the v-th view, represents S of the v-th view k,v and the subsequence of X t,m ​ The cosine distance between, d() represents calculating the cosine distance, exp() represents the natural exponential function, and α represents controlling the precision; denotes the training set the m-th sample X in t,m the φ-th subsequence in; denotes the l-th data point in;

[0022] 2.2) Calculate the set of multi-view prototypes for all task phases where denotes the set of multi-view prototypes for the t-th task phase;

[0023]

[0024] where, |() represents the indicator function, K represents the number of instances in the i-th class training set, p v,i,t denotes the class prototype for the v-th view in the t-th task phase; denotes the set of prototypes of all seen classes for the v-th view in the t-th task phase; denotes the number of all seen classes in the t-th task phase, denotes the training set for the t-th task phase, M t denotes the number of samples in the t-th task; D t,m,v denotes the S in the v-th view of the m-th sample in the t-th task phase k and X t,m cosine distance; K represents the total number of dimensions;

[0025] 2.3) Update the multi-view shapelet set by minimizing the pre-training loss function Loss t (D t ,P t ,Y t ) The pre-training loss function is as follows:

[0026]

[0027] where, |P t | represents the number of class prototypes in the t-th task phase, P t denotes the set of class prototypes in the t-th task phase, Y t denotes the set of labels in the t-th task, S represents the shapelet set, Loss t () represents the loss in the t-th task phase; V represents the number of views, Y t,mdenotes the label of the m-th sample in the training set of the t-th task phase, exp() represents the natural exponential function, D t,m,v denotes the t-th task, the distance feature mapped by the shapelet set of the v-th view for the m-th sample, and τ is the temperature parameter; M t denotes the training set X t the number of samples in;

[0028] 2.4) Repeat steps 2.1) to 2.3) until a preset number of loops is reached to obtain a pre-trained multi-view shapelet prototype network.

[0029] For further improvement, in step three, the steps for meta-training are as follows:

[0030] 3.1) Initialize the multi-view shapelet set through the pre-trained multi-view shapelet prototype network

[0031] 3.2) Sample {S 0 ,..., S 1 ,..., S t ,..., S Z ; Q 1 ,..., Q t ,..., Q Z} from D t denotes the query set of the t-th fake task, used to simulate the dataset of the t-th task phase, and Q t denotes the support set of the t-th fake task, used as a simulation of the validation set of the t-th task phase;

[0032] t = 1, 2,... Z;

[0033] 3.3) Calculate the shapelet-based representation of the query set through Equation (1)

[0034] 3.4) Calculate the multi-view class prototype of the query set through Equation (2)

[0035] 3.5) Take and as the input of the meta-correction module, and correct to obtain the outputs and

[0036] 3.6) Update and to update and {M 1 ,... M V};

[0037] 3.7) Repeat steps 3.2) to 3.6) until the preset number of cycles is reached to obtain the final multi-view shapelet prototype network and the final meta-correction module.

[0038] For further improvement, in step 3.5), the data processing method of the meta-correction module is as follows:

[0039]

[0040] Where:

[0041] is the key of the query matrix K in the adjacent view v-1 z and is the query matrix in the adjacent view v-1, is the value of the query matrix in the adjacent view v-1, MHA v () represents multi-head self-attention, FCL v represents the fully connected layer, LN v is the normalization layer, and Output is the output.

[0042] For further improvement, in step 3.6), update and by minimizing equation (5) and {M 1 ,...M V}

[0043]

[0044] is the shapelet set, M is the correction function, M V is the correction function of the Vth view, t is the t-th fake task, S t is the query set of the t-th fake task, Q t is Q t represents the support set of the t-th fake task, Loss t () represents the loss function of the t-th task, X j , Y j are respectively the j-th sample of the support set of the t-th fake task and the label of the j-th sample.

[0045] Compared with the prior art, the beneficial effects of the present invention are:

[0046] 1) The FSFIL problem is first proposed: that is, how to continuously adapt the basic model to a small number of new faults without forgetting the previously existing faults.

[0047] 2) The MSPN algorithm is proposed for FSFIL. A multi-view meta-learning framework is used to construct a generalizable feature space, and a multi-view shapelet prototype classifier is used to enhance the generalization ability of shapelets.

[0048] 3) A multi-view meta-calibration module composed of transformers is proposed to fuse multi-view information and calibrate all prototypes into a distinguishable space.

[0049] 4) Multi-view information is fused through horizontal connections. Therefore, the semantic gap between old-class prototypes and new-class prototypes is bridged. In summary, the extracted multi-view shapelets contain some key information of discriminant waveforms and can identify new faults of any duration. Description of the Drawings

[0050] Figure 1 is a false incremental fault diagnosis setup diagram;

[0051] Figure 2 False incremental task segmentation diagram;

[0052] Figure 3 is a data processing framework diagram of the multi-view shapelet prototype network. Detailed Implementation Manner

[0053] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0054] I. Problem Definition and Predefinition

[0055] A. Few-shot Fault Incremental Learning

[0056] Industrial faults can be recorded by time series segments. Therefore, in this patent, the input of data is time series segments. In the industrial data stream, a set of non-overlapping classes {C 0 ,..., C t} is continuously obtained in a step-by-step manner. The classifier in this patent is required to be able to classify all previous classes at the current stage. For example: as Figure 1 shown, C 0 = {Fault1,..., Fault 4}, C 1 = {Fault 5, Fault 6}. At task stage 0, there are a large number of training samples D 0. In the next task t phase, there are only a small number of samples in each phase. These tasks can be expressed as: {D t , t = 1, 2,...}. Once the model is trained in each t task, its performance is evaluated in all previously occurring classes C t = C 0 ∪... C t .

[0057] The dataset for each phase is composed of a training set and a test set . The training set is used to update the model, and the test set is used to test the model effect. The training set contains (X t , Y t ), train , and the test set contains (X t , Y t ). test .

[0058] denotes the input time series dataset, M t denotes the number of samples, and N denotes the length of the input time series.

[0059] denotes 's label. C t and C t respectively denote and 's label space. In the case of θ ≠ σ, Moreover, the samples of the finite data D t are reorganized into the form of W-way O-shot data, where W represents the number of classes of D t , and O represents the number of samples in each class.

[0060] B. False incremental tasks

[0061] In the meta-learning phase, we propose to use false incremental tasks to find long-term transferable embeddings. The synthetic false incremental tasks are designed to simulate the same data format as the true incremental tasks. Therefore, the shapelets optimized from the false incremental tasks can be effectively applied to the incremental fault diagnosis phase. The base class set C 0 is divided into two non-overlapping sets: where represents the false base class set, represents the false incremental class set. The false incremental class set is divided into z phases to simulate the incremental session, that is, the data in each phase is the same as Dt The data format of the W-way O-shot is similar, sampling instances from the O samples to construct a pseudo-training set. Thus, a support set is formed: S 1 , S 2 ,..., S Z and a query set Q 1 , Q 2 ,..., Q Z . The query set Q in the z-th stage z is from all visible classes to form a data set. The instances in the query set should not appear in the support set, that is Figure 2 Details the relationship between the incremental task and the pseudo-incremental task.

[0062] C. Time series shapelet

[0063] (1) shapelet: A shapelet is defined as S ∈ R K×L , where K represents the number of shapelets and L represents the length of the shapelet. In this patent, the shapelets are parameters to be optimized. They are first initialized and then obtained by optimizing the loss function. The k-th shapelet can be represented as

[0064] (2) shapelet transformer: It maps the sample X t,m to a shapelet-based latent space formed by the distance feature .

[0065]

[0066] Among them,

[0067] where α is introduced to control the precision, is the distance between the subsequence k,v of S t,m and X . φ represents the t,m -th subsequence in X Each X t,m always has N - L + 1 subsequences.

[0068] II. Overall architecture

[0069] In this section, we propose the MSPN for FSFIL in industrial processes. The entire framework is asFigure 3 As shown in the figure. First, initialize the multi-view shapelet in the pre-training stage. Second, in the multi-view meta-training stage, optimize the multi-view shapelet prototype classifier and the multi-view meta-correction module in the false increment task. The multi-view shapelet prototype classifier classifies samples using the multi-view shapelet and prototypes. The multi-view meta-correction module is used to fuse multi-view information to obtain incremental diagnostic capabilities. Third, the optimized multi-view shapelet prototype classifier and the optimized multi-view meta-correction module are used for incremental fault diagnosis in new tasks.

[0070] A. Multi-view shapelet prototype classifier

[0071] The multi-view shapelet and prototypes are the main components of the multi-view shapelet prototype classifier. After initializing the multi-view shapelet, the training dataset or the support set S p is embedded into the shapelet-based embedding space using the multi-view shapelet. Calculate the multi-view prototype in the shapelet-based embedding space. This prototype classifier classifies new samples using the similarity between the new samples and the prototypes.

[0072] 1) Shapelet initialization and embedding:

[0073] Since shapelets of different lengths contain complementary information, we consider the shapelet-based representations learned from shapelets of different lengths as different views. In this paper, denotes a set of shapelets with V different views. denotes the subset of shapes of the v-th view. Among them, L v = L×g×v, so two hyperparameters: the base length (L) and the length variation ratio (g) control the lengths of all shape subsets. For example, assume g = 0.5 and L = 20, then the length of the first view is 10. For each view v, the k-means clustering method is used to initialize the shapelet in the pre-training stage. Select the clustering centers of all subsequences with length L v as the initialized small shapes. Through the multi-view shapelet converter, we can obtain the multi-view shapelet representation where D t,v = D(S v,k ,Δ) is calculated according to formula (1). Δ represents the t-th training dataset or the support set S t .

[0074] 2) Class prototypes:

[0075] In small-sample fault diagnosis, due to limited samples, an overfitting problem is caused. Due to stronger generalization ability and requirements for less labeled data, prototype networks show competitiveness in anti-overfitting. Class prototypes represent the shape-based representative features of this class. Class prototypes of different views show representative features from different perspectives. The class prototype of the i-th class and the v-th view in the t-th task is calculated in the formula

[0076]

[0077] Here, |(·) represents the indicator function, and K represents the number of instances in the training set of the i-th class. We can easily replace the classifier belonging to the i-th class with the corresponding prototype: pv,i,t. represents the set of prototypes of all seen classes in the v-th view and the t-th task. Since equation (2) does not change the pre-trained embedding and makes it tend to fewer new instances, the overfitting phenomenon is alleviated.

[0078] B. Multi-view Meta-Calibration Module

[0079] Incremental learning needs to gradually identify new classes. After the training of a task is completed, the trained shapelets are specific descriptions of the class features of this task. Therefore, when the model faces a new task, there will be a semantic gap between the prototypes of the old classes and the prototypes of the new classes. We need to correct it. In this part, a multi-view transformer is used to fuse multi-view information and correct the prototypes of the old classes and the prototypes of the new classes in the fake incremental task. If a model can correct them in the fake incremental task, then it can also correct them in the real incremental task.

[0080] In this paper, the calibration function of the v-th view is denoted as M v (·). To find the autocorrelation between the new class features of the class prototypes P z,v and D z,v in the z-th task and the v-th view, [P z,v , D z,v is used as the input of M v (·). The equation of the correction process can be expressed as Since the Transformer has permutation invariance and is good at generating context-adjusted embeddings even when dealing with long-range dependencies. Specifically, using the self-attention mechanism in the transformer, the features of the query instance are calibrated by considering all class prototypes. The query Q z,v , key K z,v and value V z,v are three important input matrices of the transformer. In this paper, Q z,v =K z,v =Vz,v = [P z,v , D z,v . First, project these matrices linearly into a ρ-dimensional feature space. For the input matrix K z,v , the ρ-dimensional feature is equal to i.e., Q z,v and V z,v are the same. These projection matrices are calculated by formula (3).

[0081]

[0082] Here, MHA z represents multi-head self-attention, LN z is layer normalization, and FCL z is a fully connected layer. Finally,[[]]

[0083]

[0084] Inspired by the cross-view fusion method, the encoding equation in formula (3) is changed to formula (4), which fuses the information of the adjacent view v-1.[[]]

[0085]

[0086] is the key of the query matrix in the adjacent view v-1, is the query matrix in the adjacent view v-1, is the value of the query matrix in the adjacent view v-1, and MHA v () represents multi-head self-attention, FCL v represents a fully connected layer, LN v is the normalization layer, and Output is the output. As Figure 3 shown, we fuse the information between two adjacent views in turn: v-1 and v. The information in the view v-1 with a short shapelet is fused into the view v with a longer shapelet. Generally, the calibration function of the 0th view is formula (3), and the calibration function of the vth view (v > 0) is formula (4).

[0087]

[0088]

[0089] C. Optimization

[0090] This section optimizes the multi-view shapelet prototype classifier in two stages: pre-training and multi-view meta-training. The multi-view meta-correction module is only optimized in the meta-training stage. Then, the optimized multi-view shapelet and multi-view meta-correction module are used for incremental training. The pre-training and multi-view meta-training stages are given in Algorithm 1. The loss functions for the pre-training and meta-learning stages are as follows:

[0091] 1) Loss function in the pre-training stage: Since the prototype classifier classifies new data based on the similarity between prototypes and data points, an embedding network with stronger generalization ability is required. A good embedding network needs to minimize the within-class variance and maximize the between-class separation. Therefore, we enforce the shape-based embeddings to be close to their corresponding class prototypes by minimizing the prototype classification loss:

[0092]

[0093] where Loss t represents the loss of the t-th task, τ is the temperature parameter, F(·,·) is the cosine similarity, and |P t | represents the number of prototypes in the t-th task. Equation (5) learns the multi-view shapelet by averaging the losses of multiple views.

[0094] 2) Loss function in the meta-learning stage: In this paper, a meta-training method is adopted to learn the multi-view meta-correction module and generalize this module to actual incremental tasks. Therefore, the loss function in Equation (5) is changed to:

[0095]

[0096] where is the shapelet set, M is the correction function, MV is the correction function of the V-th view, t is the t-th fake task, S t is the query set of the t-th fake task, Q t is the support set of the t-th fake task denoted as Qt, Loss t () represents the loss function of the t-th task, X j , Y j are the j-th sample of the support set of the t-th fake task and the label of the j-th sample respectively. Since Equation (6) extracts the inductive bias into M v (·), M v (·) will calibrate the real incremental task well by being well-suited to the fake incremental tasks.

[0097] D: Incremental Fault Diagnosis

[0098] In the previous section, we discussed the basic task D 0The process of training the multi-view shapelet prototype network and the multi-view meta-correction module. This part discusses the incremental fault diagnosis phase in the incremental task {D t , t = 1, 2, …}. Incremental fault diagnosis includes two parts: incremental training and incremental testing. In the incremental training phase, the fault diagnosis model is updated, and in the incremental testing phase, the performance of the model in diagnosing all the seen faults is tested. It should be noted that in the incremental training phase, only the prototype set is updated without optimizing the model parameters.

[0099] 1) Incremental training: Facing the introduced task update the current classifier by enriching the prototype set in the v-th view and the (t - 1)-th task where W×(t - 1)+|C 0 | is the number of classes seen in the (t - 1)-th task.

[0100] 2) Incremental testing: We test the performance of the model on all the classes seen so far in this part. First, embed the test data set into the shape-based feature space through the multi-view learned in the multi-view shapelet prototype network. Given a test instance calculate the shape-based feature D through formula (1) v,t,m , and then assign the predicted label to the class most similar to the prototype. Since the information of all views is fused into the last v-th view, we choose the predicted value of the last view as the final predicted value. Therefore, set as the final predicted value.

[0101] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A few-shot fault incremental learning method based on multi-view shapelet prototype network, characterized in that: The steps include: Step 1: Get the initial training sample D 0 , initial training sample D 0 Including training set and validation set Step 2: Initialize the multi-view shapelet of the training set, and then input the multi-view shapelet prototype network for pre-training to obtain a pre-trained multi-view shapelet prototype network; the output of the multi-view shapelet prototype network includes the multi-view shapelet set The method for initializing the multi-view shapelet of the training set is as follows: the k-means algorithm is used to initialize the initial training sample D 0 The training set Medium v All subsequences of length L are clustered, and the obtained cluster centers are used to initialize the shapelet subset of the vth view, and finally the shapelets of multiple views after initialization are obtained; v =L×g×v, L represents the length of the time series shapelet, g represents the length change ratio; Step 3: Meta-train the pre-trained multi-view shapelet prototype network to obtain the final multi-view shapelet prototype network and meta-correction module.

2. The method for incremental learning of few-sample faults based on multi-view shapelet prototype network as claimed in claim 1, characterized in that: In step 2, the pre-training steps are as follows: 2.1) The training set Input the multi-view shapelet prototype network and convert the initialized training set through the shapelet converter The mth sample X in t,m Mapped to a multi-view shapelet set The latent space of , and the corresponding multi-view shapelet set Representation in, Where t represents the tth task stage, and t = 0 in the pre-training stage; Multi-view shapelet sets for all task stages The set of representations of Represents the multi-view shapelet set of the t-th task stage The set of representations of ; in, Represents the training set of the tth task stage The mth sample X in t,m The φth subsequence in the tth task stage; the mth sample X t,m There are N-L+1 subsequences in the time series, where N represents the length of the input time series and L represents the length of the time series shapelet; D() represents the calculation of S k,v and X t,m Calculation of the soft minimum distance; S represents k-dimensional k,v and X t,m The soft minimum distance, S k,v represents the k-th shapelet of the v-th view; represents the lth data point in the kth shapelet of the vth view, S represents the vth view k,v and X t,m subsequence of The cosine distance between them, d() represents the calculated cosine distance, exp() represents the natural exponential function, and α represents the control accuracy; express The lth data point in ; 2.2) Calculate the set of multi-view prototypes for all task stages in represents the set of multi-view prototypes at the t-th task stage; Among them, |() represents the indicator function, K represents the number of instances in the i-th class training set, and p v,i,t The class prototype representing the vth view of the tth task stage; represents the prototype set of all classes seen in the vth view at the tth task stage, represents the number of all classes seen in the t-th task stage, represents the training set of the t-th task stage, M t represents the number of samples of the tth task; D t,m,v represents the S of the vth view in the mth sample at the tth task stage k and X t,m The cosine distance of K represents the total number of dimensions; 2.3) By minimizing the pre-training loss function Loss t (D t ,P t ,Y t ) Update the multi-view shapelet set The pre-training loss function is as follows: Among them, |P t | represents the number of class prototypes in the tth task stage, P t represents the set of class prototypes in the tth task stage, Y t represents the set of labels in the tth task, S represents the shapelet set, and Loss t () represents the loss of the t-th task stage; V represents the number of views, Y t,m represents the label of the mth sample in the training set of the tth task stage, exp() represents the natural exponential function, and D t,m,v represents the distance feature of the mth sample mapped by the shapelet set of the vth view for the tth task, τ is the temperature parameter; M t Represents the training set X t The number of samples in ; 2.4) Loop step 2.1) to step 2.3) until the preset number of cycles is reached to obtain the pre-trained multi-view shapelet prototype network.

3. The method for incremental learning of few-sample faults based on multi-view shapelet prototype network as claimed in claim 2, characterized in that: In step 3, the steps of performing meta-training are as follows: 3.1) Initialize the multi-view shapelet set through the pre-trained multi-view shapelet prototype network 3.2) From D 0 Sampling 1 , ..., S t ..., S Z ;Q 1 , ..., Q t ..., Q Z },S t represents the query set of the t-th fake task, which is used to simulate the dataset of the t-th task stage. t represents the support set of the t-th fake task, which is used as the validation set for simulating the t-th task stage; t=1, 2, ... Z; 3.3) The shapelet-based representation of the query set is calculated by formula (1) 3.4) Calculate the multi-view class prototype of the query set through formula (2) 3.5) and As input to the meta-correction module, the correction output is and 3.6) According to and renew and {M 1 ,...M V }; 3.7) Repeat steps 3.2) to 3.6) until the preset number of cycles is reached to obtain the final multi-view shapelet prototype network and the final meta-correction module.

4. The method for incremental learning of small-sample faults based on a multi-view shapelet prototype network as claimed in claim 3, characterized in that: In step 3.5), the data processing method of the meta-correction module is as follows: in: is the key of the query matrix in the adjacent view v-1, is the query matrix in the adjacent view v-1, is the value of the query matrix in the adjacent view v-1, MHA v () indicates multi-head self-attention, FCL v Represents the fully connected layer, LN v is the normalization layer, and Output is the output.

5. The method for incremental learning of few-sample faults based on multi-view shapelet prototype network as claimed in claim 4, characterized in that: In step 3.6), according to and Minimize (5) to update and {M 1 ,...M V } is the shapelet set, M is the correction function, M V is the correction function of the Vth view, t is the tth false task, S t is the query set of the tth false task, Q t represents the support set of the tth false task, Loss t () represents the loss function of the tth task, X j , Y j are the j-th sample and the label of the j-th sample in the query set of the t-th fake task, respectively.

Citation Information

Patent Citations

  • Power transformation equipment fault prediction method and device, electronic equipment and readable medium

    CN118133127A