Small sample transformer fault diagnosis method based on space-time noise filtering twin network
By building a spatiotemporal noise filtering twin network, ARNFU and BiGRU are used to extract the vibration signal characteristics of the power transformer, solving the fault diagnosis problem under small sample conditions, and achieving efficient power transformer fault identification.
Patent Information
- Application Number
- CN202510353813.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-25
- Publication Date
- 2025-07-11
AI Technical Summary
Existing deep learning methods rely on a large number of training samples in power transformer fault diagnosis, resulting in low diagnostic accuracy and poor generalization under small sample conditions.
A twin network based on spatiotemporal noise filtering is built, including an adaptive residual noise filtering unit (ARNFU) and a bidirectional gated cycle unit (BiGRU), which is used to extract the timing characteristic information of the vibration signal of the power transformer, and to measure the sample similarity through the European distance, and to build a twin network for fault diagnosis.
Under small sample conditions, the accurate identification of power transformer faults is achieved, which improves diagnostic accuracy and generalization capabilities, and is suitable for limited data situations in actual projects.
Smart Images

Figure CN120296493A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of power transformer fault diagnosis, and particularly to a small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering siamese network. Background Art
[0002] Power transformers are one of the main equipment in power plants and substations. The operating state of power transformers is related to the stability of the entire new power system and the energy Internet. If a fault occurs, it will cause local or large-scale power outages, resulting in huge economic losses. Therefore, accurate and timely diagnosis of transformer faults is of great significance for the safe operation of the new power system. Mechanical faults are a common type of fault in power transformers, and vibration signal analysis can effectively identify faults. Traditional fault diagnosis methods use Fourier transform, wavelet transform, empirical mode decomposition, etc. for fault identification. These methods require manual feature extraction and have certain limitations.
[0003] With the great success of deep learning theory in the fields of computer vision, speech recognition, natural language processing, etc., deep learning methods have been applied to the field of fault diagnosis. However, deep learning-based methods rely on a large number of training samples. In production practice, due to various objective factors, it is sometimes impossible to collect sufficient power transformer fault signals. This leads to the difficulty of fully training the deep learning model, resulting in problems such as low diagnostic accuracy and poor generalization. Therefore, researching a power transformer fault diagnosis method under small-sample conditions is of great significance for accurately identifying the health status of equipment with limited training samples.
[0004] In summary, the present invention intends to propose a small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering siamese network to solve the problem of limited power transformer fault samples. Summary of the Invention
[0005] The technical problem to be solved by the present invention is to provide a small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering siamese network for the defects in the prior art.
[0006] The technical solution adopted by the present invention to solve its technical problems is as follows: The present invention provides a small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering siamese network, and the method includes: Construct an adaptive residual noise filtering unit, denoted as ARNFU, in which the noise signal is reduced by a semi-soft threshold function; construct a bidirectional gated recurrent unit, denoted as BiGRU, to mine the temporal feature information of the vibration signal of the power transformer, and capture the long-term dependence information of the vibration signal by introducing a bidirectional structure; construct a spatio-temporal noise filtering siamese network model with ARNFU and BiGRU as feature extractors for power transformer fault diagnosis under limited data conditions; Collect the fault data of various states of the power transformer, preprocess it and divide it into a training set and a test set. Use the spatio-temporal noise filtering siamese network model constructed with random parameter initialization. Randomly select two samples from the training set to form a sample pair and input it into the model for training. Use the test set to evaluate the training results to obtain a trained spatio-temporal noise filtering siamese network model; perform transformer fault diagnosis through the obtained model.
[0007] Furthermore, the method for constructing the bidirectional gated recurrent unit in the present invention specifically includes: In ARNFU, the noise signal is reduced by a semi-soft threshold function, and the input signal is first trained in two convolutional layers; the output of the second convolutional layer passes through branch 1 and branch 2 at the same time; in branch 1, after the input vector is processed by taking the absolute value, it passes through the Avg_pool layer, the BN layer and the FC layer in sequence to obtain a scaling parameter; subsequently, multiply the scaling parameter by the average value of Avg_pool to obtain a threshold; in branch 2, the input vector passes through the GAP layer and two FC layers in sequence to obtain an adaptive correction factor; finally, use the adaptive correction factor to further correct the output of the semi-soft threshold function so that the entire network can retain more effective signals.
[0008] Furthermore, the semi-soft threshold function in the present invention is expressed as: (1) wherein, x is the input; y is the output; τ is the threshold; The convolutional layer operation is expressed as: (2) wherein, and are the input and output of the l th layer respectively; is from the l -1th layer, the i th neuron to the l th layer, the i th neuron's convolution kernel, where j is the number of convolution kernels; is the number of neurons in the l -1th layer; is the bias of the l layer and the i th neuron; 1Dconv is the convolution operation; The calculation process of the BN layer is expressed as: (3) (4) (5) (6) In the formula, and are respectively the i th data value and the normalized value in the dataset; and are respectively the average value and the standard deviation of the dataset; is an infinitesimal constant; and are respectively the scaling factor and the offset factor of the BN layer operation; is the output of the i th data passing through the BN layer.
[0009] Furthermore, the method of the present invention further includes: scaling the scaling parameter to the range of (0, 1) through the sigmoid function, and using the ReLU activation function to alleviate the problem of gradient disappearance that is prone to occur in the sigmoid function; The ReLU activation function is expressed as: (7) The sigmoid function is described as: (8) In the formula, is the scaling parameter; c is the output of the FC layer in branch 1; Multiply the scaling parameter by the x average value to obtain the threshold, which is described as: (9) In the formula, is the threshold, x is the input signal.
[0010] Furthermore, the present invention further corrects the output of the semi-soft threshold function by using an adaptive correction factor, and the output is expressed as: (10) In the formula, is the adaptive correction factor.
[0011] Further, the method for constructing a spatio-temporal noise filtering Siamese network model with ARNFU and BiGRU as feature extractors specifically includes: The spatio-temporal noise filtering Siamese network model maps two samples ( X 1, X 2) to a low-dimensional feature space and calculates the Euclidean distance between the two feature vectors d ( X 1, X 2), and measures the similarity degree between samples through the distance, expressed as: (11) In the formula f ( X 1) and f ( X 2) are the low-dimensional feature vectors obtained after the samples X 1 and X 2 pass through the feature extractor respectively; is the similarity distance; represents the two-norm operation; The calculated distance is input into the FC layer using the sigmoid activation function, and the probability value is obtained to represent the similarity between the input sample pairs; the larger the probability value, the more similar the two samples are, and the greater the possibility of belonging to the same category; described as: (12) In the formula, sigmoid (·) and FC (·) represent the sigmoid function and the fully connected layer function respectively; represents X 1 and X 2 whether they belong to the same category probability.
[0012] Further, the input of the spatio-temporal noise filtering Siamese network model of the present invention is a pair of samples, and the output is the similarity between the samples; when the two samples belong to the same category, the similarity tends to 1; when the two samples belong to different categories, the similarity is close to 0.
[0013] Further, the spatio-temporal noise filtering Siamese network model of the present invention uses a contrastive loss function to optimize the training objective of the model, described as: (13) In the formula, Y is the similarity label of the sample, Y = 1 indicates that the two samples are similar; Y = 0 indicates that the two samples are not similar; N is the number of samples; m is the distance constraint threshold; According toX 1 and X Are 1 and X 1 and X 2 similar to the simplified formula (13)? When (14) When X 1 and X 2 are not similar, the loss function is simplified to: (15).
[0014] Furthermore, the method for training the spatio-temporal noise filtering siamese network model in this invention specifically includes: Data collection: Collect fault data of various states of power transformers; Data preprocessing: Normalize the collected fault data according to the input requirements of the spatio-temporal noise filtering siamese network model, and divide it into a training set and a test set; Model training: Initialize the spatio-temporal noise filtering siamese network model with random parameters, randomly select two samples from the training set to form a sample pair and input it into the model, and use the contrastive loss function to optimize the training objective of the model; Model testing: Use the test set data to evaluate the trained spatio-temporal noise filtering siamese network model, calculate its accuracy rate, and evaluate the model performance; Result analysis: Conduct visual analysis on the test results to evaluate the diagnostic ability of the spatio-temporal noise filtering siamese network model.
[0015] This invention provides a small-sample transformer fault diagnosis system based on a spatio-temporal noise filtering siamese network, including: A memory for storing executable computer programs; A processor, when executing the executable computer programs stored in the memory, implements the above-mentioned small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering siamese network.
[0016] The beneficial effects produced by this invention are: 1. This invention proposes an adaptive residual noise filtering unit (ARNFU), which can suppress the noise in the vibration signals of power transformers. This unit solves the problem of identity bias and eliminates the signal distortion caused by the soft threshold function, and is used to adaptively set the optimal threshold and further correct the output. The semi-soft threshold function included in ARNFU can retain the continuity of the soft threshold function, while eliminating the constant bias problem existing in the soft threshold function, and maximally retains the effective features.
[0017] 2. The present invention uses a bidirectional gated recurrent unit (BiGRU) to mine the temporal feature information of the vibration signals of power transformers, and captures the long-term dependence information of the past and future of the vibration signals by introducing a bidirectional structure.
[0018] 3. The present invention constructs a siamese network with ARNFU and a bidirectional gated recurrent unit (BiGRU) as feature extractors for fault diagnosis of power transformers under limited data conditions. Specifically, ARNFU extracts the spatial features of the vibration signals while eliminating noise, and BIGRU is used to extract the temporal features of the vibration signals. Finally, the similarity is measured by calculating the Euclidean distance between fault sample pairs, and then the small-sample fault diagnosis of power transformers is completed.
[0019] 4. The method of the present invention fully considers the problem that the number of certain fault samples of power transformers may be zero in actual engineering. The proposed method can excellently complete the task of fault diagnosis of power transformers under small-sample conditions, and has broad academic and engineering application prospects in other small-sample fault diagnosis fields. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] The present invention will be further described below in conjunction with the drawings and embodiments. In the drawings: Figure 1 is the flowchart of the small-sample transformer fault diagnosis method based on the spatio-temporal noise filtering siamese network proposed by the present invention; Figure 2 is the framework diagram of ARNFU designed by the present invention; Figure 3 is the BiGRU framework diagram; Figure 4 is the fault diagnosis accuracy rate of various methods in the test embodiment of the present invention; Figure 5 is the relationship between the diagnostic accuracy decline rate and the number of training samples in the test embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0021] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below in conjunction with the drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.
[0022] Embodiment 1 The small-sample transformer fault diagnosis method based on the spatio-temporal noise filtering siamese network according to the embodiment of the present invention includes: Construct an Adaptive Residual Noise Filtering Unit, denoted as ARNFU, where the noise signal is reduced by a semi-soft threshold function; construct a Bidirectional Gated Recurrent Unit, denoted as BiGRU, which is used to mine the sequential feature information of the vibration signal of the power transformer, and capture the long-term dependence information of the vibration signal by introducing a bidirectional structure; construct a spatio-temporal noise filtering Siamese network model with ARNFU and BiGRU as feature extractors for power transformer fault diagnosis under the condition of limited data. Collect the fault data of various states of the power transformer, preprocess it, divide it into a training set and a test set, use the spatio-temporal noise filtering Siamese network model constructed with random parameter initialization, randomly select two samples from the training set to form a sample pair and input it into the model for training, and use the test set to evaluate the training results to obtain a trained spatio-temporal noise filtering Siamese network model; perform transformer fault diagnosis through the obtained model.
[0023] In a preferred embodiment of the present invention, the method includes five steps: data collection, data preprocessing, model training, model testing, and result analysis, which are specifically as follows: a) Data acquisition: Collect the fault data of various states of the power transformer.
[0024] b) Data preprocessing: Normalize the collected data according to the input requirements of the model, and divide it into a training set and a test set.
[0025] c) Model training: Initialize the proposed model structure with random parameters, randomly select two samples from the training set to form a sample pair and input it into the model, and use a contrastive loss function to optimize the training objective of the model.
[0026] d) Model testing: Use the test set data to evaluate the trained model, calculate its accuracy rate, and evaluate the model performance.
[0027] e) Result analysis: Perform visual analysis on the test results to evaluate the diagnostic ability of the model.
[0028] Embodiment 2 Based on Embodiment 1, the embodiment of the present invention gives the specific structures of the Adaptive Residual Noise Filtering Unit (ARNFU) and the Bidirectional Gated Recurrent Unit (BiGRU), and gives the specific construction method of the Siamese network with ARNFU and BiGRU as feature extractors.
[0029] The small-sample transformer fault diagnosis method based on the spatio-temporal noise filtering Siamese network in the embodiment of the present invention includes the following steps: 1) Construct ARNFU. ARNFU is a special structure of the residual neural network, which reduces noise signals through a semi-soft threshold function. In ARNFU, the input signal is first trained in two convolutional layers. Then, the output of the second convolutional layer passes through Branch 1 and Branch 2 simultaneously.
[0030] In Branch 1, after the input vector is processed by the absolute value, it passes through the Avg_pool layer, the Batch Normalization (BN) layer, and the Fully Connected (FC) layer in sequence to obtain the scaling parameter. Subsequently, the scaling parameter is multiplied by the average value of Avg_pool to obtain the threshold.
[0031] In Branch 2, the input vector passes through the GAP layer and two FC layers in sequence to obtain the adaptive correction factor. Finally, the output of the semi-soft threshold function is further corrected using the adaptive correction factor, enabling the entire network to retain more effective signals.
[0032] 2) The semi-soft threshold function included in ARNFU can retain the continuity of the soft threshold function, while eliminating the constant bias problem existing in the soft threshold function, and retaining the effective features to the greatest extent. The semi-soft threshold function can be expressed as: (1) where x is the input; y is the output; τ is the threshold.
[0033] 3) Further, the convolutional layer is used for feature extraction. The convolutional operation can be expressed as: (2) where and are the input and output of the l th layer respectively; is the convolutional kernel from the l -1)th layer, the i th neuron to the l th layer, the i th neuron, where j is the number of convolutional kernels; is the number of neurons in the l -1)th layer; is the bias of the l th layer, the i th neuron; 1Dconv is the convolutional operation.
[0034] 4) Further, the BN layer normalizes the data into normally distributed data, simplifies the difficulty of subsequent data processing, speeds up the network's calculation time, and prevents the model from overfitting during training. The calculation process of the BN layer can be expressed as: (3) (4) (5) (6) In the formula and are respectively the i th data value and the normalized value in the dataset; and are respectively the average value and the standard deviation of the dataset; is an infinitesimal constant; and are respectively the scaling factor and the offset factor of the BN layer operation; is the output of the i th data passing through the BN layer.
[0035] 5) Further, the ReLU activation function is adopted to alleviate the problem of gradient disappearance easily occurring in functions such as Sigmoid. The ReLU activation function can be expressed as: (7) 6) Further, the sigmoid function scales the scaling parameter to the range of (0, 1), which can be described as: (8) In the formula is the scaling parameter; c is the output of the FC layer in branch 1.
[0036] 7) Further, multiply the scaling parameter by the average value of x to obtain the threshold, which can be described as: (9) In the formula is the threshold, x is the input signal.
[0037] 8) Further, after the threshold calculation, an adaptive correction factor is used to further correct the output of the semi-soft threshold function, and the output can be expressed as: (10) In the formula is the adaptive correction factor. From the above process, it can be seen that the proposed ARNFU can remove the features related to noise and retain the effective vibration signals of power transformers.
[0038] 9) Further, a bidirectional gated recurrent unit (BiGRU) is used to mine the temporal feature information of the vibration signals of power transformers, and the past and future long-term dependence information of the vibration signals is captured by introducing a bidirectional structure.
[0039] 10) Further, a siamese network with ARNFU and BiGRU as feature extractors is constructed. The siamese network uses two sub-networks with shared weights to receive two input samples simultaneously, and the output result is the similarity of the two samples. The network first maps the two samples ( X 1, X 2) to a low-dimensional feature space, and then calculates the Euclidean distance between the two feature vectors d ( X 1, X 2). The similarity degree between samples is measured by the distance, which can be expressed as: (11) In the formula f ( X 1) and f ( X 2) are the low-dimensional feature vectors obtained after the samples X 1 and X 2 pass through the feature extractor respectively; is the similarity distance; represents the two-norm operation.
[0040] 11) Further, the calculated distance is input into the FC layer using the sigmoid activation function, and the probability value is obtained to represent the similarity between the input sample pairs. The larger the probability value, the more similar the two samples are, and the greater the possibility of belonging to the same category. This process can be described as: (12) In the formula, sigmoid (·) and FC (·) represent the sigmoid function and the fully connected layer function respectively; represents X 1 and X 2 whether they belong to the same category probability.
[0041] 12) Further, the input of the Siamese network is a pair of samples, and the output is the similarity between the samples. When the two samples belong to the same category, the similarity tends to 1; when the two samples belong to different categories, the similarity is close to 0. The Siamese network uses a contrastive loss function to optimize the training objective of the model, which can be described as: (13) where Y is the similarity label of the samples, Y = 1 indicates that the two samples are similar; Y = 0 indicates that the two samples are not similar; N is the number of samples; m is the distance constraint threshold.
[0042] 13) Further, the formula (13) can be simplified according to whether X 1 and X 2 are similar. When X 1 and X 2 are similar, the loss function can be simplified to: (14) 14) Further, when X 1 and X 2 are not similar, the loss function can be simplified to: (15) Example 3 To comprehensively demonstrate the fault diagnosis performance of the proposed method under different numbers of training samples, the present invention randomly selects 30, 60, 120, 240, 480, 960, and 1280 training samples from the training set containing various mechanical faults of power transformers for training. To avoid random errors, the selection process is repeated 5 times to obtain 5 training sets. The present invention uses the average accuracy of the 5 training sets as the final result of the fault diagnosis, and uses the best model in each training set for one-time and five-time tests. The compared methods include SVM, MS-1DCNN, ResNet+SE, and ARNFU+BiGRU, and the comparison results are as Figure 4 shown.
[0043] From Figure 4It can be seen that in the case of different numbers of training samples, the fault diagnosis accuracy rate of SNFSN is higher than that of the other four methods. In the case of a large number of training samples, the accuracy rate of SNFSN in the first and fifth tests is significantly higher than that of SVM, and only slightly higher than that of MS-1DCNN, ResNet+SE, and ARNFU+BiGRU. As the number of training samples decreases, the accuracy rates of the five methods gradually decrease. However, compared with SVM, MS-1DCNN, ResNet+SE, and ARNFU+BiGRU, the decline rate of SNFSN is slower, as Figure 5 shown. When there are only 30 training samples, the diagnostic accuracy rate of SNFSN can reach 78%, which is 56.7%, 36.8%, 24.7%, and 25.4% higher than that of SVM, MS-1DCNN, ResNet+SE, and ARNFU+BiGRU, respectively. The above results indicate that the proposed SNFSN significantly reduces the dependence on large-scale training data and can still achieve good diagnostic results even in the case of only dozens of fault samples. At the same time, it can be observed that the accuracy of SNFSN in the fifth test is higher than that in the first test. Thus, it can be seen that increasing the number of query sets can enhance the diagnostic ability of the model to a certain extent.
[0044] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined by its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present application.
[0045] It should be understood that those of ordinary skill in the art can make improvements or transformations according to the above description, and all such improvements and transformations should fall within the protection scope of the appended claims of the present invention.
Claims
1. A small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network, characterized in that, The method includes: Construct an adaptive residual noise filtering unit, denoted as ARNFU, where the noise signal is reduced by a semi-soft threshold function; construct a bidirectional gated recurrent unit, denoted as BiGRU, for mining the temporal feature information of the vibration signal of the power transformer, and capture the long-term dependence information of the vibration signal by introducing a bidirectional structure; construct a spatio-temporal noise filtering siamese network model with ARNFU and BiGRU as feature extractors for fault diagnosis of power transformers under limited data conditions; Collect fault data of various states of the power transformer, preprocess it, divide it into a training set and a test set, use the spatio-temporal noise filtering siamese network model constructed with random parameter initialization, randomly select two samples from the training set to form a sample pair and input it into the model for training, use the test set to evaluate the training results, and obtain a trained spatio-temporal noise filtering siamese network model; perform transformer fault diagnosis through the obtained model.
2. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 1, characterized in that The method for constructing the bidirectional gated recurrent unit specifically includes: Reduce the noise signal by a semi-soft threshold function in ARNFU, and the input signal is first trained in two convolutional layers; the output of the second convolutional layer passes through branch 1 and branch 2 simultaneously; in branch 1, after the input vector is processed by the absolute value, it passes through the Avg_pool layer, BN layer, and FC layer in sequence to obtain a scaling parameter; subsequently, multiply the scaling parameter by the average value of Avg_pool to obtain a threshold; in branch 2, the input vector passes through the GAP layer and two FC layers in sequence to obtain an adaptive correction factor; finally, use the adaptive correction factor to further correct the output of the semi-soft threshold function so that the entire network can retain more effective signals.
3. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 2, wherein The semi-soft threshold function is expressed as: (1) In the formula, x is the input; y is the output; τ is the threshold value; The operation of the convolutional layer is expressed as: (2) In the formula, and are the input and output of the l th layer respectively; is the convolution kernel from the l -1th layer, the i th neuron to the l th layer, the i th neuron, where j is the number of convolution kernels; is the number of neurons in the l -1th layer; is the bias of the l th layer, the i th neuron; 1Dconv is the convolution operation; The calculation process of the BN layer is expressed as: (3) (4) (5) (6) Wherein, and are the i -th data value and the normalized value in the dataset respectively; and are the average value and the standard deviation of the dataset respectively; is an infinitesimal constant; and are the scaling factor and the offset factor of the BN layer operation respectively; is the output of the i -th data passing through the BN layer.
4. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 3, wherein This method also includes: Scale the scaling parameter to the range of (0, 1) through the sigmoid function, and use the ReLU activation function to alleviate the problem that the sigmoid function is prone to gradient disappearance; The ReLU activation function is expressed as: (7) The sigmoid function is described as: (8) In the formula, is the scaling parameter; c is the output of the FC layer in Branch 1; Multiply the scaling parameter by x the average value of to obtain a threshold, described as: (9) Wherein, is the threshold value, x is the input signal.
5. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 4, wherein Use the adaptive correction factor to further correct the output of the semi-soft threshold function, and the output is expressed as: (10) In the formula, is the adaptive correction factor.
6. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 5, wherein The method for constructing the spatio-temporal noise filtering siamese network model with ARNFU and BiGRU as feature extractors specifically includes: The spatio-temporal noise filtering Siamese network model maps two samples ( X 1, X 2) to a low-dimensional feature space and calculates the Euclidean distance between the two feature vectors d ( X 1, X 2). The similarity degree between samples is measured by the distance and is expressed as: (11) where f ( X 1)and f ( X 2)are the low-dimensional feature vectors obtained after the samples X 1and X 2pass through the feature extractor respectively; is the similarity distance; denotes the two-norm operation; The calculated distance is input into the FC layer using the sigmoid activation function, and a probability value is obtained to represent the similarity between input sample pairs; the larger the probability value, the more similar the two samples are and the greater the likelihood of belonging to the same category; described as: (12) Wherein, sigmoid(·) and FC(·) respectively represent the sigmoid function and the fully connected layer function; represents X the probability that X 1 and 2 belong to the same class.
7. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 1, wherein The input of the spatio-temporal noise filtering siamese network model is a pair of samples, and the output is the similarity between the samples; when the two samples belong to the same category, the similarity tends to 1; when the two samples belong to different categories, the similarity is close to 0.
8. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 6, wherein, The spatio-temporal noise filtering siamese network model uses a contrastive loss function to optimize the training objective of the model, which is described as: (13) In the formula, Y is the similarity label of the sample, Y = 1 indicates that the two samples are similar; Y = 0 indicates that the two samples are not similar; N is the number of samples; m is the distance constraint threshold; According to X 1 and X 2 are similar to the simplified formula (13). When X 1 and X 2 are similar, the loss function is simplified to: (14) When X 1 and X 2 are not similar, the loss function is simplified to: (15)。 9. The small-sample transformer fault diagnosis method based on a spatio-temporal noise filtering Siamese network according to claim 1, characterized in that The method for training the spatio-temporal noise filtering siamese network model in this method specifically includes: Data collection: Collect fault data of various states of the power transformer; Data preprocessing: Normalize the collected fault data according to the input requirements of the spatio-temporal noise filtering siamese network model, and divide it into a training set and a test set; Model training: Initialize the spatio-temporal noise filtering Siamese network model with random parameters, randomly select two samples from the training set to form a sample pair and input it into the model, and use the contrastive loss function to optimize the training objective of the model; Model testing: Use the test set data to evaluate the trained spatio-temporal noise filtering Siamese network model, calculate its accuracy rate, and evaluate the model performance; Result analysis: Conduct visual analysis on the test results to evaluate the diagnostic ability of the spatio-temporal noise filtering Siamese network model.
10. A small-sample transformer fault diagnosis system based on a spatio-temporal noise filtering Siamese network, characterized in that, Including: A memory for storing executable computer programs; A processor, when executing the executable computer programs stored in the memory, implements the small-sample transformer fault diagnosis method based on the spatio-temporal noise filtering Siamese network according to any one of claims 1 to 9.