Transformer fault diagnosis method and device based on self-attention heterogeneous network
By combining the self-attention LSTM network and the self-attention residual network, a complete transformer fault feature vector was constructed, which solved the problem of inaccurate diagnosis results caused by incomplete feature extraction in the existing technology and achieved higher fault diagnosis accuracy.
Patent Information
- Application Number
- CN202310857743.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-12
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2043-07-12
AI Technical Summary
The existing transformer fault diagnosis methods have incomplete feature extraction, resulting in inaccurate diagnosis results.
A self-attention-based heterogeneous network is adopted to combine thermal infrared images and dissolved gas content in oil. The self-attention LSTM network focuses on the global time scale information, the self-attention residual network focuses on the local spatial information, and a complete fault feature vector is constructed through the feature fusion layer for fault diagnosis.
The accuracy of transformer fault diagnosis is improved, redundant information interference is reduced, and the accuracy of fault diagnosis results is improved.
Smart Images

Figure CN116955951B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of transformer fault diagnosis, and more specifically to a transformer fault diagnosis method and device based on a self-attention heterogeneous network. Background Art
[0002] Transformers are the most expensive equipment in power systems. Their safe operation not only affects the production safety and economic benefits of power companies, but also has an immeasurable impact on society. Therefore, in-depth research on transformer fault diagnosis is crucial. Transformers are to power systems what the heart is to the human body; their proper operation determines the efficient transmission of energy. However, transformers typically operate under heavy loads. Environmental changes and external forces can cause a range of transformer failures, leading to widespread power outages and other problems with varying degrees of impact. Therefore, the safe operation of transformers has become a critical issue in my country's power industry.
[0003] Currently, the mainstream method for detecting transformer faults is dissolved gas analysis (DGA). However, detecting the nature of internal transformer faults (overheating or discharge) solely based on the dissolved gas composition in transformer oil has some shortcomings. For example, DGA analysis often detects transformer faults after the transformer has already exhibited significant abnormalities, often by which point the fault is already quite serious. If the dissolved gas content is not clearly abnormal, the insulation condition of the transformer cannot be accurately determined, potentially delaying repairs and leading to more serious failures. Infrared imaging technology can accurately and contactlessly capture the surface temperature distribution of operating power equipment from a distance, enabling detection of latent thermal defects in large power transformers. It has been widely used in power equipment fault detection. Therefore, combining infrared imaging with DGA is crucial for rapid and accurate transformer fault diagnosis in practical applications.
[0004] For example, Chinese patent publication number CN114462508A discloses a transformer health assessment method based on a multimodal neural network. This method collects thermal infrared images of the transformer and the dissolved gas content in the oil. Wavelet threshold denoising is used to process the collected multimodal information. A one-dimensional convolutional neural network is used to extract text features from the gas-in-oil data, and a deep residual neural network is used to extract image features from the infrared images. This method combines infrared imaging technology with dissolved gas analysis to perform transformer fault diagnosis. However, this method is unable to extract the spatiotemporal characteristics of the transformer data, resulting in incomplete extracted features and, consequently, inaccurate diagnostic results. Summary of the Invention
[0005] The technical problem to be solved by the present invention is that the transformer fault diagnosis method in the prior art has an inaccurate final diagnosis result due to incomplete extracted features.
[0006] The present invention solves the above technical problems by the following technical means: a transformer fault diagnosis method based on a self-attention heterogeneous network, the method comprising:
[0007] Step a: When the transformer is in a faulty state, obtain a thermal infrared image of the transformer and perform preprocessing to obtain a first training set T1;
[0008] Step b: When the transformer is in a faulty state, the dissolved gas content in the transformer oil is collected and preprocessed to obtain a second training set T2;
[0009] Step c: constructing a heterogeneous fusion network model based on self-attention, wherein the model includes a self-attention LSTM network focusing on global time scale information, a self-attention residual network focusing on local spatial information, a feature fusion layer, and a Softmax classifier, wherein the first training set T1 is input into the self-attention residual network, and the second training set T2 is input into the self-attention LSTM network. The output ends of the self-attention residual network and the self-attention LSTM network are both connected to the input end of the feature fusion layer, and the output end of the feature fusion layer is connected to the Softmax classifier;
[0010] Step d: Train the self-attention-based heterogeneous fusion network model and use the trained model for fault diagnosis.
[0011] Beneficial effects: The self-attention LSTM network of the present invention focuses on global time scale information, while the self-attention residual network focuses on local spatial information. The information of the two is fused through the feature fusion layer to construct a complete transformer fault feature vector, which fully explores the local spatial information and global time scale information. The complete spatiotemporal information helps to output more accurate fault diagnosis results, and the model is trained and used to perform fault diagnosis, further improving the accuracy of the fault diagnosis results.
[0012] Furthermore, the step a includes:
[0013] Obtain L time periods when the transformer is operating under h fault conditions; record any lth time period in the L time periods as T l , the lth time period T l Divide into N equally spaced moments; obtain the thermal infrared image of the oil-immersed transformer at any nth equally spaced moment and perform preprocessing and normalization to obtain a set of transformer fault image samples containing h×N×L pieces as the first training set T1.
[0014] Furthermore, the step b includes:
[0015] At any nth equally spaced moment, m dissolved gas contents in the oil of the oil-immersed transformer are collected and used as fault feature variables; thereby, h transformer fault time series samples with N×L rows and m columns are formed and normalized to obtain the normalized transformer fault time series samples, where N>m; the sliding window size is set to w×m and the step size is 1, and the transformer fault time series samples after the normalization operation are longitudinally slid and valued to obtain h×N×L w×m transformer fault matrices as the second training set T2.
[0016] Furthermore, the self-attention residual network includes three identical self-attention residual modules, and the three identical self-attention residual modules are connected in series.
[0017] Furthermore, the processing process of the self-attention residual module is:
[0018] 1) The transformer fault image Input1 is sequentially input into two 2D convolutions for residual mapping to obtain a feature map f with a dimension of C×H×W; the transformer fault image Input1 is the data in the first training set T1;
[0019] 2) After the feature map f is activated by 1×1 convolution and Softmax function, an intermediate feature vector with dimension HW×1×1 is obtained. The dot product is performed with the feature map f to obtain the channel attention feature vector f1 with dimension C×1×1.
[0020] 3) After the global feature vector f1 is input into the two fully connected layers, it is activated by the Sigmoid function, dot-producted with the feature map f, and then short-circuited with the transformer fault image Input1 to obtain the output Y.
[0021] Furthermore, the self-attention LSTM network includes three identical self-attention LSTM modules, and the three identical self-attention LSTM modules are connected in series.
[0022] Furthermore, the processing process of the self-attention LSTM module is as follows:
[0023] 1) Input the transformer fault matrix Input2 to each LSTM unit respectively, and obtain the hidden layer output h={h1,h2,...,h t}, where h t The dimension of is hs×1, and the dimension of h is hs×t; the transformer fault matrix Input2 is the data in the second training set T2;
[0024] 2) After activating the hidden layer output h at each moment with 1D convolution and softmax function, an intermediate feature vector with a dimension of 1×hs can be obtained, and a dot product with h is performed to obtain a time-scale attention feature vector with a dimension of 1×t
[0025] 3) The time-scale attention feature vector After being sequentially input into two fully connected layers, the Sigmoid function is used for activation and the dot product is performed with h to obtain the output H.
[0026] Furthermore, the step d comprises:
[0027] Step d1: define the current number of iterations of the neural network as μ and initialize μ=1; the maximum number of iterations is μ max ; Perform μ-th random initialization on the parameters of each layer in the network;
[0028] Step d2, initializing i=1;
[0029] Step d3: Select the i-th transformer fault image x from the first training set T1 i , input the self-attention residual network to obtain the feature vector space feature vector F i,μ 1 , with a dimension of p×1; select the i-th transformer fault matrix y from the second training set T2 i , input the self-attention LSTM network to obtain the time series feature vector F i,μ 2 , the dimension is q×1;
[0030] Step d4: convert the spatial feature vector F i,μ 1 With the time series feature vector F i,μ 2 Input feature fusion layer to perform head-to-tail splicing to obtain transformer fault state feature vector F i,μ =[F i,μ 1 ,F i,μ 2 ] T , the dimension is (p+q)×1; the transformer fault state feature vector F i,μ Input Softmax classifier to obtain the transformer fault image x of the current network input i and transformer fault matrix y i The fault classification result H i ;
[0031] Step d5: After assigning i+1 to i, determine whether i>h×N×L. If so, proceed to step d6; otherwise, return to step d3.
[0032] Step d6: Calculate the root mean square error e of the output of the heterogeneous fusion network based on self-attention in the μ-th iteration μ ; Judge whether e μ <error or μ> μ max holds. If it holds, save the current network model A μ ; Otherwise, assign μ + 1 to μ and then return to step d2; where error is a preset error value.
[0033] The present invention also provides a transformer fault diagnosis device based on a heterogeneous network with self-attention. The device includes:
[0034] A first training set acquisition module, configured to obtain the thermal infrared images of the transformer under the transformer fault state and perform preprocessing to obtain a first training set T1;
[0035] A second training set acquisition module, configured to collect the content of dissolved gases in the oil of the transformer under the transformer fault state and perform preprocessing to obtain a second training set T2;
[0036] A model construction module, configured to construct a heterogeneous fusion network model based on self-attention. The model includes a self-attention LSTM network that focuses on global time-scale information, a self-attention residual network that focuses on local spatial information, a feature fusion layer, and a Softmax classifier. The first training set T1 is input into the self-attention residual network, and the second training set T2 is input into the self-attention LSTM network. The output ends of the self-attention residual network and the self-attention LSTM network are both connected to the input end of the feature fusion layer, and the output end of the feature fusion layer is connected to the Softmax classifier;
[0037] A model training module, configured to train the heterogeneous fusion network model based on self-attention and perform fault diagnosis using the trained model.
[0038] Furthermore, the first training set acquisition module is further configured to:
[0039] Obtain L time periods during the operation of the transformer in h fault states; Denote any l-th time period in the L time periods as T l , and divide the l-th time period T l into N equally spaced moments; Obtain the thermal infrared images of the oil-immersed transformer at any n-th equally spaced moment and perform preprocessing and normalization to obtain a set of transformer fault image samples containing h × N × L as the first training set T1.
[0040] Even further, the second training set acquisition module is further configured to:
[0041] At any nth equally spaced moment, m dissolved gas contents in the oil of the oil-immersed transformer are collected and used as fault feature variables; thereby, h transformer fault time series samples with N×L rows and m columns are formed and normalized to obtain the normalized transformer fault time series samples, where N>m; the sliding window size is set to w×m and the step size is 1, and the transformer fault time series samples after the normalization operation are longitudinally slid and valued to obtain h×N×L w×m transformer fault matrices as the second training set T2.
[0042] Furthermore, the self-attention residual network includes three identical self-attention residual modules, and the three identical self-attention residual modules are connected in series.
[0043] Furthermore, the processing process of the self-attention residual module is:
[0044] 1) The transformer fault image Input1 is sequentially input into two 2D convolutions for residual mapping to obtain a feature map f with a dimension of C×H×W; the transformer fault image Input1 is the data in the first training set T1;
[0045] 2) After the feature map f is activated by 1×1 convolution and Softmax function, an intermediate feature vector with dimension HW×1×1 is obtained. The dot product is performed with the feature map f to obtain the channel attention feature vector f1 with dimension C×1×1.
[0046] 3) After the global feature vector f1 is input into the two fully connected layers, it is activated by the Sigmoid function, dot-producted with the feature map f, and then short-circuited with the transformer fault image Input1 to obtain the output Y.
[0047] Furthermore, the self-attention LSTM network includes three identical self-attention LSTM modules, and the three identical self-attention LSTM modules are connected in series.
[0048] Furthermore, the processing process of the self-attention LSTM module is as follows:
[0049] 1) Input the transformer fault matrix Input2 to each LSTM unit respectively, and obtain the hidden layer output h={h1,h2,...,h t}, where h t The dimension of is hs×1, and the dimension of h is hs×t; the transformer fault matrix Input2 is the data in the second training set T2;
[0050] 2) After activating the hidden layer output h at each moment with 1D convolution and softmax function, an intermediate feature vector with a dimension of 1×hs can be obtained, and a dot product with h is performed to obtain a time-scale attention feature vector with a dimension of 1×t
[0051] 3) The time-scale attention feature vector After being sequentially input into two fully connected layers, the Sigmoid function is used for activation and the dot product is performed with h to obtain the output H.
[0052] Furthermore, the model training module is also used to:
[0053] Step d1: define the current number of iterations of the neural network as μ and initialize μ = 1; the maximum number of iterations is μ max ; Perform μ-th random initialization on the parameters of each layer in the network;
[0054] Step d2, initializing i=1;
[0055] Step d3: Select the i-th transformer fault image x from the first training set T1 i , input the self-attention residual network to obtain the feature vector space feature vector F i,μ 1 , with a dimension of p×1; select the i-th transformer fault matrix y from the second training set T2 i , input the self-attention LSTM network to obtain the time series feature vector F i,μ 2 , the dimension is q×1;
[0056] Step d4: convert the spatial feature vector F i,μ 1 With the time series feature vector F i,μ 2 Input feature fusion layer to perform head-to-tail splicing to obtain transformer fault state feature vector F i,μ =[F i,μ 1 ,F i,μ 2 ] T , the dimension is (p+q)×1; the transformer fault state feature vector F i,μ Input Softmax classifier to obtain the transformer fault image x of the current network input i and transformer fault matrix y i The fault classification result H i ;
[0057] Step d5: After assigning i+1 to i, determine whether i>h×N×L. If so, proceed to step d6; otherwise, return to step d3.
[0058] Step d6: Calculate the root mean square error e of the output of the heterogeneous fusion network based on self-attention at the μ-th iteration μ ; Judge whether e μ <error or μ> μ max holds. If it holds, save the current network model A μ ; Otherwise, assign μ + 1 to μ and return to step d2; where error is a preset error value.
[0059] The advantages of the present invention are as follows:
[0060] (1) The self-attention LSTM network of the present invention focuses on global time-scale information, while the self-attention residual network focuses on local spatial information. The information of the two is fused through the feature fusion layer to construct a complete transformer fault feature vector, fully exploring local spatial information and global time-scale information. The spatio-temporal information is complete, which helps to output a more accurate fault diagnosis result. And the model is trained, and the trained model is used for fault diagnosis to further improve the accuracy of the fault diagnosis result.
[0061] (2) The present invention constructs a self-attention LSTM network to learn its own information and applies different attention sizes to different time-scale information; by constructing a self-attention residual network to process feature maps of different channels, highlighting transformer fault feature information and reducing the interference of redundant information, thereby improving the accuracy of fault diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0062] Figure 1 is a flowchart of the transformer fault diagnosis method based on a heterogeneous network with self-attention disclosed in Embodiment 1 of the present invention;
[0063] Figure 2 is a schematic structural diagram of the self-attention residual module in the transformer fault diagnosis method based on a heterogeneous network with self-attention disclosed in Embodiment 1 of the present invention;
[0064] Figure 3 is a schematic structural diagram of the self-attention LSTM module in the transformer fault diagnosis method based on a heterogeneous network with self-attention disclosed in Embodiment 1 of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0065] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0066] Example 1
[0067] like Figure 1 As shown, the present invention provides a transformer fault diagnosis method based on a self-attention heterogeneous network, the method comprising the following steps:
[0068] Step a: Obtain L time periods when the transformer is operating under h fault conditions; any lth time period in the L time periods is recorded as T l , the lth time period T l Divide into N equally spaced moments; obtain a thermal infrared image of the oil-immersed transformer at any nth equally spaced moment and perform preprocessing and normalization to obtain a set of transformer fault image samples containing h×N×L pieces as the first training set T1;
[0069] Step b, collecting m dissolved gas contents in oil of the oil-immersed transformer at any nth equally spaced time and using them as fault characteristic variables; thereby forming h transformer fault time series samples with N×L rows and m columns and performing a normalization operation to obtain normalized transformer fault time series samples, where N>m;
[0070] Set the sliding window size to w×m and the step size to 1, perform longitudinal sliding value sampling on the transformer fault time series samples after the normalization operation, and obtain h×N×L w×m transformer fault matrices as the second training set T2;
[0071] Step c: Construct a heterogeneous fusion network model based on self-attention;
[0072] The self-attention heterogeneous fusion network model includes a self-attention LSTM network, a self-attention residual network, a feature fusion layer, and a Softmax classifier; wherein the output dimension of the self-attention LSTM network is p×1, and the output dimension of the self-attention residual network is q×1;
[0073] Step c1: The self-attention residual network includes three identical self-attention residual modules connected in series, wherein the self-attention residual module includes:
[0074] 1) The transformer fault image Input1 is sequentially input into two 2D convolutions for residual mapping to obtain a feature map f with a dimension of C×H×W;
[0075] 2) After the feature map f is activated by 1×1 convolution and Softmax function, an intermediate feature vector with dimension HW×1×1 is obtained. The dot product is performed with the feature map f to obtain the channel attention feature vector f1 with dimension C×1×1.
[0076] 3) After the global feature vector f1 is input into two fully connected layers, it is activated by the Sigmoid function, and dot product is performed with the feature map f, and then short-circuited with the transformer fault image Input1 to obtain the output Y. The structure of the self-attention residual module is as follows Figure 2 shown.
[0077] Step c2: The self-attention LSTM network includes three identical self-attention LSTM modules connected in series, wherein the self-attention LSTM module content includes:
[0078] 1) Input the transformer fault matrix Input2 to each LSTM unit respectively, and obtain the hidden layer output h={h1,h2,...,h t}, where h t The dimension of is hs×1, and the dimension of h is hs×t;
[0079] 2) After applying 1D convolution and softmax function activation to h, we can obtain the intermediate feature vector with dimension 1×hs, and perform dot product with h to obtain the time-scale attention feature vector with dimension 1×t
[0080] 3) After being sequentially input into two fully connected layers, the Sigmoid function is used for activation and the dot product is performed with h to obtain the output H. The structure of the self-attention LSTM module is as follows Figure 3 shown.
[0081] Step d: Train the self-attention-based heterogeneous fusion network model and use the trained model for fault diagnosis. The specific process is as follows:
[0082] Step d1: define the current number of iterations of the neural network as μ and initialize μ = 1; the maximum number of iterations is μ max ; Performing the μth random initialization on the parameters of each layer in the perturbation neural network, thereby obtaining the μth iteration of the self-attention-based heterogeneous fusion network;
[0083] Step d2, initializing i=1;
[0084] Step d3: Select the i-th transformer fault image x from the first training set T1 i , input the self-attention residual network to obtain the feature vector space feature vector F i,μ 1 , with a dimension of p×1; select the i-th transformer fault matrix y from the second training set T2 i , input the self-attention LSTM network to obtain the time series feature vector F i,μ 2 , the dimension is q×1;
[0085] Step d4: Concatenate the spatial feature vector F i,μ 1 and the temporal feature vector F i,μ 2 in the input feature fusion layer in the head - tail order to obtain the transformer fault state feature vector F i,μ = [F i,μ 1 , F i,μ 2 , with the dimension of (p + q)×1; Input the transformer fault state feature vector F T into the Softmax classifier to obtain the fault classification result H i,μ of the current network input transformer fault image x i and the transformer fault matrix y i . i .
[0086] Step d5: After assigning i + 1 to i, determine whether i>h×N×L holds; if it holds, continue to execute Step d6, otherwise return to Step d3;
[0087] Step d6: Calculate the root - mean - square error e μ of the output of the μ - th iteration of the heterogeneous fusion network based on self - attention; Determine whether e μ [[ID=�6]]<error or μ>μ max holds. If it holds, save the current network model A μ ; otherwise, assign μ + 1 to μ and return to Step d2; where error is a preset error value.
[0088] Through the above technical solutions, a transformer intelligent fault diagnosis method based on a heterogeneous fusion network with self - attention is proposed, aiming to quickly and accurately diagnose transformer faults and meet the safety requirements of power equipment. By constructing a self - attention LSTM network to learn its own information and applying different magnitudes of attention to information at different time scales; by constructing a self - attention residual network to process feature maps of different channels, the transformer fault feature information is highlighted, reducing the interference of redundant information and improving the accuracy of fault diagnosis. The dual - branch fusion network is used to obtain complete transformer fault features. The self - attention LSTM network pays more attention to global time - scale information, while the self - attention residual network focuses on local spatial information. A complete transformer fault feature vector is constructed through the feature fusion layer, making the method of the present invention closer to the human cognitive method from local to global, fully exploring local spatial information and global time - scale information, and endowing the model with excellent generalization ability.
[0089] Example 2 [[ID=?]]
[0090] Based on Example 1, Example 2 of the present invention further provides a transformer fault diagnosis device based on a self-attention heterogeneous network, the device comprising:
[0091] A first training set acquisition module is used to acquire a thermal infrared image of the transformer in a fault state and perform preprocessing to obtain a first training set T1;
[0092] A second training set acquisition module is used to collect the dissolved gas content in the transformer oil and perform preprocessing to obtain a second training set T2 when the transformer is in a fault state;
[0093] A model construction module is used to construct a self-attention-based heterogeneous fusion network model, wherein the model includes a self-attention LSTM network that focuses on global time scale information, a self-attention residual network that focuses on local spatial information, a feature fusion layer, and a Softmax classifier. The first training set T1 is input into the self-attention residual network, and the second training set T2 is input into the self-attention LSTM network. The output ends of the self-attention residual network and the self-attention LSTM network are both connected to the input end of the feature fusion layer, and the output end of the feature fusion layer is connected to the Softmax classifier.
[0094] The model training module is used to train the self-attention-based heterogeneous fusion network model and use the trained model for fault diagnosis.
[0095] Specifically, the first training set acquisition module is further used to:
[0096] Obtain L time periods when the transformer is operating under h fault conditions; record any lth time period in the L time periods as T l , the lth time period T l Divide into N equally spaced moments; obtain the thermal infrared image of the oil-immersed transformer at any nth equally spaced moment and perform preprocessing and normalization to obtain a set of transformer fault image samples containing h×N×L pieces as the first training set T1.
[0097] More specifically, the second training set acquisition module is further used to:
[0098] At any nth equally spaced moment, m dissolved gas contents in the oil of the oil-immersed transformer are collected and used as fault feature variables; thereby, h transformer fault time series samples with N×L rows and m columns are formed and normalized to obtain the normalized transformer fault time series samples, where N>m; the sliding window size is set to w×m and the step size is 1, and the transformer fault time series samples after the normalization operation are longitudinally slid and valued to obtain h×N×L w×m transformer fault matrices as the second training set T2.
[0099] More specifically, the self-attention residual network includes three identical self-attention residual modules, and the three identical self-attention residual modules are connected in series.
[0100] More specifically, the processing process of the self-attention residual module is:
[0101] 1) The transformer fault image Input1 is sequentially input into two 2D convolutions for residual mapping to obtain a feature map f with a dimension of C×H×W; the transformer fault image Input1 is the data in the first training set T1;
[0102] 2) After the feature map f is activated by 1×1 convolution and Softmax function, an intermediate feature vector with dimension HW×1×1 is obtained. The dot product is performed with the feature map f to obtain the channel attention feature vector f1 with dimension C×1×1.
[0103] 3) After the global feature vector f1 is input into the two fully connected layers, it is activated by the Sigmoid function, dot-producted with the feature map f, and then short-circuited with the transformer fault image Input1 to obtain the output Y.
[0104] More specifically, the self-attention LSTM network includes three identical self-attention LSTM modules, and the three identical self-attention LSTM modules are connected in series.
[0105] More specifically, the processing process of the self-attention LSTM module is as follows:
[0106] 1) Input the transformer fault matrix Input2 to each LSTM unit respectively, and obtain the hidden layer output h={h1,h2,...,h t}, where h t The dimension of is hs×1, and the dimension of h is hs×t; the transformer fault matrix Input2 is the data in the second training set T2;
[0107] 2) After activating the hidden layer output h at each moment with 1D convolution and softmax function, an intermediate feature vector with a dimension of 1×hs can be obtained, and a dot product with h is performed to obtain a time-scale attention feature vector with a dimension of 1×t
[0108] 3) The time-scale attention feature vector After being sequentially input into two fully connected layers, the Sigmoid function is used for activation and the dot product is performed with h to obtain the output H.
[0109] More specifically, the model training module is further used to:
[0110] Step d1: Define the current iteration number of the neural network as μ, and initialize μ = 1; the maximum iteration number is μ max ; perform the μ-th random initialization on the parameters of each layer in the network;
[0111] Step d2: Initialize i = 1;
[0112] Step d3: Select the i-th transformer fault image x from the first training set T1 i , and input it into the self-attention residual network to obtain the spatial feature vector F of the feature vector space i,μ 1 , with the dimension of p×1; select the i-th transformer fault matrix y from the second training set T2 i , and input it into the self-attention LSTM network to obtain the temporal feature vector F i,μ 2 , with the dimension of q×1;
[0113] Step d4: Concatenate the spatial feature vector F i,μ 1 and the temporal feature vector F by sequential concatenation at the head and tail in the input feature fusion layer, to obtain the transformer fault status feature vector F i,μ 2 = [F i,μ i,μ 1 , F i,μ 2 T , with the dimension of (p + q)×1; input the transformer fault status feature vector F i,μ into the Softmax classifier to obtain the fault classification result H of the current network input transformer fault image x i and the transformer fault matrix y i i ; i ;
[0114] Step d5: After assigning i + 1 to i, determine whether i > h×N×L holds; if it holds, continue to execute Step d6, otherwise return to Step d3;
[0115] Step d6: Calculate the root mean square error e of the output of the μ-th iteration of the heterogeneous fusion network based on self-attention μ ; determine whether e μ <error or μ>μ max holds. If it holds, save the current network model A μ ; otherwise, assign μ + 1 to μ and return to Step d2; where error is a preset error value.
[0116] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit the same. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the various embodiments of the present invention.
Claims
1. A transformer fault diagnosis method based on a self-attention heterogeneous network, characterized in that: The method comprises: Step a: Under the transformer fault state, obtain the thermal infrared image of the transformer and preprocess it to obtain the first training set ; Step b: Under the transformer fault state, the dissolved gas content in the transformer oil is collected and preprocessed to obtain the second training set ; Step c: Construct a heterogeneous fusion network model based on self-attention, which includes a self-attention LSTM network focusing on global time scale information, a self-attention residual network focusing on local spatial information, a feature fusion layer and a Softmax classifier. The first training set Input into the self-attention residual network, the second training set The input is fed into the self-attention LSTM network. The output of the self-attention residual network and the output of the self-attention LSTM network are both connected to the input of the feature fusion layer. The output of the feature fusion layer is connected to the Softmax classifier. The self-attention residual network includes a self-attention residual module, and the processing process of the self-attention residual module is: 1) Transformer fault image Sequentially input two 2D convolutions for residual mapping, and obtain the dimension Feature map ; Transformer fault image The first training set Data in; 2) Feature map use After the convolution and Softmax functions are activated, the dimension can be obtained as The intermediate feature vector of Perform dot product and get the dimension as The channel attention feature vector ; 3) The global eigenvector After inputting into two fully connected layers, the Sigmoid function is used for activation and combined with the feature map Perform dot product and then compare it with the transformer fault image Short-circuit the connection to get the output ; The self-attention LSTM network includes a self-attention LSTM module. The processing process of the self-attention LSTM module is as follows: 1) Transformer fault matrix Input into each LSTM unit separately to obtain the hidden layer output at each moment ,in, The dimension is , The dimension is ; Transformer fault matrix For the second training set Data in; 2) Output of the hidden layer at each moment After using 1D convolution and softmax function activation, the dimension can be obtained as The intermediate eigenvector of Perform dot product and get the dimension as The time-scale attention feature vector ; 3) The time-scale attention feature vector After being sequentially input into two fully connected layers, the Sigmoid function is used for activation and Perform dot product to get output ; Step d: Train the self-attention-based heterogeneous fusion network model and use the trained model for fault diagnosis.
2. The transformer fault diagnosis method based on self-attention heterogeneous network according to claim 1 is characterized in that: The step a comprises: Get Transformer When running in a fault state time period; Any time period The time period is recorded as , will Time period Divided into equally spaced moments; at any The thermal infrared images of the oil-immersed transformer are obtained at equal intervals and preprocessed and normalized to obtain the image data. The transformer fault image sample set is used as the first training set .
3. The transformer fault diagnosis method based on self-attention heterogeneous network according to claim 2 is characterized in that: The step b comprises: In any The oil-immersed transformer is collected at equal intervals. The dissolved gas content in the oil is used as the fault characteristic variable; thus forming The number of rows is , the number of columns is The transformer fault time series samples are normalized to obtain the normalized transformer fault time series samples, where ; Set the sliding window size to , the step size is , perform longitudinal sliding value sampling on the transformer fault time series samples after the normalization operation, and obtain indivual The transformer fault matrix is used as the second training set .
4. The transformer fault diagnosis method based on self-attention heterogeneous network according to claim 3 is characterized in that: The self-attention residual network includes three identical self-attention residual modules, and the three identical self-attention residual modules are connected in series.
5. The transformer fault diagnosis method based on self-attention heterogeneous network according to claim 4 is characterized in that: The self-attention LSTM network includes three identical self-attention LSTM modules, and the three identical self-attention LSTM modules are connected in series in sequence.
6. The transformer fault diagnosis method based on self-attention heterogeneous network according to claim 1 is characterized in that: The step d comprises: Step d1, define the current number of iterations of the neural network as , and initialize ; The maximum number of iterations is ; The parameters of each layer in the network are Random initialization; Step d2: Initialization ; Step d3: From the first training set Select the Transformer fault image , input the self-attention residual network to obtain the feature vector space feature vector , the dimension is ; From the second training set Select the Transformer fault matrix , input the self-attention LSTM network to obtain the time series feature vector , the dimension is ; Step d4: convert the spatial feature vector With the time series feature vector Input feature fusion layer to perform head-to-tail splicing to obtain transformer fault state feature vector , the dimension is ; Transformer fault state feature vector Input Softmax classifier to obtain the transformer fault image of the current network input and transformer fault matrix Fault classification results ; Step d5: Assign to After that, judge Is it true? If so, continue to step d6, otherwise return to step d3; Step d6, calculate the self-attention based heterogeneous fusion network The root mean square error of the iterative output ;judge or Is it true? If so, save the current network model Otherwise, Assign to Then, return to step d2; wherein, is the preset error value.
7. A transformer fault diagnosis device based on a self-attention heterogeneous network, characterized in that: The device comprises: The first training set acquisition module is used to obtain the thermal infrared image of the transformer in the transformer fault state and pre-process it to obtain the first training set ; The second training set acquisition module is used to collect the dissolved gas content in the transformer oil and pre-process it to obtain the second training set when the transformer is in a fault state. ; A model building module is used to build a heterogeneous fusion network model based on self-attention, which includes a self-attention LSTM network focusing on global time scale information, a self-attention residual network focusing on local spatial information, a feature fusion layer and a Softmax classifier. Input into the self-attention residual network, the second training set The input is fed into the self-attention LSTM network. The output of the self-attention residual network and the output of the self-attention LSTM network are both connected to the input of the feature fusion layer. The output of the feature fusion layer is connected to the Softmax classifier. The self-attention residual network includes a self-attention residual module, and the processing process of the self-attention residual module is: 1) Transformer fault image Sequentially input two 2D convolutions for residual mapping, and obtain the dimension Feature map ; Transformer fault image The first training set Data in; 2) Feature map use After the convolution and Softmax functions are activated, the dimension can be obtained as The intermediate feature vector of Perform dot product and get the dimension as The channel attention feature vector ; 3) The global eigenvector After inputting into two fully connected layers, the Sigmoid function is used for activation and combined with the feature map Perform dot product and then compare it with the transformer fault image Short-circuit the connection to get the output ; The self-attention LSTM network includes a self-attention LSTM module. The processing process of the self-attention LSTM module is as follows: 1) Transformer fault matrix Input into each LSTM unit separately to obtain the hidden layer output at each moment ,in, The dimension is , The dimension is ; Transformer fault matrix For the second training set Data in 2) Output of the hidden layer at each moment After using 1D convolution and softmax function activation, the dimension can be obtained as The intermediate eigenvector of Perform dot product and get the dimension as The time-scale attention feature vector ; 3) The time-scale attention feature vector After being sequentially input into two fully connected layers, the Sigmoid function is used for activation and Perform dot product to get output ; The model training module is used to train the self-attention-based heterogeneous fusion network model and use the trained model for fault diagnosis.
8. The transformer fault diagnosis device based on self-attention heterogeneous network according to claim 7 is characterized in that: The first training set acquisition module is further configured to: Get Transformer When running in a fault state time period; Any time period The time period is recorded as , will Time period Divided into equally spaced moments; at any The thermal infrared images of the oil-immersed transformer are obtained at equal intervals and preprocessed and normalized to obtain the image data. The transformer fault image sample set is used as the first training set .
Citation Information
Patent Citations
Power transformer health state assessment method based on multi-mode neural network
CN114462508A
Deep parallel fault diagnosis method and system for dissolved gas in transformer oil
CN111337768A
Power transformer fault diagnosis method and system of multi-modal information fusion network
CN116310551A