On-line transformer monitoring system and method based on deep learning and fusing sound signals and temperature data

By setting up microphones and infrared temperature cameras on both sides of the transformer, acquiring sound signals and temperature images, and using deep learning models for feature extraction and fusion, the problem of low accuracy in transformer fault diagnosis in the prior art is solved, and higher fault detection accuracy and robustness are achieved.

CN120125944APending Publication Date: 2025-06-10南京玄创智能装备有限公司
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510175698.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The prior art has a problem of low accuracy in transformer fault diagnosis, especially the diagnostic methods based on a single data source are difficult to fully capture the operating status of the transformer.

Method used

A transformer online monitoring system based on deep learning is adopted to combine sound signals and temperature data. By setting a microphone and infrared temperature camera on both sides of the transformer, it collects sound signals and temperature images in real time, and uses a deep learning model to extract and fuse these data to achieve fault monitoring.

Benefits of technology

Through multimodal data fusion, the operating status of the transformer can be captured more comprehensively, the accuracy and robustness of fault detection can be improved, and the time and cost of traditional manual tests and inspections can be reduced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120125944A_ABST
    Figure CN120125944A_ABST
Patent Text Reader

Abstract

The invention provides a transformer on-line monitoring system and method fusing sound signals and temperature data based on deep learning, and the system comprises two microphones and two infrared temperature cameras which are arranged at the two sides of a power transformer to be monitored, the microphones and the infrared temperature cameras are respectively connected with a controller, and the controller is connected with the sound signals and the temperature data. The controller analyzes the input sound signals and temperature images in real time through deep learning to complete online monitoring; the method adopts the system to monitor the power transformer to be monitored in real time. The feature fusion mechanism is used for obtaining fusion features and inputting the fusion features into the deep learning model, the accuracy of fault diagnosis is further improved, time and cost brought by manual inspection and testing can be effectively reduced, losses caused by equipment faults are reduced, stable operation of a power system is guaranteed through intelligent online monitoring, and the fault diagnosis accuracy is improved. And efficient and reliable technical support is provided for health management of the transformer.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a transformer monitoring system and method, in particular to an on-line monitoring system and method for transformers that integrates sound signals and temperature data based on deep learning. Background Art

[0002] The information provided in this part is only background information related to the present disclosure, and it is not necessarily prior art.

[0003] As a core device in the power system, the operating state of a transformer directly affects the stability and reliability of the entire power grid. Since a transformer is affected by various factors such as overload, aging, and environmental factors during long-term operation, internal mechanical failures or abnormalities may occur, such as loose transformer components, short-circuit impacts on transformers, and DC bias of transformers.

[0004] Early expert system designs and diagnostic methods for transformer fault diagnosis were mainly based on fault signal acquisition and feature extraction. That is, the feature analysis of electrical signals (such as voltage, current, power, etc.) during transformer faults, combined with signal processing algorithms and fault feature extraction algorithms, was used to achieve transformer fault diagnosis. However, this diagnostic method often requires shutdown detection or an external circuit. Shutdown detection causes economic losses, and the external circuit will also affect the normal operation of the transformer.

[0005] In recent years, some non-contact transformer fault diagnosis methods have been proposed, such as temperature-based diagnosis, sound-based diagnosis, etc. The sound signals of a transformer, as an important feature reflecting the operating state of the device, can provide effective early warning information in the early stage of a fault. These sound signals usually originate from internal mechanical components of the transformer, such as oil pumps, cooling fans, iron cores, etc., and their characteristic changes can reflect the abnormal state of the device. In addition, the temperature of the transformer is also closely related to the working state of the transformer. For example, local overheating may indicate problems with the insulation system of the transformer, too high oil temperature may mean there are electrical faults inside the transformer, or the cooling system may malfunction.

[0006] However, most of the emerging non-contact transformer fault diagnosis technologies only collect single data of the transformer, and there is a certain gap in the detection accuracy compared with traditional transformer fault diagnosis technologies.

[0007] It should be noted that the information disclosed in the above background art section is only used to enhance the understanding of the background of the present disclosure, and thus may include information that does not constitute prior art known to those of ordinary skill in the art. Summary of the Invention

[0008] Objective of the Invention: The technical problem to be solved by the present invention is to provide an on-line monitoring system and method for transformers that integrates sound signals and temperature data based on deep learning in view of the deficiencies of the prior art.

[0009] To solve the above technical problem, the present invention discloses an on-line monitoring system and method for transformers that integrates sound signals and temperature data based on deep learning. Among them, the system includes:

[0010] Two microphones and two infrared temperature cameras are arranged on both sides of the power transformer to be monitored. The microphones and infrared temperature cameras are respectively connected to the controller. The controller performs real-time analysis on the input sound signals and temperature images through deep learning to complete on-line monitoring.

[0011] Furthermore, the microphones and infrared temperature cameras are respectively installed on both sides of the power transformer to be monitored, collect the sound signals and temperature images emitted during the operation of the monitored power transformer, and transmit them to the controller in real time.

[0012] The present invention also proposes an on-line monitoring method for transformers that integrates sound signals and temperature data based on deep learning. The aforementioned system is used to perform real-time monitoring on the power transformer to be monitored, including the following steps:

[0013] Step 1: Collect data of the power transformer to be monitored in normal working conditions and abnormal working conditions to form a data set;

[0014] Step 2: Build a deep learning model for monitoring faults during the operation of the power transformer to be monitored;

[0015] Step 3: Use the data set formed in Step 1 to train the deep learning model built in Step 2;

[0016] Step 4: Use the trained deep learning model to perform real-time on-line monitoring during the operation of the power transformer to be monitored.

[0017] Furthermore, the collection of data of the power transformer to be monitored in normal working conditions and abnormal working conditions in Step 1 includes:

[0018] Step 1-1: Collect the sound signals and infrared temperature images synchronously recorded by the microphones and infrared temperature cameras arranged on both sides of the power transformer to be monitored in normal working conditions and abnormal working conditions;

[0019] Among them, the abnormal working conditions include: loose transformer components, short-circuit impact, DC bias, or overheating;

[0020] Step 1-2: Perform data enhancement on the sound signals and infrared temperature images respectively;

[0021] Steps 1-3, perform time alignment on the sound signal and the infrared temperature image.

[0022] Furthermore, the deep learning model in step 2 includes:

[0023] An input module, a sound signal feature extraction module, a temperature image feature extraction module, and a feature fusion monitoring module; among them,

[0024] The input module is used to receive the externally input sound signal and infrared temperature image;

[0025] The sound signal feature extraction module and the temperature image feature extraction module respectively extract the input sound signal features and infrared temperature image features;

[0026] The feature fusion monitoring module fuses the sound signal features and the infrared temperature image features, and performs fault monitoring to obtain the final result.

[0027] Furthermore, the temperature image feature extraction module includes:

[0028] Reset the image size of the infrared temperature image received by the input module, and adjust the infrared temperature image to a preset size;

[0029] Through the initial convolutional layer, use a 7×7 convolutional kernel to extract the preliminary features of the infrared temperature image;

[0030] Successively pass through 4 consecutive 3×3 residual modules, and their number of channels are 64, 128, 256, and 512 respectively;

[0031] Through the ordinary convolutional layer, convert the 512-dimensional feature vector into a 256-dimensional one;

[0032] Through global average pooling, reduce the spatial size.

[0033] Furthermore, set the temperature image feature extraction module to T, input T consecutive infrared temperature images into the same temperature image feature extraction module, and perform splicing to finally obtain the infrared temperature image features.

[0034] Furthermore, the sound signal feature extraction module obtains audio features through Mel-frequency cepstral coefficients to obtain the sound signal features.

[0035] Furthermore, the feature fusion monitoring module includes:

[0036] A feature fusion module and a feature monitoring module; among them,

[0037] The input infrared temperature image features and sound signal features are fused by a feature fusion module, which is expressed as follows:

[0038]

[0039] Among them, is the fused feature matrix, indicating that the fused feature matrix belongs to a 336-dimensional feature space and has T time steps, X image is the infrared temperature image feature, X audio is the sound signal feature;

[0040] The fused feature matrix X fusion is input into the feature monitoring module to obtain a detection result;

[0041] Among them, the feature monitoring module includes a time-delay neural network model based on an enhanced attention mechanism, and its specific structure is as follows:

[0042] The fused feature matrix X fusion is used as the input;

[0043] One-dimensional ordinary convolution, the rectified linear unit (ReLU) activation function, and batch normalization are used for processing in sequence;

[0044] Three residual blocks with a squeeze-and-excitation mechanism are used for processing. The results passing through 1, 2, and 3 of the residual blocks are respectively input into one-dimensional ordinary convolution and the rectified linear unit for processing;

[0045] A statistical pooling layer with attention and batch normalization are used for processing;

[0046] Finally, a fully connected layer and batch normalization are used. Mapping is performed through the fully connected layer, and the dimension is adjusted and then the cosine similarity is calculated with the candidate target categories. The detection result is determined according to the cosine similarity, that is, the detection result of the feature monitoring module is obtained;

[0047] The time-delay neural network model based on the enhanced attention mechanism is optimized using an angular margin cross-entropy loss function.

[0048] Furthermore, the training steps in step 4 include:

[0049] Step 4-1: Using the training set and validation set obtained in step 2, a K-fold cross-training deep learning model is adopted;

[0050] Step 4-2: Using the test set obtained in step 2, the trained model is tested;

[0051] Step 4-3: According to the results feedback in 4-2, fine-tune the size, stride, and T value of the convolutional layer of the residual block with the squeeze-and-excitation mechanism to optimize the model performance. By continuously fine-tuning the model, gradually obtain the best training results. The T value is 20. Repeat Step 4-1 and Step 4-2 again until the prediction accuracy of the deep convolutional neural network model is higher than the threshold, and the training is completed.

[0052] Beneficial effects:

[0053] 1. By fusing the features of sound signals and infrared temperature images, the present invention constructs a multi-modal deep learning model, which can comprehensively detect and accurately identify mechanical and electrical faults that may occur during the operation of transformers. Combining the multi-modal features of sound signals and temperature data, the model can more comprehensively capture the working state of transformers, thereby providing more accurate information for fault diagnosis.

[0054] 2. The present invention innovatively constructs a deep learning model based on ECAPA-TDNN and applies it to transformer fault detection, and realizes online fault monitoring through the voiceprint features of transformers in different working states. Considering the limitations of a single data source, the present invention further proposes a method of multi-modal data fusion, combining sound signals and infrared temperature image data. Through multi-modal feature fusion, the operating state of transformers can be more comprehensively captured, thereby improving the accuracy and robustness of fault detection.

[0055] 3. The present invention not only improves the accuracy of transformer fault detection, but also can effectively reduce the time and cost of traditional manual tests and inspections, and reduce the losses caused by equipment failures. Through intelligent fault detection, the system can ensure the stable operation of the power system, and at the same time provide a more efficient and reliable technical means for the health management of transformers, promoting the development of power industry technology. Description of the drawings

[0056] The following further specifically describes the present invention in conjunction with the drawings and specific embodiments, and the above and / or other advantages of the present invention will become clearer.

[0057] Figure 1 is the sound frequency spectrum diagram of the transformer collected by the microphone in the present invention.

[0058] Figure 2 is the structural schematic diagram of the deep learning model in the present invention.

[0059] Figure 3 is the schematic diagram of the convergence process during the training of the deep learning model in the present invention. Specific embodiments

[0060] In view of the deficiencies in the traditional fault diagnosis of transformers, the present invention proposes a method applying deep learning, which fuses sound signals and infrared temperature images to achieve non-contact online fault monitoring of transformers, helping to improve the accuracy of fault detection and reduce the operation cost of power equipment.

[0061] The technical solution of the present invention is as follows:

[0062] S1. Build an online transformer fault monitoring system;

[0063] S2. Construct a data set;

[0064] S3. Build a deep learning model. On the basis of introducing the deep convolutional neural network ResNet18 with residual connections (Reference: He K, Zhang X, Ren S, et al. Deep residual learning for image recognition [C]. Proceedings of the IEEE conference on computer vision and pattern recognition. 2016: 770 - 778.) and the enhanced convolutional attention time-domain neural network ECAPA-TDNN, the output of the feature layer of ResNet18 is fused with the audio features extracted by Mel-frequency cepstral coefficients (MFCC), and then the fused features are input into the subsequent neural network.

[0065] S4. Train the deep learning model;

[0066] S41. Use the training set and validation set obtained in S2 to train the deep learning model by K-fold cross-training;

[0067] S42. Use the test set obtained in S2 to test the trained model;

[0068] S43. According to the results feedback by the test set, fine-tune the model parameters to find the optimal model parameters

[0069] S5. Import the final deep learning model into the computer in the online transformer fault monitoring system to monitor the operation status of the transformer in real time and remind the technicians when a fault occurs.

[0070] The online transformer fault monitoring system in step S1 includes a power transformer, two microphones, two infrared temperature cameras, a computer and other connecting components. The microphones and infrared cameras are respectively arranged at the middle positions on both sides of the transformer, and the horizontal distance from the transformer is 2 - 3m.

[0071] In step S2, the proportions of the training set, validation set, and test set in all datasets are as follows: training set 70%, validation set 20%, and test set 10%.

[0072] The process of constructing the dataset in step S2 is as follows:

[0073] S21, Record the sound signal and infrared temperature image of the transformer under normal working conditions through a microphone and an infrared camera.

[0074] S22, Simulate the faults that often occur during the operation of the transformer, including loose transformer components, short - circuit impact of the transformer, DC bias of the transformer, and over - temperature of the transformer; record the sound signal and infrared temperature image of the transformer under fault conditions.

[0075] S23, Perform data augmentation by adding noise to the audio data and adjusting the brightness and contrast of the images to expand the dataset, thereby improving the accuracy and robustness of the deep learning model.

[0076] S24, Data alignment; the frame rate of the infrared camera is 50 frames per second, while the audio frames of the sound data usually have 50% overlap when input into the deep learning model. Therefore, the sound signal uses 40 - ms audio data as one frame to achieve synchronization with the image signal.

[0077] S25, Finally, divide the dataset into a training set, a validation set, and a test set.

[0078] As Figure 2 shown, the specific structure of the model in S3 is as follows:

[0079] S31, The model first receives T infrared temperature images and T audio signal frames captured from a microphone and an infrared camera. Among them, the image data is input into a residual neural network to extract T groups of image features, and the T audio frames extract features through Mel - Frequency Cepstral Coefficients.

[0080] Among them, the structure of the residual neural network for extracting image features is as follows:

[0081] S311, Image size reset. Here, the size of the infrared temperature image is adjusted to 224×224×3 to meet the input requirements of the convolutional neural network.

[0082] S312, Initial convolutional layer. Use a relatively large 7×7 convolutional kernel to extract the preliminary features of the image to capture the context information of a larger area

[0083] S313, Max - pooling layer. Reduce the size of the feature map and enhance the key parts in the feature map.

[0084] S314, Four consecutive 3×3 residual modules with the number of channels being 64, 128, 256, and 512 respectively. The introduction of residual connections can solve the problem of gradient disappearance in the training of deep neural networks. Among them, each residual block is usually composed of two 3x3 convolutional layers.

[0085] S315, Ordinary convolution and global average pooling. Ordinary convolution converts the 512-dimensional feature vector into a 256-dimensional one. Global average pooling compresses the spatial dimension of each feature map into a scalar, reducing the number of model parameters and preventing overfitting.

[0086] S32, The structure in S31 will be replicated T times to extract features from T consecutive infrared temperature images, thereby obtaining a 256×T feature vector.

[0087] S33, The part for processing audio signals first is audio signal feature extraction. This module fuses the 80×T feature vector of the audio signal with the infrared temperature image features obtained in the previous step to form a 336×T feature vector. The fusion formula is as follows:

[0088]

[0089] Among them, is the finally fused feature matrix.

[0090] S34, After feature fusion, it will be input into a model based on ECAPA-TDNN (Emphasized Channel Attention, Propagation and Aggregation Time Delay Neural Network, which can be translated as Emphasized Attention Mechanism Time Delay Neural Network, reference: Desplanques B, Thienpondt J, Demuynck K. Ecapa-tdnn: Emphasized channel attention, propagation and aggregation in tdnn based speaker verification[J]. arXiv preprint arXiv:2005.07143, 2020.). Its structure is as follows:

[0091] S341, One-dimensional ordinary convolution, Rectified Linear Unit (ReLU), and Batch Normalization (Reference: Batch Normalization, Ioffe S. Batch normalization: Accelerating deep network training by reducing internal covariate shift[J]. arXiv preprint arXiv:1502.03167, 2015.). This module performs a convolution operation on the input features to extract local features; the ReLU activation function introduces non-linearity to enhance the model's expressive power; Batch Normalization is used to accelerate training and stabilize the model.

[0092] S342, Three residual blocks with a squeeze-and-excitation mechanism. 3x3 convolution is used to extract features, and the residual connection is used to solve the vanishing gradient problem; the SE (Reference: Squeeze-and-Excitation, Hu J, Shen L, Sun G. Squeeze-and-excitation networks[C] / / Proceedings of the IEEE conference on computer vision and pattern recognition. 2018:7132-7141.) mechanism can adaptively adjust the channel weights, strengthen important features, and suppress unimportant features. Different strides d are used for feature extraction to further improve the model's expressive power and robustness.

[0093] Among them, the calculation of the SE mechanism includes a squeeze operation and an excitation operation. The squeeze operation compresses the spatial dimension of each channel into a scalar through global average pooling:

[0094]

[0095] where x h,w,c is the value of the input feature map at position (h, w) and channel C, and H and W are the height and width of the feature map respectively, and z c is the compressed information of this channel.

[0096] The excitation operation generates channel weights through a fully connected layer and then normalizes them using the Sigmoid function:

[0097] s c = σ(W2δ(W1z)) = σ(W2δ(W1z))

[0098] Among them, z is the compressed information obtained by the Squeeze operation, σ is the ReLU activation function, W1 and W2 are the weights of the fully connected layer, σ is the Sigmoid function, and s c is the excitation coefficient of the channel.

[0099] S343, one-dimensional ordinary convolution + rectified linear unit. Among them, the one-dimensional convolution performs a convolution operation on the input features to extract local features; the ReLU activation function is used to introduce non-linearity to enhance the model's expressive ability.

[0100] S344, statistical pooling layer with attention + batch normalization. Feature extraction is performed through statistical pooling, and the attention mechanism (reference: Vaswani A. Attention is all you need[J]. Advances in Neural Information Processing Systems, 2017.) is combined to highlight the features of important time steps. The implementation of the attention mechanism is as follows:

[0101]

[0102] Among them, f t is the attention score for time step t, and a t is the normalized attention weight.

[0103] S345, fully connected layer + batch normalization. The output of the pooling layer is mapped through the fully connected layer to adjust the dimension;

[0104] S346, angular margin cross-entropy loss function (AAM-softmax, Wang F, Cheng J, Liu W, et al. Additive margin softmax for face verification[J]. IEEE Signal Processing Letters, 2018, 25(7): 926-930.). AAM-softmax is an improved cross-entropy loss function that adds an angular margin to enhance the model's discrimination ability, especially in the case of similar classes, making the classification more accurate. Its calculation is as follows:

[0105]

[0106] Among them, θ y is the angle of the correct class, θ i is the angle of other classes, m is the angular margin, and C is the total number of classes.

[0107] In step S5, the trained deep learning model is deployed to the computer control module of the on-line transformer fault monitoring system. By receiving audio signals and image data from the microphone and infrared camera in real time, the computer will input these data into the deep learning model for real-time analysis. Once the system detects a possible fault in the transformer, it will notify the technical staff through an alarm, facilitating the timely handling of the fault and avoiding accidents.

[0108] Embodiment:

[0109] The system monitors the operating status of the transformer in real time by combining sound signals and infrared temperature images, and gives fault diagnosis and treatment suggestions in a timely manner when a fault occurs. This technology includes the following steps:

[0110] S1. Build an on-line transformer fault monitoring system. The system includes a power transformer, two microphones, two infrared temperature cameras and a computer. The microphones and infrared cameras are installed on both sides of the transformer respectively to ensure that the sound signals and temperature images can comprehensively collect the status of the transformer. The horizontal distance between the microphones and infrared cameras and the transformer is 2 to 3 meters to ensure the quality and range of the signals. In the system, the sound signals and image data are transmitted to the computer through the interface, and the computer is responsible for running the deep learning model for analysis and diagnosis.

[0111] S2. Build a data set.

[0112] S21. Collect data under normal and abnormal operating conditions of the transformer. Among them, when the transformer is in normal operation, the microphones and infrared cameras record audio signals and image data at the same time. These data are used to train the normal operation characteristics of the model. In addition, simulate common faults, such as loose transformer components, short-circuit impact, DC bias and overheating, and record the corresponding sound signals and infrared temperature images. These data are used to train the model to identify various fault types. Figure 1 It is the spectrogram of the transformer operation sound within a certain period of time.

[0113] S22. Data augmentation. To enhance the robustness of the model, the audio data is augmented by adding noise, and the image data is adjusted by brightness and contrast to simulate different working environments. These data augmentation techniques help to improve the accuracy of the model in practical applications.

[0114] S23. Data alignment. Since the frame rate of the infrared camera is 50 frames per second, and the input frames of the sound signal usually have 50% overlap per second, the audio signal is processed into audio data frames of 40 milliseconds to ensure the time alignment of the audio data and image data.

[0115] S24. Divide the data. Among all the datasets, the training set accounts for 70%, the validation set accounts for 20%, and the test set accounts for 10%. Ensure that the model can be evaluated on different datasets to optimize its generalization ability.

[0116] S3. Build a deep learning neural network model. As Figure 3 shown, its network structure is as follows:

[0117] S31. The model first receives T infrared temperature images and T audio signal frames captured from the microphone and the infrared camera. Among them, the image data is input into the residual neural network to extract T groups of image features, and the T audio frames extract features through Mel-frequency cepstral coefficients.

[0118] Among them, the structure of the residual neural network for extracting image features is as follows:

[0119] S311. Reset the image size. Here, the size of the infrared temperature image is adjusted to 224×224×3 to meet the input requirements of the convolutional neural network.

[0120] S312. Initial convolutional layer. Use a **7×7 convolution** to extract the preliminary features of the image and capture the context information of a larger area.

[0121] S313. Four consecutive 3×3 residual modules with the number of channels being 64, 128, 256, and 512 respectively. Introduce **residual connections** to solve the problem of gradient disappearance in the training of deep neural networks. Each residual block is usually composed of two 3x3 convolutional layers. Behind each convolutional layer, there is batch normalization and ReLU activation.

[0122] S314. Ordinary convolution and global average pooling. Ordinary convolution converts the 512-dimensional feature vector into a 256-dimensional one. Global average pooling compresses the spatial dimension of each feature map into a scalar, reducing the number of model parameters and preventing overfitting.

[0123] S32. The structure in S31 will be replicated T times to achieve the feature extraction of T consecutive infrared temperature images, thereby obtaining a 256×T feature vector.

[0124] S33. The part of the audio signal is first audio signal feature extraction. Extract 80×T audio features through MFCC.

[0125] S34. Feature fusion. This module fuses the 80×T feature vector of the audio signal with the feature vector of the image to form a 336×T feature vector. The fusion formula is as follows:

[0126]

[0127] Among them, is the finally fused feature matrix.

[0128] S35, after feature fusion, it will be input into the ECAPA-TDNN-based model, and its structure is as follows:

[0129] S351, one-dimensional ordinary convolution, rectified linear unit (Relu), and batch normalization. One-dimensional convolution performs convolution operations on the input features to extract local features; the ReLU activation function introduces non-linearity to enhance the model's expressive ability; batch normalization is used to accelerate training and stabilize the model.

[0130] S352, three residual blocks with squeeze-and-excitation mechanism. 3x3 convolution kernels are used to extract features, and the gradient vanishing problem is solved through residual connections; the SE (Squeeze-and-Excitation) mechanism adaptively adjusts the channel weights, strengthens important features, and compresses unimportant features. Different strides d are used for feature extraction to further improve the model's expressive ability and robustness.

[0131] Among them, the calculation of the SE mechanism includes squeeze operation and excitation operation. The squeeze operation compresses the spatial dimension of each channel into a scalar through global average pooling:

[0132]

[0133] where x h,w,c is the value of the input feature map at position (h, w) and channel C, H and W are the height and width of the feature map respectively, and z c is the compressed information of this channel.

[0134] The excitation operation generates channel weights through a fully connected layer and then normalizes them using the Sigmoid function:

[0135] s c = σ(W2δ(W1z)) = σ(W2δ(W1z))

[0136] where z is the compressed information obtained from the squeeze operation, σ is the ReLU activation function, W1 and W2 are the weights of the fully connected layer, σ is the Sigmoid function, and s c is the excitation coefficient of the channel.

[0137] S353, one-dimensional ordinary convolution and rectified linear unit.

[0138] S354, statistical pooling layer with attention and batch normalization. Features are extracted through statistical pooling, and the attention mechanism is combined to highlight the features of important time steps.

[0139] Among them, the calculation of the attention layer is as follows:

[0140]

[0141] Among them, f t is the attention score for time step t, and a t is the normalized attention weight.

[0142] S355, fully connected layer + batch normalization. Map the output of the pooling layer through the fully connected layer to adjust the dimension;

[0143] S356, angular margin cross-entropy loss function. AAM-softmax is an improved cross-entropy loss function that adds an angular margin to enhance the discrimination ability of the model, especially in the case of similar classes, making the classification more accurate. Its calculation is as follows:

[0144]

[0145] Among them, θ y is the angle of the correct class, θ i is the angle of other classes, m is the angular margin, and C is the total number of classes.

[0146] S6, training the deep learning model. Use the K-fold cross-validation method for training to ensure that the model can be optimized on different data subsets, avoid overfitting, and improve the generalization ability of the model. The steps are as follows:

[0147] S61. Use the constructed dataset for training to ensure that the model can learn various fault features that may occur during the operation of the transformer from multiple angles.

[0148] S62. Use the validation set and test set to verify and evaluate the model performance. The evaluation metrics are the recognition accuracy and recall rate of the faults that occur during the operation of the transformer.

[0149] S63. According to the results of S52, adjust the parameters of the model, including the size and stride of the convolutional layer of the residual block with the squeeze-and-excitation mechanism and the T value, to optimize the model performance. By continuously fine-tuning the model, gradually obtain the best training results as Figure 2 shown, where the T value is 20. The process of model training convergence is as Figure 3 shown. The model reaches convergence after about 600,000 iterations, and the accuracy rate is about 92%.

[0150] S7. Application of the model in the on-line monitoring system for transformer faults. The trained deep learning model is deployed to the computer control module of the on-line monitoring system for transformer faults. By receiving audio signals and image data from the microphone and infrared camera in real time, the computer inputs these data into the deep learning model for real-time analysis. Once the system detects a possible fault in the transformer, it will notify the technical personnel through an alarm to facilitate timely handling of the fault and avoid accidents.

[0151] In specific implementation, the present application provides a computer storage medium and a corresponding data processing unit. Among them, the computer storage medium can store a computer program, and when the computer program is executed by the data processing unit, it can run the inventive content of an on-line monitoring system and method for transformers that fuse sound signals and temperature data provided by the present invention and some or all of the steps in each embodiment. The storage medium can be a magnetic disk, an optical disk, a read-only memory (ROM), or a random access memory (RAM), etc.

[0152] Those skilled in the art can clearly understand that the technical solutions in the embodiments of the present invention can be implemented by means of a computer program and its corresponding general hardware platform. Based on such an understanding, the essence of the technical solutions in the embodiments of the present invention, or the part that contributes to the prior art, can be embodied in the form of a computer program, that is, a software product. This computer program software product can be stored in a storage medium and includes several instructions to enable a device including a data processing unit (which can be a personal computer, a server, a single-chip microcomputer, an MCU, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments of the present invention.

[0153] The present invention provides an idea and method for an on-line monitoring system and method for transformers that fuse sound signals and temperature data based on deep learning. There are many methods and ways to specifically implement this technical solution. The above description is only the preferred embodiment of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present invention, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present invention. Each component not clearly defined in this embodiment can be implemented by using the prior art.

Claims

1. A transformer online monitoring system based on deep learning that integrates sound signals and temperature data, characterized in that: include: Two microphones and two infrared temperature cameras are set on both sides of the power transformer to be monitored. The microphones and infrared temperature cameras are respectively connected to the controller. The controller performs real-time analysis of the input sound signals and temperature images through deep learning to complete online monitoring.

2. According to claim 1, a transformer online monitoring system based on deep learning integrating sound signals and temperature data is characterized in that: The microphone and the infrared temperature camera are respectively installed on both sides of the power transformer to be monitored, collect the sound signal and temperature image emitted when the monitored power transformer is running, and transmit them to the controller in real time.

3. A transformer online monitoring method based on deep learning integrating sound signals and temperature data, characterized in that: Using any of the systems described in claims 1 to 2 to perform real-time monitoring on a power transformer to be monitored comprises the following steps: Step 1, collecting data of the power transformer to be monitored under normal working conditions and abnormal working conditions to form a data set; Step 2, building a deep learning model to monitor faults during the operation of the power transformer to be monitored; Step 3, using the data set formed in step 1 to train the deep learning model constructed in step 2; Step 4: Use the trained deep learning model to perform real-time online monitoring during the operation of the power transformer to be monitored.

4. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 3 is characterized in that: The data collected in step 1 under the normal working state and the abnormal working state of the power transformer to be monitored include: Step 1-1, collecting sound signals and infrared temperature images synchronously recorded by microphones and infrared temperature cameras installed on both sides of the power transformer to be monitored when the power transformer to be monitored is in normal working state and abnormal working state; The abnormal working conditions include loose transformer components, short circuit shock, DC bias or over-temperature; Step 1-2, perform data enhancement on the sound signal and infrared temperature image respectively; Step 1-3, time-aligning the sound signal and the infrared temperature image.

5. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 4 is characterized in that: The deep learning model described in step 2 includes: Input module, sound signal feature extraction module, temperature image feature extraction module and feature fusion monitoring module; wherein, The input module is used to receive external sound signals and infrared temperature images; The sound signal feature extraction module and the temperature image feature extraction module extract input sound signal features and infrared temperature image features respectively; The feature fusion monitoring module fuses the sound signal features and the infrared temperature image features, and performs fault monitoring to obtain a final result.

6. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 5 is characterized in that: The temperature image feature extraction module comprises: Resizing the infrared temperature image received by the input module to adjust the infrared temperature image to a preset size; Through the initial convolution layer, a 7×7 convolution kernel is used to extract preliminary features of the infrared temperature image; It passes through four consecutive 3×3 residual modules, with the number of channels being 64, 128, 256, and 512 respectively; Through the ordinary convolution layer, the 512-dimensional feature vector is converted to 256 dimensions; Through global average pooling, the spatial size will be reduced.

7. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 6 is characterized in that: The number of the temperature image feature extraction modules is T, and T consecutive infrared temperature images are input into the same temperature image feature extraction module and spliced ​​to finally obtain infrared temperature image features.

8. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 7 is characterized in that: The sound signal feature extraction module obtains audio features through Mel-frequency cepstral coefficients to obtain sound signal features.

9. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 8 is characterized in that: The feature fusion monitoring module includes: Feature fusion module and feature monitoring module; among them, The input infrared temperature image features and sound signal features are fused through the feature fusion module, as shown below: in, is the fused feature matrix, It means that the fused feature matrix belongs to a 336-dimensional feature space and has T time steps, X image is the infrared temperature image feature, X audio is the sound signal feature; The fused feature matrix X fusion Input feature monitoring module to obtain detection results; The feature monitoring module includes a time-delay neural network model based on an attention-emphasized mechanism, and its specific structure is as follows: The fused feature matrix X fusion As input; Use one-dimensional ordinary convolution, linear rectification function, i.e. ReLU activation function, and batch normalization in turn for processing; Using three residual blocks with compression excitation mechanism for processing, the results of passing through one residual block, two residual blocks and three residual blocks respectively are input into one-dimensional ordinary convolution and linear rectification function for processing; Processed using statistical pooling layers with attention and batch normalization; Finally, a fully connected layer and batch normalization are used to map the image through the fully connected layer, and the cosine similarity is calculated with the candidate target category after adjusting the dimension. The detection result is determined according to the cosine similarity, that is, the detection result of the feature monitoring module is obtained; The time-delay neural network model based on the emphasis on attention mechanism is optimized using an angled marginal cross entropy loss function.

10. The transformer online monitoring method based on deep learning and fusion of sound signals and temperature data according to claim 9 is characterized in that: The training steps described in step 4 include: Step 4-1, using the training set and validation set obtained in step 2, adopt K-fold cross-training deep learning model; Step 4-2, using the test set obtained in step 2, test the trained model; Step 4-3, according to the feedback result of 4-2, fine-tune the size, step size and T value of the convolution layer of the residual block with the compression excitation mechanism, where the T value is 20; repeat steps 4-1 and 4-2 again until the prediction accuracy of the deep convolutional neural network model is higher than the threshold, and the training is completed.

Citation Information

Cited By

  • System and method for dispatching power equipment based on equipment operation state simulation

    CN120725374A