Transformer fault intelligent detection method and device based on deep learning multi-mode fusion, computer equipment and readable storage medium
Through deep learning multimodal fusion methods, image, infrared and sound sensor data are used to detect transformer faults, which solves the problems of incomplete and inaccurate detection in existing technologies, achieves fast and accurate fault identification, and improves the operation and maintenance efficiency and reliability of the power system.
Patent Information
- Application Number
- CN202510670983.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-23
- Publication Date
- 2025-09-12
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing technologies have difficulty in comprehensively and accurately detecting transformer faults, resulting in inefficient operation and maintenance of power systems and insufficient reliability.
A multimodal fusion method based on deep learning is adopted. RGB image data is collected by image sensors, infrared thermal imaging image data is collected by infrared sensors, and sound signals are collected by sound sensors. Fault detection is performed by combining preprocessing and a multimodal fault recognition model. The model includes a perception layer, a fusion layer, a backbone network, and an output layer.
It significantly improves the comprehensiveness and accuracy of transformer fault detection, shortens fault detection time from minutes to seconds, and improves operation and maintenance efficiency and system reliability.
Smart Images

Figure CN120632670A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence, and in particular to a method, device, computer equipment and readable storage medium for intelligent detection of transformer faults based on deep learning multimodal fusion. Background Art Summary of the Invention
[0002] The object of the present invention is to provide a transformer fault intelligent detection method, device, computer equipment and readable storage medium based on deep learning multimodal fusion.
[0003] In a first aspect, an embodiment of the present invention provides a transformer fault intelligent detection method based on deep learning multimodal fusion, comprising:
[0004] The image sensor collects RGB image data of the target transformer, the infrared sensor collects infrared thermal imaging image data of the target transformer, and the sound sensor collects sound signals of the target transformer;
[0005] Preprocessing the RGB image data, the infrared thermal imaging image data, and the sound signal, and using the preprocessed data as multimodal detection data;
[0006] The multimodal detection data is input into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer. The multimodal fault recognition model includes a perception layer, a fusion layer, a backbone network and an output layer that are cascaded in sequence.
[0007] In a possible implementation, inputting the multimodal detection data into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer includes:
[0008] Inputting the multimodal detection data into the perception layer to obtain feature map information corresponding to RGB image features, infrared thermal imaging features, and sound spectrum features;
[0009] Inputting feature map information corresponding to the RGB image features, the infrared thermal imaging features, and the sound spectrum map features into the fusion layer to obtain a fused feature map;
[0010] The fused feature map is processed by the backbone network and input into the output layer, and the output layer outputs the fault detection result corresponding to the target transformer.
[0011] In a possible implementation, the perception layer includes an RGB image perception module, an infrared thermal imaging perception module, and a sound information perception module;
[0012] The RGB image perception module is used to perceive the spatial information of the RGB image and capture the characteristics of the color and texture details in the scene. The infrared thermal imaging perception module is used to perceive the infrared radiation of the object and provide temperature distribution information to identify heat sources and cold spots. The sound information perception module is used to convert the sound into a visual spectrogram, which contains information in the frequency and time dimensions.
[0013] The RGB image perception module, the infrared thermal imaging perception module and the sound information perception module all adopt a three-layer convolution structure, and the step size of the first convolution structure of the three-layer convolution structure is set to 2, the step size of the second convolution structure is set to 1, and the step size of the third convolution structure is set to 2.
[0014] In a possible implementation, the fusion layer includes a fusion module, the fusion module includes a global pooling layer, a fully connected layer, a nonlinear activation layer and a normalization layer, and an attention module;
[0015] The global pooling layer, the fully connected layer, and the nonlinear activation layer are used to splice the feature map information corresponding to the RGB image features, the infrared thermal imaging features, and the sound spectrum map features in the depth direction. The attention module is used to multiply the original fused feature map and the channel attention coefficient element by element according to the channel to obtain the fused feature map.
[0016] In a possible implementation, the backbone network includes four residual modules, which implement back propagation by introducing residual connections and alleviate gradient vanishing and gradient exploding problems.
[0017] In one possible embodiment, the output layer includes two cascaded fully connected layers, wherein the first fully connected layer is a feature conversion layer for capturing the complex relationship between different features, and the second fully connected layer is associated with the classification or regression task, and the number of its neurons is equal to the number of categories of the task.
[0018] In a possible implementation, the multimodal fault recognition model is trained based on a cross entropy loss function; the training data of the multimodal fault recognition model includes pre-labeled training data including insulation status labels, temperature status labels, and oil level status labels;
[0019] The multimodal fault identification model deploys a model weight file and corresponding reasoning program code in a detection operation environment where the target transformer is located.
[0020] In a second aspect, an embodiment of the present invention provides a transformer fault intelligent detection device based on deep learning multimodal fusion, comprising:
[0021] an acquisition unit, configured to acquire RGB image data of a target transformer through an image sensor, acquire infrared thermal imaging image data of the target transformer through an infrared sensor, and acquire a sound signal of the target transformer through a sound sensor; preprocess the RGB image data, the infrared thermal imaging image data, and the sound signal, and use the preprocessed data as multimodal detection data;
[0022] The detection unit is used to input the multimodal detection data into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer. The multimodal fault recognition model includes a perception layer, a fusion layer, a backbone network and an output layer cascaded in sequence.
[0023] In a third aspect, an embodiment of the present invention provides a computer device, comprising a processor and a non-volatile memory storing computer instructions, wherein when the computer instructions are executed by the processor, the computer device executes the method described in the first aspect.
[0024] In a fourth aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect.
[0025] Compared to existing technologies, the present invention offers the following advantages: Using the disclosed method, device, computer equipment, and readable storage medium for intelligent transformer fault detection based on deep learning multimodal fusion, an image sensor collects RGB image data of the target transformer, an infrared sensor collects infrared thermal imaging image data, and an acoustic sensor collects acoustic signals. This preprocessed data is used as multimodal detection data and fed into a pre-trained multimodal fault recognition model consisting of a perception layer, a fusion layer, a backbone network, and an output layer in cascade order, thereby obtaining the target transformer fault detection result. This method integrates multimodal data, leverages the advantages of deep learning, improves the comprehensiveness and accuracy of transformer fault detection, and provides an effective means for intelligent power system operation and maintenance. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.
[0027] Figure 1 A schematic diagram of the steps of a transformer fault intelligent detection method based on deep learning multimodal fusion provided by an embodiment of the present invention;
[0028] Figure 2 A schematic diagram of the overall architecture of a multimodal fault identification model provided by an embodiment of the present invention;
[0029] Figure 3 A schematic diagram of the structure of a perception module provided in an embodiment of the present invention;
[0030] Figure 4 A schematic diagram of the structure of the fusion module provided in an embodiment of the present invention;
[0031] Figure 5 A schematic diagram of the structure of a residual module provided in an embodiment of the present invention;
[0032] Figure 6 A schematic diagram of the structure of an output module provided in an embodiment of the present invention;
[0033] Figure 7 A schematic block diagram of the structure of a transformer fault intelligent detection device based on deep learning multimodal fusion provided by an embodiment of the present invention;
[0034] Figure 8 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0035] To make the objectives, technical solutions, and advantages of the embodiments of the present invention more apparent, the technical solutions of the embodiments of the present invention will be described clearly and completely below in conjunction with the accompanying drawings of the embodiments of the present invention. It should be understood that the described embodiments are only a portion of the embodiments of the present invention, not all of them. Generally, the components of the embodiments of the present invention described and illustrated in the drawings herein may be arranged and designed in a variety of different configurations.
[0036] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0037] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of a transformer fault intelligent detection method based on deep learning multimodal fusion provided by an embodiment of the present disclosure. The transformer fault intelligent detection method based on deep learning multimodal fusion is introduced in detail below.
[0038] Step S201, collecting RGB image data of the target transformer through an image sensor, collecting infrared thermal imaging image data of the target transformer through an infrared sensor, and collecting sound signals of the target transformer through a sound sensor;
[0039] Step S202: preprocessing the RGB image data, the infrared thermal imaging image data, and the sound signal, and using the preprocessed data as multimodal detection data;
[0040] In step S203, the multimodal detection data is input into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer. The multimodal fault recognition model includes a perception layer, a fusion layer, a backbone network, and an output layer that are cascaded in sequence.
[0041] In this embodiment of the present invention, for example, a core oil-immersed main transformer within a 220kV substation is used as the monitoring target. To ensure its safe operation, the substation monitoring center deploys an edge server as the execution body. This server connects to multimodal sensors deployed on and around the transformer via Industrial Ethernet, enabling real-time data collection, preprocessing, and fault detection. First, the server collects multimodal data from the target transformer using multiple sensors. Among them, RGB image data is collected by an industrial camera installed 3 meters from the side of the transformer with a depression angle of 30 degrees. The camera has a resolution of 1920×1080 pixels and a frame rate of 15 frames per second, covering key parts such as the transformer body, bushing, and oil level gauge. Infrared thermal imaging image data is collected by a FLIRA50 thermal imager at the same location (640×480 pixels, temperature resolution 0.05°C, wavelength 7.5-13μm) to obtain the surface temperature distribution of the transformer. Sound signals are collected by a B&K4958 microphone (frequency response 20Hz-20kHz) installed on the bottom bracket of the transformer, 1 meter away from the body, to capture the mechanical vibration sound during operation. The server communicates with each sensor through the MQTT protocol and receives data every 5 seconds: RGB images are stored in JPEG format (such as T3_RGB_20240615_142000.jpg), infrared thermal images are stored in TIFF format (including the original temperature value, which needs to be converted to °C using the formula temperature = 0.04 × pixel value - 273.15, such as T3_IR_20240615_142000.tiff), and sound signals are stored in WAV format (sampling rate 44.1kHz, 16-bit quantization, duration 5 seconds, such as T3_Audio_20240615_142000.wav).
[0042] Secondly, the server preprocesses the collected multimodal data to form valid input. For RGB images, the server first checks the integrity and removes blurred images (for example, an image was discarded due to dithering) by calculating the image gradient variance (threshold 100). Valid images are filtered using a non-local mean filter (7×7 neighborhood, 21×21 search window, h=10) to remove Gaussian noise. The pixel values are then scaled from [0, 255] to [0, 1] (pixel value = original pixel value / 255) and resized to 224×224 using bilinear interpolation (maintaining the aspect ratio and filling the edges with black borders). For infrared thermal imaging, the server checks for bad pixels (temperature > 200°C or < -50°C), replaces outliers using a 3×3 median filter (for example, a pixel temperature of -100°C in a certain oil pillow area is corrected to the surrounding mean of 32°C), smoothes noise using a bilateral filter (spatial kernel σ = 5, range kernel σ = 0.1), maps the temperature values to [0, 1] (formula: normalized value = (temperature - minimum temperature) / (maximum temperature - minimum temperature)), and resizes them to 224×224 (aligned with the RGB image). For sound signals, the server detects silent segments (amplitude <-50dB) and extracts valid audio (for example, the ambient noise in the first 1 second of a certain audio is removed, and the last 4 seconds are retained). High-frequency (such as motor sound) and low-frequency (such as wind sound) interference are removed through a bandpass filter (20Hz-5kHz). The power spectrum is then calculated through a short-time Fourier transform (frame length 1024 points, frame shift 512 points, and Hamming window added). The Mel spectrum is weighted using 40 Mel filter banks (20Hz-8kHz). After taking the logarithm, a 128×224 Mel spectrum diagram is generated (224 frames on the time axis and 128 Mel bands on the frequency axis). After preprocessing, the three types of data are packaged into T3_20240615_142000_processed.pkl (containing a 224×224×3 RGB image, a 224×224×1 infrared thermal image, and a 128×224×1 Mel spectrum image) and input into the model as multimodal detection data.
[0043] Finally, the server invokes a pre-trained multimodal fault recognition model (deployed on an NVIDIA Tesla T4 GPU, with a single-sample processing time of ≤100ms) for inference. The model consists of a perception layer, a fusion layer, a backbone network, and an output layer. The perception layer processes the multimodal data through three modules: RGB images undergo three layers of convolution (3×3 kernels with strides of 2, 1, and 2) to output a 56×56×256 feature map (to extract details such as casing cracks); infrared thermal images undergo the same convolutional structure to output a 56×56×256 feature map (to extract temperature features such as localized high temperatures); and the sound spectrogram undergoes three layers of convolution (3×3 kernels with strides of (2, 2), (1, 1), and (2, 2)) to output a 32×56×256 feature map, which is then upsampled to 56×56×256 using bilinear interpolation to align the spatial dimensions. The fusion layer concatenates the three feature maps along the depth (56×56×750) and enhances key features using a channel-wise attention module. The 56×56×750 feature map is first spatially averaged pooled to obtain a 1×1×750 vector. This vector is then passed through two fully connected layers (750-128-750, ReLU+Sigmoid activation) to generate attention coefficients, which are then multiplied channel-by-channel with the original features (e.g., enhancing casing cracks, localized high temperatures, and high-frequency abnormal noise). The backbone network consists of a cascade of four residual modules, each containing a main path (3×3 convolution-BN-ReLU-3×3 convolution-BN) and a short-circuit path (1×1 convolution-BN). The main and short-circuit paths are summed and then activated with ReLU, outputting 14×14×512 deep features (covering both global and local correlations). The output layer is classified through two fully connected layers: the first layer flattens the 14×14×512 features and reduces their dimensionality to 512. The second layer inputs three independent fully connected submodules (corresponding to insulation, temperature, and oil level faults). Each submodule outputs a 2D vector (Sigmoid activation, with a threshold of 0.5 for abnormality). For example, if a sample outputs an insulation abnormality probability of 0.85 (abnormal), a temperature abnormality probability of 0.32 (normal), and an oil level abnormality probability of 0.67 (abnormal), the server generates an "insulation abnormality, oil level abnormality" warning message and pushes it to the operation and maintenance personnel via SMS or the monitoring screen.
[0044] In practice, this method has successfully detected numerous faults. For example, on May 10, 2024, a server detected an insulation anomaly with a probability of 0.92. On-site maintenance personnel discovered tiny cracks in the casing (shown in RGB images) and a 5°C temperature rise due to partial discharge (shown in infrared thermal images). The casing was promptly replaced to prevent breakdown. On June 2, 2024, an oil level anomaly with a probability of 0.78 was detected. Combining the oil level gauge image (oil level below the minimum mark) and the abnormal vibration of the oil pump (sound signal) confirmed a pipe blockage, which returned to normal after cleaning. Field tests have shown that this method has a fault detection accuracy of over 95% (F1 score of 0.96 on the test set), reducing the time required for manual inspections from 30 minutes to seconds, significantly improving maintenance efficiency and system reliability.
[0045] In an embodiment of the present invention, the step of inputting the multimodal detection data into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer may be implemented through the following example.
[0046] Inputting the multimodal detection data into the perception layer to obtain feature map information corresponding to RGB image features, infrared thermal imaging features, and sound spectrum features;
[0047] Inputting feature map information corresponding to the RGB image features, the infrared thermal imaging features, and the sound spectrum map features into the fusion layer to obtain a fused feature map;
[0048] The fused feature map is processed by the backbone network and input into the output layer, and the output layer outputs the fault detection result corresponding to the target transformer.
[0049] In an embodiment of the present invention, illustratively, the server first inputs the multimodal detection data into the perception layer of the model. The perception layer contains three independent feature extraction modules: for RGB images, the server processes them through a three-layer convolutional network (3×3 convolution kernel, with step sizes of 2, 1, and 2, respectively). The first layer reduces the 224×224×3 image to a 112×112×64 feature map, extracting preliminary features of the transformer's appearance contour (such as bushings and heat sinks); the second layer maintains a 112×112 size and captures details such as the oil level gauge scale and bushing surface cracks through 128 convolution kernels; the third layer further reduces it to a 56×56×256 feature map to complete the spatial feature extraction of the RGB image. The processing flow of the infrared thermal image is consistent with that of the RGB module. After inputting a 224×224×1 temperature map, the same convolution operation is performed to output a 56×56×256 feature map, focusing on extracting the temperature distribution features of local high-temperature areas (such as bushing connection points). The sound spectrogram is processed by a dedicated convolution module: the 128×224×1 Mel-spectrogram undergoes three layers of convolution (with strides of (2,2), (1,1), and (2,2). The first layer reduces the size to a 64×112×64 feature map, capturing the initial time-frequency correlation of the sound. The second layer maintains the 64×112 size and extracts the temporal features of abnormal sounds (such as high-frequency discharges). The third layer outputs a 32×56×256 feature map, which is then upsampled to 56×56×256 through bilinear interpolation to align the spatial dimensions of the first two feature maps. At this point, the perception layer outputs 56×56×256 feature maps for RGB, infrared, and sound.
[0050] The server then inputs the three feature maps into the fusion layer. The fusion layer first concatenates the three features depthwise to form a 56×56×750 (256×3) fused feature map. It then enhances key information through a channel-wise attention mechanism: the server performs spatial average pooling on the 56×56×750 feature map to obtain a 1×1×750 global statistical vector. This vector is passed through two fully connected layers (750-128-750, ReLU+Sigmoid activation) to calculate the attention coefficient (between 0 and 1) for each channel. For example, if a channel corresponds to the RGB features of a casing crack, the coefficient may be as high as 0.9, while the channel coefficient for background noise may be only 0.2. Finally, the server multiplies the attention coefficients with the original fused features channel by channel, enhancing fault-related features and suppressing redundant information, resulting in a weighted 56×56×750 fused feature map.
[0051] The fused feature map is further processed by the backbone network. The backbone network consists of a cascade of four residual modules, each of which contains a main path and a short-circuit path. The main path extracts deep features through two 3×3 convolutions (64-128-256) and batch normalization (BN). The short-circuit path uses a 1×1 convolution (64-256) and BN to match the dimensions. The two are summed and the output is activated using a Reluctant Unit (ReLU) to prevent vanishing gradients. For example, the first residual module reduces the 56×56×750 feature map to 28×28×512, and the second module reduces it to 14×14×512. The final output is a 14×14×512 deep feature map that integrates multimodal global and local correlations (such as the visual characteristics of the casing crack, the high temperature characteristics at the corresponding location, and the accompanying discharge acoustic characteristics).
[0052] Finally, the server inputs the deep features into the output layer. The output layer first flattens the 14×14×512 feature map into a 100,352-dimensional vector. This is then reduced to 512 dimensions through a fully connected layer to extract high-level abstract features. This 512-dimensional vector is then fed into three independent fully connected submodules (one for insulation, one for temperature, and one for oil level faults), each of which outputs a 2-dimensional vector (mapped to a 0-1 probability value after Sigmoid activation). The server sets a threshold of 0.5, and if a submodule outputs a probability ≥ 0.5, the system is considered abnormal. For example, if, at a certain moment, the model processes the input data, the insulation submodule outputs 0.85 (abnormal), the temperature submodule outputs 0.32 (normal), and the oil level submodule outputs 0.67 (abnormal), the server generates a detection result of "insulation abnormality, oil level abnormality" and pushes it to the operation and maintenance terminal through the monitoring system, prompting prompt repairs.
[0053] In an embodiment of the present invention, the perception layer includes an RGB image perception module, an infrared thermal imaging perception module, and a sound information perception module;
[0054] The RGB image perception module is used to perceive the spatial information of the RGB image and capture the characteristics of the color and texture details in the scene. The infrared thermal imaging perception module is used to perceive the infrared radiation of the object and provide temperature distribution information to identify heat sources and cold spots. The sound information perception module is used to convert the sound into a visual spectrogram, which contains information in the frequency and time dimensions.
[0055] The RGB image perception module, the infrared thermal imaging perception module and the sound information perception module all adopt a three-layer convolution structure, and the step size of the first convolution structure of the three-layer convolution structure is set to 2, the step size of the second convolution structure is set to 1, and the step size of the third convolution structure is set to 2.
[0056] In an embodiment of the present invention, exemplarily, the server first calls the RGB image perception module to process the preprocessed RGB image. The module adopts a three-layer convolution structure: the first convolution layer uses a 3×3 convolution kernel and a step size of 2 to downsample the 224×224×3 input image and output a 112×112×64 feature map. This step quickly reduces the spatial size through a larger step size while capturing the overall contour features of the transformer (such as the position and shape of the bushing, heat sink, and oil pillow); the second convolution layer maintains a 3×3 convolution kernel and a step size of 1 to perform fine-grained convolution on the 112×112×64 feature map. Feature extraction outputs a 112×112×128 feature map, focusing on capturing color and texture details (such as tiny cracks on the casing surface, color differences in the oil level gauge scale, and texture changes caused by paint peeling on the heat sink). The third convolutional layer again uses a 3×3 convolution kernel with a stride of 2 to downsample the 112×112×128 feature map to a 56×56×256 feature map, further compressing information redundancy while retaining key spatial features (such as the edge gradient of the crack and the outline of the oil level gauge scale).
[0057] The server then calls the infrared thermal imaging perception module to process the 224×224×1 infrared thermal image. This module has the same structure as the RGB module and also uses three layers of convolution (stride lengths of 2, 1, and 2): the first layer downsamples the input to a 112×112×64 feature map to extract macroscopic features of the transformer's overall temperature distribution (such as the temperature difference between the oil pillow and the transformer body, and the base temperature of the heat sink); the second layer maintains a stride length of 1 and outputs a 112×112×128 feature map to capture local temperature details (such as high temperatures at bushing connection points and low temperatures at blocked heat sink areas); the third layer, with a stride length of 2, outputs a 56×56×256 feature map to focus on key heat sources and cold spots (such as abnormally high temperatures at discharge locations and low temperatures at oil leaks). For example, when a bushing is locally overheated due to poor contact, the pixel values at the corresponding location in the module's third layer feature map will be significantly higher than those of the surrounding area, forming a clear high-temperature signature.
[0058] Finally, the server calls the sound information perception module to process the 128×224×1 Mel-spectrogram. This module also uses a three-layer convolutional structure (stride lengths of 2, 1, and 2). However, because the input is two-dimensional time-frequency data (128 Mel-frequency bands × 224 time frames), the convolution kernel operates separately in the time and frequency dimensions: the first layer, with a stride length of (2, 2), reduces the 128×224×1 input to a 64×112×64 feature map, extracting the basic time-frequency correlation of the sound (such as the distribution of a 50Hz fundamental frequency "hum"). The second layer, with a stride length of (1, 1), outputs a 64×112×128 feature map, capturing abnormal frequency components (such as transient pulses of high-frequency discharge sound and periodic noise caused by oil pump jams). The third layer, with a stride length of (2, 2), outputs a 32×56×256 feature map, which is then upsampled to 56×56×256 through bilinear interpolation (aligning the spatial dimensions of the first two feature maps). For example, when partial discharge occurs inside a transformer, a continuous energy peak will appear in the high frequency band (3kHz-5kHz) of the Mel-spectrogram. After processing by this module, the eigenvalue of the corresponding frequency-time region in the feature map will be significantly enhanced.
[0059] Through the processing of the above three-layer convolutional structure, the server ultimately obtains three types of 56×56×256 feature maps from the perception layer: spatial texture features output by the RGB module, temperature distribution features output by the infrared module, and time-frequency correlation features output by the sound module. These features correspond to the transformer's appearance, thermal anomaly location, and mechanical vibration anomaly, respectively, providing key input for multimodal information complementation and fault identification in the subsequent fusion layer. For example, when the RGB module detects a bushing crack, the infrared module detects high temperature at the crack site, and the sound module detects an abnormal discharge sound, the feature values of the corresponding areas in the three feature maps will increase significantly, providing multi-dimensional evidence support for the model's subsequent judgment of "insulation anomaly."
[0060] In an embodiment of the present invention, the fusion layer includes a fusion module, and the fusion module includes a global pooling layer, a fully connected layer, a nonlinear activation layer and a normalization layer, and an attention module;
[0061] The global pooling layer, the fully connected layer, and the nonlinear activation layer are used to splice the feature map information corresponding to the RGB image features, the infrared thermal imaging features, and the sound spectrum map features in the depth direction. The attention module is used to multiply the original fused feature map and the channel attention coefficient element by element according to the channel to obtain the fused feature map.
[0062] In an embodiment of the present invention, illustratively, the server first splices the three types of feature maps output by the perception layer in the depth direction. Since the spatial size of the RGB, infrared, and sound feature maps are all 56×56, and the number of channels is 256, a 56×56×750 (256×3) original fusion feature map is formed after splicing. For example, when a tiny crack appears in the transformer bushing, the gradient feature of the crack edge in the RGB feature map (corresponding to the 100th-200th channels), the high temperature feature at the crack in the infrared feature map (corresponding to the 300th-400th channels), and the high-frequency feature of the abnormal discharge sound in the sound feature map (corresponding to the 600th-700th channels) will be concentrated in the corresponding channel position of the fusion feature map.
[0063] The server then calls the global pooling layer to process the original fused feature map. The global pooling layer performs average pooling on the 56×56×750 feature map along the spatial dimension (56×56), averaging the 56×56 pixel values in each channel to generate a 1×1×750 global statistical vector. This vector compresses spatial information while preserving the global importance statistics of each channel. For example, the pixel average value of the RGB feature channel corresponding to the crack (channel 150) is 0.8 (indicating that this feature is widely present in the image), while the average value of the background noise channel (channel 50) is only 0.1 (indicating that this feature is sparsely distributed).
[0064] The global statistical vector is processed by the fully connected layer and the nonlinear activation layer to generate the channel attention coefficient. The server first inputs the 1×1×750 vector into the first fully connected layer (750-128), and uses the ReLU activation function (max(0,x)) to filter negative information and enhance nonlinear expression; then, the 128-dimensional vector is input into the second fully connected layer (128-750), and the sigmoid activation function (1 / (1+e -x )) maps the output to the interval [0, 1], resulting in a channel attention coefficient of 1×1×750. For example, the RGB feature channel corresponding to cracks (channel 150) has a coefficient of 0.9 (high weight), the infrared channel corresponding to high temperature (channel 350) has a coefficient of 0.85 (second highest weight), the sound channel corresponding to discharge sound (channel 650) has a coefficient of 0.8 (medium weight), and the background noise channel (channel 50) has a coefficient of only 0.2 (low weight).
[0065] Finally, the server uses the attention module to multiply the original fused feature map by the channel attention coefficient channel by channel to obtain the final fused feature map. The specific operation is: the 56×56 pixel values of each channel in the 56×56×750 original feature map are multiplied by the attention coefficient of the corresponding channel. For example, each pixel value of the 150th channel (crack feature) is multiplied by 0.9 to further enhance the gradient characteristics of the crack edge; the pixel value of the 350th channel (high temperature feature) is multiplied by 0.85 to highlight the temperature difference in the high temperature area; the pixel value of the 650th channel (discharge sound feature) is multiplied by 0.8 to retain the time-frequency correlation of high-frequency abnormal sounds; and the pixel value of the 50th channel (background noise) is multiplied by 0.2 to suppress the interference of irrelevant information.
[0066] Through the above process, the fusion layer ultimately outputs a 56×56×750 fused feature map, in which fault-related multimodal features (such as the visual characteristics of cracks, the thermal characteristics of high temperatures, and the acoustic characteristics of discharges) are significantly enhanced and redundant background noise is suppressed, providing more focused input for the subsequent deep feature extraction and fault classification of the backbone network. For example, when the server detects a casing crack, the channel feature values corresponding to the crack, high temperature, and discharge sound in the fused feature map are significantly higher than those of other areas. The model can then accurately determine "insulation abnormality" faults based on these enhanced features, achieving complementary and collaborative recognition of multimodal information.
[0067] In an embodiment of the present invention, the backbone network includes four residual modules, which implement back propagation by introducing residual connections and alleviate gradient vanishing and gradient exploding problems.
[0068] In an embodiment of the present invention, illustratively, the server inputs the fused feature map into the first residual module. This module includes a main path and a short-circuit path: the main path reduces the 56×56×750 input to a 28×28×512 feature map (extracting local correlation features of casing cracks, high-temperature areas, and discharge sounds) through two 3×3 convolutions (step size 2, padding 1), batch normalization (BN), and ReLU activation; the short-circuit path synchronously adjusts the input size to 28×28×512 (matching the output dimension of the main path) through 1×1 convolution (step size 2, padding 0) and a BN layer. The outputs of the main path and the short-circuit path are element-by-element added and then activated by ReLU to form a 28×28×512 residual feature map. This design allows the gradient to be directly back-propagated through the short-circuit path, avoiding gradient attenuation caused by multi-layer nonlinear transformations in the main path.
[0069] The second residual module processes a 28×28×512 feature map: the main path maintains a stride of 1 (3×3 convolution) and outputs a 28×28×512 feature map (capturing the spatial overlap of hot spots and cracks); the short-circuit path is directly connected across layers (without resizing) and activated after addition to the main path output to further enhance gradient transfer.
[0070] The third and fourth residual modules, respectively, reduce the feature maps to 14×14×512 (for the third module) and 7×7×512 (for the fourth module). Residual connections are used at each step to preserve the original feature information. For example, when the fused feature map contains multimodal features such as casing cracks (RGB), localized high temperatures (infrared), and abnormal discharge noise (sound), residual connections ensure that these key features are not overwhelmed during deep transmission. Gradients can be directly propagated back along a short-circuit path, avoiding the vanishing or exploding gradients caused by the nonlinear operations of multi-layer convolution. This preserves detailed information such as crack edges, high-temperature zone boundaries, and the time-frequency points of abnormal noise.
[0071] By cascading four residual modules, the server ultimately outputs a 7×7×512 deep feature map. This feature integrates multimodal global and local correlations, providing highly discriminative input for fault classification in the output layer, ensuring that the model can still accurately identify transformer faults (such as insulation abnormalities and oil level abnormalities) in complex scenarios.
[0072] In an embodiment of the present invention, the output layer includes two cascaded fully connected layers, wherein the first fully connected layer is a feature conversion layer for capturing the complex relationship between different features, and the second fully connected layer is associated with the classification or regression task, and the number of its neurons is equal to the number of categories of the task.
[0073] In an embodiment of the present invention, for example, the server first flattens the 14×14×512 feature map output by the backbone network into a one-dimensional vector. The flattening operation directly multiplies the spatial dimension (14×14) of the three-dimensional feature map with the channel dimension (512) to obtain a long vector of 14×14×512=100352 dimensions. This vector contains global and local features after multimodal fusion, such as the edge gradient of the casing crack (from RGB), the high temperature distribution at the crack (from infrared), the time-frequency correlation of the discharge abnormal sound (from sound), and other information, all of which are encoded as different dimensional values in the vector.
[0074] The flattened 100,352-dimensional vector is input to the first fully connected layer (feature conversion layer). This layer contains 512 neurons, which use fully connected operations to reduce the dimensionality of high-dimensional features and extract abstract relationships. Each neuron receives a linear combination of the 100,352-dimensional input (with a weight matrix of 100,352×512). After batch normalization (BN) and ReLU activation (max(0,x)), it outputs a 512-dimensional abstract feature vector. The key to this step is to capture the complex relationships between multimodal features. For example, one dimension of the vector may encode the relationship between the "casing crack edge gradient" (RGB feature) and the "5°C temperature increase at the crack" (infrared feature), while another dimension may encode the relationship between the "high-frequency discharge sound duration" (acoustic feature) and the "oil level gauge scale drop" (RGB feature). This integrates the scattered multimodal information into a more discriminative abstract representation.
[0075] The 512-dimensional abstract feature vector is then input into the second fully connected layer (classification layer). Since the transformer fault detection task includes three categories (three categories), namely insulation abnormality, temperature abnormality, and oil level abnormality, the second fully connected layer is configured with three neurons, each corresponding to one type of fault. The server maps the 512-dimensional input to a 3-dimensional output vector through a fully connected operation (the weight matrix is 512×3), and then passes it through the Sigmoid activation function (1 / (1+e -x ))Map the value of each dimension to the [0,1] interval and output the probability values of the three types of failures.
[0076] For example, when a transformer experiences partial discharge due to a bushing crack, the deep feature map output by the backbone network shows a significant increase in the gradient of the crack edge, the temperature difference in the high-temperature zone, and the high-frequency energy of the discharge sound. The corresponding dimensions in the flattened 100,352-dimensional vector are relatively high. After processing by the first fully connected layer, the dimensions associated with "insulation anomaly" in the 512-dimensional abstract feature vector are strengthened (for example, the dimension encoding the crack-high-temperature-abnormal sound association is 0.7). After processing by the second fully connected layer, the output vector is [0.85, 0.32, 0.67], corresponding to an insulation anomaly probability of 0.85 (≥0.5, considered abnormal), a temperature anomaly probability of 0.32 (<0.5, considered normal), and an oil level anomaly probability of 0.67 (≥0.5, considered abnormal). Based on this, the server generates detection results for "insulation anomaly and oil level anomaly" and pushes them to the operation and maintenance terminal through the monitoring system, prompting the operator to inspect the bushing and check the oil level.
[0077] Through the two fully connected layers of the output layer, the server converts the deep features of multimodal fusion into specific fault probabilities, realizing end-to-end mapping from images, thermal imaging, and sound information to fault types, ensuring the accuracy and interpretability of the detection results.
[0078] In an embodiment of the present invention, the multimodal fault recognition model is trained based on a cross entropy loss function; the training data of the multimodal fault recognition model includes pre-labeled training data including insulation status labels, temperature status labels, and oil level status labels;
[0079] The multimodal fault identification model deploys a model weight file and corresponding reasoning program code in a detection operation environment where the target transformer is located.
[0080] In an embodiment of the present invention, for example, the server first collects and organizes multimodal training data. The data source is the historical operation records of the substation in the past three years, including multimodal data (RGB images, infrared thermal images, and sound signals) of normal status and three types of faults: insulation abnormality, temperature abnormality, and oil level abnormality. For example, the insulation abnormality data includes corresponding samples of casing cracks (RGB image display), local discharge high temperature (infrared thermal image shows a temperature increase of 5-10°C), and high-frequency discharge sound (3-5kHz energy surge in the sound signal); the temperature abnormality data includes samples of heat sink blockage (infrared thermal image shows a local low temperature area) and abnormal oil pillow temperature (higher than 80°C); the oil level abnormality data includes samples of the oil level gauge scale being below the lowest line (RGB image display) and abnormal vibration sound of the oil pump (periodic abnormal sound appears in the sound signal).
[0081] All data is manually labeled by operations and maintenance experts. Labels are formatted as three-dimensional binary vectors (insulation status, temperature status, and oil level status). Normal status is labeled [0,0,0], insulation anomalies are [1,0,0], temperature anomalies are [0,1,0], oil level anomalies are [0,0,1], and when multiple faults coexist, labels are [1,1,0] and other combinations. The server divides the labeled multimodal data into training, validation, and test sets in an 8:1:1 ratio (e.g., the training set contains 10,000 samples, the validation set 1,250, and the test set 1,250). These data are stored in the server's local database (path: / data / training_set / ).
[0082] The server uses a GPU (NVIDIA Tesla T4) to initiate the training process. The initial model weights are randomly initialized (normal distribution, mean 0, standard deviation 0.01). During training, the server inputs preprocessed multimodal data (224×224×3 RGB images, 224×224×1 infrared thermal images, and 128×224×1 Mel-spectrograms) into the model. The data is processed sequentially through the perception layer, fusion layer, backbone network, and output layer, outputting predicted probability vectors (e.g., [0.75, 0.2, 0.6]) for the three types of faults.
[0083] The model uses a cross-entropy loss function to calculate the difference between the predicted and true labels. The total loss is the sum of the losses for the three fault categories. For example, for an insulation anomaly sample, the true label is [1, 0, 0], and the model outputs a predicted probability of [0.85, 0.1, 0.05]. The cross-entropy loss for the insulation anomaly is -1×log(0.85)≈0.163, while the cross-entropy loss for temperature and oil level is -0×log(0.1)-0×log(0.05)=0, for a total loss of 0.163. The server updates model weights (such as the weights of the convolution kernels in the perception layer and the biases of the fully connected layers in the fusion layer) through backpropagation (using the Adam optimizer and a learning rate of 0.001), gradually reducing the total loss. During training, the server validates every 100 batches (using the validation set). When the validation loss stops decreasing for five consecutive rounds (e.g., from 0.25 to 0.22 and then stabilizes), the early stopping mechanism is triggered and the current optimal weights file (best_model_weights.pth) is saved.
[0084] After training is complete, the server deploys the optimal weight file and inference code (inference.py) to the transformer's detection runtime environment (i.e., the edge server in the monitoring center). The deployment steps are as follows:
[0085] 1. Weight file migration: Copy best_model_weights.pth from the training server to the / model / directory of the edge server to ensure file integrity (check the MD5 hash value to ensure it is consistent with the training server).
[0086] 2. Inference code configuration: Configure the input path of inference.py (pointing to the real-time preprocessing data directory / data / processed / ) and the output path (pointing to the detection result directory / result / ), and set the threshold (0.5, with a probability ≥ 0.5 considered an anomaly).
[0087] 3. Environment dependency check: The server verifies whether dependencies such as the CUDA version (11.7) and PyTorch version (2.0.0) match, ensuring that the model can call GPU-accelerated inference (single sample processing time ≤ 100ms).
[0088] 4. Post-deployment verification: The server uses test set data (e.g., 1,250 groups of samples) for deployment verification. The test accuracy reaches 95% (1,188 groups are correctly classified), and the F1 score is 0.96 (0.97 for insulation anomaly, 0.95 for temperature anomaly, and 0.96 for oil level anomaly), confirming that the model performance meets the requirements.
[0089] After deployment, the server receives preprocessed data in real time and invokes model inference. For example, when a transformer experiences insulation anomaly due to a bushing crack, the model outputs a predicted probability of [0.92, 0.28, 0.15]. The server identifies an "insulation anomaly" and generates an alert (timestamp: 2024-07-15 14:30:00), which is then pushed to the operation and maintenance terminal. Field verification confirmed the presence of tiny cracks on the bushing surface (as shown in the RGB image), a temperature of 78°C at the crack (50°C higher than normal), and a 3kHz discharge sound was detected in the acoustic signal, consistent with the model's detection results, validating the effectiveness of the training and deployment.
[0090] In order to more clearly describe the solution provided by the embodiment of the present invention, a more complete implementation method is provided below.
[0091] Multimodal data acquisition and processing:
[0092] Collect multimodal data: Use an image sensor to collect RGB image data of the transformer; use an infrared sensor to collect infrared thermal imaging image data of the transformer; use an acoustic sensor to collect the acoustic signal of the transformer; in addition, preprocess the collected multimodal data, including data cleaning, denoising, normalization and other steps, to improve the overall data quality.
[0093] The multimodal data is manually labeled according to the actual situation, and labeled with multiple labels such as normal insulation, abnormal insulation, normal temperature, abnormal temperature, normal oil level, and abnormal oil level.
[0094] The manually annotated multimodal data is divided into a training set and a test set in a 7:3 ratio. The training set is used to train the model, and the test set is used to evaluate model performance. This data is only used during model training; after model training, it is no longer needed for subsequent detection, recognition, and testing. In a real-world deployment, only the model weight file and the corresponding inference program code are required to perform prediction tasks.
[0095] Multimodal fault identification model:
[0096] This solution uses a deep learning approach to classify and identify common transformer faults, such as insulation anomalies, temperature anomalies, and oil level anomalies. Since the fault input source contains multiple modal information, the information of different modalities needs to be sensed separately and the modal features extracted, and then the modal features are fused before classification and identification. To this end, the present invention proposes and constructs a multimodal fault recognition model (Multimodal Exception Recognition Model), referred to as MERM; the model network structure includes a perception layer, a fusion layer, a backbone network, and an output layer. The detailed description is as follows, please refer to Figure 2 , Figure 2A schematic diagram of the overall architecture of a multimodal fault identification model provided by an embodiment of the present invention.
[0097] Perception layer: As the input layer of the MERM model, it consists of three perception modules, which are used to perceive and process data from different input sources and extract spatial and temporal features of different modalities.
[0098] RGB Image Perception Module: This module is used to perceive the spatial information of RGB images and capture the color and texture details of the scene. Since these are digital images, they can be directly perceived using the convolution module.
[0099] Infrared thermal imaging perception module: This module senses infrared radiation from objects, provides temperature distribution information, and helps identify heat sources and cold spots. Since these images are digital, they can be directly perceived using the convolution module.
[0100] The Sound Information Perception Module converts sound into a visual spectrogram, containing information in both frequency and time dimensions, suitable for analyzing the time series characteristics of audio events. Sound information is a one-dimensional signal and requires a short-time Fourier transform (STFT) to convert it into a spectrogram. This solution uses a common industry approach, calculating the Mel-frequency representation of the sound signal and generating a Mel-frequency map, which can then be used for perception using the convolution module.
[0101] See also Figure 3 , Figure 3 A schematic diagram of the perception module structure provided in an embodiment of the present invention, wherein each perception module adopts a three-layer convolution structure: Conv-BN-Relu-Conv-BN-Relu-Conv-BN-Relu. This structure can not only effectively extract key features in the image, but also enable the model to learn more complex mapping relationships by introducing the nonlinear activation function Relu. In the three-layer convolution structure, a different stride is set for each layer of convolution. The stride of the first layer of convolution is set to stride = 2, which means that during the convolution process, the filter moves with a step size of 2 pixels. This setting helps to quickly reduce the size of the image, reduce the amount of calculation, and retain the key information in the image. The stride of the second layer of convolution is set to stride = 1, that is, the filter only moves 1 pixel each time it moves, which helps to further extract the detailed features of the image. The stride of the third layer of convolution is set to stride = 2 again to further reduce the image size and prepare to pass the feature map to the subsequent fusion module.
[0102] Fusion layer: It consists of a fusion module (Fusion Module) that is used to fuse the feature information of RGB images, infrared thermal imaging, and sound spectrograms. In order to achieve cross-modal information fusion, the feature stacking method is used to splice the feature maps extracted from each modality in the depth direction. In addition, the fusion module adds a channel attention module (AttentionMechanism) to enhance the weight of key features and improve the fusion effect. This module dynamically assigns different weights according to the importance of each channel to the current task. In this way, the model can focus more on features that are critical to the task and ignore less relevant parts, thereby enhancing the overall feature representation ability and decision accuracy of the model. The channel attention module is mainly composed of key components such as GlobalPooling, fully connected layer (FC), nonlinear activation function (such as sigmoid), and attention application. Please refer to Figure 4 , Figure 4 This is a schematic diagram of the structure of the fusion module provided in an embodiment of the present invention. The specific structure is described as follows:
[0103] Global Pooling: This step performs a global pooling operation on the feature maps of each channel, reducing the two-dimensional feature maps into a one-dimensional vector. This vector concisely summarizes the feature information at all locations within the corresponding channel and is crucial for subsequent feature weight learning. Global pooling allows for a holistic understanding of the feature distribution of each channel, providing a comprehensive feature description for subsequent steps.
[0104] The fully connected layer (FC) fully connects the vectors output by the GlobalPooling layer and learns the weights for each channel through linear transformations and nonlinear activation functions. These weights essentially reflect the importance of different channels to the current task and form the core of the attention mechanism. Through the processing of the fully connected layer, the dependencies between channels can be captured, providing a basis for subsequent attention allocation.
[0105] Non-linear activation and normalization: After the fully connected layer, non-linear activation functions such as sigmoid are usually used to normalize the weights to ensure that the weight values of all channels are between 0 and 1 and the sum of the weights is 1. This process converts the weights into probabilistic attention coefficients, which facilitates accurate feature weighting in subsequent network layers.
[0106] Attention Application: The resulting channel attention coefficient vector is element-wise multiplied with the original feature map. This step adaptively weights each channel's features. The weighted feature map not only preserves the original spatial structure but also more precisely regulates the influence of each channel. This allows the network to focus more closely on feature channels that contribute most to the task, thereby improving overall feature representation and task performance.
[0107] Backbone network: It consists of 4 residual modules, which are constructed by introducing residual connections. Residual connections allow input data to bypass one or more layers and be added to the output of deeper layers. This not only simplifies the information flow path, but also ensures that gradients can be effectively back-propagated even when the network depth increases. This mechanism effectively alleviates the common problems of gradient vanishing and gradient exploding during deep neural network training. For more information on the residual module structure, please refer to Figure 5 ,.
[0108] Output Module:
[0109] This solution considers common transformer faults, such as insulation abnormalities, temperature abnormalities, and oil level abnormalities for classification and identification. Since the categories are not mutually exclusive, multiple states may appear at the same time, which makes it impossible to directly use the image multi-classification recognition method. The currently commonly used deep learning CNN network (such as MobileNet, VGG16, ResNet) is mainly used for multi-class image recognition tasks and cannot support multi-label multi-task image recognition. In order to solve this problem, the present invention constructs an output module that supports multi-label multi-task, please refer to Figure 6 .
[0110] The output module consists of two fully connected layers. The first fully connected layer is typically used as a feature transformation layer, capturing the complex relationships between different features and preparing the input for the final fully connected layer. The second fully connected layer is directly associated with the classification or regression task, and its number of neurons is typically equal to the number of categories in the task. Since only the presence of each fault needs to be determined, the final fully connected layer has two output channels, where 0 indicates the fault is absent and 1 indicates the fault is present. Please refer to Table 1 for a table of fault states provided in embodiments of the present invention.
[0111] Table 1
[0112] Label 0 1 Insulation status Normal insulation Insulation abnormality Temperature status Normal temperature Abnormal temperature Oil level status Oil level normal Abnormal oil level
[0113] Model training: The MERM model was trained and optimized using the training set; the MERM model was tested and the model performance was evaluated using the test set. During model training, the input images were uniformly scaled to 224×224. To improve data diversity and the generalization ability of the model, the training data image enhancement methods used a combination of random flipping, random cropping, random rotation, and random color transformation (such as brightness, contrast, saturation, and hue adjustment). In terms of training parameters, the batch size batch_size = 64, the optimization algorithm used the Adam optimizer, the initial learning rate was set to lr = 0.001, the cross-entropy loss function (Cross-Entropy Loss) was used, and the entire data set was iterated epoch = 200 times. During the training process, after each complete data iteration, the program calculated the accuracy of the model on the test set and saved the model weights with the highest current accuracy. In subsequent actual use, only the model weight file and the corresponding inference program code need to be retained to achieve efficient recognition of new images.
[0114] System deployment: The trained deep learning model is deployed into the power monitoring system to monitor the status of the transformer, analyze the fault type, and automatically issue early warning information when an anomaly is detected, so that operation and maintenance personnel can take timely measures.
[0115] In summary, the present invention proposes a transformer fault intelligent detection method based on deep learning multimodal fusion. In specific implementation, the method is mainly used in the daily operation and maintenance of the power system, and can monitor the status of the transformer in real time and detect potential fault hazards in a timely manner. By comprehensively analyzing data from multiple sensors, operation and maintenance personnel can more accurately judge the working status of the transformer and take corresponding maintenance measures, thereby improving the operating efficiency and reliability of the power system. The present invention has the following advantages and innovations: First, a multimodal fault recognition model (MERM) is proposed, and its network structure covers a perception layer, a fusion layer, a backbone network and an output layer. The model can perceive and fuse monitoring information from multiple sensors to achieve high-precision classification and identification of transformer faults; second, multimodal feature fusion is realized. The present invention fuses multiple monitoring information such as image sensors, infrared sensors and sound sensors to achieve more comprehensive fault feature extraction and improve fault detection accuracy. At the same time, an attention module is added to feature fusion. By assigning different importance weights to different channels, the model can focus on features that are more critical to the current task, improving the model's overall feature representation capabilities. Third, the present invention applies deep learning algorithms for feature extraction and classification, avoiding the limitations of traditional methods that rely on manual experience and improving the automation and intelligence level of diagnosis. In addition, the present invention also combines real-time and accuracy, can quickly and accurately diagnose transformer faults, and provide strong guarantees for the safe and stable operation of the power system.
[0116] Please refer to Figure 7 , Figure 7 An embodiment of the present invention provides a transformer fault intelligent detection device 110 based on deep learning multimodal fusion, including:
[0117] An acquisition unit 1101 is configured to acquire RGB image data of a target transformer through an image sensor, acquire infrared thermal imaging image data of the target transformer through an infrared sensor, and acquire sound signals of the target transformer through a sound sensor; preprocess the RGB image data, the infrared thermal imaging image data, and the sound signals, and use the preprocessed data as multimodal detection data;
[0118] The detection unit 1102 is used to input the multimodal detection data into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer. The multimodal fault recognition model includes a perception layer, a fusion layer, a backbone network and an output layer cascaded in sequence.
[0119] It should be noted that the implementation principles of the aforementioned intelligent transformer fault detection device 110 based on deep learning multimodal fusion can be referenced from the implementation principles of the aforementioned intelligent transformer fault detection method based on deep learning multimodal fusion, and will not be elaborated upon here. It should be understood that the division of the various modules of the aforementioned device is merely a division of logical functions. In actual implementation, they can be fully or partially integrated into a single physical entity, or physically separated. Furthermore, these modules can be implemented entirely in the form of software invoked by a processing element; or entirely in the form of hardware; or some modules can be implemented in the form of software invoked by a processing element, while others are implemented in hardware. For example, the intelligent transformer fault detection device 110 based on deep learning multimodal fusion can be a separate processing element, or it can be integrated into a chip of the aforementioned device. Furthermore, it can be stored in the form of program code in the memory of the aforementioned device, and invoked and executed by a processing element of the aforementioned device. The implementation of the other modules is similar. Furthermore, these modules can be fully or partially integrated together, or implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. During implementation, each step of the above method or each module above may be completed by an integrated logic circuit of hardware in a processor element or by instructions in the form of software.
[0120] For example, the above modules may be one or more integrated circuits configured to implement the above methods, such as one or more application specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field programmable gate arrays (FPGAs). For another example, when a module is implemented by scheduling program code on a processing element, the processing element may be a general-purpose processor, such as a central processing unit (CPU) or other processor that can call program code. For another example, these modules may be integrated together and implemented in the form of a system-on-a-chip (SOC).
[0121] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned transformer fault intelligent detection device 110 based on deep learning multimodal fusion. Figure 8 As shown, Figure 8 This is a block diagram of a computer device 100 according to an embodiment of the present invention. The computer device 100 includes a transformer fault intelligent detection device 110 based on deep learning multimodal fusion, a memory 111 , a processor 112 , and a communication unit 113 .
[0122] In order to realize the transmission or interaction of data, the memory 111, the processor 112 and the communication unit 113 are electrically connected to each other directly or indirectly. For example, the electrical connection between these components can be realized through one or more communication buses or signal lines. The intelligent detection device 110 for transformer faults based on deep learning multimodal fusion includes at least one software function module that can be stored in the memory 111 in the form of software or firmware or solidified in the operating system (OS) of the computer device 100. The processor 112 is used to execute the intelligent detection device 110 for transformer faults based on deep learning multimodal fusion stored in the memory 111, such as the software function modules and computer programs included in the intelligent detection device 110 for transformer faults based on deep learning multimodal fusion.
[0123] An embodiment of the present invention provides a readable storage medium, which includes a computer program. When the computer program is running, it controls the computer device where the readable storage medium is located to execute the aforementioned transformer fault intelligent detection device 110 based on deep learning multimodal fusion.
[0124] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.
Claims
1. A transformer fault intelligent detection method based on deep learning multimodal fusion, characterized in that: include: The image sensor collects RGB image data of the target transformer, the infrared sensor collects infrared thermal imaging image data of the target transformer, and the sound sensor collects sound signals of the target transformer; Preprocessing the RGB image data, the infrared thermal imaging image data, and the sound signal, and using the preprocessed data as multimodal detection data; The multimodal detection data is input into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer. The multimodal fault recognition model includes a perception layer, a fusion layer, a backbone network and an output layer that are cascaded in sequence.
2. The method according to claim 1, characterized in that Inputting the multimodal detection data into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer includes: Inputting the multimodal detection data into the perception layer to obtain feature map information corresponding to RGB image features, infrared thermal imaging features, and sound spectrum features; Inputting feature map information corresponding to the RGB image features, the infrared thermal imaging features, and the sound spectrum map features into the fusion layer to obtain a fused feature map; The fused feature map is processed by the backbone network and input into the output layer, and the output layer outputs the fault detection result corresponding to the target transformer.
3. The method according to claim 2, characterized in that The perception layer includes an RGB image perception module, an infrared thermal imaging perception module, and a sound information perception module; The RGB image perception module is used to perceive the spatial information of the RGB image and capture the characteristics of the color and texture details in the scene. The infrared thermal imaging perception module is used to perceive the infrared radiation of the object and provide temperature distribution information to identify heat sources and cold spots. The sound information perception module is used to convert the sound into a visual spectrogram, which contains information in the frequency and time dimensions. The RGB image perception module, the infrared thermal imaging perception module and the sound information perception module all adopt a three-layer convolution structure, and the step size of the first convolution structure of the three-layer convolution structure is set to 2, the step size of the second convolution structure is set to 1, and the step size of the third convolution structure is set to 2.
4. The method according to claim 2, characterized in that The fusion layer includes a fusion module, and the fusion module includes a global pooling layer, a fully connected layer, a nonlinear activation layer and a normalization layer, and an attention module; The global pooling layer, the fully connected layer, and the nonlinear activation layer are used to splice the feature map information corresponding to the RGB image features, the infrared thermal imaging features, and the sound spectrum map features in the depth direction. The attention module is used to multiply the original fused feature map and the channel attention coefficient element by element according to the channel to obtain the fused feature map.
5. The method according to claim 2, characterized in that The backbone network includes four residual modules, which implement back propagation by introducing residual connections and alleviate gradient vanishing and gradient exploding.
6. The method according to claim 2, characterized in that The output layer includes two cascaded fully connected layers, wherein the first fully connected layer is a feature conversion layer for capturing the complex relationship between different features, and the second fully connected layer is associated with the classification or regression task, and the number of its neurons is equal to the number of categories of the task.
7. The method according to claim 1, characterized in that The multimodal fault recognition model is trained based on a cross entropy loss function; the training data of the multimodal fault recognition model includes pre-labeled training data including insulation status labels, temperature status labels, and oil level status labels; The multimodal fault identification model deploys a model weight file and corresponding reasoning program code in a detection operation environment where the target transformer is located.
8. A transformer fault intelligent detection device based on deep learning multimodal fusion, characterized in that: include: an acquisition unit, configured to acquire RGB image data of a target transformer through an image sensor, acquire infrared thermal imaging image data of the target transformer through an infrared sensor, and acquire a sound signal of the target transformer through a sound sensor; preprocess the RGB image data, the infrared thermal imaging image data, and the sound signal, and use the preprocessed data as multimodal detection data; The detection unit is used to input the multimodal detection data into a pre-trained multimodal fault recognition model to obtain a fault detection result corresponding to the target transformer. The multimodal fault recognition model includes a perception layer, a fusion layer, a backbone network and an output layer cascaded in sequence.
9. A computer device, characterized in that: The computer device includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device executes the method according to any one of claims 1 to 7.
10. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 7.
Citation Information
Cited By
Lightweight RGB-T salient target detection method
CN120912871A
Power equipment on-line detection system based on lightweight model
CN120915005A
Automobile part quality online detection method based on multi-sensor fusion
CN121030684A
Inspection robot based on multi-modal data fusion analysis
CN121170741A
A patrol robot based on multi-modal data fusion analysis
CN121170741B