Typical power equipment fault identification method based on cross-source image fusion of YOLOv11

By combining cross-source image fusion with the Transformer architecture and deep learning, this method utilizes visible light, infrared, and ultraviolet images for power equipment fault detection. This solves the problems of insufficient detection efficiency and accuracy in traditional methods and achieves efficient and robust fault identification.

CN121170513APending Publication Date: 2025-12-19INNER MONGOLIA UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202511125472.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-12
Publication Date
2025-12-19

AI Technical Summary

Technical Problem

Traditional power equipment fault detection methods rely on single-source data, making it difficult to adapt to scenarios with diverse equipment types and complex fault modes. They are also susceptible to environmental interference, resulting in insufficient detection efficiency and accuracy, and failing to meet the demands of modern inspection for high efficiency and intelligence.

Method used

A cross-source image fusion method based on YOLOv11 is adopted, which combines visible light, infrared and ultraviolet images, performs feature extraction and fusion through the CSFT model, and utilizes the Transformer architecture to improve detection accuracy and robustness, thereby achieving efficient recognition and detection.

Benefits of technology

It improves the accuracy and robustness of power equipment fault detection, reduces the risk of power outages, adapts to complex backgrounds and weak targets, and supports the safe operation of smart grids.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121170513A_ABST
    Figure CN121170513A_ABST
Patent Text Reader

Abstract

The invention discloses a typical power equipment fault identification method based on cross-source image fusion of YOLOv11, which is used for power equipment fault detection. The specific principle is as follows: firstly, denoising, normalizing, enhancing and other preprocessing are performed on a multi-source image, and accurate registration is performed to eliminate spatial distribution difference; deep features of the registered image are extracted through a CSFT model, efficient fusion of global information among sources is realized by using a multi-head attention mechanism, and complementarity and correlation of multi-source features are enhanced in combination with source alignment loss and a position coding technology in fusion; and finally, the fused multi-source features are transmitted to a YOLOv11 network, and accurate positioning and classification of a fault target are realized by using an optimized feature pyramid structure and a decoupling detection head module of the YOLOv11 network. The method provides reference for combination of power equipment fault detection, cross-source data fusion and target detection, and supports safe operation of an intelligent power grid.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The application belongs to the technical field of image fusion and fault identification, and relates to cross-source image fusion, in particular to a typical power equipment fault identification method based on cross-source image fusion of YOLOv11. BACKGROUND

[0002] Traditional power equipment fault detection methods rely on single-source data, and in the scene of multiple types of equipment and complex fault modes, the detection efficiency and accuracy are insufficient, and are easily disturbed by the environment, which cannot meet the modern inspection efficiency and intelligentization needs. Cross-source image fusion combines multiple types of images from different sources, and extracts features autonomously through deep learning, which can adapt to complex fault scenes, improve detection reliability, accuracy and robustness, and reduce power outage risks.

[0003] However, the existing image fusion model for electric appliance fault identification mostly uses only visible light and infrared data as input, ignoring the fault types such as arc discharge that cannot be captured by visible light and infrared that power equipment may generate. In addition, traditional feature extraction and image fusion methods are usually performed separately, resulting in low model efficiency and difficulty in capturing complex nonlinear relationships. SUMMARY

[0004] In order to overcome the above-mentioned shortcomings of the prior art, the purpose of the present application is to provide a typical power equipment fault identification method based on cross-source image fusion of YOLOv11, which aims to use multi-source image data to efficiently identify and detect power equipment faults by combining the global model of the advanced CSFT model with YOLOv11.

[0005] In order to achieve the above-mentioned purpose, the technical solution adopted by the present application is:

[0006] A typical power equipment fault identification method based on cross-source image fusion of YOLOv11, comprising the following steps:

[0007] Step 1, acquiring visible light, infrared and ultraviolet images of the power equipment, which are respectively used to represent the appearance, temperature and discharge information of the power equipment;

[0008] Step 2, constructing a typical power equipment fault identification model based on cross-source image fusion of YOLOv11, the backbone network is composed of two parts: a three-branch feature extraction network and three cross-source Fusion Transformer modules, the three branches of the three-branch feature extraction network respectively perform convolution feature extraction on the visible light, infrared and ultraviolet images to obtain a visible light feature vector sequence I V , an infrared feature vector sequence I I and an ultraviolet feature vector sequence I U ;

[0009] The three cross-source Fusion Transformer modules connect the respective feature vector sequences and add a learnable positive position embedding To obtain the input sequence of the Transformer Then, the input sequence I is projected onto three weight matrices w Q , W K , W V to calculate a set of queries Q, keys K and values V;

[0010] Step 3, train the model of step 2 using the image data obtained in step 1, and based on the trained model, simultaneously use visible light, infrared and ultraviolet images for fault identification.

[0011] Compared with the prior art, the method of the present application is suitable for power system target detection, and through cross-source image fusion and combination of the Transformer architecture, the features of visible light and other multi-source images are deeply fused to improve the detection accuracy and robustness, and at the same time, the YOLOv11 optimized target detection network is used to realize efficient feature extraction and real-time detection, and the adaptability to complex background and weak targets is enhanced. BRIEF DESCRIPTION OF DRAWINGS

[0012] Figure 1 The main flowchart of the present application.

[0013] Figure 2 The model structure diagram of the present application. DETAILED DESCRIPTION

[0014] The embodiments of the present application will be described in detail below in conjunction with the drawings and examples.

[0015] The application is a typical power equipment fault identification method based on YOLOv11 cross-source image fusion, mainly aiming at fault detection of three types of power equipment such as transformers, combining deep learning to develop a cross-source image fusion fault detection model, analyzing equipment state and fault characteristics through cross-source image registration fusion and feature extraction technology, optimizing the diagnosis process, and realizing power equipment fault identification. Compared with the traditional identification method, the application combines Cross-Source Fusion Transformer (CSFT) and YOLOv11, which can fully exploit the characteristics of multi-source image data to adapt to fault identification in complex backgrounds. The specific principle is as follows: first, the multi-source images are preprocessed such as denoising, normalization, enhancement, and accurately registered to eliminate spatial distribution differences; then the deep features of the registered images are extracted through the CSFT model, and the multi-head attention mechanism is used to realize efficient fusion of global information between sources, and the source alignment loss and position encoding technology are combined in the fusion to enhance the complementarity and correlation of multi-source features; finally, the fused multi-source features are transmitted to the YOLOv11 network, and the feature pyramid structure and decoupling detection head module are used to realize accurate positioning and classification of fault targets. This method provides a reference for power equipment fault detection and the combination of cross-source data fusion and target detection, supporting the safe operation of smart grids.

[0016] Specifically, as shown in Figure 1 The method of the application comprises the following main steps:

[0017] Step 1, obtaining multi-source image data of power equipment and fault identification preparation.

[0018] The application needs to collect multi-source image data related to typical equipment of the power system, which includes visible light, infrared and ultraviolet images of power equipment.

[0019] The sample collection method is as follows: a partition collection method is used to design sampling points in the distribution area of power equipment, and visible light images of typical power equipment such as current transformers, voltage transformers, circuit breakers and disconnectors under different operating conditions are collected, covering various working conditions, environmental conditions and fault types (such as equipment overheating, surface damage, corona discharge, etc.), to ensure the diversity characteristics of the samples.

[0020] In this embodiment, the power equipment image data is collected from Xiangshawan 500 kV substation in Inner Mongolia Autonomous Region. In the station, a high-resolution visible light camera, an infrared thermal imager and an ultraviolet imager are deployed to collect the appearance, temperature and discharge information of the power equipment, respectively. During the shooting process, the visible light camera is used to obtain clear visible light images of the equipment appearance to present the equipment structure and state, then the infrared thermal imager is used to shoot infrared images at the same position to record the temperature distribution of the equipment, finally, the ultraviolet imager is used to capture the ultraviolet images generated by the partial discharge of the equipment at the same angle and distance. By keeping the consistency of the shooting angle and distance, the accuracy and spatial correspondence of the three types of image data are ensured, which provides reliable data support for subsequent cross-source image fusion analysis.

[0021] Specifically, the embodiment obtains multi-source images of the power equipment through the equipment acquisition hardware system (taking DJI JINGLI M300 RTK unmanned aerial vehicle as the core flight platform, carrying Zenmuse H20T multi-sensor load system, the three-axis gimbal stabilizing mechanism (vibration compensation accuracy ± 0.01°) carried thereby ensures that clear images can still be obtained at a speed of 15 m / s, and the laser range finder marking function (accuracy ± 0.2 m) realizes the spatial coordinate calibration of the defective components, the double-remote controller mode can display three-way sensor pictures at the same time, and supports visible light and infrared split-screen comparison analysis), and selects the part without occlusion and interference for subsequent processing. Among them, the pixel size of the visible light image is 640×480, which can clearly present the appearance structure and surface state of the equipment; the pixel size of the infrared image is 640×480, which can effectively capture the temperature distribution information of the equipment; the pixel size of the ultraviolet image is 720×576, which can accurately record the ultraviolet characteristics such as corona discharge of the equipment.

[0022] Before cross-source image fusion, the three types of images need to be preprocessed and registered to improve the data quality and lay a foundation for subsequent multi-source image fusion.

[0023] For visible light images, mainly Gaussian denoising, illumination normalization, color space conversion are carried out. Among them, illumination normalization is to weaken or eliminate the brightness difference in the image caused by uneven illumination, so that the visible light images under different regions and different acquisition conditions are more consistent in overall brightness and contrast. At the same time, since its noise is mainly expressed as salt and pepper noise, such noise is usually caused by impulse interference in the image sensor or burst error in the transmission process, and presents as random black and white points. Therefore, the color space conversion adopts the median filtering method, which can effectively suppress salt and pepper noise while maximizing the retention of image edge details and texture features by replacing the value of each pixel in the image with the median value of its neighborhood pixels. The advantage of this method is that the isolated noise points are effectively filtered out, and the image edge is not blurred like mean filtering, which provides a clear image basis for subsequent feature extraction of device appearance structure and surface state.

[0024] For infrared images, generally wavelet denoising, removing thermal noise, temperature normalization, target enhancement are needed. And since it is easily affected by temperature fluctuations, sensor thermal noise and electronic component interference, etc. to produce Gaussian noise, its noise distribution conforms to the normal distribution characteristics, which shows the smoothness interference of the whole image, so Gaussian filtering can be used, which is weighted average of neighborhood pixels according to Gaussian distribution weight, smooth the image and suppress noise, and the neighborhood pixels closer to the target pixel are given higher weight. This processing method can effectively smooth the mutation of image gray value, suppress the gray fluctuation caused by Gaussian noise, and better preserve the overall temperature distribution trend of the image, providing reliable temperature field information for accurate identification of device thermal abnormal areas. Among them, wavelet or non-local mean (NL-Means) denoising is commonly used for infrared imaging; removing thermal noise generally refers to removing hot spots and fixed pattern noise (FPN); temperature normalization (linear stretching or pseudo-color mapping) can unify the temperature range of different sessions; target enhancement such as contrast limited adaptive histogram equalization (CLAHE) or local gradient enhancement.

[0025] For the ultraviolet image, generally, median filtering, white balance, geometric correction, texture enhancement are needed. And due to the interference of environmental light (such as ultraviolet component in sunlight), atmospheric scattering and sensor dark current, the noise presents complex mixed characteristics, and is often accompanied by high-frequency component interference. Therefore, the wavelet denoising method can be used. The wavelet denoising method is based on the theory of multi-scale analysis. The image is decomposed into different frequency subbands through wavelet transform. The noise is mainly concentrated in the high-frequency subband. By setting a reasonable threshold value for threshold processing (such as hard threshold or soft threshold method) of the high-frequency subband, the high-frequency component where the noise is located is selectively removed. The effective information of the image reflecting the effective features such as corona discharge is retained after reconstruction, and the image quality is improved. The ultraviolet image obtained after wavelet reconstruction can completely retain the spatial distribution and intensity characteristics of the corona discharge while suppressing noise, laying a foundation for subsequent accurate detection of discharge phenomena. Among them, white balance generally refers to correction of atmospheric scattering; geometric correction (such as distortion correction) is related to projection relationship; texture enhancement (such as Laplacian sharpening or local binary pattern LBP) can highlight the discharge traces under ultraviolet.

[0026] Further, the multi-source images also need to be image registration processed to correct the spatial deviation caused by differences in device type, shooting angle and lighting condition, so as to ensure that all images have consistent spatial positioning under the same coordinate system. This is of great significance for fault detection and monitoring of power equipment. In the embodiment of the present application, the visible light, infrared and ultraviolet images are implicitly implemented by an end-to-end neural network learning method, as follows:

[0027] The visible light, infrared and ultraviolet images are normalized to the [0, 1] interval by ToTensor transformation to ensure data distribution consistency, laying a foundation for model learning registration features;

[0028] After the AGNet model receives the images, it extracts deep features of the images through a multi-layer neural network, constructs a spatial transformation mapping relationship from the input image to the target image, and outputs the reconstructed results of the visible light, infrared and ultraviolet images after registration. The essence is to align the reconstructed image and the original image in spatial position by learning.

[0029] Further, the difference between the input image and the original image is quantified by the MSE loss, and the model parameters are updated by back propagation. The minimization process of the loss function forces the model to learn the transformation rule (such as translation, rotation, scaling and other composite transformations) that can eliminate the spatial deviation of the input image, and finally realizes the implicit registration of the two images. Through multiple rounds of epoch training, the model continuously adjusts the parameters of feature extraction and transformation mapping, so that the spatial alignment error between the reconstructed image and the original image gradually decreases. When the loss tends to be stable, the model has the registration capability for similar images, thereby achieving the core goal of eliminating the spatial deviation of the two images and keeping the corresponding targets consistent in spatial position.

[0030] The process is based on feature points for image registration, by extracting the salient features in the image, the precise alignment of multi-source images is realized. This process ensures the spatial consistency of multi-source images, providing a reliable premise for subsequent fusion and detection.

[0031] The formula involved in this step registration process is as follows:

[0032] The mean square error (MSE) is used as the loss calculation criterion to quantify the deviation between the model output and the input data, and its mathematical expression is: Where y is the input data, is the model reconstruction output, and N is the total number of sample elements.

[0033] loss1 is the weighted sum of the MSE loss of the two input channels, and the formula is as follows:

[0034]

[0035] Optimization algorithm and parameter update: The Adam algorithm is used for model parameter optimization, which combines first-order momentum (momentum term) and second-order momentum (adaptive learning rate) to realize parameter update, and the iteration formula is:

[0036] Where θ t is the model parameter at time t, η is the initial learning rate (set to 10 -4 ), and are the bias correction values of the first and second moments, respectively, where is the momentum decay coefficient (default 0.9) and is the momentum decay coefficient (0.999), ∈ (default 10 -8 ) is a small value to prevent the denominator from being zero, where γ is the decay factor (0.9) and s is the decay step (30 epochs), which reduces the learning rate to avoid model oscillation in the later stage of convergence.

[0037] Data transformation and normalization: The image pixel value is normalized from [0, 255] to [0, 1], and the mathematical expression is:

[0038]

[0039] where x is the original pixel value, and the normalization operation can eliminate the data scale difference, speed up the model convergence and stabilize the training process. In addition, by fixing the random seed (seed) to control the random number generation, the experimental repeatability is ensured, involving the sampling principles of uniform distribution and normal distribution of pseudo-random numbers.

[0040] Forward and back propagation of neural network: the forward propagation of AGNet model is based on the composite operation of linear transformation and nonlinear activation function, and the output of a certain layer network can be expressed as:

[0041] h = f(Wx + b), wherein x is an input feature vector, W and b are weight matrix and bias vector respectively, and f(...) is a nonlinear activation function. The model outputs (p_x, p_y and p_z and x_hat, y_hat and z_hat) are all realized by such multi-layer operation to realize feature extraction and mapping.

[0042] The back propagation process is based on the chain rule to calculate the gradient of the loss function to the parameters:

[0043] Wherein z = Wx + b, the gradient is updated to the parameters of each layer through back propagation to realize the minimization of the loss function.

[0044] Step 2, constructing a typical power equipment fault identification model based on YOLOv11 cross-source image fusion.

[0045] The model of the application captures the structural features of the visible light image, the temperature features of the infrared image and the discharge features of the ultraviolet image through the feature extraction module, and then realizes the effective fusion of multi-dimensional features through the fusion module, and finally realizes the positioning and classification of the typical power equipment fault through the detection head. The model makes full use of the complementary information of multi-source images, improves the accuracy and robustness of fault identification, and provides a reliable basis for the state evaluation of power equipment.

[0046] The main network of the model is composed of two parts: a three-branch feature extraction network and three cross-source FusionTransformer modules.

[0047] Among them, the three branches of the three-branch feature extraction network respectively perform convolution feature extraction on the visible light, infrared and ultraviolet images to obtain the visible light feature vector sequence the infrared feature vector sequence and the ultraviolet feature vector sequence

[0048] Reference Figure 2 , specifically, the three branches respectively take the visible light, infrared and ultraviolet images as input, and obtain the feature map of the visible light image the feature map of the infrared image and the feature map of the ultraviolet image Subsequently, F V is flattened and the matrix permutation sequence is obtained I V , F I is flattened and the matrix permutation sequence is obtained I I , F UI is obtained by flattening and permuting the matrix sequence U Here, the flattening and permuting the matrix sequence refers to exchanging the dimension order of the matrix obtained after flattening (i.e., the transpose or permute operation) to match the input format required by the Transformer.

[0049] The three cross-source Fusion Transformer modules connect the respective feature vector sequences and add a learnable positive position embedding to obtain the input sentence of the Transformer as follows:

[0050]

[0051] where I cat is the connection result of each feature vector sequence.

[0052] Then, the input sequence I is projected onto three weight matrices W Q , W K , W V to calculate a set of queries Q, keys K, and values V, represented as:

[0053] Q = IW Q , K = IW K , V = IW V

[0054] where, is the weight matrix, D Q = D K = D V = C, C is the length of each feature sequence, D Q , D K , D V are the dimensions of the Q, key K, and value V vectors, respectively.

[0055] Finally, the self-attention layer of the model uses the scaled dot product between the matrix query Q and the key K to calculate the attention weights, and then multiplies the value V to infer the refined output Z, represented as:

[0056]

[0057] where, is a scaling factor used to prevent the softmax function from falling into a region with minimal gradient when the size of the dot product becomes large.

[0058] Further, in order to encapsulate multiple complex relationships from different representation subspaces at different positions, the present application adopts a multi-head attention mechanism, represented by the formula:

[0059] Z' = Concat (Z1,..., Zh W O

[0060] Where Z′ is the intermediate result obtained by concatenating the outputs along the channel dimension after performing attention calculations on all h heads. This means Concat(Z1,…,Z) h The projection matrix of ).

[0061] Subsequently, the Transformer works using a two-layer fully connected feedforward network (MLP) with a GELU activation in between to compute the output sequence O, represented as:

[0062] O = MLP(Z″) + Z″

[0063] Where Z″=Z′+I.

[0064] Finally, the output sequence O is converted into the recalibration result F′. V F′ I and F′ U As supplementary information, it is added to the original feature extraction branch, specifically F′ V For F V The feature map F″ of the visible light image after information supplementation is obtained. V F′ I For F I The feature map F″ of the infrared image after information supplementation is obtained. I F′ U For F U The feature map F″ of the infrared image after information supplementation is obtained. U F″ V 、F″ I and F″ U After splicing, the final fused multi-source feature F is obtained. S , will F S The data is then passed to the YOLOv11 network, where its feature pyramid structure and decoupled detection head module are used to achieve accurate localization and classification of faulty targets.

[0065] Step 3, Model Training and Evaluation.

[0066] The model from step 2 is trained using the image data obtained in step 1, focusing on three core objectives: detection accuracy, localization accuracy, and classification reliability. Improvements are made using bounding box loss, focal loss, and IOU-related loss functions.

[0067] The expression for the bounding box loss is as follows:

[0068]

[0069] where IoU is the intersection over union, ρ is the distance between the center of the predicted box and the ground truth, c is the diagonal length of the minimum rectangle containing both boxes, v measures the consistency of the width-height ratio, α is the weight coefficient, and the new parameter is used to balance the positioning accuracy of targets of different scales.

[0070] Dynamic scale weighting mechanism, for small target detection difficulties, introduce adaptive scale weight w s , the expression is:

[0071]

[0072] where area gt is the area of the real box, S is the scale of the feature map, and small targets correspond to higher weights.

[0073] Target confidence loss (Object Confidence Loss), the expression is:

[0074] L conf =-(1 obj ·(1-p) γ logp+1 noobj ·p γ log(1-p))

[0075] where 1 obj and 1 noobj indicate the indicator function of the existence and non-existence of the target, respectively; γ is the focusing parameter, and the new dynamic adjustment factor changes adaptively with the training stage.

[0076] Joint IoU confidence prediction, the expression is:

[0077] target conf =IoU+λ·GIoU

[0078] where λ is the weight coefficient, balancing the influence of IoU and generalized GIoU.

[0079] Class classification loss (Class Classification Loss), the expression is:

[0080]

[0081] where, is the smoothed label C is the number of classes, and ∈ is the smoothing coefficient, usually taking 0.1.

[0082] Long-tailed distribution adaptive loss, to solve the problem of data imbalance, introduce class frequency weight w c, the expression is:

[0083]

[0084] Wherein, N is the total number of samples, n c The number of samples of the category C, the rare category corresponds to a higher weight.

[0085] Cross-layer feature consistency loss forces the prediction results of different scale feature maps to keep consistency, and improves the stability of multi-scale detection, and the expression is:

[0086]

[0087] Where f m Is the output of the mth layer feature map.

[0088] Attention mechanism regularization, the expression is:

[0089]

[0090] Where A a Is the attention weight matrix, λ is the regularization strength, and the redundant attention allocation is inhibited.

[0091] The total loss function can be represented as:

[0092] L total =λ box L box +λ conf L conf +λ class L class +λ feat L feat +λ attn L attn

[0093] Performance evaluation: when evaluating the performance of the improved network model, the selected indicators include precision (P), recall (R) and mean average precision (mAP).

[0094] The accuracy P measures the proportion of samples correctly predicted as positive by the model among all samples predicted as positive, and the recall R reflects the proportion of positive samples correctly predicted by the model among all actual positive samples, and the mAP evaluates the average accuracy of the model at different recall levels by calculating the area under the precision-recall (P-R) curve, which is an important parameter for measuring the overall detection ability of the model. Through the three evaluation indexes, the size, accuracy and comprehensiveness of the model are comprehensively judged. The formulas of these indexes are as follows:

[0095] Wherein TP is the number of correct detections; FP is the number of false detections.

[0096] Wherein FN is the number of missed detections.

[0097] Wherein AP is the average detection accuracy of each type of fault, and mAP50 represents the total class average accuracy when IoU is 0.5.

[0098] Step 4, image fusion and target detection

[0099] Based on the trained model, visible light, infrared and ultraviolet images are used for fault identification.

[0100] The present application can integrate the appearance structure features of visible light images, temperature distribution information reflected by infrared images and corona discharge and other features captured by ultraviolet images, deeply analyze the fault evolution law of typical power equipment (current transformers, voltage transformers, circuit breakers, disconnectors, etc.) under different working conditions and environmental conditions, and explore the causes and potential impact mechanisms of equipment faults. Through the identification result, the operation and maintenance management strategy of the power system can be optimized, such as targeted device inspection plan, fault warning and repair resource allocation, etc., to improve the stability and reliability of the power system operation. Combined with the actual operation status of the power network, a scientific device full life cycle management scheme is constructed to ensure the safe and efficient operation of the power system, and decision support is provided for the intelligent operation and risk prevention and control of the power industry.

Claims

1. A typical power equipment fault identification method based on YOLOv11 cross-source image fusion, characterized in that, The method comprises the following steps: Step 1, acquiring visible light, infrared and ultraviolet images of the power equipment, which are used to represent the appearance, temperature and discharge information of the power equipment respectively; Step 2, a typical power equipment fault identification model based on YOLOv11 cross-source image fusion is constructed, and the backbone network is composed of two parts: a three-branch feature extraction network and three cross-source Fusion Transformer modules, the three branches of the three-branch feature extraction network respectively perform convolution feature extraction on visible light, infrared and ultraviolet images to obtain visible light feature vector sequence I V , infrared feature vector sequence I I and ultraviolet feature vector sequence I U ; The three cross-source Fusion Transformer modules connect the respective feature vector sequences and add a learnable positive position embedding E POS to obtain an input sequence I of the Transformer, and then a set of queries Q, keys K and values V are calculated using the input sequence I; Step 3, training the model of step 2 by using the image data acquired in step 1, and based on the trained model, using the visible light, infrared and ultraviolet images to identify faults.

2. The method according to claim 1, wherein the method is based on YOLOv11 cross-source image fusion of typical power equipment fault recognition. The step 1, acquiring visible light, infrared and ultraviolet images of the power equipment, is implemented as follows: First, a visible light camera is used to acquire a visible light image of the appearance of the equipment to present the structure and state of the equipment; Then, an infrared thermal imager is used to shoot an infrared image at the same position to record the temperature distribution of the equipment; Finally, an ultraviolet imager is used to capture an ultraviolet image generated by partial discharge of the equipment at the same angle and distance.

3. The method according to claim 1, wherein the method is based on YOLOv11-based cross-source image fusion of typical power equipment fault recognition. For the visible light image, Gaussian denoising is performed, and light normalization is used to weaken or eliminate the brightness difference in the image caused by uneven light, so that the visible light images under different conditions are more consistent in overall brightness and contrast; and a median filter method is used to replace the value of each pixel in the image with the median value of its neighborhood pixels to remove salt and pepper noise, retain image edge details, and realize color space conversion; For the infrared image, Gaussian filtering is used to weight and average the neighborhood pixels according to the Gaussian distribution weight to smooth the image and suppress noise, wherein the neighborhood pixels closer to the target pixel are given higher weight; And wavelet denoising, thermal noise removal, temperature normalization and target enhancement are performed; For the ultraviolet image, a wavelet denoising method is used to decompose the image into different frequency subbands by wavelet transform, remove the high-frequency components where the noise is located through threshold processing, retain the effective information of the image after reconstruction, and perform median filtering, white balance, geometric correction and texture enhancement.

4. The method according to claim 1, wherein the method is based on YOLOv11-based cross-source image fusion of typical power equipment fault recognition. For the visible light, infrared and ultraviolet images, image registration is implicitly realized through an end-to-end neural network learning method as follows: The visible light, infrared and ultraviolet images are normalized to the [0, 1] interval through ToTensor transformation to ensure data distribution consistency; The AGNet model extracts deep features of the images through multiple layers of neural networks, constructs a spatial transformation mapping relationship from the input image to the target image, and outputs the reconstructed results of the registered visible light, infrared and ultraviolet images, and the reconstructed images are aligned with the original images in spatial position.

5. The method according to claim 1, wherein the method is based on YOLOv11-based cross-source image fusion of typical power equipment fault recognition. The three-branch feature extraction network extracts feature maps F of the visible light image through convolution V , feature maps F of the infrared image I , and feature maps F of the ultraviolet image U , respectively, flattens and sequentially permutes the obtained F V , F I , and F U to obtain the visible light feature vector sequence I V , the infrared feature vector sequence I I , and the ultraviolet feature vector sequence I U .

6. The method according to claim 1, wherein the method is based on YOLOv11-based cross-source image fusion of typical power equipment fault recognition. The calculation formula of the input sequence I is as follows: where I cat is the concatenation result of each feature vector sequence.

7. The method according to claim 1, wherein the method is based on YOLOv11-based cross-source image fusion of typical power equipment fault recognition. projecting the input sequence I onto three weight matrices W Q , W K , W V to compute a set of queries Q, keys K, and values V, denoted as: Q = IW Q K = IW K V = IW V wherein, is a weight matrix, D Q = D K = D V = C, C is the length of each feature sequence, D Q , D K , D V are the dimensions of Q, the key K, and the value V vectors, respectively.

8. The method according to claim 1, wherein the method is based on YOLOv11-based cross-source image fusion of typical power equipment fault recognition. The self-attention layer of the typical power equipment fault identification model uses scaled dot-product between query Q and key K to calculate attention weight, and then multiplies value V to deduce refined output Z, wherein, is a scaling factor to prevent the softmax function from falling into a region with minimal gradient when the size of the dot product becomes large.

9. The method according to claim 8, wherein the method is based on YOLOv11 cross-source image fusion of typical power equipment fault recognition. Multiple heads of attention mechanism are adopted to encapsulate multiple complex relationships of different representation subspaces from different positions, which is formulated as Z' = Concat(Z1,...,Z h )W O , where Z' is the intermediate result of concatenating the outputs by channel dimension after performing attention computation of all h heads, , where W represents the projection matrix of Concat(Z1,...,Z h ).

10. The method according to claim 9, wherein the method is based on YOLOv11 cross-source image fusion of typical power equipment fault recognition. The Transformer uses a two-layer fully connected feedforward network MLP to calculate an output sequence O, O = MLP(Z") + Z", where Z" = Z' + I, and finally converts the output sequence O into a recalibration result F' V , F' I , and F' U , as supplementary information added to the original feature extraction branch, to obtain the feature map F" of the visible light image after adding the supplementary information V , the feature map F" of the infrared image I , and the feature map F" of the ultraviolet image U , and finally obtain the fused multi-source feature F S , and pass F S to the YOLOv11 network to realize accurate positioning and classification of the fault target by using the feature pyramid structure and decoupled detection head module thereof.

Citation Information

Cited By

  • Rotary mechanical part collaborative diagnosis method considering new fault mode

    CN117972523A

  • Partial discharge identification system algorithm based on ultraviolet sensor array

    CN121500044A