A power equipment diagnostic method based on three-light fusion

The power equipment diagnostic method based on the fusion of three light sources (visible light, infrared, and ultraviolet) combines multi-frequency domain collaborative feature decomposition with physically constrained image reconstruction, which solves the problem of blind spots in single-modal information and improves the accuracy and interpretability of power equipment defect detection.

CN120876458BActive Publication Date: 2025-12-02EAST CHINA JIAOTONG UNIVERSITY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511366543.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-09-24
Publication Date
2025-12-02
Estimated Expiration
2045-09-24

AI Technical Summary

Technical Problem

Existing methods for detecting defects in power equipment rely on information obtained from a single spectrum, which has a blind spot for single-mode information, resulting in low accuracy in defect detection under complex environments.

Method used

The original images are acquired using a visible light camera, an infrared thermal imager, and an ultraviolet camera. Multi-physics coupling enhancement features are extracted through a dual-stream deep feature extraction network. After multi-scale high-resolution texture features are combined and multi-modal registration is performed, multi-frequency domain collaborative feature decomposition and multi-modal image reconstruction under physical constraints are used for fusion. Finally, the fusion is performed in an improved convolutional neural network to output the defect detection results.

Benefits of technology

It achieves full-spectrum coverage of the physical field of defects, improves defect detection accuracy, overcomes high and low frequency interference in complex scenarios, and significantly improves detection accuracy and interpretability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120876458B_ABST
    Figure CN120876458B_ABST
Patent Text Reader

Abstract

This invention provides a power equipment diagnostic method based on three-spectrum fusion, comprising: acquiring raw visible light images, raw infrared images, and raw ultraviolet images of the power equipment; preprocessing the images to obtain preprocessed images; inputting the preprocessed images into a dual-stream deep feature extraction network to obtain multimodal registered images; obtaining low-frequency components containing shared attribute feature maps and high-frequency components containing unique attribute feature maps through a multimodal image reconstruction and fusion module, thereby obtaining a final three-spectrum fusion feature map; and inputting the final three-spectrum fusion feature map into a convolutional neural network containing an improved softmax operation module to obtain the power equipment defect detection result. This invention can solve the problem that existing technologies rely on information acquired from a single spectrum, resulting in blind spots in single-modal information and low defect detection accuracy in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of power equipment testing, and more specifically to a power equipment diagnostic method based on three-light fusion. Background Technology

[0002] The large-scale installation of power equipment, while meeting societal electricity demands, also poses serious challenges to the reliability and safe operation of the power grid. Some power equipment is installed in complex geographical environments or frequently exposed to extreme weather conditions. These factors can easily accelerate the aging, wear, and corrosion of power equipment, becoming potential points of failure and threatening the stable operation of the power grid. Therefore, implementing regular and efficient defect detection and maintenance of power equipment, accurately grasping its operating status, and preventing accidents are crucial links in ensuring the reliability of power transmission.

[0003] Defects in power equipment are one of the main causes of high-voltage transmission line faults, typically manifesting as early signs of degradation in electrical or mechanical components. These defects, during their evolution, are accompanied by diverse physical phenomena, such as abnormal heating, localized corona discharge or arcing, and abnormal changes in the equipment's surface structure. However, these phenomena belong to different electromagnetic spectra and are often concealed when occurring individually, requiring specific detection techniques for discovery. Existing power equipment defect detection methods often rely on information acquired from a single spectrum, resulting in blind spots in single-mode information. For example, while visible light images can clearly show external structural damage, they struggle to detect internal abnormal heating or early, weak discharges; infrared thermal imaging excels at capturing temperature anomalies but lacks sufficient resolution for minute dirt or mechanical damage on the equipment surface; and while ultraviolet imaging is sensitive to corona discharge, it lacks detailed support from the equipment's background structure. These issues ultimately lead to low accuracy in defect detection under complex environments. Summary of the Invention

[0004] In view of this, the present invention provides a power equipment diagnostic method based on three-spectrum fusion to solve the problem that the existing technology relies on information obtained from a single spectrum, which has a blind spot of single modal information and results in low defect detection accuracy in complex environments.

[0005] A power equipment diagnostic method based on three-light fusion includes:

[0006] Step S1: Use a visible light camera, an infrared thermal imager, and an ultraviolet camera to acquire images of the power equipment, obtaining the original visible light image, the original infrared image, and the original ultraviolet image of the power equipment.

[0007] Step S2: Preprocess the original visible light image, the original infrared image, and the original ultraviolet image to obtain the preprocessed visible light image, the processed infrared image, and the processed ultraviolet image.

[0008] Step S3: Input the preprocessed visible light image, the processed infrared image, and the processed ultraviolet image into the dual-stream depth feature extraction network. First, based on the multi-scale discharge spot features and the multi-scale hot spot features, obtain the multi-physics field coupling enhancement features. Then, combine the multi-scale high-resolution texture features to obtain the multi-modal registered image.

[0009] Step S4: Based on the image after multimodal registration, the image is fused by multi-frequency domain collaborative feature decomposition and physical constraint multimodal image reconstruction fusion module to obtain low-frequency components containing common attribute feature maps and high-frequency components containing unique attribute feature maps. The high-frequency components and low-frequency components are fused to obtain the final three-light fusion feature map.

[0010] Step S5: Input the final three-light fusion feature map into a convolutional neural network containing an improved softmax operation module to obtain the power equipment defect detection result.

[0011] The power equipment diagnostic method based on three-light fusion provided by the present invention has the following beneficial effects:

[0012] (1) This invention inputs the preprocessed visible light image, infrared image and ultraviolet image into the dual-flow depth feature extraction network. First, based on the multi-scale discharge spot features and multi-scale hot spot features, multi-physics field coupling enhancement features are obtained. Then, combined with the multi-scale high-resolution texture features, multi-modal registered images are obtained, realizing full spectrum coverage of the defect physical field, solving the problem of blind spots in single-modal information, and improving the defect detection accuracy.

[0013] (2) This invention fuses multi-frequency domain collaborative feature decomposition and physical constraint multimodal image reconstruction fusion module to obtain low-frequency components containing common attribute feature maps and high-frequency components containing unique attribute feature maps. The common attribute feature maps represent cross-modal coupling steady-state features, and the unique attribute feature maps represent modality-specific transient features. By fusing the high-frequency components and low-frequency components, the final three-light fusion feature map is obtained, which breaks through the feature space heterogeneity and strengthens the feature correlation of different feature maps.

[0014] (3) Inputting the final three-light fusion feature map into a convolutional neural network containing an improved softmax operation module can overcome high and low frequency interference in complex scenes and significantly improve the accuracy and interpretability of defect detection in complex environments. Attached Figure Description

[0015] Figure 1 This is a flowchart illustrating the power equipment diagnostic method based on three-light fusion provided in an embodiment of the present invention. Detailed Implementation

[0016] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain embodiments of the present invention, and should not be construed as limiting the present invention.

[0017] Please see Figure 1 The embodiments of the present invention provide a power equipment diagnostic method based on three-light fusion, including steps S1 to S5:

[0018] Step S1: Use a visible light camera, an infrared thermal imager, and an ultraviolet camera to acquire images of the power equipment, obtaining the original visible light image, the original infrared image, and the original ultraviolet image of the power equipment.

[0019] The visible light camera, infrared thermal imager, and ultraviolet camera are all equipped with positioning modules, and the original infrared and original ultraviolet images are marked with the acquisition area and acquisition time.

[0020] Step S2 involves preprocessing the original visible light image, the original infrared image, and the original ultraviolet image to obtain the preprocessed visible light image, the processed infrared image, and the processed ultraviolet image.

[0021] Specifically, the original visible light image, original infrared image, and original ultraviolet image are cropped respectively, and the size of the original visible light image, original infrared image, and original ultraviolet image is adjusted to 256×64×64. Invalid edge areas are removed to reduce interference from irrelevant factors, thereby obtaining the preprocessed visible light image, processed infrared image, and processed ultraviolet image.

[0022] Step S3: Input the preprocessed visible light image, the processed infrared image, and the processed ultraviolet image into the dual-stream depth feature extraction network. First, based on the multi-scale discharge spot features and the multi-scale hot spot features, multi-physics field coupling enhancement features are obtained. Then, combined with the multi-scale high-resolution texture features, the multi-modal registered image is obtained.

[0023] The dual-stream deep feature extraction network includes visible light branch, infrared branch and ultraviolet branch, association coding module and multimodal registration module.

[0024] The visible light branch is used to learn the semantic features of the preprocessed visible light image in order to obtain high-resolution texture features at multiple scales.

[0025] The infrared branch is used to learn the semantic features of the preprocessed infrared image to obtain multi-scale discharge spot features.

[0026] The ultraviolet branch is used to learn the semantic features of the preprocessed ultraviolet image to obtain multi-scale hotspot features.

[0027] The correlation coding module is used to correlate the discharge spot features and hot spot features based on the relationship between temperature and discharge intensity to obtain multi-physics field coupling enhancement features.

[0028] The multimodal registration module is used to obtain the multimodal registered image based on multi-physics field coupling enhancement features and multi-scale high-resolution texture features through cross-attention weight allocation.

[0029] In this embodiment, the visible light branch deploys a VIT-Hybrid network to extract component-level RIO features and captures the global topological relationships of various components (such as suspension devices, positioning clamps, etc.) in the power equipment through a multi-head attention mechanism, and extracts local texture features in conjunction with the backbone network of the convolutional neural network.

[0030] The infrared branch design temperature calibration residual network introduces an adaptive temperature compensation layer to eliminate the influence of ambient temperature fluctuations, and employs a hot spot feature enhancement unit to increase the channel attention weights in abnormal temperature rise regions.

[0031] The ultraviolet branch employs a deformable convolutional discharge intensity quantization network, utilizing deformable convolutional kernels to capture the gradient characteristics of irregular discharge spots.

[0032] Specifically, the associated encoding module satisfies the following formula:

[0033]

[0034]

[0035] in, This is a feature that enhances multiphysics coupling. For ELU activation function, , For trainable weight matrix, For normalization function, To enhance the hot spot characteristics, To quantify the discharge intensity based on the characteristics of the discharge spot, To enhance the characteristics of the hot spot before enhancement, It is a non-linear activation function. For 1×1 convolution, Indicates global average pooling. For convolution operations, This is element-wise addition.

[0036] This invention replaces traditional QK queries with physically guided feature interactions, enabling infrared thermal distribution to guide visible light detail enhancement. Specifically, the multimodal registration module satisfies the following formula:

[0037]

[0038] in, This represents the image after multimodal registration. For normalization function, For high-resolution texture features, For querying the matrix, The key matrix, Indicates transpose. For feature dimensions.

[0039] In this embodiment, an improved infoNCE loss is used to optimize the dual-stream deep feature extraction network. Specifically, the expression for the loss function of the dual-stream deep feature extraction network is as follows:

[0040]

[0041] in, This represents the topology-aware contrast loss value. For positive sample pairs, For negative sample pairs, For the sample set, For natural index, The cosine similarity function is used. For the sample The corresponding component feature vector, For the sample The corresponding component feature vector, For the sample The corresponding component feature vector; This is the temperature coefficient used to control sample discrimination; the default value is 0.07.

[0042] Step S4: Based on the image after multimodal registration, the image is fused by multi-frequency domain collaborative feature decomposition and physical constraint multimodal image reconstruction fusion module to obtain low-frequency components containing common attribute feature maps and high-frequency components containing unique attribute feature maps. The high-frequency components and low-frequency components are fused to obtain the final three-light fusion feature map.

[0043] Specifically, step S4 includes:

[0044] Step S41: The multimodal image reconstruction and fusion module decomposes low-frequency coefficients and high-frequency coefficients based on the multimodal registered image;

[0045] Step S42: Based on the low-frequency coefficients, the multi-physics field equations are embedded as soft constraints into the multimodal image reconstruction and fusion module. The residual loss function is used to make the fusion features conform to physical laws. The low-frequency coefficients are fused and inversely transformed to obtain the low-frequency components containing the shared attribute feature map.

[0046] Among them, the frequency component mainly carries the overall contrast and temperature distribution information of the image. Multi-physics equations, such as the Joule thermal equation and photon energy relations, are embedded as soft constraints into the multimodal image reconstruction and fusion module. The residual loss function ensures that the fused features conform to physical laws, and the low-frequency coefficients are fused and inversely transformed. The common attribute feature map is obtained from the decomposition of low-frequency coefficients and represents the cross-modal coupling steady-state features.

[0047] Step S43: Based on the high-frequency coefficients, a differentiable physics engine is developed. By combining multi-scale high-resolution texture features, discharge spot features and hot spot features, the high-frequency coefficients are fused and inversely transformed to obtain high-frequency components containing unique attribute feature maps.

[0048] Among them, the high-frequency components focus on transient features such as discharge pulses and surface cracks, which can significantly improve the spatial consistency of the fused image and avoid geometric deformation caused by multimodal superposition. It can capture the abnormal temperature rise contours of infrared thermal images and remove artifacts by leveraging the high resolution of visible light, thus achieving multimodal image reconstruction. A differentiable physics engine is developed to jointly optimize image features and physical parameters, such as crack depth and material conductivity, to achieve bidirectional driving of mechanism and data. Using a joint loss function, combined with physically compliant features such as hot spot features, discharge spot features, and high-resolution texture features obtained from feature extraction, the high-frequency components are inversely transformed. The unique attribute feature map is decomposed from the high-frequency coefficients to represent mode-specific transient features.

[0049] Specifically, the following equation is satisfied during the fusion and inverse transformation of high-frequency coefficients:

[0050]

[0051] in, Indicates the fusion of the first High-frequency components of the scale, Indicates the first Scale edge mask, This indicates taking the maximum value. Indicates the first The characteristics of the discharge spot at a certain scale. Indicates the first Scale-based hot spot characteristics, Indicates the first High-resolution texture features at scale.

[0052] Step S44: Fuse the high-frequency components and low-frequency components to obtain the final three-light fusion feature map.

[0053] Step S5: Input the final three-light fusion feature map into a convolutional neural network containing an improved softmax operation module to obtain the power equipment defect detection result.

[0054] This system builds upon traditional convolutional neural networks by incorporating an improved softmax module into the fully connected layers, ultimately outputting the network's prediction for each sample category. Each convolutional block contains one convolutional layer and one max-pooling layer, doubling the number of channels and halving the feature size. Implemented using the PyTorch framework, the network weights and bias parameters are initialized using the Xavier method, cross-entropy loss is used as the loss function, and temporary backoff is applied to the fully connected layers. Multiple cascaded CNN network modules predict different infrared and ultraviolet samples, outputting the prediction results to achieve fault diagnosis of power equipment.

[0055] Specifically, the improved softmax operation module satisfies the following formula:

[0056]

[0057]

[0058]

[0059]

[0060] in, The output vector is represented by the first... One element, Represents the first input vector. One element, Represents the first input vector. One element, For vector dimensions, , , These are the 1st, 2nd, and 3rd elements of the input vector, respectively. One element;

[0061] The gradients of a convolutional neural network containing an improved softmax operation module have an efficient form during backpropagation. Specifically, the improved softmax operation module also satisfies the following equation:

[0062]

[0063] in, It is the Kronecker function, when hour, ,on the contrary, ; The output vector is represented by the first... One element, Indicates partial derivative, This represents the output gradient.

[0064] In summary, the power equipment diagnostic method based on three-light fusion according to the above embodiments has the following beneficial effects:

[0065] (1) This invention inputs the preprocessed visible light image, infrared image and ultraviolet image into the dual-flow depth feature extraction network. First, based on the multi-scale discharge spot features and multi-scale hot spot features, multi-physics field coupling enhancement features are obtained. Then, combined with the multi-scale high-resolution texture features, multi-modal registered images are obtained, realizing full spectrum coverage of the defect physical field, solving the problem of blind spots in single-modal information, and improving the defect detection accuracy.

[0066] (2) This invention fuses multi-frequency domain collaborative feature decomposition and physical constraint multimodal image reconstruction fusion module to obtain low-frequency components containing common attribute feature maps and high-frequency components containing unique attribute feature maps. The common attribute feature maps represent cross-modal coupling steady-state features, and the unique attribute feature maps represent modality-specific transient features. By fusing the high-frequency components and low-frequency components, the final three-light fusion feature map is obtained, which breaks through the feature space heterogeneity and strengthens the feature correlation of different feature maps.

[0067] (3) Inputting the final three-light fusion feature map into a convolutional neural network containing an improved softmax operation module can overcome high and low frequency interference in complex scenes and significantly improve the accuracy and interpretability of defect detection in complex environments.

[0068] The embodiments described above are merely illustrative of several implementations of the present invention, and while the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the invention patent. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these all fall within the protection scope of the present invention. Therefore, the protection scope of this invention patent should be determined by the appended claims.

Claims

1. A method for diagnosing power equipment based on three-light fusion, characterized in that, include: Step S1: Use a visible light camera, an infrared thermal imager, and an ultraviolet camera to acquire images of the power equipment, obtaining the original visible light image, the original infrared image, and the original ultraviolet image of the power equipment. Step S2: Preprocess the original visible light image, the original infrared image, and the original ultraviolet image to obtain the preprocessed visible light image, the processed infrared image, and the processed ultraviolet image. Step S3: Input the preprocessed visible light image, the processed infrared image, and the processed ultraviolet image into the dual-stream depth feature extraction network. First, based on the multi-scale discharge spot features and the multi-scale hot spot features, obtain the multi-physics field coupling enhancement features. Then, combine the multi-scale high-resolution texture features to obtain the multi-modal registered image. Step S4: Based on the image after multimodal registration, the image is fused by multi-frequency domain collaborative feature decomposition and physical constraint multimodal image reconstruction fusion module to obtain low-frequency components containing common attribute feature maps and high-frequency components containing unique attribute feature maps. The high-frequency components and low-frequency components are fused to obtain the final three-light fusion feature map. Step S5: Input the final three-light fusion feature map into a convolutional neural network containing an improved softmax operation module to obtain the power equipment defect detection result; The dual-stream deep feature extraction network includes visible light branch, infrared branch and ultraviolet branch, association coding module and multimodal registration module; The visible light branch is used to learn the semantic features of the preprocessed visible light image in order to obtain high-resolution texture features at multiple scales; The infrared branch is used to learn the semantic features of the preprocessed infrared image to obtain multi-scale discharge spot features; The ultraviolet branch is used to learn the semantic features of the preprocessed ultraviolet image in order to obtain multi-scale hot spot features; The correlation coding module is used to correlate the discharge spot features and hot spot features based on the relationship between temperature and discharge intensity to obtain multi-physics field coupling enhancement features; The multimodal registration module is used to obtain the multimodal registered image based on multi-physics field coupling enhancement features and multi-scale high-resolution texture features through cross-attention weight allocation; The associated encoding module satisfies the following formula: in, This is a feature that enhances multiphysics coupling. For ELU activation function, , For trainable weight matrix, For normalization function, To enhance the hot spot characteristics, To quantify the discharge intensity based on the characteristics of the discharge spot, To enhance the characteristics of the hot spot before enhancement, It is a non-linear activation function. For 1×1 convolution, Indicates global average pooling. For convolution operations, This is element-wise addition; The multimodal registration module satisfies the following formula: in, This represents the image after multimodal registration. For normalization function, For high-resolution texture features, For querying the matrix, The key matrix, Indicates transpose. For feature dimensions; Step S4 specifically includes: Step S41: The multimodal image reconstruction and fusion module decomposes low-frequency coefficients and high-frequency coefficients based on the multimodal registered image; Step S42: Based on the low-frequency coefficients, the multi-physics field equations are embedded as soft constraints into the multimodal image reconstruction and fusion module. The residual loss function is used to make the fusion features conform to physical laws. The low-frequency coefficients are fused and inversely transformed to obtain the low-frequency components containing the shared attribute feature map. Step S43: Based on the high-frequency coefficients, a differentiable physics engine is developed. By combining multi-scale high-resolution texture features, discharge spot features and hot spot features, the high-frequency coefficients are fused and inversely transformed to obtain high-frequency components containing unique attribute feature maps. Step S44: Fuse the high-frequency components and low-frequency components to obtain the final three-light fusion feature map.

2. The power equipment diagnostic method based on three-light fusion according to claim 1, characterized in that, The visible light branch deploys a VIT-Hybrid network to extract component-level RIO features and captures the global topological relationships of each component in the power equipment through a multi-head attention mechanism, while cooperating with the backbone network of a convolutional neural network to extract local texture features. The infrared branch design temperature calibration residual network introduces an adaptive temperature compensation layer to eliminate the influence of ambient temperature fluctuations, and employs a hot spot feature enhancement unit to increase the channel attention weights in abnormal temperature rise regions. The ultraviolet branch employs a deformable convolutional discharge intensity quantization network, utilizing deformable convolutional kernels to capture the gradient characteristics of irregular discharge spots.

3. The power equipment diagnostic method based on three-light fusion according to claim 2, characterized in that, The expression for the loss function of the dual-stream deep feature extraction network is: in, This represents the topology-aware contrast loss value. For positive sample pairs, For negative sample pairs, For the sample set, For natural index, The cosine similarity function is used. For the sample The corresponding component feature vector, For the sample The corresponding component feature vector, For the sample The corresponding component feature vector, This is the temperature coefficient.

4. The power equipment diagnostic method based on three-light fusion according to claim 3, characterized in that, In step S43, the following equation is satisfied during the fusion and inverse transformation of the high-frequency coefficients: in, Indicates the fusion of the first High-frequency components of the scale, Indicates the first Scale edge mask, This indicates taking the maximum value. Indicates the first The characteristics of the discharge spot at a specific scale. Indicates the first Scale-based hot spot characteristics, Indicates the first High-resolution texture features at scale.

5. The power equipment diagnostic method based on three-light fusion according to claim 1, characterized in that, The improved softmax operation module satisfies the following formula: in, The output vector is represented by the first... One element, Represents the first input vector. One element, Represents the first input vector. One element, For vector dimensions, , , These are the 1st, 2nd, and 3rd elements of the input vector, respectively. One element; The improved softmax operation module also satisfies the following formula: in, It is the Kronecker function. The output vector is represented by the first... One element, This represents the partial derivative.

6. The power equipment diagnostic method based on three-light fusion according to claim 1, characterized in that, The visible light camera, infrared thermal imager, and ultraviolet camera are all equipped with positioning modules, and the original infrared images and original ultraviolet images are marked with the acquisition area and acquisition time.

Citation Information

Patent Citations

  • Transformer substation power equipment fault detection method based on multi-source fusion

    CN115661044A

  • Ultraviolet light and visible light fusion method for electrical equipment detection

    CN115937268A