A multi-modal image visualization display method for discharge defects of a power transmission line
By using multi-layer cascaded and residual convolutional structures for feature extraction and spatial alignment, combined with feature purification and cross-modal fusion, the problem of identifying weak discharge and heating defects in power transmission line inspections has been solved. This has enabled high-precision, automated multimodal image visualization detection, improving the stability and readability of the detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- MAANSHAN POWER SUPPLY COMPANY STATE GRID ANHUI ELECTRIC POWER
- Filing Date
- 2026-05-14
- Publication Date
- 2026-07-03
AI Technical Summary
Existing single-mode detection methods are difficult to identify early weak discharge and hidden heating defects in transmission line inspections, and multi-mode data cannot be accurately spatially aligned, resulting in high false negative rates, weak anti-interference capabilities, and insufficient positioning accuracy, which cannot meet the needs of all-weather, high-precision automated field inspections.
Feature extraction is performed using a multi-layered cascaded lightweight convolutional structure and a multi-level residual convolutional structure. Spatial dimension alignment is achieved by combining a shared downsampling strategy. A feature purification module and a trimodal cross-attention fusion module are constructed. Multi-scale feature extraction and fusion are performed through a feature pyramid network. An improved YOLOv8 detection module is used for defect detection. A multimodal micro-defect detection model is iteratively trained using an SGD optimizer to generate multimodal visualization images.
It enables automated detection of minute discharge and heat defects in complex field environments, significantly reducing the false detection rate, improving detection sensitivity and recall rate, and realizing a three-in-one visual display of defect location, discharge intensity and heat intensity, thereby improving the readability of inspection results and the efficiency of on-site handling.
Smart Images

Figure SMS_219 
Figure QLYQS_18 
Figure QLYQS_28
Abstract
Description
Technical Field
[0001] This invention belongs to the field of power transmission line inspection technology, specifically relating to a multimodal visual fusion method for detecting discharge defects in power transmission lines. Background Technology
[0002] In current inspection work, single-mode detection methods such as visible light, ultraviolet light, and infrared are commonly used: visible light imaging can clearly display equipment structure and location information, but it is difficult to identify early weak discharges and hidden [discharges]. Current inspection operations primarily employ single-modal detection methods using visible light, ultraviolet (UV), and infrared (IR): Visible light imaging clearly presents equipment structure and location information, but struggles to identify early, weak discharges and hidden heating defects; UV imaging can capture corona and arc signals, but is susceptible to interference from solar UV background, leading to misjudgments; Infrared imaging can reflect temperature distribution differences within equipment, but is easily affected by ambient high temperatures and solar radiation, lacking sensitivity to subtle anomalies. Furthermore, the three camera modalities differ in resolution, field of view, and imaging space, making precise spatial alignment of multimodal data impossible. Consequently, corona and heating features cannot be effectively fused and jointly determined under a unified view.
[0003] In addition, early defects in transmission lines are usually manifested as weak anomalies at the pixel level. Traditional detection methods lack targeted anti-interference mechanisms and anomaly enhancement strategies, and are prone to losing minute defect features in complex background noise. They suffer from problems such as high missed detection rate, weak anti-interference ability, and insufficient positioning accuracy, making it difficult to meet the needs of all-weather, high-precision, and automated field inspection.
[0004] Furthermore, early defects in transmission lines typically manifest as pixel-level anomalies. Traditional detection methods lack targeted anti-interference mechanisms and anomaly enhancement strategies, making it difficult to preserve minute defect features amidst complex background noise. These methods suffer from high false negative rates, weak anti-interference capabilities, and insufficient positioning accuracy, failing to meet the requirements for all-weather, high-precision, and automated field inspections. Summary of the Invention
[0005] This invention addresses the detection of high-voltage arc discharge defects in transmission lines by providing a multimodal image visualization method for such defects. The method aims to effectively suppress background interference, achieve accurate multimodal data registration, anti-interference feature purification, cross-modal adaptive fusion, and visualization, thereby enabling automated detection of minute high-voltage arc discharge defects in transmission lines and improving the accuracy and reliability of early defect identification.
[0006] To achieve the above-mentioned objectives, the present invention adopts the following technical solution: The present invention provides a multimodal image visualization method for discharging defects in transmission lines, characterized by the following steps: Step 1: Acquire one ultraviolet image each using an ultraviolet camera, an infrared camera, and a visible light camera. An infrared image A visible light image ;exist The true label information for the defects defined above includes: the true bounding box parameters of the defects and the true defect category; Step 2: Construct a feature extraction module, including: multi-layer cascaded lightweight convolutional structures and multi-level residual convolutional structures; each lightweight convolutional structure includes: convolution operation unit, batch normalization processing unit, and non-linear activation function; Step 2.1: The multi-layered cascaded lightweight convolutional structures are respectively applied to... and The ultraviolet branch features are processed and output by the lightweight convolutional structure in the last cascade. and infrared branching features ; Step 2.2, Multi-level residual convolution structure pair Feature extraction was performed to obtain visible light branching features. ; Step 2.3: Employ a shared downsampling strategy for... , and Spatial dimension alignment is performed to obtain the aligned ultraviolet branch features. Aligned infrared branch features and aligned visible light branch features ; Step 3: Construct a feature purification module to process the... , and Branch features after alignment of any i-th mode Processing is performed to obtain the purified features in the i-th mode. , i The modal index corresponds to the ultraviolet mode, infrared mode, and visible light mode, respectively. Step 4: Construct a trimodal cross-attention fusion module and process the purified feature vectors. The process is performed to obtain multimodal fusion features. Thus, a cross-modal consistency loss function is constructed. ; Step 5: Construct a feature pyramid network for... Multi-scale feature extraction and feature fusion processing are performed to obtain a multi-scale fused feature set. ={ | l },in, Indicates the first Fusion characteristics at different scales This indicates the highest level of cross-scale fusion; Step 6: Build an improved YOLOv8 detection module and perform... The data is processed to obtain the prediction results, including: the predicted bias parameters, the prediction box parameters of the defects, the predicted defect categories, and the confidence level. Step 7: Based on the real label information and prediction results, construct the total loss function of the multimodal small defect detection network, which consists of a feature purification module, a three-modal cross-attention fusion module, and a prediction module. ; Step 8: Iteratively train the multimodal small defect detection network using the SGD optimizer in conjunction with the backpropagation algorithm, and calculate the total loss function. until the total loss function The training process converges to obtain a well-trained multimodal micro-defect detection model, which is used to detect multimodal micro-defects on transmission line images. Step 9: Calculate the UV characteristics after purification. UV normalized intensity and purified infrared characteristics infrared normalized intensity Used to calculate severity level ; Step 10, Mapped to the purple color gamut and overlaid. Generate intermediate images. Then Overlay mapping to the red color gamut and overlay to This generates an intermediate image by overlaying ultraviolet and infrared thermograms. ; Step 11, based on the predicted absolute coordinates of the defect bounding box, in The corresponding defect bounding boxes are drawn on top, and the prediction results and severity levels are converted into pixel text and overlaid on top of the defect bounding boxes as the final output visualization image. This allows for the simultaneous visualization of defect location, discharge characteristics, and heating characteristics.
[0007] The multimodal image visualization method for discharge defects in transmission lines described in this invention is also characterized in that step 3 includes: Step 3.1 will After being expanded by pixel position, a one-dimensional feature vector with the same dimension as the number of channels is obtained. ,in, Represents a one-dimensional eigenvector in the i-th mode. The Middle j A pixel feature vector, where N represents the total number of pixels in the feature vector after downsampling; Step 3.2: Construct the normal background feature dictionary for the i-th modality. ; Step 3.3, for Perform sparse coding to obtain sparsity coefficient And utilize the sparsity coefficient and normal background feature dictionary right Reconstruction is performed to obtain the reconstructed feature vector. = ; Step 3.4, Calculation and Reconstruction error between Thus, the reconstruction error vector under the i-th mode is obtained. Standard deviation and mean Used to calculate the adaptive anomaly threshold in the i-th mode. ; Step 3.5: Use equation (1) to obtain the purified features in the i-th mode.
[0008] (1) In equation (1), express The feature value of the c-th channel in the m-th row, n-th column, is... express The feature value of the m-th row, n-th column, and c-th channel in the image; (m, n) are the pixel coordinates, and c is the channel index.
[0009] Furthermore, step 4 includes: Step 4.1, Perform channel-dimensional aggregation to generate joint features. ; Step 4.2, for Perform global average pooling to obtain the global feature vector. Then use a 1×1 convolution to... The number of channels is reduced from 3C to C, thus obtaining the dimensionality-reduced feature vector. ; Step 4.3, will The input is processed through a Sigmoid activation function to obtain a range of values. attention weight vector ; Step 4.4, use equation (2) to... Weighted fusion is performed to obtain multimodal fusion features. : (2) In equation (2), express The attention weight vector in the i-th modality; Step 4.5, construct the fusion consistency loss function using equation (3). : (3).
[0010] Furthermore, step 5 includes: Step 5.1, when At that time, the order was made Low-resolution features at scale Therefore, equation (4) is used for... Scale feature extraction is performed to obtain the first Low-resolution features at scale Thus, the first Fusion features at different scales ; (4) In equation (4), Indicates a downsampling operation; Step 5.2, when At that time, the order was made Fusion features at different scales Therefore, equation (5) is used to... Perform cross-scale fusion to obtain the first Fusion features at different scales ; (5) In equation (5), Indicates an upsampling operation. This indicates element-wise addition.
[0011] Furthermore, step 6 includes: Step 6.1, will The input is fed into the regression branch convolutional layer of the YOLOv8 detector head to obtain the first... Offset features at scale ;in, Indicates the first The x-coordinate skewness of the center point of the defect prediction box at this scale. Indicates the first The y-coordinate offset of the center point of the defect prediction box at the specified scale. Indicates the first Width deviation of the pre-defect measurement frame under the specified scale. Indicates the first The height bias of the defect prediction box at the scale; Step 6.2, using equation (5) to obtain the first... Relative coordinates of the defect prediction box at different scales ( Then, through coordinate transformation, the first... The absolute coordinates of the defect prediction box at different scales ( ): (6) In equation (6), For the Sigmoid function; , express Width and height, , They represent the first The x-coordinate, y-coordinate, width, and height of the center point of the defect prediction bounding box at the specified scale. , They represent the first The x-coordinate of the center point, y-coordinate of the center point, width, and length of the actual bounding box at the specified scale; Step 6.3, apply the Softmax activation function to... The process is performed to obtain the predicted defect category; Step 6.4, apply the Sigmoid activation function to... The data is processed to obtain the confidence level of the prediction.
[0012] Furthermore, step 7 includes: Step 7.1: Construct the feature purification loss as the average value of the purification losses of the three modes of ultraviolet, infrared, and visible light; Step 7.2, construct the detection loss, including: bounding box regression loss, category loss, and confidence loss; Step 7.3, Construct the total loss function It is a weighted combination of feature purification loss, detection loss, and cross-modal consistency loss.
[0013] Furthermore, step 9 includes: Step 9.1: Use bilinear interpolation to extract the purified UV features. and purified infrared characteristics Sampled to visible light image The scale is determined to obtain the scale-aligned ultraviolet features. Infrared features after scale alignment ; Step 9.2, respectively for and Performing global extremum normalization, the corresponding normalized ultraviolet intensity in the [0,1] interval is obtained. and infrared normalized intensity ; Step 9.3, calculate the severity level using equation (6). : (6) In equation (6), This represents the floor function, and N represents the total number of grades.
[0014] The present invention provides an electronic device, including a memory and a processor, characterized in that the memory is used to store a program supporting the processor in performing the method described therein, and the processor is configured to execute the program stored in the memory.
[0015] The present invention discloses a computer-readable storage medium storing a computer program, characterized in that the computer program is executed by a processor to perform the steps of the method described thereon.
[0016] Compared with the prior art, the beneficial effects of the present invention are as follows: 1. This invention proposes a modal feature purification and anti-interference mechanism. By adaptively fitting the normal background distribution of each modality through dictionary learning, it can effectively filter out solar ultraviolet interference, environmental high temperature uniform background interference and natural background clutter, significantly reducing the false detection rate of single-modal detection and improving the detection stability in complex field environments.
[0017] 2. This invention solves the problems of interference from solar ultraviolet background and high ambient temperature background, and realizes the automated detection of tiny discharge / heating defects, identifying anomalies through pixel-level feature changes.
[0018] 3. This invention amplifies and highlights pixel-level weak defect features by reconstructing error judgment and enhancing abnormal features, solving the problem of difficulty in identifying small defects such as early arc discharge and slight equipment heating, and greatly improving the detection sensitivity and recall rate of small defects.
[0019] 4. This invention enables a three-in-one visual display of defect location, discharge intensity, and heat intensity. Maintenance personnel can directly observe the defect location, discharge area, heat area, and defect severity level on a clear visible light image, which greatly improves the readability of inspection results and the efficiency of on-site handling. Detailed Implementation
[0020] In this embodiment, a multimodal image visualization method for transmission line discharge defects is a three-modal fusion detection scheme that can achieve accurate registration of multimodal data, effectively suppress background interference, and sensitively detect pixel-level minute defects. Specifically, it includes the following steps: Step 1, using a drone equipped with a resolution of Ultraviolet camera, resolution: Infrared camera, resolution: A visible light camera simultaneously acquires an ultraviolet image during flight. An infrared image A visible light image ;exist The true label information for defining defects includes: the true bounding box of the defect and the true defect category.
[0021] Step 2: Construct a feature extraction module, including: multi-layer cascaded lightweight convolutional structures and multi-level residual convolutional structures; each lightweight convolutional structure includes: convolution operation unit, batch normalization processing unit, and non-linear activation.
[0022] Step 2.1, the multi-layered cascaded lightweight convolutional structures are respectively applied to... and The ultraviolet branch features are processed and output by the lightweight convolutional structure in the last cascade. and infrared branching features Ultraviolet branching characteristics Extract local features and infrared branching features of the discharge spot. Extract gradient features of temperature anomalies.
[0023] Step 2.2, Multi-level residual convolutional structure pairs Feature extraction was performed to obtain visible light branching features. Extract the texture and structural features of transmission lines and insulators.
[0024] Step 2.3: Employ a shared downsampling strategy for... , and Spatial dimension alignment is performed to obtain the aligned ultraviolet branch features. Aligned infrared branch features and aligned visible light branch features ; By adopting a shared downsampling strategy, the scale consistency of different modal features can be guaranteed. At the same time, by controlling the number of convolution kernels, the number of model parameters can be kept within a reasonable range, making it suitable for deployment of drone edge devices.
[0025] Step 3: Construct a feature purification module to process the... , and Branch features after alignment of any i-th mode Processing is performed to obtain the purified features in the i-th mode. , i This is the modal index, corresponding to the ultraviolet mode, infrared mode, and visible light mode in sequence.
[0026] Step 3.1, will After being expanded by pixel position, a one-dimensional feature vector with the same dimension as the number of channels is obtained. ,in, Represents a one-dimensional eigenvector in the i-th mode. The Middle j N is a pixel feature vector, where N represents the total number of pixels in the feature vector after downsampling.
[0027] Step 3.2: Construct the normal background feature dictionary for the i-th modality using equation (1). ; (1) In equation (1), For the updated dictionary, the number of iterations is set to 50, until the reconstruction error converges, and the convergence threshold is [value missing]. .
[0028] Step 3.3, for Perform sparse coding to obtain sparsity coefficient And utilize the sparsity coefficient and normal background feature dictionary right Reconstruction is performed to obtain the reconstructed feature vector. = .
[0029] Step 3.4, calculate using equation (2) and Reconstruction error between Thus, the reconstruction error vector under the i-th mode is obtained. Standard deviation and mean Used to calculate the adaptive anomaly threshold in the i-th mode. ; (2) In equation (2), These are the original feature vector and the reconstructed feature vector, respectively. Each channel value.
[0030] Step 3.5: Using the formula shown in equation (3) as the basis for pixel category determination, the purified features in the i-th modality are obtained.
[0031] (3) In equation (3), express The feature value of the channel in the m-th row, n-th column, and c-th column; (m, n) are the pixel coordinates of the feature, and c is the channel index; It is an anomalous feature enhancement coefficient, which amplifies minute anomalous features to meet the needs of pixel-level minute defect detection.
[0032] Step 4: Construct a trimodal cross-attention fusion module and process the purified features. The process is performed to obtain multimodal fusion features. Thus, a cross-modal consistency loss function is constructed.
[0033] Step 4.1, use formula (4) to purify the characteristic Perform channel-dimensional aggregation to generate joint features. ; (4) In equation (4), C = 32 and 3C = 96.
[0034] Step 4.2, use formula (5) to... Global average pooling is performed to compress the two-dimensional spatial information into a channel-dimensional global feature vector. Then use a 1×1 convolution to... The number of channels is reduced from 3C to C, resulting in the dimensionality-reduced feature vector. ;
[0035] (5) In equation (5), This is a global average pooling operation.
[0036] Step 4.3, will The input is processed through a Sigmoid activation function to obtain a range of values. attention weight vector ; Step 4.4, use equation (6) to... Weighted fusion is performed to obtain multimodal fusion features. : (6) In equation (6), express The attention weight vector in the i-th modality satisfies: This enables adaptive normalization of weights.
[0037] Step 4.5, construct the consistency loss function using equation (7). : (7) In equation (7), This represents the square of the L2 norm, used to calculate the Euclidean distance between features, measuring the difference between the fused feature and each individual modal feature.
[0038] Step 5: Construct a feature pyramid network for... Multi-scale feature extraction and feature fusion processing are performed to obtain a multi-scale fused feature set. ={ | l },in, Indicates the first Fusion characteristics at different scales This indicates the highest level of cross-scale fusion. Defects of different sizes can be extracted through multi-scale feature extraction and feature fusion processing.
[0039] Step 5.1, when At that time, the order was made Low-resolution features at scale Using formula (4) Scale feature extraction is performed to obtain the first Scale features at scale Thus, the first Fusion features at different scales ; (8) In equation (8), This indicates a downsampling operation, which reduces the feature resolution by half to generate the first... Low-resolution features of the layer .
[0040] Step 5.2, when At that time, the order was made Fusion features at different scales Therefore, equation (9) is used to... Perform cross-scale fusion to obtain the first Fusion features at different scales : (9) In equation (9), Indicates hierarchy; This indicates an upsampling operation, which amplifies deep features to the same level as... Same resolution; For the first The deep fusion characteristics of the layers; The feature is added element by element to achieve cross-layer fusion of deep semantic features and shallow detail features; For the final result Multi-scale fusion features, which take into account both high-resolution detail information and low-resolution semantic information, provide multi-scale feature support for the detection of minute defects.
[0041] Step 6: Build an improved YOLOv8 detection module and perform... The data is processed to obtain the prediction results, including: the predicted bias parameters, the prediction box parameters of the defects, the predicted defect categories, and the confidence level. Step 6.1, will The input is fed into the regression branch convolutional layer of the YOLOv8 detector head to obtain the offset features. ,in, Indicates the first The x-coordinate skewness of the center point of the defect prediction box at this scale. Indicates the first The y-coordinate offset of the center point of the defect prediction box at the specified scale. Indicates the first Width offset of the defect prediction box at different scales. Indicates the first The height bias of the defect prediction box at the scale.
[0042] Step 6.2, using equation (5) to obtain the first... Relative coordinates of the defect prediction box at different scales ( Then, through coordinate transformation, the first... The absolute coordinates of the defect prediction box at different scales ( ): (10) In equation (10), Use the Sigmoid activation function; , They represent the first Fusion features at different scales Width and height, , They represent the first The x-coordinate, y-coordinate, width, and height of the center point of the defect prediction bounding box at the specified scale. , They represent the first The x-coordinate, y-coordinate, width, and length of the anchor frame at this scale are obtained by performing k-means clustering on the labeled boxes in the training set to match the size range of the defect targets to be detected at this scale.
[0043] Step 6.2, apply the Softmax activation function shown in equation (11) to... Processing is performed to obtain the first... Defect categories predicted at different scales; (11) In equation (11), For the first Category prediction features at different scales; For the first The first scale Class predictive features; For the first The probability that a defect belongs to class c at a given scale; the class corresponding to the maximum probability is taken as the class c. Final defect category at scale ,Right now: ,in, This represents the total number of categories.
[0044] Step 6.3, apply the Sigmoid activation function shown in equation (12) to... Processing is performed to obtain the first... Confidence of prediction at different scales ; (12) In equation (12), Indicates the first Confidence prediction features at different scales ,when If a defect is identified, it is considered a background element.
[0045] Step 7: Based on the real labels and prediction results, construct the multi-task total loss function of the multimodal small defect detection network, which consists of a feature extraction module, a feature purification module, and a prediction module. The total loss function comprehensively considers defect localization loss, category loss, confidence loss, feature purification loss, and fusion alignment loss, achieving synergistic optimization of detection accuracy, feature purification effect, and fusion alignment. This ensures that the model achieves optimal performance in anti-interference, detection of minor defects, and multimodal fusion, adapting to the real-time detection needs of UAV edge devices.
[0046] Step 7.1, calculate the purification loss of the i-th mode using equation (13). : (13) In equation (13), express The total number of pixels in the feature after downsampling; This is an indicator function; it takes the value 1 if the condition is met, and 0 otherwise. Indicates the i-th mode Upper The reconstruction error is 1 pixel.
[0047] The average purification loss of the three modes was calculated using equation (14). : (14) Step 7.2, define the first The total number of defect prediction boxes at the scale is ; Define the first The number of positive samples at the scale is The number of predicted bounding boxes that successfully match the ground truth defect bounding boxes and contain the defect target after IoU matching is denoted as . ; Define the first Negative samples at scale are The number of predicted bounding boxes that successfully match the true defect bounding box after IoU matching and do not contain the defect target is denoted as . ; The detection loss for the prediction module is calculated using equation (15). : (15) In equation (15), Indicates the first The bounding box regression loss at the scale is calculated using equation (16); Indicates the first The class loss at the scale is calculated using equation (17); Indicates the first The confidence loss under the scale is calculated by equation (18).
[0048] (16) In equation (16), For the first The absolute coordinates of the defect prediction box at the specified scale; For intersection, union, and comparison; For the first Absolute coordinates of the defect prediction box at different scales With the The corresponding scale is the first A true bounding box The squared Euclidean distance between the absolute coordinates, The length of the diagonal of the smallest bounding rectangle enclosing the two frames. For balance coefficient, The aspect ratio consistency coefficient is denoted as , where . , Indicates the first A true bounding box The width and height are represented.
[0049] (17) In equation (17), For the first The true class label of a positive sample For the first The first scale The prediction for the positive sample is the _th The probability of a class.
[0050] (18) In equation (18); Indicates the first The first scale Confidence prediction values for each positive sample; For the first The first scale Confidence prediction values for each negative sample.
[0051] Step 7.3, using Equation (19) to weighted integrate feature purification loss Detection loss of the prediction module and fusion alignment loss Thus obtain : (19) In equation (19), , , The weight coefficients for each loss component are determined based on the convergence of model training, ensuring a balanced optimization of detection accuracy, feature purification effect, and fusion alignment.
[0052] Step 8: Iteratively train the multimodal small defect detection network using the SGD optimizer in conjunction with the backpropagation algorithm, and calculate the total loss function. until the total loss function The training process converges to obtain a well-trained multimodal micro-defect detection model, which is used to detect multimodal micro-defects on transmission line images. Step 8, Calculation UV normalized intensity map and Normalized intensity map Used to obtain the severity level ; Step 8.1: Using bilinear interpolation as shown in equation (18), the purified UV features are... and purified infrared characteristics Sampling to the visible light image scale yields scale-aligned ultraviolet features. Infrared features after scale alignment ; (20) In equation (20), , indicating the upsampling factor, Indicates the first Features after scale alignment under each modality This is the modal index, corresponding to the ultraviolet mode and the infrared mode, respectively.
[0053] Step 8.2, using the formula shown in equation (21), respectively for... and Performing global extremum normalization yields the corresponding normalized intensity map for the [0,1] interval. and : (twenty one) In equation (21), Indicates the first Normalized intensity maps for each mode.
[0054] Step 8.3, calculate the severity level using equation (22). : (twenty two) In equation (22), This represents the floor function, where N represents the total number of grades. The value range is [1,5], corresponding to the defect severity levels: Level 1 (slight discharge / heating), Level 2 (mild defect), Level 3 (moderate defect), Level 4 (relatively severe defect), and Level 5 (severe defect), realizing the quantitative output of defect severity.
[0055] Step 9, use equation (23) to... Mapping to the purple color gamut to generate an ultraviolet thermal map Using formula (24) Superimposed onto the red domain to generate an ultraviolet thermal map Using formula (25), first... Overlay Generate intermediate images Then Overlay Generate an intermediate image after overlaying ultraviolet and infrared thermograms. This achieves the visualization fusion of multimodal defect heatmaps: (twenty three) In equation (21), This represents the RGB values of the corresponding pixels in the UV thermal map, with a purple gradient; the higher the intensity, the closer it is to pure purple.
[0056] (twenty four) In equation (22), This represents the RGB values of the corresponding pixels in the infrared thermal image, with a red gradient; the higher the intensity, the closer it is to pure red. The two different color gradients allow for differentiated visualization of various defect types.
[0057] (25) In equation (23), This indicates the superimposed transparency coefficient. express or Any image in the image.
[0058] Step 10, based on the predicted first Absolute coordinates of the discharge defect bounding box at the scale Using formula (26) The corresponding defect bounding box is drawn on top, and the prediction result and severity level are converted into pixel text and superimposed on the defect bounding box using Equation (27) as the final output visualization image. This allows for the simultaneous visualization of defect location, discharge characteristics, and heating characteristics.
[0059] (26) In equation (24), For the first The coordinates of the top-left corner of the predicted bounding box for defects at this scale. For the first The coordinates of the lower right corner of the predicted bounding box at different scales are ensured to be within the resolution range of the visible light image.
[0060] (27) In equation (27), This represents a category mapping function. This represents the string concatenation symbol. This indicates that the category, confidence level, and rating are separated into three sections. This indicates an explanation or clarification. Indicates text that is displayed permanently, and includes: (28) In this embodiment, an electronic device includes a memory and a processor. The memory stores a program that supports the processor in executing the above-described method, and the processor is configured to execute the program stored in the memory.
[0061] In this embodiment, a computer-readable storage medium stores a computer program, which is executed by a processor to perform the steps of the above method.
Claims
1. A method for multi-modal image visualization display of discharge defects of a power transmission line, characterized in that, Includes the following steps: Step 1, acquire one ultraviolet image respectively by using ultraviolet camera, infrared camera and visible light camera , one infrared image , one visible light image ; define the real label information of the defect on , which includes real bounding box parameters of the defect and real defect category Step 2: Construct a feature extraction module, including: multi-layer cascaded lightweight convolutional structures and multi-level residual convolutional structures; each lightweight convolutional structure includes: convolution operation unit, batch normalization processing unit, and non-linear activation function; Step 2.1: The multi-layered cascaded lightweight convolutional structures are respectively applied to... and The ultraviolet branch features are processed and output by the lightweight convolutional structure in the last cascade. and infrared branching features ; Step 2.2, Multi-level residual convolution structure pair Feature extraction was performed to obtain visible light branching features. ; Step 2.3: Employ a shared downsampling strategy for... , and Spatial dimension alignment is performed to obtain the aligned ultraviolet branch features. Aligned infrared branch features and aligned visible light branch features ; Step 3: Construct a feature purification module to process the... , and Branch features after alignment of any i-th mode Processing is performed to obtain the purified features in the i-th mode. , i The modal index corresponds to the ultraviolet mode, infrared mode, and visible light mode, respectively. Step 4: Construct a trimodal cross-attention fusion module and process the purified feature vectors. The process is performed to obtain multimodal fusion features. Thus, a cross-modal consistency loss function is constructed. ; Step 5: Construct a feature pyramid network for... Multi-scale feature extraction and feature fusion processing are performed to obtain a multi-scale fused feature set. ={ | l },in, Indicates the first Fusion characteristics at different scales This indicates the highest level of cross-scale fusion; Step 6: Build an improved YOLOv8 detection module and perform... The data is processed to obtain the prediction results, including: the predicted bias parameters, the prediction box parameters of the defects, the predicted defect categories, and the confidence level. Step 7: Based on the real label information and prediction results, construct the total loss function of the multimodal small defect detection network, which consists of a feature purification module, a three-modal cross-attention fusion module, and a prediction module. ; Step 8: Iteratively train the multimodal small defect detection network using the SGD optimizer in conjunction with the backpropagation algorithm, and calculate the total loss function. until the total loss function The training process converges to obtain a well-trained multimodal micro-defect detection model, which is used to detect multimodal micro-defects on transmission line images. Step 9: Calculate the UV characteristics after purification. UV normalized intensity and purified infrared characteristics infrared normalized intensity Used to calculate severity level ; Step 10, Mapped to the purple color gamut and overlaid. Generate intermediate images. Then Overlay mapping to the red color gamut and overlay to This generates an intermediate image by overlaying ultraviolet and infrared thermograms. ; Step 11, based on the predicted absolute coordinates of the defect bounding box, in The corresponding defect bounding boxes are drawn on top, and the prediction results and severity levels are converted into pixel text and overlaid on top of the defect bounding boxes as the final output visualization image. This allows for the simultaneous visualization of defect location, discharge characteristics, and heating characteristics.
2. The multimodal image visualization method for discharge defects in transmission lines according to claim 1, characterized in that, Step 3 includes: Step 3.1 will After being expanded by pixel position, a one-dimensional feature vector with the same dimension as the number of channels is obtained. ,in, Represents a one-dimensional eigenvector in the i-th mode. The Middle j A pixel feature vector, where N represents the total number of pixels in the feature vector after downsampling; Step 3.2: Construct the normal background feature dictionary for the i-th modality. ; Step 3.3, for Perform sparse coding to obtain sparsity coefficient And utilize the sparsity coefficient and normal background feature dictionary right Reconstruction is performed to obtain the reconstructed feature vector. = ; Step 3.4, Calculation and Reconstruction error between Thus, the reconstruction error vector under the i-th mode is obtained. Standard deviation and mean Used to calculate the adaptive anomaly threshold in the i-th mode. ; Step 3.5: Use equation (1) to obtain the purified features in the i-th mode. : (1) In equation (1), express The feature value of the c-th channel in the m-th row, n-th column, is... express The feature value of the m-th row, n-th column, and c-th channel in the image; (m, n) are the pixel coordinates, and c is the channel index.
3. The multimodal image visualization method for discharge defects in transmission lines according to claim 2, characterized in that, Step 4 includes: Step 4.1, Perform channel-dimensional aggregation to generate joint features. ; Step 4.2, for Perform global average pooling to obtain the global feature vector. Then use a 1×1 convolution to... The number of channels is reduced from 3C to C, thus obtaining the dimensionality-reduced feature vector. ; Step 4.3, will The input is processed through a Sigmoid activation function to obtain a range of values. attention weight vector ; Step 4.4, use equation (2) to... Weighted fusion is performed to obtain multimodal fusion features. : (2) In equation (2), express The attention weight vector in the i-th modality; Step 4.5, construct the fusion consistency loss function using equation (3). : (3)。 4. The multimodal image visualization method for discharge defects in transmission lines according to claim 3, characterized in that, Step 5 includes: Step 5.1, when At that time, the order was made Low-resolution features at scale Therefore, equation (4) is used for... Scale feature extraction is performed to obtain the first Low-resolution features at scale Thus, the first Fusion features at different scales ; (4) In equation (4), Indicates a downsampling operation; Step 5.2, when At that time, the order was made Fusion features at different scales Therefore, equation (5) is used to... Perform cross-scale fusion to obtain the first Fusion features at different scales ; (5) In equation (5), Indicates an upsampling operation. This indicates element-wise addition.
5. The multimodal image visualization method for discharge defects in transmission lines according to claim 3, characterized in that, Step 6 includes: Step 6.1, will The input is fed into the regression branch convolutional layer of the YOLOv8 detector head to obtain the first... Offset features at scale ;in, Indicates the first The x-coordinate skewness of the center point of the defect prediction box at this scale. Indicates the first The y-coordinate offset of the center point of the defect prediction box at the specified scale. Indicates the first Width deviation of the pre-defect measurement frame under the specified scale. Indicates the first The height bias of the defect prediction box at the scale; Step 6.2, using equation (5) to obtain the first... Relative coordinates of the defect prediction box at different scales ( Then, through coordinate transformation, the first... The absolute coordinates of the defect prediction box at different scales ( ): (6) In equation (6), For the Sigmoid function; , express Width and height, , They represent the first The x-coordinate, y-coordinate, width, and height of the center point of the defect prediction bounding box at the specified scale. , They represent the first The x-coordinate of the center point, y-coordinate of the center point, width, and length of the actual bounding box at the specified scale; Step 6.3, apply the Softmax activation function to... The process is performed to obtain the predicted defect category; Step 6.4, apply the Sigmoid activation function to... The data is processed to obtain the confidence level of the prediction.
6. The multimodal image visualization method for discharge defects in transmission lines according to claim 1, characterized in that, Step 7 includes: Step 7.1: Construct the feature purification loss as the average value of the purification losses of the three modes of ultraviolet, infrared, and visible light; Step 7.2, construct the detection loss, including: bounding box regression loss, category loss, and confidence loss; Step 7.3, Construct the total loss function It is a weighted combination of feature purification loss, detection loss, and cross-modal consistency loss.
7. The multimodal image visualization method for discharge defects in transmission lines according to claim 1, characterized in that, Step 9 includes: Step 9.1: Use bilinear interpolation to extract the purified UV features. and purified infrared characteristics Sampled to visible light image The scale is determined to obtain the scale-aligned ultraviolet features. Infrared features after scale alignment ; Step 9.2, respectively for and Performing global extremum normalization, the corresponding normalized ultraviolet intensity in the [0,1] interval is obtained. and infrared normalized intensity ; Step 9.3, calculate the severity level using equation (6). : (6) In equation (6), This represents the floor function, and N represents the total number of grades.
8. An electronic device, comprising a memory and a processor, characterized in that, The memory is used to store a program that supports the processor in executing the method of any one of claims 1-7, the processor being configured to execute the program stored in the memory.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The computer program is executed by the processor to perform the steps of the method according to any one of claims 1-7.