Visual power transmission line inspection monitoring method based on AI
By performing preprocessing, stretching transformation, adaptive dehazing, and deblurring operations on transmission line images, combined with feature extraction and anomaly detection, the problem of decreased edge detail clarity after dehazing of transmission line images was solved, achieving efficient and accurate inspection results.
Patent Information
- Application Number
- CN202511043636.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-28
- Publication Date
- 2025-10-28
AI Technical Summary
In existing technologies, dehazing of transmission line images can easily lead to a decrease in the clarity of edge details, affecting the efficiency and accuracy of the inspection process.
An AI-based visualization method for power transmission line inspection is adopted. By acquiring inspection images, preprocessing, stretching transformation, adaptive dehazing and deblurring operations are performed. Combined with feature extraction and anomaly detection, a lightweight feature recognition is performed using an improved YOLOv8-seg model.
It enhances the texture details of power transmission lines, removes image occlusion, corrects loss of detail, improves the efficiency and accuracy of inspections, and provides high-quality image support.
Smart Images

Figure CN120852784A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing technology, and specifically to an AI-based visualization method for power transmission line inspection and monitoring. Background Technology
[0002] As the core carrier of power transmission, power transmission lines have developed alongside the growth of electricity demand and technological advancements. From early simple lines to complex power grids, from low voltage to ultra-high voltage, transmission capacity and coverage have continuously improved. Currently, line detection in complex environments has become crucial, driving continuous innovation in related extraction and identification technologies.
[0003] Patent No. CN118134803A discloses a method and system for dehazing images of power transmission line inspections. The method includes: S1, denoising the inspection images collected from the power transmission line and then performing image enhancement on the denoised images to obtain a preprocessed image; S2, processing the image using a deep learning-based single-image dehazing network model, introducing an exponentially growing curve to improve the activation function, and calculating a clear image using the improved model; S3, removing fog from the image using a dark channel prior dehazing algorithm to obtain the dehazed image. In the deep learning-based single-image dehazing network model, the introduction of an exponentially growing curve improves the activation function, helps solve the gradient vanishing problem, improves the model's convergence speed, accelerates the entire dehazing process, enhances image processing efficiency, and can more effectively extract image features, thereby generating a clearer image.
[0004] Existing technologies still have some problems. For example, after dehazing the images of power transmission lines, the clarity of edge details is easily reduced, which leads to a decrease in efficiency and loss of accuracy in subsequent inspection processes. Summary of the Invention
[0005] The purpose of this invention is to solve the problems mentioned in the background art above, and to propose an AI-based visualization method for power transmission line inspection and monitoring.
[0006] A first aspect of this invention provides an AI-based visual transmission line inspection and monitoring method, the method comprising:
[0007] Acquire an inspection image of the target area, preprocess the inspection image to obtain a preprocessed inspection image, and perform a stretching transformation on the preprocessed inspection image to obtain a first enhanced image;
[0008] An adaptive dehazing operation is performed on the first enhanced image to obtain a second enhanced image, and a deblurring operation is performed on the second enhanced image to obtain a third enhanced image;
[0009] The third enhanced image is subjected to feature extraction using a preset model to obtain detection results, and anomaly judgment is made on the target region based on the detection results.
[0010] Optionally, the preprocessed inspection image is stretched to obtain a first enhanced image, including:
[0011] The preprocessed inspection image is binarized to obtain a grayscale image, and the grayscale image is filtered to obtain a filtered image.
[0012] A two-dimensional Fourier transform is performed on the filtered image to obtain a complex matrix in the frequency domain, and a kernel function is determined. The kernel function is then updated to obtain the target kernel function.
[0013] The target kernel function is multiplied point-by-point with the frequency domain complex matrix to obtain the enhanced frequency domain complex matrix. The enhanced frequency domain complex matrix is then subjected to a two-dimensional inverse Fourier transform to obtain a spatial domain complex matrix. The first enhanced image is then calculated based on the spatial domain complex matrix.
[0014] Optionally, updating the kernel function to obtain the target kernel function includes:
[0015] The initial parameter combination is determined according to the kernel function, and multiple chromosomes are randomly generated according to the initial parameter combination. The chromosomes are then combined to obtain the original population.
[0016] The original population is initialized to obtain an initial population. The initial population is then iteratively updated to obtain a target population. This process continues until a preset condition is met, at which point the optimal chromosome is output.
[0017] The optimal parameter combination is determined based on the optimal chromosome, and the kernel function is updated based on the optimal parameter combination to obtain the target kernel function.
[0018] Optionally, performing an adaptive dehazing operation on the first enhanced image to obtain a second enhanced image includes:
[0019] The preprocessed inspection image is sequentially input into a convolutional layer and a feature attention layer to obtain a first dehazed feature map. The first dehazed feature map is sequentially input into a convolutional layer and a feature attention layer to obtain a second dehazed feature map. The second dehazed feature map is sequentially input into a convolutional layer and a feature attention layer to obtain a third dehazed feature map.
[0020] The third dehazing feature map is sequentially input into the convolutional layer, the feature attention layer, and the feature attention layer to obtain the fourth dehazing feature map. The first enhanced image, the fourth dehazing feature map, and the third dehazing feature map are fused to obtain the target fourth dehazing feature map.
[0021] The target fourth dehazing feature map is sequentially input into the deconvolution layer and the feature attention layer to obtain the fifth dehazing feature map. The first enhanced image, the second dehazing feature map and the fifth dehazing feature map are fused to obtain the target fifth dehazing feature map.
[0022] The target fifth dehazing feature map is sequentially input into the deconvolution layer and the feature attention layer to obtain the sixth dehazing feature map. The first enhanced image, the sixth dehazing feature map and the first dehazing feature map are fused to obtain the target sixth dehazing feature map. The target sixth dehazing feature map and the preprocessed inspection image are stitched together to obtain the second enhanced image.
[0023] Optionally, the working principle of the feature attention layer includes:
[0024] Obtain the original features, input the original features into 1×1Conv to obtain the first feature, input the original features into 3×3Conv to obtain the second feature, input the original features into ConvD2 to obtain the third feature, and input the original features into ConvD3 to obtain the fourth feature;
[0025] The first feature, the second feature, the third feature, and the fourth feature are concatenated to obtain the fifth feature. The fifth feature is then sequentially input into 1×1Conv and 3×3Conv to obtain the sixth feature. The sixth feature and the original feature are then concatenated to obtain the seventh feature.
[0026] The seventh feature is sequentially input into the SE module and the 3×3Conv to obtain the eighth feature, which is then used as the output of the feature attention layer.
[0027] Optionally, a deblurring operation is performed on the second enhanced image to obtain a third enhanced image, including:
[0028] The second enhanced image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the first deblurred image. The first deblurred image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the second deblurred image. The second deblurred image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the third deblurred image.
[0029] The third deblurred image is sequentially input into the convolutional layer, the target neural gradient layer, and the convolutional layer to obtain the fourth deblurred image. The preprocessed inspection image, the fourth deblurred image, and the third deblurred image are then fused to obtain the target fourth deblurred image.
[0030] The target fourth deblurred image is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the fifth deblurred image. The preprocessed inspection image, the fifth deblurred image and the second deblurred image are fused to obtain the target fifth deblurred image.
[0031] The target fifth deblurred image is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the sixth deblurred image. The preprocessed inspection image, the sixth deblurred image and the first deblurred image are fused to obtain the target sixth deblurred image.
[0032] The sixth deblurred image of the target is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the seventh deblurred image, and the seventh deblurred image is used as the third enhancement image.
[0033] Optionally, the detection result is obtained by extracting features from the third enhanced image using a preset model, wherein the preset model, based on improvements to YOLOv8-seg, includes:
[0034] In the YOLOv8-seg model, a target attention module is added between the skip connections of the C2f module in layer 4 and the splicing module in layer 14, a target attention module is added between the skip connections of the C2f module in layer 6 and the splicing module in layer 11, and a target attention module is added between the skip connections of the SPPF module in layer 9 and the splicing module in layer 20. The upsampling module in the neck structure is replaced with a target upsampling module. The YOLOv8-seg model includes a backbone network and a neck structure.
[0035] Optionally, the working principle of the target attention module includes:
[0036] Obtain the original feature map, input the original feature map into the height attention module to obtain the height-weighted feature map, input the original feature map into the width attention module to obtain the width-weighted feature map, and input the original feature map into the channel attention module to obtain the channel-weighted feature map;
[0037] The height-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first height-weighted feature map. The height-weighted feature map is then input into the Conv module to obtain the second height-weighted feature map. The height-weighted feature map, the first height-weighted feature map, and the second height-weighted feature map are multiplied together to obtain the target height-weighted feature map.
[0038] The width-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first width-weighted feature map. The width-weighted feature map is then input into the Conv module to obtain the second width-weighted feature map. The width-weighted feature map, the first width-weighted feature map, and the second width-weighted feature map are multiplied together to obtain the target width-weighted feature map.
[0039] The channel-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first channel-weighted feature map. The channel-weighted feature map is then input into the Conv module to obtain the second channel-weighted feature map. The channel-weighted feature map, the first channel-weighted feature map, and the second channel-weighted feature map are multiplied together to obtain the target channel-weighted feature map.
[0040] The target height-weighted feature map, the target width-weighted feature map, and the target channel-weighted feature map are fused to obtain an output feature map, which is then used as the output of the target attention module.
[0041] Optionally, the operation of the target upsampling module includes:
[0042] An initial feature tensor is obtained, and the initial feature tensor is channel-compressed to obtain a first feature tensor. The first feature tensor is input into the kernel prediction module to obtain an upsampled kernel. The upsampled kernel is normalized to obtain a target upsampled kernel.
[0043] The initial feature tensor is upsampled to obtain a second feature tensor. The second feature tensor is fused with the target upsampling kernel to obtain an output feature tensor. The output feature tensor is used as the output of the target upsampling module.
[0044] Optionally, based on the detection results, anomaly detection is performed on the target area, including:
[0045] If the detection result is greater than the preset threshold, the target area is determined to be abnormal and a fault is reported.
[0046] If the detection result is not greater than the preset threshold, the target area is determined to be normal, and the inspection continues.
[0047] The beneficial effects of this invention are:
[0048] This invention proposes an AI-based visualization-based method for monitoring and inspecting power transmission lines. The method involves acquiring inspection images of the target area, preprocessing these images to obtain preprocessed inspection images, performing a stretching transformation on the preprocessed images to obtain a first enhanced image, applying adaptive dehazing to the first enhanced image to obtain a second enhanced image, and then performing deblurring on the second enhanced image to obtain a third enhanced image. A preset model is used to extract features from the third enhanced image to obtain detection results, and anomaly detection is performed on the target area based on these results. The stretching transformation enhances the texture details of the power transmission lines, the adaptive dehazing removes image occlusion, and the deblurring corrects lost details, providing high-quality images for power transmission line inspection. The preset model enables lightweight feature recognition, improving inspection efficiency and accuracy. Attached Figure Description
[0049] Figure 1 A flowchart of an AI-based visualization-based transmission line inspection and monitoring method is provided as an embodiment of the present invention;
[0050] Figure 2 This invention provides a schematic diagram of a model structure for an AI-based visualization-based transmission line inspection and monitoring method. Detailed Implementation
[0051] To further illustrate the technical means and effects of the present invention in achieving its intended purpose, the following detailed description of the specific implementation methods, structures, features, and effects of the present invention, in conjunction with the accompanying drawings and preferred embodiments, is provided below.
[0052] This invention provides an AI-based visual transmission line inspection and monitoring method. See also... Figure 1 , Figure 1 A flowchart illustrating an AI-based visualization-based transmission line inspection and monitoring method provided in this embodiment of the invention. The method includes the following steps:
[0053] S101, acquire the inspection image of the target area, preprocess the inspection image to obtain the preprocessed inspection image, and stretch the preprocessed inspection image to obtain the first enhanced image.
[0054] S102, perform an adaptive dehazing operation on the first enhanced image to obtain a second enhanced image, and perform a deblurring operation on the second enhanced image to obtain a third enhanced image;
[0055] S103, the detection results are obtained by extracting features from the third enhanced image through a preset model, and anomaly judgment is made in the target area based on the detection results.
[0056] The present invention provides an AI-based visualization method for power transmission line inspection and monitoring. This method enhances the texture details of the power transmission line through stretching transformation, removes image occlusion through adaptive dehazing, and corrects the loss of details through deblurring, providing high-quality images for power transmission line inspection. It also improves the efficiency and accuracy of inspection by performing lightweight feature recognition through a preset model.
[0057] In one implementation, stretching the image can enhance the features of the power transmission lines in the image, that is, make the texture details more obvious. After the image has undergone stretching, the occluded parts in the image are removed by dehazing, and then the image is enhanced by deblurring to correct the loss of details caused by dehazing.
[0058] In one implementation, the preprocessed inspection image is processed by stretching transformation, which can significantly enhance the texture details of the transmission line in the image, making the transmission line more prominent in the image; by adjusting the contrast and brightness of the image, the outline and structural features of the transmission line are highlighted.
[0059] In one implementation, an adaptive dehazing operation is performed on the stretched image to remove fog or other obscured parts from the image. The dehazing operation can improve the image's clarity and contrast, reduce fog obscuring key parts of the line, and prevent the inability to accurately determine the line's condition during inspection.
[0060] In one implementation, after the dehazing operation, performing a deblurring operation can further enhance the image quality and correct the loss of details that may be introduced during the dehazing process. Although the dehazing operation can remove occlusions, it may also cause some details of the image to become blurred. The deblurring operation enhances the above details again by optimizing the sharpness and clarity of the image, making the texture and structure of the transmission line clearer.
[0061] In one embodiment, stretching a preprocessed inspection image to obtain a first enhanced image includes:
[0062] The preprocessed inspection image is binarized to obtain a grayscale image, and the grayscale image is filtered to obtain a filtered image.
[0063] A two-dimensional Fourier transform is performed on the filtered image to obtain a complex matrix in the frequency domain, and the kernel function is determined. The kernel function is then updated to obtain the target kernel function.
[0064] The target kernel function is multiplied point-by-point with the frequency domain complex matrix to obtain the enhanced frequency domain complex matrix. The enhanced frequency domain complex matrix is then subjected to a two-dimensional inverse Fourier transform to obtain the spatial domain complex matrix. The first enhanced image is then calculated based on the spatial domain complex matrix.
[0065] In one implementation, the preprocessed inspection image is binarized to obtain a grayscale image. The preprocessed inspection image contains information such as complex background, noise, and power lines with weak visual features. First, binarization is performed to convert the color image into a grayscale image. By simplifying the color information of the image, the brightness features are preserved, the amount of data is reduced, and a foundation is provided for subsequent filtering, feature extraction and other processing.
[0066] In one implementation, a grayscale image is filtered to obtain a filtered image. A local smoothing low-pass filter is applied to the grayscale image to smooth high-frequency noise in the image, such as background and random noise, thereby reducing the roughness of the image, reducing interference in the subsequent feature extraction process, and making the electric field line contours in the image easier to identify.
[0067] In one implementation, a two-dimensional Fourier transform is performed on the filtered image to convert it from the spatial domain to the frequency domain, obtaining the frequency domain representation of the image (a frequency domain complex matrix); a pure amplitude stretching kernel function is introduced, for example: K(r,S,W)=S*(1+W*r) / (1+W*r) max) Where r is the polar radius in the frequency plane polar coordinates, S is the tensile strength parameter, and W is the adjustment parameter. max This is the maximum radius of the frequency plane. This kernel function enhances high-frequency components in the image, such as the edges of electric field lines and the amplitude information of curved structures, by adjusting the power K and the tensile strength S.
[0068] In one implementation, the target kernel function is multiplied by a frequency domain complex matrix to obtain an enhanced frequency domain complex matrix. A two-dimensional inverse Fourier transform is performed on the enhanced frequency domain complex matrix, and its angle is taken to generate an angle image (spatial domain complex matrix), for example: A(m,n)=atan2(b,a), where the imaginary part b is the y coordinate and the real part a is the x coordinate. The output is an angle value in the range [−π,π] or [0,2π), and A(m,n) is the first enhanced image. By stretching the amplitude field of the image, the edge features of linear targets such as electric field lines are highlighted, and their distinction from the background is improved.
[0069] In one embodiment, updating the kernel function to obtain the target kernel function includes:
[0070] The initial parameter combination is determined based on the kernel function. Multiple chromosomes are then randomly generated based on the initial parameter combination. The chromosomes are then combined to obtain the original population.
[0071] The initial population is obtained by initializing the original population. The initial population is then iteratively updated to obtain the target population. The process continues until a preset condition is met, at which point the optimal chromosome is output.
[0072] The optimal parameter combination is determined based on the optimal chromosome, and the kernel function is updated based on the optimal parameter combination to obtain the target kernel function.
[0073] In one implementation, the initial parameter combination [S,K] is globally searched through the above operations. The S value given by the optimal chromosome can achieve adaptive stretching intensity adjustment in the transmission line image: when the scene texture is rich and the contrast of the conductor edge is insufficient, the algorithm converges to S>1, and the high-frequency components are enhanced to strengthen the conductor outline; in high-noise conditions, such as backlight and haze, the algorithm tends to S<1, effectively suppressing the amplification of background noise, thereby obtaining the optimal data-driven balance between edge sharpening and noise control.
[0074] In one implementation, the K value corresponding to the optimal chromosome can be automatically matched to the spatial frequency distribution of key components of the transmission line after genetic iteration: a larger K makes the kernel function more concentrated in the high frequency band, which can enhance the texture details of small components such as vibration dampers and spacers; a smaller K expands and enhances the frequency band, which is suitable for large-scale icing or overall conductor morphology detection, avoiding the blindness of manually setting parameters, ensuring that the enhanced frequency band is always consistent with the feature scale of the current task target, thereby improving the accuracy of defect identification and increasing the efficiency of task optimization.
[0075] In one implementation, the update steps for the initial parameters mentioned above can be: metaheuristic algorithms such as genetic algorithms and particle swarm algorithms.
[0076] In one embodiment, performing an adaptive dehazing operation on a first enhanced image to obtain a second enhanced image includes:
[0077] The preprocessed inspection image is sequentially input into the convolutional layer and the feature attention layer to obtain the first dehazing feature map. The first dehazing feature map is sequentially input into the convolutional layer and the feature attention layer to obtain the second dehazing feature map. The second dehazing feature map is sequentially input into the convolutional layer and the feature attention layer to obtain the third dehazing feature map.
[0078] The third dehazing feature map is sequentially input into the convolutional layer, the feature attention layer, and the feature attention layer to obtain the fourth dehazing feature map. The first enhanced image, the fourth dehazing feature map, and the third dehazing feature map are then fused to obtain the target fourth dehazing feature map.
[0079] The target fourth dehazing feature map is sequentially input into the deconvolution layer and the feature attention layer to obtain the fifth dehazing feature map. The first enhanced image, the second dehazing feature map and the fifth dehazing feature map are then fused to obtain the target fifth dehazing feature map.
[0080] The fifth dehazed feature map of the target is sequentially input into the deconvolution layer and the feature attention layer to obtain the sixth dehazed feature map. The first enhanced image, the sixth dehazed feature map and the first dehazed feature map are fused to obtain the sixth dehazed feature map of the target. The sixth dehazed feature map of the target and the preprocessed inspection image are stitched together to obtain the second enhanced image.
[0081] In one implementation, the preprocessed inspection image is sequentially input into a convolutional layer and a feature attention layer to gradually obtain a first dehazing feature map, a second dehazing feature map, and a third dehazing feature map. By extracting key features from the image step by step, and with the feature attention layer focusing on feature regions that are more valuable for dehazing, the fog component in the image is removed. For example, in a transmission line image, the detailed features of the line itself and the fog features in the surrounding environment are effectively separated and processed, making the outline and details of the line clearer after dehazing.
[0082] In one implementation, the third dehazing feature map is processed again through a convolutional layer and two feature attention layers to obtain a fourth dehazing feature map. This fourth dehazing feature map is then fused with the first enhanced image and the third dehazing feature map to obtain the target fourth dehazing feature map. The above fusion operation can combine the feature advantages of different stages. The first enhanced image itself already has a certain image quality improvement effect. After fusion with the dehazing feature map, the details and overall quality of the image are further enhanced. For example, in the image of a transmission line, the details of the line insulators, hardware, etc. can be presented more clearly after fusion. At the same time, the contrast and color of the overall image are also optimized, which helps to more accurately identify the operating status and potential defects of the line.
[0083] In one implementation, the target fourth dehazing feature map is sequentially input into a deconvolution layer and a feature attention layer to obtain a fifth dehazing feature map. This fifth dehazing feature map is then fused with the first enhanced image and the second dehazing feature map to obtain the target fifth dehazing feature map. Finally, the target sixth dehazing feature map is obtained and stitched with the preprocessed inspection image to obtain the second enhanced image. The deconvolution layer can upsample the feature map to restore the image resolution, while the feature attention layer further optimizes the features in this process, enabling better preservation and enhancement of important information during image restoration. For example, in transmission line images, after the above operations, the image clarity and details can be restored, while enhancing features related to transmission line inspection, such as the line's direction and surrounding obstacles, thereby providing higher-quality image data support for intelligent inspection and fault diagnosis of transmission lines.
[0084] In one embodiment, the feature attention layer works by:
[0085] Obtain the original features, input the original features into 1×1Conv to obtain the first feature, input the original features into 3×3Conv to obtain the second feature, input the original features into ConvD2 to obtain the third feature, and input the original features into ConvD3 to obtain the fourth feature;
[0086] The first, second, third, and fourth features are concatenated to obtain the fifth feature. The fifth feature is then input into 1×1Conv and 3×3Conv sequentially to obtain the sixth feature. The sixth feature is then concatenated with the original features to obtain the seventh feature.
[0087] The seventh feature is sequentially input into the SE module and the 3×3Conv to obtain the eighth feature, which is then used as the output of the feature attention layer.
[0088] In one implementation, by inputting the original features into different types of convolutional layers (1×1 Conv, 3×3 Conv, ConvD2, ConvD3), features at different scales can be extracted. 1×1 convolutions are used for dimensionality reduction or expansion, balancing the number of channels; 3×3 convolutions extract local spatial features; dilated convolutions (ConvD2 and ConvD3) expand the receptive field through different dilation rates, capturing medium-scale and large-scale contextual information respectively. The fifth feature, obtained by concatenating these features, integrates local details and multi-scale contextual information, making the features richer and more comprehensive.
[0089] In one implementation, the fifth feature is processed through a 1×1 Conv to adjust its channels, and then further refined through a 3×3 Conv to obtain the sixth feature. The sixth feature is then concatenated with the original features and input into the SE module for channel attention optimization. The SE module learns the dependencies between channels, recalibrates the feature channels, suppresses unimportant channels, and enhances the response of key features. The seventh feature, optimized by the SE module, is then spatially integrated through a 3×3 convolution to obtain the eighth feature. This process, through the channel attention mechanism, strengthens key information in the feature map and suppresses irrelevant information, thereby improving the feature representation ability.
[0090] In one embodiment, performing a deblurring operation on the second enhanced image to obtain a third enhanced image includes:
[0091] The second enhanced image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the first deblurred image. The first deblurred image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the second deblurred image. The second deblurred image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the third deblurred image.
[0092] The third deblurred image is sequentially input into the convolutional layer, the target neural gradient layer and the convolutional layer to obtain the fourth deblurred image. The preprocessed inspection image, the fourth deblurred image and the third deblurred image are fused to obtain the target fourth deblurred image.
[0093] The fourth deblurred image of the target is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the fifth deblurred image. The preprocessed inspection image, the fifth deblurred image and the second deblurred image are fused to obtain the fifth deblurred image of the target.
[0094] The fifth deblurred image of the target is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the sixth deblurred image. The preprocessed inspection image, the sixth deblurred image and the first deblurred image are fused to obtain the sixth deblurred image of the target.
[0095] The sixth deblurred image of the target is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the seventh deblurred image, and the seventh deblurred image is used as the third enhancement image.
[0096] In one implementation, the second enhanced image is sequentially input into a convolutional layer and a target neural gradient layer to gradually obtain a first deblurred image, a second deblurred image, and a third deblurred image. The above process extracts image features through the convolutional layer and further optimizes gradient information by combining the target neural gradient layer, thereby gradually removing the blurry components in the image.
[0097] In one implementation, after obtaining the third deblurred image, it is sequentially input into a convolutional layer, a target neural gradient layer, and another convolutional layer to obtain a fourth deblurred image. This fourth deblurred image is then fused with the preprocessed inspection image and the third deblurred image to obtain the target fourth deblurred image. This fusion operation combines the feature advantages of different stages. The preprocessed inspection image provides background information of the original image, while the deblurred images at each stage provide optimized detail information. The fusion operation improves the overall quality and detail of the image. For example, in a transmission line image, the texture of the line and the details of the surrounding vegetation are clearer after fusion, while the contrast and color of the overall image are also optimized, which helps to more accurately identify the operating status and potential defects of the line.
[0098] In one implementation, the target attention layer is an attention mechanism module for image enhancement, such as a dual-path attention layer. This allows the model to adaptively adjust its attention to features in different spatial regions and channels, enhancing feature extraction capabilities. It comprises two attention paths. The first path performs average pooling and max pooling on all channels of the feature map—average pooling captures overall image features, while max pooling captures prominent features such as edges and textures; the two are complementary. Subsequently, attention weights are generated through convolution operations and the sigmoid activation function, and these weights are then applied to the feature map to complete feature weighting.
[0099] In one implementation, the target neural gradient layer is a neural network module with specific computational logic, such as a neural gradient algorithm layer. In this module, the total number of neural network layers, the number of neurons in each layer, and the input and output are all clearly defined. Input data is processed step-by-step through multiple layers of neurons. The output of each neuron is first obtained by weighted summation of the input from the previous layer and its corresponding weights, superimposed with a Gaussian perturbation that allows for optimized noise levels, and then processed by an activation function. The module measures the difference between the output and the label using a loss function. Based on the backpropagation mechanism, it first calculates the error term of the output layer, then defines the residuals layer by layer to estimate the gradient of the loss function with respect to the weights and the Gaussian perturbation. Subsequently, the network weights and noise levels are updated based on the gradient, and the process of forward calculation of output and loss, backpropagation estimation of gradient, and parameter update is repeated until the model parameters meet the requirements.
[0100] In one embodiment, the detection result is obtained by extracting features from the third enhanced image using a preset model. The preset model, based on improvements to YOLOv8-seg, includes:
[0101] In the YOLOv8-seg model, a target attention module is added between the skip connections of the C2f module in layer 4 and the splicing module in layer 14, between the C2f module in layer 6 and the splicing module in layer 11, and between the SPPF module in layer 9 and the splicing module in layer 20. The upsampling module in the neck structure is replaced with a target upsampling module. The YOLOv8-seg model includes a backbone network and a neck structure.
[0102] In one implementation, see [link to implementation details]. Figure 2 , Figure 2 This invention provides a schematic diagram of the model structure of an AI-based visualization-based transmission line inspection and monitoring method. A target attention module is added between the jump connections of the 4th layer C2f module and the 14th layer stitching module, the 6th layer C2f module and the 11th layer stitching module, and the 9th layer SPPF module and the 20th layer stitching module to enhance the transmission effect of transmission line features. The low-to-mid-level features output by the 4th and 6th layer C2f modules contain detailed information such as the edges and local textures of the transmission line, while the high-level features output by the 9th layer SPPF module contain semantic information such as the overall outline and spatial distribution of the line. The target attention module enables these features at different levels to highlight transmission line-related features, such as conductors and tower connection points, while suppressing background interference from trees and buildings. This makes the stitched features more focused on the line target, thereby improving the detection accuracy of transmission lines in complex scenarios, such as distinguishing between the line and the shadows of intersecting trees.
[0103] In one implementation, replacing the upsampling module in the neck structure with a target upsampling module can optimize the spatial reconstruction effect of transmission line features. The target upsampling module can combine the semantic information of line features, such as the continuity of conductors and the geometry of towers, to generate an adaptive upsampling kernel. When enlarging the feature map, it preserves the slender shape of the line, node connections and other details, avoiding edge blurring or shape distortion that may be caused by traditional upsampling. If the target upsampling module adopts a lightweight design, the amount of computation can be reduced while ensuring accuracy, so that the model can balance detection speed and detail capture capability when dealing with large-scale transmission line scenarios.
[0104] In one embodiment, the target attention module operates as follows:
[0105] Obtain the original feature map, input the original feature map into the height attention module to obtain the height-weighted feature map, input the original feature map into the width attention module to obtain the width-weighted feature map, and input the original feature map into the channel attention module to obtain the channel-weighted feature map;
[0106] The height-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first height-weighted feature map. The height-weighted feature map is then input into the Conv module to obtain the second height-weighted feature map. The height-weighted feature map, the first height-weighted feature map, and the second height-weighted feature map are multiplied together to obtain the target height-weighted feature map.
[0107] The width-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first width-weighted feature map. The width-weighted feature map is then input into the Conv module to obtain the second width-weighted feature map. The width-weighted feature map, the first width-weighted feature map, and the second width-weighted feature map are multiplied together to obtain the target width-weighted feature map.
[0108] The channel-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first channel-weighted feature map. The channel-weighted feature map is then input into the Conv module to obtain the second channel-weighted feature map. The channel-weighted feature map, the first channel-weighted feature map, and the second channel-weighted feature map are multiplied together to obtain the target channel-weighted feature map.
[0109] The target height-weighted feature map, target width-weighted feature map, and target channel-weighted feature map are fused to obtain the output feature map, which is then used as the output of the target attention module.
[0110] In one implementation, the original feature map is input into three independent attention modules: height, width, and channel. Each module focuses on the feature interaction in a single dimension, avoiding the information coupling problem caused by multi-dimensional mixed calculations in the original attention module. This allows each attention branch to capture key features of the transmission line in the vertical direction (height), horizontal direction (width), and spectral / semantic direction (channel), such as the spatial distribution characteristics of facilities like conductor sag or insulator strings, thereby improving the targeting of the features.
[0111] In one implementation, the weighted feature map for each dimension is processed using "Z-Pool+Conv" and "individual Conv": Z-Pool retains cross-dimensional interaction information by compressing dimensions, that is, the height-channel joint feature, while independent Conv strengthens the local features within the same dimension, that is, the continuity of the conductor in the height direction; feature map × first weight × second weight realizes dynamic weighting of cross-dimensional and local distribution, suppresses background noise of sky or vegetation texture, and enables the main features of the transmission line to obtain higher contrast in a single dimension.
[0112] In one implementation, by fusing target weighted feature maps of three dimensions, the complementary attention weights of different dimensions are achieved. Channel attention enhances the spectral features of conductor corrosion, while height / width attention constrains its spatial range. The output feature map enhances the distinction between the transmission line facilities and the background while maintaining the geometric integrity of the transmission line, providing a smoother edge transition and a more accurate mask prediction basis for the subsequent segmentation network.
[0113] In one embodiment, the operation of the target upsampling module includes:
[0114] Obtain the initial feature tensor, perform channel compression on the initial feature tensor to obtain the first feature tensor, input the first feature tensor into the kernel prediction module to obtain the upsampled kernel, and perform normalization processing on the upsampled kernel to obtain the target upsampled kernel.
[0115] The initial feature tensor is upsampled to obtain the second feature tensor. The second feature tensor is fused with the target upsampling kernel to obtain the output feature tensor. The output feature tensor is used as the output of the target upsampling module.
[0116] In one implementation, the initial feature tensor is channel compressed to obtain the first feature tensor. By reducing the number of feature channels, the amount of input data for the subsequent kernel prediction module can be reduced. Since the computational complexity of the kernel prediction module is positively correlated with the number of input feature channels, channel compression can reduce the computational complexity of the module, such as the number of multiplications and accumulations in the convolution operation and the required storage resources, thus achieving lightweight computation.
[0117] In one implementation, the upper-level sampling kernel is normalized, for example, by softmax processing, to ensure that the sum of the weights of the target upsampled kernel is 1, thus avoiding deviations in the fusion result caused by excessive differences in weight values. The initial feature tensor is upsampled to obtain a second feature tensor. The features are first mapped to the output size. By preserving the global structure and basic information of the initial features, for example, by using nearest neighbor upsampling to reduce early information loss, the feature integrity loss caused by kernel weighting is avoided.
[0118] In one implementation, the second feature tensor is fused with the target upsampling kernel. By weighting the local regions of the second feature tensor with the kernel, key local details, such as target edges and textures, are enhanced while preserving the basic features.
[0119] In one embodiment, anomaly detection of the target area based on the detection results includes:
[0120] If the detection result is greater than the preset threshold, the target area is determined to be abnormal and a fault is reported.
[0121] If the detection result is not greater than the preset threshold, the target area is determined to be normal, and the inspection continues.
[0122] In one implementation, by setting a preset threshold and comparing the detection result with the threshold, the state of the target area can be quickly determined. When the detection result is greater than the preset threshold, the system automatically determines that the target area is abnormal and reports a fault, and then continues the inspection; otherwise, it is determined to be normal and the inspection continues. The automated determination process reduces human intervention, avoids the subjectivity and uncertainty of human judgment, and improves detection accuracy and inspection efficiency.
[0123] In one implementation, a fixed threshold is used as a quantitative judgment standard, replacing the traditional subjective evaluation method. This avoids judgment bias caused by differences in human experience and ensures the consistency and objectivity of judgment logic in different inspection scenarios.
[0124] The above description is merely a preferred embodiment of the present invention and is not intended to limit the present invention in any way. Although the present invention has been disclosed above with reference to preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art can make some modifications or alterations to the above-disclosed technical content to create equivalent embodiments without departing from the scope of the present invention. Any simple modifications, equivalent changes and alterations made to the above embodiments based on the technical essence of the present invention without departing from the scope of the present invention shall still fall within the scope of the present invention.
Claims
1. A method for inspecting and monitoring transmission lines based on AI visualization, characterized in that, The method includes: Acquire an inspection image of the target area, preprocess the inspection image to obtain a preprocessed inspection image, and perform a stretching transformation on the preprocessed inspection image to obtain a first enhanced image; An adaptive dehazing operation is performed on the first enhanced image to obtain a second enhanced image, and a deblurring operation is performed on the second enhanced image to obtain a third enhanced image; The third enhanced image is subjected to feature extraction using a preset model to obtain detection results, and anomaly judgment is made on the target region based on the detection results.
2. The AI-based visual transmission line inspection and monitoring method according to claim 1, characterized in that, The preprocessed inspection image is stretched to obtain a first enhanced image, including: The preprocessed inspection image is binarized to obtain a grayscale image, and the grayscale image is filtered to obtain a filtered image. A two-dimensional Fourier transform is performed on the filtered image to obtain a complex matrix in the frequency domain, and a kernel function is determined. The kernel function is then updated to obtain the target kernel function. The target kernel function is multiplied point-by-point with the frequency domain complex matrix to obtain the enhanced frequency domain complex matrix. The enhanced frequency domain complex matrix is then subjected to a two-dimensional inverse Fourier transform to obtain a spatial domain complex matrix. The first enhanced image is then calculated based on the spatial domain complex matrix.
3. The AI-based visualization-based transmission line inspection and monitoring method according to claim 2, characterized in that, The target kernel function is obtained by updating the kernel function, including: The initial parameter combination is determined according to the kernel function, and multiple chromosomes are randomly generated according to the initial parameter combination. The chromosomes are then combined to obtain the original population. The original population is initialized to obtain an initial population. The initial population is then iteratively updated to obtain a target population. This process continues until a preset condition is met, at which point the optimal chromosome is output. The optimal parameter combination is determined based on the optimal chromosome, and the kernel function is updated based on the optimal parameter combination to obtain the target kernel function.
4. The AI-based visualization-based transmission line inspection and monitoring method according to claim 2, characterized in that, Performing an adaptive dehazing operation on the first enhanced image to obtain a second enhanced image includes: The preprocessed inspection image is sequentially input into a convolutional layer and a feature attention layer to obtain a first dehazed feature map. The first dehazed feature map is sequentially input into a convolutional layer and a feature attention layer to obtain a second dehazed feature map. The second dehazed feature map is sequentially input into a convolutional layer and a feature attention layer to obtain a third dehazed feature map. The third dehazing feature map is sequentially input into the convolutional layer, the feature attention layer, and the feature attention layer to obtain the fourth dehazing feature map. The first enhanced image, the fourth dehazing feature map, and the third dehazing feature map are fused to obtain the target fourth dehazing feature map. The target fourth dehazing feature map is sequentially input into the deconvolution layer and the feature attention layer to obtain the fifth dehazing feature map. The first enhanced image, the second dehazing feature map and the fifth dehazing feature map are fused to obtain the target fifth dehazing feature map. The target fifth dehazing feature map is sequentially input into the deconvolution layer and the feature attention layer to obtain the sixth dehazing feature map. The first enhanced image, the sixth dehazing feature map and the first dehazing feature map are fused to obtain the target sixth dehazing feature map. The target sixth dehazing feature map and the preprocessed inspection image are stitched together to obtain the second enhanced image.
5. The AI-based visualization-based transmission line inspection and monitoring method according to claim 4, characterized in that, The working principle of the feature attention layer includes: Obtain the original features, input the original features into 1×1Conv to obtain the first feature, input the original features into 3×3Conv to obtain the second feature, input the original features into ConvD2 to obtain the third feature, and input the original features into ConvD3 to obtain the fourth feature; The first feature, the second feature, the third feature, and the fourth feature are concatenated to obtain the fifth feature. The fifth feature is then sequentially input into 1×1Conv and 3×3Conv to obtain the sixth feature. The sixth feature and the original feature are then concatenated to obtain the seventh feature. The seventh feature is sequentially input into the SE module and the 3×3Conv to obtain the eighth feature, which is then used as the output of the feature attention layer.
6. The AI-based visualization-based transmission line inspection and monitoring method according to claim 1, characterized in that, Performing a deblurring operation on the second enhanced image to obtain a third enhanced image includes: The second enhanced image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the first deblurred image. The first deblurred image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the second deblurred image. The second deblurred image is sequentially input into the convolutional layer and the target neural gradient layer to obtain the third deblurred image. The third deblurred image is sequentially input into the convolutional layer, the target neural gradient layer, and the convolutional layer to obtain the fourth deblurred image. The preprocessed inspection image, the fourth deblurred image, and the third deblurred image are then fused to obtain the target fourth deblurred image. The target fourth deblurred image is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the fifth deblurred image. The preprocessed inspection image, the fifth deblurred image and the second deblurred image are fused to obtain the target fifth deblurred image. The target fifth deblurred image is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the sixth deblurred image. The preprocessed inspection image, the sixth deblurred image and the first deblurred image are fused to obtain the target sixth deblurred image. The sixth deblurred image of the target is sequentially input into the target attention layer, convolutional layer and target neural gradient layer to obtain the seventh deblurred image, and the seventh deblurred image is used as the third enhancement image.
7. The AI-based visualization-based transmission line inspection and monitoring method according to claim 1, characterized in that, The detection result is obtained by extracting features from the third enhanced image using a preset model. The preset model, based on improvements to YOLOv8-seg, includes: In the YOLOv8-seg model, a target attention module is added between the skip connections of the C2f module in layer 4 and the splicing module in layer 14, a target attention module is added between the skip connections of the C2f module in layer 6 and the splicing module in layer 11, and a target attention module is added between the skip connections of the SPPF module in layer 9 and the splicing module in layer 20. The upsampling module in the neck structure is replaced with a target upsampling module. The YOLOv8-seg model includes a backbone network and a neck structure.
8. The AI-based visualization-based transmission line inspection and monitoring method according to claim 1, characterized in that, The working principle of the target attention module includes: Obtain the original feature map, input the original feature map into the height attention module to obtain the height-weighted feature map, input the original feature map into the width attention module to obtain the width-weighted feature map, and input the original feature map into the channel attention module to obtain the channel-weighted feature map; The height-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first height-weighted feature map. The height-weighted feature map is then input into the Conv module to obtain the second height-weighted feature map. The height-weighted feature map, the first height-weighted feature map, and the second height-weighted feature map are multiplied together to obtain the target height-weighted feature map. The width-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first width-weighted feature map. The width-weighted feature map is then input into the Conv module to obtain the second width-weighted feature map. The width-weighted feature map, the first width-weighted feature map, and the second width-weighted feature map are multiplied together to obtain the target width-weighted feature map. The channel-weighted feature map is sequentially input into the Z-Pool pooling module and the Conv module to obtain the first channel-weighted feature map. The channel-weighted feature map is then input into the Conv module to obtain the second channel-weighted feature map. The channel-weighted feature map, the first channel-weighted feature map, and the second channel-weighted feature map are multiplied together to obtain the target channel-weighted feature map. The target height-weighted feature map, the target width-weighted feature map, and the target channel-weighted feature map are fused to obtain an output feature map, which is then used as the output of the target attention module.
9. The AI-based visualization-based transmission line inspection and monitoring method according to claim 1, characterized in that, The working process of the target upsampling module includes: An initial feature tensor is obtained, and the initial feature tensor is channel-compressed to obtain a first feature tensor. The first feature tensor is input into the kernel prediction module to obtain an upsampled kernel. The upsampled kernel is normalized to obtain a target upsampled kernel. The initial feature tensor is upsampled to obtain a second feature tensor. The second feature tensor is fused with the target upsampling kernel to obtain an output feature tensor. The output feature tensor is used as the output of the target upsampling module.
10. The AI-based visualization-based transmission line inspection and monitoring method according to claim 1, characterized in that, Based on the detection results, anomaly detection is performed on the target area, including: If the detection result is greater than the preset threshold, the target area is determined to be abnormal and a fault is reported. If the detection result is not greater than the preset threshold, the target area is determined to be normal, and the inspection continues.
Citation Information
Patent Citations
Defogging method and system for inspection image of power transmission line
CN118134803A