An unmanned aerial vehicle power inspection image intelligent analysis method and system based on deep learning and multi-modal fusion, and a medium thereof

By using a multi-rotor drone equipped with a multi-sensor system and deep learning multimodal fusion technology, the problems of multi-source data fusion and insufficient positioning accuracy in drone power line inspection have been solved, enabling efficient and accurate detection of power equipment defects and generation of reports.

CN120726041BActive Publication Date: 2025-11-11STATE GRID SICHUAN YAAN ELECTRIC POWER (GRP) CO LTD YUCHENG POWER SUPPLY CO +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202511203539.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-11-11
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing UAV power line inspection technologies suffer from problems such as insufficient multi-source data fusion, poor environmental adaptability, low automation, weak real-time performance, and insufficient positioning accuracy, making it difficult to achieve efficient and accurate power equipment defect detection.

Method used

A multi-rotor UAV equipped with a three-axis stabilized gimbal system is used, integrating a visible light camera, an infrared thermal imager, and a lidar. The path is planned using the A* algorithm, and multi-angle scanning and multi-sensor data acquisition are performed. Image processing and feature fusion are performed by combining deep learning and multimodal fusion technology. The ORB-SLAM3 algorithm is used to achieve sub-pixel level registration, and a multi-scale feature fusion network is constructed for defect detection and report generation.

Benefits of technology

It significantly improves the automation level and detection accuracy of power inspection, enabling the identification of power equipment defects in complex environments, achieving sub-pixel-level positioning, and generating inspection reports that include defect type, location, and maintenance suggestions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726041B_ABST
    Figure CN120726041B_ABST
Patent Text Reader

Abstract

This invention discloses an intelligent image analysis method, system, and medium for UAV power line inspection based on deep learning and multimodal fusion, relating to the field of power equipment inspection technology. The method includes: using the A* algorithm to plan the optimal inspection path and achieve synchronous data acquisition from multiple sensors; performing adaptive histogram equalization and dehazing on visible light images, non-uniformity correction and temperature calibration on infrared images, and filtering and registration on point cloud data; constructing a multi-scale feature fusion network based on an improved VGGNet-16, introducing deformable convolution and cross-modal attention mechanisms to achieve multi-source data fusion; performing defect detection based on a three-level template library and feature map cross-correlation algorithm, combined with non-maximum suppression and sub-pixel localization to improve accuracy; and finally generating an inspection report containing defect type, location, and maintenance suggestions. This invention significantly improves the automation level and detection accuracy of power line inspection.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent inspection technology for power equipment, specifically to an intelligent image analysis method, system, and medium for UAV power inspection based on deep learning and multimodal fusion, applicable to automated defect detection and condition assessment of power facilities such as overhead power lines and substation equipment. Background Technology

[0002] With the rapid development of smart grid construction, the requirements for the accuracy and efficiency of power equipment inspection are increasing. Traditional manual inspection methods suffer from low efficiency, high risk, and limited coverage, while drone inspection, due to its advantages of high flexibility, wide coverage, and low cost, is gradually becoming an important means of power equipment inspection. However, existing drone inspection technology still faces the following challenges:

[0003] Insufficient multi-source data fusion: Existing methods often rely on single sensors (such as visible light or infrared), making it difficult to comprehensively utilize multimodal data such as visible light, infrared thermal imaging, and laser point clouds, resulting in insufficient comprehensiveness and accuracy in defect detection. Poor environmental adaptability: Power equipment is often located in complex environments (such as changing lighting, smog, tree obstruction, etc.), and existing image processing methods show significant performance degradation under low contrast, blur, or noise interference, affecting defect identification. Low automation: Traditional inspections rely on manual route planning and manual adjustment of shooting parameters, and defect detection often uses manual interpretation or simple threshold segmentation methods, which are inefficient and highly subjective. Insufficient real-time performance: Existing algorithms have high computational complexity, making it difficult to achieve real-time processing on UAVs, resulting in delays in data transmission and processing, and failing to meet the needs of large-scale inspections. Insufficient sub-pixel level positioning accuracy: Defects in power equipment (such as broken insulators, broken conductor strands, etc.) are often tiny, and existing methods cannot achieve sub-pixel level precise positioning, affecting the accuracy of maintenance decisions.

[0004] To address the aforementioned issues, there is an urgent need for an intelligent analysis method based on deep learning and multimodal fusion. This method can improve the automation, accuracy, and real-time performance of UAV power line inspection by enhancing sensor data registration, multi-scale feature fusion, and adaptive defect detection technologies. Summary of the Invention

[0005] This invention proposes an intelligent analysis method, system, and medium for UAV power line inspection images based on deep learning and multimodal fusion, aiming to solve problems such as insufficient multi-source data fusion, poor environmental adaptability, low degree of automation, weak real-time performance, and insufficient positioning accuracy in existing technologies.

[0006] In a first aspect, embodiments of this application provide an intelligent analysis method for UAV power line inspection images based on deep learning and multimodal fusion, the method comprising:

[0007] S1. A multi-rotor UAV equipped with a three-axis stabilized gimbal system is used. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The optimal inspection path is planned through the A* algorithm to achieve multi-angle full-coverage scanning and multi-sensor synchronous data acquisition.

[0008] S2. Adaptive histogram equalization and dehazing are performed on the acquired visible light images, non-uniformity correction and temperature calibration are performed on the infrared images, outlier removal and voxel grid filtering are performed on the point cloud data, and the ORB-SLAM3 algorithm is used to achieve spatiotemporal synchronization and sub-pixel level registration of multi-source sensors.

[0009] S3. Construct a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing deformable convolutional modules, multi-level pyramid pooling structures, and cross-modal attention mechanisms, adaptive feature fusion of visible light, infrared, and point cloud data is achieved.

[0010] S4. Based on the constructed three-level template library, a feature map cross-correlation algorithm is used for multi-scale dynamic template matching. High-confidence defect areas are screened through non-maximum suppression and sub-pixel-level accurate localization is achieved.

[0011] S5. By combining structural similarity indicators, thermal anomaly analysis, and three-dimensional deformation measurement results, an inspection report is automatically generated that includes the defect type, precise location, and maintenance recommendations.

[0012] Optionally, in one implementation of the first aspect of the present invention, step S1 specifically includes:

[0013] S11. It adopts a six-rotor UAV platform and is equipped with a three-axis stabilized gimbal system. The gimbal system integrates a global shutter visible light camera, an uncooled infrared thermal imager, and a mechanical lidar. The three-axis stabilized gimbal system is driven by a high-precision servo motor.

[0014] S12. The A* path planning algorithm introduces a dynamic weight adjustment mechanism, which can automatically optimize the inspection path according to real-time environmental changes.

[0015] S13. Multi-sensor data acquisition synchronization is achieved through a hardware synchronization trigger module, and real-time data transmission is achieved using 5G wireless transmission technology.

[0016] Optionally, in one implementation of the first aspect of the present invention, step S2 specifically includes:

[0017] S21. Visible light image processing: A multi-scale adaptive histogram equalization algorithm based on Retinex theory is adopted to enhance the image through local contrast-limited processing; a dark channel prior dehazing algorithm is applied, which combines atmospheric scattering physical models and deep learning prior knowledge to perform image dehazing.

[0018] S22. Infrared Image Processing: Non-uniformity correction is performed using a two-point correction method based on a reference blackbody. The non-uniformity error of the corrected image is less than 0.1℃. A high-precision temperature calibration model is established based on Planck's blackbody radiation law to achieve accurate conversion of pixel grayscale values ​​to temperature values.

[0019] S23. Point cloud data processing: A statistical outlier detection algorithm based on k-nearest neighbors is adopted, and a distance threshold is dynamically calculated to filter out noise points; an adaptive resolution voxel grid filtering method is used to compress point cloud data while maintaining feature integrity.

[0020] S24. Multi-source sensor registration: The ORB-SLAM3 algorithm achieves accurate registration of multi-source data in the following ways: constructing a multimodal feature descriptor and fusing visual, infrared and geometric features; introducing IMU pre-integration constraints to optimize the initial pose estimation; and designing a joint optimization objective function to minimize reprojection error and geometric consistency error.

[0021] Optionally, in one implementation of the first aspect of the present invention, step S3 specifically includes:

[0022] S31. Network architecture improvements:

[0023] Deformable convolutional modules are inserted after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. The modules include: a deformable convolutional kernel group that supports dynamic receptive field adjustment from 3×3 to 7×7; a two-layer cascaded offset prediction network for generating feature sampling position offsets; and a bilinear interpolation feature resampling unit.

[0024] S32. Multi-scale feature extraction: Construct a multi-scale feature pyramid with four levels: Level 1: Extract low-level features using 3×3 standard convolution; Level 2: Combine dilated convolution with a dilation rate of 2 with 2×2 max pooling for mid-level feature fusion; Level 3: Achieve high-level feature aggregation through cross-layer skip connections; Level 4: Apply the SE module to implement channel attention weighting.

[0025] S33, Cross-modal feature fusion: Design a three-branch parallel feature extraction network: Visible light branch: extracts RGB spatial texture features through 5 layers of convolution; Infrared branch: extracts temperature distribution features using 4 layers of convolution; Point cloud branch: extracts 3D geometric features based on the PointNet++ architecture;

[0026] S34. Adaptive Feature Fusion: A three-level fusion strategy is adopted: Primary fusion: Unifying the feature dimension to 256 dimensions through 1×1 convolution; Intermediate fusion: Applying 4-level spatial pyramid pooling to extract multi-scale contextual information; Advanced fusion: Weighted concatenation of features based on an attention mechanism; The final output resolution is the original...Figure 1 / 8 Multiscale Multimodal Fusion Feature Map.

[0027] Optionally, in one implementation of the first aspect of the present invention, step S4 includes the following steps:

[0028] S41. Construct a three-level template library, including a standard template library, a local feature template library, and a dynamically updated template library;

[0029] S42. Employing a feature map cross-correlation algorithm for multi-scale dynamic template matching, specifically including:

[0030] S421. Scale-adaptive feature extraction based on VGGNet: For the input image and templates Extracting the first digit through a convolutional neural network Layer feature map and template feature map Among them, feeling the wild Satisfy constraints Where H and W represent the height and width of the image, respectively, C represents the number of color channels in the visible light image, and h and w represent the height and width of the template, respectively;

[0031] S422. Feature Map Cross-Correlation Calculation: The similarity matrix S is calculated using the cross-correlation formula.

[0032] ,

[0033] in, express In position place Subregion, For position coordinates, Denotes the Frobenius norm. These represent the height and width of the sampled feature map, respectively. The height and width of the template feature map , This involves extracting features and mapping the image to a matching feature space. This represents the improved cross-correlation function, where i, j, and k represent the index variables used to traverse the feature map data in the height, width, and channel directions, respectively.

[0034] S423, Multi-scale matching: Setting the set of scaling factors R and r are scaling factors, respectively. Number and its corresponding index, for each Execution steps S421 - S422 The matching results at each scale are unified by bilinear interpolation. Once back to the original resolution;

[0035] S43, Non-maximum suppression is used to screen candidate regions: multi-scale matching results are fused, and regions with IOU ≥ 1 are retained. Candidate regions with a confidence level of 0.7 and ranking in the top K;

[0036] S44 Subpixel-level precise positioning: center point of candidate region Perform Taylor expansion optimization:

[0037] ,

[0038] The sub-pixel level offset is determined by solving the eigenvectors of the Hessian matrix. Indicates subpixel level offset;

[0039] S45. Defect Area Verification and Template Library Update: Store the confirmed defect area features into the dynamically updated template library and periodically optimize the local feature template library.

[0040] Optionally, in one implementation of the first aspect of the present invention, the improved cross-correlation formula in step S422 introduces dynamic attention weights for each channel. This method can more effectively measure the similarity between feature maps, as shown in the following formula:

[0041]

[0042] in, By assigning different attention weights to each channel, the model can pay more attention to channel information relevant to the current task, thereby improving the accuracy of matching.

[0043] Optionally, in one implementation of the first aspect of the present invention, the non-maximum suppression in step S43 employs an adaptive threshold strategy, specifically as follows:

[0044] ,

[0045] in, and These are the mean and standard deviation of the scores for all candidate regions in the current image, respectively. It is an adjustable parameter. This is the non-maximum suppression value.

[0046] Optionally, in one implementation of the first aspect of the present invention, the sub-pixel level positioning in step S44 employs a quadratic surface fitting method, including:

[0047] At candidate point Establish a quadratic surface model within a 3×3 neighborhood;

[0048] The parameters of the quadratic surface model are solved by the least squares method, and a set of parameters is found that minimizes the sum of squared errors between the model's predicted similarity values ​​and the actual observed values.

[0049] The optimal offset is obtained by solving for the extreme points of the quadratic surface. ;

[0050] The calculated optimal offset Add to candidate points The above yields sub-pixel level precise positioning coordinates. .

[0051] Secondly, embodiments of this application provide an intelligent image analysis system for UAV power line inspection based on deep learning and multimodal fusion, applied to the intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion as described in the first aspect, the system comprising:

[0052] The data acquisition module is used to employ a multi-rotor UAV equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. It uses the A* algorithm to plan the optimal inspection path and achieve multi-angle full-coverage scanning and synchronous data acquisition from multiple sensors.

[0053] The data preprocessing module is used to perform adaptive histogram equalization and dehazing on the acquired visible light images, non-uniformity correction and temperature calibration on the infrared images, outlier removal and voxel grid filtering on the point cloud data, and to use the ORB-SLAM3 algorithm to achieve spatiotemporal synchronization and sub-pixel level registration of multi-source sensors.

[0054] The multimodal feature fusion module is used to construct a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing deformable convolution modules, multi-level pyramid pooling structures and cross-modal attention mechanisms, it achieves adaptive feature fusion of visible light, infrared and point cloud data.

[0055] The defect detection module is used to perform multi-scale dynamic template matching based on the constructed three-level template library and the feature map cross-correlation algorithm. It filters high-confidence defect areas through non-maximum suppression and achieves sub-pixel-level accurate positioning.

[0056] The report generation module is used to automatically generate an inspection report containing defect type, precise location, and maintenance recommendations by combining structural similarity indicators, thermal anomaly analysis, and three-dimensional deformation measurement results.

[0057] Thirdly, embodiments of this application provide an electronic device, including:

[0058] processor;

[0059] Memory used to store processor-executable instructions;

[0060] The processor is configured to implement the intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion as described in the first aspect when executing the instructions.

[0061] Fourthly, embodiments of this application provide a computer-readable storage medium storing a program that instructs a device to execute the intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion as described in the first aspect.

[0062] This invention discloses an intelligent image analysis method, system, and medium for power equipment inspection using unmanned aerial vehicles (UAVs) based on deep learning and multimodal fusion, relating to the field of power equipment inspection technology. The method includes: using a multi-rotor UAV equipped with a three-axis stabilized gimbal system integrating a visible light camera, an infrared thermal imager, and a lidar; employing the A* algorithm to plan the optimal inspection path and achieve simultaneous data acquisition from multiple sensors; performing adaptive histogram equalization and dehazing on the acquired visible light images; performing non-uniformity correction and temperature calibration on the infrared images; and filtering and registering the point cloud data; constructing a multi-scale feature fusion network based on an improved VGGNet-16, introducing deformable convolution and cross-modal attention mechanisms to achieve multi-source data fusion; performing defect detection based on a three-level template library and feature map cross-correlation algorithm, combined with non-maximum suppression and sub-pixel localization to improve accuracy; and finally generating an inspection report containing defect type, location, and maintenance suggestions. This invention significantly improves the automation level and detection accuracy of power equipment inspection and is suitable for defect identification in distribution network equipment under complex environments.

[0063] This technology integrates a visible light camera (texture), an infrared thermal imager (temperature), and a LiDAR (3D geometry) simultaneously to achieve data complementarity. It introduces deformable convolution and cross-modal attention mechanisms to dynamically adjust feature fusion weights, addressing the heterogeneity of different modal data (RGB, infrared, point cloud) and significantly improving the comprehensiveness and accuracy of defect detection. Adaptive image enhancement technology and a dynamic exposure adjustment mechanism ensure high-quality image data acquisition even in complex environments. Intelligent path planning and fully autonomous processing enable a highly automated inspection process. Lightweight algorithm design and an edge-cloud collaborative computing architecture guarantee real-time processing performance. Sub-pixel-level positioning technology achieves a positioning accuracy of ≤0.3 pixels, accurately identifying minute defects and automatically generating standardized reports containing defect type, location, and maintenance recommendations. This technology addresses the shortcomings of traditional inspection methods in multi-source data fusion, environmental adaptability, automation, and positioning accuracy, providing an efficient and reliable solution for intelligent inspection of power equipment. Attached Figure Description

[0064] Figure 1 This is a schematic diagram of the process for intelligent image analysis of UAV power line inspection based on deep learning and multimodal fusion, provided as an embodiment of this application.

[0065] Figure 2 The overall architecture of the ORB-SLAM3 algorithm provided in one embodiment of this application.

[0066] Figure 3 This is a diagram of a multi-scale feature fusion network structure based on the VGGNet-16 architecture provided in an embodiment of this application.

[0067] Figure 4 This application provides an embodiment of an intelligent image analysis system architecture for UAV power line inspection based on deep learning and multimodal fusion.

[0068] Figure 5 A schematic diagram of an electronic device provided in an embodiment of this application. Detailed Implementation

[0069] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them.

[0070] It should be noted that, in the embodiments of this application, "at least one" refers to one or more, and "more than one" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in the specification of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application.

[0071] It should be noted that in the embodiments of this application, the terms "first," "second," etc., are used only for descriptive purposes and should not be construed as indicating or implying relative importance, nor as indicating or implying order. Features specified as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of this application should not be construed as being more preferred or advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.

[0072] Based on the embodiments described in this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0073] Example 1.

[0074] Figure 1 This is a schematic diagram of a method for intelligent image analysis of power line inspection by UAV based on deep learning and multimodal fusion, provided as an embodiment of this application. The deep learning network architecture is the Evo-VGGNet architecture, which includes: a basic feature extraction layer (fixed), an evolvable feature fusion layer (dynamically adjustable), and an online knowledge distillation module.

[0075] like Figure 1 As shown, a method for intelligent analysis of UAV power line inspection images based on deep learning and multimodal fusion includes:

[0076] S1. A multi-rotor UAV equipped with a three-axis stabilized gimbal system is adopted. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The optimal inspection path is planned through the A* algorithm to achieve multi-angle full-coverage scanning and multi-sensor synchronous data acquisition. This design ensures that the drone can fly stably in complex terrain and acquire high-quality data. Multimodal data. A As a heuristic search algorithm, the algorithm can effectively reduce path planning time and improve inspection efficiency. Furthermore, the synchronous acquisition by multiple sensors provides a foundation for subsequent multimodal data fusion.

[0077] Specifically, in this embodiment, step S1 includes:

[0078] S11. Utilizing a hexacopter UAV platform, it is equipped with a three-axis stabilized gimbal system. This gimbal system integrates a global shutter visible light camera, an uncooled infrared thermal imager, and a mechanical lidar. The three-axis stabilized gimbal system is driven by high-precision servo motors. This design ensures that the UAV can fly stably in complex terrain and acquire high-quality multimodal data.

[0079] S12. The A* path planning algorithm introduces a dynamic weight adjustment mechanism, which can automatically optimize the inspection path according to real-time environmental changes. A As a heuristic search algorithm, the algorithm can effectively reduce path planning time and improve inspection efficiency.

[0080] S13. Multi-sensor data acquisition synchronization is achieved through a hardware synchronization trigger module, and real-time data transmission is achieved using 5G wireless transmission technology. The high bandwidth and low latency characteristics of 5G technology enable drones to transmit real-time high-definition images and videos, enhancing the real-time performance and accuracy of inspection data.

[0081] S2. Adaptive histogram equalization and dehazing are performed on the acquired visible light images, non-uniformity correction and temperature calibration are performed on the infrared images, outlier removal and voxel grid filtering are performed on the point cloud data, and the ORB-SLAM3 algorithm (Oriented Rapid and Brief Simultaneous Localization And Mapping 3, which combines the FAST corner detection ORB algorithm with SLAM synchronous localization and mapping technology) is used to achieve spatiotemporal synchronization and sub-pixel level registration of multi-source sensors.

[0082] In the data preprocessing stage, adaptive histogram equalization and dehazing are performed on visible light images to enhance image contrast and eliminate environmental interference; non-uniformity correction and temperature calibration are performed on infrared images to improve the accuracy of thermal imaging; outlier removal and voxel mesh filtering are performed on point cloud data to remove noise and retain effective structural information. Subsequently, the ORB-SLAM3 algorithm is used to achieve spatiotemporal synchronization and sub-pixel level registration of multi-source sensors. ORB-SLAM3 is a visual-inertial odometry-based SLAM algorithm that can achieve high-precision pose estimation and image registration, suitable for dynamic scenarios in UAV inspection. These preprocessing steps lay the foundation for subsequent multimodal feature fusion and defect detection.

[0083] Specifically, in this embodiment, step S2 includes:

[0084] S21. Visible Light Image Processing: A multi-scale adaptive histogram equalization algorithm based on Retinex theory is employed to enhance the image through local contrast-limited processing. A dark channel prior dehazing algorithm is applied, combining an atmospheric scattering physics model and deep learning prior knowledge for image dehazing. These methods effectively improve the contrast and clarity of visible light images, thereby enhancing the accuracy of subsequent analysis.

[0085] S22. Infrared Image Processing: Non-uniformity correction is performed using a two-point correction method based on a blackbody reference, resulting in a non-uniformity error of less than 0.1℃. A high-precision temperature calibration model is established based on Planck's blackbody radiation law to achieve accurate conversion between pixel grayscale values ​​and temperature values. These processing methods ensure the accuracy and reliability of thermal imaging data from infrared images, providing a foundation for subsequent thermal anomaly analysis.

[0086] S23. Point Cloud Data Processing: A k-nearest neighbor-based statistical outlier detection algorithm is employed, dynamically calculating distance thresholds to filter out noise points. An adaptive resolution voxel grid filtering method is used to compress point cloud data while maintaining feature integrity. These methods effectively remove noise from point cloud data and retain key structural information while compressing the data.

[0087] S24. Multi-source sensor registration: The ORB-SLAM3 algorithm achieves accurate registration of multi-source data by: constructing a multimodal feature descriptor that fuses visual, infrared, and geometric features; introducing IMU (Inertial Measurement Unit) pre-integration constraints to optimize the initial pose estimation; and designing a joint optimization objective function that minimizes both reprojection error and geometric consistency error. These methods effectively remove noise from point cloud data while preserving key structural information during data compression.

[0088] Figure 2 This application presents an embodiment of the ORB-SLAM3 algorithm architecture. The present invention makes several key improvements to the original ORB-SLAM3 framework, forming an enhanced version for multimodal inspection of power equipment. Its core structure can be divided into the following parts: a multimodal feature extraction layer, a spatiotemporal synchronization optimization module, and a sub-pixel level registration module.

[0089] The multimodal feature extraction layer includes cross-modal ORB features, which fuse visible light texture ORB points, infrared temperature gradient points, and LiDAR geometric feature points to improve the feature matching success rate under complex lighting conditions; a dynamic feature selector, which automatically adjusts feature weights based on environmental conditions (e.g., emphasizing LiDAR features in hazy weather), improving robustness and reducing trajectory loss rate; and a dedicated dictionary for power equipment, pre-trained with a visual vocabulary tree containing power equipment such as insulators and conductors, to improve relocalization speed.

[0090] Spatiotemporal synchronization optimization can introduce hardware synchronization triggers, such as 5G timestamp alignment, and add multimodal timing consistency constraints to address sensor sampling rate differences and reduce synchronization errors. Subpixel-level registration includes initial matching, fine optimization, and dynamic reference surface correction. Initial matching can employ an improved RANSAC algorithm: prior geometric constraints on power equipment (such as the spatial distribution of insulators). Fine optimization combines three-stage optimization with IMU pre-integration: such as 1. Visual BA → 2. LiDAR ICP → 3. Joint nonlinear optimization. Dynamic reference surface correction can construct local three-dimensional reference surfaces for power equipment in real time (such as conductor fitting planes) to eliminate vibration errors.

[0091] S3. Construct a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing deformable convolutional modules, multi-level pyramid pooling structures, and cross-modal attention mechanisms, adaptive feature fusion of visible light, infrared, and point cloud data is achieved.

[0092] Figure 3 This is a diagram of a multi-scale feature fusion network structure based on the VGGNet-16 architecture provided in one embodiment of this application. Figure 3 As shown, a multi-scale feature fusion network based on the VGGNet-16 architecture was constructed, introducing deformable convolutional modules, multi-level pyramid pooling structures, and cross-modal attention mechanisms to achieve adaptive feature fusion of visible light, infrared, and point cloud data. VGGNet-16 is a classic deep convolutional neural network with strong feature extraction capabilities. By introducing deformable convolutions and pyramid pooling structures, VGGNet-16 can better capture multi-scale features and achieve information interaction between different modalities through cross-modal attention mechanisms. This multi-modal feature fusion method can effectively improve the accuracy and robustness of defect detection.

[0093] The visible light branch is an improvement on VGGNet-16, inserting deformable convolutions after the Conv3 / Conv4 convolutional layers and adding an SE attention module to dynamically adjust the receptive field and enhance adaptability to the deformation of power equipment. The infrared branch includes 4 convolutional layers + dilated convolutions (dilation=2) to expand the receptive field and capture a wide range of correlation features in temperature distribution. The point cloud branch uses the PointNet++ (Deep Hierarchical Feature Learning on Point Sets in a MetricSpace) architecture + voxel mesh feature extraction to convert 3D point clouds into regularized features, adapting to 2D convolutional fusion. The deformable convolution module uses an offset prediction network, enabling the convolution kernel to adapt to changes in the shape of power equipment (such as twisted wires). In the cross-modal attention fusion module, Q / K / V features are feature projections from different modal branches, with infrared features dominating the temperature-sensitive area, visible light features dominating texture analysis, and point cloud features assisting in geometric verification.

[0094] The structure of the multi-scale pyramid (SPP) is shown in Table 1:

[0095] Table 1

[0096]

[0097] Specifically, in this embodiment, step S3 includes:

[0098] S31. Network Architecture Improvement: Deformable convolutional modules are inserted after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. These modules include: a deformable convolutional kernel group supporting dynamic receptive field adjustment from 3×3 to 7×7 to adapt to target feature extraction at different scales; a two-layer cascaded offset prediction network for generating feature sampling position offsets, thus achieving more flexible feature extraction; and a bilinear interpolation feature resampling unit for resampling the offset features, ensuring feature integrity and accuracy.

[0099] S32. Multi-scale feature extraction: Construct a multi-scale feature pyramid with four levels: Level 1: Use 3×3 standard convolution to extract low-level features, suitable for detecting small targets; Level 2: Combine dilated convolution with a dilation rate of 2 with 2×2 max pooling to fuse mid-level features, enhancing the ability to recognize medium-scale targets; Level 3: Achieve high-level feature aggregation through cross-layer skip connections, improving adaptability to complex backgrounds; Level 4: Apply the SE module to implement channel attention weighting, optimizing feature representation capabilities.

[0100] S33, Cross-modal feature fusion: Design a three-branch parallel feature extraction network: Visible light branch: Extracts RGB spatial texture features through 5 layers of convolution, suitable for target detection in regular images; Infrared branch: Extracts temperature distribution features using 4 layers of convolution, suitable for target detection at night or in low visibility environments; Point cloud branch: Extracts 3D geometric features based on the PointNet++ architecture, suitable for target detection in 3D space.

[0101] S34. Adaptive Feature Fusion: A three-level fusion strategy is adopted: Primary fusion: The feature dimension is unified to 256 dimensions through 1×1 convolution for easier subsequent processing; Intermediate fusion: 4-level spatial pyramid pooling is applied to extract multi-scale contextual information, enhancing adaptability to targets of different scales; Advanced fusion: Feature weighted concatenation based on an attention mechanism; The final output resolution is the original... Figure 1 The multi-scale, multi-modal fusion feature map of / 8 is used to highlight key features and suppress redundant information, resulting in a final output resolution equal to the original. Figure 1 / 8 Multiscale Multimodal Fusion Feature Map.

[0102] This embodiment improves the VGGNet-16 backbone network by introducing deformable convolutional modules, constructing a multi-scale feature pyramid, and combining multi-modal feature fusion and adaptive feature fusion strategies, which significantly improves the accuracy and robustness of target detection.

[0103] As shown in Table 2, the advantages compared to the traditional VGGNet-16 are:

[0104] Table 2

[0105]

[0106] S4. Based on the constructed three-level template library, a feature map cross-correlation algorithm is used for multi-scale dynamic template matching. High-confidence defect areas are screened through non-maximum suppression and sub-pixel-level accurate positioning is achieved.

[0107] Based on the constructed three-level template library, a feature map cross-correlation algorithm is used for multi-scale dynamic template matching. Non-maximum suppression (NMS) is employed to filter high-confidence defect regions and achieve sub-pixel-level precise localization. The feature map cross-correlation algorithm effectively matches local features in the image, while NMS removes redundant detection results, improving detection efficiency. Furthermore, the three-level template library design enables the system to adapt to different types of defects, improving its detection generalization ability.

[0108] Specifically, in this embodiment, step S4 includes the following steps:

[0109] S41. Construct a three-level template library, including a standard template library, a local feature template library, and a dynamically updated template library.

[0110] S42. Multi-scale dynamic template matching is performed using a feature map cross-correlation algorithm. Cross-correlation is a commonly used image matching method to calculate the similarity between two images. In this step, the cross-correlation formula is used to calculate the similarity matrix, which helps improve the efficiency and accuracy of matching. Specifically, it includes:

[0111] S421. Scale-adaptive feature extraction based on VGGNet: For the input image and templates Extracting the first digit through a convolutional neural network Layer feature map and template feature map Among them, feeling the wild Satisfy constraints Here, H and W represent the height and width of the image, respectively, C represents the number of color channels in the visible light image, and h and w represent the height and width of the template, respectively. VGGNet is a highly efficient deep convolutional neural network used for image recognition tasks. It extracts features by using multiple 3×3 convolutional layers, effectively capturing local features in the image. In this step, the feature maps extracted by VGGNet can be used in the subsequent matching process to improve matching accuracy.

[0112] S422. Feature Map Cross-Correlation Calculation: The similarity matrix S is calculated using the cross-correlation formula.

[0113] ,

[0114] in, express In position place Subregion, For position coordinates, Denotes the Frobenius norm. These represent the height and width of the sampled feature map, respectively. The height and width of the template feature map , This involves extracting features and mapping the image to a matching feature space. This represents the improved cross-correlation function, where i, j, and k represent the index variables used to traverse the feature map data in the height, width, and channel directions, respectively. Cross-correlation is an image matching method used to calculate the similarity between two images. In this step, the cross-correlation formula is used to calculate the similarity matrix, which helps improve the efficiency and accuracy of matching.

[0115] Specifically, in this embodiment, the improved cross-correlation formula in step S422 introduces dynamic attention weights for each channel. This method can more effectively measure the similarity between feature maps, as shown in the following formula:

[0116]

[0117] in, By assigning different attention weights to each channel, the model can focus more on channel information relevant to the current task, thereby improving matching accuracy. The calculation is performed in stages: first, local similarity is calculated in the spatial dimension (i,j), and then weighted fusion is performed in the channel dimension (k).

[0118] S423, Multi-scale matching: Setting the set of scaling factors R and r are scaling factors, respectively. Number and its corresponding index, for each Execution steps S421 - S422 The matching results at each scale are unified by bilinear interpolation. Once it reaches the original resolution. To accommodate targets at different scales, multi-scale matching is necessary. By setting different scaling factors and performing a matching step at each scale, it is ensured that the best match is found at each scale. Bilinear interpolation is used to unify the matching results from each scale to the original resolution, ensuring consistency and accuracy of the results.

[0119] S43, Non-maximum suppression is used to screen candidate regions: multi-scale matching results are fused, and regions with IOU ≥ 1 are retained. Candidate regions with a confidence level of 0.7 and ranking in the top K;

[0120] Specifically, in this embodiment, the non-maximum suppression in step S43 adopts an adaptive threshold strategy, and the specific formula is as follows:

[0121] ,

[0122] in, This is the non-maximum suppression value. and These are the mean and standard deviation of the scores for all candidate regions in the current image, respectively. This is an adjustable parameter. Non-Maximum Suppression (NMS) is a post-processing technique used to select the most likely region from multiple candidate regions. In this step, by retaining candidate regions with an IOU (Intersection of Union) ≥ 0.7 and a confidence level in the top K, redundant regions can be effectively reduced, improving detection efficiency and accuracy.

[0123] S44 Subpixel-level precise positioning: center point of candidate region Perform Taylor expansion optimization:

[0124] ,

[0125] The sub-pixel level offset is determined by solving the eigenvectors of the Hessian matrix. This represents the sub-pixel offset. Sub-pixel-level precise localization is a crucial step in improving detection accuracy. Optimizing the position of the candidate region's center point through Taylor expansion can further improve localization accuracy. Sub-pixel-level localization can be achieved by solving for the eigenvectors of the Hessian matrix to determine the offset, thus enabling more precise localization.

[0126] Specifically, in this embodiment, the sub-pixel level positioning in step S44 employs a quadratic surface fitting method, including:

[0127] At candidate point A quadratic surface model is established within a 3×3 neighborhood; the parameters of the quadratic surface model are solved using the least squares method to find a set of parameters that minimizes the sum of squared errors between the model's predicted similarity values ​​and the actual observed values; the optimal offset is obtained by solving for the extreme points of the quadratic surface. The calculated optimal offset Add to candidate points The above yields sub-pixel level precise positioning coordinates. .

[0128] S45. Defect Region Verification and Template Library Update: Confirmed defect region features are stored in a dynamically updated template library, and the local feature template library is periodically optimized. This method helps improve the adaptability and robustness of the template library, especially when dealing with complex scenarios and interfering factors.

[0129] S5. By combining structural similarity indicators, thermal anomaly analysis, and three-dimensional deformation measurement results, an inspection report is automatically generated that includes the defect type, precise location, and maintenance recommendations.

[0130] Among them, the structural similarity index can be used to compare the local structural differences between the current image and the standard template, as shown in the following formula:

[0131] ,

[0132] in, , This represents the mean of the x and y coordinates of the image, reflecting the overall brightness level of the image. , This represents the standard deviation of the image's x and y values. The covariance of images x and y is used to describe the linear correlation of the local structures of two images. , This represents the stability constant.

[0133] Thermal anomaly analysis can be performed using an infrared temperature calibration model, setting absolute temperature thresholds (e.g., abnormal wire joints >90℃) and relative temperature differences (ΔT >15K between adjacent phases).

[0134] In power line inspection, point cloud deformation is calculated by comparing the current scanned point cloud with a standard reference model. The core formula is:

[0135] ,

[0136] This represents the three-dimensional coordinates of the currently scanned point cloud, and can be used for real-time acquisition of device surface point cloud data by the drone's LiDAR. The three-dimensional coordinates representing the reference point cloud can be either the standard model of the device or point cloud data from its historical normal state. This represents the Euclidean norm (L2 distance), used to calculate the straight-line distance in three-dimensional space between corresponding points in two point clouds. This represents the scalar value of the deformation, and the final output is the magnitude of the deformation (positive values ​​indicate outward convexity, and negative values ​​indicate inward concavity).

[0137] Finally, by combining the Structural Similarity Index (SSIM), thermal anomaly analysis, and 3D deformation measurement results, an inspection report is automatically generated, containing the defect type, precise location, and maintenance recommendations. SSIM is used to assess image quality, thermal anomaly analysis to identify areas of abnormal temperature, and 3D deformation measurement to detect structural deformation. These multi-dimensional analysis results provide comprehensive defect information for power line inspections, helping maintenance personnel quickly locate problems and develop maintenance plans. Specifically, defect types can be generated by comparing a defect type decision tree with SSIM thresholds.

[0138] This invention achieves efficient and accurate analysis of UAV power line inspection images through multi-sensor data acquisition, multi-modal feature fusion, dynamic template matching, and automatic defect identification. This method not only improves the automation level of inspections but also significantly enhances the accuracy and efficiency of defect detection, providing strong support for the intelligent operation and maintenance of power systems.

[0139] Example 2.

[0140] like Figure 4 As shown, this application provides an architecture diagram of an intelligent analysis system for UAV power line inspection images based on deep learning and multimodal fusion, which is applied to the intelligent analysis system for UAV power line inspection images based on deep learning and multimodal fusion as described in Embodiment 1. It includes a data acquisition module 11, a data preprocessing module 12, a multimodal feature fusion module 13, a defect detection module 14, and a report generation module 15.

[0141] The data acquisition module 11 is used to employ a multi-rotor UAV equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. It plans the optimal inspection path through the A* algorithm to achieve multi-angle full-coverage scanning and synchronous data acquisition from multiple sensors.

[0142] The data preprocessing module 12 is used to perform adaptive histogram equalization and dehazing on the acquired visible light images, perform non-uniformity correction and temperature calibration on the infrared images, remove outliers and filter voxel grids on the point cloud data, and use the ORB-SLAM3 algorithm to achieve spatiotemporal synchronization and subpixel-level registration of multi-source sensors.

[0143] The multimodal feature fusion module 13 is used to construct a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing deformable convolution modules, multi-level pyramid pooling structures and cross-modal attention mechanisms, it achieves adaptive feature fusion of visible light, infrared and point cloud data.

[0144] The defect detection module 14 is used to perform multi-scale dynamic template matching based on the constructed three-level template library and the feature map cross-correlation algorithm. It filters high-confidence defect areas through non-maximum suppression and achieves sub-pixel-level accurate positioning.

[0145] The report generation module 15 is used to automatically generate an inspection report containing defect type, precise location, and maintenance recommendations by combining structural similarity index, thermal anomaly analysis, and three-dimensional deformation measurement results.

[0146] Figure 5 This is an electronic device provided in one embodiment of this application. For example... Figure 5As shown, the electronic device includes at least the following components: processor 101 and memory 100, communication interface 103, and bus 102.

[0147] In this embodiment of the application, memory 100 is used to store executable instructions of processor 101, which, when configured to execute instructions, implements the method as described in the first aspect.

[0148] In embodiments of this application, a computer-readable storage medium includes instructions that instruct a device to perform the method as described in the first aspect. For example, the instructions instruct the device to perform... Figure 1 The method is shown in the process steps.

[0149] In one embodiment of this application, the program operating in the electronic device may be a program that controls a central processing unit (CPU) or similar device to achieve the functions of the above-described embodiments of the present invention (a program that enables the computer to function). Information processed by these devices is then temporarily stored in random access memory (RAM) during processing, and subsequently stored in various ROMs such as read-only memory (FlashROM) and hard disk drives (HDDs), and read, corrected, and written by the CPU as needed.

[0150] It should be noted that a portion of the electronic device described above can also be implemented using a computer. In this case, the program for implementing the control function can be recorded on a computer-readable recording medium, and the program recorded on the recording medium can be read into the computer and executed.

[0151] It should be noted that the term "computer" as used here refers to a computer built into an electronic device, employing hardware including an operating system and peripheral devices. Furthermore, "computer-readable recording media" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard drives built into a computer.

[0152] Furthermore, a "computer-readable recording medium" can include: a medium that dynamically stores a program for a short period of time, such as a communication line used when transmitting a program via a network such as the Internet or a communication line such as a telephone line; or a medium that stores a program for a fixed period of time, such as volatile memory inside a computer that serves as a server or client in this case. In addition, the aforementioned program can be a program used to implement the above-mentioned functions, or it can be a program that can implement the above-mentioned functions by combining with programs already recorded in the computer.

[0153] Furthermore, the electronic device in the above embodiments can also be implemented as an assembly (device group) composed of multiple devices. Each device constituting the device group can possess some or all of the functions or functional blocks of the electronic device in the above embodiments. As a device group, it is sufficient to have all the functions or functional blocks of the electronic device.

[0154] Those skilled in the art should recognize that the above embodiments are only used to illustrate this application and are not intended to limit this application. Any appropriate changes and variations made to the above embodiments within the essential spirit and scope of this application fall within the scope of protection claimed in this application.

Claims

1. A method for intelligent analysis of UAV power line inspection images based on deep learning and multimodal fusion, characterized in that, The method includes: S1. A multi-rotor UAV equipped with a three-axis stabilized gimbal system is used. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The optimal inspection path is planned through the A* algorithm to achieve multi-angle full-coverage scanning and multi-sensor synchronous data acquisition. S2. Adaptive histogram equalization and dehazing are performed on the acquired visible light images, non-uniformity correction and temperature calibration are performed on the infrared images, outlier removal and voxel grid filtering are performed on the point cloud data, and the ORB-SLAM3 algorithm is used to achieve spatiotemporal synchronization and sub-pixel level registration of multi-source sensors. S3. Construct a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing deformable convolutional modules, multi-level pyramid pooling structures, and cross-modal attention mechanisms, adaptive feature fusion of visible light, infrared, and point cloud data is achieved. S4. Based on the constructed three-level template library, a feature map cross-correlation algorithm is used for multi-scale dynamic template matching. High-confidence defect areas are screened through non-maximum suppression and sub-pixel-level accurate localization is achieved. S5. By combining structural similarity indicators, thermal anomaly analysis, and three-dimensional deformation measurement results, an inspection report is automatically generated that includes the defect type, precise location, and maintenance recommendations. Step S2 specifically includes: S21. Visible light image processing: A multi-scale adaptive histogram equalization algorithm based on Retinex theory is adopted to enhance the image through local contrast-limited processing; a dark channel prior dehazing algorithm is applied, which combines atmospheric scattering physical models and deep learning prior knowledge to perform image dehazing. S22. Infrared Image Processing: Non-uniformity correction is performed using a two-point correction method based on a reference blackbody. The non-uniformity error of the corrected image is less than 0.1℃. A high-precision temperature calibration model is established based on Planck's blackbody radiation law to achieve accurate conversion of pixel grayscale values ​​to temperature values. S23. Point cloud data processing: A statistical outlier detection algorithm based on k-nearest neighbors is adopted, and the distance threshold is dynamically calculated to filter out noise points; A voxel grid filtering method with adaptive resolution is used to compress point cloud data while maintaining feature integrity. S24. Multi-source sensor registration: The ORB-SLAM3 algorithm achieves accurate registration of multi-source data in the following ways: constructing a multimodal feature descriptor and fusing visual, infrared and geometric features; introducing IMU pre-integration constraints to optimize the initial pose estimation; and designing a joint optimization objective function to minimize reprojection error and geometric consistency error. Step S3 specifically includes: S31. Network architecture improvements: Deformable convolutional modules are inserted after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. The modules include: a deformable convolutional kernel group that supports dynamic receptive field adjustment from 3×3 to 7×7; a two-layer cascaded offset prediction network for generating feature sampling position offsets; and a bilinear interpolation feature resampling unit. S32. Multi-scale feature extraction: Construct a multi-scale feature pyramid with four levels: Level 1: Extract low-level features using 3×3 standard convolution; Level 2: Combine dilated convolution with a dilation rate of 2 with 2×2 max pooling for mid-level feature fusion; Level 3: Achieve high-level feature aggregation through cross-layer skip connections; Level 4: Apply the SE module to implement channel attention weighting. S33, Cross-modal feature fusion: Design a three-branch parallel feature extraction network: Visible light branch: extracts RGB spatial texture features through 5 layers of convolution; Infrared branch: extracts temperature distribution features using 4 layers of convolution; Point cloud branch: extracts 3D geometric features based on the PointNet++ architecture; S34. Adaptive Feature Fusion: A three-level fusion strategy is adopted: Primary fusion: The feature dimension is unified to 256 dimensions through 1×1 convolution; Intermediate fusion: Multi-scale contextual information is extracted by applying 4-level spatial pyramid pooling; Advanced fusion: Feature weighted concatenation based on attention mechanism; The final output is a multi-scale, multi-modal fusion feature map with a resolution of 1 / 8 of the original image. S4 includes the following steps: S41. Construct a three-level template library, including a standard template library, a local feature template library, and a dynamically updated template library; S42. Employing a feature map cross-correlation algorithm for multi-scale dynamic template matching, specifically including: S421. Scale-adaptive feature extraction based on VGGNet: For the input image and templates Extracting the first digit through a convolutional neural network Layer feature map and template feature map Among them, feeling the wild Satisfy constraints Where H and W represent the height and width of the image, respectively, C represents the number of color channels in the visible light image, and h and w represent the height and width of the template, respectively; S422. Feature Map Cross-Correlation Calculation: The similarity matrix is ​​calculated using the cross-correlation formula. : , in, express In position place Subregion, For position coordinates, Denotes the Frobenius norm. These represent the height and width of the sampled feature map, respectively. The height and width of the template feature map , This involves extracting features and mapping the image to a matching feature space. This represents the improved cross-correlation function, where i, j, and k represent the index variables used to traverse the feature map data in the height, width, and channel directions, respectively. S423、 Multi-scale matching: Setting the set of scaling factors R and r are the scaling factor exponents and respectively. The corresponding index, for each Execution steps S421 - S422 Bilinear interpolation is used to unify the matching results at each scale to [a certain value]. Original resolution; S43、 Non-maximum suppression is used to screen candidate regions: multi-scale matching results are fused, retaining regions with IOU ≥ 0.7 and... Candidate regions with the highest confidence levels (K). S44、 Subpixel-level precise positioning: center point of candidate region Perform Taylor expansion optimization: , The sub-pixel level offset is determined by solving the eigenvectors of the Hessian matrix. Indicates subpixel level offset; S45. Defect Area Verification and Template Library Update: Store the confirmed defect area features into the dynamically updated template library and periodically optimize the local feature template library.

2. The intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion as described in claim 1, characterized in that, Step S1 specifically includes: S11. It adopts a six-rotor UAV platform and is equipped with a three-axis stabilized gimbal system. The gimbal system integrates a global shutter visible light camera, an uncooled infrared thermal imager, and a mechanical lidar. The three-axis stabilized gimbal system is driven by a high-precision servo motor. S12. The A* algorithm, by introducing a dynamic weight adjustment mechanism, can automatically optimize the inspection path according to real-time environmental changes. S13. Multi-sensor data acquisition synchronization is achieved through a hardware synchronization trigger module, and real-time data transmission is achieved using 5G wireless transmission technology.

3. The intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion according to claim 1, characterized in that, The improved cross-correlation formula in step S422 introduces dynamic attention weights for each channel. This method can more effectively measure the similarity between feature maps, as shown in the following formula: ; in, By assigning different attention weights to each channel, the model can pay more attention to channel information relevant to the current task, thereby improving the accuracy of matching.

4. The intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion according to claim 1, characterized in that, The non-maximum suppression in step S43 employs an adaptive threshold strategy, specifically as follows: , in, and These are the mean and standard deviation of the scores for all candidate regions in the current image, respectively. It is an adjustable parameter. This is the non-maximum suppression value.

5. The intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion according to claim 1, characterized in that, The sub-pixel level positioning in step S44 employs a quadratic surface fitting method, including: At candidate point Establish a quadratic surface model within a 3×3 neighborhood; The parameters of the quadratic surface model are solved by the least squares method, and a set of parameters is found that minimizes the sum of squared errors between the model's predicted similarity values ​​and the actual observed values. The optimal offset is obtained by solving for the extreme points of the quadratic surface. ; The calculated optimal offset Add to candidate points The above yields sub-pixel level precise positioning coordinates. .

6. A UAV power line inspection image intelligent analysis system based on deep learning and multimodal fusion, applied to the UAV power line inspection image intelligent analysis method based on deep learning and multimodal fusion as described in any one of claims 1 to 5, characterized in that, The system includes: The data acquisition module is used to employ a multi-rotor UAV equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. It uses the A* algorithm to plan the optimal inspection path and achieve multi-angle full-coverage scanning and synchronous data acquisition from multiple sensors. The data preprocessing module is used to perform adaptive histogram equalization and dehazing on the acquired visible light images, non-uniformity correction and temperature calibration on the infrared images, outlier removal and voxel grid filtering on the point cloud data, and to use the ORB-SLAM3 algorithm to achieve spatiotemporal synchronization and sub-pixel level registration of multi-source sensors. The multimodal feature fusion module is used to construct a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing deformable convolution modules, multi-level pyramid pooling structures and cross-modal attention mechanisms, it achieves adaptive feature fusion of visible light, infrared and point cloud data. The defect detection module is used to perform multi-scale dynamic template matching based on the constructed three-level template library and the feature map cross-correlation algorithm. It filters high-confidence defect areas through non-maximum suppression and achieves sub-pixel-level accurate positioning. The report generation module is used to automatically generate an inspection report containing defect type, precise location, and repair recommendations by combining structural similarity indexes, thermal anomaly analysis, and three-dimensional deformation measurement results. The data preprocessing module specifically includes: Visible light image processing module: It adopts a multi-scale adaptive histogram equalization algorithm based on Retinex theory to enhance the image through local contrast-limited processing; it applies a dark channel prior dehazing algorithm, combining atmospheric scattering physics model and deep learning prior knowledge to perform image dehazing. Infrared image processing module: It adopts a two-point correction method based on a reference blackbody for non-uniformity correction, and the non-uniformity error of the corrected image is less than 0.1℃; it establishes a high-precision temperature calibration model based on Planck's blackbody radiation law to achieve accurate conversion of pixel grayscale values ​​to temperature values. Point cloud data processing module: Employs a k-nearest neighbor-based statistical outlier detection algorithm and dynamically calculates distance thresholds to filter out noise points; A voxel grid filtering method with adaptive resolution is used to compress point cloud data while maintaining feature integrity. Multi-source sensor registration module: The ORB-SLAM3 algorithm achieves accurate registration of multi-source data in the following ways: constructing a multimodal feature descriptor and fusing visual, infrared and geometric features; introducing IMU pre-integration constraints to optimize the initial pose estimation; and designing a joint optimization objective function to minimize reprojection error and geometric consistency error. The multimodal feature fusion module specifically includes: Network architecture improvement module: Deformable convolutional modules are inserted after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. The modules include: a deformable convolutional kernel group that supports dynamic receptive field adjustment from 3×3 to 7×7; a two-layer cascaded offset prediction network for generating feature sampling position offsets; and a bilinear interpolation feature resampling unit. Multi-scale feature extraction module: Constructs a multi-scale feature pyramid with four levels: Level 1: Extracts low-level features using 3×3 standard convolution; Level 2: Combines dilated convolution with a dilation rate of 2 with 2×2 max pooling for mid-level feature fusion; Level 3: Aggregates high-level features through cross-layer skip connections; Level 4: Applyes the SE module to implement channel attention weighting. Cross-modal feature fusion module: Design a three-branch parallel feature extraction network: Visible light branch: extracts RGB spatial texture features through 5 layers of convolution; Infrared branch: extracts temperature distribution features using 4 layers of convolution; Point cloud branch: extracts 3D geometric features based on PointNet++ architecture; The adaptive feature fusion module employs a three-level fusion strategy: primary fusion: unifying the feature dimension to 256 dimensions through 1×1 convolution; intermediate fusion: applying 4-level spatial pyramid pooling to extract multi-scale contextual information; and advanced fusion: weighted concatenation of features based on an attention mechanism. The final output is a multi-scale, multi-modal fusion feature map with a resolution of 1 / 8 of the original image. The defect detection module includes: The three-level template library module includes a standard template library, a local feature template library, and a dynamically updated template library. The multi-scale dynamic template matching module employs a feature map cross-correlation algorithm for multi-scale dynamic template matching, specifically including: The adaptive feature extraction module, based on VGGNet's scale-adaptive feature extraction, is used for the input image. and templates Extracting the first digit through a convolutional neural network Layer feature map and template feature map Among them, feeling the wild Satisfy constraints Where H and W represent the height and width of the image, respectively, C represents the number of color channels in the visible light image, and h and w represent the height and width of the template, respectively; Feature map cross-correlation calculation module: Calculates the similarity matrix using the cross-correlation formula. : , in, express In position place Subregion, For position coordinates, Describing the Frobenius norm, These represent the height and width of the sampled feature map, respectively. The height and width of the template feature map , This involves extracting features and mapping the image to a matching feature space. This represents the improved cross-correlation function, where i, j, and k represent the index variables used to traverse the feature map data in the height, width, and channel directions, respectively. Multi-scale matching module: Sets the set of scaling factors R and r are the scaling factor exponents and respectively. The corresponding index, for each Execution steps S421 - S422 Bilinear interpolation is used to unify the matching results at each scale to [a certain value]. Original resolution; Candidate region filtering module: Employs non-maximum suppression to filter candidate regions; fuses multi-scale matching results to preserve... Retain candidate regions with IOU ≥ 0.7 and the top K confidence scores; Subpixel-level precise positioning module: for the center point of the candidate region Perform Taylor expansion optimization: , The sub-pixel level offset is determined by solving the eigenvectors of the Hessian matrix. Indicates subpixel level offset; Defect Area Verification and Template Library Update Module: Stores confirmed defect area features into a dynamically updated template library and periodically optimizes the local feature template library.

7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program that instructs the device to execute the intelligent image analysis method for UAV power line inspection based on deep learning and multimodal fusion as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Heterogenous image matching quality evaluation method and device based on GDB-ICP algorithm

    CN117456317A

  • Monorail crane inspection robot intelligent test method based on data analysis

    CN120429803A