Unmanned aerial vehicle electric power inspection image intelligent analysis method and system based on deep learning and multi-modal fusion and medium of unmanned aerial vehicle electric power inspection image intelligent analysis method and system

By using a multi-rotor drone equipped with a three-axis stabilized gimbal system and deep learning multimodal fusion technology, the problems of multi-source data fusion and insufficient positioning accuracy in drone power inspections have been solved, achieving efficient and accurate power equipment defect detection and automated report generation.

CN120726041AActive Publication Date: 2025-09-30STATE GRID SICHUAN YAAN ELECTRIC POWER (GRP) CO LTD YUCHENG POWER SUPPLY CO +1

Patent Information

Application Number
CN202511203539.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-08-27
Publication Date
2025-09-30
Estimated Expiration
2045-08-27

AI Technical Summary

Technical Problem

Existing drone power inspection technology has problems such as insufficient multi-source data fusion, poor environmental adaptability, low degree of automation, insufficient real-time performance and insufficient positioning accuracy, making it difficult to achieve efficient and accurate power equipment defect detection.

Method used

A multi-rotor drone equipped with a three-axis stabilized gimbal system integrates a visible light camera, infrared thermal imager and lidar. It uses the A* algorithm to plan the path, perform multi-angle scanning and multi-sensor data collection; combined with deep learning and multimodal fusion technology, it performs image processing and feature fusion to achieve sub-pixel precise positioning and defect detection.

Benefits of technology

It significantly improves the automation level and detection accuracy of power inspections, is suitable for identifying defects in distribution network equipment in complex environments, achieves sub-pixel positioning accuracy and real-time data processing, and generates inspection reports containing defect type, location, and repair recommendations.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120726041A_ABST
    Figure CN120726041A_ABST
Patent Text Reader

Abstract

The invention discloses an unmanned aerial vehicle electric power inspection image intelligent analysis method and system based on deep learning and multi-modal fusion and a medium thereof, and relates to the technical field of electric power equipment detection. The method comprises the following steps: planning an optimal inspection path by adopting an A * algorithm to realize multi-sensor synchronous data acquisition; adaptive histogram equalization and defogging processing are carried out on the visible light image, non-uniformity correction and temperature calibration are carried out on the infrared image, and filtering and registration are carried out on point cloud data; constructing a multi-scale feature fusion network based on improved VGGNet-16, and introducing deformable convolution and a cross-modal attention mechanism to realize multi-source data fusion; defect detection is carried out based on a three-level template library and a feature map cross-correlation algorithm, and the precision is improved in combination with non-maximum suppression and sub-pixel positioning; and finally generating a detection report containing defect types, positions and maintenance suggestions. According to the invention, the automation level and the detection precision of power inspection are obviously improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent inspection of power equipment, and specifically to a method, system and medium for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion, which is suitable for automated defect detection and status assessment of power facilities such as distribution network overhead lines and substation equipment. Background Art

[0002] With the rapid development of smart grid construction, the requirements for the accuracy and efficiency of power equipment inspections are increasing. Traditional manual inspection methods have problems such as low efficiency, high risk, and limited coverage. Drone inspections, due to their advantages such as strong flexibility, wide coverage, and low cost, are gradually becoming an important means of power equipment inspection. However, existing drone inspection technology still faces the following challenges: Inadequate multi-source data fusion: Existing methods often rely on a single sensor (such as visible light or infrared), making it difficult to comprehensively utilize multimodal data such as visible light, infrared thermal imaging, and laser point clouds. This results in insufficient comprehensiveness and accuracy in defect detection. Poor environmental adaptability: Power equipment often operates in complex environments (such as fluctuating lighting, haze, and tree obstruction). Existing image processing methods significantly degrade in low contrast, blur, or noise, impacting defect identification. Low automation: Traditional inspections rely on manual route planning and image parameter adjustment, and defect detection often relies on manual interpretation or simple threshold segmentation methods, which are inefficient and highly subjective. Inadequate real-time performance: Existing algorithms have high computational complexity, making real-time processing difficult on drones. This leads to delays in data transmission and processing, making them unsuitable for large-scale inspections. Inadequate sub-pixel positioning accuracy: Power equipment defects (such as broken insulators and broken conductors) are often tiny, making it difficult for existing methods to accurately locate them at the sub-pixel level, impacting the accuracy of maintenance decisions.

[0003] To address the above problems, there is an urgent need for an intelligent analysis method based on deep learning and multimodal fusion. By improving sensor data registration, multi-scale feature fusion and adaptive defect detection technology, the automation, accuracy and real-time performance of UAV power inspections can be improved. Summary of the Invention

[0004] The present invention proposes a method, system and medium for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion, aiming to solve the problems of insufficient multi-source data fusion, poor environmental adaptability, low degree of automation, weak real-time performance and insufficient positioning accuracy in the existing technology.

[0005] In a first aspect, an embodiment of the present application provides a method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion, the method comprising: S1. Use a multi-rotor drone equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The A* algorithm is used to plan the optimal inspection path, achieving multi-angle full coverage scanning and multi-sensor synchronous data collection. S2. Adaptive histogram equalization and dehazing are performed on the collected visible light images, non-uniformity correction and temperature calibration are performed on the infrared images, outlier removal and voxel grid filtering are performed on the point cloud data, and the ORB-SLAM3 algorithm is used to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors; S3. Build a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing a deformable convolution module, a multi-level pyramid pooling structure, and a cross-modal attention mechanism, it achieves adaptive feature fusion of visible light, infrared, and point cloud data. S4. Based on the constructed three-level template library, the feature map cross-correlation algorithm is used for multi-scale dynamic template matching. High-confidence defect areas are screened through non-maximum suppression to achieve sub-pixel precision positioning. S5. Combining structural similarity indicators, thermal anomaly analysis, and 3D deformation measurement results, an inspection report is automatically generated that includes defect type, precise location, and repair recommendations.

[0006] Optionally, in an implementation of the first aspect of the present invention, step S1 specifically includes: S11, using a six-rotor UAV platform equipped with a three-axis stabilized gimbal system that integrates a global shutter visible light camera, an uncooled infrared thermal imager, and a mechanical laser radar. The three-axis stabilized gimbal system is driven by a high-precision servo motor. S12. The A* path planning algorithm can automatically optimize the inspection path according to real-time environmental changes by introducing a dynamic weight adjustment mechanism; S13. Multi-sensor data acquisition synchronization is achieved through the hardware synchronization trigger module, and 5G wireless transmission technology is used to achieve real-time data backhaul.

[0007] Optionally, in an implementation of the first aspect of the present invention, step S2 specifically includes: S21. Visible light image processing: A multi-scale adaptive histogram equalization algorithm based on Retinex theory is used to achieve image enhancement through local contrast-limited processing. A dark channel prior dehazing algorithm is applied, combining the atmospheric scattering physical model and deep learning prior knowledge for image dehazing. S22. Infrared image processing: A two-point correction method based on a reference blackbody is used for non-uniformity correction. The non-uniformity error of the corrected image is less than 0.1°C. A high-precision temperature calibration model is established based on Planck's blackbody radiation law to achieve accurate conversion of pixel grayscale values ​​to temperature values. S23. Point cloud data processing: Using a statistical outlier detection algorithm based on k-nearest neighbors, dynamically calculating distance thresholds to filter out noise points; using an adaptive resolution voxel grid filtering method to achieve point cloud data compression while maintaining feature integrity; S24. Multi-source sensor registration: The ORB-SLAM3 algorithm achieves accurate multi-source data registration by: constructing a multimodal feature descriptor that integrates visual, infrared, and geometric features; introducing IMU pre-integration constraints to optimize the initial pose estimation; and designing a joint optimization objective function to simultaneously minimize the reprojection error and geometric consistency error.

[0008] Optionally, in an implementation of the first aspect of the present invention, step S3 specifically includes: S31. Network architecture improvements: A deformable convolution module is inserted after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. The module includes: a deformable convolution kernel group that supports dynamic receptive field adjustment from 3×3 to 7×7; a two-layer cascaded offset prediction network for generating feature sampling position offsets; and a bilinear interpolation feature resampling unit. S32, Multi-scale Feature Extraction: Construct a multi-scale feature pyramid with four levels: Level 1: Use 3×3 standard convolution to extract low-level features; Level 2: Combine dilated convolution with a dilation rate of 2 and 2×2 max pooling to fuse mid-level features; Level 3: Aggregate high-level features through cross-layer skip connections; Level 4: Apply the SE module to implement channel attention weighting; S33, Cross-modal Feature Fusion: Design a three-branch parallel feature extraction network: visible light branch: extracts RGB spatial texture features through 5-layer convolution; infrared branch: extracts temperature distribution features using 4-layer convolution; point cloud branch: extracts 3D geometric features based on the PointNet++ architecture; S34, Adaptive feature fusion: adopt a three-level fusion strategy: primary fusion: unify the feature dimension to 256 dimensions through 1×1 convolution; intermediate fusion: apply 4-level spatial pyramid pooling to extract multi-scale context information; high-level fusion: feature weighted splicing based on attention mechanism; the final output resolution is the original Figure 1 / 8 multi-scale multi-modal fusion feature map.

[0009] Optionally, in an implementation of the first aspect of the present invention, S4 includes the following steps: S41. Construct a three-level template library, including a standard template library, a local feature template library, and a dynamic update template library; S42. Using a feature map cross-correlation algorithm to perform multi-scale dynamic template matching, specifically including: S421, scale-adaptive feature extraction based on VGGNet: for the input image and templates , extract the first Layer feature map and template feature map , where the receptive field Satisfy constraints , where H and W represent the height and width of the image, respectively, C represents the number of color channels of the visible light image, and h and w represent the height and width of the template, respectively; S422, feature graph cross-correlation calculation: use the cross-correlation formula to calculate the similarity matrix S: , in, express In position at sub-region, is the position coordinate, represents the Frobenius norm, Represent the height and width of the sampling feature map respectively, The height and width of the template feature map, 、 Indicates feature extraction, mapping the image to a matching feature space, represents the improved cross-correlation function, i, j, k represent the index variables used to traverse the feature map data in the height direction, width direction, and channel direction respectively; S423, Multi-scale matching: setting a set of scaling factors , R and r are scaling factors respectively Number and corresponding index, for each Execution steps S421 - S422 , and the matching results of each scale are unified by bilinear interpolation One to native resolution; S43, Use non-maximum suppression to filter candidate regions: fuse multi-scale matching results and retain IOU≥ 0.7 and the top K candidate regions with confidence; S44, Sub-pixel precise positioning: the center point of the candidate area Perform Taylor expansion optimization: , The sub-pixel offset is determined by solving the eigenvector of the Hessian matrix. Indicates sub-pixel offset; S45. Defective area verification and template library update: store the confirmed defective area features into the dynamic update template library, and regularly optimize the local feature template library.

[0010] Optionally, in an implementation of the first aspect of the present invention, the improved cross-correlation formula in step S422 introduces dynamic channel attention weights : It can more effectively measure the similarity between feature maps. The formula is as follows:

[0011] in, ,By assigning different attention weights to each channel, the model can pay more attention to the channel information relevant to the current task, thereby improving the matching accuracy.

[0012] Optionally, in an implementation of the first aspect of the present invention, the non-maximum suppression in step S43 adopts an adaptive threshold strategy, and the specific formula is: , in, and are the mean and standard deviation of the scores of all candidate regions in the current image, is an adjustable parameter, is the non-maximum suppression value.

[0013] Optionally, in an implementation of the first aspect of the present invention, the sub-pixel positioning in step S44 adopts a quadratic surface fitting method, including: At the candidate point A quadratic surface model is established in the 3×3 neighborhood of ; Solving the parameters of the quadratic surface model by the least square method to find a set of parameters that minimizes the sum of squared errors between the similarity values ​​predicted by the model and the actual observed values; By solving the extreme points of the quadratic surface, the optimal offset is obtained ; The calculated optimal offset Add to candidate points Get the sub-pixel precise positioning coordinates .

[0014] In a second aspect, an embodiment of the present application provides a UAV power inspection image intelligent analysis system based on deep learning and multimodal fusion, which is applied to the UAV power inspection image intelligent analysis method based on deep learning and multimodal fusion as described in the first aspect. The system includes: The data acquisition module is used to use a multi-rotor drone equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The A* algorithm is used to plan the optimal inspection path, achieving multi-angle full coverage scanning and multi-sensor synchronous data acquisition; The data preprocessing module is used to perform adaptive histogram equalization and dehazing on the collected visible light images, perform non-uniformity correction and temperature calibration on the infrared images, remove outliers and perform voxel grid filtering on the point cloud data, and use the ORB-SLAM3 algorithm to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors; The multimodal feature fusion module is used to build a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing a deformable convolution module, a multi-level pyramid pooling structure, and a cross-modal attention mechanism, it achieves adaptive feature fusion of visible light, infrared, and point cloud data. The defect detection module is used to perform multi-scale dynamic template matching based on the constructed three-level template library, using the feature map cross-correlation algorithm to screen high-confidence defect areas through non-maximum suppression and achieve sub-pixel precision positioning; A report generation module combines structural similarity indicators, thermal anomaly analysis, and 3D deformation measurement results to automatically generate inspection reports that include defect type, precise location, and repair recommendations.

[0015] In a third aspect, an embodiment of the present application provides an electronic device, including: processor; a memory for storing processor-executable instructions; In which, the processor is configured to implement the UAV power inspection image intelligent analysis method based on deep learning and multimodal fusion as described in the first aspect when executing the instructions.

[0016] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, which stores a program, and the program instructs a device to execute the intelligent analysis method of drone power inspection images based on deep learning and multimodal fusion as described in the first aspect.

[0017] The present invention discloses a method, system and medium for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion, and relates to the technical field of power equipment detection. The method comprises: using a multi-rotor UAV equipped with a three-axis stabilized gimbal system that integrates a visible light camera, an infrared thermal imager and a lidar, and using the A* algorithm to plan the optimal inspection path to achieve multi-sensor synchronous data acquisition; performing adaptive histogram equalization and defogging on the collected visible light images, performing non-uniformity correction and temperature calibration on the infrared images, and filtering and aligning the point cloud data; constructing a multi-scale feature fusion network based on an improved VGGNet-16, introducing deformable convolution and cross-modal attention mechanisms to achieve multi-source data fusion; performing defect detection based on a three-level template library and feature map cross-correlation algorithm, and combining non-maximum suppression and sub-pixel positioning to improve accuracy; and finally generating a test report containing defect type, location and maintenance recommendations. The present invention significantly improves the automation level and detection accuracy of power inspections, and is suitable for defect identification of distribution network equipment in complex environments.

[0018] This technology integrates a visible light camera (texture), an infrared thermal imager (temperature), and a lidar (3D geometry) to achieve data complementarity. It introduces deformable convolution and a cross-modal attention mechanism to dynamically adjust feature fusion weights, addressing the heterogeneity of data from different modalities (RGB, infrared, and point cloud), significantly improving the comprehensiveness and accuracy of defect detection. Adaptive image enhancement technology and a dynamic exposure adjustment mechanism ensure high-quality image data acquisition even in complex environments. Intelligent path planning and fully autonomous processing enable a highly automated inspection process. A lightweight algorithm design and an end-to-end collaborative computing architecture ensure real-time processing performance. Sub-pixel positioning technology achieves positioning accuracy of ≤0.3 pixels, accurately identifying tiny defects and automatically generating standardized reports containing defect type, location, and repair recommendations. This technology addresses the shortcomings of traditional inspection methods in multi-source data fusion, environmental adaptability, automation, and positioning accuracy, providing an efficient and reliable solution for intelligent inspection of power equipment. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] Figure 1 A flowchart of a method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion is provided in accordance with one embodiment of the present application.

[0020] Figure 2 The overall architecture of the ORB-SLAM3 algorithm provided in one embodiment of the present application.

[0021] Figure 3 This is a diagram of the multi-scale feature fusion network structure based on the VGGNet-16 architecture provided in one embodiment of the present application.

[0022] Figure 4An architecture diagram of an intelligent analysis system for drone power inspection images based on deep learning and multimodal fusion is provided in accordance with one embodiment of the present application.

[0023] Figure 5 A schematic diagram of an electronic device provided in one embodiment of the present application. DETAILED DESCRIPTION

[0024] The technical solutions in the embodiments of the present application will be described clearly and completely below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments.

[0025] It should be noted that, in the embodiments of the present application, "at least one" refers to one or more, and "more" refers to two or more. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art in the art to which this application relates. The terms used in the specification of this application are only for the purpose of describing specific embodiments and are not intended to limit this application.

[0026] It should be noted that, in the embodiments of the present application, words such as "first" and "second" are only used for the purpose of distinguishing descriptions, and cannot be understood as indicating or implying relative importance, nor can they be understood as indicating or implying an order. Features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations or explanations. Any embodiment or design scheme described as "exemplary" or "for example" in the embodiments of the present application should not be interpreted as being more preferred or more advantageous than other embodiments or design schemes. Specifically, the use of words such as "exemplary" or "for example" is intended to present related concepts in a concrete way.

[0027] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.

[0028] Example 1.

[0029] Figure 1 This is a flowchart of a method for intelligent analysis of drone-generated power inspection images based on deep learning and multimodal fusion, according to one embodiment of the present application. The deep learning network architecture is the Evo-VGGNet architecture, comprising a fixed basic feature extraction layer, an evolvable feature fusion layer (dynamically adjusted), and an online knowledge distillation module.

[0030] like Figure 1 As shown in FIG, a method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion includes: S1. Use a multi-rotor drone equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The A* algorithm is used to plan the optimal inspection path, achieving multi-angle full coverage scanning and multi-sensor synchronous data collection. This design ensures that the UAV can fly stably and obtain high-quality Multimodal data. As a heuristic search algorithm, the algorithm can effectively reduce the time of path planning and improve inspection efficiency. In addition, the synchronous acquisition of multiple sensors provides a basis for subsequent multimodal data fusion.

[0031] Specifically, in this embodiment, step S1 specifically includes: The S11 utilizes a six-rotor drone platform equipped with a three-axis stabilized gimbal system. This system integrates a global shutter visible light camera, an uncooled infrared thermal imager, and a mechanical lidar. Driven by high-precision servo motors, this design ensures stable flight in complex terrain and captures high-quality multimodal data.

[0032] S12. The A* path planning algorithm can automatically optimize the inspection path according to real-time environmental changes by introducing a dynamic weight adjustment mechanism. A As a heuristic search algorithm, the algorithm can effectively reduce the time of path planning and improve inspection efficiency.

[0033] S13. Synchronize multi-sensor data collection through a hardware synchronization trigger module, and use 5G wireless transmission technology for real-time data transmission. The high bandwidth and low latency of 5G technology enable drones to transmit real-time high-definition images and videos, enhancing the real-time and accuracy of inspection data.

[0034] S2. Adaptive histogram equalization and dehazing are performed on the collected visible light images. Infrared images are corrected for non-uniformity and temperature calibrated. Point cloud data is filtered for outliers and filtered using a voxel grid. The ORB-SLAM3 algorithm (Oriented Rapid and Brief Simultaneous Localization And Mapping 3, combining the FAST corner detection ORB algorithm with SLAM simultaneous positioning and mapping technology) is used to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors.

[0035] During the data preprocessing stage, adaptive histogram equalization and dehazing are performed on visible light images to enhance image contrast and eliminate environmental interference. Infrared images undergo non-uniformity correction and temperature calibration to improve thermal imaging accuracy. Point cloud data undergoes outlier removal and voxel grid filtering to remove noise and preserve effective structural information. Subsequently, the ORB-SLAM3 algorithm is used to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors. ORB-SLAM3 is a visual-inertial odometry-based SLAM algorithm that enables high-precision pose estimation and image registration, making it suitable for dynamic scenes used in drone inspections. These preprocessing steps lay the foundation for subsequent multimodal feature fusion and defect detection.

[0036] Specifically, in this embodiment, step S2 specifically includes: S21. Visible light image processing: A multi-scale adaptive histogram equalization algorithm based on Retinex theory is used to enhance image quality through local contrast restriction. A dark channel prior dehazing algorithm is used, combined with a physical model of atmospheric scattering and deep learning prior knowledge, to dehaze images. A multi-scale adaptive histogram equalization algorithm based on Retinex theory is used to enhance image quality through local contrast restriction. A dark channel prior dehazing algorithm is used, combined with a physical model of atmospheric scattering and deep learning prior knowledge, to dehaze images. These methods can effectively enhance the contrast and clarity of visible light images, thereby improving the accuracy of subsequent analysis.

[0037] S22. Infrared Image Processing: A two-point calibration method based on a reference blackbody is used for non-uniformity correction. The non-uniformity error of the corrected image is less than 0.1°C. A high-precision temperature calibration model based on Planck's blackbody radiation law is established to accurately convert pixel grayscale values ​​to temperature values. These processing methods ensure accurate and reliable thermal imaging data for infrared images, providing a foundation for subsequent thermal anomaly analysis.

[0038] S23. Point cloud data processing: A statistical outlier detection algorithm based on k-nearest neighbors is used to dynamically calculate distance thresholds to filter out noise points. A voxel grid filtering method with adaptive resolution is used to compress point cloud data while maintaining feature integrity. These methods effectively remove noise from point cloud data and preserve key structural information while compressing the data.

[0039] S24. Multi-source sensor registration: The ORB-SLAM3 algorithm achieves accurate multi-source data registration through the following methods: constructing a multimodal feature descriptor that integrates visual, infrared, and geometric features; optimizing the initial pose estimate using IMU (Inertial Measurement Unit) pre-integration constraints; and designing a joint optimization objective function that simultaneously minimizes reprojection error and geometric consistency error. These methods effectively remove noise from point cloud data and compress the data while preserving key structural information.

[0040] Figure 2 This is the overall architecture of the ORB-SLAM3 algorithm provided in one embodiment of this application. This invention makes several key improvements to the original ORB-SLAM3 framework, creating an enhanced version for multimodal inspection of power equipment. Its core architecture is divided into the following components: a multimodal feature extraction layer, a spatiotemporal synchronization optimization module, and a sub-pixel registration module.

[0041] Among them, the multimodal feature extraction layer includes cross-modal ORB features, which are used to fuse visible light texture ORB points + infrared temperature gradient points + lidar geometric feature points to improve the feature matching success rate under complex lighting conditions; a dynamic feature selector is used to automatically adjust feature weights based on environmental conditions (such as focusing on lidar features in haze weather), improve robustness, and reduce track loss rate; a special dictionary for power equipment, pre-training a visual vocabulary tree containing power equipment such as insulators and wires, is used to improve repositioning speed.

[0042] Spatiotemporal synchronization optimization can introduce hardware synchronization triggers, such as 5G timestamp alignment, and add multimodal timing consistency constraints to address sensor sampling rate differences and reduce synchronization errors. Subpixel registration includes initial matching, fine-tuning, and dynamic reference plane correction. Initial matching can utilize a modified RANSAC algorithm with geometric prior constraints on power equipment (such as the spatial distribution of insulators). Fine-tuning combines a three-stage optimization process with IMU pre-integration: 1. Visual BA → 2. LiDAR ICP → 3. Joint nonlinear optimization. Dynamic reference plane correction constructs a local 3D reference plane for power equipment (such as a conductor fitting plane) in real time to eliminate vibration errors.

[0043] S3. Build a multi-scale feature fusion network based on the VGGNet-16 architecture, and realize adaptive feature fusion of visible light, infrared and point cloud data by introducing deformable convolution modules, multi-level pyramid pooling structure and cross-modal attention mechanism.

[0044] Figure 3 This is a diagram of a multi-scale feature fusion network structure based on the VGGNet-16 architecture provided in one embodiment of the present application. Figure 3As shown in the figure, a multi-scale feature fusion network based on the VGGNet-16 architecture was constructed, incorporating a deformable convolution module, a multi-level pyramid pooling structure, and a cross-modal attention mechanism to achieve adaptive feature fusion of visible light, infrared, and point cloud data. VGGNet-16 is a classic deep convolutional neural network with strong feature extraction capabilities. By introducing deformable convolution and pyramid pooling, VGGNet-16 can better capture multi-scale features and enable information exchange between different modalities through a cross-modal attention mechanism. This multimodal feature fusion approach can effectively improve the accuracy and robustness of defect detection.

[0045] The visible light branch is based on improvements to VGGNet-16, inserting deformable convolutions after the Conv3 / Conv4 layers and adding a SE attention module to dynamically adjust the receptive field and enhance adaptability to deformation of power equipment. The infrared branch includes four convolutional layers plus a dilated convolution (dilation=2) to expand the receptive field and capture wide-range correlated features of temperature distribution. The point cloud branch uses the PointNet++ (Deep Hierarchical Feature Learning on Point Sets in a MetricSpace) architecture and voxel grid feature extraction to convert 3D point clouds into regularized features suitable for 2D convolution fusion. The deformable convolution module uses an offset prediction network to enable the convolution kernel to adapt to shape changes in power equipment (such as twisted wires). The Q / K / V features in the cross-modal attention fusion module project features from different modal branches. Infrared features dominate temperature-sensitive areas, while visible light features lead texture analysis. Point cloud features assist in geometry verification.

[0046] The multi-scale pyramid (SPP) structure is shown in Table 1: Table 1

[0047] Specifically, in this embodiment, step S3 specifically includes: S31. Network Architecture Improvements: Deformable convolution modules are added after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. These modules include: a deformable convolution kernel group that supports dynamic receptive field adjustment from 3×3 to 7×7 to accommodate target feature extraction at different scales; a two-layer cascaded offset prediction network that generates feature sampling position offsets, enabling more flexible feature extraction; and a bilinear interpolation feature resampling unit that resamples the offset features to ensure feature integrity and accuracy.

[0048] S32. Multi-scale feature extraction: Construct a multi-scale feature pyramid with four levels: First level: Use 3×3 standard convolution to extract underlying features, which is suitable for detecting small targets; Second level: Combine the dilated convolution with a dilation rate of 2 with 2×2 maximum pooling to fuse mid-level features to enhance the recognition ability of medium-scale targets; Third level: Achieve high-level feature aggregation through cross-layer jump connections to improve adaptability to complex backgrounds; Fourth level: Apply the SE module to implement channel attention weighting and optimize feature expression capabilities.

[0049] S33. Cross-modal feature fusion: Design a three-branch parallel feature extraction network: Visible light branch: Extracts RGB spatial texture features through 5-layer convolution, which is suitable for target detection in conventional images; Infrared branch: Uses 4-layer convolution to extract temperature distribution features, which is suitable for target detection at night or in low-visibility environments; Point cloud branch: Extracts 3D geometric features based on the PointNet++ architecture, which is suitable for target detection in 3D space.

[0050] S34, Adaptive feature fusion: adopt a three-level fusion strategy: primary fusion: unify the feature dimension to 256 dimensions through 1×1 convolution to facilitate subsequent processing; intermediate fusion: apply 4-level spatial pyramid pooling to extract multi-scale context information and enhance adaptability to targets of different scales; advanced fusion: feature weighted splicing based on attention mechanism; the final output resolution is the original Figure 1 / 8 multi-scale multi-modal fusion feature map, highlighting key features, suppressing redundant information, and the final output resolution is the original Figure 1 / 8 multi-scale multi-modal fusion feature map.

[0051] This embodiment significantly improves the accuracy and robustness of target detection by improving the VGGNet-16 backbone network, introducing a deformable convolution module, constructing a multi-scale feature pyramid, and combining multimodal feature fusion with adaptive feature fusion strategies.

[0052] As shown in Table 2, the advantages compared with the traditional VGGNet-16 are: Table 2

[0053] S4. Based on the constructed three-level template library, the feature map cross-correlation algorithm is used for multi-scale dynamic template matching, and high-confidence defect areas are screened through non-maximum suppression to achieve sub-pixel precise positioning.

[0054] Based on a constructed three-level template library, the system employs a feature map cross-correlation algorithm for multi-scale dynamic template matching. Non-maximum suppression (NMS) is used to screen high-confidence defect areas and achieve sub-pixel precision localization. The feature map cross-correlation algorithm effectively matches local features in the image, while NMS removes redundant detection results, improving detection efficiency. Furthermore, the three-level template library design enables the system to adapt to different defect types, enhancing detection generalization capabilities.

[0055] Specifically, in this embodiment, S4 includes the following steps: S41. Construct a three-level template library, including a standard template library, a local feature template library and a dynamic update template library.

[0056] S42. Perform multi-scale dynamic template matching using a feature graph cross-correlation algorithm. Cross-correlation is a commonly used image matching method used to calculate the similarity between two images. In this step, a cross-correlation formula is used to calculate the similarity matrix, which helps improve the efficiency and accuracy of matching. Specifically, the following steps are performed: S421, scale-adaptive feature extraction based on VGGNet: for the input image and templates , extract the first Layer feature map and template feature map , where the receptive field Satisfy constraints , where H and W represent the height and width of the image, respectively; C represents the number of color channels in the visible light image; and h and w represent the height and width of the template, respectively. VGGNet is a highly effective deep convolutional neural network for image recognition tasks. It uses multiple 3×3 convolutional layers to extract features, effectively capturing local features in the image. In this step, the feature maps extracted by VGGNet can be used in the subsequent matching process to improve matching accuracy.

[0057] S422, feature graph cross-correlation calculation: use the cross-correlation formula to calculate the similarity matrix S: , in, express In position at sub-region, is the position coordinate, represents the Frobenius norm, Represent the height and width of the sampling feature map respectively, The height and width of the template feature map, 、 Indicates feature extraction, mapping the image to a matching feature space, represents the improved cross-correlation function, where i, j, and k represent the index variables used to traverse the feature map data in the height, width, and channel directions, respectively. Cross-correlation is an image matching method used to calculate the similarity between two images. In this step, the cross-correlation formula is used to calculate the similarity matrix, which helps improve matching efficiency and accuracy.

[0058] Specifically, in this embodiment, the improved cross-correlation formula in step S422 introduces dynamic channel attention weights : It can more effectively measure the similarity between feature maps. The formula is as follows:

[0059] in, By assigning different attention weights to each channel, the model can pay more attention to the channel information related to the current task, thereby improving the matching accuracy. The calculation is divided into stages, first calculating the local similarity in the spatial dimension (i, j), and then performing weighted fusion in the channel dimension (k).

[0060] S423, Multi-scale matching: setting a set of scaling factors , R and r are scaling factors respectively Number and corresponding index, for each Execution steps S421 - S422 , and the matching results of each scale are unified by bilinear interpolation One to native resolution. Multi-scale matching is necessary to accommodate objects of varying scales. By setting different scaling factors and performing the matching step for each scale, we ensure that the best match is found at each scale. Bilinear interpolation is used to unify the matching results at each scale to the original resolution, ensuring consistency and accuracy.

[0061] S43, Use non-maximum suppression to filter candidate regions: fuse multi-scale matching results and retain IOU≥ 0.7 and the top K candidate regions with confidence; Specifically, in this embodiment, the non-maximum suppression in step S43 adopts an adaptive threshold strategy, and the specific formula is: , in, is the non-maximum suppression value, and are the mean and standard deviation of the scores of all candidate regions in the current image, is an adjustable parameter. Non-Maximum Suppression (NMS) is a post-processing technique used to select the most likely region from multiple candidate regions. In this step, by retaining candidate regions with an Intersection of Union (IOU) ≥ 0.7 and a confidence score in the top K, redundant regions can be effectively reduced, improving detection efficiency and accuracy.

[0062] S44, Sub-pixel precise positioning: the center point of the candidate area Perform Taylor expansion optimization: , The sub-pixel offset is determined by solving the eigenvector of the Hessian matrix. Represents the sub-pixel offset. Sub-pixel precise positioning is an important step in improving detection accuracy. Optimizing the center point of the candidate region using Taylor expansion can further improve positioning accuracy. Sub-pixel positioning can determine the offset by solving the eigenvectors of the Hessian matrix, achieving more precise positioning.

[0063] Specifically, in this embodiment, the sub-pixel positioning in step S44 adopts a quadratic surface fitting method, including: At the candidate point A quadratic surface model is established in a 3×3 neighborhood of ; the parameters of the quadratic surface model are solved by the least squares method to find a set of parameters that minimizes the sum of squared errors between the similarity values ​​predicted by the model and the actual observed values; the optimal offset is obtained by solving the extreme points of the quadratic surface ; The calculated optimal offset Add to candidate points Get the sub-pixel precise positioning coordinates .

[0064] S45. Defect Area Verification and Template Library Update: Confirmed defect area features are stored in the dynamically updated template library, and the local feature template library is regularly optimized. Confirmed defect area features are stored in the dynamically updated template library, and the local feature template library is regularly optimized. This approach helps improve the adaptability and robustness of the template library, especially when dealing with complex scenarios and interference factors.

[0065] S5. Combining structural similarity indicators, thermal anomaly analysis, and 3D deformation measurement results, an inspection report is automatically generated that includes defect type, precise location, and repair recommendations.

[0066] Among them, the structural similarity index can be combined to compare the local structural differences between the current image and the standard template. The formula is as follows: , in, 、 Represents the mean value of image x,y, reflecting the overall brightness level of the image, 、 Represents the standard deviation of image x,y, Represents the covariance of image x, y, which is used to describe the linear correlation of the local structures of the two images. 、 represents the stability constant.

[0067] Thermal anomaly analysis can be performed based on an infrared temperature calibration model, setting absolute temperature thresholds (such as wire joints >90°C abnormality) and relative temperature differences (ΔT between adjacent phases >15K).

[0068] In power inspection, point cloud deformation calculation is achieved by comparing the difference between the current scanned point cloud and the standard reference model. The core formula is: , Indicates the three-dimensional coordinates of the current scan point cloud, which can be the device surface point cloud data collected in real time by the UAV lidar. Indicates the three-dimensional coordinates of the reference point cloud, which can be the point cloud data of the equipment standard model or historical normal state. Represents the Euclidean norm (L2 distance), which is used to calculate the three-dimensional straight-line distance between corresponding points of two point clouds. Represents the scalar value of the deformation variable, the final output deformation size (positive value is convex, negative value is concave).

[0069] Finally, the system automatically generates an inspection report that includes the defect type, precise location, and repair recommendations by combining the Structural Similarity Index (SSIM), thermal anomaly analysis, and 3D deformation measurement results. SSIM is used to assess image quality, thermal anomaly analysis identifies areas of temperature anomalies, and 3D deformation measurement detects structural deformation. These multi-dimensional analysis results provide comprehensive defect information for power inspections, helping operators quickly locate problems and develop repair plans. Specifically, defect types can be generated by comparing them with SSIM thresholds using a defect type decision tree.

[0070] This invention achieves efficient and accurate analysis of drone-generated power inspection images through multi-sensor data acquisition, multimodal feature fusion, dynamic template matching, and automatic defect identification. This method not only improves the automation level of inspections but also significantly enhances the accuracy and efficiency of defect detection, providing strong support for intelligent power system operation and maintenance.

[0071] Example 2.

[0072] like Figure 4As shown, the present application provides an architecture diagram of an UAV power inspection image intelligent analysis system based on deep learning and multimodal fusion, which is applied to the UAV power inspection image intelligent analysis system based on deep learning and multimodal fusion as described in Example 1, including a data acquisition module 11, a data preprocessing module 12, a multimodal feature fusion module 13, a defect detection module 14, and a report generation module 15.

[0073] The data acquisition module 11 is used to use a multi-rotor drone equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager and a lidar, and plans the optimal inspection path through the A* algorithm to achieve multi-angle full coverage scanning and multi-sensor synchronous data acquisition.

[0074] The data preprocessing module 12 is used to perform adaptive histogram equalization and dehazing on the collected visible light images, perform non-uniformity correction and temperature calibration on the infrared images, remove outliers and perform voxel grid filtering on the point cloud data, and use the ORB-SLAM3 algorithm to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors.

[0075] The multimodal feature fusion module 13 is used to build a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing a deformable convolution module, a multi-level pyramid pooling structure and a cross-modal attention mechanism, it realizes the adaptive feature fusion of visible light, infrared and point cloud data.

[0076] The defect detection module 14 is used to perform multi-scale dynamic template matching based on the constructed three-level template library, adopt the feature map cross-correlation algorithm, screen high-confidence defect areas through non-maximum suppression, and achieve sub-pixel precise positioning.

[0077] The report generation module 15 is used to automatically generate an inspection report including defect type, precise location and repair suggestions by combining the structural similarity index, thermal anomaly analysis and three-dimensional deformation measurement results.

[0078] Figure 5 This is an electronic device provided by an embodiment of the present application. Figure 5 As shown, the electronic device includes at least the following parts: a processor 101 and a memory 100 , a communication interface 103 , and a bus 102 .

[0079] In an embodiment of the present application, the memory 100 is used to store instructions executable by the processor 101, and the processor 101 is configured to implement the method of the first aspect when executing the instructions.

[0080] In an embodiment of the present application, a computer-readable storage medium includes instructions, and the instructions instruct a device to execute the method of the first aspect. For example, the instructions instruct the device to execute Figure 1 The method is shown in the process steps.

[0081] The program running in the electronic device involved in one embodiment of the present application may be a program that controls a central processing unit (CPU) and the like to implement the functions of the above-mentioned embodiment involved in one embodiment of the present invention (a program that causes a computer to function). The information processed by these devices is temporarily stored in random access memory (RAM) while being processed, and is then stored in various ROMs such as read-only memory (Flash ROM) and a hard disk drive (HDD), where it is read, modified, and written as needed by the CPU.

[0082] It should be noted that a portion of the electronic device of the above embodiment may also be implemented by a computer. In this case, a program for implementing the control function may be recorded on a computer-readable recording medium, and the program recorded on the recording medium may be read into a computer and executed.

[0083] It should be noted that the "computer" mentioned here refers to a computer built into an electronic device, employing hardware including an operating system (OS) and peripheral devices. Furthermore, "computer-readable recording medium" refers to removable media such as floppy disks, magneto-optical disks, ROMs, and CD-ROMs, as well as storage devices such as hard disks built into computers.

[0084] Furthermore, "computer-readable recording media" may include: media that dynamically store programs for a short period of time, such as communication lines when transmitting programs via networks such as the Internet or communication lines such as telephone lines; and media that store programs for a fixed period of time, such as volatile memory within computers acting as servers or clients in this context. Furthermore, the aforementioned program may be a program for implementing a portion of the aforementioned functions, or a program that can achieve the aforementioned functions by combining with a program already stored in a computer.

[0085] Furthermore, the electronic device in the above-described embodiments can also be implemented as a collection of multiple devices (a device group). Each device comprising the device group may include some or all of the functions or functional blocks of the electronic device in the above-described embodiments. A device group only needs to include all of the functions or functional blocks of the electronic device.

[0086] Those skilled in the art should recognize that the above embodiments are merely intended to illustrate the present application and are not intended to limit the present application. As long as they are within the spirit of the present application, appropriate changes and modifications to the above embodiments are within the scope of protection claimed in the present application.

Claims

1. An intelligent analysis method for UAV power inspection images based on deep learning and multimodal fusion, characterized in that: The method comprises: S1. Use a multi-rotor drone equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The A* algorithm is used to plan the optimal inspection path, achieving multi-angle full coverage scanning and multi-sensor synchronous data collection. S2. Adaptive histogram equalization and dehazing are performed on the collected visible light images, non-uniformity correction and temperature calibration are performed on the infrared images, outlier removal and voxel grid filtering are performed on the point cloud data, and the ORB-SLAM3 algorithm is used to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors; S3. Build a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing a deformable convolution module, a multi-level pyramid pooling structure, and a cross-modal attention mechanism, it achieves adaptive feature fusion of visible light, infrared, and point cloud data. S4. Based on the constructed three-level template library, the feature map cross-correlation algorithm is used for multi-scale dynamic template matching. High-confidence defect areas are screened through non-maximum suppression to achieve sub-pixel precision positioning. S5. Combining structural similarity indicators, thermal anomaly analysis, and 3D deformation measurement results, an inspection report is automatically generated that includes defect type, precise location, and repair recommendations.

2. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 1 is characterized in that: The step S1 specifically includes: S11, using a six-rotor UAV platform equipped with a three-axis stabilized gimbal system that integrates a global shutter visible light camera, an uncooled infrared thermal imager, and a mechanical laser radar. The three-axis stabilized gimbal system is driven by a high-precision servo motor. S12. The A* algorithm can automatically optimize the inspection path according to real-time environmental changes by introducing a dynamic weight adjustment mechanism; S13. Multi-sensor data acquisition synchronization is achieved through the hardware synchronization trigger module, and 5G wireless transmission technology is used to achieve real-time data backhaul.

3. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 1 is characterized in that: The step S2 specifically includes: S21. Visible light image processing: A multi-scale adaptive histogram equalization algorithm based on Retinex theory is used to achieve image enhancement through local contrast-limited processing. A dark channel prior dehazing algorithm is applied, combining the atmospheric scattering physical model and deep learning prior knowledge for image dehazing. S22. Infrared image processing: A two-point correction method based on a reference blackbody is used for non-uniformity correction. The non-uniformity error of the corrected image is less than 0.1°C. A high-precision temperature calibration model is established based on Planck's blackbody radiation law to achieve accurate conversion of pixel grayscale values ​​to temperature values. S23, point cloud data processing: using a statistical outlier detection algorithm based on k-nearest neighbors, dynamically calculating the distance threshold to filter out noise points; Use voxel grid filtering with adaptive resolution to compress point cloud data while maintaining feature integrity; S24. Multi-source sensor registration: The ORB-SLAM3 algorithm achieves accurate multi-source data registration by: constructing a multimodal feature descriptor that integrates visual, infrared, and geometric features; introducing IMU pre-integration constraints to optimize the initial pose estimation; and designing a joint optimization objective function to simultaneously minimize the reprojection error and geometric consistency error.

4. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 2 is characterized in that: The step S3 specifically includes: S31. Network architecture improvements: A deformable convolution module is inserted after the Conv3 and Conv4 layers of the VGGNet-16 backbone network. The module includes: a deformable convolution kernel group that supports dynamic receptive field adjustment from 3×3 to 7×7; a two-layer cascaded offset prediction network for generating feature sampling position offsets; and a bilinear interpolation feature resampling unit. S32, Multi-scale Feature Extraction: Construct a multi-scale feature pyramid with four levels: Level 1: Use 3×3 standard convolution to extract low-level features; Level 2: Combine dilated convolution with a dilation rate of 2 and 2×2 max pooling to fuse mid-level features; Level 3: Aggregate high-level features through cross-layer skip connections; Level 4: Apply the SE module to implement channel attention weighting; S33, Cross-modal Feature Fusion: Design a three-branch parallel feature extraction network: visible light branch: extracts RGB spatial texture features through 5-layer convolution; infrared branch: extracts temperature distribution features using 4-layer convolution; point cloud branch: extracts 3D geometric features based on the PointNet++ architecture; S34. Adaptive feature fusion: adopts a three-level fusion strategy: primary fusion: unify the feature dimension to 256 dimensions through 1×1 convolution; intermediate fusion: apply 4-level spatial pyramid pooling to extract multi-scale context information; high-level fusion: feature weighted splicing based on attention mechanism; the final output is a multi-scale multimodal fusion feature map with a resolution of 1 / 8 of the original image.

5. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 2 is characterized in that: The S4 comprises the following steps: S41. Construct a three-level template library, including a standard template library, a local feature template library, and a dynamic update template library; S42. Using a feature map cross-correlation algorithm to perform multi-scale dynamic template matching, specifically including: S421, scale-adaptive feature extraction based on VGGNet: for the input image and templates , extract the first Layer feature map and template feature map , where the receptive field Satisfy constraints , where H and W represent the height and width of the image, respectively, C represents the number of color channels of the visible light image, and h and w represent the height and width of the template, respectively; S422, feature graph cross-correlation calculation: use the cross-correlation formula to calculate the similarity matrix S: , in, express In position at sub-region, is the position coordinate, represents the Frobenius norm, Represent the height and width of the sampling feature map respectively, The height and width of the template feature map, 、 Indicates feature extraction, mapping the image to a matching feature space, represents the improved cross-correlation function, i, j, k represent the index variables used to traverse the feature map data in the height direction, width direction, and channel direction respectively; S423、 Multi-scale matching: setting a set of scaling factors , R and r are scaling factor exponents and The corresponding index, for each Execution steps S421 - S422 , and unify the matching results of each scale into Native resolution; S43、 Use non-maximum suppression to filter candidate regions: fuse the multi-scale matching results, retain IOU ≥ 0.7 and Top K candidate regions with the highest confidence level; S44、 Sub-pixel precise positioning: the center point of the candidate area Perform Taylor expansion optimization: , The sub-pixel offset is determined by solving the eigenvector of the Hessian matrix. Indicates sub-pixel offset; S45. Defective area verification and template library update: store the confirmed defective area features into the dynamic update template library, and regularly optimize the local feature template library.

6. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 5 is characterized in that: The improved cross-correlation formula in step S422 introduces dynamic channel attention weights : It can more effectively measure the similarity between feature maps. The formula is as follows: ; in, ,By assigning different attention weights to each channel, the model can pay more attention to the channel information relevant to the current task, thereby improving the matching accuracy.

7. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 5 is characterized in that: The non-maximum suppression in step S43 adopts an adaptive threshold strategy, and the specific formula is: , in, and are the mean and standard deviation of the scores of all candidate regions in the current image, is an adjustable parameter, is the non-maximum suppression value.

8. The method for intelligent analysis of UAV power inspection images based on deep learning and multimodal fusion according to claim 5 is characterized in that: The sub-pixel positioning in step S44 adopts a quadratic surface fitting method, including: At the candidate point A quadratic surface model is established in the 3×3 neighborhood of ; Solving the parameters of the quadratic surface model by the least square method to find a set of parameters that minimizes the sum of squared errors between the similarity values ​​predicted by the model and the actual observed values; By solving the extreme points of the quadratic surface, the optimal offset is obtained ; The calculated optimal offset Add to candidate points Get the sub-pixel precise positioning coordinates .

9. A UAV power inspection image intelligent analysis system based on deep learning and multimodal fusion, applied to the UAV power inspection image intelligent analysis method based on deep learning and multimodal fusion as described in any one of claims 1 to 8, characterized in that: The system comprises: The data acquisition module is used to use a multi-rotor drone equipped with a three-axis stabilized gimbal system. The gimbal system integrates a high-resolution visible light camera, an infrared thermal imager, and a lidar. The A* algorithm is used to plan the optimal inspection path, achieving multi-angle full coverage scanning and multi-sensor synchronous data acquisition; The data preprocessing module is used to perform adaptive histogram equalization and dehazing on the collected visible light images, perform non-uniformity correction and temperature calibration on the infrared images, remove outliers and perform voxel grid filtering on the point cloud data, and use the ORB-SLAM3 algorithm to achieve spatiotemporal synchronization and sub-pixel registration of multi-source sensors; The multimodal feature fusion module is used to build a multi-scale feature fusion network based on the VGGNet-16 architecture. By introducing a deformable convolution module, a multi-level pyramid pooling structure, and a cross-modal attention mechanism, it achieves adaptive feature fusion of visible light, infrared, and point cloud data. The defect detection module is used to perform multi-scale dynamic template matching based on the constructed three-level template library, using the feature map cross-correlation algorithm to screen high-confidence defect areas through non-maximum suppression and achieve sub-pixel precision positioning; A report generation module combines structural similarity indicators, thermal anomaly analysis, and 3D deformation measurement results to automatically generate inspection reports that include defect type, precise location, and repair recommendations.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a program, and the program instructs the device to execute the UAV power inspection image intelligent analysis method based on deep learning and multimodal fusion as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Heterogenous image matching quality evaluation method and device based on GDB-ICP algorithm

    CN117456317A

  • Transformer substation fault early warning method and system based on multi-modal data sensing equipment

    CN118314485A

  • Radiator strength detection equipment and detection method

    CN120105357A

  • Monorail crane inspection robot intelligent test method based on data analysis

    CN120429803A

  • Glass lens surface scratch detection method and system

    CN120446165A

Cited By

  • Online monitoring method and system for power transmission and distribution line based on unmanned aerial vehicle platform

    CN121236703A

  • Online monitoring method and system for power transmission and distribution lines based on a drone platform

    CN121236703B

  • Post-disaster line inspection method and system based on acoustic imaging

    CN121256714A

  • Automatic board selection method and system based on magnesium oxide board surface defect recognition

    CN121305193A

  • Method for automatically drawing inspection point position by inspection robot based on multiple sensors

    CN121325873A