Iron tower structure defect intelligent detection system based on multi-view image fusion
The intelligent tower structure defect detection system, which utilizes multi-view image fusion and differential geometric feature extraction, solves the problems of low efficiency and insufficient accuracy in tower inspection. It achieves high-precision defect detection and assessment, reduces false detection rate, improves detection coverage and small defect identification capability, and provides scientific maintenance recommendations.
Patent Information
- Application Number
- CN202511485632.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-10-17
- Publication Date
- 2026-01-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
In existing technologies, defect detection of transmission line towers relies on manual inspection, which has problems such as high labor intensity, low detection efficiency, high safety risks, and strong subjectivity of detection results. Furthermore, single-view image acquisition methods are difficult to fully capture the tower structure, lack the ability to effectively identify defects of different scales, and lack a multi-source information fusion mechanism.
An intelligent detection system based on multi-view image fusion is adopted. Multi-view images of the tower are acquired through a drone platform. An efficient feature extraction and fusion mechanism is constructed by combining differential geometry theory. The system includes image acquisition, preprocessing, multi-view image fusion, feature extraction based on differential geometry and defect detection and evaluation, generating a three-dimensional visual model and identifying structural defects.
It achieves high-precision detection of structural defects in iron towers, with a coverage rate of 98.7%, a 15% improvement in defect detection accuracy, an 18% reduction in false detection rate, a 23% improvement in the detection rate of small defects, and provides scientific assessment of defect severity and risk level classification, saving maintenance costs.
Smart Images

Figure CN121437397A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of power facility inspection technology, and more particularly to an intelligent detection system for tower structure defects based on multi-view image fusion, which is used to automatically detect and evaluate various defects in transmission line tower structures. Background Technology
[0002] As a crucial infrastructure component of the power grid, the safe operation of transmission line towers is vital to the stability of the power system. Exposed to a complex natural environment for extended periods, these towers are susceptible to erosion from wind and rain, aging, and mechanical stress, leading to structural defects such as rust, fractures, deformation, and loose bolts. Traditional tower defect detection relies primarily on manual inspections, which suffers from high labor intensity, low efficiency, high safety risks, and strong subjectivity in the results.
[0003] In recent years, with the development of computer vision and artificial intelligence technologies, image-based intelligent detection methods for iron tower defects have gradually emerged. Existing technologies mainly adopt single-view image acquisition and processing methods. Although they can achieve a certain degree of automated detection, they have the following problems: 1) A single viewpoint is difficult to fully capture the complex structure of the iron tower, resulting in blind spots in detection; 2) They lack the ability to effectively identify defects of different scales, especially the detection accuracy of small defects such as loose bolts is insufficient; 3) The feature extraction and representation methods are relatively simple and difficult to cope with interference and occlusion in complex environments; 4) They lack an effective multi-source information fusion mechanism and cannot fully utilize the complementary information from different viewpoints and modalities.
[0004] Therefore, there is an urgent need to develop an intelligent detection system for tower structure defects that can make full use of multi-view image information and combine advanced feature extraction and fusion technologies to improve the comprehensiveness, accuracy and robustness of detection. Summary of the Invention
[0005] The purpose of this invention is to provide an intelligent detection system for structural defects in iron towers based on multi-view image fusion. The system acquires multi-view images of iron towers through an unmanned aerial vehicle (UAV) platform and constructs an efficient feature extraction and fusion mechanism by combining differential geometry theory, thereby achieving high-precision detection and evaluation of structural defects in iron towers.
[0006] This invention proposes an intelligent detection system for defects in iron tower structures based on multi-view image fusion, comprising:
[0007] The image acquisition module is used to acquire multi-view image data of the iron tower, which includes images of the tower top, tower body and tower base from different angles.
[0008] An image preprocessing module, connected to the image acquisition module, is used to preprocess the multi-view image data to generate standardized image data;
[0009] A multi-view image fusion module, connected to the image preprocessing module, is used to construct a three-dimensional visual model of the tower based on the standardized image data;
[0010] A feature extraction module based on differential geometry, connected to the multi-view image fusion module, is used to extract feature representations from the 3D visual model. The feature extraction module based on differential geometry includes:
[0011] Manifold feature space construction unit, used to map image features onto Riemannian manifolds and define the geometric structure of the feature manifolds;
[0012] Spatial feature pyramid building blocks are used to construct multi-scale feature representations on feature manifolds;
[0013] Global-local manifold guidance unit, used to guide the extraction and enhancement of local features using global structural information;
[0014] Multi-scale manifold fusion unit is used to fuse feature manifolds of different scales to generate a unified feature representation;
[0015] The defect detection module, connected to the differential geometry-based feature extraction module, is used to identify structural defects of the tower based on the unified feature representation and output defect information, including the location, type, and confidence level of the defect.
[0016] The defect assessment module, connected to the defect detection module, is used to assess the severity of defects in the tower structure based on the defect information and generate a risk level assessment result.
[0017] The early warning module, connected to the defect assessment module, is used to generate a tower defect detection report and maintenance recommendations based on the risk level assessment results.
[0018] Preferably, the image acquisition module includes:
[0019] The drone platform is used to carry imaging equipment and fly around the iron tower at multiple angles.
[0020] Imaging equipment, installed on the UAV platform, is used to acquire visible light and thermal infrared images of the tower;
[0021] A data acquisition and control unit, connected to the imaging device, is used to control the imaging device to use different shooting angles and parameters for different parts of the tower, wherein:
[0022] The top of the tower was photographed from above at a distance of 300m.
[0023] The middle section of the tower is oriented in a forward orientation.
[0024] The tower base is photographed from an upward angle, with a height angle of 30° and a tilt angle of 5°.
[0025] Preferably, the image preprocessing module includes:
[0026] An image quality assessment unit is used to perform sharpness detection, illumination uniformity analysis, and noise level assessment on the multi-view image data.
[0027] An image enhancement processing unit, connected to the image quality assessment unit, is used to perform adaptive histogram equalization, dehazing, and edge enhancement on multi-view image data based on the image quality assessment results.
[0028] A geometric correction unit, connected to the image enhancement processing unit, is used to perform lens distortion correction and angle normalization on the enhanced image;
[0029] The data augmentation unit, connected to the geometric correction unit, is used to expand the image dataset through operations such as rotation, mirroring, scaling, adding noise, and brightness changes.
[0030] Preferably, the multi-view image fusion module includes:
[0031] The feature point extraction unit is used to extract key feature points from images at various angles based on the SIFT algorithm.
[0032] The feature matching unit, connected to the feature point extraction unit, is used to match feature points using the RANSAC algorithm and remove mismatched points.
[0033] A geometric reconstruction unit, connected to the feature matching unit, is used to construct a three-dimensional geometric model of the tower based on the matching points;
[0034] The view synthesis unit, connected to the geometric reconstruction unit, is used to fuse multi-angle images into a unified three-dimensional visual model.
[0035] Preferably, the manifold feature space building unit includes:
[0036] The backbone network is used to extract preliminary feature maps from the input image;
[0037] A manifold mapper, connected to the backbone network, is used to map feature maps to a high-dimensional feature space;
[0038] A metric definer, connected to the manifold mapper, is used to define Riemannian metrics in the feature space and construct the feature manifold;
[0039] The geodesic calculator, connected to the metric definer, is used to calculate the geodesic distance between points on the feature manifold, describing the intrinsic geometric relationships between features.
[0040] Preferably, the spatial feature pyramid building block includes:
[0041] Multi-scale convolutional units are used to apply convolutional operations with different receptive fields to feature manifolds to generate multi-level feature representations.
[0042] A channel compressor, connected to the multi-scale convolution, is used to compress the channel dimension using 1×1 convolution, preserving the essential characteristics of the manifold structure.
[0043] A nonlinear converter, connected to the channel compressor, is used to enhance the expressive power of manifold features through a nonlinear activation function;
[0044] A residual connector, connected to the nonlinear converter, is used to introduce residual connections to maintain the topological invariance of the characteristic manifold.
[0045] Preferably, the global-local manifold guidance unit includes:
[0046] A global manifold extractor is used to extract a global manifold representation from the feature map of the top layer of the pyramid, capturing the overall structural information of the tower.
[0047] A local manifold generator is used to apply multi-scale pooling operations to the feature maps of each layer in the pyramid to construct multiple local manifold representations.
[0048] A mapping relationship builder, connected to the global manifold extractor and the local manifold generator, is used to establish mapping functions from the global manifold to each local manifold space;
[0049] An attention guide, connected to the mapping relationship builder, is used to generate an attention weight map based on global information to guide the extraction and enhancement of local features.
[0050] Preferably, the multi-scale manifold fusion unit includes:
[0051] A scale aligner is used to project feature manifolds of different scales onto a unified feature space.
[0052] A similarity calculator, connected to the scale aligner, is used to measure the correlation of features at different scales based on geodesic distance;
[0053] A weighted fusion unit, connected to the similarity calculator, is used to achieve effective integration of multi-scale features based on a weighting strategy of geodesic distance;
[0054] The feature enhancer, connected to the weighted fusion unit, is used to post-process the fused features to enhance their expressive power.
[0055] Preferably, the defect detection module includes:
[0056] The detection head generator is used to construct three detection heads for detecting targets at different scales, corresponding to small-scale defects, medium-scale defects, and large-scale defects, respectively.
[0057] A location regressor, connected to the detection head generator, is used to predict the location coordinates and size information of defects in the iron tower;
[0058] A category classifier, connected to the detection head generator, is used to identify the types of defects in the iron tower, including rust, fracture, deformation, loose bolts, loose nuts, defects, and irregularities.
[0059] A nonmaximum suppression processor, connected to the location regressor and the category classifier, is used to merge overlapping detection boxes to generate the final defect detection result.
[0060] Preferably, the defect assessment module includes:
[0061] Conditional random field analyzer is used to analyze the correlation between defect locations and tower structure;
[0062] A mechanical impact assessor, connected to the conditional random field analyzer, is used to assess the impact of defects on structural integrity based on the structural mechanical model of the iron tower.
[0063] A severity rating classifier, connected to the mechanical effect assessor, is used to classify defects into four severity levels: normal, moderate, severe, and hazardous.
[0064] A risk priority sorter, connected to the severity level classifier, is used to prioritize detected defects based on their severity level and location importance.
[0065] The present invention has the following beneficial effects:
[0066] 1. By using multi-view image fusion technology, a complete three-dimensional visual model of the tower is constructed, eliminating detection blind spots and achieving a detection coverage rate of up to 98.7%;
[0067] 2. A feature extraction method based on differential geometry is adopted, which treats image features as points on a Riemannian manifold and measures feature similarity by geodesic distance, thereby capturing the intrinsic structure of the feature space more accurately and improving the defect detection accuracy by 15%.
[0068] 3. A global-local manifold guidance mechanism is designed to utilize global structural information to guide the extraction and enhancement of local features, effectively addressing detection scenarios with complex tower structures and numerous background interferences, and reducing the false detection rate by approximately 18%.
[0069] 4. Construct a multi-scale feature fusion strategy that is adaptive to manifold metrics, and optimize it for defects of different scales, especially improving the detection rate of small defects by about 23%;
[0070] 5. It enables automatic assessment of defect severity and risk level classification, providing a scientific basis for tower maintenance and saving approximately 300,000 yuan in maintenance costs per 100km of transmission line per year. Attached Figure Description
[0071] Figure 1 This is the overall architecture diagram of the intelligent detection system for defects in iron tower structures based on multi-view image fusion, as proposed in this invention.
[0072] Figure 2 This is a schematic diagram of the multi-view image acquisition path of the UAV of the present invention;
[0073] Figure 3 This is a schematic diagram of the feature extraction module based on differential geometry of the present invention;
[0074] Figure 4 This is a flowchart of the spatial feature pyramid construction unit of the present invention;
[0075] Figure 5 This is a schematic diagram of the global-local manifold guidance unit of the present invention;
[0076] Figure 6 This is a flowchart of the multi-scale manifold fusion unit of the present invention;
[0077] Figure 7 This is a flowchart of the defect severity assessment and risk level classification of the present invention. Detailed Implementation
[0078] Please refer to Figures 1-7 The present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Those skilled in the art should understand that these embodiments are for illustrative purposes only and are not intended to limit the scope of the invention.
[0079] like Figure 1 As shown, the intelligent detection system for defects in iron tower structures based on multi-view image fusion of the present invention includes an image acquisition module 1, an image preprocessing module 2, a multi-view image fusion module 3, a feature extraction module based on differential geometry 4, a defect detection module 5, a defect assessment module 6, and an early warning module 7.
[0080] Image acquisition module 1 is used to acquire multi-view image data of the tower, which includes images of the tower top, tower body, and tower base from different angles. Preferably, the multi-view images include visible light images and thermal infrared images to provide richer structural information about the tower.
[0081] Image preprocessing module 2 is connected to image acquisition module 1 and is used to preprocess the multi-view image data to generate standardized image data. Preprocessing operations include, but are not limited to, image quality assessment, image enhancement, geometric correction, and data augmentation.
[0082] The multi-view image fusion module 3 is connected to the image preprocessing module 2 and is used to construct a three-dimensional visual model of the tower based on the standardized image data. This module fuses images from different perspectives into a complete representation of the tower through techniques such as feature point extraction and matching, and three-dimensional reconstruction.
[0083] The feature extraction module 4 based on differential geometry is connected to the multi-view image fusion module 3, and is used to extract feature representations from the three-dimensional visual model. This module is the core innovation of this system, mapping image features onto a Riemannian manifold and improving the expressive power of features through multi-scale feature representation and global-local feature fusion.
[0084] The defect detection module 5 is connected to the feature extraction module 4 based on differential geometry. It is used to identify the structural defects of the tower based on the unified feature representation and output defect information, which includes the location, type and confidence level of the defect.
[0085] The defect assessment module 6 is connected to the defect detection module 5 and is used to assess the severity of defects in the tower structure based on the defect information and generate a risk level assessment result.
[0086] The early warning module 7 is connected to the defect assessment module 6 and is used to generate a tower defect detection report and maintenance recommendations based on the risk level assessment results.
[0087] The implementation methods of each module of the system are described in detail below.
[0088] like Figure 2 As shown, the image acquisition module 1 includes a drone platform 11, an imaging device 12, and an acquisition control unit 13.
[0089] The drone platform 11 is used to carry imaging equipment and fly around the tower from multiple angles. In this embodiment, a six-axis industrial-grade drone is preferably used, with a payload capacity of not less than 2 kg and a flight time of not less than 30 minutes, to ensure that the task of acquiring images of the tower from all directions can be completed.
[0090] Imaging device 12 is mounted on UAV platform 11 and is used to acquire visible light and thermal infrared images of the tower. Preferably, the visible light camera is a 4K high-definition camera with optical image stabilization and 30x optical zoom capability; the thermal infrared camera uses an infrared sensor with a resolution of 640×512 and a temperature resolution of 0.05°C, so as to detect abnormal temperature areas in the tower structure.
[0091] The acquisition and control unit 13 is connected to the imaging device 12 and is used to control the imaging device to use different shooting angles and parameters for different parts of the tower. Based on the characteristics of the tower structure, the following shooting strategy is adopted: a top-down shooting method is used for the top of the tower, with an aerial shooting distance of 300m; a frontal shooting method is used for the middle of the tower; and a bottom-up shooting method is used for the base of the tower, with a base elevation angle of 30° and a pitch angle of 5°. This multi-angle acquisition strategy can effectively cover all parts of the tower and reduce blind spots in detection.
[0092] In practical applications, drones typically fly around the tower multiple times, capturing 4-6 images from different angles at each location to ensure comprehensive coverage of the tower structure. The image acquisition resolution is set to 3840×2160 pixels to guarantee the ability to capture detailed information.
[0093] The image preprocessing module 2 includes an image quality assessment unit 21, an image enhancement processing unit 22, a geometric correction unit 23, and a data augmentation unit 24.
[0094] Image quality assessment unit 21 is used to perform sharpness detection, illumination uniformity analysis, and noise level assessment on multi-view image data. Sharpness detection uses the Laplacian operator to calculate the image gradient variance. When the gradient variance is lower than the threshold of 25 (an empirical value, within the range of pixel values 0-255), the image is judged to be of low sharpness. Illumination uniformity is assessed by calculating the standard deviation of the image histogram. A standard deviation less than 40 (an empirical value, based on a large number of experimental statistics) indicates uneven illumination. Noise level is estimated using the mean square error. A noise estimate greater than 15 is judged to be a high-noise image.
[0095] Image enhancement processing unit 22 is connected to image quality evaluation unit 21 and is used to perform adaptive histogram equalization, dehazing, and edge enhancement on multi-view image data based on image quality evaluation results. For images with uneven illumination, the adaptive histogram equalization (CLAHE) algorithm is applied, where the contrast limit parameter is set to 3.0 and the grid size is 8×8; for images with haze or smoke interference, a dehazing algorithm based on dark channel prior is used; for images with unclear edge details, a nonlinear sharpening filter is used for edge enhancement, with the enhancement intensity coefficient set to 1.5.
[0096] The geometric correction unit 23 is connected to the image enhancement processing unit 22 and is used to perform lens distortion correction and viewpoint normalization on the enhanced image. Distortion correction adopts a distortion model based on camera intrinsic parameters, including correction of radial distortion and tangential distortion; viewpoint normalization transforms images taken from different angles to a unified reference coordinate system through perspective transformation, which facilitates subsequent image fusion processing.
[0097] The data augmentation unit 24 is connected to the geometric correction unit 23 and is used to expand the image dataset through operations such as rotation, mirroring, scaling, adding noise, and brightness changes. Specifically, the original image is rotated by ±15° and ±30°, horizontally and vertically mirrored, randomly scaled by 80%-120%, and its brightness is changed by ±10%. Gaussian noise (mean 0, standard deviation 0.01) is also added. Through these data augmentation operations, the original dataset can be expanded by 6 times, effectively enhancing the model's generalization ability.
[0098] The multi-view image fusion module 3 includes a feature point extraction unit 31, a feature matching unit 32, a geometric reconstruction unit 33, and a view synthesis unit 34.
[0099] The feature point extraction unit 31 is used to extract key feature points from images at various angles based on the SIFT algorithm. The SIFT algorithm can extract feature points that are invariant to changes in rotation, scaling, and illumination, making it suitable for registration of images of the Eiffel Tower from different perspectives. In this embodiment, approximately 2000 feature points are extracted from each image, and each feature point is represented by a 128-dimensional vector. To improve processing efficiency, a GPU-accelerated version of the SIFT algorithm can be used.
[0100] Feature matching unit 32 is connected to feature point extraction unit 31 and is used to match feature points using the RANSAC algorithm and remove false matches. First, the k-nearest neighbor (k=2) algorithm is used for preliminary matching, and then the ratio test (threshold set to 0.75) is used to filter candidate matching pairs. Next, the RANSAC algorithm is applied to further remove outliers. The number of iterations is set to 1000, and the inlier threshold is set to 3.0 pixels. Generally, 70%-80% of the valid matching points can be retained.
[0101] The geometric reconstruction unit 33 is connected to the feature matching unit 32 to construct a 3D geometric model of the tower based on the matching points. A structure from motion (SfM) algorithm based on matching points is used here, comprising two steps: camera pose estimation and sparse point cloud reconstruction. Camera pose estimation is optimized using the Bundle Adjustment algorithm with 100 iterations and a convergence threshold of 1e-6. Sparse point cloud reconstruction employs triangulation, calculating 3D coordinates based on known camera intrinsic and extrinsic parameters and the matching points.
[0102] The view synthesis unit 34 is connected to the geometry reconstruction unit 33 to fuse multi-angle images into a unified 3D visual model. This unit employs a mesh-based multi-view texture mapping method to map image information from different perspectives onto the surface of the 3D model. To handle texture conflicts in overlapping view areas, a view weight fusion strategy is used, with the weight factor proportional to the cosine of the angle between the viewing direction and the surface normal. The final generated 3D visual model achieves a resolution of up to 80% of the original image, preserving rich detail information.
[0103] The feature extraction module 4 based on differential geometry is the core innovative part of this system, such as... Figure 3 As shown, it includes a manifold feature space construction unit 41, a spatial feature pyramid construction unit 42, a global-local manifold guidance unit 43, and a multi-scale manifold fusion unit 44.
[0104] Manifold Feature Space Building Unit
[0105] The manifold feature space building unit 41 includes a backbone network 411, a manifold mapper 412, a metric definer 413, and a geodesic calculator 414.
[0106] The backbone network 411 is used to extract preliminary feature maps from the input image. In this embodiment, an improved ResNet50 is used as the backbone network. The input is a 2D projection image of a 3D visual model (size 640×640×3), and the output is feature maps at five different scales: C1 (320×320×64), C2 (160×160×128), C3 (80×80×256), C4 (40×40×512), and C5 (20×20×1024). The convolutional layers in the backbone network use 3×3 convolutional kernels with a stride of 2 and the ReLU activation function.
[0107] Manifold mapper 412 is connected to backbone network 411 and is used to map feature maps to a high-dimensional feature space. The specific implementation is as follows: For feature maps... Through nonlinear transformation function Map it to a higher-dimensional space:
[0108] ,
[0109] in: For the mapped high-dimensional features, the dimension and same; It is a nonlinear mapping function; For the first Layer feature map; The transformation matrix has dimensions that depend on the specific implementation. For bias terms, dimensions and Correspondingly, ReLU is the modified linear unit activation function, defined as follows: Transformation matrix The values were obtained through backpropagation, and the initial values were initialized using the He method.
[0110] The metric definer 413 is connected to the manifold mapper 412 and is used to define Riemannian metrics in the feature space and construct the feature manifold. In the feature space, a Riemannian metric tensor is defined. :
[0111] ,
[0112] in: The Riemannian metric tensor is a positive definite matrix that defines the distance metric in the feature space. For mapping functions At point The Jacobian matrix at a point represents the local linear approximation of the mapping function at that point; the superscript... This represents the matrix transpose operation. This definition makes the feature space a Riemannian manifold, which can more accurately describe the intrinsic geometric relationships of the features.
[0113] The geodesic calculator 414 is connected to the metric definer 413 and is used to calculate the geodesic distance between points on a feature manifold, describing the intrinsic geometric relationships between features. On a Riemannian manifold, the geodesic distance between two points is defined as:
[0114] ,
[0115] in: For point and Geodesic distance between them; For connection points and The path satisfies and ; For path At point The tangent vector at the point; For point Riemannian metric at the location; The infimum represents the shortest distance among all possible paths; the integral represents the distance along the path. The length of the geodesic distance. In actual calculations, the FastMarching Method is used to approximate the geodesic distance, with a computational complexity of O(n log n). ,in This represents the number of feature points.
[0116] like Figure 4As shown, the spatial feature pyramid building block 42 includes a multi-scale convolutional unit 421, a channel compressor 422, a nonlinear converter 423, and a residual connector 424.
[0117] The multi-scale convolutional unit 421 is used to apply convolutional operations with different receptive fields to the feature manifold, generating multi-level feature representations. The specific implementation is as follows: For the feature manifold... Three different receptive fields (1×1, 3×3, 5×5) are applied to the convolution operation:
[0118] ,
[0119] in: Indicates use Feature maps obtained from convolution kernels; Indicates the size is Convolution operations; The kernel size is denoted by 1, 3, or 5. Different kernel sizes can capture spatial information of different ranges, enriching the representation of features.
[0120] Channel compressor 422 is connected to multi-scale convolutional unit 421 to achieve channel-dimensional compression using 1×1 convolution, preserving the essential characteristics of the manifold structure. The channel compression operation is defined as follows:
[0121] ,
[0122] in: This is the compressed feature map; This involves concatenating the channel dimensions of multi-scale features, i.e. ; This is a 1×1 convolution operation; This is a channel-dimensional concatenation operation, which combines multiple feature maps into a single feature map along the channel dimension. In this embodiment, the number of channels is compressed to 1 / 4 of the original, effectively reducing computational load while preserving key feature information.
[0123] The nonlinear converter 423 is connected to the channel compressor 422 and is used to enhance the expressive power of manifold features through a nonlinear activation function. The nonlinear converter is defined as:
[0124] ,
[0125] in: The feature map after nonlinear transformation; To modify the activation function of the linear unit, it is defined as follows: Negative values are truncated to zero, while positive values remain unchanged. This nonlinear transformation enhances the expressive power of features, enabling the network to learn more complex feature representations.
[0126] Residual connector 424 connects to nonlinear converter 423 to introduce residual connections, preserving the topological invariance of the characteristic manifold. A residual connection is defined as:
[0127] ,
[0128] in: Features for the final output; The feature map after nonlinear transformation; This represents the original feature map; the plus sign indicates element-wise addition. Residual connections effectively alleviate the vanishing gradient problem in deep networks while preserving the structural information of the original features, thus improving the expressive power of the features and training stability.
[0129] Through the above operations, a four-layer feature pyramid is constructed: P2 (160×160×256), P3 (80×80×256), P4 (40×40×256), and P5 (20×20×256). Each layer integrates information from different levels through upsampling and lateral connections.
[0130] like Figure 5 As shown, the global-local manifold guidance unit 43 includes a global manifold extractor 431, a local manifold generator 432, a mapping relation builder 433, and an attention guide 434.
[0131] The global manifold extractor 431 is used to extract a global manifold representation from the top-level feature map of the pyramid, capturing the overall structural information of the tower. The extraction process of the global manifold representation is as follows:
[0132] ,
[0133] ,
[0134] ,
[0135] in: This is the result of global average pooling, with a dimension of 1×1×256; This is the result of global max pooling, with a dimension of 1×1×256; This represents the top-level feature map of the feature pyramid, with dimensions of 20×20×256; GlobalAvgPool represents the global average pooling operation, which calculates the average value of each channel of the feature map; GlobalMaxPool represents the global max pooling operation, which calculates the maximum value of each channel of the feature map; Concat represents the channel dimension concatenation operation; FC represents the fully connected layer operation. This is the global context vector, with dimensions 1×1×256.
[0136] The local manifold generator 432 is used to apply multi-scale pooling operations to the feature maps of each layer in the pyramid to construct multiple local manifold representations. For each layer of the feature pyramid... Four different pooling scales are applied:
[0137] ,
[0138] in: This represents the corresponding local manifold. For the characteristic pyramid number Layer feature map; Indicates the size is Pooling operations; This represents the pooling window size, which can be 3, 5, 7, or 9. Pooling operations at different scales can capture structural and texture features within different ranges.
[0139] The mapping relation builder 433 is connected to the global manifold extractor 431 and the local manifold generator 432, and is used to establish mapping functions from the global manifold to the spaces of each local manifold. Mapping functions Defined as:
[0140] ,
[0141] in: These are the mapped global features; For global context vectors; The function will use the global context vector Remodeled to the same spatial dimensions as the local feature map ; This is a 1×1 convolution operation; and The first The height and width of the layer feature map.
[0142] Attention guide 434 is connected to mapping relationship builder 433 to generate attention weight map based on global information, guiding the extraction and enhancement of local features. The generation process of attention weight map is as follows:
[0143] ,
[0144] ,
[0145] in: Given a similarity matrix, calculate the dot product similarity between the global mapping features and the local features; These are the mapped global features; Local features; express Transpose of; and Representing global mapping features and local features respectively Norm; The temperature parameter controls the smoothness of the softmax distribution, set to 0.07; softmax is the softmax normalization function, defined as softmax... ; This is the attention weight map. Then, local features are enhanced based on the attention weights:
[0146] ,
[0147] ,
[0148] in: These are attention-weighted local features; This is an enhanced feature resulting from integration with global guidance information; The plus sign indicates element-wise multiplication; the plus sign indicates element-wise addition.
[0149] like Figure 6 As shown, the multi-scale manifold fusion unit 444 includes a scale aligner 441, a similarity calculator 442, a weighted fusion unit 443, and a feature enhancer 444.
[0150] The scale aligner 441 is used to project feature manifolds of different scales onto a unified feature space. The specific implementation is as follows:
[0151] ,
[0152] in: The aligned feature map; For the characteristic pyramid number Layer feature map; for Convolution operations are used to adjust the number of channels; The operation of adjusting the spatial dimensions for bilinear interpolation; and For the target size, the size of the intermediate layer is usually taken, such as... In this way, feature maps of different scales are aligned to a uniform spatial resolution.
[0153] A similarity calculator 442 is connected to a scale aligner 441 to measure the correlation of features at different scales based on geodesic distance. For the aligned feature map... and Calculate their geodesic distance matrix:
[0154] ,
[0155] in: These are the element values of the distance matrix; This is the geodesic distance function; and Representing feature maps respectively At position x and feature map The feature vector at position y. Based on the distance matrix, calculate the feature similarity:
[0156] ,
[0157] in: This is a similarity matrix; It is a distance matrix; It is a natural exponential function; The scaling parameter controls the sensitivity of similarity and is set to 1 / 3 of the average distance.
[0158] The weighted fusion unit 443 is connected to the similarity calculator 442 to achieve efficient integration of multi-scale features using a geodesic distance-based weighting strategy. The fusion weights are calculated as follows:
[0159] ,
[0160] in: The fusion weights are for the i-th scale feature; Let be the geodesic distance matrix between the i-th scale feature and the reference feature; It is a natural exponential function; The temperature parameter controls the concentration of the weight distribution, which is set to 10.0. The number of feature scales is 4 in this embodiment; the denominator is a normalization factor to ensure that the sum of all weights is 1. The multi-scale feature fusion process is as follows:
[0161] ,
[0162] in: This represents the fused features. The fusion weights are the features at the i-th scale. This represents the i-th scale feature after alignment. This represents element-wise multiplication; This represents element-wise summation; This represents the number of feature scales.
[0163] Feature enhancer 444 is connected to weighted fusion unit 443 and is used to post-process the fused features to enhance their expressive power. The feature enhancement process is as follows:
[0164] ,
[0165] in For the enhanced feature representation; This represents the fused feature representation; ReLU is the modified linear unit activation function; BN is the batch normalization operation, used to stabilize the training process and accelerate convergence; Conv... These are 3×3 convolution operations. Through these operations, the final unified feature representation is generated. The dimensions are 80×80×512, which serves as the input for defect detection.
[0166] The defect detection module 5 includes a detection head generator 51, a position regressor 52, a category classifier 53, and a non-maximum suppression processor 54.
[0167] The detection head generator 51 is used to construct three detection heads for target detection at different scales, corresponding to small-scale defects (≤32×32 pixels), medium-scale defects (32×32~96×96 pixels), and large-scale defects (≥96×96 pixels), respectively. Each detection head contains a series of convolutional layers and upsampling operations to adjust the feature resolution to the corresponding scale. The output feature map size of the small-scale detection head is 80×80×256, the medium-scale is 40×40×256, and the large-scale is 20×20×256.
[0168] The location regressor 52 is connected to the detection head generator 51 and is used to predict the location coordinates and size information of defects in the tower. The location regression is implemented using a convolutional layer, outputting four channels, each corresponding to the center coordinate of the bounding box. and size To improve the accuracy of small target detection, a higher weight (weight factor set to 1.5) is used for the position regression of the small-scale detection head. The position coordinates are normalized to the range [0,1] using the sigmoid function, and the size information is transformed using an exponential function to ensure positive values.
[0169] The classifier 53 is connected to the detection head generator 51 to identify the types of defects in the iron tower, including seven categories: rust, fracture, deformation, loose bolts, loose nuts, defects, and irregular shapes. The classifier is implemented using 1×1 convolution, and the number of output channels is the number of categories (7) plus 1 (background class). The category probabilities are obtained by normalization using the softmax function. To handle the class imbalance problem, the focal loss function is used, with parameters... Set it to 0.25. Set it to 2.0.
[0170] The non-maximum suppression processor 54 is connected to the location regressor 52 and the category classifier 53 to merge overlapping detection boxes and generate the final defect detection result. The non-maximum suppression process is as follows: First, all detection boxes are sorted in descending order of confidence; then, the box with the highest confidence is selected, and other boxes with an IoU (Intersection over Union) greater than a threshold are removed; this process is repeated until all boxes have been processed. In this embodiment, the IoU threshold is set to 0.5, and the minimum confidence threshold is set to 0.3. Through non-maximum suppression, redundant detections are effectively eliminated, resulting in accurate defect localization results.
[0171] The defect assessment module 6 includes a conditional random field analyzer 61, a mechanical impact assessor 62, a severity level classifier 63, and a risk priority sorter 64.
[0172] The Conditional Random Field (CRF) analyzer 61 is used to analyze the correlation between defect locations and the tower structure. The CRF model can model the spatial relationships between tower structural elements, improving the contextual understanding of defect detection. In the CRF model, the potential function is defined as:
[0173] ,
[0174] in: It is an energy function; For observation variables, represent image features; This is a latent variable representing the defect label; It is a univariate potential function, representing the label probability at a single location; This is a binary potential function, representing the interaction between adjacent labels; and The pixel position in the image; and Representing positions respectively and The label variable; This represents summing over all pixel positions; This represents the summation over all adjacent pixel pairs. The optimal defect segmentation result is obtained by minimizing the energy function.
[0175] The mechanical impact assessor 62 is connected to the conditional random field analyzer 61 to evaluate the impact of defects on structural integrity based on the tower structure's mechanical model. The evaluation process considers the following factors: the severity coefficient of the defect type, the defect size (proportion to the component size), the structural importance of the defect location, and the defect's development trend. For different types of defects, the following severity coefficients are set: corrosion (0.6), fracture (1.0), deformation (0.8), bolt loosening (0.7), nut loosening (0.7), defects (0.9), and irregularities (0.5). Based on these factors, a comprehensive impact score is calculated.
[0176] ,
[0177] in: The overall impact score represents the degree to which defects affect structural integrity. The defect type weight is determined based on the severity of different defect types. The base severity coefficient is determined based on the defect type. The defect area; The area of the component; This represents the ratio of the defect area to the component area. This is the location importance coefficient, ranging from 1.0 to 2.0, determined based on the importance of the defect location in the structure. The development trend factor, ranging from 1.0 to 1.5, reflects the potential risk of defects developing over time.
[0178] The severity grading unit 63 is connected to the mechanical impact assessor 62 and is used to classify defects into four severity levels: normal, moderate, severe, and hazardous. The grading is based on a comprehensive impact score. The thresholds were 0.3, 0.6, and 0.8, respectively. Normal level ( ) indicates no immediate action is required; general level ( ) indicates that regular monitoring is required; severity level ( ) indicates that it needs to be dealt with as soon as possible; hazard level ( () indicates that immediate action is required. Each level is marked with a different color: Normal (green), Moderate (yellow), Severe (orange), and Dangerous (red).
[0179] Risk prioritization unit 64 is connected to severity rating unit 63 to prioritize detected defects based on their severity level and location importance. The prioritization considers the following factors: defect level, structural importance of defect location, combined effect of multiple defects, and confidence level of defect detection. The overall priority score is calculated as follows:
[0180] ,
[0181] in: The priority score indicates the priority of defect repair. The numerical values represent the severity level (Normal = 1, Average = 2, Severe = 3, Dangerous = 4). This is the location importance coefficient, ranging from 1.0 to 2.0; The number of defects on the same component; To determine the confidence level, the range is 0.0-1.0; This is a combined effect factor, reflecting the cumulative impact of multiple defects on the structure. The higher the priority score, the higher the maintenance priority.
[0182] The early warning module 7 is used to generate tower defect detection reports and maintenance recommendations based on the risk level assessment results. The early warning module includes three main functions: early warning information generation, maintenance recommendation formulation, and detection report generation.
[0183] The warning information generation function automatically generates warning information of corresponding levels based on the severity of the defect. For general defects, a low-level warning is generated, suggesting handling it during the next scheduled maintenance; for severe defects, a medium-level warning is generated, suggesting handling it within one month; for critical defects, a high-level warning is generated, suggesting immediate handling. Warning information is sent to relevant maintenance personnel via the system interface, email, or SMS.
[0184] The maintenance recommendation feature provides tailored maintenance suggestions based on the type and severity of the defect. For example, for rust defects, it recommends rust removal and replacement of the anti-corrosion coating; for loose bolt defects, it recommends retightening the bolts and installing anti-loosening devices; for fracture defects, it recommends replacing the fractured component and reinforcing the surrounding structure. Maintenance recommendations include recommended repair methods, required tools and materials, estimated workload, and safety precautions.
[0185] The inspection report generation function automatically generates a comprehensive report on tower defects, including basic tower information, inspection time and conditions, defect statistics, severity assessment results, image evidence of key defects, maintenance recommendations, and follow-up monitoring plans. The report uses a standardized format for easy management and archiving, and supports multiple output formats such as PDF, Word, and HTML.
[0186] The intelligent tower structure defect detection system of this invention achieves a complete workflow from data acquisition to defect early warning through effective collaboration between modules. The data flow paths between modules are as follows:
[0187] Multi-view image data is acquired by image acquisition module 1 and transmitted to image preprocessing module 2 for preprocessing to generate standardized image data. The standardized image data is then processed by multi-view image fusion module 3 to construct a 3D visual model of the tower. The 3D visual model is processed by feature extraction module 4 based on differential geometry to extract unified feature representations. These feature representations are processed by defect detection module 5 to identify structural defects in the tower and output defect information. The defect information is processed by defect assessment module 6 to generate a risk level assessment result. The risk level assessment result is processed by early warning module 7 to generate a tower defect detection report and maintenance recommendations.
[0188] Standardized interface design between modules ensures consistency and compatibility in data transmission. For example, the image data interface uniformly adopts the standard RGB format with a resolution of 3840×2160 pixels; the feature representation interface specifies the size and number of channels of the feature map; and the defect information interface defines a unified defect representation format, including location, type, and confidence fields. This standardized interface design enables seamless connection and collaborative work among the various modules of the system.
[0189] The system of this invention can be deployed in a hardware environment consisting of a drone platform and a back-end processing server. The drone platform is a six-axis industrial-grade drone equipped with a 4K high-definition camera and a thermal infrared camera, and features an NVIDIA Jetson Xavier NX edge computing unit for front-end image processing and preliminary feature extraction. The back-end processing server is configured with 64GB of RAM and an NVIDIA RTX A6000 GPU for complex model inference and defect assessment.
[0190] The system adopts a distributed processing architecture, rationally allocating computing tasks to the edge and cloud to optimize the utilization of computing resources. Image preprocessing and preliminary feature extraction are completed at the edge, while complex feature fusion, defect detection, and evaluation tasks are performed in the cloud. The system supports an offline working mode, enabling basic data acquisition and caching in environments without network access, with data uploaded to the cloud for processing once the network is restored.
[0191] In practical applications, the system demonstrates excellent performance: the defect detection accuracy reaches 92.8%, a 15% improvement over traditional methods; the small target defect detection rate is improved by 23%; the false detection rate is reduced by 18%; the single image processing time is less than 200ms (NVIDIA Jetson platform); and the detection coverage reaches 98.7%. The system has been tested on transmission towers in 10 substations across 5 different provinces, reducing manual inspection time by 75%, increasing the early detection of potential risk points by 40%, and saving approximately 300,000 yuan in maintenance costs per 100km of transmission line annually.
[0192] The embodiments described above are merely illustrative of specific implementations of the present invention, and while the descriptions are detailed, they should not be construed as limiting the scope of the present invention. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of the present invention, and these modifications and improvements all fall within the scope of protection of the present invention.
Claims
1. An intelligent detection system for tower structure defects based on multi-view image fusion, characterized in that, The method comprises the following steps: An image acquisition module is used to acquire multi-view image data of the tower, which includes different angle images of the top, body and base of the tower; An image preprocessing module is connected with the image acquisition module and is used to preprocess the multi-view image data to generate standardized image data; A multi-view image fusion module is connected with the image preprocessing module and is used to construct a three-dimensional visual model of the tower based on the standardized image data; A differential geometry-based feature extraction module is connected with the multi-view image fusion module and is used to extract feature representations from the three-dimensional visual model, which comprises: A manifold feature space construction unit is used to map image features to a Riemannian manifold and define the geometric structure of the feature manifold; A spatial feature pyramid construction unit is used to construct multi-scale feature representations on the feature manifold; A global-local manifold guidance unit is used to guide the extraction and enhancement of local features using global structure information; A multi-scale manifold fusion unit is used to fuse feature manifolds of different scales to generate a unified feature representation; A defect detection module is connected with the differential geometry-based feature extraction module and is used to identify structural defects of the tower based on the unified feature representation and output defect information, which includes the location, type and confidence of the defects; A defect evaluation module is connected with the defect detection module and is used to evaluate the severity of the tower structure defects based on the defect information and generate a risk level evaluation result; An early warning module is connected with the defect evaluation module and is used to generate a tower defect detection report and maintenance recommendations based on the risk level evaluation result. 2.The intelligent tower structure defect detection system of claim 1, wherein, The image acquisition module comprises: A UAV platform is used to carry an imaging device to fly around the tower at multiple angles; An imaging device is installed on the UAV platform and is used to collect visible light images and thermal infrared images of the tower; A collection control unit is connected with the imaging device and is used to control the imaging device to adopt different shooting angles and parameters for different parts of the tower, wherein: A downward-looking mode is adopted for the top part with a flight distance of 300m; A front-looking mode is adopted for the middle part of the tower body; An upward-looking mode is adopted for the base part with a base height angle of 30° and a pitch angle of 5°. 3.The intelligent tower structure defect detection system of claim 1, wherein, The image preprocessing module comprises: An image quality assessment unit is used to detect the sharpness, analyze the lighting uniformity and evaluate the noise level of the multi-view image data; An image enhancement processing unit is connected with the image quality assessment unit and is used to perform adaptive histogram equalization, defogging processing and edge enhancement on the multi-view image data based on the image quality assessment result; A geometric correction unit is connected with the image enhancement processing unit and is used to perform lens distortion correction and perspective normalization on the enhanced images; A data augmentation unit is connected with the geometric correction unit and is used to expand the image data set through rotation, mirroring, scaling, adding noise and brightness changes, etc.
4. The intelligent detection system for tower structure defects according to claim 1, characterized in that, The multi-view image fusion module comprises: A feature point extraction unit is used to extract key feature points of each angle image based on the SIFT algorithm; The feature matching unit, connected with the feature point extraction unit, is configured to match the feature points by using a RANSAC algorithm and to eliminate the mis-matching points; The geometric reconstruction unit, connected with the feature matching unit, is configured to construct a three-dimensional geometric model of the tower based on the matching points; The view synthesis unit, connected with the geometric reconstruction unit, is configured to fuse the multi-angle images into a unified three-dimensional visual model.
5. The intelligent detection system for tower structure defects according to claim 1, characterized in that, The manifold feature space construction unit comprises: A backbone network configured to extract a preliminary feature map of an input image; A manifold mapper connected with the backbone network and configured to map the feature map to a high-dimensional feature space; A metric definer connected with the manifold mapper and configured to define a Riemannian metric in the feature space to construct a feature manifold; A geodesic line calculator connected with the metric definer and configured to calculate the geodesic distance between points on the feature manifold to describe the intrinsic geometric relationship between features.
6. The intelligent tower structure defect detection system of claim 1, wherein The spatial feature pyramid construction unit comprises: A multi-scale convolver configured to apply convolution operations with different receptive fields to the feature manifold to generate multi-level feature representations; A channel compressor connected with the multi-scale convolver and configured to use 1*1 convolution to compress the channel dimension and retain the essential features of the manifold structure; A nonlinear converter connected with the channel compressor and configured to enhance the expression ability of the manifold features by using a nonlinear activation function; A residual connector connected with the nonlinear converter and configured to introduce a residual connection to maintain the topological structure invariance of the feature manifold.
7. The intelligent tower structure defect detection system of claim 1, wherein The global-local manifold guidance unit comprises: A global manifold extractor configured to extract a global manifold representation from the feature map at the top layer of the pyramid to capture the overall structural information of the tower; A local manifold generator configured to apply multi-scale pooling operations to the feature maps at each layer of the pyramid to construct multiple local manifold representations; A mapping relationship constructor connected with the global manifold extractor and the local manifold generator and configured to establish a mapping function from the global manifold to each local manifold space; An attention guider connected with the mapping relationship constructor and configured to generate an attention weight map based on the global information to guide the extraction and enhancement of local features. 8.The intelligent tower structure defect detection system of claim 1, wherein, The multi-scale manifold fusion unit comprises: A scale aligner configured to project the feature manifolds at different scales to a unified feature space; A similarity calculator connected with the scale aligner and configured to measure the relevance of features at different scales based on the geodesic distance; A weighted fusioner connected with the similarity calculator and configured to effectively integrate the multi-scale features based on a weighting strategy of the geodesic distance; A feature enhancer connected with the weighted fusioner and configured to post-process the fused features to enhance the feature expression ability. 9.The intelligent tower structure defect detection system of claim 1, wherein, The defect detection module comprises: A detection head generator configured to construct three detection heads for detecting defects at different scales, corresponding to small-scale defects, medium-scale defects and large-scale defects, respectively; A position regressor connected with the detection head generator and configured to predict the position coordinates and size information of the tower defects; A class classifier connected with the detection head generator and configured to identify the types of the tower defects, including rust, fracture, deformation, loose bolt, loose nut, damage and abnormal shape. A non-maximum suppression processor, connected with the position regressor and the category classifier, is configured to merge overlapping bounding boxes to generate a final defect detection result.
10. The intelligent tower structure defect detection system of claim 1, wherein, The defect evaluation module comprises: A conditional random field analyzer configured to analyze the relevance of the defect position to the tower structure; A mechanical influence evaluator, connected with the conditional random field analyzer, configured to evaluate the influence of the defect on the structural integrity based on a tower structure mechanical model; A severity level classifier, connected with the mechanical influence evaluator, configured to classify the defect into four severity levels: normal, general, severe, and dangerous; A risk priority sorter, connected with the severity level classifier, configured to prioritize the detected defects based on the severity level and the position importance of the defect.
Citation Information
Cited By
Underwater structure volume measurement method and system based on improved MVSnet
CN122176038A