A tower material recognition method and system based on machine vision
Patent Information
- Application Number
- CN202610771616.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-01
- Publication Date
- 2026-08-28
AI Technical Summary
[0005]本发明针对现有输电杆塔塔材识别方法易受到户外复杂环境变化的影响,令塔材特征提取失真,从而导致识别精度低的问题,提供了一种基于机器视觉的塔材识别方法及系统,通过对输电杆塔进行多视角自适应图像采集,消除了检测盲区,并对图像进行去噪和自适应光照增强处理,从而可以自适应适配强光、弱光、阴影、光照不均等复杂户外光照场景,解决了传统视觉识别对光照敏感的问题,再通过几何特征和语义特征提取,丰富了特征维度,并根据双特征融合机制进行特征融合,兼顾了塔材结构形态特征与深层语义特征,两者互补,从而规避了单一特征识别鲁棒性差的问题,显著提升了识别精度,且识别过程全程自动化,无需人工参与,大幅提升了巡检效率,消除了人工主观性误差和漏检问题,同时可以以户外复杂环境杆塔缺陷识别的大数据为依据,指引杆塔生产工艺提升方向
所述识别模块将融合特征向量与预设的标准塔材特征模板进行匹配度分析得到模板匹配度,基于模板匹配度对塔材目标图像中的对应塔材进行识别得到识别结果。
Smart Images

Figure CN122657571A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent detection technology for power equipment, specifically to a method and system for identifying tower materials based on machine vision. Background Technology
[0002] Transmission towers are the core supporting equipment of the power grid system, and the integrity and stability of their tower materials directly determine the operational safety and service life of transmission lines. Transmission towers are exposed to the complex outdoor environment for extended periods, and are susceptible to damage from factors such as wind and rain erosion, temperature differences from sunlight, icing and strong winds, construction errors, and long-term mechanical fatigue. This makes them highly prone to defects such as tower material wear, missing angle steel, component deformation, and damage to connections. If these defects are not detected and repaired in a timely manner, they can easily lead to major power safety accidents such as tower tilting, collapse, and short circuits.
[0003] Currently, tower material defect detection in the industry is mainly divided into two categories: manual inspection and traditional machine vision inspection. Both methods have significant technical shortcomings. Traditional manual inspection relies on inspectors climbing or observing from the ground to identify tower material defects, resulting in extremely low inspection efficiency, significant subjective differences, high risks associated with high-altitude operations, easy omission of small defects at long distances, and inability to achieve full-area, full-coverage detection. This makes it difficult to meet the needs of modern power grids for large-scale, routine, and high-precision inspections. Existing machine vision-based tower material detection technologies mostly use single feature matching recognition algorithms, which are not optimized for complex outdoor environments. The recognition process is easily affected by changes in lighting intensity, shadow occlusion, drone shooting angle deviations, and differences in shooting scale, leading to failure in tower material feature extraction, low recognition accuracy, and high false / false detection rates. This makes it impossible to achieve stable, high-precision automated detection in complex outdoor scenarios. Therefore, there is an urgent need for an intelligent tower material recognition method that can adapt to lighting interference, resist shooting angle distortion, and achieve high recognition accuracy.
[0004] Chinese patent, publication number CN121053109A, discloses a method and device for detecting missing tower materials on power transmission towers based on 3D point cloud data. The method involves acquiring 3D point cloud data of the power transmission tower and performing noise reduction and thinning processing; classifying the 3D point cloud; using a clustering algorithm to separate different tower material components based on the tower point cloud categories, disassembling the tower into independent tower material components and labeling their geometric parameters; locating the intersection points of the tower material components, selecting density peak points as candidate nodes, and verifying the number of connected tower material components and geometric angle conditions; converting the 3D point cloud into a topology graph connecting nodes and edges, and verifying the topology graph to ensure it conforms to engineering standards through rule verification; detecting the topology graph based on preset rules; and for detected abnormal nodes / edges, deriving the theoretical location and parameters of missing tower material components, generating a 3D visualization report, and labeling the missing locations. While this method can identify tower material missing anomalies to a certain extent, it struggles to identify anomalies such as surface erosion and structural wear that are not obvious. Summary of the Invention
[0005] This invention addresses the problem that existing methods for identifying transmission tower materials are susceptible to changes in complex outdoor environments, leading to distorted feature extraction and low identification accuracy. It provides a machine vision-based tower material identification method and system. By performing multi-view adaptive image acquisition of transmission towers, blind spots are eliminated. The images undergo denoising and adaptive lighting enhancement, enabling adaptive adaptation to complex outdoor lighting scenarios such as strong light, weak light, shadows, and uneven lighting. This solves the problem of traditional visual recognition being sensitive to lighting. Furthermore, geometric and semantic features are extracted to enrich the feature dimensions. Feature fusion is performed using a dual-feature fusion mechanism, taking into account both the structural morphology and deep semantic features of the tower materials. This complementary approach avoids the poor robustness of single-feature recognition, significantly improving identification accuracy. The entire identification process is automated, requiring no human intervention, greatly improving inspection efficiency and eliminating subjective errors and missed detections. Simultaneously, the large-scale data on tower defect identification in complex outdoor environments can guide improvements in tower production processes.
[0006] In a first aspect, one technical solution provided in this embodiment of the invention is: a tower material identification method based on machine vision, comprising the following steps: S1. Perform multi-view adaptive image acquisition on the transmission tower to be inspected to obtain the original image of the tower material; S2. Denoising and adaptive illumination enhancement are performed on the original image to obtain the target image of the tower material; S3. Based on the geometric feature extraction mechanism, geometric features are extracted from the tower material target image to obtain stable geometric features of the tower material; based on the semantic feature extraction mechanism, deep semantic features are extracted from the tower material target image to obtain multi-dimensional deep semantic features; S4. Perform feature fusion on the stable geometric features and multi-dimensional deep semantic features of the tower material to obtain a fused feature vector; S5. Perform a matching degree analysis between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree. Based on the template matching degree, identify the corresponding tower material in the tower material target image to obtain the identification result.
[0007] This solution utilizes multi-view adaptive image acquisition to eliminate the problems of occlusion and blind spots in single-view shooting, thereby achieving full-coverage image acquisition of tower materials. Image denoising and adaptive illumination enhancement effectively suppress interference from outdoor noise, strong light, shadows, and uneven illumination, restoring the true texture of the tower materials and overcoming the light sensitivity limitations of traditional visual algorithms. A dual-feature parallel extraction mechanism simultaneously acquires rotation- and scale-distortion-resistant geometric features, as well as deep semantic features characterizing tower component properties, topology, and defect textures, enriching the feature dimensions. Furthermore, a data-driven optimization-based weighted fusion method complements the advantages of both types of features, avoiding the poor robustness of single-feature recognition. Through feature fusion and standard tower material template matching analysis, accurate tower material identification can be achieved, thoroughly improving the low efficiency, high subjectivity, and easy omissions of manual inspection, as well as the insufficient accuracy of traditional image recognition. This significantly enhances the stability and accuracy of tower material detection in complex outdoor scenarios.
[0008] Preferably, in step S1, the original image of the transmission tower to be inspected is obtained by multi-view adaptive image acquisition, including the following steps: Based on the preset three-dimensional trajectory of the transmission tower, a drone is used to collect panoramic images of the transmission tower from the front, back, left, right, and top views as the original images of the tower material.
[0009] In this solution, a drone is controlled along a preset three-dimensional trajectory to collect panoramic images of the transmission tower from multiple perspectives, including front, back, left, right, and overhead views. This comprehensively covers all tower components, including main members, diagonal members, crossarms, and tower feet, effectively solving the problems of component occlusion and blind spots in single-view photography. At the same time, it significantly reduces visual interference caused by shooting angle shift and perspective distortion, providing a complete and reliable raw data source for subsequent image preprocessing and geometric and semantic feature extraction.
[0010] Preferably, in step S2, the original image is subjected to denoising and adaptive illumination enhancement processing to obtain the tower material target image, including the following steps: The pixels of the tower material and its fixed edge area in the original image are retained, and the pixels of the remaining areas are removed to obtain a denoised image; If the light intensity in the light and shadow regions of the denoised image is greater than or equal to the light threshold, then the corresponding region is light weakened; if the light intensity in the light and shadow regions of the denoised image is less than the light threshold, then the corresponding region is light enhanced to obtain the tower material target image.
[0011] This solution achieves precise denoising by selectively retaining pixels in key areas of the tower material and its edges while removing invalid background pixels. This effectively filters out complex background interference such as sky, vegetation, and mountains in outdoor scenes, thereby removing irrelevant image noise and accurately locking the effective detection area of the tower material. It avoids redundant background information interfering with subsequent feature extraction. By implementing zoned adaptive illumination adjustment on the denoised image, the illumination of overly bright areas is weakened while the illumination of shadowy and weakly lit areas is strengthened, dynamically balancing the overall image illumination distribution. This completely solves the defects of traditional visual algorithms that are easily affected by strong light, shadows, and uneven illumination, thereby optimizing image quality, restoring the real texture details of the tower material, and ensuring the accuracy and stability of subsequent geometric feature and deep semantic feature extraction.
[0012] Preferably, in step S3, geometric feature extraction is performed on the target image of the tower material based on a geometric feature extraction mechanism to obtain stable geometric features of the tower material, including the following steps: A multi-scale spatial image is obtained by constructing a Gaussian scale space from the tower material target image, and a set of key feature points of the tower material is obtained by performing spatial localization on the multi-scale spatial image. Calculate the gradient magnitude and gradient direction of all pixels in the neighborhood of the multi-scale spatial image corresponding to all tower material feature key points in the tower material feature key point set, construct an orientation histogram, and take the peak direction as the main direction of the key point. Centered on the key points of the tower material features, rotate the corresponding multi-scale spatial image according to the main direction of the key points to extract a neighborhood window of fixed pixel size, and divide it according to a fixed ratio to obtain key sub-regions. For each key sub-region, calculate the gradient histogram of n directions to obtain the geometric feature descriptor. The geometric feature descriptor is normalized to obtain the stable geometric features of the tower material.
[0013] This solution constructs a Gaussian scale space and a difference space to perform multi-scale feature detection and key point localization. It can adaptively adapt to image scale changes caused by different drone shooting distances, thereby accurately locating core structural key points such as tower angle steel, truss nodes, and edge contours. By calculating the gradient information of the key point's neighborhood and generating a direction histogram, the main direction of the key point is determined, making the geometric features rotationally invariant and effectively solving the feature failure problem caused by the tilt of the shooting angle and the rotation of the image during outdoor inspections. By rotating along the main direction to extract the neighborhood window and statistically analyzing the gradient features in different regions, the local structural morphology of the tower can be finely characterized, thus completely preserving the rigid geometric feature information of the tower. Finally, normalization processing is used to eliminate interference from weak lighting and pixel amplitude, outputting highly robust and stable geometric features of the tower. This mechanism overcomes the distortion sensitivity of traditional recognition algorithms from multiple dimensions of scale, angle, and detail, providing stable and reliable structural feature support for subsequent feature fusion and accurate matching of tower materials.
[0014] Preferably, in S3, deep semantic features are extracted from the tower material target image based on the semantic feature extraction mechanism to obtain multi-dimensional deep semantic features, including the following steps: The tower material target image is normalized to obtain a tower material normalized image, and a two-dimensional convolution operation is performed on the tower material normalized image to obtain a low-order semantic feature map; The low-order semantic feature map is normalized by residual units to obtain a deep two-dimensional feature map. The scalar feature values of all deep two-dimensional feature maps are then processed by global average pooling. All scalar feature values are then concatenated in the pooling order to obtain multi-dimensional deep semantic features.
[0015] In this scheme, the target image of the tower material is normalized, and then a low-order semantic feature map is extracted through two-dimensional convolution. This not only eliminates the interference caused by differences in image amplitude, but also effectively captures shallow visual information such as surface texture, minor wear marks, and edge details of the tower material. By using residual unit standardization to extract deep two-dimensional feature maps layer by layer, the problem of gradient vanishing in deep networks can be solved. This enables efficient learning of high-order semantic features such as tower material component properties, rod topology, and defect morphology, making up for the shortcomings of geometric features in identifying the relationship between subtle defects and the structure. Finally, global average pooling is used to compress the two-dimensional feature map into scalar feature values and concatenate them sequentially to generate lightweight and high-dimensional stable semantic features. This mechanism can effectively resist the interference of outdoor residual light and complex backgrounds, and accurately identify defect features such as tower material corrosion, wear, and missing components. It can complement geometric features, greatly enriching the feature expression dimension and providing reliable deep semantic support for subsequent feature fusion and high-precision tower material identification and defect detection.
[0016] Preferably, in S4, the stable geometric features and multi-dimensional deep semantic features of the tower material are fused to obtain a fused feature vector, including the following steps: Based on a preset search step size, corresponding candidate weighted optimization coefficients are generated sequentially. The candidate weighted optimization coefficients are identified and evaluated to obtain the corresponding identification score. The value range of the weighted optimization coefficients is between 0 and 1. The fusion feature vector is obtained by weighting and summing the tower material stability geometric features and multidimensional deep semantic features using the candidate weighted optimization coefficient corresponding to the highest recognition score as the target optimization coefficient.
[0017] In this scheme, candidate weighted optimization coefficients are generated in batches within the 0-1 range based on a fixed search step size. The recognition score is then quantitatively output based on the tower material recognition effect. This abandons the traditional crude method of manually setting weights based on experience and achieves data-driven intelligent optimization with fused weights. By selecting the coefficient corresponding to the highest recognition score as the target optimization coefficient, the weight ratio of geometric features and semantic features can be adaptively matched according to the complex outdoor scene characteristics of transmission towers. The dual-feature fusion is completed by weighted summation of the optimal coefficients. This fully combines the advantages of geometric features in resisting scale distortion and shooting rotation interference with the advantages of semantic features in resisting lighting interference and recognizing component topology and subtle defects, thereby making up for the shortcomings of insufficient single feature representation capabilities.
[0018] Preferably, the candidate feature vectors are identified and evaluated to obtain the corresponding identification score, including the following steps: Acquire images of normal and abnormal tower materials under different lighting conditions, shooting angles, and backgrounds; For normal tower material images, label the tower material location and component category; for abnormal tower material images, label the tower material location, component category, and abnormality type. Normal tower material images and abnormal tower material images are randomly mixed to form the total dataset. Geometric and semantic features are extracted from the total dataset, and the extracted features are weighted and summed according to weighted optimization coefficients to obtain the fused test vector. The test matching degree is obtained by performing a matching degree analysis between the fused test vector and the preset standard tower material feature template. Based on the test matching degree, the total dataset is identified to obtain the number of correct identifications, false detections, and missed detections. The identification score is determined based on the number of correct identifications, false detections, and missed detections.
[0019] In this scheme, a total dataset is constructed by randomly mixing a large number of known samples, which effectively improves data diversity and avoids evaluation bias caused by data solidification. By extracting features in batches and generating fusion test vectors by combining candidate weighting coefficients, template matching tests are completed, and the number of correct identifications, false detections, and false negatives are accurately counted. Recognition scores are generated by quantifying detection indicators, which realizes an objective and quantitative comparison of the performance of each candidate coefficient. This completely eliminates the subjectivity of manually setting weights based on experience, and can accurately select the optimal weighting coefficients that are suitable for complex scenarios. It can maximize the advantages of dual feature fusion and effectively improve the algorithm's scenario generalization ability and tower material defect recognition accuracy.
[0020] As a preferred method, the identification score is determined based on the number of correct identifications, false positives, and false negatives, including the following steps: The sum of the correct identifications and false positives is used as the result sample, and the sum of the correct identifications and false negatives is used as the recall sample. The accuracy is obtained by comparing the number of correctly identified samples with the number of results samples, and the recall is obtained by comparing the number of correctly identified samples with the number of recalled samples. A recognition score is obtained by balancing the recognition precision and recognition recall.
[0021] In this scheme, the accuracy and recall are calculated by statistically analyzing the number of correct identifications, false positives, and false negatives. This allows for the quantification of tower material identification performance from two core dimensions: false positives and false negatives. By balancing the accuracy and recall to obtain the identification score, the scheme avoids the one-sidedness of a single evaluation indicator. This enables an objective and accurate quantitative evaluation of the performance of the candidate weighted coefficients, and also balances the overall identification accuracy and defect detection capability. This effectively improves the scenario adaptability and robustness of the dual-feature fusion strategy, providing a reliable evaluation basis for high-precision tower material identification in complex outdoor scenarios.
[0022] Preferably, in step S5, the fused feature vector is compared with a preset standard tower material feature template to obtain a template matching degree. Based on the template matching degree, the corresponding tower material in the tower material target image is identified to obtain the identification result, including the following steps: The cosine similarity between the fused feature vector and the standard tower material feature template is calculated as the template matching degree. If the template matching degree is greater than the matching threshold, the tower material in the transmission tower to be tested is normal. If the template matching degree is less than or equal to the matching threshold, the tower material in the transmission tower to be tested is abnormal.
[0023] In this solution, feature matching degree is calculated by cosine similarity, which avoids interference from feature amplitude and ensures stable and reliable matching results. At the same time, based on the preset matching threshold for the judgment standard, the normal and abnormal states of tower materials can be clearly distinguished. It also combines the advantages of dual feature fusion, which can quickly complete tower material identification and defect judgment. The logic is simple and efficient, thereby further reducing the probability of misjudgment and improving the overall detection accuracy.
[0024] Secondly, one technical solution provided in this embodiment of the invention is: a tower material identification system based on machine vision, including an image acquisition module, an image processing module, a feature extraction module, a feature fusion module, and an identification module; The image acquisition module performs multi-view adaptive image acquisition on the transmission tower to be inspected to obtain the original image of the tower material; The image processing module performs noise reduction and adaptive illumination enhancement on the original image to obtain the target image of the tower material. The feature extraction module performs geometric feature extraction on the tower material target image based on the geometric feature extraction mechanism to obtain stable geometric features of the tower material, and performs deep semantic feature extraction on the tower material target image based on the semantic feature extraction mechanism to obtain multi-dimensional deep semantic features; The feature fusion module performs feature fusion on the stable geometric features and multi-dimensional deep semantic features of the tower material to obtain a fused feature vector; The recognition module performs a matching degree analysis between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree, and identifies the corresponding tower material in the tower material target image based on the template matching degree to obtain the recognition result.
[0025] In this solution, a corresponding system is built to integrate the tower material identification method, thereby enabling human-computer interaction and improving the user experience.
[0026] The beneficial effects of this invention are as follows: By performing multi-view adaptive image acquisition on transmission towers, this invention eliminates detection blind spots and performs noise reduction and adaptive illumination enhancement processing on the images. This allows it to adapt to complex outdoor lighting scenarios such as strong light, weak light, shadows, and uneven illumination, solving the problem of traditional visual recognition being sensitive to light. Furthermore, by extracting geometric and semantic features, the feature dimensions are enriched, and feature fusion is performed according to a dual-feature fusion mechanism, taking into account both the structural morphology features of the tower material and deep semantic features. The two complement each other, thereby avoiding the problem of poor robustness of single feature recognition, significantly improving recognition accuracy. Moreover, the recognition process is fully automated, requiring no human intervention, greatly improving inspection efficiency and eliminating the problems of subjective human error and missed detection.
[0027] The above description of the invention is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description
[0028] Other features, objects, and advantages of the invention will become more apparent from the following detailed description of non-limiting embodiments with reference to the accompanying drawings. The drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings.
[0029] Figure 1 This is a flowchart of a tower material identification method based on machine vision according to the present invention; Figure 2 This is a schematic diagram of a tower material identification system based on machine vision according to the present invention. Detailed Implementation
[0030] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only one preferred embodiment of this invention and are only used to explain this invention. They do not limit the scope of protection of this invention. All other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this invention.
[0031] Before discussing the exemplary embodiments in more detail, it should be mentioned that some exemplary embodiments are described as processes or methods depicted as flowcharts. Although the flowcharts describe the operations (or steps) as sequential processes, many of the operations (or steps) can be performed in parallel, concurrently, or simultaneously. Furthermore, the order of the operations can be rearranged. The process can be terminated when its operation is completed, but it may also have additional steps not included in the figures; the process may correspond to a method, function, procedure, subroutine, subroutine, etc.
[0032] Example 1: To address the problem that existing methods for identifying transmission tower materials are easily affected by changes in the complex outdoor environment, leading to distorted feature extraction and low identification accuracy, this example provides a machine vision-based method for identifying tower materials, such as... Figure 1 As shown, it includes the following steps: S1: Perform multi-view adaptive image acquisition on the transmission tower to be inspected to obtain the original image of the tower material.
[0033] In this embodiment, the original image of the transmission tower to be inspected is obtained by multi-view adaptive image acquisition, including the following steps: Based on the preset three-dimensional trajectory of the transmission tower, a drone is used to collect panoramic images of the transmission tower from the front, back, left, right, and top views as the original images of the tower material.
[0034] Specifically, in this embodiment, a drone equipped with a high-definition visible light camera can be used to collect panoramic images of the tower from five dimensions: front, back, left, right, and top, along a preset three-dimensional trajectory of the tower. This fully covers all tower components, including the main tower body, diagonal members, crossarms, tower head, and tower feet, achieving image acquisition without blind spots and providing a complete data source for subsequent feature extraction.
[0035] This embodiment uses a drone to capture panoramic images of the power transmission tower from multiple perspectives, including front, back, left, right, and overhead views, by controlling the drone along a preset three-dimensional trajectory. This can comprehensively cover all tower components, including main members, diagonal members, crossarms, and tower feet, effectively solving the problems of component occlusion and blind spots in single-view shooting. At the same time, it also significantly reduces visual interference caused by shooting angle offset and perspective distortion, providing a complete and reliable original data source for subsequent image preprocessing and geometric and semantic feature extraction.
[0036] S2: Denoising and adaptive illumination enhancement are performed on the original image to obtain the target image of the tower material.
[0037] In this embodiment, the target image of the tower material is obtained by performing denoising and adaptive illumination enhancement processing on the original image, including the following steps: The original image retains pixels within a fixed area of the tower structure and its edges, while removing pixels from other areas to obtain a denoised image. Specifically, a two-dimensional Gaussian kernel function is used to smooth the image and remove noise from the outdoor image while preserving the detailed features of the tower structure's edges. The Gaussian kernel function formula is as follows: in, , Image pixel coordinates, This is the Gaussian scaling parameter, used to control the smoothness of the filter. The filtered and denoised image is obtained through convolution operation: in, The raw images captured by the drone. This is the denoised image. This is the convolution operator.
[0038] If the light intensity in the shaded areas of the denoised image is greater than or equal to the light threshold, the corresponding areas are weakened. If the light intensity in the shaded areas of the denoised image is less than the light threshold, the corresponding areas are enhanced to obtain the target image of the tower material. Specifically, a single-scale Retinex illumination enhancement formula is used to remove the light interference components from the image, restore the inherent texture and color features of the tower material, and eliminate the influence of brightness deviation. The formula is as follows: in, The image shows the tower material target image after adaptive enhancement based on illumination. It is a logarithmic transform operator used to compress the dynamic range of an image and achieve illumination equalization.
[0039] This embodiment achieves precise denoising by selectively retaining pixels in key areas of the tower material and its edges while removing invalid background pixels. This effectively filters out complex background interference such as sky, vegetation, and mountains in outdoor scenes, thereby removing irrelevant image noise and accurately locking the effective detection area of the tower material. It avoids redundant background information interfering with subsequent feature extraction. By implementing zoned adaptive illumination adjustment on the denoised image, the illumination of overly bright areas is weakened while the illumination of shadowy and weakly lit areas is strengthened, dynamically balancing the illumination distribution of the entire image. This completely solves the defects of traditional visual algorithms that are easily affected by strong light, shadows, and uneven illumination, thereby optimizing image quality, restoring the real texture details of the tower material, and ensuring the accuracy and stability of subsequent geometric feature and deep semantic feature extraction.
[0040] S3: Based on the geometric feature extraction mechanism, geometric features are extracted from the tower material target image to obtain stable geometric features of the tower material; based on the semantic feature extraction mechanism, deep semantic features are extracted from the tower material target image to obtain multidimensional deep semantic features.
[0041] In this embodiment, geometric feature extraction is performed on the target image of the tower material based on a geometric feature extraction mechanism to obtain stable geometric features of the tower material, including the following steps: A multi-scale spatial image is obtained by constructing a Gaussian scale space from the tower material target image, and a set of key feature points of the tower material is obtained by performing spatial localization on the multi-scale spatial image. Calculate the gradient magnitude and gradient direction of all pixels in the neighborhood of the multi-scale spatial image corresponding to all tower material feature key points in the tower material feature key point set, construct an orientation histogram, and take the peak direction as the main direction of the key point. Centered on the key points of the tower material features, rotate the corresponding multi-scale spatial image according to the main direction of the key points to extract a neighborhood window of fixed pixel size, and divide it according to a fixed ratio to obtain key sub-regions. For each key sub-region, calculate the gradient histogram of n directions to obtain the geometric feature descriptor. Normalizing the geometric feature descriptors yields the stable geometric features of the tower structure. Specifically, the SIFT algorithm is used to extract these stable geometric features, which possess resistance to rotation, scaling, and viewpoint changes. This accurately characterizes the structural morphology of the tower structure, including angle steel, trusses, and crossarms. First, a Gaussian scale space construction formula is used: in, For multi-scale spatial images, For the preprocessed and enhanced target image of the tower material; Then, the Gaussian difference scale space (DOG) formula is used to locate key feature points of the tower material through difference space, avoiding scale interference: in, For fixed-scale scaling factors, The extreme point is the key characteristic of the tower material.
[0042] The process of obtaining geometric feature descriptors is as follows: First, taking the keypoints detected by DOG as the center, and at their corresponding scales... In the image, calculate the gradient magnitude of all pixels in the neighborhood. and gradient direction The gradient direction histogram is statistically analyzed, and the peak direction is taken as the principal direction of the key point. The calculation formula is as follows: Then, centering on the keypoint, rotate the window in the main direction to extract a 16×16 pixel neighborhood region (using the corresponding scale). (Image), then divide the 16×16 window into 4×4=16 small regions, and count the gradient histogram in 8 directions for each small region; 16 regions × 8 dimensions = 128-dimensional data, i.e., 128 geometric feature descriptors. Finally, the 128-dimensional data is normalized to eliminate illumination interference, resulting in the final SIFT tower material stability geometric features: .
[0043] This embodiment constructs a Gaussian scale space and a difference space to complete multi-scale feature detection and key point localization. It can adaptively adapt to image scale changes caused by different drone shooting distances, thereby accurately locating core structural key points such as tower angle steel, truss nodes, and edge contours. By calculating the gradient information of the key point's neighborhood and generating a direction histogram, the main direction of the key point is determined, making the geometric features rotationally invariant. This effectively solves the feature failure problem caused by the tilt of the shooting angle and the rotation of the image during outdoor inspections. By rotating along the main direction to extract the neighborhood window and statistically analyzing the gradient features in different regions, the local structural morphology of the tower can be finely characterized, thus completely preserving the rigid geometric feature information of the tower. Finally, normalization processing is used to eliminate interference from weak lighting and pixel amplitude, outputting highly robust stable geometric features of the tower. This mechanism overcomes the distortion sensitivity of traditional recognition algorithms from multiple dimensions of scale, angle, and detail, providing stable and reliable structural feature support for subsequent feature fusion and accurate matching of tower materials.
[0044] In this embodiment, multi-dimensional deep semantic features are obtained by performing deep semantic feature extraction on the tower material target image based on the semantic feature extraction mechanism, including the following steps: The tower material target image is normalized to obtain a tower material normalized image, and a two-dimensional convolution operation is performed on the tower material normalized image to obtain a low-order semantic feature map; The low-order semantic feature maps are normalized using residual units to obtain deep two-dimensional feature maps. Global average pooling is then applied to all deep two-dimensional feature maps to process scalar feature values. Finally, all scalar feature values are concatenated according to the pooling order to obtain multi-dimensional deep semantic features. Specifically, a ResNet network is used. As the sole input to the ResNet network, the semantic feature extraction branch first performs pixel normalization on the input image. The pixel values of the original enhanced image range from [0, 255]. The deep learning network needs to perform normalization to map the pixels to the standard range and eliminate the influence of amplitude differences. The normalization formula used is: in, This is the normalized image of the tower material, where all pixel values range from [0,1].
[0045] Then, a two-dimensional convolution operation is performed on the normalized image of the tower material to extract low-level visual features such as surface texture, local edges, and brightness distribution. The convolution operation formula is as follows: in, The first layer convolutional kernel (a learnable parameter of the network). It is a set of feature maps output by the first convolution layer, containing basic texture information such as rust, wear marks, and edges on the surface of the tower material; Then, residual unit stacking operations are performed. Lightweight ResNet is composed of multiple sets of residual units stacked in series. The residual structure solves the gradient vanishing problem in deep networks and abstracts basic features layer by layer into high-level semantics such as component morphology, topological relationships, and defect patterns. The standard formula for a single residual unit is as follows: in, This is the input feature map of the l-th layer residual unit. The main branch transformation function (composed of convolution, activation function, and batch normalization). These are the network weight parameters for this layer. This is the output feature map of the (l+1)th layer residual unit. Short connections are used to implement residual mapping and ensure stable feature propagation in deep networks.
[0046] The residual units include shallow residual units and deep residual units. Shallow residual units further combine the convolutional features of the first layer to output mid-order semantic features, representing the morphology and local splicing relationships of independent components such as single angle steel, bolts, and connecting plates. Deep residual units are used to fuse global information and output high-order semantic features, representing the overall truss topology, component categories, and defect paradigms such as missing / damaged components of the tower material. Assuming the network stacks L sets of residual units, the final deep two-dimensional feature map of the network is obtained: Due to deep output Since it is a multi-channel two-dimensional feature map, it cannot be directly used as a feature vector. Therefore, global average pooling is used for dimensionality reduction and compression. in, , These represent the height and width of a single-channel feature map, respectively. For the c-th channel at pixel eigenvalues at that location Let c be the scalar feature value obtained after pooling the c-th channel. After traversing all channels, the pooling results of each channel are concatenated in order to form a one-dimensional feature sequence: , where m is the total dimension of the feature vector, which is determined by the lightweight ResNet structure (usually 256 or 512 dimensions). Each dimension represents a semantic feature component, corresponding to different semantic information such as component attributes, topological structure, and defect type.
[0047] This embodiment normalizes the target image of the tower material and then extracts low-order semantic feature maps through two-dimensional convolution operations. This not only eliminates the interference caused by differences in image amplitude but also effectively captures shallow visual information such as surface texture, minor wear marks, and edge details of the tower material. By using residual unit standardization to extract deep two-dimensional feature maps layer by layer, the problem of gradient vanishing in deep networks can be solved. This enables efficient learning of high-order semantic features such as tower material component properties, rod topology, and defect morphology, making up for the shortcomings of geometric features in identifying the relationship between subtle defects and the structure. Finally, global average pooling is used to compress the two-dimensional feature maps into scalar feature values and concatenate them sequentially to generate lightweight and high-dimensional stable semantic features. This mechanism can effectively resist interference from outdoor residual light and complex backgrounds and accurately identify defect features such as tower material corrosion, wear, and missing components. It can complement geometric features, greatly enriching the feature expression dimension and providing reliable deep semantic support for subsequent feature fusion and high-precision tower material identification and defect detection.
[0048] S4: The stable geometric features and multi-dimensional deep semantic features of the tower material are fused to obtain the fused feature vector.
[0049] In this embodiment, the fusion of the tower material's stable geometric features and multi-dimensional deep semantic features to obtain a fused feature vector includes the following steps: Based on a preset search step size, corresponding candidate weighted optimization coefficients are generated sequentially. The candidate weighted optimization coefficients are identified and evaluated to obtain the corresponding identification score. The value range of the weighted optimization coefficients is between 0 and 1. Using the candidate weighted optimization coefficient corresponding to the highest recognition score as the target optimization coefficient, the stable geometric features and multi-dimensional deep semantic features of the tower material are weighted and summed to obtain the fused feature vector. The weighted fusion formula is as follows: in, The weighted optimization coefficients were obtained through training and optimization using a massive dataset of tower images. This is the fused feature vector ultimately used for matching and recognition.
[0050] This embodiment generates candidate weighted optimization coefficients in batches within the 0-1 range based on a fixed search step size, and outputs a quantitative recognition score based on the tower material recognition effect. This abandons the traditional crude method of manually setting weights based on experience, and realizes data-driven intelligent optimization with fused weights. By selecting the coefficient corresponding to the highest recognition score as the target optimization coefficient, it can adaptively match the weight ratio of geometric features and semantic features according to the complex outdoor scene characteristics of transmission towers. The dual-feature fusion is completed by weighted summation of the optimal coefficients, which fully combines the advantages of geometric features in resisting scale distortion and shooting rotation interference with the advantages of semantic features in resisting lighting interference and recognizing component topology and subtle defects, thereby making up for the shortcomings of insufficient single feature representation capabilities.
[0051] In this embodiment, the candidate feature vectors are identified and evaluated to obtain the corresponding identification score, including the following steps: Acquire images of normal and abnormal tower materials under different lighting conditions, shooting angles, and backgrounds; For normal tower material images, label the tower material location and component category; for abnormal tower material images, label the tower material location, component category, and abnormality type. Normal tower material images and abnormal tower material images are randomly mixed to form the total dataset. Geometric and semantic features are extracted from the total dataset, and the extracted features are weighted and summed according to weighted optimization coefficients to obtain the fused test vector. The test matching degree is obtained by performing a matching degree analysis between the fused test vector and the preset standard tower material feature template. Based on the test matching degree, the total dataset is identified to obtain the number of correct identifications, false detections, and missed detections. The identification score is determined based on the number of correct identifications, false detections, and missed detections.
[0052] Specifically, in the weighted optimization coefficients Set a fixed search step size within the range. Generate the coefficient sequence to be tested: For each candidate, the weighted optimization coefficients are... Substituting the feature fusion formula yields the corresponding fusion test vector, and the entire process of template matching and defect determination is completed sequentially, calculating the current... Corresponding correct recognition volume False detection rate and missed detections And calculate the recognition score, and finally obtain the target optimization coefficient. .
[0053] This embodiment constructs a total dataset by randomly mixing a large number of known samples, effectively improving data diversity and avoiding evaluation bias caused by data solidification. By batch extracting features and combining them with candidate weighting coefficients to generate fusion test vectors, template matching tests are completed, and the number of correct identifications, false positives, and false negatives are accurately counted. Recognition scores are generated through quantitative detection indicators, achieving an objective and quantitative comparison of the performance of each candidate coefficient. This completely eliminates the subjectivity of manually setting weights based on experience, accurately selects the optimal weighting coefficients suitable for complex scenarios, maximizes the advantages of dual-feature fusion, and effectively improves the algorithm's scenario generalization ability and tower material defect recognition accuracy.
[0054] In this embodiment, the identification score is determined based on the number of correct identifications, false detections, and false negatives, including the following steps: The sum of the correct identifications and false positives is used as the result sample, and the sum of the correct identifications and false negatives is used as the recall sample. The accuracy is obtained by comparing the number of correctly identified samples with the number of results samples, and the recall is obtained by comparing the number of correctly identified samples with the number of recalled samples. The recognition score is obtained by balancing the recognition precision and recognition recall, and the specific formula is as follows: in, To correctly identify the quantity, False detection rate This is the number of missed detections. It is the recognition accuracy. It is about identifying recall rate. It is a scoring system.
[0055] This embodiment calculates the recognition precision and recall by statistically analyzing the number of correct identifications, false positives, and false negatives, thereby quantifying the tower material recognition effect from the two core dimensions of false positives and false negatives. By balancing the precision and recall to obtain the recognition score, the one-sidedness of a single evaluation index is avoided, thus achieving an objective and accurate quantitative evaluation of the performance of the candidate weighted coefficients. It also balances the overall recognition accuracy and defect detection capability, effectively improving the scene adaptability and robustness of the dual feature fusion strategy, and providing a reliable evaluation basis for high-precision tower material recognition in complex outdoor scenarios.
[0056] S5: Perform a matching degree analysis between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree. Based on the template matching degree, identify the corresponding tower material in the tower material target image to obtain the identification result.
[0057] In this embodiment, the matching degree analysis is performed between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree. Based on the template matching degree, the corresponding tower material in the tower material target image is identified to obtain the identification result, including the following steps: The cosine similarity between the fused feature vector and the standard tower material feature template is calculated as the template matching degree. If the template matching degree is greater than the matching threshold, the tower material in the transmission tower to be tested is normal; if the template matching degree is less than or equal to the matching threshold, the tower material in the transmission tower to be tested is abnormal. The specific formula for calculating the cosine similarity is as follows: in is the feature vector of the standard tower material template; S is the template matching degree, with a value range of [0,1]; the higher the similarity, the better the integrity of the tower material in the current area, and a matching threshold is set. If the similarity is lower than the threshold, it is determined that there is an anomaly in the tower material in the area.
[0058] When faced with situations such as missing tower materials, as one implementation method, for defects such as wear, missing, or localized damage to tower materials, the defect area can be extracted through image morphology operations to achieve quantitative judgment. The core calculation formula is as follows: In the formula, The area of the extracted defect region. A preset defect judgment threshold is set. If the defect area is greater than the threshold, the tower material is judged to have wear or missing defects.
[0059] This embodiment calculates the feature matching degree using cosine similarity, which can avoid interference from feature amplitude and ensure stable and reliable matching results. At the same time, relying on the preset matching threshold to classify the judgment criteria, it can clearly distinguish between the normal and abnormal states of tower materials. It also combines the advantages of dual feature fusion, which can quickly complete tower material identification and defect judgment. The logic is simple and efficient, thereby further reducing the probability of misjudgment and improving the overall detection accuracy.
[0060] Example 2: This example also provides a tower material identification system based on machine vision, such as... Figure 2 As shown, it includes an image acquisition module, an image processing module, a feature extraction module, a feature fusion module, and a recognition module; The image acquisition module performs multi-view adaptive image acquisition on the transmission tower to be inspected to obtain the original image of the tower material; The image processing module performs noise reduction and adaptive illumination enhancement on the original image to obtain the target image of the tower material. The feature extraction module performs geometric feature extraction on the tower material target image based on the geometric feature extraction mechanism to obtain stable geometric features of the tower material, and performs deep semantic feature extraction on the tower material target image based on the semantic feature extraction mechanism to obtain multi-dimensional deep semantic features; The feature fusion module performs feature fusion on the stable geometric features and multi-dimensional deep semantic features of the tower material to obtain a fused feature vector; The recognition module performs a matching degree analysis between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree, and identifies the corresponding tower material in the tower material target image based on the template matching degree to obtain the recognition result.
[0061] This embodiment integrates the tower material identification method in this solution by constructing a corresponding system, thereby realizing human-computer interaction and improving the user experience.
[0062] As can be seen from the above embodiments, it has at least the following substantial effects: (1) This invention can eliminate the problems of single-view shooting occlusion and detection blind spots by multi-view adaptive image acquisition, thereby realizing full-area full-coverage image acquisition of tower materials; (2) This invention can effectively suppress outdoor noise, strong light, shadow, uneven lighting and other interferences through image denoising and adaptive lighting enhancement processing, thereby restoring the real texture of tower materials and solving the problem of light sensitivity of traditional visual algorithms. (3) This invention uses a dual-feature parallel extraction mechanism to simultaneously acquire geometric features that resist rotation and scale distortion, as well as deep semantic features that can characterize the properties, topology, and defect texture of tower components, thus enriching the feature dimensions. At the same time, through a data-driven optimization weighted fusion method, the advantages of the two types of features can be complemented, thus avoiding the problem of poor robustness of single feature recognition. (4) By combining features with standard tower material templates for matching analysis, this invention can accurately identify tower materials, thoroughly improve the problems of low efficiency, strong subjectivity and easy omission of manual inspection, as well as the shortcomings of insufficient accuracy of traditional image recognition, and significantly improve the stability and accuracy of tower material detection in complex outdoor scenarios.
[0063] The specific embodiments described above are preferred embodiments of the present invention and are not intended to limit the specific scope of the present invention. The scope of the present invention includes, but is not limited to, these specific embodiments. All equivalent changes made in accordance with the shape and structure of the present invention are within the protection scope of the present invention.
Claims
1. A tower material identification method based on machine vision, characterized in that: Includes the following steps: S1. Perform multi-view adaptive image acquisition on the transmission tower to be inspected to obtain the original image of the tower material; S2. Denoising and adaptive illumination enhancement are performed on the original image to obtain the target image of the tower material; S3. Based on the geometric feature extraction mechanism, geometric features are extracted from the tower material target image to obtain stable geometric features of the tower material; based on the semantic feature extraction mechanism, deep semantic features are extracted from the tower material target image to obtain multi-dimensional deep semantic features; S4. Perform feature fusion on the stable geometric features and multi-dimensional deep semantic features of the tower material to obtain a fused feature vector; S5. Perform a matching degree analysis between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree. Based on the template matching degree, identify the corresponding tower material in the tower material target image to obtain the identification result.
2. The tower material identification method based on machine vision according to claim 1, characterized in that: In S1, multi-view adaptive image acquisition is performed on the transmission tower to be inspected to obtain the original image of the tower material, including the following steps: Based on the preset three-dimensional trajectory of the transmission tower, a drone is used to collect panoramic images of the transmission tower from the front, back, left, right, and top views as the original images of the tower material.
3. The tower material identification method based on machine vision according to claim 1, characterized in that: In S2, the original image is denoised and adaptively enhanced to obtain the target image of the tower material, including the following steps: The pixels of the tower material and its fixed edge area in the original image are retained, and the pixels of the remaining areas are removed to obtain a denoised image; If the light intensity in the light and shadow regions of the denoised image is greater than or equal to the light threshold, then the corresponding region is light weakened; if the light intensity in the light and shadow regions of the denoised image is less than the light threshold, then the corresponding region is light enhanced to obtain the tower material target image.
4. The tower material identification method based on machine vision according to claim 1, characterized in that: In S3, geometric feature extraction is performed on the target image of the tower material based on the geometric feature extraction mechanism to obtain the stable geometric features of the tower material, including the following steps: A multi-scale spatial image is obtained by constructing a Gaussian scale space from the tower material target image, and a set of key feature points of the tower material is obtained by performing spatial localization on the multi-scale spatial image. Calculate the gradient magnitude and gradient direction of all pixels in the neighborhood of the multi-scale spatial image corresponding to all tower material feature key points in the tower material feature key point set, construct an orientation histogram, and take the peak direction as the main direction of the key point. Centered on the key points of the tower material features, rotate the corresponding multi-scale spatial image according to the main direction of the key points to extract a neighborhood window of fixed pixel size, and divide it according to a fixed ratio to obtain key sub-regions. For each key sub-region, calculate the gradient histogram of n directions to obtain the geometric feature descriptor. The geometric feature descriptor is normalized to obtain the stable geometric features of the tower material.
5. The tower material identification method based on machine vision according to claim 1, characterized in that: In S3, deep semantic features are extracted from the tower material target image based on the semantic feature extraction mechanism to obtain multi-dimensional deep semantic features, including the following steps: The tower material target image is normalized to obtain a tower material normalized image, and a two-dimensional convolution operation is performed on the tower material normalized image to obtain a low-order semantic feature map; The low-order semantic feature map is normalized by residual units to obtain a deep two-dimensional feature map. The scalar feature values of all deep two-dimensional feature maps are then processed by global average pooling. All scalar feature values are then concatenated in the pooling order to obtain multi-dimensional deep semantic features.
6. The tower material identification method based on machine vision according to claim 1, characterized in that: In S4, the stable geometric features and multi-dimensional deep semantic features of the tower material are fused to obtain a fused feature vector, including the following steps: Based on a preset search step size, corresponding candidate weighted optimization coefficients are generated sequentially. The candidate weighted optimization coefficients are identified and evaluated to obtain the corresponding identification score. The value range of the weighted optimization coefficients is between 0 and 1. The fusion feature vector is obtained by weighting and summing the tower material stability geometric features and multidimensional deep semantic features using the candidate weighted optimization coefficient corresponding to the highest recognition score as the target optimization coefficient.
7. The tower material identification method based on machine vision according to claim 6, characterized in that: The process of identifying and evaluating candidate feature vectors to obtain corresponding identification scores includes the following steps: Acquire images of normal and abnormal tower materials under different lighting conditions, shooting angles, and backgrounds; For normal tower material images, label the tower material location and component category; for abnormal tower material images, label the tower material location, component category, and abnormality type. Normal tower material images and abnormal tower material images are randomly mixed to form the total dataset. Geometric and semantic features are extracted from the total dataset, and the extracted features are weighted and summed according to weighted optimization coefficients to obtain the fused test vector. The test matching degree is obtained by performing a matching degree analysis between the fused test vector and the preset standard tower material feature template. Based on the test matching degree, the total dataset is identified to obtain the number of correct identifications, false detections, and missed detections. The identification score is determined based on the number of correct identifications, false detections, and missed detections.
8. The tower material identification method based on machine vision according to claim 7, characterized in that: The identification score is determined based on the number of correct identifications, false positives, and false negatives, including the following steps: The sum of the correct identifications and false positives is used as the result sample, and the sum of the correct identifications and false negatives is used as the recall sample. The accuracy is obtained by comparing the number of correctly identified samples with the number of results samples, and the recall is obtained by comparing the number of correctly identified samples with the number of recalled samples. A recognition score is obtained by balancing the recognition precision and recognition recall.
9. The tower material identification method based on machine vision according to claim 1, characterized in that: In S5, the fused feature vector is matched with a preset standard tower material feature template to obtain the template matching degree. Based on the template matching degree, the corresponding tower material in the tower material target image is identified to obtain the identification result, including the following steps: The cosine similarity between the fused feature vector and the standard tower material feature template is calculated as the template matching degree. If the template matching degree is greater than the matching threshold, the tower material in the transmission tower to be tested is normal. If the template matching degree is less than or equal to the matching threshold, the tower material in the transmission tower to be tested is abnormal.
10. A tower material identification system based on machine vision, applicable to the tower material identification method based on machine vision as described in any one of claims 1-9, characterized in that: It includes an image acquisition module, an image processing module, a feature extraction module, a feature fusion module, and a recognition module; The image acquisition module performs multi-view adaptive image acquisition on the transmission tower to be inspected to obtain the original image of the tower material; The image processing module performs noise reduction and adaptive illumination enhancement on the original image to obtain the target image of the tower material. The feature extraction module performs geometric feature extraction on the tower material target image based on the geometric feature extraction mechanism to obtain stable geometric features of the tower material, and performs deep semantic feature extraction on the tower material target image based on the semantic feature extraction mechanism to obtain multi-dimensional deep semantic features; The feature fusion module performs feature fusion on the stable geometric features and multi-dimensional deep semantic features of the tower material to obtain a fused feature vector; The recognition module performs a matching degree analysis between the fused feature vector and the preset standard tower material feature template to obtain the template matching degree, and identifies the corresponding tower material in the tower material target image based on the template matching degree to obtain the recognition result.
Citation Information
Patent Citations
Transmission tower material missing detection method and device based on three-dimensional point cloud data
CN121053109A