A method and program product for detection of missing lattice tower members

By simulating binocular vision with a monocular camera to generate depth maps and optimizing disparity maps, and combining self-supervised learning and incremental learning, the high cost and environmental adaptability problems of communication tower component missing detection are solved, achieving efficient and accurate component missing identification.

CN120932131BActive Publication Date: 2026-08-04CHINA TOWER CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
CHINA TOWER CO LTD
Filing Date
2025-07-16
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

Existing methods for detecting missing components on communication towers are costly, have poor adaptability, are greatly affected by environmental factors, and suffer from mismatches and inaccurate depth calculations in traditional visual methods, making them difficult to promote on a large scale. Furthermore, they lack the ability to identify key areas.

Method used

A monocular camera is used to simulate binocular vision, generate depth maps and optimize disparity maps. By combining self-supervised learning and incremental learning, a convolutional neural network is used to identify missing members and dynamically update the model to adapt to environmental changes.

Benefits of technology

It improves the accuracy and robustness of missing member detection, reduces detection costs, adapts to complex environments, reduces dependence on labeled data, and achieves efficient missing member identification.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120932131B_ABST
    Figure CN120932131B_ABST
Patent Text Reader

Abstract

The application belongs to the technical field of communication structure health monitoring, and provides a method and program product for detecting missing of lattice tower members, the method comprising: acquiring and processing image data shot by a monocular camera to obtain a depth map; segmenting and fusing the depth map to obtain a lattice tower picture; detecting the lattice tower picture to identify and label missing positions. The application simulates binocular vision by monocular vision, divides image matching areas according to the size of the camera field of view to reduce the amount of calculation, improves the identification ability of key areas by self-supervised learning pre-training and attention mechanism, introduces transfer learning and incremental learning to increase the ability of the model to adapt to new data, and can realize reliable evaluation and real-time analysis of the state of the members in a complex environment, thereby providing a more adaptive and economical solution for the long-term health operation of the communication tower.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of communication structure health monitoring technology, and in particular relates to a method and program product for detecting missing members of lattice towers. Background Technology

[0002] Communication towers, as the supporting structure for base stations, play a crucial role in achieving widespread signal coverage. To ensure the safety and stability of communication towers, it is essential to implement effective monitoring measures for the risk of missing components in lattice-type communication towers, facilitating subsequent structural inspections and reinforcement maintenance. However, lattice-type towers have numerous components, making online health monitoring of each individual component impossible. Therefore, developing an efficient, economical, and reliable risk monitoring system is crucial for ensuring the long-term stability of communication towers.

[0003] With the rapid development of computer vision and deep learning technologies, these technologies have shown great potential in monitoring missing structural members. By constructing deep learning-based computer vision algorithms, the actual operating status of tower structures can be captured in real time, accurately identifying missing or damaged members. Specifically, the system uses high-resolution cameras to acquire tower images, and then extracts visual features through advanced models such as convolutional neural networks (CNNs) or vision transformers (ViTs) to achieve automated detection of member defects. The detection results can be further linked with structural vulnerability analysis models to assess the overall stability of the tower in real time and predict potential risks, thus providing a scientific basis for maintenance decisions.

[0004] However, while existing methods for detecting missing components on communication towers can achieve a certain degree of automation using deep learning, image processing, and sensor technologies, they typically face challenges such as high equipment costs, complex data acquisition, and poor adaptability. These methods rely on high-resolution cameras, real-time monitoring systems, or large amounts of labeled data, the costs and technical requirements of which are difficult to implement on a large scale for the numerous and widely distributed communication towers. Furthermore, environmental factors (such as strong winds, rain, fog, and snow) significantly affect image acquisition quality, thus reducing detection accuracy. Simultaneously, communication towers are often located in different areas, including urban areas, rural areas, and complex terrains, making it difficult to balance inspection efficiency and coverage, further limiting the practicality of existing detection technologies.

[0005] Current methods for missing component detection using computer vision combined with neural networks have certain limitations. Traditional monocular vision methods suffer from misidentification of hollowed-out components, leading to mismatches and misidentifications. Traditional binocular vision methods are prone to inaccurate depth map calculations in dynamic and complex environments, with reduced depth resolution when camera spacing is small and difficulties in disparity matching when spacing is large. Traditional neural networks require extensive data annotation, lack sufficient ability to identify key regions, and have poor adaptability to different models. Therefore, a method that can accurately calculate image depth information and perform self-supervised, precise detection of missing components adapted to new data is of great significance for the health monitoring of lattice towers. Summary of the Invention

[0006] To address the problems mentioned in the background section, this invention provides a method and program product for detecting missing members in lattice towers based on computer vision and deep learning.

[0007] The present invention provides a method for detecting missing members in a lattice tower, the method comprising: Acquire and process image data captured by a monocular camera to obtain a depth map; The depth map is segmented and fused to obtain a lattice tower image; The image of the lattice tower is inspected to identify and mark the missing parts.

[0008] Furthermore, the acquisition and processing of image data captured by a monocular camera to obtain a depth map specifically includes: Acquire image data captured by a monocular camera, the image data including the current photo and the photo to be matched; Based on the image data, the photo to be matched is matched with the current photo to determine the left view and the right view; Based on the brightness information of the left and right views, a stereo camera is simulated to generate a stereo view; A disparity map is generated based on the binocular view, and the disparity map is optimized to obtain a depth map.

[0009] Further, the step of matching the photo to be matched with the current photo based on the image data to determine the left view and the right view specifically includes: The coverage area of ​​the current photo is calculated using the following formula:

[0010] in, For the drone's lateral field of view, This refers to the longitudinal field of view of the drone. For horizontal coverage width, For vertical coverage broadband, This is the current flight altitude of the drone; The formula for calculating the overlap area between two adjacent photos is as follows:

[0011] in, For horizontal coverage width, For vertical coverage broadband, The overlap width, The overlap height; Based on the overlapping region, determine the permissible horizontal displacement range for matching image data. and vertical displacement range The calculation formula is:

[0012] in, For horizontal coverage width, For vertical coverage broadband, The overlap width, The overlap height; Calculate the Euclidean distance between the drone and the center point of the current photo. The calculation formula is:

[0013] in, For the range of horizontal displacement, The vertical displacement range, The Euclidean distance between the drone and the center point of the current photo; The radius of the distance between the image data and the current photo shooting distance is determined to be... The photos to be matched are calculated, and the feature points of the photos to be matched are matched with the current photos. The matching results are used as the left and right views.

[0014] Furthermore, the step of simulating a binocular camera to generate a binocular view based on the brightness information of the left and right views specifically includes: The grayscale gradient of the left view is calculated using the following formula:

[0015] in, The grayscale value of the image. , For the image in and Gray-level gradient in direction, This represents the grayscale gradient of an image as its grayscale values ​​change over time. Based on the grayscale value changes of the pixels in the left view, constraint equations are constructed to calculate the movement velocity of each pixel in the horizontal and vertical directions. The constraint equations are as follows:

[0016] in, and For pixels in and velocity in the direction of motion , For the image in and Gray-level gradient in direction, This represents the grayscale gradient of an image as its grayscale values ​​change over time. Depth estimation is performed by calculating the motion velocity of each pixel, using the following formula:

[0017] in, and Image grayscale value and The second derivative in the direction is used to describe the pixel. and Smoothness of velocity changes in direction; It is a balance factor used to adjust the importance of optical flow constraint terms and regularization terms; The depth of each pixel is calculated using the following formula:

[0018] in, The depth value of a pixel; The focal length of the camera; The baseline distance is the distance between the left and right cameras in the virtual binocular view, estimated by the drone's displacement. This refers to the parallax of corresponding points, which is the horizontal pixel displacement of the same object in the left and right views. Based on the depth information, a virtual right view is generated from the left view:

[0019] in, These are the corresponding pixel coordinates in the right view. These are the pixel coordinates in the left view. The depth value of a pixel; The focal length of the camera; The baseline distance is the distance between the left and right cameras in the virtual binocular view, estimated by the drone's displacement.

[0020] Furthermore, the step of generating a disparity map based on the binocular view and optimizing the disparity map to obtain a depth map specifically includes: gradient The calculation formula is:

[0021] in, Represents pixels The gradient at a given point represents the rate of change of brightness in the image. and For image I in x and y The partial derivative in the direction represents the partial derivative in the direction of ... x and y Brightness changes in direction; Weighting function The calculation formula is:

[0022] in, k This is an adjustment factor used to control the effect of the gradient on the weights. Internal value; The equation for a local window of a pixel reflects the local texture complexity; c This is a control factor used to adjust the relationship between weights and texture complexity; Equation of local window of pixel The calculation formula is:

[0023] in, For the first in the window The value of each pixel. The mean, This represents the total number of pixels within the window. Based on the binocular view, the initial disparity is obtained. The calculation formula is:

[0024] in, The focal length of the camera; By incorporating a weighting function into the interpolation function, the initial disparity is corrected; the calculation formula is as follows:

[0025] in, The disparity value is the top-left corner of the image. The parallax value is the value of the top right corner of the image. The disparity value is the value at the bottom left corner of the image. The disparity value is the value in the lower right corner of the image. are interpolation coefficients and , are interpolation coefficients and 'a' is the proportional offset based on pixel position and b is the proportional offset based on the pixel position and , and For gradient correction and , ; and These are adjustment coefficients used to control the effect of the gradient on the interpolation coefficients; An energy function is constructed to optimize the disparity map, resulting in a depth map. The energy function is constructed as follows:

[0026]

[0027]

[0028] in, It is a smoothing term used to control the smoothness of the disparity map; It is a data item that reflects the similarity between disparity values ​​and image data; Used to adjust the relative importance of smoothing items and data items, in The range of values ​​is determined by the level of noise; the larger the noise level, the larger the value. and It is a location and The disparity value, The left image is in position pixel values, The right image is in position pixel values, That is the corresponding disparity value.

[0029] Furthermore, the segmentation and fusion of the depth map to obtain the lattice tower image specifically includes: Feature extraction and segmentation optimization are performed on a local area of ​​the depth map to obtain a local segmentation result; The depth map is analyzed and segmented globally to obtain a global segmentation result; The local segmentation results and the global segmentation results are fused to obtain the lattice tower image.

[0030] Furthermore, the step of performing feature extraction and segmentation optimization on a local portion of the depth map to obtain a local segmentation result specifically includes: For each pixel, a local window is selected and local features are calculated using the following formula:

[0031] in, In pixels A local window centered on the center. N The number of pixels within the window. The grayscale value of the pixels within the window is selected by choosing the maximum value from the red, green, and blue channels; Based on local features, a threshold is set. The calculation formula is:

[0032] in, This is a constant used to adjust the sensitivity of the threshold, and its value ranges from [value range missing]. ; The grayscale value of each pixel in the depth map is compared with its corresponding local threshold to perform segmentation, specifically:

[0033] in, In the segmentation result, 1 represents the foreground and 0 represents the background.

[0034] Furthermore, the global analysis and segmentation optimization of the depth map to obtain a global segmentation result specifically includes: The number of pixels at each gray level in the depth map is counted to obtain the gray-level histogram. H(g) ; Based on the preset gray level and preset gray threshold of the depth map T The depth map is divided into foreground and background, and the grayscale value of the foreground is greater than... T The grayscale value of the background is less than T ; The threshold is defined by the following formula:

[0035] in, and Weights are assigned to the pixel proportions of the foreground and background. ; and The average gray value of the foreground and background. g represents the grayscale value; T The preset grayscale threshold is used; Determine Maximize the global threshold T Based on the global threshold T The depth map is segmented:

[0036] in, In the segmentation results, 1 represents the foreground and 0 represents the background; The local segmentation results and the global segmentation results are fused to obtain a lattice tower image, specifically including: The local segmentation results and the global segmentation results are then fused:

[0037] When both the local and global segmentation results show as 1, the pixel is used as the foreground to obtain the lattice tower image.

[0038] Furthermore, the step of detecting, identifying, and marking missing parts in the lattice tower image specifically includes: Convolutional neural network model is constructed and optimized, wherein the convolutional neural network model includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer; The lattice tower image was used to extract features using a self-supervised learning method and pre-trained using augmented data. The convolutional neural network model is dynamically updated using incremental learning methods. The image of the lattice tower is input into a trained convolutional neural network model for prediction. Based on the extracted features, it is determined whether each pixel belongs to a missing region, and the missing parts are marked.

[0039] The present invention also provides a computer program product, including a computer program or instructions, characterized in that, when the computer program or instructions are executed by a processor, they implement the aforementioned method and program product for detecting missing members of a lattice tower.

[0040] Compared with the prior art, the present invention has the following advantages: This invention constructs an efficient model for detecting missing components on communication towers by combining monocular-to-binocular vision technology and deep learning methods. By generating disparity maps and optimizing depth information estimation methods, it can effectively identify missing components even in various complex environments and resolve mismatches caused by the hollow structure of lattice towers, significantly improving detection accuracy and robustness. Simultaneously, incremental learning technology is employed, introducing a dynamic data cache pool and a distillation loss function into the model, enabling it to continuously update based on real-time acquired image data without requiring complete retraining. The designed dynamic learning mechanism can simultaneously maintain memory of historical tasks and adapt to new data, thus more efficiently meeting the detection needs in dynamic environments. Attached Figure Description

[0041] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0042] Figure 1 This is a flowchart illustrating a method for detecting missing members in a lattice tower according to an embodiment of the present invention. Figure 2 This is a framework diagram of the lattice tower missing member detection technology according to an embodiment of the present invention. Detailed Implementation

[0043] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0044] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this invention are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a list of steps or methods is not necessarily limited to those explicitly listed, but may include other steps or units not explicitly listed or inherent to these processes or methods.

[0045] Missing members in lattice-type communication towers can lead to reduced structural stiffness, decreased load-bearing capacity, and altered dynamic characteristics, potentially causing tower tilting or collapse in severe cases. Due to the large number of tower members, existing online health monitoring systems are costly and difficult to implement on a large scale. To address the problem of efficient member detection, this invention combines computer vision and deep learning technologies to achieve efficient member detection, reducing detection costs while improving detection accuracy.

[0046] In practical applications, communication towers are distributed in various natural environments (such as mountainous areas, cities, and coastal areas). Strong winds, rain, and snow can interfere with image acquisition quality, affecting the detection accuracy of existing methods. Furthermore, monocular vision detection cannot avoid the influence of hollowed-out structural members. To address the issue of detection accuracy in complex environments, this invention improves traditional visual algorithms by employing a monocular-to-binocular vision-based depth estimation technique, combined with optimized disparity map generation and image segmentation methods, thereby enhancing the robustness of detection in complex environments.

[0047] Deep learning models are highly dependent on large-scale labeled data, and existing models struggle to effectively extract key features under limited sample conditions. To address the feature learning problem under limited labeled data conditions, this invention employs self-supervised learning techniques, utilizing data augmentation and contrastive learning to pre-train features on unlabeled data. This significantly reduces the need for labeled data and improves the model's sensitivity to missing components.

[0048] Because communication tower data comes from diverse sources and is updated in real time, traditional detection models require frequent retraining to adapt to new data, which is both time-consuming and computationally expensive. To address the issue of real-time adaptation to data changes, this invention introduces incremental learning techniques, designs a dynamic data cache pool and a distillation loss function, ensuring that the model can dynamically update and adapt to real-time data changes while avoiding forgetting old tasks.

[0049] This invention proposes a method for monitoring the risk of missing components in lattice-type communication towers based on computer vision combined with deep learning. This method improves upon traditional vision methods by using monocular vision to simulate binocular vision, dividing the image matching region according to the camera's field of view to reduce computation, using self-supervised learning pre-training and attention mechanisms to improve the identification ability of key regions, and introducing transfer learning and incremental learning to increase the model's ability to adapt to new data.

[0050] In one embodiment of the present invention, a method for detecting missing members in a lattice tower is provided, such as... Figure 1 As shown, the method includes the following steps: S1. Acquire and process image data captured by a monocular camera to obtain a depth map.

[0051] In this embodiment, image data is acquired using a monocular camera, a virtual binocular view is generated through motion parallax analysis, and the parallax map is optimized by combining an energy function, ultimately outputting depth information.

[0052] In this embodiment, motion parallax analysis refers to inferring the depth information of a scene by calculating the differences in motion speed of pixels in multiple frames of images.

[0053] In this embodiment, step S1 involves acquiring and processing image data captured by a monocular camera to obtain a depth map, including the following steps: S11. Acquire image data captured by a monocular camera, the image data including the current photo and the photo to be matched; In this embodiment, a drone is used to capture multiple frames of images containing the lattice tower within the target area.

[0054] S12. Based on the image data, match the photo to be matched with the current photo to determine the left view and the right view, specifically including: In this embodiment, the lateral and longitudinal coverage widths are calculated based on the drone's lateral field of view (representing the width of the field of view in the horizontal direction), longitudinal field of view (representing the width of the field of view in the vertical direction), and the current flight altitude of the drone. The formula for calculating the coverage area of ​​the current photo is as follows:

[0055] in, For the drone's lateral field of view, This refers to the longitudinal field of view of the drone. For horizontal coverage width, For vertical coverage broadband, This represents the current flight altitude of the drone.

[0056] By calculating the overlapping area between photos, the overlap width and overlap height are obtained. The formula for calculating the overlap area between two adjacent photos is as follows:

[0057] in, For horizontal coverage width, For vertical coverage broadband, The overlap width, The overlap height; The overlapping region is determined by the field of view during drone photography and is used to determine the allowable horizontal and vertical displacement ranges for matching image data. Based on the overlapping region, the horizontal displacement range is calculated. and vertical displacement range The calculation formula is:

[0058] in, For horizontal coverage width, For vertical coverage broadband, The overlap width, The overlap height; Calculate the Euclidean distance between the drone and the center point of the current photo. Euclidean distance It is derived by taking the square root of the sum of the squares of the current flight altitude and the horizontal displacement. The calculation formula is as follows:

[0059] in, For the range of horizontal displacement, The vertical displacement range, The distance between the drone and the center point of the current photo is the Euclidean distance.

[0060] The radius of the distance between the image data and the current photo shooting distance is determined to be... The photos to be matched within a certain range are used to extract feature points from the photos to be matched using the Scale Invariant Feature Theory (SIFT) algorithm. These feature points are then matched with the current photo to find the photo with the highest matching degree to the current image. These two photos together form the left and right views.

[0061] S13. Based on the brightness information of the left and right views, simulate a binocular camera to generate a binocular view, specifically including: In this embodiment, based on the brightness change information of objects in the two frames of images (left and right views), a virtual right view is generated using the left view to simulate the imaging effect of a binocular camera. The left view is the current photograph. The grayscale gradient of the left view, i.e., the rate of change of the image's grayscale values ​​in the horizontal, vertical, and temporal directions, is calculated using the following formula:

[0062] in, The grayscale value of the image. , For the image in and Gray-level gradient in direction, This represents the grayscale gradient of an image as its grayscale values ​​change over time. Based on the grayscale value changes of the pixels in the left view, constraint equations are constructed to calculate the movement velocity of each pixel in the horizontal and vertical directions. The constraint equations are as follows:

[0063] in, and For pixels in and velocity in the direction of motion , For the image in and Gray-level gradient in direction, This represents the grayscale gradient of an image as its grayscale values ​​change over time. Depth estimation is performed by calculating the motion velocity of each pixel, using the following formula:

[0064] in, and Image grayscale value and The second derivative in the direction is used to describe the pixel. and Smoothness of velocity changes in direction; It is a balance factor used to adjust the importance of optical flow constraint terms and regularization terms. The larger the value, the more the model tends to smooth the optical flow field, and the smaller the value, the more it tends to strictly satisfy the optical flow constraint equation. For lattice towers, which have large dynamic changes and rich texture information, the default value is 5.

[0065] The depth of each pixel is calculated using the following formula:

[0066] in, The depth value of a pixel; The focal length of the camera; The baseline distance is the distance between the left and right cameras in the virtual binocular view, estimated by the drone's displacement. This refers to the parallax of corresponding points, which is the horizontal pixel displacement of the same object in the left and right views. Based on the depth information, a virtual right view is generated from the left view:

[0067] in, These are the corresponding pixel coordinates in the right view. These are the pixel coordinates in the left view. The depth value of a pixel; The focal length of the camera; The baseline distance is the distance between the left and right cameras in the virtual binocular view, estimated by the drone's displacement.

[0068] S14. Generate a disparity map based on the binocular view, and optimize the disparity map to obtain a depth map, specifically including: In this embodiment, the initial disparity is determined by the gradient change rate of the image and the texture complexity of the local region where the pixel is located. The texture complexity reflects the gray-scale changes in the local region and is determined by calculating the deviation between the mean pixel value within the local window and the local features.

[0069] gradient The calculation formula is:

[0070] in, Represents pixels The gradient at a given point represents the rate of change of brightness in the image. and For image I exist x and y The partial derivative in the direction represents the partial derivative in the direction of ... x and y Brightness changes in direction; Weighting function The calculation formula is:

[0071] in, k This is an adjustment factor used to control the effect of the gradient on the weights. The value can be set internally, with a default value of 3. The equation for a local window of a pixel reflects the local texture complexity; c This is a control factor, set to 1 by default, used to adjust the relationship between weights and texture complexity; Equation of local window of pixel The calculation formula is:

[0072] in, For the first in the window The value of each pixel. The mean, This represents the total number of pixels within the window. Based on the binocular view, the initial disparity is obtained. The calculation formula is:

[0073] in, The focal length of the camera; Represents pixels The gradient at a given point represents the rate of change in the image's brightness.

[0074] By incorporating a weighting function into the interpolation function, the initial disparity is corrected; the calculation formula is as follows:

[0075] in, The disparity value is the top-left corner of the image. The parallax value is the value of the top right corner of the image. The disparity value is the value at the bottom left corner of the image. The disparity value is the value in the lower right corner of the image. are interpolation coefficients and , are interpolation coefficients and 'a' is the proportional offset based on pixel position and , b For pixel position-based proportional offset and , and For gradient correction and , ; and The adjustment coefficients are used to control the effect of the gradient on the interpolation coefficients, because the image contrast is high. Take 3. Taking 0.5 will make the interpolation closer to the true value in the high-frequency region, while retaining a high degree of continuity in the smooth region.

[0076] An energy function is constructed to optimize the disparity map, resulting in a depth map. The energy function is constructed as follows:

[0077]

[0078]

[0079] in, It is a smoothing term used to control the smoothness of the disparity map; It is a data item that reflects the similarity between disparity values ​​and image data; Used to adjust the relative importance of smoothing items and data items, in The range of values ​​is determined by the level of noise; the larger the noise level, the larger the value. The default value is 5. and It is a location and The disparity value, The left image is in position pixel values, The right image is in position pixel values, That is the corresponding disparity value.

[0080] In this embodiment, during the optimization process, these weights are combined with the interpolation function to ensure that the disparity values ​​are closer to the true values ​​in the high-frequency region and maintain continuity in the low-frequency smooth region. Simultaneously, an energy function is introduced, which includes a data term and a smoothness term. The data term reflects the similarity between the disparity values ​​and the actual image data, while the smoothness term controls the smoothness of the disparity map. In noisy scenes, the smoothness term has a higher weight to suppress noise interference with the disparity map. Finally, an optimized depth map is generated for subsequent analysis.

[0081] S2. The depth map is segmented and fused to obtain a lattice tower image.

[0082] In this embodiment, the depth map is segmented and fused based on a threshold segmentation algorithm to obtain a lattice tower image. The threshold segmentation algorithm refers to dividing the pixels of the image into different categories (such as foreground and background) by setting one or more grayscale thresholds, thereby achieving separation between the target and the background.

[0083] In this embodiment, step S2 involves segmenting and fusing the depth map to obtain a lattice tower image, including the following steps: S21. Perform feature extraction and segmentation optimization on the local area of ​​the depth map to obtain local segmentation results, specifically including: For each pixel, a fixed-size local window is selected for analysis, and the grayscale value of the pixels within that window is calculated as a local feature. The calculation formula is as follows:

[0084] in, In pixels A local window centered on the center. N The number of pixels within the window. The grayscale value of the pixels within the window is selected by choosing the maximum value from the red, green, and blue channels; Based on local features, a segmentation threshold is set. The threshold sensitivity is controlled by a constant whose value is adjusted according to the noise level; the higher the noise, the higher the threshold. The calculation formula is as follows:

[0085] in, This is a constant used to adjust the sensitivity of the threshold, and its value ranges from [value range missing]. The greater the noise, the larger the value.

[0086] The grayscale value of each pixel in the depth map is compared with its corresponding local threshold. If the grayscale value of a pixel is greater than the corresponding local threshold, the image is marked as foreground; otherwise, it is marked as background, thus completing local segmentation. Specifically:

[0087] in, In the segmentation result, 1 represents the foreground and 0 represents the background.

[0088] S22. Perform global analysis and segmentation optimization on the depth map to obtain a global segmentation result, specifically including: In this embodiment, after the initial image is divided into several sub-images, the optimal segmentation threshold for the entire image is determined based on the grayscale information of the entire image.

[0089] The number of pixels at each gray level in the depth map is counted to obtain the gray-level histogram. H(g) , where g is the grayscale value.

[0090] Based on the preset gray level and preset gray threshold of the depth map T The depth map is divided into foreground and background, and the grayscale value of the foreground is greater than... T The grayscale value of the background is less than T ; The optimal threshold is defined by the following formula:

[0091] in, and Weights are assigned to the pixel proportions of the foreground and background. ; and The average gray value of the foreground and background. g is the grayscale value; T is the preset grayscale threshold.

[0092] By iterating through all possible thresholds T , determine Maximize the global threshold T Based on the global threshold T The depth map is segmented:

[0093] in, In the segmentation result, 1 represents the foreground and 0 represents the background.

[0094] S23. The local segmentation results and the global segmentation results are fused to obtain a lattice tower image, specifically including: The local segmentation results and the global segmentation results are then fused:

[0095] When both the local and global segmentation results show as 1, the pixel is used as the foreground to obtain the lattice tower image.

[0096] S3. Detect the lattice tower image, identify and mark the missing parts.

[0097] In this embodiment, a deep learning algorithm is used to identify missing components in the lattice tower image. This invention combines monocular-to-binocular vision technology with deep learning methods to construct an efficient model suitable for detecting missing components on communication towers. By generating a disparity map and optimizing the depth information estimation method, it can effectively identify missing components even in various complex environments (such as strong winds, rain, snow, and obstructions) and resolve the mismatch effects caused by the hollow structure of the lattice tower. The improved model fully considers the impact of tower structural features and environmental changes on the detection results, significantly improving detection accuracy and robustness.

[0098] In this embodiment, step S3 involves detecting the lattice tower image, identifying and marking missing parts, including the following steps: S31. Construct and optimize a convolutional neural network model, wherein the convolutional neural network model includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer, specifically: In this embodiment, a convolutional neural network model is constructed, including an input layer, a convolutional layer, a pooling layer, and a fully connected layer.

[0099] The input layer receives a size of 100,000. (pixel size is) An RGB image (with 3 color channels, namely red, green, and blue) needs to be input and normalized to a certain standard. scope.

[0100] Convolutional layers extract deep features from images. The first convolutional layer consists of 32 convolutional kernels, and the next two layers consist of 64 and 128 convolutional kernels, respectively, to extract richer feature information.

[0101] Convolutional layer 1:

[0102] in, Here is the feature map of the first layer, and here is the input image; The bias term is initialized to 0; For activation functions, retain non-negative values; It is the size of The convolutional kernel has 32 elements and is initialized as follows:

[0103] in, This represents the number of input channels for the previous layer.

[0104] The second convolutional layer has 64 kernels, each 3×3 in size, and the output feature size is [missing value]. .

[0105] The third convolutional layer has 128 kernels, each 3×3 in size, and the output feature size is... .

[0106] Pooling layers: A max pooling layer is added after every two convolutional layers. Max pooling is used to reduce the size of the feature maps. The pooling window size is [size missing]. Step size 2.

[0107] Fully connected layer: Flattened feature map Input a fully connected layer with 128 neurons; use the Softmax activation function to generate a three-class classification output, in the following form.

[0108] in, These are the weights of the fully connected layer, with a default value of 1. This is the bias of the fully connected layer, with a default value of 0. This is a feature map of a certain layer.

[0109] Optimization is performed using the cross-entropy loss function:

[0110] in, It is a real label, using one-hot encoding; Predict probabilities for the model.

[0111] Training parameter probabilities are set as follows: Optimizer: Adam, Initial learning rate: Momentum parameters: Batch size: 32, Maximum number of training rounds: 50.

[0112] S32. Features are extracted from the lattice tower image using a self-supervised learning method, and pre-trained using augmented data, specifically as follows: In this embodiment, the input image is subjected to two random transformations through data augmentation to generate positive sample pairs. These transformations include random cropping, horizontal flipping, Gaussian blurring, and grayscale conversion. The similarity between feature representations is calculated using the augmented samples, and optimized through contrastive learning loss, thereby improving the model's feature extraction capability.

[0113] By analyzing the input image Perform two random boosts to generate positive sample pairs. and The specific steps for each enhancement are as follows: Random cropping: Retain 8% of image content, pixel size scaled to [value missing]. ; Horizontal Flip: Flips the cropped image with a 50% probability; Gaussian blur: using a 50% probability The blur kernel blurs the image; Random grayscale conversion: Converts the image to grayscale with a 20% probability.

[0114] The enhanced image is input into the CNN to obtain:

[0115] Define contrastive learning loss:

[0116] in, This is the temperature parameter, with a default value of 0.12. Let be the cosine similarity.

[0117] Training settings: Optimizer: Adam, Initial learning rate: Batch size: 256; Maximum number of training rounds: 2000.

[0118] S33. Utilize incremental learning methods to dynamically update the convolutional neural network model, specifically as follows: In this embodiment, traditional detection methods struggle to adapt to the real-time changes in newly added data from communication towers, often requiring the entire model to be retrained. This invention employs incremental learning technology, introducing a dynamic data cache pool and a distillation loss function into the model. This allows the model to continuously update based on real-time acquired image data without requiring complete retraining. Simultaneously, the designed dynamic learning mechanism maintains both a memory of historical tasks and adaptability to new data, thus more efficiently meeting the detection needs in dynamic environments.

[0119] The dynamic data cache pool selects the most important samples from the old data and adds them to the cache pool based on the sample gradient contribution calculation formula:

[0120] in, For the gradient of sample loss, retain the one that contributes the most. One sample, with a cache pool size of 1000.

[0121] When training on new data, distillation loss is used to maintain consistency between the old and new models, thereby avoiding forgetting historical tasks.

[0122] in, Indicates the features of the old model, This represents the features of the new model.

[0123] Merge newly added data and cache pools, perform incremental training, establish the total loss function, and optimize model parameters to complete the incremental update:

[0124] Output the detection image results, mark the missing rod parts with a red box, and label them as "lack".

[0125] The image of the lattice tower is input into a trained convolutional neural network model for prediction, and the output detection image result is used to determine whether each pixel belongs to the missing region based on the extracted features. The missing parts are marked with a red box and labeled as "lack".

[0126] Figure 2 This is a framework diagram of the lattice tower member missing member detection technology according to an embodiment of the present invention, as shown below. Figure 2 As shown, this invention provides a method for detecting missing lattice tower members based on computer vision and deep learning. First, based on computer vision and deep learning technologies, the detection area is determined through image matching range calculation. Multi-view data is generated using binocular vision simulation, and disparity maps are generated and optimized. Then, local detail extraction and segmentation, and global information analysis and segmentation are combined to fuse the local and global segmentation results. Next, a deep learning network architecture is built, pre-trained using self-supervised learning, and the model is dynamically updated using incremental learning. Finally, through binocular vision simulation and target recognition, accurate detection using deep learning is achieved, identifying missing lattice tower members and providing visual annotation.

[0127] This invention proposes a method for detecting missing structural members based on binocular vision simulation, deep learning, and incremental learning. This method eliminates the need for expensive monitoring equipment; it utilizes only a conventional UAV image acquisition system to obtain multi-angle image data of the tower, enabling efficient detection of missing members through algorithmic processing. By combining contrastive learning and self-supervised learning techniques with the UAV-acquired image data, feature extraction can be performed even in the absence of large-scale labeled data. Furthermore, the designed incremental learning mechanism dynamically updates the model, adapting to the ever-changing inspection data requirements and providing technical support for subsequent refined inspection and maintenance plans.

[0128] Furthermore, traditional pole detection algorithms typically rely on fixed image processing models, failing to adequately consider the dynamic changes in communication tower performance over time and due to environmental influences. This invention, through a dynamically optimized model architecture, enables reliable assessment and real-time analysis of pole conditions in complex environments, providing a more adaptable and economical solution for the long-term healthy operation of communication towers.

[0129] The present invention also provides a computer program product, including a computer program or instructions, which, when executed by a processor, implement the aforementioned method for detecting missing lattice tower members based on computer vision and deep learning.

[0130] Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. A method for detecting missing members in a lattice tower, characterized in that, The method includes: Acquire and process image data captured by a monocular camera to obtain a depth map; The depth map is segmented and fused to obtain a lattice tower image; The image of the lattice tower is inspected to identify and mark the missing parts; The acquisition and processing of image data captured by a monocular camera to obtain a depth map specifically includes: Acquire image data captured by a monocular camera, the image data including the current photo and the photo to be matched; Based on the image data, the photo to be matched is matched with the current photo to determine the left view and the right view; Based on the brightness information of the left and right views, a stereo camera is simulated to generate a stereo view; A disparity map is generated based on the binocular view, and a depth map is obtained by optimizing the disparity map. The step of matching the photo to be matched with the current photo based on the image data to determine the left view and the right view specifically includes: The coverage area of ​​the current photo is calculated using the following formula: in, For the drone's lateral field of view, This refers to the longitudinal field of view of the drone. For horizontal coverage width, The vertical coverage width, This is the current flight altitude of the drone; The formula for calculating the overlap area between two adjacent photos is as follows: in, For horizontal coverage width, The vertical coverage width, The overlap width, The overlap height; Based on the overlapping region, determine the permissible horizontal displacement range for matching image data. and vertical displacement range The calculation formula is: in, For horizontal coverage width, The vertical coverage width, The overlap width, The overlap height; Calculate the Euclidean distance between the drone and the center point of the current photo. The calculation formula is: in, For the range of horizontal displacement, The vertical displacement range, The Euclidean distance between the drone and the center point of the current photo; The radius of the distance between the image data and the current photo shooting distance is determined to be... The photos to be matched are calculated and matched with the current photo by calculating the feature points of the photos to be matched, and the matching results are used as the left and right views; The process of segmenting and fusing the depth map to obtain the lattice tower image specifically includes: Feature extraction and segmentation optimization are performed on a local area of ​​the depth map to obtain a local segmentation result; The depth map is analyzed and segmented globally to obtain a global segmentation result; The local segmentation results and the global segmentation results are fused to obtain a lattice tower image; The process of detecting, identifying, and marking missing parts in the lattice tower image specifically includes: Convolutional neural network model is constructed and optimized, wherein the convolutional neural network model includes an input layer, a convolutional layer, a pooling layer, and a fully connected layer; The lattice tower image was used to extract features using a self-supervised learning method and pre-trained using augmented data. The convolutional neural network model is dynamically updated using incremental learning methods. The image of the lattice tower is input into a trained convolutional neural network model for prediction. Based on the extracted features, it is determined whether each pixel belongs to a missing region, and the missing parts are marked.

2. The method according to claim 1, characterized in that, The step of simulating a binocular camera to generate a binocular view based on the brightness information of the left and right views specifically includes: The grayscale gradient of the left view is calculated using the following formula: in, The grayscale value of the image. , For the image in and Gray-level gradient in direction, This represents the grayscale gradient of an image as its grayscale values ​​change over time. Based on the grayscale value changes of the pixels in the left view, constraint equations are constructed to calculate the movement velocity of each pixel in the horizontal and vertical directions. The constraint equations are as follows: in, and For pixels in and velocity in the direction of motion , For the image in and Gray-level gradient in direction, This represents the grayscale gradient of an image as its grayscale values ​​change over time. Depth estimation is performed by calculating the motion velocity of each pixel, using the following formula: in, and Image grayscale value and The second derivative in the direction is used to describe the pixel. and Smoothness of velocity changes in direction; It is a balance factor used to adjust the importance of optical flow constraint terms and regularization terms; The depth of each pixel is calculated using the following formula: in, The depth value of a pixel; The focal length of the camera; The baseline distance is the distance between the left and right cameras in the virtual binocular view, estimated by the drone's displacement. This refers to the parallax of corresponding points, which is the horizontal pixel displacement of the same object in the left and right views. Based on the depth information, a virtual right view is generated from the left view: in, These are the corresponding pixel coordinates in the right view. These are the pixel coordinates in the left view. The depth value of a pixel; The focal length of the camera; The baseline distance is the distance between the left and right cameras in the virtual binocular view, estimated by the drone's displacement.

3. The method according to claim 1, characterized in that, The step of generating a disparity map based on the binocular view and optimizing the disparity map to obtain a depth map specifically includes: gradient The calculation formula is: in, Represents pixels The gradient at a given point represents the rate of change of brightness in the image. and For image I in x and y The partial derivative in the direction represents the partial derivative in the direction of ... x and y Brightness changes in direction; Weighting function The calculation formula is: in, k This is an adjustment factor used to control the effect of the gradient on the weights. Internal value; The equation for a local window of a pixel reflects the local texture complexity; c This is a control factor used to adjust the relationship between weights and texture complexity; Equation of local window of pixel The calculation formula is: in, For the first in the window The value of each pixel. The mean, This represents the total number of pixels within the window. Based on the binocular view, the initial disparity is obtained. The calculation formula is: in, The focal length of the camera; By incorporating a weighting function into the interpolation function, the initial disparity is corrected; the calculation formula is as follows: in, The disparity value is the top-left corner of the image. The parallax value is the value of the top right corner of the image. The disparity value is the value at the bottom left corner of the image. The disparity value is the value in the lower right corner of the image. are interpolation coefficients and , are interpolation coefficients and 'a' is the proportional offset based on pixel position and b is the proportional offset based on the pixel position and , and For gradient correction and , ; and These are adjustment coefficients used to control the effect of the gradient on the interpolation coefficients; An energy function is constructed to optimize the disparity map, resulting in a depth map. The energy function is constructed as follows: in, It is a smoothing term used to control the smoothness of the disparity map; It is a data item that reflects the similarity between disparity values ​​and image data; Used to adjust the relative importance of smoothing items and data items, in The range of values ​​is determined by the level of noise; the larger the noise level, the larger the value. and It is a location and The disparity value, The left image is in position pixel values, The right image is in position pixel values, That is the corresponding disparity value.

4. The method according to claim 1, characterized in that, The step of performing feature extraction and segmentation optimization on a local area of ​​the depth map to obtain a local segmentation result specifically includes: For each pixel, a local window is selected and local features are calculated using the following formula: in, In pixels A local window centered on the center. N The number of pixels within the window. The grayscale value of the pixels within the window is selected by choosing the maximum value from the red, green, and blue channels; Based on local features, a threshold is set. The calculation formula is: in, This is a constant used to adjust the sensitivity of the threshold, and its value ranges from [value range missing]. ; The grayscale value of each pixel in the depth map is compared with its corresponding local threshold to perform segmentation, specifically: in, In the segmentation result, 1 represents the foreground and 0 represents the background.

5. The method according to claim 1, characterized in that, The global analysis and segmentation optimization of the depth map to obtain the global segmentation result specifically includes: The number of pixels at each gray level in the depth map is counted to obtain the gray-level histogram. H(g) ; Based on the preset gray level and preset gray threshold of the depth map T The depth map is divided into foreground and background, and the grayscale value of the foreground is greater than... T The grayscale value of the background is less than T ; The threshold is defined by the following formula: in, and Weights are assigned based on the pixel proportions of the foreground and background. ; and The average gray value of the foreground and background. g represents the grayscale value; T The preset grayscale threshold is used; Determine Maximize the global threshold T Based on the global threshold T The depth map is segmented: in, In the segmentation results, 1 represents the foreground and 0 represents the background; The local segmentation results and the global segmentation results are fused to obtain a lattice tower image, specifically including: The local segmentation results and the global segmentation results are then fused: When both the local and global segmentation results show as 1, the pixel is used as the foreground to obtain the lattice tower image.

6. A computer program product, comprising a computer program or instructions, characterized in that, When the computer program or instructions are executed by the processor, they implement the method for detecting missing lattice tower members as described in any one of claims 1-5.