Communication tower with image recognition device

By employing optical flow calculation, key feature extraction, and matching modules, combined with deep learning models and loss function optimization, the problems of high sensor monitoring costs and low image monitoring efficiency have been solved. This enables real-time and accurate monitoring of communication towers and timely detection of abnormal vibrations, ensuring the stability and safety of the towers.

CN120071250BActive Publication Date: 2025-10-28HEBEI CENTURY METAL STRUCTURE
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145825.0
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-10-28
Estimated Expiration
2045-02-10

AI Technical Summary

Technical Problem

Existing methods for monitoring communication towers rely on sensors, which have high installation and maintenance costs and can only acquire information from limited locations, making it difficult to comprehensively reflect the overall operating status of the tower. Image monitoring methods have low efficiency and accuracy in calculating image motion information, extracting features, and matching features, and cannot meet the needs of real-time and accurate monitoring.

Method used

An optical flow calculation module is used to calculate the optical flow field of the target tower image sequence. A key feature extraction module extracts key feature points from the first frame image. A feature matching module performs matching based on the optical flow field. A displacement judgment module calculates the vibration displacement of the tower. The efficiency of feature extraction and matching is improved by using a target model and a deep learning model. The optical flow estimation is optimized by combining a multi-scale average absolute error loss function.

Benefits of technology

It enables real-time and precise monitoring of communication towers, timely detection of abnormal vibrations, ensuring the stable and safe operation of towers, and improving the efficiency and accuracy of image recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120071250B_ABST
    Figure CN120071250B_ABST
Patent Text Reader

Abstract

The present disclosure provides a communication tower with an image recognition device, belonging to the field of image recognition technology. The communication tower with an image recognition device includes: an optical flow calculation module for inputting a target tower image sequence into a target model to obtain an optical flow field between adjacent target tower images; a key feature extraction module for extracting features from a first target tower image to obtain key feature points; a feature matching module for matching the key feature points with a second target tower image based on the optical flow field to obtain the feature coordinates of the key feature points on the second target tower image; and a displacement judgment module for obtaining the vibration displacement of the tower based on the coordinates of the key feature points between adjacent frames of target tower images. The present disclosure can meet the demand for real-time and accurate monitoring of communication towers.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of image recognition technology, and more particularly to a communication tower with an image recognition device. Background Technology

[0002] As a crucial infrastructure of communication networks, the stability of communication towers directly impacts communication quality and security. Traditional tower monitoring methods primarily rely on installing various sensors, such as accelerometers and displacement sensors. However, these methods have several limitations. On one hand, the installation and maintenance costs of these sensors are high, and they are susceptible to malfunctions due to harsh environments. On the other hand, sensors can only acquire information from limited locations, making it difficult to comprehensively reflect the overall operational status of the tower.

[0003] With the development of image recognition technology, using images to monitor the status of communication towers has become a new approach. However, existing image monitoring methods are inefficient and inaccurate in terms of calculating image motion information, extracting features, and matching features, and cannot meet the needs of real-time and precise monitoring of communication towers. Summary of the Invention

[0004] This disclosure provides a communication tower with an image recognition device to meet the need for real-time and accurate monitoring of communication towers.

[0005] This disclosure provides a communication tower with an image recognition device, comprising:

[0006] The optical flow calculation module is used to input the target tower image sequence into the target model to obtain the optical flow field between adjacent target tower images, wherein the adjacent frame target tower images are adjacent target tower images in the target tower image sequence;

[0007] The key feature extraction module is used to extract features from the first target tower image to obtain key feature points. The first target tower image is the first frame of the target tower image in the target tower image sequence.

[0008] The feature matching module is used to match the key feature points with the second target tower image based on the optical flow field to obtain the feature coordinates of the key feature points on the second target tower image. The second target tower image is any frame of the target tower image sequence other than the first frame of the target tower image.

[0009] The displacement determination module is used to obtain the vibration displacement of the tower based on the coordinates of the key feature points between adjacent frame images of the target tower.

[0010] In one exemplary embodiment of this disclosure, it further includes: a target model construction module;

[0011] The target model construction module is specifically used to: extract features from any two frames of target tower images to obtain a first feature map and a second feature map;

[0012] An initial optical flow estimate is obtained based on the similarity between the first feature map and the second feature map;

[0013] The initial optical flow estimate is evaluated based on the target loss function;

[0014] The target model is determined based on the evaluation results.

[0015] In one exemplary embodiment of this disclosure, it further includes: a similarity calculation module;

[0016] The similarity calculation module is specifically used to: calculate the similarity between the first feature map and the second feature map based on the first formula;

[0017] The first formula is:

[0018]

[0019] in, This represents the similarity between the first feature map and the second feature map. This represents the eigenvector of the first feature map. This represents the feature vector of the second feature map. This represents the Euclidean distance between the first feature map and the second feature map. This represents the weighting coefficients corresponding to the Euclidean distance. The cosine similarity between the first and second feature maps is represented. This represents the weight coefficient corresponding to the cosine similarity. Represents the feature distribution weights.

[0020] In one exemplary embodiment of this disclosure, the key feature extraction module is specifically used for:

[0021] Spatially align the first target tower image and the first infrared tower image to obtain a preprocessed image;

[0022] An initial feature map is obtained based on the association weights between each pixel in the preprocessed image and its surrounding pixels;

[0023] The key feature points are determined based on the importance of each target feature in the initial feature map.

[0024] In one exemplary embodiment of this disclosure, it further includes: a feature enhancement module;

[0025] The feature enhancement module is used for:

[0026] In response to the fact that the importance of the target feature in the initial feature map is greater than a first preset value, the resolution of the target feature is increased;

[0027] In response to the fact that the importance of the target feature in the initial feature map is less than or equal to a first preset value, the resolution of the target feature is reduced.

[0028] In one exemplary embodiment of this disclosure, the feature matching module is specifically used for:

[0029] Based on the coordinates of each key feature point in the first target tower image, the motion vector corresponding to each key feature point is obtained from the optical flow field;

[0030] Based on the motion vector, the initial coordinates of key feature points in the second target tower image are determined;

[0031] The search window is determined based on the vibration amplitude of the communication tower, with the initial coordinates as the center.

[0032] In response, within the search window range, descriptors for all key feature points of the first target tower image and the second target tower image are calculated respectively within the corresponding search window range;

[0033] The descriptors of key feature points in the first target tower image are matched with the descriptors of key feature points in the second target tower image to obtain the feature coordinates of the key feature points on the second target tower image.

[0034] In one exemplary embodiment of this disclosure, the feature matching module is further configured to:

[0035] In response to the search window range, the feature point closest to the descriptor of each key feature point in the first target tower image to each key feature point in the corresponding second target tower image is taken as the matching point.

[0036] In one exemplary embodiment of this disclosure, the displacement determination module is specifically used for:

[0037] Calculate the displacement of the key feature points between adjacent frames of the target tower image;

[0038] The vibration displacement of the tower is obtained based on the average displacement of all the key feature points.

[0039] In one exemplary embodiment of this disclosure, the displacement determination module is further configured to:

[0040] The average displacement of all the key feature points is calculated based on the second formula;

[0041] The second calculation formula is:

[0042]

[0043] in, This represents the average displacement of all key feature points. This represents the importance weight of the i-th key feature point. This represents the magnitude of the displacement vector of the i-th key feature point. This represents the cosine of the angle between the displacement vector and the reference vector at each key feature point. This represents the total number of key feature points.

[0044] In one exemplary embodiment of this disclosure, it further includes:

[0045] The vibration analysis module is used to analyze the tower based on its vibration displacement to obtain the tower's operating status.

[0046] The beneficial effects of the communication tower with image recognition device provided in this embodiment are as follows: The optical flow calculation module in this embodiment, with the help of the target model, can calculate the optical flow field between adjacent target tower images, clearly presenting the tower's motion information between adjacent frames, providing a basis for subsequent analysis. The key feature extraction module extracts key feature points from the first frame image and establishes an initial feature reference. These representative and unique points are easy to identify and track in different frames. The feature matching module utilizes the motion information of the optical flow field to narrow the matching search range, greatly improving the efficiency and accuracy of key feature point matching. The displacement judgment module obtains the displacement of each point by comparing the coordinates of key feature points in adjacent frames, and then obtains the tower vibration displacement through comprehensive analysis. These vibration displacement data can effectively monitor the tower's operating status, promptly detect abnormal vibrations, and thus ensure the stable and safe operation of the communication tower. Attached Figure Description

[0047] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0048] Figure 1 This is a schematic diagram of the structure of a communication tower with an image recognition device according to an embodiment of this disclosure;

[0049] Figure 2 This is a schematic diagram of the structure of a communication tower with an image recognition device provided in another embodiment of this disclosure. Detailed Implementation

[0050] To enable those skilled in the art to better understand this solution, the technical solutions in the embodiments of this solution will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this solution, not all of them. Based on the embodiments of this solution, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this solution.

[0051] The term "comprising" and any other variations thereof in the specification, claims, and accompanying drawings of this invention mean "including but not limited to," and are intended to cover a non-exclusive inclusion, not limited to the examples listed herein. Furthermore, the terms "first" and "second," etc., are used to distinguish different objects, not to describe a specific order.

[0052] The implementation of this disclosure will be described in detail below with reference to the specific accompanying drawings:

[0053] Figure 1 This is a schematic diagram of the structure of a communication tower with an image recognition device, provided as an embodiment of this disclosure. (Refer to...) Figure 1 The communication tower with an image recognition device includes:

[0054] The optical flow calculation module is used to input the target tower image sequence into the target model to obtain the optical flow field between adjacent target tower images. The adjacent frame target tower images are the adjacent target tower images in the target tower image sequence.

[0055] The key feature extraction module is used to extract features from the first target tower image to obtain key feature points. The first target tower image is the first frame of the target tower image in the target tower image sequence.

[0056] The feature matching module is used to match key feature points with the second target tower image based on the optical flow field, and obtain the feature coordinates of the key feature points on the second target tower image. The second target tower image is any frame of the target tower image sequence other than the first frame of the target tower image.

[0057] The displacement determination module is used to obtain the vibration displacement of the tower based on the coordinates of key feature points between adjacent frame images of the target tower.

[0058] In this embodiment, an image recognition device can be applied to communication towers. The optical flow calculation module takes the target tower image sequence as input and calculates the optical flow field between adjacent target tower images. The optical flow field is the motion information of the communication tower in the image between adjacent frames, and the optical flow vector of each point represents the displacement direction and magnitude of that point between two image frames.

[0059] The target tower image sequence is input into the target model. The target model can be a trained optical flow estimation algorithm model, such as a deep learning model like RAFT or PWC-Net. The target model can analyze the brightness changes and correlations of pixels in adjacent frames. Based on assumptions such as constant brightness (i.e., the brightness of the same object remains unchanged in adjacent frames) and temporal continuity (the object moves slowly), it can calculate the motion vector of each pixel. The set of these motion vectors constitutes the optical flow field.

[0060] In this embodiment, the key feature extraction module uses the first frame of the target tower image sequence as input to extract key feature points and establish an initial feature reference. Key feature points are representative and unique points in the image, such as corner points and edge points of communication towers. Key feature points are relatively easy to identify and track in different image frames.

[0061] Algorithms such as scale-invariant feature transformation, accelerated robust feature transformation, or Harris corner detection can be used to find representative and stable key feature points in the first target tower image.

[0062] In this embodiment, after obtaining the optical flow field between adjacent target tower images and the key feature points of the first frame image, the feature matching module matches the key feature points in the first frame image with the second target tower image based on the motion information provided by the optical flow field. The motion vectors in the optical flow field can help predict the approximate location of the key feature points in the second target tower image, thereby reducing the search range for matching and improving matching efficiency and accuracy. Through the matching process, the corresponding positions of the key feature points on the second target tower image are found, and their feature coordinates are recorded. These feature coordinates represent the positional changes of the key feature points in different image frames.

[0063] In this embodiment, the displacement determination module compares the coordinates of key feature points between adjacent frames of the target tower image. For each key feature point, the difference in its coordinates between two adjacent frames is calculated, and this difference is the displacement of the feature point within that time period.

[0064] By analyzing and synthesizing the displacements of multiple key feature points, such as calculating the average and maximum values ​​of the displacements of all key feature points, the vibration displacement of the entire tower or a specific part can be obtained. Vibration displacement data can be used to monitor the operating status of the tower and determine whether there is any abnormal vibration.

[0065] As can be seen from the above, the optical flow calculation module in this embodiment, with the help of the target model, can calculate the optical flow field between adjacent target tower images, clearly presenting the tower's motion information between adjacent frames, providing a foundation for subsequent analysis. The key feature extraction module extracts key feature points from the first frame image, establishing an initial feature reference. These representative and unique points are easy to identify and track in different frames. The feature matching module utilizes the motion information of the optical flow field to narrow the matching search range, greatly improving the efficiency and accuracy of key feature point matching. The displacement judgment module obtains the displacement of each point by comparing the coordinates of key feature points in adjacent frames, and then obtains the tower vibration displacement through comprehensive analysis. This vibration displacement data can effectively monitor the tower's operating status, promptly detect abnormal vibrations, and thus ensure the stable and safe operation of the communication tower.

[0066] refer to Figure 2 In one embodiment of this disclosure, a communication tower with an image recognition device further includes: a target model construction module;

[0067] The target model construction module is specifically used to: extract features from any two frames of target tower images to obtain a first feature map and a second feature map;

[0068] The initial optical flow estimate is obtained based on the similarity between the first feature map and the second feature map;

[0069] The initial optical flow estimate is evaluated based on the target loss function;

[0070] The target model is determined based on the evaluation results.

[0071] In this embodiment, a convolutional neural network can be used to process the two input frames of images. and Feature extraction can automatically extract features meaningful for optical flow estimation from images. Let the feature extraction network be... The extracted first feature map and second feature map are represented as follows:

[0072]

[0073]

[0074] in, Represents the first feature map. This represents the second feature map. The parameters of the feature extraction network are continuously learned and optimized during training to ensure that the extracted features better reflect the essential information of the image, providing a foundation for subsequent similarity calculation and optical flow estimation.

[0075] In this embodiment, the correlation volume is used to measure the similarity of features between the first and second feature maps. The correlation volume is obtained by calculating the dot product of the feature vectors at each location in the feature maps and dividing by the vector norm. The correlation volume reflects the degree of similarity between features at different locations. By considering both the direction and magnitude of the feature vectors, the similarity between features can be measured more accurately.

[0076] For the first feature map Each position in Second feature map Each position in Correlation volume The calculation formula is:

[0077]

[0078] in, Represents the vector dot product. The norm of a vector.

[0079] Based on the similarity between the first and second feature maps, the motion of each point in the image is preliminarily inferred, resulting in an initial optical flow estimate. High-similarity regions within the correlation volume correspond to the motion trajectories of objects in the image, thus providing a basis for optical flow estimation.

[0080] In this embodiment, the optical flow estimation results are updated based on the cyclic update module. Let... Let be the hidden state at the t-th iteration. Let be the optical flow estimate at the t-th iteration. The calculation steps of the iterative update module are as follows:

[0081] Current optical flow Bilinear sampling is performed onto the correlation volume C to obtain the sampled correlation features. Then , and spliced ​​together as input :

[0082]

[0083]

[0084] In this embodiment, a gated loop unit is used to update the hidden state. The calculation formula for the gated loop unit is as follows:

[0085]

[0086]

[0087]

[0088]

[0089] in, This indicates that the door is being reset. Indicates an update to the door. This represents the Sigmoid function. and This represents the learnable weight matrix. Indicates the candidate hidden state. This represents the hyperbolic tangent function.

[0090] In this embodiment, a small convolutional network is used. Based on the hidden state Update optical flow :

[0091]

[0092]

[0093] in, Represents a convolutional network The parameters.

[0094] Convolutional networks can learn the mapping relationship between hidden states and optical flow updates, thereby continuously optimizing the optical flow estimation results.

[0095] In this embodiment, the initial optical flow estimate is evaluated based on the multi-scale mean absolute error loss function;

[0096] Assume that N iterations are performed to obtain N optical flow estimation results. The corresponding real optical flow is The loss function L is defined as:

[0097]

[0098] in, The weight representing the optical flow loss at the i-th scale can be adjusted according to the importance of optical flow at different scales, allowing the model to focus more on optical flow estimation at important scales during training. The mean absolute error loss function measures the error between the optical flow estimation result and the true optical flow; minimizing the loss function can make the model's optical flow estimation more accurate.

[0099] Based on the evaluation results of the loss function, optimization algorithms (such as stochastic gradient descent, Adam, etc.) can be used to optimize the model parameters. and Adjustments are made. Parameters are iteratively updated until the loss function converges to a small value; the resulting model is the target model. The target model can accurately estimate the optical flow between adjacent target tower images, providing support for subsequent image recognition and tower vibration monitoring.

[0100] As can be seen from the above, this embodiment can accurately capture key information of the image by extracting feature maps from any two frames of target tower images. An initial optical flow estimate is obtained based on the similarity of the feature maps, providing a basis for analyzing tower motion. Using a target loss function to evaluate the initial optical flow estimate allows for timely detection of deviations. Determining the target model based on the evaluation results effectively improves the accuracy of optical flow estimation, thereby enabling more precise monitoring of tower vibration and ensuring the stable operation of communication towers.

[0101] In one embodiment of this disclosure, a communication tower with an image recognition device further includes: a similarity calculation module;

[0102] The similarity calculation module is specifically used to: calculate the similarity between the first feature map and the second feature map based on the first formula;

[0103] The first formula is:

[0104]

[0105] in, This represents the similarity between the first feature map and the second feature map. This represents the eigenvector of the first feature map. This represents the feature vector of the second feature map. This represents the Euclidean distance between the first feature map and the second feature map. This represents the weighting coefficients corresponding to the Euclidean distance. The cosine similarity between the first and second feature maps is represented. This represents the weight coefficient corresponding to the cosine similarity. Represents the feature distribution weights.

[0106] In this embodiment, let the feature vectors extracted from the two frames of images be respectively and , where n is the dimension of the feature vector. These feature vectors are numerical representations of image features, containing key information about the image, such as texture and shape.

[0107] Define a composite distance D, and the formula for calculating the composite distance D is:

[0108]

[0109] in, Representing the Euclidean distance, it reflects the cumulative numerical differences across the various dimensions of the feature vectors. A larger Euclidean distance indicates that the two feature vectors are farther apart in space, and the greater the potential difference between their images. This represents cosine similarity, which focuses on the directional similarity between two feature vectors. This is then converted to cosine distance; the larger the value, the greater the difference in vector direction.

[0110] The composite distance D combines Euclidean distance and cosine distance, using weighted coefficients. and This approach balances the importance of both factors in comprehensive distance calculation. It allows for the simultaneous consideration of the spatial location and orientation information of feature vectors, avoiding the limitations of a single distance metric.

[0111] In this embodiment, to further consider the distribution of features throughout the feature space, a weight W based on the feature distribution is introduced. Assume the feature vector... and The overall distribution of the characteristics can be represented by the mean vector. Covariance Matrix To describe.

[0112] First, calculate the Mahalanobis distance from the eigenvectors to the population mean:

[0113]

[0114]

[0115] Then, define the weights W of the feature distribution:

[0116]

[0117] in, It is an adjustable parameter used to control the degree of influence of Mahalanobis distance difference on the weights.

[0118] When the difference between the Mahalanobis distances of two feature vectors to the population mean is small, the weight W of the feature distribution is close to 1, indicating that the two feature vectors are relatively similar in the feature space. When the difference is large, the weight W of the feature distribution will decrease, reducing the influence of the overall distance in similarity calculation.

[0119] Finally, the combined distance D and the weights W of the feature distribution are combined to obtain the final similarity metric S:

[0120] As can be seen from the above, this embodiment combines Euclidean distance and cosine distance, taking into account both the spatial location and orientation information of feature vectors, thus avoiding the limitations of a single distance metric. By utilizing Mahalanobis distance and overall feature distribution information, the similarity calculation is adjusted according to the distribution of feature vectors in the feature space, enabling a more accurate measurement of the similarity of features between two image frames even when feature distribution is uneven.

[0121] In one embodiment of this disclosure, the key feature extraction module is specifically used for:

[0122] The first target tower image and the first infrared tower image are spatially aligned to obtain a preprocessed image;

[0123] An initial feature map is obtained based on the association weights between each pixel in the preprocessed image and its surrounding pixels;

[0124] Key feature points are determined based on the importance of each target feature in the initial feature map.

[0125] In this embodiment, the first infrared tower image can be acquired by an infrared camera.

[0126] Acquire high-resolution visible light images of the primary target tower, ensuring clarity and a complete representation of the tower's structure. Simultaneously acquire infrared images of the same tower, as these images reflect the temperature distribution on the tower's surface, helping to identify potential abnormal heating points that may correspond to structural defects.

[0127] Feature-based registration algorithms, such as SIFT combined with random sampling consensus, are used to accurately register visible light and infrared images, aligning them spatially to obtain a preprocessed image.

[0128] Traditional convolutional neural networks often focus only on local information when extracting features. This embodiment introduces a context-aware convolutional neural network to calculate the association weights between each pixel and its surrounding pixels, thereby enhancing the perception of global information. For example, when extracting features of an iron tower, the relationship between the tower and its surrounding environment (such as terrain and other buildings) can be taken into account.

[0129] The preprocessed image is input into a context-aware convolutional neural network, which outputs an initial feature map. This initial feature map not only contains local features of the tower but also incorporates contextual information, helping to more accurately locate and identify the tower's key structures.

[0130] Each target feature in the initial feature map is evaluated. Importance assessment can be based on multiple factors, such as feature variance, information entropy, and location in the image. Features with higher variance typically contain more variation information, while features with higher information entropy have greater uncertainty and uniqueness. Features located in critical structural parts of the tower (such as the top or corners) can also be assigned higher importance weights. Based on the importance scores obtained from the evaluation, a threshold is set, and pixels corresponding to features with importance scores higher than the threshold are identified as key feature points.

[0131] As can be seen from the above, the key feature extraction module in this embodiment fully utilizes the information of multimodal images by aligning multimodal images spatially, generating initial feature maps based on correlation weights, and determining key feature points. It also mines the contextual relationships between pixels, thereby accurately extracting representative key feature points, providing strong support for subsequent image analysis and tower status monitoring.

[0132] In one embodiment of this disclosure, a communication tower with an image recognition device further includes: a feature enhancement module;

[0133] The feature enhancement module is used for:

[0134] In response to the fact that the importance of the target feature in the initial feature map is greater than a first preset value, the resolution of the target feature is increased;

[0135] In response to the fact that the importance of the target feature in the initial feature map is less than or equal to a first preset value, the resolution of the target feature is reduced.

[0136] In this embodiment, the feature enhancement module is used to perform targeted resolution adjustment on the target features in the initial feature map to highlight important features and suppress unimportant features, thereby improving the accuracy of subsequent image analysis and recognition.

[0137] Before performing feature enhancement operations, it is necessary to evaluate the importance of the target features in the initial feature map. Importance evaluation can be based on various factors, such as feature variance, information entropy, and contribution to the judgment of the tower's state, assigning an importance score to each target feature.

[0138] In this embodiment, when the importance of a target feature in the initial feature map is detected to be greater than a first preset value, it indicates that the target feature has high value for image recognition and tower status monitoring. The first preset value is a pre-set threshold used to distinguish between important and unimportant features.

[0139] To more clearly capture and analyze the detailed information of these important features, it is necessary to increase the resolution of the target features. This can be achieved through deep learning-based super-resolution convolutional neural networks or enhanced super-resolution generative adversarial networks. By enlarging and enhancing the details of low-resolution target feature images, their resolution is improved, making the texture, edges, and other information of important features more clearly discernible.

[0140] When the importance of a target feature is less than or equal to a first preset value, it indicates that the feature has a relatively small impact on image recognition and tower condition monitoring, such as noise or irrelevant information. To reduce the interference of these unimportant features on subsequent analysis, the resolution of the target feature can be reduced. Image downsampling methods, such as average pooling and max pooling, can be used. This reduces the detailed information of the features, decreases the data volume, and highlights the dominant role of important features.

[0141] As can be seen from the above, this embodiment effectively improves image quality and feature discrimination by adjusting the resolution of target features of different importance in the initial feature map through the feature enhancement module. This enhances the expressive power of important features, enabling subsequent image recognition devices to more accurately identify the state and features of the tower; simultaneously, it suppresses interference from unimportant features, reducing computational load and the possibility of misjudgment, thereby improving the performance and efficiency of the entire communication tower image recognition system.

[0142] In one embodiment of this disclosure, the feature matching module is specifically used for:

[0143] Based on the coordinates of each key feature point in the first target tower image, the motion vector corresponding to each key feature point is obtained from the optical flow field;

[0144] The initial coordinates of key feature points in the image of the second target tower are determined based on motion vectors.

[0145] The search window is determined based on the vibration amplitude of the communication tower, with the initial coordinates as the center.

[0146] In response, within the search window range, descriptors for all key feature points of the first target tower image and the second target tower image are calculated respectively within the corresponding search window range;

[0147] The descriptors of key feature points in the first target tower image are matched with the descriptors of key feature points in the second target tower image to obtain the feature coordinates of the key feature points on the second target tower image.

[0148] In this embodiment, the optical flow field describes the motion information of each point in the image between adjacent frames, including motion vectors that represent the direction and magnitude of displacement of a point between two frames. The feature matching module first finds the corresponding motion vector from the pre-calculated optical flow field based on the coordinates of each key feature point in the first target tower image. These motion vectors reflect the possible motion of the key feature points from the first frame to the second frame, providing a basis for subsequently predicting the position of the key feature points in the second target tower image.

[0149] In this embodiment, based on the acquired motion vectors, the positions of key feature points in the second target tower image can be preliminarily predicted. Since the motion vector represents the displacement of the feature point, its approximate position in the second target tower image can be estimated by adding the corresponding motion vector to the coordinates of the feature point in the first target tower image.

[0150] Let the coordinates of a key feature point in the image of the first target iron tower be... Its corresponding motion vector is The initial coordinates of this key feature point in the image of the second target tower are then determined. Through formula , The calculation yielded the result.

[0151] In this embodiment, due to potential errors in the optical flow field calculation and the actual vibration of communication towers, the actual positions of key feature points in the second target tower image may deviate from their initial coordinates. Therefore, a search window centered on the initial coordinates is needed for more accurate feature matching within this range.

[0152] The size of the search window can be determined based on the vibration amplitude of the communication tower. If the tower vibration amplitude is large, the search window needs to be set larger to ensure that the locations where key feature points may appear are covered; conversely, if the vibration amplitude is small, the search window can be reduced accordingly. This ensures the accuracy of the matching while reducing unnecessary computation.

[0153] In this embodiment, the feature descriptor is a vector that describes the local features of key feature points. It has properties such as rotation invariance and scale invariance, and can accurately represent the feature information of feature points under different image conditions. By calculating the feature descriptor, the features of feature points can be quantified, which facilitates subsequent matching operations.

[0154] Within the search window, feature descriptors are calculated for all key feature points in both the first and second target tower images. Feature descriptors can be calculated using scale-invariant feature transformation, accelerated robust feature generation, or the ORB algorithm.

[0155] In this embodiment, by comparing the similarity between the descriptors of key feature points in the first target tower image and the descriptors of key feature points in the second target tower image, the correspondence between the first target tower image and the second target tower image can be found, thereby determining the accurate feature coordinates of the key feature points on the second target tower image.

[0156] The descriptors of key feature points in the first target tower image are matched with the descriptors of key feature points in the second target tower image. If a match is successful, the coordinates of the corresponding key feature point in the second target tower image are recorded; these are the feature coordinates of that key feature point in the second target tower image.

[0157] As can be seen from the above, the feature matching module predicts the initial position of key feature points by utilizing optical flow field information, determines the search window by combining the vibration amplitude of the communication tower, and then calculates and matches feature descriptors, thereby achieving the goal of accurately matching key feature points in the first target tower image to the second target tower image, providing key data for subsequent calculation of the tower's vibration displacement.

[0158] In one embodiment of this disclosure, the feature matching module is further configured to:

[0159] In response to the search window range, the feature point closest to the descriptor of each key feature point in the first target tower image to each key feature point in the corresponding second target tower image is taken as the matching point.

[0160] In this embodiment, the search window is a region determined based on the initial coordinates of key feature points in the first target tower image and the vibration amplitude of the communication tower in the second target tower image. Since there may be errors in the optical flow field calculation, and the tower vibrates, the actual position of the key feature points in the second target tower image may deviate from the initial coordinates. The search window is set to find accurate matching positions of key feature points within a reasonable range, reducing unnecessary global searches and improving matching efficiency.

[0161] For each key feature point in the first target tower image, the feature point with the closest descriptor distance within the search window of the second target tower image is found as the matching point. In this embodiment, Euclidean distance can be used to calculate the descriptor distance.

[0162] Let the descriptor of a key feature point in the image of the first target tower be a vector. The descriptor of a feature point within the second target tower image search window is a vector. Then the Euclidean distance between them is:

[0163]

[0164] in, The descriptor represents the Euclidean distance between the descriptor of a key feature point in the first target tower image and the descriptor of a key feature point in the second target tower image, where n is the dimension of the descriptor vector. and They are vectors and The i-th component.

[0165] Within the search window of the second target tower image, calculate the distance between the descriptor of each key feature point in the first target tower image and the descriptors of all feature points within the search window. Select the feature point with the smallest distance as the matching point for that key feature point. If the descriptors of two feature points are very close, it indicates that they correspond to the same physical point in the image, meaning they are a match.

[0166] As can be seen from the above, this embodiment finds matching points by calculating the distance between its descriptor and the descriptors of all feature points within the search window, based on the nearest neighbor matching principle. This allows for accurate mapping of key feature points in the first target tower image to the second target tower image, providing crucial correspondence for subsequent calculations of the tower's vibration displacement, thereby achieving effective monitoring of the communication tower's status.

[0167] In one embodiment of this disclosure, the displacement determination module is specifically used for:

[0168] Calculate the displacement of key feature points between adjacent frames of the target tower image;

[0169] The vibration displacement of the tower is obtained based on the average displacement of all key feature points.

[0170] In this embodiment, for each key feature point, its displacement between adjacent frames of the target tower image can be obtained by calculating the difference between the coordinates of the feature point in the two frames. Assume the coordinates of a key feature point in the first frame image are... The coordinates in the second frame image are Then the displacement vector of the key feature point can be expressed as:

[0171]

[0172] The magnitude (modulus) of the displacement is:

[0173]

[0174] In this way, the displacement of each key feature point between adjacent frames can be obtained.

[0175] In this embodiment, the displacement of a single key feature point may be affected by factors such as local noise and image acquisition errors, and cannot completely and accurately represent the vibration of the entire tower. However, the displacement data of all key feature points contains motion information of different parts of the tower. By calculating their average value, the displacement of each key feature point can be comprehensively considered, reducing the interference of individual abnormal displacements, thereby obtaining a value that better reflects the overall vibration state of the tower.

[0176] In one embodiment of this disclosure, the displacement determination module is further configured to:

[0177] The average displacement of all key feature points is calculated based on the second formula;

[0178] The second calculation formula is:

[0179]

[0180] in, This represents the average displacement of all key feature points. This represents the importance weight of the i-th key feature point. This represents the magnitude of the displacement vector of the i-th key feature point. This represents the cosine of the angle between the displacement vector and the reference vector at each key feature point. This represents the total number of key feature points.

[0181] Traditional methods for calculating the average displacement of all key feature points typically involve simply adding the displacement vectors of each key feature point across adjacent frames and then dividing by the number of key feature points. This method fails to consider the varying importance of different key feature points in reflecting the tower's vibration characteristics, nor does it take into account the comprehensive impact of displacement direction on the overall vibration.

[0182] In this embodiment, suppose we have n key feature points. In two adjacent image frames (the first frame and the second frame), the coordinates of the i-th key feature point in the first frame are... In the second frame, the coordinates are .

[0183] The displacement vector of the i-th key feature point It can be represented as:

[0184]

[0185] The magnitude of its displacement is:

[0186]

[0187] Key feature points at different locations contribute differently to reflecting the overall vibration characteristics of the tower. Traditional methods using simple averaging cannot capture this difference, potentially leading to calculation results that do not accurately reflect the actual vibration of the tower. Therefore, this embodiment defines an importance weight for each key feature point. Importance weight The weight of a key feature point can be determined based on its location within the tower structure and its sensitivity to vibration. For example, key feature points located at the top of the tower are more sensitive to vibration and can therefore have a higher weight; while key feature points located in the relatively stable area at the bottom of the tower can have a lower weight. The weight must meet certain conditions. .

[0188] To comprehensively consider the direction of displacement, the dot product operation of vectors can be introduced. A reference vector is selected. Reference vector This can be determined based on the main vibration direction or structural characteristics of the tower. Calculate each displacement vector. With reference vector cosine value of the included angle , the calculation formula is:

[0189]

[0190] Finally, the average displacement of all key feature points is calculated. The formula is:

[0191]

[0192] As can be seen from the above, this embodiment, by multiplying the displacement magnitude of each key feature point by its importance weight and the cosine of the included angle, and then summing the values ​​for all key feature points, obtains an average value that can more comprehensively and accurately reflect the overall vibration of the tower. It considers both the differences in importance of different key feature points and the influence of displacement direction on the overall vibration, resulting in higher accuracy and reliability compared to traditional simple averaging methods.

[0193] refer to Figure 2 In one embodiment of this disclosure, a communication tower with an image recognition device further includes:

[0194] The vibration analysis module is used to analyze the tower based on its vibration displacement to obtain the tower's operating status.

[0195] In this embodiment, multiple vibration displacement safety thresholds of different levels can be preset based on the tower's design standards, its environment, and past operating experience. For example, a slight vibration threshold, a moderate vibration threshold, and a severe vibration threshold can be set.

[0196] The currently acquired tower vibration displacement is compared in real time with these preset thresholds. If the vibration displacement is below the minor vibration threshold, it can be preliminarily determined that the tower is in a normal and stable operating state, with a solid structure and no obvious abnormal vibration. If the vibration displacement exceeds the minor vibration threshold but is below the moderate vibration threshold, it indicates that the tower may have some minor anomalies, such as being disturbed by minor wind or surrounding construction, but it is still within an acceptable operating range and requires continuous monitoring. When the vibration displacement exceeds the moderate vibration threshold or even reaches the severe vibration threshold, it means that the tower may have more serious problems, such as structural loosening or component damage, requiring immediate further inspection and maintenance.

[0197] As can be seen from the above, the vibration analysis module can comprehensively assess the operating status of the tower and draw specific assessment conclusions, such as "operating normally," "minor anomalies exist, requiring attention," and "serious problems exist, requiring immediate repair." These conclusions can be displayed to maintenance personnel through a visual interface and can also generate corresponding reports, providing strong support for communication tower operation and maintenance decisions, ensuring the stable operation of the tower and the reliability of communication services.

[0198] The above embodiments are only used to illustrate the technical solutions of this disclosure, and are not intended to limit it. Although this disclosure has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this disclosure.

Claims

1. A communication tower with an image recognition device, characterized in that, include: The optical flow calculation module is used to input the target tower image sequence into the target model to obtain the optical flow field between adjacent target tower images. The adjacent frame target tower images are the adjacent target tower images in the target tower image sequence. The key feature extraction module is used to extract features from the first target tower image to obtain key feature points. The first target tower image is the first frame of the target tower image in the target tower image sequence. The key feature extraction module is specifically used for, The first infrared tower image is acquired by an infrared camera and is used to reflect the temperature distribution on the surface of the tower. The first target tower image is a visible light image; Feature-based registration algorithms register visible light images and infrared images, aligning them spatially to obtain a preprocessed image. The preprocessed image is input into a context-aware convolutional neural network, and an initial feature map is obtained based on the association weights between each pixel in the preprocessed image and its surrounding pixels. The initial feature map contains local features of the tower. For each target feature in the initial feature map, the importance score is evaluated, and the pixels corresponding to the features with an importance score higher than the threshold are identified as key feature points. The feature matching module is used to match the key feature points with the second target tower image based on the optical flow field to obtain the feature coordinates of the key feature points on the second target tower image. The second target tower image is any frame of the target tower image sequence other than the first frame of the target tower image. The displacement determination module is used to obtain the vibration displacement of the tower based on the coordinates of the key feature points between adjacent frame images of the target tower.

2. A communication tower with an image recognition device as described in claim 1, characterized in that, Also includes: Target model building module; The target model construction module is specifically used to: extract features from any two frames of target tower images to obtain a first feature map and a second feature map; An initial optical flow estimate is obtained based on the similarity between the first feature map and the second feature map; The initial optical flow estimate is evaluated based on the target loss function; The target model is determined based on the evaluation results.

3. A communication tower with an image recognition device as described in claim 2, characterized in that, Also includes: Similarity calculation module; The similarity calculation module is specifically used to: calculate the similarity between the first feature map and the second feature map based on the first formula; The first formula is: in, This represents the similarity between the first feature map and the second feature map. This represents the eigenvector of the first feature map. This represents the feature vector of the second feature map. This represents the Euclidean distance between the first feature map and the second feature map. This represents the weighting coefficients corresponding to the Euclidean distance. The cosine similarity between the first and second feature maps is represented. This represents the weight coefficient corresponding to the cosine similarity. Represents the feature distribution weights.

4. A communication tower with an image recognition device as described in claim 1, characterized in that, It also includes: a feature enhancement module; The feature enhancement module is used for: In response to the fact that the importance of the target feature in the initial feature map is greater than a first preset value, the resolution of the target feature is increased; In response to the fact that the importance of the target feature in the initial feature map is less than or equal to a first preset value, the resolution of the target feature is reduced.

5. A communication tower with an image recognition device as described in claim 1, characterized in that, The feature matching module is specifically used for: Based on the coordinates of each key feature point in the first target tower image, the motion vector corresponding to each key feature point is obtained from the optical flow field; Based on the motion vector, the initial coordinates of key feature points in the second target tower image are determined; The search window is determined based on the vibration amplitude of the communication tower, with the initial coordinates as the center. In response, within the search window range, descriptors for all key feature points of the first target tower image and the second target tower image are calculated respectively within the corresponding search window range; The descriptors of key feature points in the first target tower image are matched with the descriptors of key feature points in the second target tower image to obtain the feature coordinates of the key feature points on the second target tower image.

6. A communication tower with an image recognition device as described in claim 5, characterized in that, The feature matching module is also specifically used for: In response to the search window range, the feature point closest to the descriptor of each key feature point in the first target tower image to each key feature point in the corresponding second target tower image is taken as the matching point.

7. A communication tower with an image recognition device as described in claim 1, characterized in that, The displacement determination module is specifically used for: Calculate the displacement of the key feature points between adjacent frames of the target tower image; The vibration displacement of the tower is obtained based on the average displacement of all the key feature points.

8. A communication tower with an image recognition device as described in claim 7, characterized in that, The displacement determination module is also specifically used for: The average displacement of all the key feature points is calculated based on the second formula; The second calculation formula is: in, This represents the average displacement of all key feature points. This represents the importance weight of the i-th key feature point. This represents the magnitude of the displacement vector of the i-th key feature point. This represents the cosine of the angle between the displacement vector and the reference vector at each key feature point. This represents the total number of key feature points.

9. A communication tower with an image recognition device as described in claim 1, characterized in that, Also includes: The vibration analysis module is used to analyze the tower based on its vibration displacement to obtain the tower's operating status.