Power tower recognition method and system based on character symbols

By constructing a multimodal mapping relationship between text symbols and tower components, and utilizing radar arrays and dynamic noise suppression models, a three-dimensional recognition model of power transmission towers resistant to environmental interference was reconstructed. This solved the problems of low recognition accuracy and high implementation difficulty in existing technologies, and achieved high-precision and robust three-dimensional reconstruction.

CN120689859BActive Publication Date: 2026-01-27TCXY (TIANJIN) MOULD FRAME CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510908933.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2026-01-27
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing power tower identification technologies are sensitive to environmental factors, have low identification accuracy, are difficult to implement, and are hard to promote and apply in resource-limited areas.

Method used

The method for identifying power transmission towers based on text symbols constructs a multimodal mapping relationship between text symbols and tower components, uses radar arrays to acquire reflected signals, analyzes phase difference data, generates a set of penetration features, constructs a dynamic noise suppression model, and combines a shared weight encoder to perform cross-modal association learning to reconstruct a three-dimensional recognition model resistant to environmental interference.

Benefits of technology

It achieves high-precision 3D reconstruction in complex environments, enhances the interpretability and automation of component identification, improves identification accuracy and robustness, has strong adaptability, and reduces implementation difficulty.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120689859B_ABST
    Figure CN120689859B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of intelligent detection of power facilities, and provides a power tower identification method and system based on character symbols, which are used to solve the problems of insufficient resolution and lack of semantic correlation of models in the prior art. The application comprises the following steps: constructing a multi-modal correlation rule of a mapping relationship between character symbols on the surface of a power tower and tower body components, screening a target reflection area, and analyzing and generating a penetrating feature set; using a dynamic noise suppression model to perform electromagnetic attenuation compensation on the penetrating features, and generating an anti-interference compensated feature set; inputting the compensated features and coating geometric identification features into a shared weight encoder, fusing character symbol semantic vectors, and extracting internal support structure features and external coating texture features; and reconstructing a three-dimensional identification model conforming to load characteristics through bolt hole topological distribution constraints and main beam curvature continuity constraints. The technical scheme provided by the application can realize high-precision three-dimensional identification of power towers and anti-interference structure restoration in a complex environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent detection technology for power facilities, and in particular to a method and system for identifying power transmission towers based on text symbols. Background Technology

[0002] With the rapid development of the power industry, the number and types of power transmission towers are increasing daily, and their safety and stability are directly related to the reliable operation of the power grid. Traditional manual inspection methods are inefficient and easily affected by environmental factors, making it difficult to meet the needs of modern power grid management. Therefore, developing a power transmission tower identification method based on text symbols that can automatically, quickly, and accurately identify tower types and statuses has become a pressing technical challenge for the power industry.

[0003] Currently, power transmission tower identification mainly relies on image processing and machine learning techniques. Common approaches include deep learning-based image classification and feature extraction-based pattern recognition. For example, convolutional neural networks (CNNs) are used to classify tower images, or features such as the tower's shape and color are extracted for matching. In addition, some studies have attempted to combine GPS and remote sensing technologies to achieve tower location and identification.

[0004] However, while existing solutions have improved the automation level of power tower identification to some extent, they still have many shortcomings. First, image processing methods are sensitive to environmental factors such as lighting and weather, making it difficult to guarantee identification accuracy. Second, machine learning methods require a large amount of labeled data for training, resulting in high data acquisition and labeling costs. Furthermore, existing solutions largely rely on complex algorithms and hardware, making implementation difficult and hindering their widespread application in resource-limited areas. Therefore, a more efficient, low-cost, and adaptable method for identifying power towers is urgently needed. Summary of the Invention

[0005] This application provides a method and system for identifying power transmission towers based on text symbols, in order to solve the problems of insufficient resolution and lack of semantic relevance in existing technologies.

[0006] Firstly, this application provides a method for identifying power transmission towers based on text symbols, including:

[0007] Based on the text symbols obtained from the surface of the power tower, a multimodal mapping relationship between the text symbols and the tower components is constructed;

[0008] The original reflected signals of the power tower are obtained by radar array, the target reflection area is filtered based on the multimodal mapping relationship, the phase difference data in the original reflected signals is analyzed, and a set of penetrating features containing the outline of the tower skeleton and the positioning information of the corresponding text symbols is generated.

[0009] A dynamic noise suppression model is constructed based on the aforementioned penetration feature set. The dynamic noise suppression model is used to compensate for the attenuation of multipath reflection signals, thereby generating a set of compensation features to resist electromagnetic interference.

[0010] The compensation feature set and the coating geometric identification features are input into a shared weight encoder for cross-modal association learning. The shared weight encoder fuses the semantic vectors of text symbols based on an attention mechanism, and simultaneously extracts the internal support structure features and the external anti-corrosion coating texture features of the tower body. The geometric constraint relationship between the support structure and the coating texture is established by alternately updating the convolution kernel parameters.

[0011] Based on the aforementioned geometric constraints, the missing tower connection components are reconstructed, generating a three-dimensional recognition model of the power tower that is resistant to environmental interference.

[0012] Optionally, the compensation feature set and the coating geometric identification features are input into a shared weight encoder, and the compensation feature set and the coating geometric identification features are extracted through the multi-layer network structure in the shared weight encoder to generate a preliminary feature vector;

[0013] The initial feature vector and the semantic vector of the text symbol are weighted and fused based on the attention mechanism to generate a fused feature vector;

[0014] By performing multi-scale decomposition on the fused feature vector, the internal support structure features and the external anti-corrosion coating texture features of the tower body are extracted respectively. The support structure features are extracted by high-frequency components, and the coating texture features are extracted by low-frequency components.

[0015] Based on the local consistency between the internal support structure features and the external anti-corrosion coating texture features, the convolution kernel parameters are updated alternately, and a geometric constraint relationship between the support structure features and the coating texture features is established.

[0016] Optionally, the preliminary feature vector and the semantic vector of the text symbol are respectively input into independent fully connected layers for feature dimension alignment to generate aligned preliminary feature vector and semantic vector of the text symbol;

[0017] Based on the spatial distribution of the aligned preliminary feature vector and the semantic vector of the text symbol, local symbol association features and global structure association features are constructed, wherein the local symbol association features are generated by the mean of features in the symbol neighborhood, and the global structure association features are generated by cross-region feature difference.

[0018] By cross-projecting the local symbol association features and the global structure association features, a symbol structure association mapping vector is generated, and a dynamic correlation score between the aligned preliminary feature vector and the semantic vector of the text symbols is calculated based on the symbol structure association mapping vector.

[0019] Based on the dynamic relevance score, channel-level weights are assigned to the aligned preliminary feature vector, and symbol-level weights are assigned to the text symbol semantic vector, generating channel-weighted preliminary feature vectors and symbol-weighted semantic vectors;

[0020] The channel-weighted preliminary feature vector and the symbol-weighted semantic vector are concatenated using multi-scale context features to generate a multi-scale context feature concatenation result. The multi-scale context features are extracted by sliding the context window of the symbol semantics, and the window size is positively correlated with the density of text symbols on the surface of the power tower.

[0021] Based on the topological constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale context feature splicing results to generate a fused feature vector.

[0022] Optionally, the local symbol association features are spatially encoded and channel-dimensional encoded respectively to generate spatially encoded local symbol association features and channel-encoded global structural association features;

[0023] By interactively superimposing feature channels, the spatially encoded local symbol association features and the channel-encoded global structure association features are bidirectionally projected to generate a symbol structure association mapping vector.

[0024] The spatial weight matrix and the channel weight matrix are extracted from the symbol structure association mapping vector, wherein the spatial weight matrix is ​​generated based on the local similarity of symbol semantics, and the channel weight matrix is ​​generated based on the global difference of structural features;

[0025] By parametrically weighting the spatial topological constraints of the global structural association features, and fusing the spatial weight matrix and the channel weight matrix, dynamic correlation weight coefficients are generated.

[0026] Based on the dynamic relevance weight coefficient, the cross-modal association strength between the aligned preliminary feature vector and the semantic vector of the text symbol is quantified to generate a dynamic relevance score.

[0027] Optionally, the internal support structure features are decomposed into the spatial distribution features of the connectors and the morphological features of the main beam, and the texture features of the external anti-corrosion coating are decomposed into the distribution features of the corrosion area and the variation features of the coating thickness.

[0028] Based on the proximity constraint between bolt hole locations and coating defect areas, spatial topological encoding is performed on the spatial distribution characteristics of the connectors and the distribution characteristics of corrosion areas to generate corrosion correlation characteristics of the connectors. Based on the synchronicity constraint between tower curvature changes and coating thickness gradients, morphological-thickness coupling encoding is performed on the morphological characteristics of the main beam and the coating thickness variation characteristics to generate main beam thickness coupling characteristics.

[0029] Construct a multi-branch convolutional kernel group, wherein the first branch performs spatial constraint modeling on the corrosion correlation features of the connector, and the second branch performs texture constraint modeling on the thickness coupling features of the main beam;

[0030] Based on the local consistency loss of the spatial constraint modeling results and the texture constraint modeling results, the parameters of the multi-branch convolution kernel group are alternately updated to generate the geometric constraint relationship between the supporting structure features and the coating texture features.

[0031] Optionally, the power tower is subjected to multi-angle signal acquisition by a ring-shaped or linearly deployed radar array to obtain the original set of reflected signals, wherein the direction of the radar beam is aligned with the spatial distribution of the text symbols on the tower surface during signal acquisition.

[0032] Based on the multimodal mapping relationship, beamforming processing is performed on the original set of reflected signals to filter target reflection areas containing text symbol positioning information, wherein the beamforming parameters are dynamically adjusted according to the size of the text symbol and the geometric dimensions of the tower components.

[0033] Phase difference interferometry processing is performed on the original reflection signal of the target reflection area to analyze the phase difference data between adjacent radar channels and generate an initial three-dimensional point cloud of the tower skeleton outline.

[0034] Based on the component correlation of text symbols in multimodal mapping, spatial filtering is performed on the initial three-dimensional point cloud to remove scattering points that are irrelevant to the positioning of text symbols, thereby generating a symbol-constrained tower skeleton point cloud.

[0035] The point cloud of the tower skeleton constrained by the symbol is fused with the multipath reflection intensity from the previous phase difference interferometry processing to generate a set of penetrating features containing the positioning information of the text symbols and the geometric parameters of the tower skeleton.

[0036] Optionally, based on the association between text symbols and tower components in the multimodal mapping relationship, the semantically associated region and the non-associated region of the symbols are divided;

[0037] Within the symbol semantic association area, the scattering points are processed according to the component association of the text symbols. Within the non-symbolic semantic association area, the scattering points that are unrelated to the rust expansion characteristics are eliminated according to the defect boundary distribution of the coating corrosion symbols.

[0038] The initial three-dimensional point cloud is marked with symbolic semantic association regions. During the marking process, the lightning protection symbol region radiates outward from the center of the symbol, the support symbol region extends along the axis of the main beam, and the fixed symbol region extends according to the bolt distribution density matrix, thereby generating a region-marked point cloud.

[0039] Based on the multi-scale accuracy parameters of the region marker point cloud, the scattering points after processing the symbolic semantic association region and the remaining scattering points after removing the non-association region are aggregated at multiple scales to generate a symbol-constrained tower skeleton point cloud.

[0040] Optionally, symbolic noise suppression parameters are extracted by the spatial mapping relationship between the text symbol location information in the penetration feature set and the tower components, and a dynamic noise suppression model is constructed.

[0041] In the dynamic noise suppression model, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower skeleton profile.

[0042] Based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, an interference source spatial weight map is generated, and the interference source spatial weight map is integrated into the dynamic noise suppression model to perform azimuth-related attenuation compensation on the multipath reflection signal and extract the azimuth compensation parameters.

[0043] The symbolic semantic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are weighted and fused in the tower area of ​​the dynamic noise suppression model to update the dynamic noise suppression model.

[0044] The updated dynamic noise suppression model is used to perform iterative path attenuation compensation on the multipath reflection signals in the penetration feature set, thereby generating a compensation feature set for electromagnetic interference resistance.

[0045] Optionally, based on the curvature continuity constraint of the main beam of the tower and the topological distribution constraint of the bolt hole positions in the geometric constraint relationship, the area with missing connecting components is spatially located to generate the bounding box of the missing area;

[0046] The scope is locked within the boundary box of the missing area, and candidate connection component parameters are generated based on the angular constraint relationship between the main beam and the support component.

[0047] Based on the topological distribution constraints of bolt hole positions, the parameters of the candidate connection components are spatially matched and verified. Candidate parameters whose axial deviation from the bolt hole positions exceeds a preset threshold are eliminated, and a preliminary set of connection components is generated.

[0048] Based on the curvature continuity constraint of the main beam of the tower, the initial set of connecting components is morphologically optimized, and the installation angle and contact surface curvature of the connecting components are adjusted to meet the curvature difference threshold of the main beam skeleton point cloud.

[0049] Based on the load distribution characteristics of the support components and the tower frame, the environmental interference resistance reinforcement model is performed on the optimized connection components, wherein the reinforcement strength is dynamically adjusted according to the connection density between the support components and the main beam.

[0050] The reinforced and modeled connecting components are geometrically matched with the point cloud of the tower skeleton to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.

[0051] Secondly, this application provides a power transmission tower identification system based on text symbols, including:

[0052] The module constructs a multimodal mapping relationship between the text symbols on the surface of the power tower and the tower components based on the obtained text symbols.

[0053] The analysis module acquires the original reflected signal of the power tower through the radar array, filters the target reflection area based on the multimodal mapping relationship, analyzes the phase difference data in the original reflected signal, and generates a set of penetrating features containing the outline of the tower skeleton and the positioning information of the corresponding text symbols.

[0054] The compensation module constructs a dynamic noise suppression model based on the penetration feature set, and performs attenuation compensation on the multipath reflection signal through the dynamic noise suppression model to generate a compensation feature set for electromagnetic interference resistance.

[0055] The update module inputs the compensation feature set and the coating geometric identification features into the shared weight encoder for cross-modal association learning. The shared weight encoder fuses the semantic vectors of text symbols based on the attention mechanism, and simultaneously extracts the internal support structure features and the external anti-corrosion coating texture features of the tower body. The geometric constraint relationship between the support structure and the coating texture is established by alternately updating the convolution kernel parameters.

[0056] The generation module reconstructs the missing tower connection components based on the geometric constraints, generating a three-dimensional recognition model of the power tower that is resistant to environmental interference.

[0057] In this embodiment, based on the acquired text symbols on the surface of the power tower, a multimodal mapping relationship between the text symbols and tower components is constructed. The original reflected signals of the power tower are acquired using a radar array. Based on the multimodal mapping relationship, target reflection areas are filtered, and the phase difference data in the original reflected signals is analyzed to generate a set of penetrating features containing the tower skeleton outline and the corresponding text symbol positioning information. A dynamic noise suppression model is constructed based on the penetrating feature set. This model attenuates multipath reflected signals, generating a set of compensation features to resist electromagnetic interference. The compensation feature set and coating geometric identification features are input into a shared weight encoder for cross-modal association learning. The shared weight encoder, based on an attention mechanism, fuses the semantic vectors of the text symbols, simultaneously extracting the internal support structure features and the external anti-corrosion coating texture features. Geometric constraints between the support structure and the coating texture are established by alternately updating the convolution kernel parameters. Based on these geometric constraints, missing tower connection components are reconstructed, generating a three-dimensional recognition model of the power tower resistant to environmental interference.

[0058] The technical solution of this application has the following beneficial effects:

[0059] This application achieves high-precision 3D reconstruction in complex environments through multimodal data fusion and anti-interference mechanisms. First, based on the mapping relationship between surface text symbols and internal components of power transmission towers, semantic identifiers are digitally associated with the physical structure, significantly enhancing the interpretability and automation of component identification. Second, phase difference data is analyzed using radar array penetration detection technology, and the tower skeleton outline is accurately reconstructed by combining spatial positioning of text symbols, breaking through the dependence of traditional visual methods on surface coatings. To address electromagnetic interference and multipath reflection issues, a dynamic noise suppression model improves signal quality through an adaptive compensation mechanism, providing high-fidelity input for feature fusion. A shared weight encoder performs cross-modal association between radar features and coating geometric identifiers, and an attention mechanism fuses text semantics and structural features, establishing geometric topological constraints between the support system and the anti-corrosion coating, significantly improving the accuracy of internal and external feature matching. Finally, an anti-interference 3D recognition model is generated based on multimodal complementary features, and generative repair of missing components is achieved through geometric constraints. It maintains strong robustness even in harsh environments such as rain, fog, and corrosion, forming a full-link solution from physical detection to semantic analysis, providing reliable technical support for intelligent inspection of power facilities.

[0060] These or other aspects of this application will become more apparent in the following description of the embodiments. Attached Figure Description

[0061] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0062] Figure 1 A flowchart of a power tower identification method based on text symbols provided in this application is shown;

[0063] Figure 2 A schematic diagram of the structure of a power tower identification system based on text symbols provided in this application is shown. Detailed Implementation

[0064] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings.

[0065] In some of the processes described in the specification, claims, and accompanying drawings of this application, multiple operations appearing in a specific order are included. However, it should be clearly understood that these operations may not be executed in the order they appear herein, or may be executed in parallel. The operation numbers, such as 101, 102, etc., are merely used to distinguish different operations and do not themselves represent any execution order. Furthermore, these processes may include more or fewer operations, and these operations may be executed sequentially or in parallel. It should be noted that the descriptions such as "first," "second," etc., in this document are used to distinguish different messages, devices, modules, etc., and do not represent a chronological order, nor do they limit "first" and "second" to different types.

[0066] This application aims to develop an intelligent identification and 3D reconstruction system for power transmission towers based on multimodal semantic fusion, achieving a breakthrough in the entire process from text symbol parsing to interference-resistant 3D modeling. The system integrates text symbol semantics with radar penetration detection data to construct a multimodal mapping relationship, realizing the digital association between the semantics of tower components and their physical structure. Finally, based on multimodal feature complementarity and generative repair algorithms, it achieves intelligent reconstruction of missing connection components under complex environmental interference, significantly improving the accuracy of power transmission tower structure identification and the environmental robustness of the 3D model, providing a highly reliable digital solution for intelligent inspection of power grid facilities.

[0067] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0068] Figure 1 A flowchart illustrating a method for identifying power transmission towers based on text symbols is provided in this application embodiment. Figure 1 As shown, the method includes:

[0069] 101. Based on the obtained text symbols on the surface of the power transmission tower, construct a multimodal mapping relationship between the text symbols and the tower components;

[0070] In this step, textual symbols refer to the markings on the surface of the power tower, such as numbers, models, or component names; tower components are the physical parts of the tower, such as crossarms, tower body, insulators, etc.

[0071] Tower components refer to the physical parts of a power transmission tower, such as crossarms, tower body, insulators, ground wire supports, etc., which are the basic units that make up the tower.

[0072] Multimodal mapping refers to establishing a correspondence between text symbols and tower components by associating them with the spatial location and semantic information of the components.

[0073] In this embodiment, firstly, a photoelectric character recognition sensor array captures pre-set text symbols on the surface of the power transmission tower. A multi-scale semantic segmentation algorithm is then used to analyze the morphological features and spatial distribution patterns of the symbols, encoding the recognition results into structured semantic vectors to construct a multimodal mapping relationship between the text symbols and tower components. The system captures images of the power transmission tower surface using a high-resolution camera or drone and extracts the location and semantic information of the text symbols using optical character recognition technology. Next, based on the semantic information of the text symbols, the system associates them with tower components and records their physical location and structural features. Finally, the system stores the multimodal mapping relationship as structured data for subsequent radar signal analysis.

[0074] In a power transmission tower inspection scenario, the system uses a drone to capture images of the tower's surface and identifies the text symbol "Support Rod B". Next, the system associates "Support Rod B" with the tower's support rod components, recording their location, dimensions, and material information. The system then analyzes the structural features of the support rod using a deep learning model and binds them to the text symbol "Support Rod B", constructing a multimodal mapping relationship between the text symbol and tower components. This step provides an accurate foundation for subsequent radar signal analysis by mapping the text symbol to components.

[0075] 102. Obtain the original reflection signal of the power tower through the radar array, filter the target reflection area based on the multimodal mapping relationship, analyze the phase difference data in the original reflection signal, and generate a set of penetrating features containing the outline of the tower skeleton and the positioning information of the corresponding text symbols.

[0076] In this step, a radar array refers to an array of devices used to transmit and receive radar signals. It typically consists of multiple radar sensors and is capable of detecting the position, shape, and motion of objects.

[0077] The original reflected signal refers to the electromagnetic wave signal reflected by the power tower received by the radar array, which contains information about the location, shape and material of the target object.

[0078] Phase difference data refers to the phase difference of radar signals between different receiving points, and is typically used to calculate the outline and position of an object. Phase difference data can reflect the geometric features and spatial distribution of a target object.

[0079] The penetrating feature set refers to the feature set containing the outline of the tower's skeleton and the location information of text symbols, generated through radar signal analysis. This feature set can reflect the internal structure and external markings of the tower.

[0080] In this embodiment, the system first acquires the original reflected signal of the power transmission tower using a radar array, and then uses Doppler radar technology to collect electromagnetic wave reflection data from the tower. Next, the system filters the target reflection area based on multimodal mapping relationships. Then, the system uses orthogonal polarization interferometry to analyze the phase difference data in the original reflected signal, generating the tower skeleton outline and text symbol positioning information. The system calculates the outline and position of the crossbeam components using a phase difference algorithm and binds them to the text symbols. Finally, the system stores the penetration feature set as structured data for subsequent noise suppression.

[0081] In the scenario of power transmission tower inspection, the system acquires the raw reflected signals of the power transmission tower through a radar array. Based on multimodal mapping, the system locates the reflection area corresponding to the text symbol "Support Rod B" and extracts the radar signal from that area. Next, the system analyzes the phase difference data in the raw reflected signal to generate the skeleton outline of the support rod and the positioning information of the text symbol. The system calculates the geometry and position of the support rod using a phase difference algorithm and binds it to the text symbol "Support Rod B". Finally, the system stores the penetration feature set as structured data for subsequent noise suppression. This step provides an accurate set of penetration features for subsequent noise suppression.

[0082] 103. Construct a dynamic noise suppression model based on the penetration feature set, and use the dynamic noise suppression model to attenuate the multipath reflection signal to generate a compensation feature set for electromagnetic interference resistance.

[0083] In this step, the dynamic noise suppression model refers to a model used to suppress noise in radar signals, typically implemented using signal processing techniques. This model can dynamically identify and eliminate multipath reflections and other interference.

[0084] Multipath reflection signals refer to interference signals generated by radar signals due to multiple reflections during propagation, which usually leads to distortion of the outline and position information of the target object.

[0085] The compensated feature set refers to the noise-free feature set generated by the dynamic noise suppression model, which can reflect the true contour and position information of the target object.

[0086] In this embodiment, firstly, the system constructs a dynamic noise suppression model based on a penetration feature set. The system then identifies and suppresses multipath reflection signals using an adaptive filtering algorithm. Next, based on the correlation between the tower material's dielectric constant and the millimeter-wave scattering path length, the system performs attenuation compensation on the multipath reflection signals, generating a compensation feature set to resist electromagnetic interference. The system then eliminates noise interference in the support rod region using a signal reconstruction algorithm, generating distortion-free support rod contour information. Finally, the system stores the compensation feature set as structured data for subsequent cross-modal correlation learning.

[0087] When other metal structures exist near the power transmission tower, the system eliminates noise interference in multipath reflection signals using a dynamic noise suppression model. For example, electromagnetic wave signals reflected from the ground are identified as noise and attenuated using a filtering algorithm. Ultimately, the system generates a set of compensated features to ensure the accuracy of the crossarm frame contour information. This data provides a reliable foundation for subsequent cross-modal correlation learning. In a practical application, the system also identified multipath reflection signals in the tower area, eliminated interference using the dynamic noise suppression model, and generated accurate tower contour information. This data provides crucial support for subsequent 3D reconstruction.

[0088] 104. Input the compensation feature set and the coating geometric identification features into the shared weight encoder for cross-modal association learning, so as to extract the internal support structure features and the external anti-corrosion coating texture features of the tower body simultaneously by fusing the semantic vector of the text symbol through the shared weight encoder based on the attention mechanism, and establish the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters.

[0089] In this step, the coating geometry identification features refer to the geometry and identification features of the coating on the surface of the power tower, which are typically used to assess the integrity and corrosion resistance of the coating.

[0090] A shared-weight encoder is a neural network encoder used for cross-modal data association learning, typically implemented based on an attention mechanism. This encoder can fuse data from different modalities to generate a unified feature representation.

[0091] Geometric constraint relationships refer to the geometric relationship between the internal support structure of the tower and the texture of the external anti-corrosion coating, which can reflect the integrity of the tower structure and the distribution characteristics of the coating.

[0092] In this application example, firstly, the compensation feature set and the coating geometric identification features are input into a shared weight encoder for cross-modal association learning. The shared weight encoder then fuses semantic vectors of text symbols based on an attention mechanism. Next, the attention mechanism dynamically adjusts the model's focus on the internal support structure features and the external anti-corrosion coating texture features. For example, when the coating texture information is relatively vague, the model will focus more on the support structure features. Then, by alternately updating the convolution kernel parameters, the geometric constraint relationship between the support structure and the coating texture is established. For example, the spatial distance and relative position between the coating texture and the support structure are determined to ensure their spatial consistency. Finally, feature data containing the association relationship between the support structure and the coating texture is generated, providing a basis for 3D reconstruction.

[0093] In power tower inspection, the system fuses the crossarm frame outline and coating texture features using a shared weight encoder. For example, the system determines the spatial distance between the crossarm frame and the coating to be 0.1 meters and establishes the geometric constraints between them. This information provides crucial support for subsequent 3D reconstruction. In a practical application, the system also identified the coating texture in the tower body area, determined its spatial relationship with the tower frame through cross-modal association learning, and further refined the geometric constraints. This data provides an important basis for the final 3D recognition model.

[0094] 105. Based on the geometric constraints, reconstruct the missing tower connection components to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.

[0095] In this step, missing tower connection components refer to components that are missing due to corrosion or external damage, such as missing bolts or broken insulators.

[0096] The environmentally resistant 3D recognition model of power towers refers to a building information model that integrates multimodal data, supporting dynamic updates such as rust diffusion simulation and interference-resistant rendering such as removing rain and fog noise.

[0097] In this application example, firstly, a multimodal fusion decision-making mechanism is triggered based on meteorological sensor data. When the rain and fog concentration exceeds a preset threshold, spatial topological calibration is performed on the internal support structure features of the tower, the texture features of the external anti-corrosion coating, and the positioning information of text symbols. Based on geometric constraints, missing parts of the tower's connecting components are identified. Using 3D reconstruction technology, combined with the tower skeleton outline and coating texture information, the missing tower connecting components are reconstructed. Simultaneously, an environmental noise suppression algorithm is used to eliminate the impact of environmental interference on the model. For example, natural factors such as wind and rain may affect the accuracy of the 3D model; the system eliminates these interferences through filtering algorithms and signal enhancement techniques. Finally, an environmentally resistant 3D recognition model of the power tower is generated, providing reliable technical support for inspection and maintenance.

[0098] When a crossarm connecting component of a power transmission tower is missing, the system restores its complete structure using 3D reconstruction technology. The system reconstructs the length and shape of the crossarm connecting component based on geometric constraints and generates a noise suppression algorithm to generate a 3D recognition model resistant to environmental interference. This model provides inspection personnel with accurate information about the tower's structure, improving inspection efficiency. In a practical application, the system also identified missing parts of the tower's connecting components, restored their structure through 3D reconstruction, and generated a complete 3D recognition model using an environmental noise suppression algorithm. This data provides crucial support for the maintenance and repair of power transmission towers.

[0099] In summary, steps 101 to 105 achieved multimodal mapping between surface text symbols on power transmission towers and tower components. A set of penetrating features was generated using radar arrays and signal processing techniques, and noise interference in multipath reflection signals was eliminated through a dynamic noise suppression model. Geometric constraints were established through cross-modal association learning to reconstruct missing tower connection components, ultimately generating a three-dimensional recognition model of power transmission towers resistant to environmental interference. This model can accurately identify tower structures, providing reliable technical support for the inspection and maintenance of power transmission towers and significantly improving the safety and stability of the power system. For example, in an actual inspection, the system successfully identified and reconstructed a crossarm connection component missing due to obstruction, providing maintenance personnel with accurate structural information and avoiding potential safety hazards.

[0100] To address the cross-modal association learning problem between the internal support structure and the texture features of the external anti-corrosion coating of the tower, in some embodiments, step 103 involves inputting the compensated feature set and the coating geometric identifier features into a shared weight encoder for cross-modal association learning. This shared weight encoder, based on an attention mechanism, fuses semantic vectors of text symbols to simultaneously extract the features of the internal support structure and the texture features of the external anti-corrosion coating. Furthermore, it establishes the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters. This includes:

[0101] 201. Input the compensation feature set and the coating geometric identification features into the shared weight encoder, and extract features from the compensation feature set and the coating geometric identification features through the multi-layer network structure in the shared weight encoder to generate a preliminary feature vector;

[0102] In step 201, the compensation feature set refers to the feature data related to coating geometric deviations collected by sensors, including physical property data such as coating thickness, surface roughness, and coating adhesion. The coating geometric identification features refer to the geometric shape features of the coating extracted through image processing techniques, such as coating edges, contours, and surface curvature. The shared weight encoder is a multi-layer neural network structure whose weights are shared between the inputs of the compensation feature set and the coating geometric identification features, used to extract common features from both. The preliminary feature vector refers to the fused feature representation extracted by the shared weight encoder, used for subsequent feature fusion and decomposition.

[0103] In this embodiment, the compensation feature set and the coating geometric identification features are first input into a shared weight encoder. The shared weight encoder employs a multi-layer fully connected network structure, with each layer transforming the input features through a non-linear activation function. Due to weight sharing, the encoder can simultaneously learn the common information of the compensation features and the coating geometric identification features, avoiding information redundancy during feature extraction. Specifically, the encoder's input layer concatenates the compensation feature set and the coating geometric identification features, and then progressively extracts higher-order features through multiple hidden layers. Finally, the output of the encoder's last layer generates a preliminary feature vector, which contains comprehensive information from the compensation features and the coating geometric identification features, laying the foundation for subsequent feature fusion.

[0104] 202. Based on the attention mechanism, the preliminary feature vector and the semantic vector of the text symbol are weighted and fused to generate a fused feature vector;

[0105] In step 202, the attention mechanism refers to a technique that achieves feature fusion by assigning different weights to different features, dynamically adjusting the weights according to the importance of the features. The text symbol semantic vector refers to textual descriptive features related to the coating generated through natural language processing techniques, such as "uneven coating thickness" or "excessively high surface roughness." The fused feature vector refers to the feature representation generated after weighted fusion of the initial feature vector and the text symbol semantic vector through the attention mechanism, used for subsequent multi-scale decomposition.

[0106] In this embodiment, the initial feature vector and the semantic vector of the text symbols are first input into the attention mechanism module. The attention mechanism calculates the similarity matrix between the two to obtain the weight coefficient of each feature. Specifically, a self-attention mechanism is used to interactively calculate the initial feature vector and the semantic vector of the text symbols to generate an attention weight matrix. Then, the initial feature vector and the semantic vector of the text symbols are weighted and fused according to the weight coefficients to generate a fused feature vector. This vector not only includes compensation features and coating geometric identification features, but also incorporates textual description information, enhancing the semantic expressiveness of the features. For example, when the textual description mentions "uneven coating thickness," the attention mechanism will assign higher weights to thickness-related features, thereby improving the targeting of feature fusion.

[0107] 203. By performing multi-scale decomposition on the fused feature vector, the internal support structure features and the external anti-corrosion coating texture features of the tower body are extracted respectively, wherein the support structure features are extracted by high-frequency components and the coating texture features are extracted by low-frequency components.

[0108] In step 203, multi-scale decomposition refers to a technique that uses wavelet transform or convolutional neural networks to perform multi-scale analysis on feature vectors, enabling the simultaneous capture of both high-frequency and low-frequency information of the features. The internal support structure features of the tower body refer to features related to the support structure extracted through high-frequency components, such as the geometry, connection method, and stress distribution of the support rods. The external anti-corrosion coating texture features refer to features related to the surface texture of the coating extracted through low-frequency components, such as the particle distribution, color uniformity, and surface smoothness of the coating.

[0109] In this embodiment, the fused feature vector is first decomposed into high-frequency and low-frequency components using wavelet transform or convolutional neural network. Specifically, discrete wavelet transform is used to decompose the fused feature vector, yielding high-frequency and low-frequency components. The high-frequency components typically contain detailed information about the supporting structure; therefore, the internal supporting structure features of the tower are extracted using the high-frequency components. The low-frequency components contain overall texture information of the coating; therefore, the texture features of the external anti-corrosion coating are extracted using the low-frequency components. For example, in the detection scenario of a wind turbine tower, the high-frequency components can capture minute cracks or deformations in the support rods, while the low-frequency components can reflect the particle distribution and color uniformity of the coating surface. Finally, the supporting structure features and coating texture features are obtained, providing a foundation for subsequent optimization and constraints.

[0110] 204. Based on the local consistency between the internal support structure features and the external anti-corrosion coating texture features, optimize the alternating update of the convolution kernel parameters and establish the geometric constraint relationship between the support structure features and the coating texture features.

[0111] In step 204, the internal support structure features of the tower body refer to the features related to the support structure extracted through high-frequency components, such as the geometry, connection method, and stress distribution of the support rods. The external anti-corrosion coating texture features refer to the features related to the surface texture of the coating extracted through low-frequency components, such as the particle distribution, color uniformity, and surface smoothness of the coating. Local consistency optimization is a technique that achieves feature smoothing by constraining parameter changes in adjacent regions, which can improve the continuity and consistency of features. Alternating update of convolution kernel parameters refers to the process of improving feature extraction performance by iteratively optimizing convolution kernel parameters. Geometric constraint relationships refer to establishing the spatial relationship between the support structure features and the coating texture features through mathematical modeling, ensuring their geometric compatibility.

[0112] In this embodiment, the convolution kernel parameters are first optimized based on the local consistency of the internal support structure features and the external anti-corrosion coating texture features. Specifically, the alternating direction multiplier method is used to iteratively optimize the convolution kernel parameters, ensuring that the extracted features remain consistent within local regions. For example, in the detection scenario of wind turbine towers, optimizing the convolution kernel parameters ensures that the geometric features of the support rod and the coating texture features are spatially consistent. Then, a geometric constraint relationship is established between the support structure features and the coating texture features to ensure their spatial coordination. Specifically, mathematical modeling methods are used to establish the geometric relationship between the two, such as the relationship between the geometry of the support rod and the distribution of the coating thickness. Finally, optimized support structure features and coating texture features are obtained, providing a high-precision feature representation for coating detection.

[0113] Here is a specific example:

[0114] In the coating inspection scenario of wind turbine towers, the system first collects a set of compensation features of the tower coating, such as thickness, roughness, and adhesion, using sensors. Then, image processing techniques are used to extract geometric features of the coating, such as edges, contours, and surface curvature. Both are input into a shared-weight encoder to generate a preliminary feature vector. Next, an attention mechanism is used to weightedly fuse the preliminary feature vector with semantic vectors such as "uniform coating thickness" and "excessively high surface roughness," generating a fused feature vector. Then, multi-scale decomposition is used to extract features of the internal support structure of the tower, such as the shape and connection method of the support rods, and texture features of the external anti-corrosion coating, such as the particle distribution and color uniformity of the coating. Finally, local consistency optimization is used to alternately update the convolution kernel parameters, and a geometric constraint relationship is established between the support structure features and the coating texture features to ensure the accuracy and consistency of the inspection results.

[0115] In summary, steps 201 to 204 achieve efficient fusion of compensation features, coating geometric identification features, and textual semantic features, enhancing the expressive power of the features. Through multi-scale decomposition and local consistency optimization, accurate extraction of support structure features and coating texture features is achieved, and a geometric constraint relationship between the two is established, improving the accuracy and robustness of coating detection. This method can not only effectively identify geometric deviations and texture defects in the coating but also ensure the spatial coordination between the support structure and the coating, providing reliable technical support for coating detection in complex industrial scenarios.

[0116] To address the difficulty in effectively fusing the initial feature vector and the semantic vector of the text symbol, in some embodiments, step 202 involves weighted fusing of the initial feature vector and the semantic vector of the text symbol based on an attention mechanism to generate a fused feature vector, including:

[0117] 301. Input the preliminary feature vector and the semantic vector of the text symbol into independent fully connected layers to align the feature dimensions, and generate the aligned preliminary feature vector and semantic vector of the text symbol;

[0118] In step 301, the preliminary feature vector refers to the fused feature representation extracted by the shared weight encoder, which contains multi-level visual information of the power tower surface image, such as texture, edges, and shape. The text symbol semantic vector refers to the semantic features related to the text symbols on the power tower surface generated by natural language processing technology, which captures the semantic meaning and contextual relationships of the text symbols. The fully connected layer is a neural network layer used to map input features to the target dimension, which adjusts the feature dimension through linear transformation and non-linear activation functions. Feature dimension alignment refers to adjusting the dimensions of the preliminary feature vector and the text symbol semantic vector to the same size through the fully connected layer, so as to facilitate subsequent feature fusion and correlation analysis.

[0119] In this embodiment, the preliminary feature vector and the semantic vector of the text symbols are first input into separate fully connected layers for feature dimension alignment. Specifically, the output dimension of the fully connected layer is set according to the larger of the two dimensions to ensure the integrity and consistency of the feature information. Secondly, through linear transformation and nonlinear activation functions of the fully connected layer, the preliminary feature vector and the semantic vector of the text symbols are mapped to the same dimensional space. Finally, the aligned preliminary feature vector and semantic vector of the text symbols are generated, providing a foundation for subsequent feature association analysis. This step ensures the dimensional consistency between the preliminary feature vector and the semantic vector of the text symbols, laying the technical foundation for subsequent feature fusion and association analysis.

[0120] 302. Based on the spatial distribution of the aligned preliminary feature vector and the semantic vector of the text symbol, construct local symbol association features and global structure association features, wherein the local symbol association features are generated by the mean of features in the symbol neighborhood, and the global structure association features are generated by cross-region feature difference.

[0121] In step 302, local symbol association features refer to local association information generated through the mean of features within the symbol's neighborhood. This captures the local relationship between the text symbol and its neighboring regions, reflecting the spatial distribution and feature variations of the text symbol within a local area. Global structural association features refer to global association information generated through cross-regional feature difference. This captures the global relationship between the text symbol and the overall structure, reflecting the position and feature differences of the text symbol within the overall structure. Spatial distribution refers to the distribution of the aligned preliminary feature vector and the text symbol semantic vector in the feature space. It reflects the spatial relationship between the text symbol and the preliminary features through the numerical distribution of the feature vectors.

[0122] In this embodiment, firstly, based on the spatial distribution of the aligned preliminary feature vectors and the semantic vectors of the text symbols, local symbol association features and global structural association features are constructed. Secondly, local symbol association features are generated by calculating the mean of features within the neighborhood of the symbol. Specifically, with each text symbol as the center, the mean of features in the surrounding area is calculated, thereby capturing the local relationship between the text symbol and its neighboring area. Global structural association features are generated through cross-regional feature difference. Specifically, the feature difference between the region where the text symbol is located and the overall structural region is calculated, thereby capturing the global relationship between the text symbol and the overall structure. Finally, local symbol association features and global structural association features are generated to provide input for subsequent cross-projection. This step, through the construction of local and global features, comprehensively reflects the spatial relationship between the text symbols and the preliminary features, providing rich information support for subsequent feature fusion.

[0123] 303. By cross-projecting the local symbol association features and the global structure association features, a symbol structure association mapping vector is generated, and the dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector is calculated based on the symbol structure association mapping vector.

[0124] In step 303, cross-projection refers to a technique that maps local symbolic association features and global structural association features to the same space through matrix operations. It achieves feature fusion through linear transformations and nonlinear activation functions. The symbolic structure association mapping vector is the fused feature representation generated through cross-projection, which contains comprehensive information from both local symbolic association features and global structural association features. The dynamic relevance score is a score obtained by calculating the similarity between the symbolic structure association mapping vector and the aligned initial feature vector and text symbol semantic vector. It measures the relevance between features and reflects the dynamic association relationship between the initial feature vector and the text symbol semantic vector.

[0125] In this embodiment, firstly, a cross-projection is performed on local symbol association features and global structural association features to generate a symbol-structure association mapping vector. Specifically, matrix multiplication is used to map both to the same space, and normalization is employed to ensure numerical stability. Secondly, a dynamic relevance score is calculated between the aligned preliminary feature vector and the semantic vector of the text symbols based on the symbol-structure association mapping vector. Specifically, cosine similarity is used to calculate the relevance between both features and the symbol-structure association mapping vector to obtain the dynamic relevance score. Finally, the symbol-structure association mapping vector and the dynamic relevance score are generated to provide a basis for subsequent weight allocation. This step, through cross-projection and the calculation of dynamic relevance scores, achieves a comprehensive analysis of local and global features, providing a scientific basis for subsequent weight allocation.

[0126] 304. Based on the dynamic relevance score, channel-level weight allocation is performed on the aligned preliminary feature vector, and symbol-level weight allocation is performed on the text symbol semantic vector to generate channel-weighted preliminary feature vector and symbol-weighted semantic vector;

[0127] In step 304, channel-level weight allocation refers to the process of assigning different weights to each channel of the initial feature vector based on the dynamic relevance score. This weight adjustment highlights feature channels that are highly correlated with the semantic vector of the text symbols. Symbol-level weight allocation refers to the process of assigning different weights to each symbol of the semantic vector of the text symbols based on the dynamic relevance score. This weight adjustment highlights text symbols that are highly correlated with the initial feature vector. The channel-weighted initial feature vector and the symbol-weighted semantic vector refer to the feature representations after weight allocation, respectively. The weight adjustment optimizes the expression effect of the features.

[0128] In this embodiment, firstly, channel-level weights are assigned to the aligned preliminary feature vectors based on the dynamic relevance score. Specifically, the dynamic relevance score is converted into weight values, and each channel of the preliminary feature vector is weighted. Secondly, symbol-level weights are assigned to the semantic vectors of text symbols based on the dynamic relevance score; specifically, different weights are assigned to the semantic features of each text symbol. Finally, channel-weighted preliminary feature vectors and symbol-weighted semantic vectors are generated, providing input for subsequent multi-scale context feature concatenation. This step optimizes the feature representation through weight allocation, providing high-quality input for subsequent feature fusion.

[0129] 305. The channel weighted preliminary feature vector and the symbol weighted semantic vector are concatenated with multi-scale context features to generate a multi-scale context feature concatenation result. The multi-scale context features are extracted by sliding the context window of symbol semantics. The window size is positively correlated with the density of text symbols on the surface of the power tower.

[0130] In step 305, multi-scale context feature concatenation refers to the process of fusing channel-weighted preliminary feature vectors and symbol-weighted semantic vectors at different scales. It captures the multi-level relationship between text symbols and preliminary features through multi-scale analysis. Context window sliding refers to extracting the context features of symbol semantics through a sliding window. It achieves multi-scale feature extraction through window sliding. The multi-scale context feature concatenation result refers to the multi-scale fused feature representation generated by concatenation, which contains multi-level comprehensive information of text symbols and preliminary features.

[0131] In this embodiment, the channel-weighted preliminary feature vector and the symbol-weighted semantic vector are first concatenated using multi-scale contextual features. Specifically, contextual features of symbolic semantics are extracted by sliding a context window, with the window size dynamically adjusted according to the density of text symbols on the power tower surface. Next, the extracted multi-scale contextual features are concatenated to generate a multi-scale contextual feature concatenation result. Finally, this provides input for subsequent bidirectional feature propagation. This step, through multi-scale contextual feature concatenation, achieves multi-level fusion of text symbols and preliminary features, providing rich feature representations for subsequent feature propagation.

[0132] 306. Based on the topological relationship constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale context feature splicing result to generate a fused feature vector.

[0133] In step 306, the topological relationship constraint of the power tower skeleton refers to constraining feature propagation based on the skeleton structure of the power tower, which ensures the rationality and consistency of feature propagation through topological relationships; bidirectional feature propagation refers to the process of propagating the multi-scale context feature concatenation result under the skeleton topological relationship constraint through a bidirectional graph neural network, which achieves global optimization of features through bidirectional propagation; the fused feature vector refers to the final feature representation generated through bidirectional feature propagation, which contains comprehensive information of text symbol semantics and preliminary feature vectors.

[0134] In this embodiment, firstly, based on the topological constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale contextual feature stitching results. Specifically, a bidirectional graph neural network is used to propagate the multi-scale contextual features under the skeleton topological constraints. Secondly, the skeleton structure of the power tower is modeled as a graph, where nodes represent key points of the skeleton and edges represent skeleton connections. Finally, the multi-scale contextual features are bidirectionally propagated in the graph through the bidirectional graph neural network to generate a fused feature vector. This step optimizes the global representation of features through bidirectional feature propagation, providing a high-precision feature representation for the recognition of text symbols on the surface of power towers.

[0135] Here is a specific example:

[0136] In the inspection of power transmission towers, high-resolution cameras are first used to capture images of the tower surface. A shared-weight encoder is then used to extract preliminary feature vectors. Simultaneously, a BERT model is used to perform semantic analysis on the textual symbols on the tower surface (such as the number "T-102" and the warning "High Voltage Danger"), generating semantic vectors for these symbols. Both are then input into independent fully connected layers for feature dimension alignment to ensure consistency. Next, using the textual symbol "T-102" as the center, the mean feature value of a 3×3 region surrounding it is calculated to generate local symbol association features. The feature difference between the region containing the symbol and the overall tower structure is then calculated to generate global structure association features. Finally, the local and global features are cross-projected to generate a symbol structure association mapping vector. The dynamic correlation score between this vector and the preliminary feature vector and the textual symbol semantic vector is calculated; for example, the correlation score between "T-102" and the preliminary feature vector is 0.82. Based on the correlation scores, channel-level weights are assigned to the preliminary feature vector, and symbol-level weights are assigned to the textual symbol semantic vector, highlighting important features. Subsequently, the weighted feature vectors are concatenated using multi-scale contextual features: features are extracted using a 5×5 window in the numbered region and a 3×3 window in the warning text region, generating a multi-scale contextual feature concatenation result. Finally, based on the tower's skeletal topology (such as crossarms, tower body, and tower feet), a bidirectional graph neural network is used to propagate the multi-scale features bidirectionally, generating a fused feature vector.

[0137] In summary, steps 301 to 306 achieved efficient fusion of the preliminary feature vector and the semantic vector of the text symbol; by constructing local symbol association features and global structure association features, the relationship between the text symbol and its neighboring region and the overall structure was captured; by multi-scale context feature splicing and bidirectional feature propagation, the feature extraction and propagation process was optimized, and finally, a high-precision fused feature vector was generated, providing reliable technical support for the recognition of text symbols on the surface of power transmission towers.

[0138] To address the challenges of efficiently fusing local symbol association features and global structural association features, and the difficulty in quantifying dynamic correlation, some embodiments include step 303, which involves cross-projecting the local symbol association features and global structural association features to generate a symbol structure association mapping vector, and calculating the dynamic correlation score between the aligned preliminary feature vector and the semantic vector of the text symbols based on the symbol structure association mapping vector.

[0139] 401. The local symbol association features are spatially encoded and channel-dimensional encoded respectively to generate spatially encoded local symbol association features and channel-encoded global structural association features;

[0140] In step 401, local symbol-related features refer to the local features related to symbols extracted from the image. Spatial location encoding is used to capture the spatial distribution information of symbols in the image, and channel dimension encoding is used to capture the global structural information of symbols in the feature channels. Spatially encoded local symbol-related features refer to the local features after spatial location encoding, and channel encoded global structural-related features refer to the global features after channel dimension encoding.

[0141] In this embodiment, the local symbol association features are first spatially encoded by using a position encoder (such as sine coding) to embed the spatial coordinate information of the symbols into the features, generating spatially encoded local symbol association features. Next, the local symbol association features are channel-dimensionally encoded by using fully connected layers or convolutional layers to globally model the feature channels, generating channel-encoded global structural association features. Finally, spatially encoded local symbol association features and channel-encoded global structural association features are obtained respectively.

[0142] 402. By interactively superimposing the feature channels, the spatially encoded local symbol association features and the channel-encoded global structure association features are bidirectionally projected to generate a symbol structure association mapping vector.

[0143] In step 402, the interactive overlay of feature channels refers to fusing the spatially encoded local symbolic association features with the channel-encoded global structural association features along the channel dimension. Bidirectional projection refers to mapping the two types of features to the same space using a projection matrix, generating a symbolic structure association mapping vector. The symbolic structure association mapping vector is a feature vector containing both local semantic and global structural information of the symbols.

[0144] In this embodiment, the spatially encoded local symbol association features and the channel-encoded global structural association features are first concatenated or weighted and summed along the channel dimension to achieve interactive overlay of feature channels. Next, a bidirectional projection matrix, such as a learnable fully connected layer, maps the two types of features to the same space, generating a symbol structure association mapping vector. Finally, a symbol structure association mapping vector containing both local semantic and global structural information is obtained.

[0145] 403. Extract the spatial weight matrix and the channel weight matrix from the symbol structure association mapping vector, wherein the spatial weight matrix is ​​generated based on the local similarity of symbol semantics, and the channel weight matrix is ​​generated based on the global difference of structural features;

[0146] In step 403, the spatial weight matrix is ​​used to quantify the local similarity of symbols in space, and the channel weight matrix is ​​used to quantify the differences of feature channels in the global structure. The local similarity of symbol semantics refers to the degree of semantic association between symbols in spatial location, while the global difference of structural features refers to the distribution differences of feature channels in the global structure.

[0147] In this embodiment, firstly, a spatial weight matrix is ​​extracted from the symbol structure association mapping vector. A convolutional layer or self-attention mechanism is then used to calculate the local similarity of symbols in space, generating the spatial weight matrix. Next, a channel weight matrix is ​​extracted from the symbol structure association mapping vector. Global average pooling or a fully connected layer is used to calculate the differences in the global structure of feature channels, generating the channel weight matrix. Finally, the spatial weight matrix and the channel weight matrix are obtained respectively.

[0148] 404. By parametrically weighting the spatial topological constraints of the global structural association features, the spatial weight matrix and the channel weight matrix are fused to generate dynamic correlation weight coefficients;

[0149] In step 404, spatial topological constraints refer to the spatial distribution relationship of global structural association features, typically implemented through Graph Convolutional Networks (GCNs) or Transformers. Parameterized weighting refers to dynamically adjusting the spatial weight matrix and channel weight matrix using learnable parameters to generate dynamic correlation weight coefficients. Learnable parameters are typically implemented through fully connected layers or convolutional layers. Dynamic correlation weight coefficients are weight parameters used to adjust the strength of cross-modal feature alignment. A fully connected layer is a network layer that maps input features to output features, while a convolutional layer is a network layer used to extract local features. Dynamic correlation weight coefficients are weight parameters used to adjust the strength of cross-modal feature alignment.

[0150] In this embodiment, spatial topological constraints are first modeled for the global structural correlation features, and a graph convolutional network or Transformer is used to capture the spatial distribution relationship of the features. Next, the spatial weight matrix and channel weight matrix are weighted and fused using learnable parameters (such as fully connected layers) to generate dynamic correlation weight coefficients. Finally, weight coefficients that can dynamically adjust cross-modal feature alignment are obtained.

[0151] 405. Based on the dynamic relevance weight coefficient, the cross-modal association strength between the aligned preliminary feature vector and the semantic vector of the text symbol is quantified to generate a dynamic relevance score.

[0152] In step 405, the aligned preliminary feature vector refers to the image features that have undergone preliminary alignment processing, typically achieved through feature mapping or feature transformation. Cross-modal correlation strength refers to the correlation between image features and text features, usually calculated using dot product or cosine similarity. Dynamic correlation score is used to quantify the degree of alignment of cross-modal features, typically achieved through weighted summation or normalization.

[0153] In this embodiment, the aligned preliminary feature vector is first compared with the semantic vector of the text symbol by a dot product or cosine similarity calculation to obtain a preliminary cross-modal association strength. Then, the preliminary association strength is weighted and adjusted using a dynamic relevance weight coefficient to generate a dynamic relevance score. Finally, a dynamic relevance score that accurately reflects the degree of cross-modal feature alignment is obtained.

[0154] Here is a specific example:

[0155] Suppose that the text symbols "A", "B", and "C" are obtained from the surface of a power transmission tower, located at different positions on the tower. Spatial encoding vectors and channel encoding vectors are generated through spatial location encoding and channel dimension encoding. Symbol structure association mapping vectors are generated through the interactive superposition and bidirectional projection of feature channels. Spatial weight matrices and channel weight matrices are extracted from the symbol structure association mapping vectors, and dynamic correlation weight coefficients are generated through parameterized weighting. Finally, based on the dynamic correlation weight coefficients, the cross-modal correlation strength between the initial feature vectors and the semantic vectors of the text symbols is quantified to generate a dynamic correlation score, which is used for the identification and reconstruction of power transmission towers. This embodiment effectively improves the accuracy and anti-interference capability of power transmission tower identification, providing strong support for the maintenance and management of power facilities.

[0156] In summary, steps 401 to 405 achieve high precision, robustness, and anti-interference capabilities for the 3D recognition of power transmission towers. This method can accurately align text symbols with tower component features, efficiently suppress multipath reflection signals and environmental noise, deeply fuse semantic vectors of text symbols with tower support structures and anti-corrosion coating texture features, accurately reconstruct missing tower connection components, and achieve high-precision 3D recognition and positioning of power transmission towers, providing reliable technical support for the intelligent detection and maintenance of power infrastructure.

[0157] To address the challenges of collaborative modeling of internal support structure features and external anti-corrosion coating texture features, as well as the difficulty in establishing geometric constraints, in some embodiments, step 204, which involves optimizing and alternately updating the convolution kernel parameters based on the local consistency of the internal support structure features and the external anti-corrosion coating texture features, and establishing geometric constraints between the support structure features and the coating texture features, includes:

[0158] 501. The internal support structure features are decomposed into the spatial distribution features of the connectors and the morphological features of the main beam, and the external anti-corrosion coating texture features are decomposed into the corrosion area distribution features and the coating thickness variation features.

[0159] In step 501, the internal support structure features refer to the support structure information extracted from the inside of the power transmission tower. The spatial distribution features of the connectors refer to the spatial distribution of the connectors within the support structure. The main beam morphology features refer to the geometric shape and curvature variations of the main beam. The external anti-corrosion coating texture features refer to the coating texture information extracted from the outside of the power transmission tower. The corrosion area distribution features refer to the distribution of corrosion areas within the coating. The coating thickness variation features refer to the gradient variation of the coating thickness.

[0160] In this embodiment, the internal support structure features are first decomposed into the spatial distribution features of the connectors and the morphological features of the main beam. Spatial segmentation and morphological analysis techniques are used to extract the positional information of the connectors and the geometric shape of the main beam, respectively. Next, the texture features of the external anti-corrosion coating are decomposed into the distribution features of the corrosion area and the coating thickness variation features. Image segmentation and thickness measurement techniques are used to extract the distribution information of the corrosion area and the gradient change of the coating thickness, respectively. Finally, the spatial distribution features of the connectors, the morphological features of the main beam, the distribution features of the corrosion area, and the coating thickness variation features are obtained, respectively.

[0161] 502. Based on the proximity constraint between bolt hole locations and coating defect areas, spatial topological encoding is performed on the spatial distribution characteristics of the connectors and the distribution characteristics of corrosion areas to generate connector corrosion correlation characteristics. Based on the synchronization constraint between tower curvature changes and coating thickness gradients, morphological-thickness coupling encoding is performed on the main beam morphological characteristics and coating thickness change characteristics to generate main beam thickness coupling characteristics.

[0162] In step 502, the proximity constraint between bolt hole locations and coating defect areas refers to the spatial proximity relationship between these locations. Spatial topology encoding refers to modeling the spatial relationships of features based on proximity constraints. The synchronicity constraint between tower curvature changes and coating thickness gradients refers to the morphological synchronicity between these two changes. Morphological-thickness coupling encoding refers to modeling the morphological relationships of features based on synchronicity constraints. Connector corrosion association features refer to the spatial association information between the spatial distribution features of connectors and the distribution features of corrosion areas. Main beam thickness coupling features refer to the morphological association information between the main beam morphological features and the coating thickness variation features.

[0163] In this embodiment, firstly, based on the proximity constraint between bolt hole locations and coating defect areas, spatial topological encoding is performed on the spatial distribution characteristics of connectors and the distribution characteristics of corrosion areas. A graph convolutional network or spatial attention mechanism is used to model the spatial correlation, generating connector corrosion correlation features. Next, based on the synchronization constraint between tower curvature changes and coating thickness gradients, morphological-thickness coupling encoding is performed on the main beam morphological characteristics and coating thickness variation characteristics. A morphological analysis network or synchronization modeling technique is used to model the morphological correlation, generating main beam thickness coupling features. Finally, connector corrosion correlation features and main beam thickness coupling features are obtained respectively.

[0164] 503. Construct a multi-branch convolutional kernel group, wherein the first branch performs spatial constraint modeling on the corrosion correlation features of the connector, and the second branch performs texture constraint modeling on the thickness coupling features of the main beam;

[0165] In step 503, the multi-branch convolutional kernel group refers to a set of convolutional kernels composed of multiple branches. The first branch is used to perform spatial constraint modeling of the corrosion-related features of the connectors, and the second branch is used to perform texture constraint modeling of the thickness coupling features of the main beam. Spatial constraint modeling refers to modeling the spatial relationship of the corrosion-related features of the connectors, while texture constraint modeling refers to modeling the texture relationship of the thickness coupling features of the main beam.

[0166] In this embodiment, a multi-branch convolutional kernel group is first constructed. The first branch uses convolutional layers or graph convolutional networks to perform spatial constraint modeling on the corrosion-related features of the connector, capturing the spatial relationship between the connector and the corrosion area. The second branch uses convolutional layers or texture analysis networks to perform texture constraint modeling on the thickness coupling features of the main beam, capturing the texture relationship between the main beam and the coating thickness. Finally, the spatial constraint modeling results and texture constraint modeling results are obtained respectively.

[0167] 504. Based on the local consistency loss of the spatial constraint modeling result and the texture constraint modeling result, the parameters of the multi-branch convolution kernel group are alternately updated to generate the geometric constraint relationship between the supporting structure features and the coating texture features.

[0168] In step 504, local consistency loss refers to the consistency difference between spatial constraint modeling results and texture constraint modeling results in local regions, and alternating update refers to the alternating adjustment of parameters of multi-branch convolutional kernel groups through iterative optimization methods. Geometric constraint relationship refers to the spatial and morphological correlation between supporting structural features and coating texture features.

[0169] In this embodiment, firstly, based on the local consistency loss between the spatial constraint modeling results and the texture constraint modeling results, a loss function (such as mean squared error loss or cross-entropy loss) is used to calculate the consistency difference. Next, the parameters of the multi-branch convolutional kernel group are iteratively optimized using an alternating update method (such as gradient descent or backpropagation) to generate the geometric constraint relationship between the support structure features and the coating texture features. Finally, a geometric constraint relationship that accurately reflects the correlation between the support structure and the coating texture is obtained.

[0170] Here is a specific example:

[0171] In power transmission tower identification, taking a specific tower as an example, radar scanning revealed that its internal connectors were concentrated at the four corners, the main beam had a grid structure, and external coating corrosion was mostly observed at the bottom, with the thickness gradually decreasing from bottom to top. Based on the proximity of bolt hole locations to corrosion areas, corrosion correlation features of connectors were generated; for example, corrosion areas were mostly distributed around bolt hole locations. Simultaneously, utilizing the synchronicity between tower curvature and coating thickness, coupling features of main beam thickness were generated; for example, the coating thickness changed significantly at tower corners. Subsequently, multi-branch convolutional kernel groups were used to spatially model the corrosion correlation features of connectors, extracting corrosion distribution patterns, and to texture-model the coupling features of main beam thickness, capturing thickness variations. By optimizing the consistency of spatial and texture modeling, geometric constraints between the support structure and coating texture were generated; for example, ensuring that connector corrosion and main beam thickness features were aligned spatially and texturally. Finally, based on this relationship, missing tower connection components were reconstructed, such as repairing corrosion defects, generating an interference-resistant 3D identification model, providing accurate support for facility maintenance.

[0172] In summary, steps 501 to 504 decompose the internal support structure and external coating texture features, encode them based on spatial proximity and morphological synchronicity constraints, construct multi-branch convolutional kernel groups for spatial and texture modeling, and alternately update parameters through local consistency loss, ultimately generating a precise geometric constraint relationship between the support structure and coating texture, thereby efficiently associating and analyzing the features of the tower's internal support and external coating.

[0173] To address the challenges of accurately analyzing the original reflected signals of power transmission towers and extracting the location information of text symbols, some embodiments involve, in step 103, acquiring the original reflected signals of the power transmission tower using a radar array, filtering target reflection areas based on the multimodal mapping relationship, analyzing the phase difference data in the original reflected signals, and generating a set of penetrating features containing the tower's skeleton outline and corresponding text symbol location information, including:

[0174] 601. Multi-angle signal acquisition of power transmission towers using a ring-shaped or linearly deployed radar array to obtain the original set of reflected signals, wherein the direction of the radar beam is aligned with the spatial distribution of text symbols on the tower surface during signal acquisition.

[0175] In step 601, a radar array refers to a detection system composed of multiple radar units, typically employing the principle of transmitting and receiving electromagnetic waves for target detection. Ring or linear deployment refers to the arrangement of radar units in a ring or linear pattern; ring deployment is suitable for omnidirectional detection, while linear deployment is suitable for detection in a specific direction. Multi-angle signal acquisition refers to detecting power transmission towers with radar signals from different angles to obtain more comprehensive target information. The original set of reflected signals refers to the set of reflected signal data acquired by the radar array, containing information such as the reflection intensity and phase of the target object. The radar beam direction refers to the direction of radar signal transmission and reception, typically controlled by adjusting the pointing of the radar antenna. The spatial distribution of text symbols on the tower surface refers to the position and arrangement of text symbols on the tower surface, typically used to identify tower components or provide other information.

[0176] In this embodiment, a multi-angle signal acquisition is first performed on the power transmission tower using a ring-shaped or linearly deployed radar array. This ensures that the radar beam direction is aligned with the spatial distribution of the text symbols on the tower surface, thereby obtaining a raw set of reflected signals containing the text symbol information. The deployment method of the radar array is selected based on actual detection needs; a ring-shaped deployment is suitable for omnidirectional detection, while a linear deployment is suitable for detection in a specific direction. During signal acquisition, the radar beam direction is aligned with the spatial distribution of the text symbols on the tower surface to ensure that the text symbol information can be effectively captured. Finally, a raw set of reflected signals containing multi-angle detection data is obtained.

[0177] 602. Based on the multimodal mapping relationship, beamforming processing is performed on the original set of reflected signals to filter target reflection areas containing text symbol positioning information, wherein the beamforming parameters are dynamically adjusted according to the size of the text symbol and the geometric dimensions of the tower components;

[0178] In step 602, the multimodal mapping relationship refers to the association between text symbols and tower components, typically achieved through a predefined mapping table or machine learning model. Beamforming processing refers to focusing and enhancing the original reflected signal to improve the signal-to-noise ratio and target resolution. The target reflection region refers to the signal region containing text symbol positioning information, usually obtained through signal processing algorithms. Beamforming parameters are parameters used to control beamforming processing, such as beamwidth, direction, and gain, typically dynamically adjusted according to the characteristics of the detected target. Text symbol size refers to the size of the text symbol, usually measured using image processing techniques. Tower component geometry refers to the shape and size of the tower components, typically obtained through 3D modeling or measurement techniques.

[0179] In this embodiment, beamforming processing is first performed on the original set of reflected signals based on multimodal mapping relationships. Beamforming algorithms (such as delay summation algorithms or adaptive beamforming algorithms) are used to filter target reflection areas containing text symbol positioning information. Beamforming parameters are dynamically adjusted according to the size of the text symbols and the geometric dimensions of the tower components to ensure signal focusing accuracy. The delay summation algorithm achieves beam focusing by adjusting signal delay, while the adaptive beamforming algorithm optimizes the beam through adaptive filtering. During beamforming processing, multimodal mapping relationships guide the selection of target reflection areas, ensuring that text symbol positioning information can be effectively extracted. Finally, target reflection areas containing text symbol positioning information are obtained.

[0180] 603. Perform phase difference interferometry processing on the original reflection signal of the target reflection area, analyze the phase difference data between adjacent radar channels, and generate an initial three-dimensional point cloud of the tower skeleton outline.

[0181] In step 603, phase difference interferometry processing refers to processing the phase difference data between adjacent radar channels to extract the geometric information of the target object. Phase difference data refers to the phase difference between signals received by adjacent radar channels, which is usually related to the distance and shape of the target object. Analyzing the phase difference data refers to extracting the geometric information of the target object from the phase difference, typically achieved through interferometric measurement algorithms. The initial 3D point cloud refers to the 3D point cloud data of the tower skeleton outline generated through phase difference interferometry processing, containing the spatial coordinate information of the target object.

[0182] In this embodiment, the original reflected signal of the target reflection area is first subjected to phase difference interferometry processing. A phase difference interferometry algorithm (such as synthetic aperture radar interferometry) is used to analyze the phase difference data between adjacent radar channels to generate an initial 3D point cloud of the tower skeleton outline. The synthetic aperture radar interferometry algorithm improves the resolution and geometric accuracy of the target object by synthesizing signals from multiple radar channels. During the phase difference interferometry processing, the phase difference data between adjacent radar channels is used to calculate the distance and shape of the target object, generating the initial 3D point cloud. Finally, an initial 3D point cloud containing the geometric information of the tower skeleton is obtained.

[0183] 604. Based on the component correlation of text symbols in multimodal mapping, spatial filtering is performed on the initial three-dimensional point cloud to remove scattering points that are irrelevant to the positioning of text symbols, and a symbol-constrained tower skeleton point cloud is generated.

[0184] In step 604, the component correlation of text symbols in the multimodal mapping relationship refers to the correspondence between text symbols and tower components, which is usually achieved through a predefined mapping table or machine learning model. Spatial filtering refers to the screening and culling operation on the initial 3D point cloud to remove noise and irrelevant data. Scattering points refer to the point cloud data generated by radar signal reflection, which usually contains information about the target object and noise. The symbol-constrained tower skeleton point cloud refers to the point cloud data related to the localization of text symbols retained after spatial filtering, which is usually used for subsequent 3D reconstruction and analysis.

[0185] In this embodiment, the initial 3D point cloud is first spatially filtered based on the component correlation of text symbols in the multimodal mapping relationship. Spatial filtering algorithms (such as distance-based filtering or density-based filtering) are used to remove scattering points unrelated to the text symbol positioning, generating a symbol-constrained tower skeleton point cloud. Distance-based filtering removes noise points by setting a distance threshold, while density-based filtering removes sparse points by setting a density threshold. During spatial filtering, the component correlation of text symbols in the multimodal mapping relationship guides the point cloud selection, ensuring that point cloud data related to text symbol positioning is preserved. Finally, a symbol-constrained tower skeleton point cloud containing text symbol positioning information is obtained.

[0186] 605. The point cloud of the tower skeleton constrained by the symbol is fused with the multipath reflection intensity in the previous phase difference interferometry processing to generate a set of penetrating features containing the positioning information of the text symbols and the geometric parameters of the tower skeleton.

[0187] In step 605, multipath reflection intensity refers to the intensity information of radar signals reflected multiple times during propagation, which is usually related to the material and shape of the target object. Feature fusion refers to combining the symbol-constrained tower skeleton point cloud with multipath reflection intensity to generate more comprehensive feature data. The penetration feature set refers to the feature data set containing text symbol positioning information and tower skeleton geometric parameters, which is usually used for subsequent 3D reconstruction and analysis.

[0188] In this embodiment, the symbol-constrained tower skeleton point cloud is first fused with the multipath reflection intensity from the previous phase difference interferometry processing. A feature fusion algorithm (such as weighted fusion or convolutional fusion) is then used to generate a penetration feature set containing the location information of the text symbols and the geometric parameters of the tower skeleton. Weighted fusion combines different features by setting weight coefficients, while convolutional fusion fuses different features through convolution operations. During the feature fusion process, the symbol-constrained tower skeleton point cloud and multipath reflection intensity are used to generate more comprehensive feature data, ensuring the integrity and accuracy of the penetration feature set. Finally, a penetration feature set for the 3D reconstruction of power transmission towers is obtained.

[0189] Here is a specific example:

[0190] In power tower identification, the process begins by acquiring multi-angle signals from the tower using a ring or linear radar array. This ensures the radar beam aligns with text symbols (such as "A" and "B") on the tower surface, obtaining the original reflected signals. Based on the multimodal mapping relationship between the text symbols and tower components, beamforming parameters are dynamically adjusted to filter target reflection areas containing text symbol positioning information. Next, phase difference interferometry is applied to the target area signals, analyzing the phase difference data of adjacent radar channels to generate an initial 3D point cloud of the tower skeleton. Based on the component correlation of the text symbols, spatial filtering is performed on the initial point cloud to remove irrelevant scattering points, generating a symbol-constrained tower skeleton point cloud. Finally, the symbol-constrained point cloud is fused with multipath reflection intensity to generate a set of penetrating features containing text symbol positioning information and skeleton geometric parameters. This method achieves high-precision identification and 3D reconstruction of power towers, providing reliable support for facility maintenance.

[0191] In summary, steps 601 to 605 achieve multi-angle signal acquisition of power transmission towers using a ring- or linearly deployed radar array. Beamforming processing is applied to the original reflected signals based on multi-modal mapping relationships to filter target reflection areas containing text symbol positioning information. Phase difference interferometry is then used to generate an initial 3D point cloud of the tower skeleton outline. Further, spatial filtering is applied to the initial point cloud based on the component correlation of the text symbols to remove irrelevant scattering points, generating a symbol-constrained tower skeleton point cloud. Finally, the symbol-constrained point cloud is fused with multipath reflection intensity features to generate a set of penetrating features containing text symbol positioning information and tower skeleton geometric parameters. This achieves high-precision positioning and feature extraction of text symbols and skeleton geometry on power transmission towers.

[0192] To address the challenges of removing irrelevant scattering points from the initial 3D point cloud and generating symbol-constrained tower skeleton point clouds, some embodiments include step 603, which involves spatial filtering of the initial 3D point cloud based on the component correlation of text symbols in multimodal mapping relationships to remove scattering points irrelevant to text symbol positioning and generate symbol-constrained tower skeleton point clouds.

[0193] 701. Based on the correlation between text symbols and tower components in the multimodal mapping relationship, divide the symbol semantically associated region into non-associated regions;

[0194] In step 701, the multimodal mapping relationship refers to the association between text symbols and tower components, typically implemented through a predefined mapping table or machine learning model. The semantically associated region refers to the region related to the semantics of the text symbols, usually containing scattering points directly related to the tower components. The unassociated region refers to the region unrelated to the semantics of the text symbols, usually containing scattering points or noise points unrelated to the tower components.

[0195] In this embodiment, based on the correlation between text symbols and tower components in the multimodal mapping relationship, a region partitioning algorithm (such as a semantic-based segmentation algorithm or clustering algorithm) is used to divide the point cloud into symbolically related regions and unrelated regions. Symbolically related regions include areas associated with lightning protection symbols, support symbols, and fixing symbols, while unrelated regions include areas unrelated to coating corrosion symbols. During the region partitioning process, the multimodal mapping relationship is used to guide the partitioning, ensuring the accuracy of distinguishing between symbolically related and unrelated regions. Finally, the partitioning results of symbolically related and unrelated regions are obtained.

[0196] 702. Within the symbol semantic association area, the scattering points are processed according to the component association of the text symbols. Within the non-symbolic semantic association area, the scattering points that are unrelated to the rust expansion characteristics are eliminated according to the defect boundary distribution of the coating corrosion symbols.

[0197] In step 702, the component association of text symbols refers to the correspondence between text symbols and tower components, typically achieved through a predefined mapping table or machine learning model. Scattering points refer to point cloud data generated by radar signal reflection, usually containing information about the target object and noise. The defect boundary distribution of coating corrosion symbols refers to the boundary features of the coating corrosion area, typically extracted using image processing techniques. Rust expansion characteristics refer to the expansion and deformation characteristics of the coating corrosion area, usually related to the degree of corrosion.

[0198] In this embodiment, firstly, within the semantically associated region of the symbols, the scattering points are processed based on the component association of the text symbols. Point cloud processing algorithms (such as density-based filtering or distance-based filtering) are used to retain scattering points related to the tower components. Next, within the non-semantically associated region, based on the defect boundary distribution of the coating corrosion symbols, point cloud elimination algorithms (such as boundary-based filtering or feature-based filtering) are used to eliminate scattering points unrelated to the corrosion expansion characteristics. During point cloud processing, the component association of the text symbols and the defect boundary distribution of the coating corrosion symbols guide the processing of scattering points, ensuring that scattering points within the semantically associated region are retained, while irrelevant scattering points in the non-associated region are eliminated. Finally, the scattering points after processing the semantically associated region and the remaining scattering points after eliminating the non-associated region are obtained.

[0199] 703. Mark the symbol semantic association region of the initial three-dimensional point cloud. When marking, the lightning protection symbol region radiates outward with the symbol center as the sphere center, the support symbol region extends along the main beam axis, and the fixed symbol region extends according to the bolt distribution density matrix to generate the region marked point cloud.

[0200] In step 703, the initial 3D point cloud refers to the 3D point cloud data generated through radar signal processing, which typically includes the spatial coordinate information of the target object. Symbol semantic association region labeling refers to the identification and expansion of symbol semantic association regions, typically used for subsequent point cloud processing and analysis. The lightning protection symbol region refers to the region associated with the lightning protection symbol, typically radiating outwards from the symbol's center. The support symbol region refers to the region associated with the support symbol, typically extending along the main beam's axial direction. The fixed symbol region refers to the region associated with the fixed symbol, typically expanding according to the bolt distribution density matrix. The region-labeled point cloud refers to the point cloud data after symbol semantic association region labeling, typically used for subsequent multi-scale aggregation.

[0201] In this embodiment, the initial 3D point cloud is first labeled with symbolic semantic association regions. Region labeling algorithms (such as a center-based expansion algorithm or an axial extension algorithm) are used to label the lightning protection symbol region, support symbol region, and fixed symbol region. The lightning protection symbol region expands outwards from the symbol center, the support symbol region extends along the main beam axial direction, and the fixed symbol region expands according to the bolt distribution density matrix. During the region labeling process, the characteristics of the symbolic semantic association regions are used to guide the labeling, ensuring the accuracy of the region-labeled point cloud. Finally, a region-labeled point cloud containing the symbolic semantic association region labels is obtained.

[0202] 704. Based on the multi-scale accuracy parameters of the region marker point cloud, the scattering points after processing the symbol semantic association region and the remaining scattering points after removing the non-association region are aggregated at multiple scales to generate a symbol-constrained tower skeleton point cloud.

[0203] In step 704, the multi-scale accuracy parameters refer to the parameters used to control the accuracy of point cloud aggregation, which typically include point cloud density, resolution, and scale range. Multi-scale aggregation refers to fusing point cloud data at different scales to generate more comprehensive point cloud data. The symbolically constrained tower skeleton point cloud refers to the point cloud data containing symbolic semantic information generated after multi-scale aggregation, which is typically used for subsequent 3D reconstruction and analysis.

[0204] In this embodiment, firstly, based on the multi-scale accuracy parameters of the region-marked point cloud, a multi-scale aggregation algorithm (such as density-based aggregation or resolution-based aggregation) is used to aggregate the scattering points after processing the symbolically associated regions with the remaining scattering points after removing the non-associated regions, generating a symbolically constrained tower skeleton point cloud. During the multi-scale aggregation process, the multi-scale accuracy parameters are used to control the aggregation accuracy, ensuring the integrity and accuracy of the symbolically constrained tower skeleton point cloud. Finally, a symbolically constrained tower skeleton point cloud for 3D reconstruction of power transmission towers is obtained.

[0205] Here is a specific example:

[0206] In power tower identification, the tower surface is first divided into semantically associated and non-associated regions based on the multimodal mapping relationship between text symbols and tower components. For example, the text symbol "A" is associated with lightning protection components, and the area surrounding it is a semantically associated region, while other areas are non-associated regions. Within the semantically associated regions, scattering points are processed according to the component association of the text symbol; for example, scattering points related to lightning protection components are retained. In non-associated regions, scattering points unrelated to corrosion expansion characteristics are removed based on the defect boundary distribution of coating corrosion symbols; for example, noise points caused by coating peeling are removed. Next, the initial 3D point cloud is marked with semantically associated regions. Finally, based on the multi-scale accuracy parameters of the marked point cloud, the scattering points processed from the semantically associated regions are aggregated with the remaining scattering points after removing non-associated regions to generate a symbol-constrained tower skeleton point cloud. For example, through multi-scale aggregation, the scattering points from the lightning protection symbol region, the support symbol region, and the fixed symbol region are fused to form an accurate tower skeleton point cloud. This method improves the accuracy and reliability of power tower identification through symbolic semantic constraints, providing strong support for facility maintenance.

[0207] In summary, steps 701 to 704 establish the association between text symbols and tower components based on multimodal mapping, dividing the symbol semantically associated regions into non-associated regions. Within the associated regions, scattering points are processed according to the component association of the text symbols, while irrelevant scattering points are eliminated in the non-associated regions based on the defect boundary distribution of the coating corrosion symbols. Furthermore, the initial 3D point cloud is marked with symbol semantically associated regions, expanded according to the distribution characteristics of lightning protection symbols, support symbols, and fixed symbols, generating a region-marked point cloud. Finally, based on the multi-scale accuracy parameters of the region-marked point cloud, the scattering points processed in the symbol semantically associated regions are aggregated with the remaining scattering points after eliminating non-associated regions at multiple scales to generate a symbol-constrained tower skeleton point cloud, thus achieving high-precision symbol semantic constraints and region optimization processing of the tower skeleton point cloud.

[0208] To address the challenges of suppressing electromagnetic interference and insufficient noise compensation accuracy in multipath reflection signals, some embodiments include, in step 102, constructing a dynamic noise suppression model based on the penetration feature set, and using this model to compensate for attenuation of the multipath reflection signal, thereby generating an anti-electromagnetic interference compensation feature set, including:

[0209] 801. Extract symbolic noise suppression parameters by means of the spatial mapping relationship between the text symbol location information in the penetration feature set and the tower components, and construct a dynamic noise suppression model;

[0210] In step 801, the penetration feature set refers to a set of feature data containing the location information of text symbols and the geometric parameters of the tower skeleton, typically generated through radar signal processing. The location information of text symbols refers to the position information of the text symbols on the tower surface, typically extracted through image processing techniques. The spatial mapping relationship of tower components refers to the spatial correspondence between text symbols and tower components, typically implemented through a predefined mapping table or machine learning model. The symbol semantic noise suppression parameters are parameters used to suppress noise unrelated to the semantics of the text symbols, typically generated through feature extraction algorithms. The dynamic noise suppression model is a model used to suppress noise in multipath reflection signals, typically constructed through parameterization methods.

[0211] In this embodiment, symbolic noise suppression parameters are first extracted using a feature extraction algorithm (such as semantic-based filtering or mapping-based feature extraction) based on the spatial mapping relationship between the text symbol location information in the penetration feature set and the tower components. Next, a dynamic noise suppression model is constructed using a parameterization method (such as weight-based model building or optimization-based model building) based on the symbolic noise suppression parameters. During feature extraction and model building, the spatial mapping relationship between the text symbol location information and the tower components guides the extraction of noise suppression parameters and the construction of the model, ensuring the accuracy of the dynamic noise suppression model. Finally, a dynamic noise suppression model containing the symbolic noise suppression parameters is obtained.

[0212] 802. In the dynamic noise suppression model, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower skeleton outline;

[0213] In step 802, the tower skeleton contour geometry parameters refer to the shape and size information of the tower skeleton, which are typically generated through 3D point cloud processing. The structural reflection path attenuation coefficient refers to the degree of attenuation of multipath reflected signals in the tower structure, and is usually related to the geometry of the tower skeleton. Dynamic adjustment refers to adjusting the structural reflection path attenuation coefficient according to the tower skeleton contour geometry parameters, which is usually achieved through parameterization methods.

[0214] In this embodiment, the attenuation coefficient of the structural reflection path is first dynamically adjusted in the dynamic noise suppression model based on the geometric parameters of the tower skeleton profile using a parametric method (such as geometry-based attenuation coefficient adjustment or optimization-based attenuation coefficient adjustment). During the parametric adjustment process, the geometric parameters of the tower skeleton profile are used to guide the adjustment of the attenuation coefficient, ensuring the accuracy of the attenuation coefficient of the structural reflection path. Finally, a dynamic noise suppression model containing the dynamically adjusted attenuation coefficient of the structural reflection path is obtained.

[0215] 803. Based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, generate an interference source spatial weight map, integrate the interference source spatial weight map into the dynamic noise suppression model, perform azimuth-related attenuation compensation on the multipath reflection signal, and extract the azimuth compensation parameters.

[0216] In step 803, electromagnetic interference distribution characteristics refer to the distribution of electromagnetic interference around the power transmission tower, typically obtained through electromagnetic field measurement technology. The azimuth information of the high-voltage equipment refers to the direction information of the high-voltage equipment relative to the power transmission tower, typically obtained through azimuth measurement technology. The interference source spatial weight map is a graph representing the spatial weight distribution of electromagnetic interference sources, typically generated through a weight allocation algorithm. Azimuth-related attenuation compensation refers to attenuation compensation of multipath reflected signals based on the azimuth of the interference source, typically implemented through parameterization methods. Azimuth compensation parameters are parameters used for azimuth-related attenuation compensation, typically generated through feature extraction algorithms.

[0217] In this embodiment, firstly, based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of high-voltage equipment around the power tower, a weighted allocation algorithm (such as distance-based weighted allocation or direction-based weighted allocation) is used to generate an interference source spatial weight map. Next, the interference source spatial weight map is integrated into a dynamic noise suppression model, and a parameterized method (such as weight-based attenuation compensation or optimization-based attenuation compensation) is used to perform azimuth-related attenuation compensation on the multipath reflection signal, extracting azimuth compensation parameters. During the weight allocation and attenuation compensation process, the electromagnetic interference distribution characteristics and the azimuth information of the high-voltage equipment are used to guide the generation of the interference source spatial weight map and the extraction of azimuth compensation parameters, ensuring the accuracy of azimuth-related attenuation compensation. Finally, a dynamic noise suppression model containing azimuth compensation parameters is obtained.

[0218] 804. The symbolic semantic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are weighted and fused in the tower area of ​​the dynamic noise suppression model to update the dynamic noise suppression model;

[0219] In step 804, the weighted fusion of the tower area refers to the weighted fusion of symbolic semantic noise suppression parameters, structural reflection path attenuation coefficients, and azimuth compensation parameters in the dynamic noise suppression model, typically achieved through weighted summation or convolutional fusion. Updating the dynamic noise suppression model refers to adjusting the model parameters based on the weighted fusion results, usually achieved through parameterization methods.

[0220] In this embodiment, the symbolic semantic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are first weighted and fused within the tower region of the dynamic noise suppression model. A weighted fusion algorithm (such as weighted summation or convolutional fusion) is used to generate the fused parameters. Next, based on the fused parameters, a parameterized method (such as weight-based model update or optimization-based model update) is used to update the dynamic noise suppression model. During the weighted fusion and model update processes, the symbolic semantic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters guide the model update, ensuring the accuracy of the dynamic noise suppression model. Finally, the updated dynamic noise suppression model is obtained.

[0221] 805. The multipath reflection signal in the penetration feature set is iteratively attenuated by the updated dynamic noise suppression model to generate a compensation feature set for electromagnetic interference resistance.

[0222] In step 805, iterative path attenuation compensation refers to compensating for the path attenuation of multipath reflection signals through multiple iterations, typically achieved through iterative optimization methods. The anti-electromagnetic interference compensation feature set refers to the set of anti-interference feature data generated after path attenuation compensation, usually used for subsequent 3D reconstruction and analysis.

[0223] In this embodiment, firstly, iterative path attenuation compensation is performed on the multipath reflection signals in the penetration feature set using an updated dynamic noise suppression model. An iterative optimization method (such as gradient descent or backpropagation) is then employed to generate a compensation feature set for electromagnetic interference (EMI) mitigation. During the iterative optimization process, the updated dynamic noise suppression model guides the path attenuation compensation, ensuring the accuracy and robustness of the EMI mitigation compensation feature set. Finally, an EMI mitigation compensation feature set for 3D reconstruction of power transmission towers is obtained.

[0224] Here is a specific example:

[0225] In power tower identification, this method first extracts symbolic noise suppression parameters based on the spatial mapping relationship between the location information of text symbols in the penetration feature set and the tower components, thus constructing a dynamic noise suppression model. For example, the correlation between the text symbol "A" and lightning protection components is used to suppress irrelevant noise. Next, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower frame outline, for example, to enhance the signal effectiveness in the main beam area. Subsequently, by combining the electromagnetic interference distribution characteristics and the azimuth information of surrounding high-voltage equipment, a spatial weight map of interference sources is generated, and azimuth-related attenuation compensation is performed on multipath reflection signals, extracting azimuth compensation parameters. For example, signal compensation is performed for interference from high-voltage equipment. The symbolic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are weighted and fused in the model to update the dynamic noise suppression model. Finally, the updated model is used to perform iterative path attenuation compensation on multipath reflection signals, generating a compensation feature set for electromagnetic interference resistance. This method significantly improves the anti-interference capability of power tower identification through dynamic noise suppression and compensation, providing reliable support for facility maintenance.

[0226] In summary, steps 801 to 805 extract symbolic noise suppression parameters by mapping the spatial relationship between the text symbol positioning information in the penetration feature set and the tower components, construct a dynamic noise suppression model, and dynamically adjust the structural reflection path attenuation coefficient based on the tower skeleton outline geometric parameters. Subsequently, the symbolic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are weighted and fused in the tower region within the model to update the dynamic noise suppression model. Finally, the updated model is used to iteratively compensate for the multipath reflection signals in the penetration feature set, generating a compensation feature set for electromagnetic interference resistance, thereby achieving efficient noise suppression and electromagnetic interference compensation for multipath reflection signals from power transmission towers.

[0227] To address the challenges of reconstructing missing tower connection components and the insufficient environmental interference resistance of the 3D recognition model, some embodiments include, in step 105, reconstructing the missing tower connection components based on the geometric constraints to generate an environmentally resistant 3D recognition model of the power tower, which includes:

[0228] 901. Based on the curvature continuity constraint of the main beam of the tower and the topological distribution constraint of the bolt hole positions in the geometric constraint relationship, the area with missing connecting parts is spatially located to generate the boundary box of the missing area;

[0229] In step 901, geometric constraints refer to the spatial and morphological relationships between supporting structural features and coating texture features, typically generated through multimodal feature modeling. The main beam curvature continuity constraint means that the curvature change of the main beam in space should remain continuous, usually achieved through curvature analysis techniques. The bolt hole topology distribution constraint means that the spatial distribution of bolt holes should satisfy specific topological relationships, typically achieved through topology analysis techniques. Regions with missing connecting components refer to parts of the tower structure lacking connecting components, typically identified through point cloud analysis techniques. The missing region bounding box refers to the spatial bounding box used to identify regions with missing connecting components, typically achieved through bounding box generation algorithms.

[0230] In this embodiment, firstly, based on the geometric constraints of the tower main beam curvature continuity and the bolt hole topology distribution, a spatial positioning algorithm (such as curvature-based positioning or topology-based positioning) is used to spatially locate the region of the missing connecting component, generating a bounding box for the missing region. During the spatial positioning process, the tower main beam curvature continuity constraint and the bolt hole topology distribution constraint are used to guide the positioning of the missing region, ensuring the accuracy of the bounding box. Finally, the bounding box of the missing region, including the region of the missing connecting component, is obtained.

[0231] 902. Limit the scope to the boundary box of the missing area, and generate candidate connection component parameters based on the angular constraint relationship between the main beam and the support;

[0232] In step 902, the angular constraint relationship between the main beam and the support refers to the specific geometric relationship that the angle between the main beam and the support should satisfy, which is usually achieved through angle analysis techniques. Candidate connection component parameters refer to the candidate parameters used to reconstruct missing connection components, and are usually generated through parametric methods.

[0233] In this embodiment, the scope is first limited to the bounding box of the missing region. Based on the angular constraint relationship between the main beam and the support, parameterization methods (such as angle-based parameter generation or geometry-based parameter generation) are used to generate candidate connection component parameters. During the parameterization process, the angular constraint relationship between the main beam and the support is used to guide the generation of candidate connection component parameters, ensuring the rationality of the candidate connection component parameters. Finally, a set of candidate parameters containing the candidate connection component parameters is obtained.

[0234] 903. Based on the topological distribution constraints of bolt hole positions, perform spatial matching verification on the parameters of the candidate connection components, eliminate candidate parameters whose axial deviation from the bolt hole positions exceeds a preset threshold, and generate a preliminary set of connection components;

[0235] In step 903, the bolt hole topology distribution constraint refers to the requirement that the spatial distribution of bolt holes should satisfy a specific topological relationship, typically achieved through topology analysis techniques. Spatial matching verification refers to verifying the spatial matching of candidate connection component parameters, typically achieved through matching algorithms. Axial deviation refers to the axial deviation between the candidate connection component parameters and the bolt hole positions, typically achieved through deviation analysis techniques. The preset threshold is a threshold used to determine whether the axial deviation is reasonable, typically determined experimentally or empirically. The preliminary connection component set refers to the set of candidate connection component parameters retained after spatial matching verification, typically used for subsequent morphological optimization.

[0236] In this embodiment, firstly, based on the topological distribution constraints of bolt hole positions, a matching algorithm (such as topology-based matching or geometry-based matching) is used to perform spatial matching verification on the parameters of candidate connection components. Candidate parameters whose axial deviation from the bolt hole positions exceeds a preset threshold are eliminated, generating a preliminary set of connection components. During the matching verification process, the topological distribution constraints of bolt hole positions are used to guide the verification of candidate connection component parameters, ensuring the rationality of the preliminary set of connection components. Finally, a preliminary set of connection components containing reasonable candidate connection component parameters is obtained.

[0237] 904. Based on the curvature continuity constraint of the main beam of the tower, the initial set of connecting components is morphologically optimized, and the installation angle and contact surface curvature of the connecting components are adjusted to meet the curvature difference threshold of the point cloud of the main beam skeleton.

[0238] In step 904, the curvature continuity constraint of the main beam of the tower body means that the curvature change of the main beam in space should remain continuous, which is usually achieved through curvature analysis techniques. Shape optimization refers to optimizing the shape of the initial set of connecting components, which is usually achieved through optimization algorithms. Installation angle refers to the installation angle between the connecting component and the main beam, which is usually achieved through angle analysis techniques. Contact surface curvature refers to the curvature of the contact surface between the connecting component and the main beam, which is usually achieved through curvature analysis techniques. The curvature difference threshold is a threshold used to determine curvature continuity, which is usually determined experimentally or empirically.

[0239] In this embodiment, firstly, based on the curvature continuity constraint of the main beam of the tower, an optimization algorithm (such as curvature-based optimization or geometry-based optimization) is used to perform morphological optimization on the initial set of connecting components. The installation angle and contact surface curvature of the connecting components are adjusted to meet the curvature difference threshold of the main beam skeleton point cloud. During the morphological optimization process, the curvature continuity constraint of the main beam of the tower is used to guide the optimization of the connecting components, ensuring the rationality of the installation angle and contact surface curvature of the connecting components. Finally, the morphologically optimized set of connecting components is obtained.

[0240] 905. Based on the load distribution characteristics of the support component and the tower frame, the environmental interference resistance reinforcement model is performed on the optimized connection component, wherein the reinforcement strength is dynamically adjusted according to the connection density between the support component and the main beam.

[0241] In step 905, the load distribution characteristics between the support member and the tower frame refer to the specific mechanical relationship that the load distribution between the support member and the tower frame should satisfy, which is usually achieved through mechanical analysis techniques. Environmental interference resistance reinforcement modeling refers to performing environmental interference resistance reinforcement modeling on the morphologically optimized connecting components, which is usually achieved through reinforcement algorithms. Reinforcement strength refers to the reinforcement strength of the connecting components, which is usually achieved through strength analysis techniques. The connection density between the support member and the main beam refers to the connection density between the support member and the main beam, which is usually achieved through density analysis techniques.

[0242] In this embodiment, firstly, based on the load distribution characteristics of the support components and the tower frame, a reinforcement algorithm (such as load-based reinforcement or strength-based reinforcement) is used to perform environmental interference resistance reinforcement modeling on the morphologically optimized connecting components. The reinforcement strength is dynamically adjusted according to the connection density between the support components and the main beam. During the reinforcement modeling process, the load distribution characteristics of the support components and the tower frame are used to guide the reinforcement of the connecting components, ensuring the rationality of the reinforcement strength. Finally, a set of connecting components after environmental interference resistance reinforcement modeling is obtained.

[0243] 906. Perform geometric constraint matching between the reinforced and modeled connecting components and the point cloud of the tower skeleton to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.

[0244] In step 906, geometric constraint matching refers to matching the geometric constraints of the reinforced modeled connecting components with the point cloud of the tower skeleton, typically achieved through a matching algorithm. The environmentally resistant 3D recognition model of the power tower refers to the interference-resistant 3D recognition model generated after geometric constraint matching, usually used for subsequent inspection and maintenance.

[0245] In this embodiment, the reinforced modeled connecting components are first geometrically matched with the tower skeleton point cloud. A matching algorithm (such as geometry-based matching or constraint-based matching) is then used to generate a 3D recognition model of the power tower resistant to environmental interference. During the geometric constraint matching process, the reinforced modeled connecting components and the tower skeleton point cloud are used to guide the model generation, ensuring the integrity and accuracy of the 3D recognition model of the power tower resistant to environmental interference. Finally, a 3D recognition model of the power tower resistant to environmental interference is obtained for power tower inspection and maintenance.

[0246] Here is a specific example:

[0247] During the inspection of a power transmission tower, a missing connector was found in the middle of the tower due to long-term corrosion. Analysis of the main beam curvature and bolt hole distribution determined the missing area's boundary to be between 15 and 17 meters in height. Within this boundary, multiple candidate connector parameters were generated based on the standard angle between the main beam and the support (typically 45° to 60°), such as a connector with a length of 1.2 meters and a diameter of 0.3 meters. Subsequently, the candidate parameters were validated using the bolt hole topology distribution, eliminating options with axial deviations exceeding 5 millimeters from the bolt holes, ultimately selecting candidates aligned with the bolt holes. Based on the main beam curvature continuity constraint, the selected connector underwent morphological optimization, adjusting its installation angle and contact surface curvature to match the main beam curvature. For example, the connector's installation angle was adjusted to 52°, and the contact surface curvature to 0.15 to match the main beam curvature. Finally, based on the load distribution characteristics of the support and the tower frame, environmental interference resistance reinforcement modeling was performed on the optimized connector. For example, a 2mm thick reinforcement layer was added to the contact area between the connector and the main beam to improve wind load resistance. Finally, the reinforced connector was geometrically matched with the tower frame point cloud to ensure precise alignment, generating a complete 3D recognition model of the power transmission tower. This method successfully repaired the missing connector and significantly improved the tower's anti-interference capability, providing a reliable basis for subsequent maintenance.

[0248] In summary, steps 901 to 906 achieved precise localization of missing connection component regions and generation of missing region bounding boxes based on the curvature continuity constraint of the main beam of the tower and the topological distribution constraint of bolt holes. Simultaneously, candidate connection component parameters were generated and optimized to ensure they met the axial deviation threshold of bolt holes and the curvature difference threshold of the main beam. Furthermore, through environmental interference-resistant reinforcement modeling, the reinforcement strength of the connection components was dynamically adjusted to adapt to the load distribution characteristics of the support components and the main beam. Finally, an environmental interference-resistant 3D recognition model of the power tower was generated, thus achieving high-precision localization and optimization of missing connection components, as well as reliable reconstruction and improved anti-interference capabilities of the 3D model of the power tower.

[0249] Figure 2 This application provides a schematic diagram of the structure of a power tower identification system based on text symbols, as shown in the embodiment. Figure 2 As shown, the system includes:

[0250] Module 21 constructs a multimodal mapping relationship between the text symbols and tower components based on the obtained text symbols on the surface of the power tower.

[0251] The analysis module 22 acquires the original reflection signal of the power tower through the radar array, filters the target reflection area based on the multimodal mapping relationship, analyzes the phase difference data in the original reflection signal, and generates a set of penetrating features containing the outline of the tower skeleton and the positioning information of the corresponding text symbols.

[0252] Compensation module 23 constructs a dynamic noise suppression model based on the penetration feature set, and performs attenuation compensation on multipath reflection signals through the dynamic noise suppression model to generate a compensation feature set for electromagnetic interference resistance.

[0253] The update module 24 inputs the compensation feature set and the coating geometric identification features into the shared weight encoder for cross-modal association learning. The shared weight encoder fuses the semantic vectors of text symbols based on the attention mechanism, and simultaneously extracts the internal support structure features and the external anti-corrosion coating texture features of the tower body. The geometric constraint relationship between the support structure and the coating texture is established by alternately updating the convolution kernel parameters.

[0254] The generation module 25 reconstructs the missing tower connection components based on the geometric constraints, and generates a three-dimensional recognition model of the power tower that is resistant to environmental interference.

[0255] Figure 2 The aforementioned power tower identification system based on text symbols can perform... Figure 1 The implementation principle and technical effects of the power transmission tower identification method based on text symbols described in the illustrated embodiment will not be repeated here. The specific methods by which each module and unit of the power transmission tower identification system based on text symbols in the above embodiments are described in detail in the embodiments related to this method, and will not be elaborated upon here.

[0256] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of this application.

Claims

1. A method for identifying power transmission towers based on text symbols, characterized in that, include: Based on the text symbols obtained from the surface of the power tower, a multimodal mapping relationship between the text symbols and the tower components is constructed; The original reflected signals of the power tower are obtained by radar array, the target reflection area is filtered based on the multimodal mapping relationship, the phase difference data in the original reflected signals is analyzed, and a set of penetrating features containing the outline of the tower skeleton and the positioning information of the corresponding text symbols is generated. A dynamic noise suppression model is constructed based on the aforementioned penetration feature set. The dynamic noise suppression model is used to compensate for the attenuation of multipath reflection signals, thereby generating a set of compensation features to resist electromagnetic interference. The compensation feature set and the coating geometric identification features are input into a shared weight encoder for cross-modal association learning. The shared weight encoder fuses the semantic vectors of text symbols based on an attention mechanism, and simultaneously extracts the internal support structure features and the external anti-corrosion coating texture features of the tower body. The geometric constraint relationship between the support structure and the coating texture is established by alternately updating the convolution kernel parameters. Based on the aforementioned geometric constraints, the missing tower connection components are reconstructed, generating a three-dimensional recognition model of the power tower resistant to environmental interference. Specifically, the compensation feature set and the coating geometric identification features are input into a shared weight encoder for cross-modal association learning. The shared weight encoder, based on an attention mechanism, fuses semantic vectors of text symbols to simultaneously extract the internal support structure features and the external anti-corrosion coating texture features of the tower body. Furthermore, it establishes a geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters, including: The compensation feature set and the coating geometric identification feature are input into the shared weight encoder. The multi-layer network structure in the shared weight encoder is used to extract features from the compensation feature set and the coating geometric identification feature to generate a preliminary feature vector. The initial feature vector and the semantic vector of the text symbol are weighted and fused based on the attention mechanism to generate a fused feature vector; By performing multi-scale decomposition on the fused feature vector, the internal support structure features and the external anti-corrosion coating texture features of the tower body are extracted respectively. The support structure features are extracted by high-frequency components, and the coating texture features are extracted by low-frequency components. Based on the local consistency between the internal support structure features and the external anti-corrosion coating texture features, the convolution kernel parameters are updated alternately, and a geometric constraint relationship between the support structure features and the coating texture features is established.

2. The method according to claim 1, characterized in that, The initial feature vector and the semantic vector of the text symbols are weighted and fused based on an attention mechanism to generate a fused feature vector, including: The preliminary feature vector and the semantic vector of the text symbol are respectively input into independent fully connected layers for feature dimension alignment, generating aligned preliminary feature vector and semantic vector of the text symbol; Based on the spatial distribution of the aligned preliminary feature vector and the semantic vector of the text symbol, local symbol association features and global structure association features are constructed, wherein the local symbol association features are generated by the mean of features in the symbol neighborhood, and the global structure association features are generated by cross-region feature difference. By cross-projecting the local symbol association features and the global structure association features, a symbol structure association mapping vector is generated, and a dynamic correlation score between the aligned preliminary feature vector and the semantic vector of the text symbols is calculated based on the symbol structure association mapping vector. Based on the dynamic relevance score, channel-level weights are assigned to the aligned preliminary feature vector, and symbol-level weights are assigned to the text symbol semantic vector, generating channel-weighted preliminary feature vectors and symbol-weighted semantic vectors; The channel-weighted preliminary feature vector and the symbol-weighted semantic vector are concatenated using multi-scale context features to generate a multi-scale context feature concatenation result. The multi-scale context features are extracted by sliding the context window of the symbol semantics, and the window size is positively correlated with the density of text symbols on the surface of the power tower. Based on the topological constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale context feature splicing results to generate a fused feature vector.

3. The method according to claim 2, characterized in that, By cross-projecting the local symbol association features and the global structure association features, a symbol structure association mapping vector is generated. Based on this vector, a dynamic relevance score is calculated between the aligned preliminary feature vector and the semantic vector of the text symbols, including: The local symbol association features are spatially encoded and channel-dimensional encoded to generate spatially encoded local symbol association features and channel-encoded global structural association features. By interactively superimposing feature channels, the spatially encoded local symbol association features and the channel-encoded global structure association features are bidirectionally projected to generate a symbol structure association mapping vector. The spatial weight matrix and the channel weight matrix are extracted from the symbol structure association mapping vector, wherein the spatial weight matrix is ​​generated based on the local similarity of symbol semantics, and the channel weight matrix is ​​generated based on the global difference of structural features; By parametrically weighting the spatial topological constraints of the global structural association features, and fusing the spatial weight matrix and the channel weight matrix, dynamic correlation weight coefficients are generated. Based on the dynamic relevance weight coefficient, the cross-modal association strength between the aligned preliminary feature vector and the semantic vector of the text symbol is quantified to generate a dynamic relevance score.

4. The method according to claim 1, characterized in that, Based on the local consistency of the internal support structure features and the external anti-corrosion coating texture features, the convolution kernel parameters are alternately updated, and a geometric constraint relationship is established between the support structure features and the coating texture features, including: The internal support structure features are decomposed into the spatial distribution features of connectors and the morphological features of the main beam, and the external anti-corrosion coating texture features are decomposed into the corrosion area distribution features and the coating thickness variation features. Based on the proximity constraint between bolt hole locations and coating defect areas, spatial topological coding is performed on the spatial distribution characteristics of the connector and the distribution characteristics of the corrosion area to generate corrosion association features of the connector. Based on the synchronicity constraint between the tower curvature change and the coating thickness gradient, the morphological and thickness coupling coding of the main beam morphological features and coating thickness change features is performed to generate the main beam thickness coupling features. Construct a multi-branch convolutional kernel group, wherein the first branch performs spatial constraint modeling on the corrosion correlation features of the connector, and the second branch performs texture constraint modeling on the thickness coupling features of the main beam; Based on the local consistency loss of the spatial constraint modeling results and the texture constraint modeling results, the parameters of the multi-branch convolution kernel group are alternately updated to generate the geometric constraint relationship between the supporting structure features and the coating texture features.

5. The method according to claim 1, characterized in that, The original reflected signals of the power transmission towers are acquired using a radar array. Based on the multimodal mapping relationship, target reflection areas are filtered. Phase difference data in the original reflected signals is analyzed, and a set of penetrating features containing the tower's skeleton outline and corresponding text symbol positioning information is generated, including: Multi-angle signal acquisition is performed on power transmission towers using a circular or linearly deployed radar array to obtain the original set of reflected signals. During signal acquisition, the direction of the radar beam is aligned with the spatial distribution of text symbols on the tower surface. Based on the multimodal mapping relationship, beamforming processing is performed on the original set of reflected signals to filter target reflection areas containing text symbol positioning information. The beamforming parameters corresponding to the beamforming processing are dynamically adjusted according to the size of the text symbol and the geometric dimensions of the tower components. Phase difference interferometry processing is performed on the original reflection signal of the target reflection area to analyze the phase difference data between adjacent radar channels and generate an initial three-dimensional point cloud of the tower skeleton outline. Based on the component correlation of text symbols in multimodal mapping, spatial filtering is performed on the initial three-dimensional point cloud to remove scattering points that are irrelevant to the positioning of text symbols, thereby generating a symbol-constrained tower skeleton point cloud. The point cloud of the tower skeleton constrained by the symbol is fused with the multipath reflection intensity from the previous phase difference interferometry processing to generate a set of penetrating features containing the positioning information of the text symbols and the geometric parameters of the tower skeleton.

6. The method according to claim 5, characterized in that, Based on the component correlation of text symbols in multimodal mapping relationships, spatial filtering is performed on the initial 3D point cloud to remove scattering points unrelated to the positioning of text symbols, generating a symbol-constrained tower skeleton point cloud, including: Based on the correlation between text symbols and tower components in the multimodal mapping relationship, the semantically related regions and non-related regions of the symbols are divided. Within the symbol semantic association area, the scattering points are processed according to the component association of the text symbols. Within the non-symbolic semantic association area, the scattering points that are unrelated to the rust expansion characteristics are eliminated according to the defect boundary distribution of the coating corrosion symbols. The initial three-dimensional point cloud is marked with symbolic semantic association regions. During the marking process, the lightning protection symbol region radiates outward from the center of the symbol, the support symbol region extends along the axis of the main beam, and the fixed symbol region extends according to the bolt distribution density matrix, thereby generating a region-marked point cloud. Based on the multi-scale accuracy parameters of the region marker point cloud, the scattering points after processing the symbolic semantic association region and the remaining scattering points after removing the non-association region are aggregated at multiple scales to generate a symbol-constrained tower skeleton point cloud.

7. The method according to claim 1, characterized in that, A dynamic noise suppression model is constructed based on the aforementioned penetration feature set. This model is then used to compensate for the attenuation of multipath reflection signals, generating a set of electromagnetic interference-resistant compensation features, including: By extracting symbolic noise suppression parameters from the location information of text symbols in the penetration feature set and the spatial mapping relationship between them and the tower components, a dynamic noise suppression model is constructed. In the dynamic noise suppression model, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower skeleton profile. Based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, an interference source spatial weight map is generated, and the interference source spatial weight map is integrated into the dynamic noise suppression model to perform azimuth-related attenuation compensation on the multipath reflection signal and extract the azimuth compensation parameters. The symbolic semantic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are weighted and fused in the tower area of ​​the dynamic noise suppression model to update the dynamic noise suppression model. The updated dynamic noise suppression model is used to perform iterative path attenuation compensation on the multipath reflection signals in the penetration feature set, thereby generating a compensation feature set for electromagnetic interference resistance.

8. The method according to claim 1, characterized in that, Based on the aforementioned geometric constraints, the missing tower connection components are reconstructed, generating a three-dimensional recognition model of the power tower resistant to environmental interference, including: Based on the curvature continuity constraint of the main beam of the tower and the topological distribution constraint of the bolt hole positions in the geometric constraint relationship, the area with missing connecting parts is spatially located to generate the boundary box of the missing area. The scope is locked within the boundary box of the missing area, and candidate connection component parameters are generated based on the angular constraint relationship between the main beam and the support component. Based on the topological distribution constraints of bolt hole positions, the parameters of the candidate connection components are spatially matched and verified. Candidate parameters whose axial deviation from the bolt hole positions exceeds a preset threshold are eliminated, and a preliminary set of connection components is generated. Based on the curvature continuity constraint of the main beam of the tower, the initial set of connecting components is morphologically optimized, and the installation angle and contact surface curvature of the connecting components are adjusted to meet the curvature difference threshold of the main beam skeleton point cloud. Based on the load distribution characteristics of the support components and the tower frame, the environmental interference resistance reinforcement model is performed on the optimized connection components, wherein the reinforcement strength is dynamically adjusted according to the connection density between the support components and the main beam. The reinforced and modeled connecting components are geometrically matched with the point cloud of the tower skeleton to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.

9. A power transmission tower identification system based on text symbols, used to execute the power transmission tower identification method based on text symbols as described in any one of claims 1 to 8, characterized in that, include: The module constructs a multimodal mapping relationship between the text symbols on the surface of the power tower and the tower components based on the obtained text symbols. The analysis module acquires the original reflected signal of the power tower through the radar array, filters the target reflection area based on the multimodal mapping relationship, analyzes the phase difference data in the original reflected signal, and generates a set of penetrating features containing the outline of the tower skeleton and the positioning information of the corresponding text symbols. The compensation module constructs a dynamic noise suppression model based on the penetration feature set, and performs attenuation compensation on the multipath reflection signal through the dynamic noise suppression model to generate a compensation feature set for electromagnetic interference resistance. The update module inputs the compensation feature set and the coating geometric identification features into the shared weight encoder for cross-modal association learning. The shared weight encoder fuses the semantic vectors of text symbols based on the attention mechanism, and simultaneously extracts the internal support structure features and the external anti-corrosion coating texture features of the tower body. The geometric constraint relationship between the support structure and the coating texture is established by alternately updating the convolution kernel parameters. The generation module reconstructs the missing tower connection components based on the geometric constraints, generating a three-dimensional recognition model of the power tower that is resistant to environmental interference.

Citation Information

Patent Citations

  • Multi-modal data fusion for enhanced 3D perception for platforms

    US20200184718A1

  • Three-dimensional target detection method based on multimodal fusion and depth attention mechanism

    US20250037299A1