Electric power iron tower identification method and system based on character symbols
By constructing a multimodal mapping relationship between text symbols and tower components, and combining radar arrays and dynamic noise suppression models, a three-dimensional recognition model of power towers that is resistant to environmental interference was reconstructed, solving the problems of low recognition accuracy and difficulty in implementation in existing technologies, and achieving efficient and low-cost power tower recognition.
Patent Information
- Application Number
- CN202510908933.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-02
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2045-07-02
AI Technical Summary
Existing power tower identification technology is sensitive to environmental factors, has low recognition accuracy, is difficult to implement, and is difficult to promote and apply in resource-limited areas.
A text-based power tower recognition method constructs a multimodal mapping relationship between text symbols and tower components, uses a radar array to obtain reflected signals, analyzes phase difference data, generates a set of penetrating features, builds a dynamic noise suppression model, and cross-modally associates and learns the texture features of the tower's internal support structure and external coating to reconstruct a three-dimensional recognition model that is resistant to environmental interference.
It achieves high-precision three-dimensional reconstruction in complex environments, enhances the interpretability and automation capabilities of component recognition, improves recognition accuracy and robustness, has strong adaptability, and reduces implementation costs.
Smart Images

Figure CN120689859A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of intelligent detection of electric power facilities, and in particular to a method and system for identifying electric power towers based on text symbols. Background Art
[0002] With the rapid development of the power industry, the number and types of power towers are increasing. Their safety and stability are directly related to the reliable operation of the power grid. Traditional manual inspection methods are inefficient and susceptible to environmental factors, making them difficult to meet the needs of modern power grid management. Therefore, developing a text-based power tower recognition method that can automatically, quickly, and accurately identify tower type and status has become a pressing technical challenge for the power industry.
[0003] Currently, power tower identification primarily relies on image processing and machine learning technologies. Common approaches include deep learning-based image classification and feature extraction-based pattern recognition. For example, convolutional neural networks (CNNs) are used to classify tower images, or to match towers by extracting features such as their shape and color. Furthermore, some research is attempting to combine GPS and remote sensing technologies to achieve tower location and identification.
[0004] However, while existing solutions have improved the automation level of tower identification to a certain extent, they still have many drawbacks. First, image processing methods are sensitive to environmental factors such as lighting and weather, making recognition accuracy difficult to guarantee. Second, machine learning methods require large amounts of labeled data for training, which results in high data acquisition and annotation costs. Furthermore, existing solutions often rely on complex algorithms and hardware, making implementation difficult and prohibitive for widespread application in resource-limited areas. Therefore, a more efficient, cost-effective, and adaptable method for identifying power towers is urgently needed. Summary of the Invention
[0005] The present application provides a method and system for identifying electric towers based on text symbols, which are used to solve the problems of insufficient resolution and lack of semantic relevance of the model in the existing technology.
[0006] In a first aspect, the present application provides a method for identifying power towers based on text symbols, comprising: According to the text symbols obtained on the surface of the power tower, a multimodal mapping relationship between the text symbols and the tower components is constructed; Acquire the original reflection signal of the power tower through the radar array, screen the target reflection area based on the multimodal mapping relationship, analyze the phase difference data in the original reflection signal, and generate a penetration feature set including the tower skeleton outline and corresponding text symbol positioning information; Building a dynamic noise suppression model based on the penetration feature set, performing attenuation compensation on multipath reflection signals through the dynamic noise suppression model, and generating a compensation feature set for resisting electromagnetic interference; The compensation feature set and the coating geometric identification feature are input into a shared weight encoder for cross-modal association learning, so that the shared weight encoder fuses the semantic vector of the text symbol based on the attention mechanism, synchronously extracts the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, and establishes the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; The missing tower connection components are reconstructed based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0007] Optionally, the compensation feature set and the coating geometric identification feature are input into a shared weight encoder, and feature extraction is performed on the compensation feature set and the coating geometric identification feature through a multi-layer network structure in the shared weight encoder to generate a preliminary feature vector; Performing weighted fusion of the preliminary feature vector and the text symbol semantic vector based on an attention mechanism to generate a fused feature vector; By performing multi-scale decomposition on the fused feature vector, the internal support structure features of the tower body and the texture features of the external anti-corrosion coating are extracted respectively, wherein the support structure features are extracted by high-frequency components and the coating texture features are extracted by low-frequency components; The convolution kernel parameters are alternately updated based on the local consistency optimization between the internal support structure features and the external anti-corrosion coating texture features, and a geometric constraint relationship between the support structure features and the coating texture features is established.
[0008] Optionally, the preliminary feature vector and the text symbol semantic vector are respectively input into independent fully connected layers for feature dimension alignment to generate aligned preliminary feature vector and text symbol semantic vector; Based on the spatial distribution of the aligned preliminary feature vector and the text symbol semantic vector, constructing local symbol association features and global structure association features, wherein the local symbol association features are generated by the feature mean within the symbol neighborhood, and the global structure association features are generated by the cross-region feature difference; Generate a symbol-structure association mapping vector by cross-projecting the local symbol association feature and the global structure association feature, and calculate a dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector based on the symbol-structure association mapping vector; Performing channel-level weight assignment on the aligned preliminary feature vectors according to the dynamic relevance scores, and performing symbol-level weight assignment on the text symbol semantic vectors to generate a channel-weighted preliminary feature vector and a symbol-weighted semantic vector; Performing multi-scale context feature splicing on the channel-weighted preliminary feature vector and the symbol-weighted semantic vector, and generating a multi-scale context feature splicing result, wherein the multi-scale context feature is extracted by sliding a context window of symbol semantics, and the window size is positively correlated with the density of text symbols on the surface of the power tower; Based on the topological relationship constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale context feature splicing result to generate a fused feature vector.
[0009] Optionally, the local symbol association features are respectively subjected to spatial position encoding and channel dimension encoding to generate spatially encoded local symbol association features and channel-encoded global structure association features; By interactively superimposing feature channels, the local symbol association features after spatial encoding and the global structure association features after channel encoding are bidirectionally projected to generate a symbol structure association mapping vector; Extracting a spatial weight matrix and a channel weight matrix from the symbol-structure association mapping vector, wherein the spatial weight matrix is generated based on the local similarity of symbol semantics, and the channel weight matrix is generated based on the global difference of structural features; By parameterizing and weighting the spatial topological constraints of the global structural association features, the spatial weight matrix and the channel weight matrix are fused to generate a dynamic correlation weight coefficient; Based on the dynamic relevance weight coefficient, the cross-modal association strength between the aligned preliminary feature vector and the text symbol semantic vector is quantified to generate a dynamic relevance score.
[0010] Optionally, the internal support structure characteristics are decomposed into connector spatial distribution characteristics and main beam morphology characteristics, and the external anti-corrosion coating texture characteristics are decomposed into corrosion area distribution characteristics and coating thickness variation characteristics; Based on the proximity constraints between bolt hole locations and coating defect areas, the spatial distribution characteristics of the connectors and the distribution characteristics of the corrosion areas are spatially topologically encoded to generate connector corrosion correlation features. Based on the synchronization constraints between the tower body curvature change and the coating thickness gradient, the main beam morphological characteristics and the coating thickness change characteristics are morphologically coupled and encoded to generate the main beam thickness coupling features. Constructing a multi-branch convolution kernel group, wherein the first branch performs spatial constraint modeling on the corrosion-related features of the connector, and the second branch performs texture constraint modeling on the thickness coupling features of the main beam; Based on the local consistency loss of the spatial constraint modeling results and the texture constraint modeling results, the parameters of the multi-branch convolution kernel group are alternately updated to generate a geometric constraint relationship between the support structure features and the coating texture features.
[0011] Optionally, a radar array deployed in a circular or linear manner is used to collect multi-angle signals from the power tower to obtain a set of original reflected signals, wherein the direction of the radar beam is aligned with the spatial distribution of text symbols on the tower surface during signal collection; Based on the multimodal mapping relationship, beamforming processing is performed on the original reflection signal set to screen the target reflection area containing the text symbol positioning information, wherein the beamforming parameters are dynamically adjusted according to the text symbol size and the geometric size of the tower body component; Performing phase difference interferometry processing on the original reflection signal of the target reflection area, analyzing the phase difference data between adjacent radar channels, and generating an initial three-dimensional point cloud of the tower skeleton outline; Based on the component association of the text symbols in the multimodal mapping relationship, the initial three-dimensional point cloud is spatially filtered to eliminate scattered points irrelevant to the positioning of the text symbols, thereby generating a symbol-constrained tower skeleton point cloud; The symbol-constrained tower skeleton point cloud is feature-fused with the multipath reflection intensity in the previous phase difference interferometry processing to generate a penetrating feature set containing text symbol positioning information and tower skeleton geometric parameters.
[0012] Optionally, based on the association between the text symbols and the tower components in the multimodal mapping relationship, the symbol semantically associated area and the non-associated area are divided; In the symbol semantic association area, the scattered points are processed according to the component association of the text symbol. In the non-symbol semantic association area, the scattered points that are not related to the corrosion expansion characteristics are eliminated according to the defect boundary distribution of the coating corrosion symbol. Marking the initial three-dimensional point cloud with symbol semantic association areas, wherein the lightning protection symbol area radiates outward from the center of the symbol, the support symbol area extends along the main beam axis, and the fixed symbol area extends according to the bolt distribution density matrix, to generate a regional marking point cloud; Based on the multi-scale precision parameters of the regional marker point cloud, the scattered points processed in the symbolic semantic associated area and the remaining scattered points after the non-associated area is eliminated are multi-scale aggregated to generate a symbol-constrained tower skeleton point cloud.
[0013] Optionally, symbol semantic noise suppression parameters are extracted through the spatial mapping relationship between the text symbol positioning information in the penetration feature set and the tower body components to construct a dynamic noise suppression model; In the dynamic noise suppression model, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower skeleton outline; Generate an interference source spatial weight map based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, integrate the interference source spatial weight map into the dynamic noise suppression model, perform azimuth-related attenuation compensation on the multipath reflection signal, and extract azimuth compensation parameters; The symbolic semantic noise suppression parameter, the structural reflection path attenuation coefficient and the azimuth compensation parameter are weightedly integrated in the tower region in the dynamic noise suppression model to update the dynamic noise suppression model; Iterative path attenuation compensation is performed on the multipath reflection signals in the penetration feature set by using the updated dynamic noise suppression model to generate a compensation feature set for resisting electromagnetic interference.
[0014] Optionally, based on the tower main beam curvature continuity constraint and the bolt hole position topological distribution constraint in the geometric constraint relationship, the area where the connection component is missing is spatially located to generate a bounding box of the missing area; Lock the range within the missing area boundary box, and generate candidate connection component parameters based on the angle constraint relationship between the main beam and the support member; Performing spatial matching verification on the candidate connection component parameters based on the bolt hole position topological distribution constraint, eliminating candidate parameters whose axial deviation from the bolt hole position exceeds a preset threshold, and generating a preliminary connection component set; Optimizing the shape of the preliminary set of connection components according to the curvature continuity constraint of the tower main beam, adjusting the installation angle and contact surface curvature of the connection components to meet the curvature difference threshold of the main beam skeleton point cloud; Based on the load distribution characteristics of the support members and the tower frame, an anti-environmental interference reinforcement model is performed on the connection components after the shape optimization, wherein the reinforcement strength is dynamically adjusted according to the connection density between the support members and the main beam; The reinforced modeled connection components are matched with the tower skeleton point cloud through geometric constraints to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0015] In a second aspect, the present application provides a text-based power tower identification system, comprising: A construction module, which constructs a multimodal mapping relationship between text symbols and tower components based on the text symbols obtained on the surface of the power tower; An analysis module acquires the original reflection signal of the power tower through a radar array, screens the target reflection area based on the multimodal mapping relationship, analyzes the phase difference data in the original reflection signal, and generates a penetration feature set including the tower skeleton outline and the corresponding text symbol positioning information; a compensation module, constructing a dynamic noise suppression model based on the penetration feature set, performing attenuation compensation on multipath reflection signals through the dynamic noise suppression model, and generating a compensation feature set for resisting electromagnetic interference; An updating module inputs the compensation feature set and the coating geometric identification feature into a shared weight encoder for cross-modal association learning, so as to fuse the semantic vector of the text symbol through the shared weight encoder based on the attention mechanism, synchronously extract the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, and establish the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; A generation module is provided for reconstructing the missing tower connection components based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0016] In an embodiment of the present application, a multimodal mapping relationship between text symbols and tower components is constructed based on the text symbols obtained on the surface of the power tower; the original reflection signal of the power tower is obtained through a radar array, the target reflection area is screened based on the multimodal mapping relationship, the phase difference data in the original reflection signal is analyzed, and a penetration feature set including the tower skeleton outline and the corresponding text symbol positioning information is generated; a dynamic noise suppression model is constructed based on the penetration feature set, and the multipath reflection signal is attenuated and compensated by the dynamic noise suppression model to generate a compensation feature set for electromagnetic interference resistance; the compensation feature set and the coating geometric identification feature are input into a shared weight encoder for cross-modal association learning, so as to fuse the text symbol semantic vector based on the attention mechanism through the shared weight encoder, synchronously extract the internal support structure features of the tower body and the external anti-corrosion coating texture features, and establish a geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; the missing tower connection components are reconstructed based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0017] The technical solution of this application has the following beneficial effects: This application achieves high-precision 3D reconstruction in complex environments through multimodal data fusion and anti-interference mechanisms. First, based on the mapping relationship between text symbols on the surface of power towers and internal components, semantic identifiers are digitally associated with the physical structure, significantly enhancing the interpretability and automation of component recognition. Second, radar array penetrating detection technology is used to analyze phase difference data, combined with the spatial positioning of text symbols to accurately reconstruct the tower skeleton outline, breaking through the reliance of traditional visual methods on surface coatings. To address electromagnetic interference and multipath reflection, a dynamic noise suppression model uses an adaptive compensation mechanism to improve signal quality and provide high-fidelity input for feature fusion. A shared weight encoder is used to cross-modally associate radar features with coating geometric identifiers. An attention mechanism integrates text semantics and structural features to establish geometric topological constraints between the support system and the anti-corrosion coating, significantly improving the accuracy of internal and external feature matching. Finally, an interference-resistant 3D recognition model is generated based on multimodal complementary features. Geometric constraints enable generative repair of missing components, which is robust even in harsh environments such as rain, fog, and rust. This completes a full-link solution from physical detection to semantic analysis, providing reliable technical support for intelligent inspection of power facilities.
[0018] These and other aspects of the present application will become more readily apparent from the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0020] Figure 1 A flowchart of a method for identifying electric towers based on text symbols provided by the present application is shown; Figure 2 The schematic diagram shows the structure of a power tower identification system based on text symbols provided by the present application. DETAILED DESCRIPTION
[0021] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.
[0022] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this document or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this document are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to being different types.
[0023] This application aims to develop a system for intelligent identification and 3D reconstruction of power towers based on multimodal semantic fusion, achieving a technological breakthrough in the entire process, from text symbol parsing to interference-resistant 3D modeling. By fusing text symbol semantics with radar penetration detection data, the system constructs a multimodal mapping relationship, enabling a digital association between the semantics of tower components and their physical structure. Ultimately, based on multimodal feature complementarity and a generative repair algorithm, it achieves intelligent reconstruction of missing connected components in complex environmental interference, significantly improving the accuracy of power tower structure identification and the environmental robustness of 3D models, providing a highly reliable digital solution for intelligent inspection of power grid facilities.
[0024] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of this application.
[0025] Figure 1 A flowchart of a method for identifying a power tower based on text symbols is provided for an embodiment of the present application. Figure 1 As shown, the method includes: 101. Based on the acquired text symbols on the surface of the power tower, a multimodal mapping relationship between the text symbols and the tower components is constructed; In this step, text symbols refer to the identification marked on the surface of the power tower, such as numbers, models or component names; tower components are the physical components of the tower, such as crossarms, tower bodies, insulators, etc.
[0026] Tower components refer to the physical components of a power tower, such as cross arms, tower body, insulators, ground wire supports, etc., which are the basic units that make up the tower.
[0027] The multimodal mapping relationship refers to establishing an association between text symbols and the spatial position and semantic information of tower components, thereby forming a corresponding relationship between text symbols and components.
[0028] In an embodiment of the present application, first, a photoelectric character recognition sensor array is used to capture the preset text symbol logos on the surface of the power tower, and a multi-scale semantic segmentation algorithm is used to analyze the morphological characteristics and spatial distribution patterns of the symbols. The recognition results are encoded into structured semantic vectors, and a multimodal mapping relationship between the text symbols and the tower components is constructed. The system uses a high-resolution camera or drone to capture images of the surface of the power tower and uses optical character recognition technology to extract the position and semantic information of the text symbols. Then, based on the semantic information of the text symbols, the system associates them with the tower components and records their physical location and structural characteristics. Finally, the system stores the multimodal mapping relationship as structured data for subsequent radar signal analysis.
[0029] During a power tower inspection, the system used a drone to capture images of the tower's surface and identified the text symbol "Support Rod B." The system then associated "Support Rod B" with the tower's support rod components, recording their location, dimensions, and material. Using a deep learning model, the system analyzed the structural characteristics of the support rods and associated them with the text symbol "Support Rod B," establishing a multimodal mapping between the text symbol and the tower components. This step provided the foundation for accurate text-to-component mapping for subsequent radar signal analysis.
[0030] 102. Acquire the original reflection signal of the power tower through the radar array, screen the target reflection area based on the multimodal mapping relationship, analyze the phase difference data in the original reflection signal, and generate a penetration feature set including the tower skeleton outline and corresponding text symbol positioning information; In this step, the radar array refers to an array of devices used to transmit and receive radar signals, usually composed of multiple radar sensors that can detect the position, shape, and motion state of an object.
[0031] The original reflected signal refers to the electromagnetic wave signal reflected by the power tower and received by the radar array, which contains the location, shape and material information of the target object.
[0032] Phase difference data refers to the phase difference of radar signals between different receiving points and is typically used to calculate the outline and position of an object. Phase difference data can reflect the geometric characteristics and spatial distribution of the target object.
[0033] The penetrability feature set is a feature set generated by radar signal analysis that includes the tower skeleton outline and text symbol positioning information. This feature set can reflect the internal structure and external identification information of the tower.
[0034] In the embodiment of the present application, the system first acquires the original reflection signal from the power tower through a radar array. The system then uses Doppler radar technology to collect electromagnetic wave reflection data from the tower. Next, the system selects the target reflection area based on a multimodal mapping relationship. The system then uses orthogonal polarization interference mode to analyze the phase difference data in the original reflection signal to generate the tower skeleton outline and text symbol positioning information. The system calculates the outline and position of the crossbeam component using a phase difference algorithm and binds it to the text symbol. Finally, the system stores the penetration feature set as structured data for subsequent noise suppression.
[0035] In a power tower inspection scenario, the system acquires the original reflected signal from the power tower through a radar array. Based on a multimodal mapping relationship, the system locates the reflection area corresponding to the text symbol "Support Pole B" and extracts the radar signal in that area. Next, the system analyzes the phase difference data in the original reflected signal to generate the support pole's skeleton outline and text symbol location information. The system calculates the geometric shape and position of the support pole using a phase difference algorithm and binds it to the text symbol "Support Pole B." Finally, the system stores the penetration feature set as structured data for subsequent noise suppression. This step provides an accurate penetration feature set for subsequent noise suppression.
[0036] 103. Constructing a dynamic noise suppression model based on the penetration feature set, performing attenuation compensation on multipath reflection signals using the dynamic noise suppression model, and generating a compensation feature set for resisting electromagnetic interference; In this step, the dynamic noise suppression model is used to suppress noise in radar signals, usually based on signal processing technology. This model can dynamically identify and eliminate multipath reflection signals and other interference.
[0037] Multipath reflection signals refer to interference signals generated by multiple reflections of radar signals during propagation, which usually lead to distortion of the target object's outline and position information.
[0038] The compensation feature set refers to a feature set without noise interference generated by a dynamic noise suppression model, which can reflect the true contour and position information of the target object.
[0039] In the embodiment of the present application, the system first constructs a dynamic noise suppression model based on a set of penetration features. The system uses an adaptive filtering algorithm to identify and suppress multipath reflection signals. Next, the system performs attenuation compensation on the multipath reflection signals based on the correlation between the tower material's dielectric constant and the millimeter-wave scattering path length, generating a set of compensation features that resist electromagnetic interference. The system uses a signal reconstruction algorithm to eliminate noise interference in the support rod area and generate undistorted support rod profile information. Finally, the system stores the compensation feature set as structured data for subsequent cross-modal association learning.
[0040] When other metal structures are near power towers, the system uses a dynamic noise suppression model to eliminate noise interference from multipath reflection signals. For example, electromagnetic wave signals reflected from the ground are identified as noise and attenuated and compensated for using a filtering algorithm. Ultimately, the system generates a set of compensated features to ensure the accuracy of the crossarm skeleton profile information. This data provides a reliable foundation for subsequent cross-modal association learning. In a practical application, the system also identified multipath reflection signals in the tower area and used the dynamic noise suppression model to eliminate interference and generate accurate tower outline information. This data provides important support for subsequent 3D reconstruction.
[0041] 104. Input the compensation feature set and the coating geometric identification feature into a shared weight encoder for cross-modal association learning, so as to fuse the text symbol semantic vectors through the shared weight encoder based on the attention mechanism, synchronously extract the internal support structure features of the tower body and the external anti-corrosion coating texture features, and establish the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; In this step, the coating geometric identification features refer to the geometric shape and identification features of the coating on the surface of the power tower, which are usually used to evaluate the integrity and anti-corrosion performance of the coating.
[0042] A shared weight encoder is a neural network encoder used for cross-modal data association learning, typically implemented based on an attention mechanism. This encoder can fuse data from different modalities to generate a unified feature representation.
[0043] The geometric constraint relationship refers to the geometric correlation between the internal supporting structure of the tower and the texture of the external anti-corrosion coating, which can reflect the integrity of the tower structure and the distribution characteristics of the coating.
[0044] In the example of the present application, first, the compensation feature set and the coating geometric identification feature are input into a shared weight encoder for cross-modal association learning, so as to fuse the text symbol semantic vector based on the attention mechanism through the shared weight encoder, and then, the attention mechanism is used to dynamically adjust the model's attention to the internal support structure features of the tower body and the external anti-corrosion coating texture features. For example, when the coating texture information is relatively vague, the model will pay more attention to the support structure features. Then, by alternately updating the convolution kernel parameters, a geometric constraint relationship between the support structure and the coating texture is established. For example, the spatial distance and relative position between the coating texture and the support structure are determined to ensure the consistency of the two in space. Finally, feature data containing the association relationship between the support structure and the coating texture is generated to provide a basis for three-dimensional reconstruction.
[0045] During power tower inspections, the system uses a shared weight encoder to fuse the crossarm skeleton outline and coating texture features. For example, the system determines the spatial distance between the crossarm skeleton and the coating to be 0.1 meter and establishes a geometric constraint relationship between the two. This information provides critical support for subsequent 3D reconstruction. In one practical application, the system also identified the coating texture in the tower area and, through cross-modal association learning, determined its spatial relationship with the tower skeleton, further refining the geometric constraint relationship. This data provides a crucial basis for the final 3D recognition model.
[0046] 105. Reconstruct the missing tower connection components based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0047] In this step, the missing tower connection parts refer to parts that are missing due to rust or external damage, such as missing bolts, broken insulators, etc.
[0048] The environmentally resistant 3D recognition model of power towers refers to a building information model that integrates multimodal data and supports dynamic updates such as corrosion diffusion simulation and anti-interference rendering such as removing rain and fog noise.
[0049] In the example of the present application, first, a multimodal fusion decision mechanism is triggered based on meteorological sensor data. When the rain and fog concentration exceeds a preset threshold, the internal supporting structure features of the tower body, the texture features of the external anti-corrosion coating, and the text symbol positioning information are spatially topologically calibrated. Based on the geometric constraint relationship, the missing parts of the tower body connection components are identified. Using three-dimensional reconstruction technology, the missing tower body connection components are reconstructed in combination with the tower body skeleton outline and coating texture information. At the same time, the environmental noise suppression algorithm is used to eliminate the impact of environmental interference on the model. For example, natural factors such as wind and rain may affect the accuracy of the three-dimensional model. The system eliminates these interferences through filtering algorithms and signal enhancement technology. Ultimately, a three-dimensional recognition model of the power tower that is resistant to environmental interference is generated, providing reliable technical support for inspection and maintenance.
[0050] When a cross-arm connection component on a power tower is partially missing, the system uses 3D reconstruction technology to restore its complete structure. The system reconstructs the length and shape of the cross-arm connection based on geometric constraints and, combined with a noise suppression algorithm, generates a 3D recognition model that is resistant to environmental interference. This model provides inspectors with accurate tower structural information, improving inspection efficiency. In one application, the system also identified a missing section of a tower connection, restored its structure through 3D reconstruction, and, combined with an environmental noise suppression algorithm, generated a complete 3D recognition model. This data provides crucial support for the maintenance and repair of power towers.
[0051] In summary, steps 101 to 105 achieve a multimodal mapping between text symbols on the surface of a power tower and its components. Radar arrays and signal processing techniques are used to generate a penetrating feature set, and a dynamic noise suppression model is used to eliminate noise interference in multipath reflection signals. Geometric constraints are established through cross-modal association learning, and missing tower connection components are reconstructed, ultimately generating a three-dimensional recognition model of the power tower that is resistant to environmental interference. This model can accurately identify the tower structure, providing reliable technical support for the inspection and maintenance of power towers and significantly improving the safety and stability of the power system. For example, during an actual inspection, the system successfully identified and reconstructed a crossarm connection component that was missing due to obstruction, providing maintenance personnel with accurate structural information and avoiding potential safety hazards.
[0052] In order to solve the problem of cross-modal association learning between the internal support structure of the tower body and the texture features of the external anti-corrosion coating, in some embodiments, the compensation feature set and the coating geometric identification features are input into a shared weight encoder for cross-modal association learning in step 103, so that the shared weight encoder fuses the semantic vectors of text symbols based on the attention mechanism, synchronously extracts the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, and establishes a geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters, including: 201. Input the compensation feature set and the coating geometric identification feature into a shared weight encoder, perform feature extraction on the compensation feature set and the coating geometric identification feature through a multi-layer network structure in the shared weight encoder, and generate a preliminary feature vector; In step 201, the compensation feature set refers to feature data related to coating geometric deviations collected by sensors, including physical property data such as coating thickness, surface roughness, and coating adhesion. The coating geometric identification features refer to the coating geometric shape features extracted by image processing technology, such as coating edges, contours, surface curvature, etc. The shared weight encoder refers to a multi-layer neural network structure whose weights are shared between the inputs of the compensation feature set and the coating geometric identification features, and is used to extract common features between the two. The preliminary feature vector refers to the fused feature representation extracted by the shared weight encoder, which is used for subsequent feature fusion and decomposition.
[0053] In an embodiment of the present application, first, the compensation feature set and the coating geometric identification feature are respectively input into a shared weight encoder. The shared weight encoder adopts a multi-layer fully connected network structure, and each layer transforms the input features through a nonlinear activation function. Due to weight sharing, the encoder can simultaneously learn the common information of the compensation features and the coating geometric identification features, avoiding the information redundancy of the two in the feature extraction process. In a specific implementation, the input layer of the encoder splices the compensation feature set and the coating geometric identification feature, and then gradually extracts high-order features through multiple hidden layers. Finally, the final layer output of the encoder generates a preliminary feature vector, which contains comprehensive information of the compensation features and the coating geometric identification features, laying the foundation for subsequent feature fusion.
[0054] 202. Perform weighted fusion on the preliminary feature vector and the text symbol semantic vector based on an attention mechanism to generate a fused feature vector; In step 202, the attention mechanism is a technique that achieves feature fusion by assigning different weights to different features, dynamically adjusting weights based on feature importance. The text symbol semantic vector refers to a textual descriptive feature related to the coating, generated using natural language processing techniques, such as "uneven coating thickness" and "excessive surface roughness." The fused feature vector is a feature representation generated by weighted fusion of the preliminary feature vector and the text symbol semantic vector using the attention mechanism, and is used for subsequent multi-scale decomposition.
[0055] In an embodiment of the present application, the preliminary feature vector and the text symbol semantic vector are first input into the attention mechanism module. The attention mechanism obtains the weight coefficient of each feature by calculating the similarity matrix between the two. In the specific implementation, the self-attention mechanism is used to interactively calculate the preliminary feature vector and the text symbol semantic vector to generate an attention weight matrix. Then, the preliminary feature vector and the text symbol semantic vector are weighted and fused according to the weight coefficient to generate a fused feature vector. This vector not only contains the compensation features and the coating geometric identification features, but also incorporates the text description information, thereby enhancing the semantic expression ability of the features. For example, when the text description mentions "uneven coating thickness", the attention mechanism will give thickness-related features a higher weight, thereby improving the pertinence of feature fusion.
[0056] 203. Perform multi-scale decomposition on the fused feature vector to extract the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, respectively, wherein the support structure features are extracted by high-frequency components, and the coating texture features are extracted by low-frequency components; In step 203, multiscale decomposition refers to a technique that performs multiscale analysis on feature vectors using wavelet transforms or convolutional neural networks, which can simultaneously capture both high-frequency and low-frequency information. The tower's internal support structure features are those related to the support structure extracted through high-frequency components, such as the geometry, connection method, and stress distribution of the support rods. The external anti-corrosion coating texture features are those related to the coating's surface texture extracted through low-frequency components, such as the coating's particle distribution, color uniformity, and surface smoothness.
[0057] In an embodiment of the present application, the fused feature vector is first subjected to multi-scale decomposition, and the feature vector is decomposed into high-frequency components and low-frequency components using wavelet transform or convolutional neural network. In a specific implementation, the fused feature vector is decomposed using discrete wavelet transform to obtain high-frequency components and low-frequency components. The high-frequency component usually contains detailed information of the supporting structure, so the internal supporting structure characteristics of the tower body are extracted through the high-frequency component. The low-frequency component contains the overall texture information of the coating, so the texture characteristics of the external anti-corrosion coating are extracted through the low-frequency component. For example, in the detection scenario of a wind power tower, the high-frequency component can capture tiny cracks or deformations of the support rod, while the low-frequency component can reflect the particle distribution and color uniformity of the coating surface. Ultimately, the support structure characteristics and coating texture characteristics are obtained respectively, which provide a basis for subsequent optimization and constraints.
[0058] 204. Optimize and alternately update convolution kernel parameters based on the local consistency between the internal support structure features and the external anti-corrosion coating texture features, and establish a geometric constraint relationship between the support structure features and the coating texture features.
[0059] In step 204, the internal support structure features of the tower body refer to features related to the support structure extracted through high-frequency components, such as the geometric shape of the support rod, the connection method, the stress distribution, etc. The external anti-corrosion coating texture features refer to features related to the coating surface texture extracted through low-frequency components, such as the particle distribution, color uniformity, surface smoothness, etc. of the coating. Local consistency optimization refers to a technology that achieves feature smoothing by constraining parameter changes in adjacent areas, which can improve the continuity and consistency of features. Alternating updating of convolution kernel parameters refers to the process of improving feature extraction effects by iteratively optimizing convolution kernel parameters. Geometric constraint relationship refers to the establishment of a spatial relationship between support structure features and coating texture features through mathematical modeling to ensure the geometric coordination of the two.
[0060] In an embodiment of the present application, the convolution kernel parameters are first optimized based on the local consistency of the internal support structure features and the external anti-corrosion coating texture features. In a specific implementation, the alternating direction multiplier method is used to iteratively optimize the convolution kernel parameters so that the extracted features remain consistent within the local area. For example, in the detection scenario of a wind power tower, the convolution kernel parameters are optimized to ensure that the geometric features of the support rod and the coating texture features are consistent in space. Then, by establishing a geometric constraint relationship between the support structure features and the coating texture features, the spatial coordination of the two is ensured. In a specific implementation, a mathematical modeling method is used to establish the geometric relationship between the two, such as the distribution relationship between the geometric shape of the support rod and the coating thickness. Ultimately, the optimized support structure features and coating texture features are obtained, which provide a high-precision feature representation for coating detection.
[0061] Here's a specific example: In the coating inspection scenario of a wind turbine tower, sensors first collect a set of compensating features of the tower coating, such as thickness, roughness, and adhesion. Image processing techniques are then used to extract geometric features of the coating, such as edges, contours, and surface curvature. These features are then input into a shared weight encoder to generate a preliminary feature vector. Next, an attention mechanism is used to weightedly fuse this preliminary feature vector with semantic vectors of text symbols, such as "uneven coating thickness" and "excessive surface roughness," to generate a fused feature vector. Multi-scale decomposition is then used to extract features of the tower's internal support structure, such as the shape and connection of the support rods, and texture features of the external anti-corrosion coating, such as the particle distribution and color uniformity of the coating. Finally, local consistency optimization is used to alternately update the convolution kernel parameters and establish a geometric constraint relationship between the support structure features and the coating texture features, ensuring the accuracy and consistency of the inspection results.
[0062] In summary, steps 201 to 204 achieve efficient fusion of compensation features, coating geometric identification features, and text semantic features, enhancing the expressive power of features. Multi-scale decomposition and local consistency optimization enable precise extraction of support structure features and coating texture features, establishing a geometric constraint relationship between the two, and improving the accuracy and robustness of coating inspection. This method not only effectively identifies geometric deviations and texture defects in coatings but also ensures spatial coordination between the support structure and coating, providing reliable technical support for coating inspection in complex industrial scenarios.
[0063] In order to solve the problem that the preliminary feature vector and the text symbol semantic vector are difficult to be effectively fused, in some embodiments, the weighted fusion of the preliminary feature vector and the text symbol semantic vector based on the attention mechanism in step 202 to generate a fused feature vector includes: 301. Input the preliminary feature vector and the text symbol semantic vector into independent fully connected layers respectively to align feature dimensions, thereby generating aligned preliminary feature vector and text symbol semantic vector; In step 301, the preliminary feature vector refers to the fused feature representation extracted by the shared weight encoder, which contains multi-level visual information of the surface image of the power tower, such as texture, edge, shape, etc. The text symbol semantic vector refers to the semantic features related to the text symbols on the surface of the power tower generated by natural language processing technology, which captures the semantic meaning and contextual relationship of the text symbols. The fully connected layer refers to a neural network layer used to map input features to the target dimension, which realizes feature dimensionality adjustment through linear transformation and nonlinear activation function. Feature dimension alignment refers to adjusting the dimensions of the preliminary feature vector and the text symbol semantic vector to the same size through the fully connected layer to facilitate subsequent feature fusion and association analysis.
[0064] In an embodiment of the present application, the preliminary feature vector and the text symbol semantic vector are first input into independent fully connected layers for feature dimension alignment. In the specific implementation, the output dimension of the fully connected layer is set according to the larger dimension of the two to ensure the integrity and consistency of the feature information. Secondly, the preliminary feature vector and the text symbol semantic vector are mapped to the same dimensional space through the linear transformation and nonlinear activation function of the fully connected layer. Finally, the aligned preliminary feature vector and text symbol semantic vector are generated to provide a basis for subsequent feature association analysis. This step ensures the dimensional consistency of the preliminary feature vector and the text symbol semantic vector, laying a technical foundation for subsequent feature fusion and association analysis.
[0065] 302. Based on the spatial distribution of the aligned preliminary feature vector and the text symbol semantic vector, construct a local symbol association feature and a global structure association feature, wherein the local symbol association feature is generated by averaging the features within the symbol neighborhood, and the global structure association feature is generated by differencing the features across regions; In step 302, the local symbol association feature refers to the local association information generated by the feature mean within the symbol neighborhood, which captures the local relationship between the text symbol and its neighboring area, and reflects the spatial distribution and feature changes of the text symbol within the local range. The global structure association feature refers to the global association information generated by cross-regional feature differentiation, which captures the global relationship between the text symbol and the overall structure, and reflects the position and feature differences of the text symbol in the overall structure. Spatial distribution refers to the distribution of the aligned preliminary feature vector and the text symbol semantic vector in the feature space, which reflects the spatial relationship between the text symbol and the preliminary feature through the numerical distribution of the feature vector.
[0066] In the embodiment of the present application, first, based on the spatial distribution of the aligned preliminary feature vector and the semantic vector of the text symbol, local symbol association features and global structure association features are constructed. Secondly, the local symbol association features are generated by calculating the feature mean in the neighborhood of the symbol. In a specific implementation, with each text symbol as the center, the feature mean in the surrounding area is calculated, thereby capturing the local relationship between the text symbol and its neighboring area. Global structure association features are generated by cross-region feature differences. In a specific implementation, the feature difference between the area where the text symbol is located and the overall structure area is calculated, thereby capturing the global relationship between the text symbol and the overall structure. Finally, local symbol association features and global structure association features are generated to provide input for subsequent cross-projection. This step comprehensively reflects the spatial relationship between the text symbol and the preliminary features through the construction of local and global features, providing rich information support for subsequent feature fusion.
[0067] 303. Generate a symbol-structure association mapping vector by cross-projecting the local symbol association feature and the global structure association feature, and calculate a dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector based on the symbol-structure association mapping vector; In step 303, cross-projection refers to a technique for mapping local symbol-related features and global structure-related features to the same space through matrix operations, which achieves feature fusion through linear transformation and nonlinear activation function. The symbol-structure-related mapping vector refers to a fusion feature representation generated by cross-projection, which contains comprehensive information about local symbol-related features and global structure-related features. The dynamic correlation score refers to a score obtained by calculating the similarity between the symbol-structure-related mapping vector and the aligned preliminary feature vector and the text symbol semantic vector. It is used to measure the correlation between features and reflects the dynamic correlation relationship between the preliminary feature vector and the text symbol semantic vector.
[0068] In an embodiment of the present application, the local symbol association features and the global structure association features are first cross-projected to generate a symbol structure association mapping vector. In a specific implementation, matrix multiplication is used to map the two to the same space, and normalization is performed to ensure numerical stability. Secondly, the dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector is calculated based on the symbol structure association mapping vector. In a specific implementation, the correlation between the two and the symbol structure association mapping vector is calculated by cosine similarity to obtain a dynamic correlation score. Finally, the symbol structure association mapping vector and the dynamic correlation score are generated to provide a basis for subsequent weight allocation. This step realizes the comprehensive analysis of local and global features through cross-projection and the calculation of dynamic correlation scores, providing a scientific basis for subsequent weight allocation.
[0069] 304. Perform channel-level weight allocation on the aligned preliminary feature vectors according to the dynamic correlation scores, and perform symbol-level weight allocation on the text symbol semantic vectors to generate a channel-weighted preliminary feature vector and a symbol-weighted semantic vector. In step 304, channel-level weight allocation refers to the process of assigning different weights to each channel of the preliminary feature vector based on the dynamic relevance score. This process emphasizes the feature channels with a high correlation with the text symbol semantic vector through weight adjustment. Symbol-level weight allocation refers to the process of assigning different weights to each symbol of the text symbol semantic vector based on the dynamic relevance score. This process emphasizes the text symbols with a high correlation with the preliminary feature vector through weight adjustment. The channel-weighted preliminary feature vector and the symbol-weighted semantic vector refer to the feature representations after weight allocation, respectively. The weight adjustment optimizes the feature expression effect.
[0070] In an embodiment of the present application, first, channel-level weights are assigned to the aligned preliminary feature vectors based on the dynamic correlation scores. In a specific implementation, the dynamic correlation scores are converted into weight values, and each channel of the preliminary feature vector is weighted. Secondly, symbol-level weights are assigned to the text symbol semantic vectors based on the dynamic correlation scores. In a specific implementation, different weights are assigned to the semantic features of each text symbol. Finally, a channel-weighted preliminary feature vector and a symbol-weighted semantic vector are generated to provide input for subsequent multi-scale context feature splicing. This step optimizes the expression effect of the features through weight distribution, and provides high-quality input for subsequent feature fusion.
[0071] 305. Perform multi-scale context feature splicing on the channel-weighted preliminary feature vector and the symbol-weighted semantic vector, and generate a multi-scale context feature splicing result, wherein the multi-scale context feature is extracted by sliding a context window of symbol semantics, and the window size is positively correlated with the density of text symbols on the surface of the power tower; In step 305, multi-scale context feature splicing refers to the process of fusing the channel-weighted preliminary feature vector and the symbol-weighted semantic vector at different scales, which captures the multi-level relationship between the text symbols and the preliminary features through multi-scale analysis; context window sliding refers to extracting the context features of the symbol semantics through a sliding window, which realizes the extraction of multi-scale features through window sliding; the multi-scale context feature splicing result refers to the multi-scale fusion feature representation generated by splicing, which contains multi-level comprehensive information of the text symbols and the preliminary features.
[0072] In an embodiment of the present application, the channel-weighted preliminary feature vector and the symbol-weighted semantic vector are first subjected to multi-scale context feature splicing. In a specific implementation, the context features of the symbol semantics are extracted by sliding the context window, and the window size is dynamically adjusted according to the density of text symbols on the surface of the power tower. Secondly, the extracted multi-scale context features are spliced to generate a multi-scale context feature splicing result. Finally, input is provided for subsequent two-way feature propagation. This step realizes the multi-level fusion of text symbols and preliminary features through multi-scale context feature splicing, providing rich feature representation for subsequent feature propagation.
[0073] 306. Based on the topological relationship constraints of the power tower skeleton, perform bidirectional feature propagation on the multi-scale context feature splicing result to generate a fused feature vector.
[0074] In step 306, the power tower skeleton topology relationship constraint refers to the constraint on feature propagation based on the skeleton structure of the power tower, which ensures the rationality and consistency of feature propagation through topological relationships; bidirectional feature propagation refers to the process of propagating the multi-scale context feature splicing results under the skeleton topology relationship constraint through a bidirectional graph neural network, which realizes the global optimization of features through bidirectional propagation; the fused feature vector refers to the final feature representation generated by bidirectional feature propagation, which contains the comprehensive information of the text symbol semantics and the preliminary feature vector.
[0075] In an embodiment of the present application, first, based on the topological relationship constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale context feature splicing results. In the specific implementation, a bidirectional graph neural network is used to propagate the multi-scale context features under the skeleton topological relationship constraints. Secondly, the skeleton structure of the power tower is modeled as a graph, with nodes representing skeleton key points and edges representing skeleton connection relationships. Finally, the multi-scale context features are bidirectionally propagated in the graph through a bidirectional graph neural network to generate a fused feature vector. This step optimizes the global expression of features through bidirectional feature propagation, and provides high-precision feature representation for the recognition of text symbols on the surface of the power tower.
[0076] Here's a specific example: During a power tower inspection, a high-resolution camera captures the tower surface image. A shared weight encoder is used to extract a preliminary feature vector. Simultaneously, a BERT model performs semantic analysis on the tower surface text symbols (such as the number "T-102" and the warning "High Voltage Danger") to generate a text symbol semantic vector. Both are fed into separate fully connected layers for feature dimension alignment to ensure dimensionality consistency. Next, the feature mean of a 3×3 region around the symbol "T-102" is calculated to generate a local symbol association feature. The feature difference between the symbol region and the overall tower structure is calculated to generate a global structure association feature. The local and global features are then cross-projected to generate a symbol-structure association mapping vector. The dynamic correlation score between the mapping vector and the preliminary feature vector and the text symbol semantic vector is calculated. For example, the correlation score between "T-102" and the preliminary feature vector is 0.82. Based on the correlation scores, channel-level weights are assigned to the preliminary feature vector and symbol-level weights are assigned to the text symbol semantic vector to highlight important features. The weighted feature vectors were then subjected to multi-scale contextual feature concatenation: a 5×5 window was used to extract features in the numbered area, and a 3×3 window was used in the warning area to generate the multi-scale contextual feature concatenation result. Finally, based on the tower's skeleton topology (such as crossarms, tower body, and tower foot), a bidirectional graph neural network was used to propagate the multi-scale features in both directions to generate a fused feature vector.
[0077] In summary, through steps 301 to 306, the efficient fusion of the preliminary feature vector and the semantic vector of the text symbol is achieved; by constructing local symbol association features and global structure association features, the association relationship between the text symbol and its neighboring area and the overall structure is captured; through multi-scale context feature splicing and bidirectional feature propagation, the feature extraction and propagation process is optimized, and finally a high-precision fused feature vector is generated, providing reliable technical support for the recognition of text symbols on the surface of power towers.
[0078] In order to solve the problem that it is difficult to efficiently fuse local symbol association features and global structure association features and that dynamic correlation is difficult to quantify, in some embodiments, step 303 generates a symbol structure association mapping vector by cross-projecting the local symbol association features and the global structure association features, and calculates the dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector based on the symbol structure association mapping vector, including: 401. Perform spatial position encoding and channel dimension encoding on the local symbol association features to generate spatially encoded local symbol association features and channel-encoded global structure association features; In step 401, local symbol-related features refer to local features related to symbols extracted from an image. Spatial position encoding is used to capture the spatial distribution of symbols in the image, and channel dimension encoding is used to capture the global structural information of symbols across feature channels. Spatially encoded local symbol-related features refer to local features processed through spatial position encoding, and channel-coded global structural association features refer to global features processed through channel dimension encoding.
[0079] In an embodiment of the present application, the local symbol association features are first spatially encoded, and the spatial coordinate information of the symbol is embedded into the features using a position encoder (such as a sinusoidal encoder) to generate spatially encoded local symbol association features. Next, the local symbol association features are channel-dimensionally encoded, and the feature channels are globally modeled using a fully connected layer or a convolutional layer to generate channel-encoded global structural association features. Ultimately, the spatially encoded local symbol association features and the channel-encoded global structural association features are obtained, respectively.
[0080] 402. By interactively superimposing feature channels, bidirectionally project the local symbol association features after spatial encoding and the global structure association features after channel encoding to generate a symbol-structure association mapping vector; In step 402, the interactive superposition of feature channels refers to the fusion of local symbol-related features after spatial encoding and global structure-related features after channel encoding in the channel dimension. Bidirectional projection refers to mapping these two features into the same space using a projection matrix to generate a symbol-structure-related mapping vector. The symbol-structure-related mapping vector is a feature vector that contains both the local semantics and global structural information of the symbol.
[0081] In the embodiment of the present application, the local symbol association features after spatial encoding and the global structure association features after channel encoding are first spliced or weighted summed in the channel dimension to achieve interactive superposition of feature channels. Then, the two features are mapped to the same space through a bidirectional projection matrix, such as a learnable fully connected layer, to generate a symbol structure association mapping vector. Ultimately, a symbol structure association mapping vector containing the local semantics and global structural information of the symbol is obtained.
[0082] 403. Extracting a spatial weight matrix and a channel weight matrix from the symbol structure association mapping vector, wherein the spatial weight matrix is generated based on the local similarity of symbol semantics, and the channel weight matrix is generated based on the global difference of structural features; In step 403, the spatial weight matrix is used to quantify the local spatial similarity of symbols, and the channel weight matrix is used to quantify the global structural differences of feature channels. Local semantic similarity of symbols refers to the degree of semantic association between symbols in spatial positions, while global structural differences refer to the distribution differences of feature channels in the global structure.
[0083] In the embodiment of the present application, first, a spatial weight matrix is extracted from the symbol structure association mapping vector, and a convolutional layer or self-attention mechanism is used to calculate the local similarity of the symbols in space to generate a spatial weight matrix. Next, a channel weight matrix is extracted from the symbol structure association mapping vector, and a global average pooling or fully connected layer is used to calculate the difference in the global structure of the feature channels to generate a channel weight matrix. Finally, a spatial weight matrix and a channel weight matrix are obtained respectively.
[0084] 404. Perform parameterized weighting on the spatial topological constraints of the global structural association features, fuse the spatial weight matrix and the channel weight matrix, and generate a dynamic correlation weight coefficient; In step 404, spatial topology constraints refer to the spatial distribution relationship of global structural correlation features, which are usually implemented through graph convolutional networks (GCNs) or Transformers. Parameterized weighting refers to dynamically adjusting the spatial weight matrix and channel weight matrix through learnable parameters to generate dynamic correlation weight coefficients. Learnable parameters are usually implemented through fully connected layers or convolutional layers. Dynamic correlation weight coefficients refer to weight parameters used to adjust the strength of cross-modal feature alignment. Learnable parameters are usually implemented through fully connected layers or convolutional layers. A fully connected layer is a network layer that maps input features to output features, and a convolutional layer is a network layer that extracts local features. Dynamic correlation weight coefficients refer to weight parameters used to adjust the strength of cross-modal feature alignment.
[0085] In the embodiments of the present application, spatial topological constraints are first modeled for global structural correlation features, and a graph convolutional network or Transformer is used to capture the spatial distribution of features. Next, a weighted fusion of the spatial weight matrix and the channel weight matrix is performed using learnable parameters (such as a fully connected layer) to generate dynamic correlation weight coefficients. Ultimately, a weight coefficient is obtained that can dynamically adjust cross-modal feature alignment.
[0086] 405. Based on the dynamic correlation weight coefficient, quantify the cross-modal association strength between the aligned preliminary feature vector and the text symbol semantic vector to generate a dynamic correlation score.
[0087] In step 405, the aligned preliminary feature vector refers to the image features that have undergone preliminary alignment, typically achieved through feature mapping or feature transformation. Cross-modal correlation strength refers to the correlation between image and text features, typically calculated using dot product or cosine similarity. A dynamic correlation score quantifies the degree of cross-modal feature alignment, typically achieved through weighted summation or normalization.
[0088] In the present embodiment, the aligned preliminary feature vectors are first subjected to a dot product or cosine similarity calculation with the semantic vectors of the text symbols to obtain a preliminary cross-modal association strength. Next, the preliminary association strength is weighted and adjusted using a dynamic correlation weight coefficient to generate a dynamic correlation score. Ultimately, a dynamic correlation score is obtained that accurately reflects the degree of cross-modal feature alignment.
[0089] Here's a specific example: Assume that the text symbols "A", "B" and "C" are obtained on the surface of the power tower, respectively located at different positions of the tower body. Through spatial position coding and channel dimension coding, a spatial coding vector and a channel coding vector are generated. Through the interactive superposition and bidirectional projection of feature channels, a symbol structure association mapping vector is generated. The spatial weight matrix and the channel weight matrix are extracted from the symbol structure association mapping vector, and the dynamic correlation weight coefficient is generated through parameterized weighting. Finally, based on the dynamic correlation weight coefficient, the cross-modal correlation strength between the preliminary feature vector and the text symbol semantic vector is quantified to generate a dynamic correlation score for the recognition and reconstruction of the power tower. Through the above embodiments, the accuracy and anti-interference ability of power tower recognition can be effectively improved, providing strong support for the maintenance and management of power facilities.
[0090] In summary, steps 401 to 405 achieve high-precision, robust, and anti-interference capabilities for 3D recognition of power towers. This method precisely aligns text symbols with tower component features, effectively suppresses multipath reflections and environmental noise, and deeply integrates text symbol semantic vectors with the tower support structure and anti-corrosion coating texture features. It accurately reconstructs missing tower connection components and achieves high-precision 3D recognition and positioning of power towers, providing reliable technical support for intelligent inspection and maintenance of power infrastructure.
[0091] In order to solve the problem that it is difficult to collaboratively model the internal support structure features and the external anti-corrosion coating texture features and to establish a geometric constraint relationship, in some embodiments, the step 204 of optimizing and alternately updating the convolution kernel parameters based on the local consistency of the internal support structure features and the external anti-corrosion coating texture features, and establishing a geometric constraint relationship between the support structure features and the coating texture features, includes: 501. Decompose the internal support structure characteristics into connector spatial distribution characteristics and main beam morphology characteristics, and decompose the external anti-corrosion coating texture characteristics into corrosion area distribution characteristics and coating thickness variation characteristics; In step 501, the internal support structure features refer to the support structure information extracted from the interior of the power tower. The connector spatial distribution features refer to the spatial distribution of connectors within the support structure. The main beam morphology features refer to the geometry and curvature of the main beam. The external anti-corrosion coating texture features refer to the coating texture information extracted from the exterior of the power tower. The corrosion area distribution features refer to the distribution of corrosion areas within the coating. The coating thickness variation features refer to the gradient variation of the coating thickness.
[0092] In the examples of this application, the internal support structure features are first decomposed into the spatial distribution characteristics of the connectors and the morphological characteristics of the main beam. Spatial segmentation and morphological analysis techniques are then used to extract the positional information of the connectors and the geometric shape of the main beam, respectively. Next, the texture features of the external anti-corrosion coating are decomposed into the distribution characteristics of the corrosion area and the variation characteristics of the coating thickness. Image segmentation and thickness measurement techniques are then used to extract the distribution information of the corrosion area and the gradient variation of the coating thickness, respectively. Ultimately, the spatial distribution characteristics of the connectors, the morphological characteristics of the main beam, the distribution characteristics of the corrosion area, and the variation characteristics of the coating thickness are obtained.
[0093] 502. Based on the proximity constraints between bolt hole locations and coating defect areas, spatial topological encoding is performed on the spatial distribution characteristics of the connector and the distribution characteristics of the corrosion area to generate connector corrosion correlation features. Based on the synchronization constraints between the tower body curvature change and the coating thickness gradient, the main beam morphological characteristics and the coating thickness change characteristics are morphologically coupled and encoded to generate the main beam thickness coupling features. In step 502, the proximity constraint between the bolt hole position and the coating defect area refers to the spatial proximity relationship between the bolt hole position and the coating defect area. Spatial topology coding refers to modeling the spatial relationship of features based on proximity constraints. The synchronization constraint between the tower body curvature change and the coating thickness gradient refers to the morphological synchronization relationship between the tower body curvature change and the coating thickness gradient. Morphological thickness coupling coding refers to modeling the morphological relationship of features based on synchronization constraints. The corrosion association feature of the connector refers to the spatial association information between the spatial distribution features of the connector and the distribution features of the corrosion area. The main beam thickness coupling feature refers to the morphological association information between the main beam morphological features and the coating thickness change features.
[0094] In the embodiments of the present application, first, based on the proximity constraints between bolt hole locations and coating defect areas, spatial topological encoding is performed on the spatial distribution characteristics of the connectors and the distribution characteristics of the corrosion areas. A graph convolutional network or spatial attention mechanism is used to model the spatial association relationship, generating connector corrosion association characteristics. Next, based on the synchronization constraints between the tower body curvature change and the coating thickness gradient, morphological-thickness coupling encoding is performed on the main beam morphological characteristics and coating thickness change characteristics. A morphological analysis network or synchronization modeling technology is used to model the morphological association relationship, generating main beam thickness coupling characteristics. Finally, connector corrosion association characteristics and main beam thickness coupling characteristics are obtained, respectively.
[0095] 503. Construct a multi-branch convolution kernel group, wherein the first branch performs spatial constraint modeling on the corrosion-related features of the connector, and the second branch performs texture constraint modeling on the thickness coupling features of the main beam; In step 503, a multi-branch convolution kernel group is a set of convolution kernels consisting of multiple branches. The first branch is used to perform spatial constraint modeling on the connector corrosion-related features, and the second branch is used to perform texture constraint modeling on the main beam thickness coupling features. Spatially constrained modeling refers to modeling the spatial relationship of the connector corrosion-related features, while texture constraint modeling refers to modeling the texture relationship of the main beam thickness coupling features.
[0096] In this embodiment, a multi-branch convolutional kernel group is first constructed. The first branch uses a convolutional layer or graph convolutional network to perform spatially constrained modeling on the corrosion-related features of the connector, capturing the spatial relationship between the connector and the corroded area. The second branch uses a convolutional layer or texture analysis network to perform texture-constrained modeling on the coupled features of the main beam thickness, capturing the textural relationship between the main beam and the coating thickness. Ultimately, the spatially constrained modeling results and the texture-constrained modeling results are obtained, respectively.
[0097] 504. Based on the local consistency loss of the spatial constraint modeling result and the texture constraint modeling result, alternately update the parameters of the multi-branch convolution kernel group to generate a geometric constraint relationship between the support structure feature and the coating texture feature.
[0098] In step 504, local consistency loss refers to the difference in consistency between the spatially constrained modeling results and the texture-constrained modeling results in a local area. Alternating update refers to alternating adjustment of the parameters of the multi-branch convolution kernel group through iterative optimization. Geometric constraint relationships refer to the spatial and morphological associations between support structure features and coating texture features.
[0099] In the present embodiment, a loss function (such as mean squared error or cross-entropy loss) is first used to calculate the consistency difference based on the local consistency loss between the spatially constrained modeling results and the texture-constrained modeling results. Next, the parameters of the multi-branch convolution kernel group are iteratively optimized using an alternating update method (such as gradient descent or backpropagation) to generate a geometric constraint relationship between the support structure features and the coating texture features. Ultimately, a geometric constraint relationship is obtained that accurately reflects the correlation between the support structure and the coating texture.
[0100] Here's a specific example: In the identification of power towers, for example, radar scans revealed that internal connectors were concentrated at the four corners, the main beams formed a grid structure, and external coating corrosion was primarily located at the base, with thickness gradually decreasing from bottom to top. Based on the proximity of bolt holes and corroded areas, connector corrosion-related features were generated. For example, corrosion areas were primarily distributed around bolt holes. Simultaneously, the synchronization between tower curvature and coating thickness was exploited to generate main beam thickness-coupled features. For example, coating thickness varied significantly at tower corners. Subsequently, a multi-branch convolution kernel group was used to spatially model the connector corrosion-related features to extract corrosion distribution patterns, and texture model the main beam thickness-coupled features to capture thickness variations. By optimizing the consistency of spatial and texture modeling, a geometric constraint relationship between the supporting structure and coating texture was generated. For example, this ensured that connector corrosion and main beam thickness features were spatially and texturally aligned. Finally, this relationship was used to reconstruct missing tower connections, such as repairing corrosion defects. This generated an interference-resistant 3D recognition model, providing precise support for facility maintenance.
[0101] In summary, through steps 501 to 504, the internal support structure and external coating texture features are decomposed, encoded based on spatial proximity and morphological synchronization constraints, a multi-branch convolution kernel group is constructed for spatial and texture modeling, and parameters are alternately updated through local consistency loss. Finally, a precise geometric constraint relationship between the support structure and the coating texture is generated, thereby efficiently associating and analyzing the internal support and external coating features of the tower body.
[0102] In order to solve the problem that it is difficult to accurately analyze the original reflected signal of the power tower and to extract the positioning information of the text symbols, in some embodiments, the original reflected signal of the power tower is obtained by the radar array in step 103, the target reflection area is screened based on the multimodal mapping relationship, the phase difference data in the original reflected signal is analyzed, and a penetration feature set containing the tower skeleton outline and the corresponding text symbol positioning information is generated, including: 601. Perform multi-angle signal acquisition on the power tower using a circular or linearly deployed radar array to obtain a set of original reflected signals, wherein the radar beam direction is aligned with the spatial distribution of text symbols on the tower surface during signal acquisition; In step 601, the radar array refers to a detection system composed of multiple radar units, which generally uses the principle of transmitting and receiving electromagnetic waves for target detection. Annular or linear deployment refers to the arrangement of radar units in an annular or linear manner. Annular deployment is suitable for all-round detection, and linear deployment is suitable for detection in a specific direction. Multi-angle signal acquisition refers to the detection of radar signals from power towers from different angles to obtain more comprehensive target information. The original reflection signal set refers to the reflection signal data set collected by the radar array, which contains information such as the reflection intensity and phase of the target object. The radar beam direction refers to the direction of transmission and reception of the radar signal, which is usually controlled by adjusting the pointing direction of the radar antenna. The spatial distribution of text symbols on the tower surface refers to the position and arrangement of text symbols on the tower surface, which is usually used to identify tower components or provide other information.
[0103] In an embodiment of the present application, a radar array deployed in a circular or linear manner is first used to collect multi-angle signals from a power tower, ensuring that the radar beam direction is aligned with the spatial distribution of text symbols on the tower surface, so as to obtain a set of original reflected signals containing text symbol information. The deployment method of the radar array is selected based on actual detection needs. Circular deployment is suitable for all-round detection, while linear deployment is suitable for detection in a specific direction. During the signal acquisition process, the radar beam direction is aligned with the spatial distribution of text symbols on the tower surface, ensuring that the text symbol information can be effectively captured. Ultimately, a set of original reflected signals containing multi-angle detection data is obtained.
[0104] 602. Based on the multimodal mapping relationship, perform beamforming processing on the original reflected signal set to select a target reflection area containing text symbol positioning information, wherein beamforming parameters are dynamically adjusted according to the text symbol size and the geometric size of the tower body component; In step 602, the multimodal mapping relationship refers to the association relationship between the text symbols and the tower body components, which is usually implemented through a predefined mapping table or a machine learning model. Beamforming processing refers to focusing and enhancing the original reflected signal to improve the signal-to-noise ratio and target resolution of the signal. The target reflection area refers to the signal area containing the text symbol positioning information, which is usually obtained by screening through a signal processing algorithm. Beamforming parameters refer to the parameters used to control the beamforming processing, such as beam width, direction and gain, which are usually dynamically adjusted according to the characteristics of the detection target. The text symbol size refers to the size of the text symbol, which is usually measured by image processing technology. The geometric dimensions of the tower body components refer to the shape and size of the tower body components, which are usually obtained through three-dimensional modeling or measurement technology.
[0105] In an embodiment of the present application, beamforming processing is first performed on the original reflection signal set based on a multimodal mapping relationship, and a beamforming algorithm (such as a delayed-sum algorithm or an adaptive beamforming algorithm) is used to screen the target reflection area containing the text symbol positioning information. The beamforming parameters are dynamically adjusted according to the text symbol size and the geometric dimensions of the tower components to ensure signal focusing accuracy. The delayed-sum algorithm achieves beam focusing by adjusting the signal delay, and the adaptive beamforming algorithm achieves beam optimization through adaptive filtering. During the beamforming process, the multimodal mapping relationship is used to guide the screening of the target reflection area to ensure that the text symbol positioning information can be effectively extracted. Ultimately, the target reflection area containing the text symbol positioning information is obtained.
[0106] 603. Perform phase difference interferometry processing on the original reflection signal of the target reflection area, analyze the phase difference data between adjacent radar channels, and generate an initial three-dimensional point cloud of the tower skeleton outline; In step 603, phase difference interferometry processing involves processing the phase difference data between adjacent radar channels to extract the target object's geometric information. Phase difference data refers to the phase difference between the signals received by adjacent radar channels and is typically related to the target object's distance and shape. Interpreting the phase difference data involves extracting the target object's geometric information from the phase difference, typically achieved through an interferometry algorithm. The initial 3D point cloud is the 3D point cloud data of the tower skeleton outline generated through phase difference interferometry, containing the target object's spatial coordinate information.
[0107] In this embodiment, phase difference interferometry is first performed on the original reflected signals from the target reflection area. A phase difference interferometry algorithm (such as synthetic aperture radar interferometry) is used to analyze the phase difference data between adjacent radar channels to generate an initial three-dimensional point cloud of the tower skeleton outline. Synthetic aperture radar interferometry improves the resolution and geometric accuracy of the target object by synthesizing the signals from multiple radar channels. During the phase difference interferometry process, the phase difference data between adjacent radar channels is used to calculate the distance and shape of the target object, generating an initial three-dimensional point cloud. Ultimately, an initial three-dimensional point cloud containing geometric information about the tower skeleton is obtained.
[0108] 604. Based on the component association of the text symbols in the multimodal mapping relationship, spatial filtering is performed on the initial three-dimensional point cloud to remove scattered points irrelevant to the positioning of the text symbols, thereby generating a symbol-constrained tower skeleton point cloud; In step 604, the component association of the text symbol in the multimodal mapping relationship refers to the correspondence between the text symbol and the tower components, which is usually implemented through a predefined mapping table or machine learning model. Spatial filtering refers to the screening and elimination of the initial three-dimensional point cloud to remove noise and irrelevant data. Scattering points refer to the point cloud data generated by radar signal reflection, which generally contains information about the target object and noise. The symbol-constrained tower skeleton point cloud refers to the point cloud data related to the positioning of the text symbol retained after spatial filtering, and is generally used for subsequent three-dimensional reconstruction and analysis.
[0109] In an embodiment of the present application, first, based on the component association of the text symbol in the multimodal mapping relationship, the initial three-dimensional point cloud is spatially filtered, and a spatial filtering algorithm (such as distance-based filtering or density-based filtering) is used to eliminate scattered points that are not related to the positioning of the text symbol, thereby generating a symbol-constrained tower skeleton point cloud. Distance-based filtering eliminates noise points by setting a distance threshold, and density-based filtering eliminates sparse points by setting a density threshold. During the spatial filtering process, the component association of the text symbol in the multimodal mapping relationship is used to guide point cloud screening to ensure that point cloud data related to the positioning of the text symbol can be retained. Ultimately, a symbol-constrained tower skeleton point cloud containing text symbol positioning information is obtained.
[0110] 605. Perform feature fusion on the tower skeleton point cloud with the symbol constraint and the multipath reflection intensity in the previous phase difference interferometry processing to generate a penetration feature set including the text symbol positioning information and the tower skeleton geometric parameters.
[0111] In step 605, multipath reflection intensity refers to the intensity of multiple reflections of the radar signal during propagation, and is typically related to the material and shape of the target object. Feature fusion combines the symbol-constrained tower skeleton point cloud with the multipath reflection intensity to generate more comprehensive feature data. The penetrability feature set, a feature data set containing both the text symbol positioning information and the tower skeleton geometric parameters, is typically used for subsequent 3D reconstruction and analysis.
[0112] In an embodiment of the present application, the symbol-constrained tower skeleton point cloud is first subjected to feature fusion with the multipath reflection intensity from the previous phase difference interferometry processing. A feature fusion algorithm (such as weighted fusion or convolutional fusion) is then used to generate a penetration feature set containing the text symbol positioning information and the tower skeleton geometric parameters. Weighted fusion combines different features by setting weight coefficients, while convolutional fusion fuses different features through a convolution operation. During the feature fusion process, the symbol-constrained tower skeleton point cloud and the multipath reflection intensity are used to generate more comprehensive feature data, ensuring the integrity and accuracy of the penetration feature set. Ultimately, a penetration feature set is obtained for the three-dimensional reconstruction of the power tower.
[0113] Here's a specific example: To identify power towers, a circular or linear radar array is first used to collect multi-angle signals from the tower, ensuring that the radar beam is aligned with the text symbols (such as "A" and "B") on the tower surface to obtain the original reflected signal. Based on the multimodal mapping relationship between the text symbols and tower components, the beamforming parameters are dynamically adjusted to select target reflection areas containing the text symbol location information. Next, phase difference interferometry is performed on the target area signals, analyzing the phase difference data of adjacent radar channels to generate an initial 3D point cloud of the tower skeleton. Based on the component association of the text symbols, the initial point cloud is spatially filtered to remove irrelevant scattering points and generate a symbol-constrained tower skeleton point cloud. Finally, the symbol-constrained point cloud is fused with the multipath reflection intensity to generate a set of penetration features containing the text symbol location information and the skeleton geometric parameters. This method achieves high-precision recognition and 3D reconstruction of power towers, providing reliable support for facility maintenance.
[0114] In summary, through steps 601 to 605, multi-angle signal acquisition of the power tower is achieved through a circular or linearly deployed radar array, beamforming processing is performed on the original reflected signal based on the multimodal mapping relationship, the target reflection area containing the text symbol positioning information is screened, and the initial three-dimensional point cloud of the tower skeleton outline is generated through phase difference interferometry processing; the initial point cloud is further spatially filtered based on the component association of the text symbol, and irrelevant scattering points are eliminated to generate a symbol-constrained tower skeleton point cloud; finally, the symbol-constrained point cloud is fused with the multipath reflection intensity feature to generate a penetration feature set containing the text symbol positioning information and the tower skeleton geometric parameters, thereby achieving high-precision positioning and feature extraction of the text symbol and skeleton geometry of the power tower.
[0115] In order to solve the problem that it is difficult to remove scattered points irrelevant to the positioning of the text symbols in the initial three-dimensional point cloud and difficult to generate a symbol-constrained tower skeleton point cloud, in some embodiments, the initial three-dimensional point cloud is spatially filtered based on the component association of the text symbols in the multimodal mapping relationship, and scattered points irrelevant to the positioning of the text symbols are removed to generate a symbol-constrained tower skeleton point cloud, including: 701. Divide the symbol semantically associated area and the non-associated area based on the association between the text symbols and the tower components in the multimodal mapping relationship; In step 701, the multimodal mapping relationship refers to the association between the text symbol and the tower components, typically implemented through a predefined mapping table or machine learning model. The symbol semantically associated region refers to the region that is semantically relevant to the text symbol, typically including scattered points directly related to the tower components. The unassociated region refers to the region that is unrelated to the text symbol semantics, typically including scattered points or noise points unrelated to the tower components.
[0116] In an embodiment of the present application, a region division algorithm (such as a semantic-based segmentation algorithm or a clustering algorithm) is first used to divide the point cloud into symbol semantically associated regions and non-associated regions based on the association between the text symbols and the tower components in the multimodal mapping relationship. The symbol semantically associated regions include regions associated with lightning protection symbols, support symbols, and fixed symbols, and the non-associated regions include regions unrelated to coating corrosion symbols. During the region division process, the multimodal mapping relationship is used to guide the region division and ensure the accuracy of the symbol semantically associated regions and non-associated regions. Ultimately, the division result of the symbol semantically associated regions and non-associated regions is obtained.
[0117] 702. In the symbol semantic association area, the scattered points are processed according to the component association of the text symbol. In the non-symbol semantic association area, the scattered points that are not related to the corrosion expansion characteristics are eliminated according to the defect boundary distribution of the coating corrosion symbol; In step 702, the component association of the text symbol refers to the correspondence between the text symbol and the tower components, which is usually achieved through a predefined mapping table or machine learning model. Scattering points refer to the point cloud data generated by radar signal reflection, which generally contains information about the target object and noise. The defect boundary distribution of the coating corrosion symbol refers to the boundary characteristics of the coating corrosion area, which is usually extracted through image processing technology. The rust expansion characteristics refer to the expansion and deformation characteristics of the coating corrosion area, which are generally related to the degree of corrosion.
[0118] In an embodiment of the present application, first, within the symbol semantic association area, the scattering points are processed according to the component association of the text symbol, and a point cloud processing algorithm (such as density-based filtering or distance-based filtering) is used to retain the scattering points related to the tower components. Then, within the non-symbol semantic association area, based on the defect boundary distribution of the coating corrosion symbol, a point cloud elimination algorithm (such as boundary-based filtering or feature-based filtering) is used to eliminate scattering points that are not related to the corrosion expansion characteristics. During the point cloud processing process, the component association of the text symbol and the defect boundary distribution of the coating corrosion symbol are used to guide the processing of the scattering points, ensuring that the scattering points within the symbol semantic association area are retained and the irrelevant scattering points in the non-association area are eliminated. Ultimately, the scattering points after processing the symbol semantic association area and the remaining scattering points after eliminating the non-association area are obtained.
[0119] 703. Mark the initial three-dimensional point cloud with symbol semantic association areas. During marking, the lightning protection symbol area radiates outward from the center of the symbol, the support symbol area extends along the main beam axis, and the fixed symbol area extends according to the bolt distribution density matrix, thereby generating a regional marking point cloud. In step 703, the initial three-dimensional point cloud refers to the three-dimensional point cloud data generated by radar signal processing, which usually contains the spatial coordinate information of the target object. Symbol semantic association area marking refers to the identification and expansion of the symbol semantic association area, which is usually used for subsequent point cloud processing and analysis. The lightning protection symbol area refers to the area related to the lightning protection symbol, which usually radiates outward with the center of the symbol as the center of the sphere. The support symbol area refers to the area related to the support symbol, which usually extends along the axial direction of the main beam. The fixed symbol area refers to the area related to the fixed symbol, which is usually expanded according to the bolt distribution density matrix. The area marked point cloud refers to the point cloud data after the symbol semantic association area marking, which is usually used for subsequent multi-scale aggregation.
[0120] In an embodiment of the present application, the initial three-dimensional point cloud is first labeled with symbolic semantically associated regions, and a region labeling algorithm (such as a sphere-centered expansion algorithm or an axially-centered extension algorithm) is used to label the lightning protection symbol region, the support symbol region, and the fixed symbol region. The lightning protection symbol region radiates outward from the center of the symbol, the support symbol region extends along the main beam axis, and the fixed symbol region extends according to the bolt distribution density matrix. During the region labeling process, the features of the symbolic semantically associated regions are used to guide labeling and ensure the accuracy of the region labeling point cloud. Ultimately, a region labeling point cloud containing symbolic semantically associated region labels is obtained.
[0121] 704. Based on the multi-scale accuracy parameters of the regional marker point cloud, perform multi-scale aggregation on the scattered points processed in the symbolic semantic associated area and the remaining scattered points after the non-associated area is eliminated to generate a symbol-constrained tower skeleton point cloud.
[0122] In step 704, multi-scale accuracy parameters are used to control the accuracy of point cloud aggregation, typically including point cloud density, resolution, and scale range. Multi-scale aggregation involves fusing point cloud data at different scales to generate a more comprehensive point cloud. The symbolically constrained tower skeleton point cloud, generated after multi-scale aggregation and containing symbolic semantic information, is typically used for subsequent 3D reconstruction and analysis.
[0123] In this embodiment, based on the multi-scale precision parameters of the regionally labeled point cloud, a multi-scale aggregation algorithm (such as density-based or resolution-based aggregation) is first used to aggregate the scattered points after processing the semantically associated regions and the remaining scattered points after removing the non-associated regions, generating a symbolically constrained tower skeleton point cloud. During the multi-scale aggregation process, the multi-scale precision parameters are used to control the aggregation accuracy, ensuring the integrity and accuracy of the symbolically constrained tower skeleton point cloud. Ultimately, a symbolically constrained tower skeleton point cloud is obtained for 3D reconstruction of power towers.
[0124] Here's a specific example: In power tower recognition, the tower surface is first divided into semantically associated regions and unassociated regions based on the multimodal mapping relationship between text symbols and tower components. For example, the text symbol "A" is associated with the lightning protection component, and the surrounding area is the semantically associated region, while the remaining area is the unassociated region. Within the semantically associated region, scattered points are processed based on the component association of the text symbol, for example, scatter points related to the lightning protection component are retained. In the unassociated region, scatter points unrelated to the corrosion expansion characteristics are removed based on the defect boundary distribution of the coating corrosion symbol, for example, noise points caused by coating shedding are removed. Next, the initial 3D point cloud is labeled with semantically associated regions. Finally, based on the multi-scale accuracy parameters of the region-labeled point cloud, the scattered points processed in the semantically associated regions and the remaining scattered points after the unassociated regions are removed are multi-scale aggregated to generate a symbol-constrained tower skeleton point cloud. For example, through multi-scale aggregation, scattered points in the lightning protection symbol region, the support symbol region, and the fixed symbol region are fused to form an accurate tower skeleton point cloud. This method improves the accuracy and reliability of power tower identification through symbolic semantic constraints, providing strong support for facility maintenance.
[0125] In summary, steps 701 to 704 are used to divide the symbol semantically associated area and the non-associated area based on the association between the text symbols and the tower components in the multimodal mapping relationship. In the associated area, the scattering points are processed according to the component association of the text symbols, and in the non-associated area, irrelevant scattering points are eliminated according to the defect boundary distribution of the coating corrosion symbol. The initial three-dimensional point cloud is further marked with symbol semantic associated areas, and expanded according to the distribution characteristics of the lightning protection symbol, support symbol and fixed symbol to generate a regional marked point cloud. Finally, based on the multi-scale accuracy parameters of the regional marked point cloud, the scattering points after the symbol semantic associated area processing and the remaining scattering points after the non-associated area is eliminated are multi-scale aggregated to generate a symbol-constrained tower skeleton point cloud, thereby achieving high-precision symbol semantic constraint and regional optimization processing of the tower skeleton point cloud.
[0126] In order to solve the problem of difficulty in suppressing electromagnetic interference in multipath reflection signals and insufficient noise compensation accuracy, in some embodiments, step 102 constructs a dynamic noise suppression model based on the penetration feature set, performs attenuation compensation on the multipath reflection signals using the dynamic noise suppression model, and generates an anti-electromagnetic interference compensation feature set, including: 801. Extracting symbol semantic noise suppression parameters based on the spatial mapping relationship between the text symbol positioning information in the penetration feature set and the tower body components, and constructing a dynamic noise suppression model; In step 801, the penetration feature set refers to a feature data set containing text symbol positioning information and tower skeleton geometric parameters, which is usually generated by radar signal processing. Text symbol positioning information refers to the position information of text symbols on the tower surface, which is usually extracted by image processing technology. The spatial mapping relationship of tower components refers to the spatial correspondence between text symbols and tower components, which is usually implemented through a predefined mapping table or machine learning model. The symbol semantic noise suppression parameter refers to the parameter used to suppress noise unrelated to the semantics of the text symbol, which is usually generated by a feature extraction algorithm. The dynamic noise suppression model refers to a model used to suppress noise in multipath reflection signals, which is usually constructed through a parameterized method.
[0127] In this embodiment of the present application, a feature extraction algorithm (such as semantic-based filtering or mapping-based feature extraction) is first used to extract symbolic noise suppression parameters based on the spatial mapping relationship between the text symbol positioning information in the penetration feature set and the tower components. Next, a dynamic noise suppression model is constructed based on the symbolic noise suppression parameters using a parameterized approach (such as weight-based model construction or optimization-based model construction). During the feature extraction and model construction process, the spatial mapping relationship between the text symbol positioning information and the tower components is used to guide the extraction of noise suppression parameters and the construction of the model, ensuring the accuracy of the dynamic noise suppression model. Ultimately, a dynamic noise suppression model containing the symbolic noise suppression parameters is obtained.
[0128] 802. In the dynamic noise suppression model, dynamically adjust the structural reflection path attenuation coefficient based on the geometric parameters of the tower skeleton outline; In step 802, the tower frame geometry parameters refer to the shape and dimensions of the tower frame, typically generated through 3D point cloud processing. The structural reflection path attenuation coefficient refers to the attenuation of multipath reflection signals within the tower structure and is typically related to the tower frame geometry. Dynamic adjustment refers to adjusting the structural reflection path attenuation coefficient based on the tower frame geometry parameters, typically achieved through parameterization.
[0129] In this embodiment, a dynamic noise suppression model is first implemented using a parametric approach (e.g., geometry-based or optimization-based attenuation coefficient adjustment) to dynamically adjust the attenuation coefficient of the structural reflection path based on the geometric parameters of the tower frame's profile. During this parametric adjustment process, the tower frame's geometric parameters guide the adjustment of the attenuation coefficient, ensuring the accuracy of the structural reflection path attenuation coefficient. Ultimately, a dynamic noise suppression model is obtained that incorporates the dynamically adjusted structural reflection path attenuation coefficient.
[0130] 803. Generate an interference source spatial weight map based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, integrate the interference source spatial weight map into the dynamic noise suppression model, perform azimuth-dependent attenuation compensation on the multipath reflection signal, and extract azimuth compensation parameters; In step 803, the electromagnetic interference distribution characteristics refer to the distribution of electromagnetic interference around the power tower, which is usually obtained through electromagnetic field measurement technology. The azimuth information of the high-voltage equipment refers to the direction information of the high-voltage equipment relative to the power tower, which is usually obtained through azimuth measurement technology. The interference source spatial weight map refers to a map used to represent the weight distribution of electromagnetic interference sources in space, which is usually generated by a weight allocation algorithm. Azimuth-related attenuation compensation refers to attenuation compensation for multipath reflection signals based on the azimuth of the interference source, which is usually implemented through a parameterized method. The azimuth compensation parameters refer to parameters used for azimuth-related attenuation compensation, which are usually generated through a feature extraction algorithm.
[0131] In an embodiment of the present application, first, based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, a weight distribution algorithm (such as distance-based weight distribution or direction-based weight distribution) is used to generate a spatial weight map of the interference source. Next, the spatial weight map of the interference source is integrated into the dynamic noise suppression model, and a parameterized method (such as weight-based attenuation compensation or optimization-based attenuation compensation) is used to perform azimuth-related attenuation compensation on the multipath reflection signal, and extract the azimuth compensation parameters. During the weight distribution and attenuation compensation process, the electromagnetic interference distribution characteristics and the azimuth information of the high-voltage equipment are used to guide the generation of the spatial weight map of the interference source and the extraction of the azimuth compensation parameters, ensuring the accuracy of the azimuth-related attenuation compensation. Finally, a dynamic noise suppression model containing azimuth compensation parameters is obtained.
[0132] 804. Performing a tower region weighted fusion of the symbol semantic noise suppression parameter, the structural reflection path attenuation coefficient, and the azimuth compensation parameter in the dynamic noise suppression model to update the dynamic noise suppression model. In step 804, weighted fusion of the tower region involves weighted fusion of the symbolic noise suppression parameters, the structural reflection path attenuation coefficient, and the azimuth compensation parameters within the dynamic noise suppression model, typically achieved through weighted summation or convolution fusion. Updating the dynamic noise suppression model involves adjusting the model parameters based on the weighted fusion results, typically achieved through parameterization.
[0133] In this embodiment of the present application, the symbolic noise suppression parameters, the structural reflection path attenuation coefficient, and the azimuth compensation parameters are first weightedly fused within the tower region of a dynamic noise suppression model. A weighted fusion algorithm (such as weighted summation or convolution fusion) is used to generate the fused parameters. Next, based on the fused parameters, the dynamic noise suppression model is updated using a parameterized method (such as weight-based model updating or optimization-based model updating). During this weighted fusion and model update process, the symbolic noise suppression parameters, the structural reflection path attenuation coefficient, and the azimuth compensation parameters are used to guide the model update and ensure the accuracy of the dynamic noise suppression model. Ultimately, an updated dynamic noise suppression model is obtained.
[0134] 805. Perform iterative path attenuation compensation on the multipath reflection signals in the penetration feature set by using the updated dynamic noise suppression model to generate a compensation feature set for resisting electromagnetic interference.
[0135] In step 805, iterative path attenuation compensation involves performing multiple iterations of path attenuation compensation on multipath reflection signals, typically achieved through iterative optimization. The electromagnetic interference (EMI) compensation feature set is the set of EMI feature data generated after path attenuation compensation, typically used for subsequent 3D reconstruction and analysis.
[0136] In this embodiment, an updated dynamic noise suppression model is first used to iteratively compensate for multipath reflection signals in a penetration feature set using path attenuation compensation. An iterative optimization method (such as gradient descent or backpropagation) is then used to generate a compensation feature set for electromagnetic interference resistance. During this iterative optimization process, the updated dynamic noise suppression model is used to guide path attenuation compensation, ensuring the accuracy and robustness of the compensation feature set for electromagnetic interference resistance. Ultimately, a compensation feature set for electromagnetic interference resistance is obtained for 3D reconstruction of power towers.
[0137] Here's a specific example: In power tower identification, symbolic semantic noise suppression parameters are first extracted based on the spatial mapping relationship between the positioning information of text symbols in the penetration feature set and tower components, and a dynamic noise suppression model is constructed. For example, the correlation between the text symbol "A" and lightning protection components is used to suppress irrelevant noise. Next, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower skeleton outline, for example, to enhance the effectiveness of the signal in the main beam area. Subsequently, a spatial weight map of interference sources is generated by combining the electromagnetic interference distribution characteristics and the azimuth information of surrounding high-voltage equipment. Azimuth-dependent attenuation compensation is performed on multipath reflection signals, and azimuth compensation parameters are extracted. For example, signal compensation is performed for interference from high-voltage equipment. The symbolic semantic noise suppression parameters, structural reflection path attenuation coefficient, and azimuth compensation parameters are weighted and fused in the model to update the dynamic noise suppression model. Finally, the updated model is used to iteratively compensate for path attenuation of multipath reflection signals to generate a compensation feature set for electromagnetic interference resistance. This method significantly improves the anti-interference capability of power tower identification through dynamic noise suppression and compensation, providing reliable support for facility maintenance.
[0138] In summary, through steps 801 to 805, it is achieved to extract symbol semantic noise suppression parameters through the spatial mapping relationship between the text symbol positioning information in the penetrability feature set and the tower body components, construct a dynamic noise suppression model, and dynamically adjust the structural reflection path attenuation coefficient based on the geometric parameters of the tower body skeleton outline; then, the symbol semantic noise suppression parameters, the structural reflection path attenuation coefficient and the azimuth compensation parameters are weightedly fused in the tower body area in the model to update the dynamic noise suppression model; finally, the multipath reflection signal in the penetrability feature set is iteratively compensated for the path attenuation through the updated model to generate a compensation feature set for electromagnetic interference resistance, thereby achieving efficient noise suppression and electromagnetic interference compensation of the multipath reflection signal of the power tower.
[0139] In order to solve the problem of difficulty in reconstructing missing tower connection components and insufficient ability of the three-dimensional recognition model to resist environmental interference, in some embodiments, the step 105 of reconstructing the missing tower connection components based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference includes: 901. Based on the curvature continuity constraint of the tower main beam and the topological distribution constraint of the bolt hole positions in the geometric constraint relationship, spatially locate the area where the connection component is missing and generate a bounding box of the missing area; In step 901, the geometric constraint relationship refers to the spatial and morphological correlation between the supporting structure features and the coating texture features, which is usually generated through multimodal feature modeling. The curvature continuity constraint of the tower main beam means that the curvature change of the main beam in space should remain continuous, which is usually achieved through curvature analysis technology. The topological distribution constraint of the bolt hole position means that the distribution of the bolt hole position in space should satisfy a specific topological relationship, which is usually achieved through topological analysis technology. The area of missing connection parts refers to the part of the tower structure where the connection parts are missing, which is usually identified through point cloud analysis technology. The missing area bounding box refers to the spatial bounding box used to identify the area of missing connection parts, which is usually achieved through a bounding box generation algorithm.
[0140] In this embodiment, a spatial positioning algorithm (such as curvature-based or topology-based positioning) is first used to locate the region of missing connection components based on the geometric constraints of the tower beam curvature continuity and the bolt hole topology distribution. This algorithm then generates a bounding box for the missing region. During this spatial positioning process, the curvature continuity constraint of the tower beam and the topology distribution constraint of the bolt hole locations guide the location of the missing region, ensuring the accuracy of the bounding box. Ultimately, a bounding box encompassing the missing connection component region is obtained.
[0141] 902. Lock the range within the missing region boundary box and generate candidate connection component parameters based on the angle constraint relationship between the main beam and the support member; In step 902, the angle constraint between the main beam and the support member refers to the angle between the main beam and the support member must satisfy a specific geometric relationship, which is usually achieved through angle analysis technology. Candidate connection component parameters are candidate parameters used to reconstruct missing connection components, which are usually generated through parameterization methods.
[0142] In this embodiment, the scope is first locked within the missing region bounding box. Based on the angular constraints between the main beam and the support members, a parametric method (such as angle-based or geometry-based parameter generation) is used to generate candidate connection component parameters. During the parametric generation process, the angular constraints between the main beam and the support members guide the generation of candidate connection component parameters, ensuring their rationality. Ultimately, a candidate parameter set containing the candidate connection component parameters is obtained.
[0143] 903. Perform spatial matching verification on the candidate connection component parameters based on the bolt hole topological distribution constraint, eliminate candidate parameters whose axial deviation from the bolt hole exceeds a preset threshold, and generate a preliminary connection component set; In step 903, the topological distribution constraint of the bolt hole position means that the spatial distribution of the bolt hole position should satisfy a specific topological relationship, which is usually achieved through topological analysis technology. Spatial matching verification refers to the spatial matching verification of the candidate connection component parameters, which is usually achieved through a matching algorithm. The axial deviation refers to the deviation between the candidate connection component parameters and the bolt hole position in the axial direction, which is usually achieved through deviation analysis technology. The preset threshold refers to the threshold used to judge whether the axial deviation is reasonable, which is usually determined through experiments or experience. The preliminary connection component set refers to the set of candidate connection component parameters retained after spatial matching verification, which is usually used for subsequent morphological optimization.
[0144] In this embodiment, a matching algorithm (such as topology-based or geometry-based matching) is first used to perform spatial matching verification on candidate connection component parameters based on the topological distribution constraints of bolt hole locations. Candidate parameters whose axial deviation from the bolt hole locations exceeds a preset threshold are eliminated, generating a preliminary set of connection components. During this matching verification process, the topological distribution constraints of the bolt hole locations are used to guide the verification of candidate connection component parameters, ensuring the rationality of the preliminary set of connection components. Ultimately, a preliminary set of connection components containing reasonable candidate connection component parameters is obtained.
[0145] 904. Performing morphological optimization on the preliminary set of connection components based on the curvature continuity constraint of the tower main beam, adjusting the installation angle and contact surface curvature of the connection components to satisfy the curvature difference threshold of the main beam skeleton point cloud; In step 904, the curvature continuity constraint of the tower main beam means that the curvature change of the main beam in space should remain continuous, which is usually achieved through curvature analysis technology. Morphological optimization refers to the morphological optimization of the preliminary set of connection components, which is usually achieved through optimization algorithms. The installation angle refers to the installation angle between the connection component and the main beam, which is usually achieved through angle analysis technology. The contact surface curvature refers to the curvature of the contact surface between the connection component and the main beam, which is usually achieved through curvature analysis technology. The curvature difference threshold refers to the threshold used to judge the continuity of curvature, which is usually determined through experiments or experience.
[0146] In this embodiment, a preliminary set of connection components is first optimized using an optimization algorithm (such as curvature-based or geometry-based optimization) based on the curvature continuity constraint of the tower main beam. The connection component installation angles and contact surface curvatures are adjusted to meet the curvature difference threshold of the main beam skeleton point cloud. During the optimization process, the curvature continuity constraint of the tower main beam guides the optimization of the connection components, ensuring the rationality of the installation angles and contact surface curvatures. Ultimately, a morphologically optimized set of connection components is obtained.
[0147] 905. Based on the load distribution characteristics of the support members and the tower frame, perform environmental interference resistance reinforcement modeling on the connection components after the morphology optimization, wherein the reinforcement strength is dynamically adjusted according to the connection density between the support members and the main beam; In step 905, the load distribution characteristics of the support members and the tower frame refer to the load distribution between the support members and the tower frame should satisfy a specific mechanical relationship, which is usually achieved through mechanical analysis technology. Anti-environmental interference reinforcement modeling refers to the anti-environmental interference reinforcement modeling of the connection components after morphological optimization, which is usually achieved through reinforcement algorithms. Reinforcement strength refers to the reinforcement strength of the connection components, which is usually achieved through strength analysis technology. The connection density between the support members and the main beam refers to the connection density between the support members and the main beam, which is usually achieved through density analysis technology.
[0148] In this embodiment, based on the load distribution characteristics of the supports and the tower frame, a reinforcement algorithm (such as load-based or strength-based reinforcement) is first used to model the connection components after morphological optimization for environmental interference resistance. The reinforcement strength is dynamically adjusted based on the connection density between the supports and the main beam. During the reinforcement modeling process, the load distribution characteristics of the supports and the tower frame are used to guide the reinforcement of the connection components, ensuring the appropriate reinforcement strength. Ultimately, a set of connection components is obtained after the environmental interference resistance reinforcement modeling.
[0149] 906. Perform geometric constraint matching on the reinforced modeled connection components and the tower skeleton point cloud to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0150] In step 906, geometric constraint matching involves matching the reinforced modeled connection components with the tower skeleton point cloud. This is typically achieved through a matching algorithm. The environmentally-resistant 3D recognition model of the power tower is generated after geometric constraint matching and is typically used for subsequent inspection and maintenance.
[0151] In this embodiment, geometric constraint matching is first performed on the reinforced and modeled connection components and the tower skeleton point cloud. A matching algorithm (such as geometry-based or constraint-based matching) is then used to generate a 3D recognition model of the power tower that is resistant to environmental interference. During this geometric constraint matching process, the reinforced and modeled connection components and the tower skeleton point cloud are used to guide model generation, ensuring the integrity and accuracy of the 3D recognition model of the power tower that is resistant to environmental interference. Ultimately, a 3D recognition model of the power tower that is resistant to environmental interference is obtained, which can be used for power tower inspection and maintenance.
[0152] Here's a specific example: During an inspection of a power tower, a connector was discovered missing in the middle of the tower due to long-term corrosion. By analyzing the main beam curvature and bolt hole distribution, the bounding box of the missing region was determined to be between 15 and 17 meters in height. Within this bounding box, multiple candidate connector parameters were generated based on the standard angle between the main beam and the support (typically 45° to 60°). For example, a connector with a length of 1.2 meters and a diameter of 0.3 meters was used. The candidate parameters were then verified using the topological distribution of the bolt holes. Options with axial deviations exceeding 5 mm from the bolt holes were eliminated, ultimately selecting candidates that aligned with the bolt holes. Based on the main beam curvature continuity constraint, the selected connectors were morphologically optimized, adjusting their installation angle and contact surface curvature to align with the main beam curvature. For example, the connector installation angle was adjusted to 52°, and the contact surface curvature was adjusted to 0.15 to match the main beam curvature. Next, based on the load distribution characteristics of the support and the tower frame, the optimized connectors were reinforced to withstand environmental interference. For example, a 2mm-thick reinforcement layer was added to the contact area between the connector and the main beam to improve wind load resistance. Finally, the reinforced connector was geometrically matched with the tower skeleton point cloud to ensure precise alignment, generating a complete 3D recognition model of the power tower. This method successfully repaired the missing connector and significantly improved the tower's anti-interference capability, providing a reliable basis for subsequent maintenance.
[0153] In summary, steps 901 to 906 achieve precise positioning of the missing connection component area and generation of the missing region bounding box based on the curvature continuity constraint of the tower main beam and the topological distribution constraint of the bolt hole position. At the same time, the parameters of the candidate connection components are generated and optimized to ensure that they meet the thresholds of the axial deviation of the bolt hole position and the differential curvature of the main beam. Furthermore, through anti-environmental interference reinforcement modeling, the reinforcement strength of the connection components is dynamically adjusted to adapt to the load distribution characteristics of the support and the main beam. Finally, an environmentally-interference-resistant 3D recognition model of the power tower is generated, thereby achieving high-precision positioning and optimization of the missing connection components, as well as reliable reconstruction of the 3D model of the power tower and improving its anti-interference capability.
[0154] Figure 2 The present invention provides a structural diagram of a power tower identification system based on text symbols, as shown in FIG. Figure 2 As shown, the system includes: A construction module 21 constructs a multimodal mapping relationship between the text symbols and the tower body components based on the text symbols obtained on the surface of the power tower; The analysis module 22 obtains the original reflection signal of the power tower through the radar array, screens the target reflection area based on the multimodal mapping relationship, analyzes the phase difference data in the original reflection signal, and generates a penetration feature set including the tower skeleton outline and the corresponding text symbol positioning information; A compensation module 23 constructs a dynamic noise suppression model based on the penetration feature set, performs attenuation compensation on multipath reflection signals through the dynamic noise suppression model, and generates a compensation feature set for resisting electromagnetic interference; An updating module 24 inputs the compensation feature set and the coating geometric identification feature into a shared weight encoder for cross-modal association learning, so as to fuse the semantic vector of the text symbol through the shared weight encoder based on the attention mechanism, synchronously extract the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, and establish a geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; The generating module 25 reconstructs the missing tower connection components based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
[0155] Figure 2 The electric tower identification system based on text symbols can be executed Figure 1 The implementation principle and technical effects of the text-based power tower identification method described in the illustrated embodiment will not be elaborated on here. The specific manner in which each module and unit performs operations in the text-based power tower identification system in the above embodiment has been described in detail in the embodiments of the method and will not be elaborated on here.
[0156] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.
Claims
1. A method for identifying power towers based on text symbols, characterized in that: include: According to the text symbols obtained on the surface of the power tower, a multimodal mapping relationship between the text symbols and the tower components is constructed; Acquire the original reflection signal of the power tower through the radar array, screen the target reflection area based on the multimodal mapping relationship, analyze the phase difference data in the original reflection signal, and generate a penetration feature set including the tower skeleton outline and corresponding text symbol positioning information; Building a dynamic noise suppression model based on the penetration feature set, performing attenuation compensation on multipath reflection signals through the dynamic noise suppression model, and generating a compensation feature set for resisting electromagnetic interference; The compensation feature set and the coating geometric identification features are input into a shared weight encoder for cross-modal association learning. The shared weight encoder fuses the semantic vectors of text symbols based on the attention mechanism, synchronously extracts the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, and establishes the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; The missing tower connection components are reconstructed based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
2. The method according to claim 1, characterized in that The compensation feature set and the coating geometric identification feature are input into a shared weight encoder for cross-modal association learning, so as to fuse the text symbol semantic vector based on the attention mechanism through the shared weight encoder, synchronously extract the internal support structure features of the tower body and the external anti-corrosion coating texture features, and establish the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters, including: Inputting the compensation feature set and the coating geometric identification feature into a shared weight encoder, performing feature extraction on the compensation feature set and the coating geometric identification feature through a multi-layer network structure in the shared weight encoder to generate a preliminary feature vector; Performing weighted fusion of the preliminary feature vector and the text symbol semantic vector based on an attention mechanism to generate a fused feature vector; By performing multi-scale decomposition on the fused feature vector, the internal support structure features of the tower body and the texture features of the external anti-corrosion coating are extracted respectively, wherein the support structure features are extracted by high-frequency components and the coating texture features are extracted by low-frequency components; The convolution kernel parameters are alternately updated based on the local consistency optimization between the internal support structure features and the external anti-corrosion coating texture features, and a geometric constraint relationship between the support structure features and the coating texture features is established.
3. The method according to claim 2, characterized in that The preliminary feature vector and the text symbol semantic vector are weightedly fused based on the attention mechanism to generate a fused feature vector, including: Inputting the preliminary feature vector and the text symbol semantic vector into independent fully connected layers respectively to align feature dimensions, thereby generating aligned preliminary feature vectors and text symbol semantic vectors; Based on the spatial distribution of the aligned preliminary feature vector and the text symbol semantic vector, constructing local symbol association features and global structure association features, wherein the local symbol association features are generated by the feature mean within the symbol neighborhood, and the global structure association features are generated by the cross-region feature difference; Generate a symbol-structure association mapping vector by cross-projecting the local symbol association feature and the global structure association feature, and calculate a dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector based on the symbol-structure association mapping vector; Performing channel-level weight assignment on the aligned preliminary feature vectors according to the dynamic relevance scores, and performing symbol-level weight assignment on the text symbol semantic vectors to generate a channel-weighted preliminary feature vector and a symbol-weighted semantic vector; Performing multi-scale context feature splicing on the channel-weighted preliminary feature vector and the symbol-weighted semantic vector, and generating a multi-scale context feature splicing result, wherein the multi-scale context feature is extracted by sliding a context window of symbol semantics, and the window size is positively correlated with the density of text symbols on the surface of the power tower; Based on the topological relationship constraints of the power tower skeleton, bidirectional feature propagation is performed on the multi-scale context feature splicing result to generate a fused feature vector.
4. The method according to claim 3, characterized in that Generate a symbol-structure association mapping vector by cross-projecting the local symbol association feature and the global structure association feature, and calculate a dynamic correlation score between the aligned preliminary feature vector and the text symbol semantic vector based on the symbol-structure association mapping vector, including: The local symbol association features are respectively subjected to spatial position encoding and channel dimension encoding to generate spatially encoded local symbol association features and channel-encoded global structure association features; By interactively superimposing feature channels, the local symbol association features after spatial encoding and the global structure association features after channel encoding are bidirectionally projected to generate a symbol structure association mapping vector; Extracting a spatial weight matrix and a channel weight matrix from the symbol-structure association mapping vector, wherein the spatial weight matrix is generated based on the local similarity of symbol semantics, and the channel weight matrix is generated based on the global difference of structural features; By parameterizing and weighting the spatial topological constraints of the global structural association features, the spatial weight matrix and the channel weight matrix are fused to generate a dynamic correlation weight coefficient; Based on the dynamic relevance weight coefficient, the cross-modal association strength between the aligned preliminary feature vector and the text symbol semantic vector is quantified to generate a dynamic relevance score.
5. The method according to claim 2, characterized in that The convolution kernel parameters are alternately updated based on the local consistency between the internal support structure features and the external anti-corrosion coating texture features, and a geometric constraint relationship between the support structure features and the coating texture features is established, including: The internal support structure characteristics are decomposed into the spatial distribution characteristics of the connectors and the main beam morphology characteristics, and the external anti-corrosion coating texture characteristics are decomposed into the corrosion area distribution characteristics and the coating thickness variation characteristics; Based on the proximity constraint between the bolt hole position and the coating defect area, the spatial distribution characteristics of the connector and the distribution characteristics of the corrosion area are spatially topologically encoded to generate the corrosion correlation characteristics of the connector; Based on the synchronization constraint of the tower body curvature change and the coating thickness gradient, the main beam morphological characteristics and coating thickness change characteristics are coded by morphological thickness coupling to generate the main beam thickness coupling characteristics. Constructing a multi-branch convolution kernel group, wherein the first branch performs spatial constraint modeling on the corrosion-related features of the connector, and the second branch performs texture constraint modeling on the thickness coupling features of the main beam; Based on the local consistency loss of the spatial constraint modeling results and the texture constraint modeling results, the parameters of the multi-branch convolution kernel group are alternately updated to generate a geometric constraint relationship between the support structure features and the coating texture features.
6. The method according to claim 1, characterized in that The original reflection signal of the power tower is obtained through the radar array, the target reflection area is screened based on the multimodal mapping relationship, the phase difference data in the original reflection signal is analyzed, and a penetration feature set including the tower skeleton outline and the corresponding text symbol positioning information is generated, including: A circular or linearly deployed radar array collects multi-angle signals from the power tower to obtain a set of original reflected signals. During signal collection, the radar beam direction is aligned with the spatial distribution of text symbols on the tower surface. Based on the multimodal mapping relationship, beamforming processing is performed on the original reflection signal set to screen a target reflection area containing text symbol positioning information, wherein beamforming parameters corresponding to the beamforming processing are dynamically adjusted according to the size of the text symbol and the geometric size of the tower body component; Performing phase difference interferometry processing on the original reflection signal of the target reflection area, analyzing the phase difference data between adjacent radar channels, and generating an initial three-dimensional point cloud of the tower skeleton outline; Based on the component association of the text symbols in the multimodal mapping relationship, the initial three-dimensional point cloud is spatially filtered to eliminate scattered points irrelevant to the positioning of the text symbols, thereby generating a symbol-constrained tower skeleton point cloud; The symbol-constrained tower skeleton point cloud is feature-fused with the multipath reflection intensity in the previous phase difference interferometry processing to generate a penetrating feature set containing text symbol positioning information and tower skeleton geometric parameters.
7. The method according to claim 6, characterized in that Based on the component association of the text symbols in the multimodal mapping relationship, the initial three-dimensional point cloud is spatially filtered to remove scattered points irrelevant to the positioning of the text symbols, and a symbol-constrained tower skeleton point cloud is generated, including: Based on the association between the text symbols and the tower components in the multimodal mapping relationship, dividing the symbol semantically associated areas and non-associated areas; In the symbol semantic association area, the scattered points are processed according to the component association of the text symbol. In the non-symbol semantic association area, the scattered points that are not related to the corrosion expansion characteristics are eliminated according to the defect boundary distribution of the coating corrosion symbol. Marking the initial three-dimensional point cloud with symbol semantic association areas, wherein the lightning protection symbol area radiates outward from the center of the symbol, the support symbol area extends along the main beam axis, and the fixed symbol area extends according to the bolt distribution density matrix, to generate a regional marking point cloud; Based on the multi-scale precision parameters of the regional marker point cloud, the scattered points processed in the symbolic semantic associated area and the remaining scattered points after the non-associated area is eliminated are multi-scale aggregated to generate a symbol-constrained tower skeleton point cloud.
8. The method according to claim 1, characterized in that A dynamic noise suppression model is constructed based on the penetration feature set, and attenuation compensation is performed on multipath reflection signals using the dynamic noise suppression model to generate a compensation feature set for electromagnetic interference resistance, including: Extracting symbol semantic noise suppression parameters through the spatial mapping relationship between the text symbol positioning information in the penetration feature set and the tower body components, and constructing a dynamic noise suppression model; In the dynamic noise suppression model, the structural reflection path attenuation coefficient is dynamically adjusted based on the geometric parameters of the tower skeleton outline; Generate an interference source spatial weight map based on the electromagnetic interference distribution characteristics in the penetration feature set and the azimuth information of the high-voltage equipment around the power tower, integrate the interference source spatial weight map into the dynamic noise suppression model, perform azimuth-related attenuation compensation on the multipath reflection signal, and extract azimuth compensation parameters; The symbolic semantic noise suppression parameter, the structural reflection path attenuation coefficient and the azimuth compensation parameter are weightedly integrated in the tower region in the dynamic noise suppression model to update the dynamic noise suppression model; Iterative path attenuation compensation is performed on the multipath reflection signals in the penetration feature set by using the updated dynamic noise suppression model to generate a compensation feature set for resisting electromagnetic interference.
9. The method according to claim 1, characterized in that Reconstructing the missing tower connection components based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference includes: Based on the curvature continuity constraint of the tower main beam and the topological distribution constraint of the bolt hole positions in the geometric constraint relationship, the area where the connection parts are missing is spatially located to generate a bounding box of the missing area; Lock the range within the missing area boundary box, and generate candidate connection component parameters based on the angle constraint relationship between the main beam and the support member; Performing spatial matching verification on the candidate connection component parameters based on the bolt hole position topological distribution constraint, eliminating candidate parameters whose axial deviation from the bolt hole position exceeds a preset threshold, and generating a preliminary connection component set; Optimizing the shape of the preliminary set of connection components according to the curvature continuity constraint of the tower main beam, adjusting the installation angle and contact surface curvature of the connection components to meet the curvature difference threshold of the main beam skeleton point cloud; Based on the load distribution characteristics of the support members and the tower frame, an anti-environmental interference reinforcement model is performed on the connection components after the shape optimization, wherein the reinforcement strength is dynamically adjusted according to the connection density between the support members and the main beam; The reinforced modeled connection components are matched with the tower skeleton point cloud through geometric constraints to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
10. A power tower identification system based on text symbols, characterized in that: include: A construction module, which constructs a multimodal mapping relationship between text symbols and tower components based on the text symbols obtained on the surface of the power tower; An analysis module acquires the original reflection signal of the power tower through a radar array, screens the target reflection area based on the multimodal mapping relationship, analyzes the phase difference data in the original reflection signal, and generates a penetration feature set including the tower skeleton outline and the corresponding text symbol positioning information; a compensation module, constructing a dynamic noise suppression model based on the penetration feature set, performing attenuation compensation on multipath reflection signals through the dynamic noise suppression model, and generating a compensation feature set for resisting electromagnetic interference; An updating module inputs the compensation feature set and the coating geometric identification feature into a shared weight encoder for cross-modal association learning, so as to fuse the semantic vector of the text symbol through the shared weight encoder based on the attention mechanism, synchronously extract the internal support structure features of the tower body and the texture features of the external anti-corrosion coating, and establish the geometric constraint relationship between the support structure and the coating texture by alternately updating the convolution kernel parameters; A generation module reconstructs the missing tower connection components based on the geometric constraint relationship to generate a three-dimensional recognition model of the power tower that is resistant to environmental interference.
Citation Information
Patent Citations
Analytical method for design image of power transmission tower
CN119479002A
Multi-modal data fusion for enhanced 3D perception for platforms
US20200184718A1
Three-dimensional target detection method based on multimodal fusion and depth attention mechanism
US20250037299A1
Point cloud data simplification method for power line towers, and system
WO2025098506A1