Power transmission channel hidden danger inspection system based on cooperation of computer vision and unmanned aerial vehicle

By deeply fusing spatiotemporal alignment and semantic geometric constraints of multi-source sensors, the pose drift problem of UAV inspection systems in ultra-high voltage strong magnetic fields and complex field environments has been solved, achieving centimeter-level precise positioning of potential targets and improving the robustness and accuracy of the system.

CN121962264APending Publication Date: 2026-05-01NANYANG POWER SUPPLY COMPANY OF STATE GRID HENAN ELECTRIC POWER
View PDF 0 Cites 2 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANYANG POWER SUPPLY COMPANY OF STATE GRID HENAN ELECTRIC POWER
Filing Date
2026-01-27
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing UAV inspection systems are affected by high-frequency mechanical vibration and strong magnetic field interference in ultra-high voltage strong magnetic fields and complex field environments, resulting in pose estimation drift and three-dimensional coordinate deviation of potential hazards, making it difficult to meet the requirements of high-precision positioning.

Method used

By deeply fusing spatiotemporal alignment and semantic geometric constraints from multi-source sensors, a spatiotemporal alignment mechanism for multi-source heterogeneous data is constructed. This mechanism can detect electromagnetic environment anomalies in real time, generate interference state markers, and construct a weighted mask using deep semantic perception and topological geometric priors to shield dynamic feature interference. Combined with adaptive adjustment of the noise covariance matrix from inertial navigation pre-integration, the six-degree-of-freedom pose is optimized using factor graphs to achieve centimeter-level precise positioning of potential hazards.

Benefits of technology

Robust pose calculation and hazard location were achieved in harsh environments, improving the system's positioning accuracy and stability, and meeting the requirements for high-precision hazard three-dimensional coordinate calculation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962264A_ABST
    Figure CN121962264A_ABST
Patent Text Reader

Abstract

The invention provides a power transmission channel hidden danger inspection system based on cooperation of computer vision and an unmanned aerial vehicle, and relates to the technical field of automatic inspection of power lines. Firstly, a synchronous measurement data set is constructed through space-time alignment and magnetic field anomaly detection of multi-source sensor data, and dynamic feature interference such as vegetation is shielded by using a semantic perception technology; and only the rigid area of the transmission tower is locked. On the basis, truss topology prior is introduced to carry out physical node extraction on a tower structure, a high-robustness visual anchor point factor is constructed, meanwhile, noise covariance of inertial navigation pre-integration is dynamically adjusted according to the electromagnetic interference state, and robust state recursion is achieved. And finally, carrying out joint optimization solution by constructing a factor graph containing inertial navigation pre-integration, visual anchor points and tower geometric constraints, obtaining a high-precision six-degree-of-freedom pose while suppressing an accumulative error, and reversely projecting a hidden danger target in the image to a three-dimensional space by using the pose, thereby realizing accurate positioning of the hidden danger of the power transmission channel.
Need to check novelty before this filing date? Find Prior Art

Description

A power transmission corridor hazard inspection system based on computer vision and drone collaboration Technical Field

[0001] This application relates to the field of automated power line inspection technology, and more specifically, to a power transmission channel hidden danger inspection system based on computer vision and drone collaboration. Background Technology

[0002] With the increasing demand for electricity, transmission line inspection has become a core component of ensuring the safe operation of the power grid. Traditional inspection methods mainly rely on manual climbing or long-distance observation from the ground, which is not only labor-intensive and risky, but also has many blind spots and low efficiency. To improve the intelligence and automation of inspections, the development of a transmission channel hazard inspection system based on computer vision and drone collaboration has become an inevitable trend in the industry. The system aims to achieve centimeter-level accurate positioning and defect elimination decisions for hazards such as vegetation encroachment through high-precision visual perception and edge pose calculation.

[0003] However, the robustness of existing inspection systems in the face of strong magnetic fields and complex field environments under ultra-high voltage (UHV) conditions still needs improvement. Near UHV transmission lines, strong electromagnetic interference can cause magnetometer readings to become disordered or saturated, rendering the heading angle correction reference ineffective. Furthermore, the broadband mechanical vibrations generated by UAVs adjusting motor speeds at high frequencies can easily couple into the inertial navigation unit (INS) readings, leading to severe drift in pose estimation. At the visual perception level, traditional visual odometry often assumes the environment is a static rigid body. However, in actual inspection scenarios, dynamic feature points such as swaying trees caused by wind can be mistaken by the algorithm for static anchor points, resulting in a visual illusion of reverse motion compensation. A deeper problem lies in the fact that existing solutions often rely solely on pixel-level gradient responses during feature extraction, lacking prior utilization of the physical topological characteristics of the transmission tower truss structure. This leads to the extraction of a large number of non-structural feature points on the tower surface, such as rust spots and shadow boundaries, which are prone to feature drift or loss under motion blur or drastic changes in lighting. This disconnect between semantic perception and physical geometric constraints makes it difficult for the system to maintain a stable six-degree-of-freedom pose under strong interference environments. This leads to serious deviations in the calculated three-dimensional coordinates of potential hazards, failing to meet the power industry's need for high-precision inspection data. Therefore, an optimized power transmission channel hazard inspection system based on computer vision and UAV collaboration is desired. Summary of the Invention

[0004] To address the aforementioned technical problems, this application is proposed. Embodiments of this application provide a power transmission channel hazard inspection system based on computer vision and UAV collaboration, comprising: a data acquisition and synchronization module, used for spatiotemporal alignment and magnetic field anomaly detection of raw visual streams, raw inertial navigation readings, raw magnetometer readings, and raw satellite observations acquired by the UAV to obtain synchronized multi-source data frames and electromagnetic interference status markers; a semantic feature extraction module, used for semantic perception and dynamic feature mask construction of image data in the synchronized multi-source data frames to obtain a set of rigid feature points containing only tower structures; and a robust pre-integration module, used for, based on the electromagnetic interference status markers, to perform semantic perception and dynamic feature mask construction on the inertial navigation data in the synchronized multi-source data frames. The data undergoes robust integration to obtain the inertial navigation pre-integration factor, which includes the relative pose increment. A geometric constraint construction module performs spatial registration and triangulation on the rigid feature point set and a pre-set standard tower model to obtain tower geometric constraints and visual anchor point factors. A pose optimization solution module performs factor graph optimization and robust pose solution on the inertial navigation pre-integration factor, visual anchor point factor, and tower geometric constraints to obtain an optimized six-degree-of-freedom pose. A hazard location output module locates the hazard target pixel from the synchronized multi-source data frames and, combined with the optimized six-degree-of-freedom pose, inversely calculates the true three-dimensional coordinates of the hazard point to obtain the hazard's three-dimensional coordinates.

[0005] Compared with existing technologies, the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration provided in this application first performs spatiotemporal alignment and magnetic field anomaly detection on multi-source data such as vision, inertial navigation, and satellite. It then uses semantic reasoning to generate a weighted mask to shield against dynamic feature interference such as trees, and extracts rigid line segments and nodes of towers based on truss topology priors, establishing absolute geometric constraints through registration with a standard model. On this basis, it adaptively adjusts the trust weight of the inertial navigation pre-integration according to the magnetic field state, and uses factor graphs to jointly optimize visual anchor points, geometric priors, and inertial navigation data, correcting heading deviations while suppressing accumulated errors. Finally, combining the optimized six-DOF pose and camera intrinsic parameters, it performs multi-frame epipolar matching and inverse projection on hidden danger targets in the image, achieving centimeter-level accuracy in calculating the three-dimensional coordinates of hidden dangers. Attached Figure Description

[0006] The above and other objects, features, and advantages of this application will become more apparent from the more detailed description of the embodiments of this application in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this application and form part of the specification. They are used together with the embodiments of this application to explain this application and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.

[0007] Figure 1 is a system block diagram of a power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application; Figure 2 is a data flow diagram of the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application; Figure 3 is a block diagram of the semantic feature extraction module in the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application; Figure 4 is a block diagram of the feature filtering unit in the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application; Figure 5 is a block diagram of the robust pre-integration module in the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application; Figure 6 is a block diagram of the geometric constraint construction module in the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. Detailed Implementation

[0008] Hereinafter, exemplary embodiments according to this application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of this application, and not all embodiments of this application. It should be understood that this application is not limited to the exemplary embodiments described herein.

[0009] As indicated in this application and claims, unless the context clearly indicates otherwise, the words "a," "an," "an," and / or "the" are not specifically singular and may include plural forms. Generally speaking, the terms "comprising" and "including" only indicate the inclusion of explicitly identified steps and elements, which do not constitute an exclusive list, and the method or apparatus may also include other steps or elements.

[0010] While this application makes various references to certain modules of the systems according to embodiments of this application, any number of different modules can be used and run on user terminals and / or servers. The modules described are merely illustrative, and different aspects of the systems and methods may use different modules.

[0011] Flowcharts are used in this application to illustrate the operations performed by the system according to embodiments of this application. It should be understood that the preceding or following operations are not necessarily performed in exact order. Instead, various steps can be processed in reverse order or simultaneously as needed. Furthermore, other operations can be added to these processes, or one or more steps can be removed from them.

[0012] Currently, UAV inspection systems are highly susceptible to noise interference from high-frequency mechanical vibrations on inertial measurement units and interference from strong magnetic fields on magnetometer readings in ultra-high voltage strong magnetic fields and complex field environments. This leads to inaccurate estimation of camera extrinsic parameters and severe pose drift, resulting in significant deviations in the calculated 3D coordinates of potential hazards, making it difficult to meet high-precision positioning requirements. Therefore, this application proposes a power transmission channel hazard inspection system based on computer vision and UAV collaboration. This system achieves robust pose calculation and hazard localization in harsh environments through deep fusion of spatiotemporal alignment and semantic geometric constraints from multiple sensor sources. Specifically, the system first constructs a spatiotemporal alignment mechanism for multi-source heterogeneous data to detect electromagnetic environment anomalies in real time and generate interference state markers. Then, it uses deep semantic perception and topological geometric priors to construct a weighted mask, which forcibly shields dynamic feature interference such as vegetation during the feature extraction stage, locking only the rigid truss structure of the transmission tower as a reliable anchor point. At the same time, it dynamically adjusts the noise covariance matrix of the inertial navigation pre-integration in combination with the electromagnetic interference state, adaptively reducing the magnetometer weight and relying on robust state recursion when the magnetic field is disturbed. Finally, it constructs a factor graph containing inertial navigation pre-integration factors, visual anchor point factors, and tower geometric constraint factors for joint nonlinear optimization, which solves a high-precision six-degree-of-freedom pose while suppressing accumulated errors. Based on this, it inversely projects the hidden danger targets on the image plane into the real three-dimensional space, thereby achieving centimeter-level accurate positioning of hidden dangers in the transmission channel.

[0013] Figure 1 is a system block diagram of a power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. Figure 2 is a data flow diagram of the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. As shown in Figures 1 and 2, the power transmission channel hidden danger inspection system 100 based on computer vision and UAV collaboration according to an embodiment of this application includes: a data acquisition and synchronization module 110, used to perform spatiotemporal alignment and magnetic field anomaly detection on the original visual stream, original inertial navigation readings, original magnetometer readings and original satellite observations acquired by the UAV to obtain synchronized multi-source data frames and electromagnetic interference status markers; a semantic feature extraction module 120, used to perform semantic perception and dynamic feature mask construction on the image data in the synchronized multi-source data frames to obtain a set of rigid feature points containing only tower structures; and a robust pre-integration module 130, used to perform semantic perception and dynamic feature mask construction on the synchronized multi-source data based on the electromagnetic interference status markers. The inertial navigation data in the frame is robustly integrated to obtain the inertial navigation pre-integration factor containing the relative pose increment; the geometric constraint construction module 140 is used to perform spatial registration and triangulation processing on the rigid feature point set and the pre-set standard tower model to obtain the tower geometric constraints and visual anchor point factors; the pose optimization solution module 150 is used to perform factor graph optimization and robust pose solution on the inertial navigation pre-integration factor, visual anchor point factor and tower geometric constraints to obtain the optimized six-degree-of-freedom pose; the hazard location output module 160 is used to locate the hazard target pixel from the synchronized multi-source data frame, and combine it with the optimized six-degree-of-freedom pose to back-calculate the true three-dimensional coordinates of the hazard point to obtain the hazard three-dimensional coordinates.

[0014] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the data acquisition synchronization module 110 is used to perform spatiotemporal alignment and magnetic field anomaly detection on the raw visual flow, raw inertial navigation readings, raw magnetometer readings, and raw satellite observations collected by the UAV to obtain synchronized multi-source data frames and electromagnetic interference status markers. It should be noted that due to the strong electromagnetic fields generated by currents up to several thousand amperes in the ultra-high voltage power transmission channel environment, and the high-frequency mechanical vibrations generated by the UAV under outdoor airflow disturbances, these vibrations couple to the inertial measurement unit, resulting in spatial distortion and microsecond-level asynchronous deviations in the multi-source sensor data. Directly fusing the raw data would cause state divergence in the visual inertial odometry. Based on this, the technical solution of this application first constructs a hardware clock synchronization and anti-interference preprocessing mechanism. Optical distortion correction is performed on the original visual stream, and this is used as the time base axis to align the inertial navigation, magnetometer, and satellite observations (with nearest neighbor alignment used for satellite observations). Simultaneously, electromagnetic interference regions are identified through magnetic field modulus fluctuation detection, thereby establishing a spatiotemporally unified synchronous measurement dataset with environmental awareness capabilities. Through the above processing, the cumulative errors caused by sensor intrinsic parameter drift and timestamp misalignment can be eliminated, and a decision-making basis for shielding unreliable magnetometer data can be provided for subsequent optimization algorithms, thus effectively ensuring the robustness of pose calculation.

[0015] More specifically, in a specific example of this application, the data acquisition synchronization module 110 includes: an image correction unit, used to perform distortion correction processing on the original visual stream based on the camera intrinsic parameter matrix and distortion model parameters, and extract the exposure timestamp to obtain a distortion-corrected image frame and an image timestamp; a time alignment unit, used to perform linear interpolation and time alignment on the original inertial navigation readings, original magnetometer readings, and original satellite observations with the image timestamp as the main clock reference to obtain a synchronized measurement dataset containing the distortion-corrected image frame; and an anomaly detection unit, used to perform electromagnetic environment anomaly detection and status marking based on the magnetometer data in the synchronized measurement dataset to output synchronized multi-source data frames and electromagnetic interference status markings.

[0016] In other words, more specifically, the execution process follows a rigorous logic from single-sensor calibration to multi-sensor fusion. First, intrinsic parameter calibration and time stamp extraction of the image stream are performed. For the raw visual stream acquired by the UAV monocular camera, the Brown-Conrady model is used to remap the pixel coordinates of each frame of the image using a pre-calibrated camera intrinsic parameter matrix and distortion model parameters including radial and tangential distortion coefficients. This eliminates barrel or pincushion distortion caused by wide-angle lenses. During the calculation, the distortion coordinates are calculated by combining ideal pixel coordinates with an even-power polynomial of radial distance and a tangential correction term. After pixel reconstruction, the precise time of the exposure center of the frame is extracted as the image timestamp.

[0017] Subsequently, time-domain alignment of multi-source data is performed, and the image timestamp is set as the system's master clock reference. Considering that the sampling frequency of the inertial navigation unit and magnetometer is much higher than that of the visual sensor, a linear interpolation algorithm is used to unify the heterogeneous sensor data to the visual frame time. Specifically, the operation involves searching for the previous frame data and the next frame data whose timestamps are adjacent to the current image timestamp in the original inertial navigation reading sequence and the original magnetometer reading sequence, calculating the ratio of the time difference to the numerical difference between the two frames, and calculating the inertial navigation vector and magnetometer vector that are strictly aligned with the exposure time by combining the difference between the current image timestamp and the previous frame timestamp. For satellite observations with lower sampling frequencies, the observation data with the nearest time neighbor is selected, thereby generating a synchronous measurement dataset containing the distortion-free image frame and the aligned sensor data.

[0018] Finally, the system performs electromagnetic environment anomaly detection and status marking. It reads the aligned magnetometer data vector from the synchronous measurement dataset, calculates the Euclidean norm of the vector to obtain the current environmental magnetic field magnitude, and compares this magnitude with the local geomagnetic model reference value. At the same time, it calculates the statistical variance of the magnetic field magnitude within a time window. When the deviation between the magnetic field magnitude and the reference value exceeds a preset threshold or the magnetic field variance shows drastic fluctuations, it is determined that the system is in a strong electromagnetic interference area and the electromagnetic interference status mark is set to high level. Otherwise, it is marked as normal. Finally, the mark is appended to the synchronous data frame for output, so that the subsequent state estimation module can dynamically adjust the confidence weight of the heading angle observation value when constructing the noise covariance matrix.

[0019] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the semantic feature extraction module 120 is used to perform semantic perception and dynamic feature mask construction on image data in synchronized multi-source data frames to obtain a set of rigid feature points containing only the tower structure. It should be noted that, given the extremely complex visual environment of power transmission channels in the field, images are filled with non-rigid dynamic features such as trees and vegetation swaying in the wind, flowing clouds, and changing light projections. If the traditional visual odometry method is used to perform indiscriminate feature tracking on the entire image, it is very easy to misjudge dynamic interference in the environment as the movement of the UAV itself, which will lead to serious cumulative drift or even divergence in pose estimation. Based on this, the technical solution of this application further introduces a deep semantic perception mechanism. Pixel-level semantic reasoning and attribute classification are performed on the image data in the synchronized multi-source data frames. A dynamic feature mask capable of shielding dynamic noise is constructed based on a pre-set category stability mapping table. This mask is then used to force the feature extraction algorithm to search and lock features only within the absolutely static rigid structural region of the transmission tower. This achieves data cleaning and confidence grading from the source. Through the above processing, the problem of visual feature degradation in dynamic scenes can be effectively solved, ensuring that the output rigid feature point set has extremely high physical stability, thereby improving the system's localization robustness in complex field environments.

[0020] Figure 3 is a block diagram of the semantic feature extraction module in a power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. As shown in Figure 3, the semantic feature extraction module 120 includes: a semantic reasoning unit 121, used to perform image tensor preprocessing and pixel-level semantic reasoning on the image data in the synchronized multi-source data frames to obtain a semantic label map; a mask generation unit 122, used to map the pixel categories in the semantic label map to differentiated numerical weights according to a preset category stability mapping table to obtain a semantic weighted mask; and a feature filtering unit 123, used to perform feature extraction and filtering under rigid region constraints on the image data in the synchronized multi-source data frames based on the semantic weighted mask to obtain a rigid feature point set.

[0021] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the semantic reasoning unit 121 is used to perform image tensor preprocessing and pixel-level semantic reasoning on the image data in the synchronized multi-source data frames to obtain a semantic label map. It should be noted that, due to the extremely complex visual background of the power transmission channel inspection scene, the original image data only contains the light intensity values ​​of the red, green, and blue channels, lacking semantic understanding of the physical attributes of pixels. If directly used for subsequent processing, the algorithm will struggle to distinguish between key facilities such as power transmission towers and interfering elements such as trees and sky in the background. Therefore, the technical solution of this application further performs standardized tensor transformation and deep learning semantic reasoning on the image data in the synchronized multi-source data frames to map unstructured image pixels into a set of semantic labels with clear physical category attributes. Through the above processing, pixel-level prior information can be provided for subsequent dynamic feature removal and rigid structure locking, thereby effectively improving the visual perception module's ability to understand complex environments.

[0022] More specifically, in a concrete example of this application, the processing flow encompasses the entire process from data formatting to deep feature parsing. First, tensor preprocessing of the image data is performed, parsing the original red-green-blue image matrix from the synchronized data frame. Based on a predefined neural network input specification, the image matrix is ​​numerically normalized, scaling pixel intensity values ​​to a specific range and subtracting the statistical mean to eliminate distribution bias caused by differences in illumination intensity. Simultaneously, the image resolution is adjusted to fit the network's receptive field size, reshaping it into four-dimensional tensor data conforming to batch dimensions. Subsequently, pixel-level semantic inference operations are performed, inputting the preprocessed tensor into a pre-trained lightweight semantic segmentation neural network model. This model extracts texture and structural features of the image layer by layer during the encoder stage and fuses multi-scale features using skip connections during the decoder stage to recover spatial details. Finally, it outputs a probability feature map containing multiple channels, where each channel corresponds to a predefined semantic category probability response. Finally, semantic labels are generated. The maximum index operation is applied to the channel dimension of the probabilistic feature map to determine the category index with the highest probability at each spatial location, and it is mapped to a specific semantic label map to achieve pixel-level classification of different ground features such as poles, power lines, vegetation and sky.

[0023] Specifically, the construction and training process of the pre-trained lightweight semantic segmentation neural network model is as follows. This embodiment constructs a dedicated semantic segmentation dataset for power transmission channels, covering various lighting conditions (such as backlighting and sidelighting) and complex backgrounds (such as mountains and forests). This dataset contains a large number of manually annotated sample images at the pixel level, with the annotation categories strictly divided into power transmission towers, power lines, vegetation, sky, and other backgrounds. During the training phase, an end-to-end supervised learning approach is adopted, selecting a combination of weighted cross-entropy loss function and Dice coefficient loss function as the objective function to address the sample class imbalance problem caused by the small pixel proportion of fine linear targets such as power lines in the image. Simultaneously, data augmentation strategies such as random scaling, cropping, rotation, and color dithering are introduced to simulate the perspective changes and imaging differences during drone flight. Finally, multiple rounds of iterative training are performed on a high-performance computing platform using stochastic gradient descent or Adam optimization algorithms until the model's average intersection-over-union ratio on the validation set reaches the preset industrial-grade accuracy requirement. The optimal weight parameters are then saved as a pre-trained model for online inference.

[0024] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the mask generation unit 122 is used to map the pixel categories in the semantic label image to differentiated numerical weights according to a preset category stability mapping table to obtain a semantically weighted mask. It should be noted that different environmental elements in the power transmission channel inspection scenario exhibit drastically different physical motion characteristics. While the main structure of the power transmission tower remains absolutely rigid and stationary, surrounding vegetation, clouds, and water surfaces are easily deformed or displaced under natural wind loads or airflow. If all image regions are given the same level of confidence in subsequent visual processing, the unstable features generated by dynamic regions will introduce incorrect geometric constraints, thereby interfering with pose calculation. Based on this, the technical solution of this application further maps the pixel categories in the semantic label image to differentiated numerical weights according to a preset category stability mapping table to obtain a semantically weighted mask, thereby quantifying the confidence level of different semantic regions as static geometric reference anchor points. Through the above processing, a spatial attention mechanism based on physical properties can be effectively established to suppress dynamic environmental noise at the data source level, ensuring that subsequent feature extraction and pose optimization focus on highly stable rigid structural regions.

[0025] More specifically, in a concrete example of this application, the implementation process follows a logical path from rule definition to matrix generation. First, a category stability mapping table containing prior physical knowledge is established. This table defines in detail the stability weight values ​​corresponding to each semantic label category. Categories with rigid body properties, such as towers, hardware, and insulators, are mapped to the maximum weight value representing high confidence; categories with geometric regularity but potentially slight swaying, such as power lines, are mapped to intermediate weight values ​​representing medium confidence; and volatile or textureless categories, such as trees, weeds, clouds, and sky backgrounds, are mapped to the minimum weight value representing low or zero confidence. Then, a weight mapping operation is performed across the entire image, traversing each pixel in the semantic label image row by row and column by column, reading the semantic category index value of that pixel, and retrieving the corresponding stability weight from the mapping table. Finally, based on the retrieval results, a single-channel or multi-channel numerical matrix with strictly consistent dimensions with the original image is constructed. The weight values ​​of each pixel position are filled into the corresponding coordinates, thereby generating a semantically weighted mask that intuitively reflects the physical stability distribution of the scene, providing accurate pixel-level filtering credentials for feature point selection in subsequent steps.

[0026] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the feature filtering unit 123 is used to extract and filter features from image data in synchronized multi-source data frames under rigid region constraints based on semantic weighted masks to obtain a rigid feature point set. It should be noted that, given that in actual field inspection environments, the surface of power transmission towers often has interference factors such as rust spots, shadow boundaries formed by light projection, and attachments, traditional feature extraction algorithms that rely solely on pixel gradients are prone to misjudging such unstructured pseudo-corner points as feature points. Furthermore, when high-frequency vibrations of UAVs cause motion blur in the image, these texture-based feature points are easily drifted or lost, failing to meet the demand for highly robust positioning features in complex industrial environments. Based on this, the technical solution of this application further introduces a truss topological saliency field mechanism within the region of interest defined by the semantically weighted mask. Instead of directly searching for corner points, it prioritizes structured line segment detection and noise removal to identify the set of vector line segments constituting the tower skeleton. The intersections of these line segments are marked as candidate physical nodes to construct a topologically weighted response map, which is then subjected to structure-aware non-maximum suppression and feature locking. This weighted fusion of simple image visual saliency and the Euclidean distance from feature points to physical skeleton intersections forces the introduction of a more reliable topological prior closer to the skeleton intersection, achieving a deep fusion of semantic awareness and geometric constraints. Through this processing, noise interference from points with large gradients but far from structural intersections can be effectively suppressed, ensuring that each locked rigid feature point possesses both texture clarity and structural stability, firmly anchored to the physical nodes of the tower. This improves the robustness of feature points under weak textures, dynamic blur, and surface contamination interference, providing reliable underlying data support for subsequent high-precision visual odometry calculations.

[0027] Figure 4 is a block diagram of the feature screening unit in the power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. As shown in Figure 4, the feature screening unit 123 includes: a line segment detection subunit 1231, used to perform line segment detection and noise removal on the image data in the synchronized multi-source data frame based on the binarized region constraint in the semantic weighted mask to obtain the truss line segment set of the power transmission tower; a topology construction subunit 1232, used to take the intersection points of the line segments in the truss line segment set as candidate physical nodes, and construct a topology weighted response map based on the candidate physical nodes; and a feature locking subunit 1233, used to perform structure-aware nonmaximum suppression and feature locking on the topology weighted response map to obtain a set of rigid feature points.

[0028] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the line segment detection subunit 1231 is used to perform line segment detection and noise removal on the image data in the synchronized multi-source data frames based on the binarized region constraints in the semantic weighted mask to obtain the truss line segment set of the power transmission tower. It should be noted that, given that the tower structure is assembled from a steel frame, in the complex field inspection environment, rust spots on the tower surface, shadow boundaries formed by light projection, and attachments often generate strong gradient responses. If the pixel-level gradient threshold is simply relied upon to filter corner points within the semantic mask, ignoring the inherent topological geometric characteristics of the power transmission tower as a typical truss structure, it is very easy to misjudge pseudo-corner points as rigid feature points, leading to divergence in subsequent pose calculations. Based on this, the technical solution of this application further implements a structured feature extraction method based on truss topological saliency field. Within the region of interest defined by a semantically weighted mask, instead of directly searching for corner points, it prioritizes structured line segment detection and noise removal to enhance the physical stability of semantic features using geometric prior knowledge. Through the above processing, a geometric skeleton with rigid constraints can be extracted from the chaotic texture, laying the foundation for subsequent physical node finding. This effectively solves the feature drift problem in weak texture and dynamic blur scenes, meeting the demand for highly robust localization features in complex industrial environments.

[0029] More specifically, in a concrete example of this application, this step first acquires grayscale image data from the synchronized data frame, and then performs a masking operation on the grayscale image using a generated semantically weighted mask, i.e., performs a Hadamard product operation, to filter out background interference and retain texture information within the region of interest. Subsequently, a line segment detection operator is used to identify the set of vector line segments constituting the tower skeleton, specifically represented as follows: in, This represents the set of truss segments that remain after filtering and have structural support significance. To synchronize grayscale image data in data frames, For semantic weighting mask, This represents a masking operation (Hadamard product) used to filter out background interference. For line segment detection operators (such as Line Segment Detector). A minimum structural scale threshold is set to remove broken short line noise that is not part of the image skeleton. In practical applications, such as when facing a power transmission tower with localized rust and complex shadows caused by side-backlighting, the semantic weighted mask first sets the sky background and vegetation behind the tower to zero, retaining only the grayscale value of the tower area. The line segment detection operator extracts a large number of line segment candidates within the mask area, including both long straight lines at the edges of steel and short line segments at the edges of rust spots and broken shadows. By applying the aforementioned length threshold constraint, the algorithm automatically discards short line noise with a length lower than a certain pixel value. The final output set of truss line segments clearly outlines the "V" or "X" shaped steel structure skeleton of the tower, eliminating interference from non-physical structural textures and ensuring that the subsequently extracted intersections strictly correspond to the actual bolt connections or steel weld points.

[0030] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the topology construction subunit 1232 is used to select the intersection points of line segments in the truss segment set as candidate physical nodes and construct a topology weighted response map based on the candidate physical nodes. It should be noted that, given that in actual field inspection environments, rust spots on tower surfaces, shadow boundaries formed by light projections, and attachments (such as the edges of bird nests) often generate strong gradient responses, they are misjudged as rigid feature points by the original algorithm due to a lack of physical constraints. These pseudo-corner points do not possess structural stability physically, and directly using them for pose calculation will lead to error accumulation. Based on this, the technical solution of this application further establishes a structural saliency field to quantify the physical confidence of feature points. By systematically traversing the extracted set of truss segments, the mathematical intersections of the extensions of any non-parallel segments are calculated and marked as candidate physical nodes. Based on these physical nodes, a spatial weighting mechanism is constructed to fuse the simple image gradient response with the Euclidean distance from the feature point to the physical node. This suppresses noise points (such as rust) that have large gradients but are far from the structural intersections, while enhancing the weight of feature points located at the intersections of the tower steel. Through the above processing, a topological prior that is more reliable the closer to the skeleton intersection is can be forcibly introduced, achieving alignment of the physical meaning of features, thereby effectively solving the problem of pseudo-corner interference and ensuring that feature points are tightly anchored to the physical structure of the tower.

[0031] More specifically, in a concrete example of this application, this step first receives the set of truss segments generated in the previous stage. Using an analytical geometry algorithm, it performs coplanarity determination and extension calculations on each pair of segments within the set, obtaining the theoretical intersection coordinates of all non-parallel segments on the two-dimensional image plane. These coordinate points are then defined as candidate physical nodes of the transmission tower skeleton. Subsequently, for each pixel in the image, it not only calculates its original gradient response based on pixel grayscale changes but also calculates the Euclidean distance between that pixel and all candidate physical nodes. Using a Gaussian kernel function, the distance information is transformed into topological weights, constructing a structural saliency field that incorporates physical constraints. The calculation process is as follows: in, Represents pixels after incorporating topological constraints. The structural salience score, The original corner response values ​​(such as Harris or Shi-Tomasi responses) are obtained based on traditional image gradient calculations. Indicates the first The coordinate spatial location of each candidate physical node (intersection of truss segments). This is the structure enhancement coefficient, used to control the weighted gain of physical intersections on their neighboring pixels. The confidence region radius characterizes the influence range of the intersection point. The set of candidate physical nodes. The coordinates of the current pixel. This represents the square of the Euclidean distance. In this process, consider a specific UHVDC transmission tower scenario where the main tower members and diagonal members form multiple "X" or "K" shaped nodes in space. Due to years of wind and rain erosion, a high-contrast rust patch exists in the center of a certain section of the steel surface, which will produce extremely high values ​​in traditional gradient calculations. The value is easily mistakenly selected as a feature point. However, through the above formula processing, because the rust point is far from the actual line segment intersection (physical node), its exponential term weight approaches zero, resulting in the final value. Scoring was suppressed. Conversely, the points located at the intersection of steel welds, even though the texture was slightly weaker due to lighting, were extremely close to the physical intersection. ,exist Under the gain of the coefficient The score was significantly improved, thus ensuring that the feature points ultimately locked by the algorithm accurately corresponded to the physical connection points of the towers, rather than random noise on the surface.

[0032] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the feature locking subunit 1233 is used to perform structure-aware non-maximum suppression and feature locking on the topology-weighted response map to obtain a rigid feature point set. It should be noted that, given that at this stage, the non-maximum suppression strategy is used to filter out points with extremely large response values ​​in the local neighborhood, and since the input source has been changed to a response map containing topological weights, the finally locked feature points must be high-quality anchor points with both texture clarity and structural stability. A simple response map may contain locally dense high-response regions, leading to an overly concentrated or redundant spatial distribution of feature points, which not only increases the burden of subsequent calculations but may also reduce the numerical stability of pose calculation due to the strong correlation between feature points. Based on this, the technical solution of this application further performs structure-aware feature locking based on the generated topology-weighted response map to ensure that each feature point tracked by the system is firmly anchored to the physical node of the tower. Through the above processing, disordered pixel features can be transformed into ordered structured anchor points, thereby providing extremely high-precision geometric constraints in subsequent processing. This effectively solves the feature drift problem under weak texture, dynamic blur and surface stain interference, and provides reliable underlying data support for centimeter-level accurate positioning of potential hazards in power transmission channels.

[0033] More specifically, in a concrete example of this application, this step receives a topologically weighted response map obtained through prior processing as input. First, a local neighborhood window adapted to the scale of the tower structure and an adaptive filtering threshold are set. Then, the neighborhood window is slid across the response map, and non-maximum suppression is performed on each pixel within the window. A pixel is considered a valid feature point only if its response value simultaneously meets the conditions of being higher than the set filtering threshold and being the maximum value within its neighborhood. The mathematical expression of this filtering process is as follows: in, This is the final set of rigid feature points output. Represents the coordinates of pixels in an image. This represents the response value of the pixel in the topologically weighted response map. An adaptive screening threshold incorporating structural weights is used to filter out low-confidence points with weak responses. This represents the logical AND operation. Represents pixels The spatial neighborhood range is used to define the size of the region for comparing local extrema. In a real-world inspection scenario, suppose a drone is hovering beside a UHV transmission tower. Due to strong winds, the drone's fuselage experiences high-frequency micro-vibrations. In the captured image, a key node of the tower—the bolt connection between the main steel beam and the diagonal brace—forms a bright area due to light reflection. Due to motion blur, this bright area appears as a diffuse patch on the response map. Direct extraction might yield multiple clustered feature points. Through the aforementioned structure-aware non-maximum suppression, the algorithm first eliminates pixels with weak edge responses in the patch. Then, within the high-response region at the center of the patch, it precisely locates the peak point with a unique response value after topological weighting. This point physically corresponds precisely to the geometric center of the bolt. Even if the next frame image experiences a slight displacement due to vibration, the topological saliency of this physical center remains at its peak, thus ensuring stable tracking of feature points over time and eliminating random jumps in feature points within the patch caused by blurring.

[0034] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the robust pre-integration module 130 is used to robustly integrate the inertial navigation data in the synchronized multi-source data frames based on electromagnetic interference state markings to obtain an inertial navigation pre-integration factor containing relative pose increments. It should be noted that, given the extremely harsh electromagnetic environment surrounding ultra-high voltage power transmission channels, the high-frequency power currents of several thousand amperes generate strong alternating magnetic fields, causing severe distortion or even saturation of the magnetometer readings on the UAV. If the magnetometer observations are continuously trusted during the state recursion process, erroneous heading angle correction information will directly contaminate the zero-bias estimation of the inertial navigation system, thereby causing heading divergence in pose calculation. Based on this, the technical solution of this application further adjusts the heading angle variance term in the noise covariance matrix according to the electromagnetic interference state marker, and constructs an error state equation in conjunction with the relative motion increment to output the inertial navigation pre-integration factor. That is, when a strong interference signal is detected, the algorithm logic actively amplifies the variance parameter of the magnetometer observation noise, reduces or even cuts off the weight dependence of the magnetic field heading observation in the optimization algorithm, and instead uses the zero-bias estimate optimized at the previous moment to perform zero-bias compensation and measurement value sequence correction on the inertial navigation measurement sequence in the synchronized multi-source data frame, and performs pre-integration state recursion on the corrected data in the manifold space. Through the above processing, adaptive isolation of environmental magnetic field interference can be effectively achieved, ensuring that even under extreme conditions where the external compass reference fails, a high-confidence relative pose increment can still be output based on pure inertial data, providing robust dynamic constraints for subsequent global optimization.

[0035] Figure 5 is a block diagram of the robust pre-integration module in a power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. As shown in Figure 5, the robust pre-integration module 130 includes: a zero-bias compensation unit 131, used to use the zero-bias estimate optimized at the previous moment to perform zero-bias compensation and measurement value sequence correction on the inertial navigation measurement sequence in the synchronized multi-source data frame to obtain the corrected inertial navigation measurement sequence; a state recursion unit 132, used to perform pre-integration state recursion in manifold space on the corrected inertial navigation measurement sequence to obtain the relative motion increment including displacement, velocity and rotation changes; and a factor construction unit 133, used to adjust the heading angle variance term in the noise covariance matrix according to the electromagnetic interference state mark, and construct the error state equation in combination with the relative motion increment to output the inertial navigation pre-integration factor.

[0036] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the zero-bias compensation unit 131 is used to perform zero-bias compensation and measurement value sequence correction on the inertial navigation measurement sequence in the synchronized multi-source data frame using the zero-bias estimate optimized at the previous moment to obtain the corrected inertial navigation measurement sequence. It should be noted that, due to the influence of manufacturing processes and long-term operation heat generation, the original output data of the airborne MEMS inertial measurement unit inevitably contains a zero-bias error that drifts slowly over time in its accelerometer and gyroscope. If the original measurement value containing this system error is directly integrated, the error will accumulate quadratically over time, causing the pose calculation result to diverge rapidly. Based on this, the technical solution of this application further utilizes the zero-bias estimate optimized at the previous moment to perform zero-bias compensation and measurement value sequence correction on the inertial navigation measurement sequence in the synchronized multi-source data frame to obtain the corrected inertial navigation measurement sequence. This removes deterministic sensor bias components during the data preprocessing stage and restores the true motion physical quantities of the carrier. The above processing can effectively block the propagation path of error sources in the state recursion process, ensuring the accuracy and reliability of subsequent pre-integral factor construction.

[0037] More specifically, in a concrete example of this application, the execution process encompasses a complete closed loop from historical state extraction to serialized numerical correction. First, all raw inertial navigation measurement data within the time window between the current and previous frames is parsed from the synchronized multi-source data frames, including high-frequency sampled three-axis acceleration and angular velocity reading sequences. Simultaneously, the converged accelerometer and gyroscope zero-bias vectors are indexed from the back-end optimization results of the previous time step as the current error correction benchmark. Subsequently, point-by-point compensation operations are performed on the extracted measurement sequences. For each timestamp in the sequence, the original angular velocity reading is subtracted from the gyroscope zero-bias vector to obtain the corrected angular rate, and the original acceleration reading is subtracted from the accelerometer zero-bias vector to obtain the corrected specific force value. Finally, all sampling points that have undergone error stripping are repackaged into a corrected inertial navigation measurement sequence and passed as a clean dynamic input to the manifold space pre-integration module.

[0038] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the state recursion unit 132 is used to perform pre-integration state recursion in manifold space on the corrected inertial navigation measurement sequence to obtain the relative motion increment including displacement, velocity, and rotational changes. It should be noted that since traditional inertial navigation calculations typically rely on the initial state in the global coordinate system, whenever the back-end nonlinear optimization iteration causes a small update to the initial state estimate, it is necessary to re-integrate hundreds of high-frequency inertial measurement values ​​between two frames. This consumes enormous computational resources, and directly processing three-dimensional rotational motion in Euclidean space faces the problems of singularity and loss of nonlinear constraints. Based on this, the technical solution of this application further performs pre-integration state recursion in manifold space on the corrected inertial navigation measurement sequence to obtain the relative motion increment including displacement, velocity, and rotational changes. This mathematically decouples the relative motion between frames from the absolute state of the carrier in the world coordinate system and rigorously handles rotational transformations in the Lie group manifold space to avoid gimbal lock-up. Through the above processing, the high-frequency and dense raw inertial navigation data can be compressed into low-frequency relative constraint factors that are independent of the initial state. This eliminates the need to repeatedly perform integration operations in subsequent optimization processes, thereby effectively improving the real-time performance and numerical stability of edge computing devices.

[0039] More specifically, in a concrete example of this application, the recursive process strictly follows the principles of Lie groups and Lie algebra mappings in differential geometry. First, the calibrated inertial navigation measurement sequence within the time interval between two adjacent visual keyframes is read. The discrete sampling points are processed using the midpoint integral numerical method, i.e., the average of the measurements from two adjacent moments is used as the motion input for that small time interval. For the recursion of the rotation component, an exponential mapping is used to map the calibrated angular velocity vector from the tangent space (i.e., the Lie algebra space) to the space of a special orthogonal group manifold. The relative rotation matrix of the current moment relative to the previous keyframe moment is calculated cumulatively through matrix multiplication. For the recursion of the velocity and displacement components, the calibrated acceleration measurement value is projected onto the volume coordinate system of the previous keyframe using the instantaneously calculated relative rotation matrix. An integral is performed using the time step to obtain the relative velocity increment, and then the velocity increment is integrated again to obtain the relative displacement increment. The entire calculation process is completed entirely within a local reference frame, ultimately outputting a set of motion increment vectors containing relative displacement, relative velocity, and relative rotation. This vector is only related to the sensor measurements and is independent of the absolute position or attitude of the UAV in the Earth coordinate system.

[0040] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the factor construction unit 133 is used to adjust the heading angle variance term in the noise covariance matrix according to the electromagnetic interference state marker, and to construct an error state equation in conjunction with the relative motion increment to output the inertial navigation pre-integration factor. It should be noted that due to the high-intensity power frequency electromagnetic field interference around the power transmission channel, the heading data collected by the magnetometer may exhibit random jumps or unreliable phenomena in specific areas. If a constant measurement noise assumption is maintained in the state estimation, erroneous magnetic field observations will deflect the system's heading angle estimation, leading to pose divergence. Based on this, the technical solution of this application further adjusts the heading angle variance term in the noise covariance matrix according to the electromagnetic interference state marker, and constructs an error state equation in conjunction with the relative motion increment to output the inertial navigation pre-integration factor, thereby establishing an adaptive immune mechanism against environmental magnetic field interference and blocking the propagation of abnormal observation information to the state estimator at the data fusion level. Through the above processing, dynamic control of sensor confidence can be achieved, ensuring that the algorithm logic automatically reduces its dependence on the magnetometer and switches to a pure inertial recursive mode in a strong magnetic field environment, thereby effectively ensuring the continuity and accuracy of heading angle estimation.

[0041] More specifically, in a concrete example of this application, the implementation process involves the dynamic weighting of the noise model and the recursive calculation of the probability distribution. First, the electromagnetic interference status flag attached to the synchronization data frame is read, which characterizes the reliability of the magnetometer observation data at the current moment. When the flag indicates a strong interference state, variance inflation is performed during the construction of the measurement noise covariance matrix. This involves increasing the variance value corresponding to the heading angle dimension through numerical multiplication, setting it to a value significantly higher than under normal conditions, mathematically representing extremely high uncertainty in this measurement dimension. Subsequently, covariance propagation of the error state is performed. Based on micro-perturbation theory, a linearized state equation for position, velocity, and attitude errors is constructed. The state transition Jacobian matrix is ​​calculated using the relative motion increment obtained in the previous steps. Combined with the dynamically adjusted noise covariance matrix, the cumulative covariance matrix of the pre-integral is updated using a discrete-time recursive formula. Finally, the mean vector containing relative displacement, velocity, and rotation increments is encapsulated together with the anti-interference covariance matrix to generate an inertial navigation pre-integration factor containing complete uncertainty information, which is then passed to the back-end optimization stage as a high-confidence constraint edge.

[0042] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the geometric constraint construction module 140 is used to perform spatial registration and triangulation processing on the rigid feature point set and the pre-set standard tower model to obtain the tower geometric constraints and visual anchor point factors. It should be noted that, given that in the inspection of long-distance power transmission channels, relying solely on visual odometry for relative pose recursion not only suffers from monocular scale ambiguity but also generates irreversible cumulative drift errors as the flight distance increases, and considering that linear inspection paths in the field typically lack repeating textures, it is difficult to eliminate global errors using traditional loop closure detection mechanisms. Based on this, the technical solution of this application further performs feature time-series tracking and triangulation reprojection on the rigid feature point set to obtain visual anchor point factors and local sparse point clouds. The local sparse point clouds are then iteratively solved using a rigid transformation with a pre-set standard tower model to obtain the registration transformation matrix and registration score. When the registration score meets the confidence threshold, tower geometric constraints are constructed based on the registration transformation matrix. This utilizes the standardized geometric topology of the transmission tower as an absolute spatial reference to physically align the UAV's current local observation coordinate system with the tower model in the world coordinate system, constructing a priori geometric factors with strong constraints. Through the above processing, the unbounded cumulative drift error of the visual odometry can be forcibly converged to the registration residual range of the tower geometric model. This is equivalent to continuously performing absolute position calibration along the inspection path using standardized "digital landmarks," effectively ensuring that the system maintains centimeter-level high-precision positioning capabilities during long-endurance missions.

[0043] Figure 6 is a block diagram of the geometric constraint construction module in a power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to an embodiment of this application. As shown in Figure 6, the geometric constraint construction module 140 includes: a visual tracking unit 141, used to perform feature time-series tracking and triangulation reprojection on a rigid feature point set to obtain visual anchor point factors and local sparse point clouds; an iterative registration unit 142, used to perform rigid transformation iterative solution on the local sparse point cloud and the standard tower model to obtain the registration transformation matrix and registration score; and a geometric constraint construction unit 143, used to construct tower geometric constraints based on the registration transformation matrix when the registration score meets the confidence threshold.

[0044] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the visual tracking unit 141 is used to perform feature temporal tracking and triangulation reprojection on a set of rigid feature points to obtain visual anchor point factors and local sparse point clouds. It should be noted that since monocular vision sensors can only acquire pixel coordinate information in a two-dimensional plane and lack direct depth perception capabilities, and feature points in a single frame image cannot independently construct spatial geometric constraints, without restoring the three-dimensional spatial position of the feature points, it is impossible to utilize the physical structural characteristics of the power transmission tower for absolute pose calibration. Based on this, the technical solution of this application further performs feature temporal tracking and triangulation reprojection on the set of rigid feature points to obtain visual anchor point factors and local sparse point clouds, thereby utilizing the parallax generated by UAV flight to restore the depth information of the rigid feature points and construct visual observation constraints based on reprojection errors. Through the above processing, two-dimensional image features can be transformed into local sparse point clouds with three-dimensional spatial coordinates, providing a geometric data basis for subsequent registration with standard tower models. At the same time, by constraining the camera pose estimation through visual anchor point factors, the solution accuracy of visual odometry can be effectively improved.

[0045] More specifically, in a concrete example of this application, the implementation process involves multiple steps, including optical flow tracing, disparity filtering, and depth calculation. First, temporal correlation of feature points is performed. Optical flow is used to track the pixel positions of a rigid set of feature points in a continuous sequence of image frames, establishing the correspondence between feature points at different times. During the tracking process, the pixel displacement of feature points between the current frame and the reference frame is calculated, and feature point pairs whose disparity meets a preset threshold are selected to ensure sufficient geometric observation baseline for triangulation calculation. Then, triangulation of depth information is performed. Using the camera intrinsic parameter matrix and the currently estimated relative pose, line-of-sight cross-calculation is performed on feature point pairs that meet the disparity requirements to solve for the inverse depth value of the feature points in the camera coordinate system, thereby recovering the three-dimensional spatial coordinates of key parts of the transmission tower and forming a local sparse point cloud. Finally, a visual anchor factor is constructed. The observed pixel coordinates of the feature points are differiated from the reprojected pixel coordinates calculated based on the inverse depth and camera pose to construct a reprojection error term. This error term and its corresponding information matrix are encapsulated as a visual anchor factor, which is added to the factor graph as a constraint edge for the optimization problem.

[0046] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the iterative registration unit 142 is used to perform a rigid transformation iterative solution between the local sparse point cloud and the standard tower model to obtain the registration transformation matrix and registration score. It should be noted that during long-term UAV inspections, the pose calculated solely by visual inertial odometry inevitably accumulates errors, causing the recovered local point cloud to rotate and translate in the world coordinate system. Furthermore, local observation data often only includes a portion of the transmission tower's side structure, lacking globally consistent absolute spatial constraints. Therefore, the technical solution of this application further performs a rigid transformation iterative solution between the local sparse point cloud and a pre-set standard tower model to obtain the registration transformation matrix and registration score. This utilizes the absolute geometric dimensions of the standard tower model as a calibration benchmark, and mathematical optimization methods are used to find the optimal spatial transformation relationship between the local observation data and the prior model. Through the above processing, the deviation of the current visual positioning result from the real physical structure can be accurately quantified, and a high-precision correction matrix that can force the cumulative drift to be reset can be obtained, thereby effectively realizing the geometric alignment from relative positioning to absolute positioning and ensuring that the pose estimation does not diverge over time.

[0047] More specifically, in a concrete example of this application, the registration process employs an optimization strategy based on the iterative nearest neighbor algorithm, covering the entire process from point set construction to robust optimization. First, the locally sparse point cloud generated through visual triangulation is defined as the source point set, and the discretized surface points of a pre-defined standard tower model retrieved from the database are defined as the target point set. A kd-tree data structure is used to quickly search for the nearest neighbor corresponding point of each rigid feature point in the source point set within the target point set. Subsequently, a nonlinear objective function containing rotation matrices and translation vectors is constructed. This objective function aims to minimize the sum of squared Euclidean distances between all matched point pairs. During the iteration process, a Huber robust kernel function is introduced to weight the residuals, automatically reducing the influence of mismatched points or outlier noise points on the optimization results. By iteratively calculating multiple times until the descent gradient of the objective function is less than a preset threshold or the maximum number of iterations is reached, the optimal rigid transformation matrix that maximizes the overlap between the two sets of point clouds is calculated. At the same time, the root mean square error of the final matched point pair is calculated as the registration score. This score directly reflects the consistency between the current visual observation data and the standard model, providing a quantitative confidence basis for subsequent judgment on whether to add the geometric constraint to the global factor map.

[0048] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the geometric constraint construction unit 143 is used to construct tower geometric constraints based on the registration transformation matrix when the registration score meets the confidence threshold. It should be noted that in complex field environments, due to viewpoint occlusion or missing local features, the point cloud registration process may converge to a local minimum or produce a large fitting residual. If all registration results are indiscriminately introduced as strong constraints into the optimization equation, erroneous geometric information will directly contaminate the global pose calculation. Therefore, the technical solution of this application further sets strict confidence gating logic, only using the calculated registration transformation matrix to construct mathematical tower geometric constraints when the registration score meets the preset confidence threshold, thereby adaptively filtering the reliability of the absolute position correction information. Through the above processing, the propagation of low-quality or erroneous registration results to the state estimator can be effectively blocked, ensuring that only high-precision prior information with physical authenticity is introduced, thus guaranteeing that the convergence direction of the factor graph optimization process always tends towards the true value.

[0049] More specifically, in a concrete example of this application, the process first performs a numerical determination of registration quality. The root mean square error score output by the iterative nearest-point algorithm is read and compared numerically with a pre-set empirical confidence threshold, which represents the maximum allowable geometric deviation tolerance. The current registration result is deemed valid only if the score is less than this threshold. Then, using the registration transformation matrix combined with the absolute coordinates of the standard model, the prior pose matrix of the UAV camera center in the world coordinate system is calculated. Subsequently, a univariate prior factor is constructed. This factor defines the error residual term between the predicted pose in the current state estimator and the aforementioned prior pose matrix in the Lie algebra space, and assigns this residual term an information matrix with minimal covariance to give it a higher optimization weight. Finally, this univariate factor containing strong geometric constraints is encapsulated and loaded onto the pose node corresponding to the factor graph, forcing the backend optimizer to prioritize pulling the pose trajectory closer to this absolute geometric reference when minimizing the total energy function.

[0050] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the pose optimization and calculation module 150 is used to perform factor graph optimization and robust pose calculation on the inertial navigation pre-integration factor, visual anchor point factor, and tower geometric constraints to obtain the optimized six-degree-of-freedom pose. It should be noted that, given that single-sensor observation data inevitably contains noise and system errors, high-frequency mechanical vibrations introduce random walk noise into the inertial recursion, long-distance flight causes unbounded cumulative drift in the visual odometry, and the geometric constraints obtained from local registration exhibit sparsity and discontinuity, relying solely on loose coupling or linear filtering methods is insufficient to guarantee globally consistent positioning accuracy in complex ultra-high voltage electromagnetic interference environments. Based on this, the technical solution of this application further constructs a sliding window containing state nodes, and loads the inertial navigation pre-integration factor, visual anchor point factor, and tower geometric constraints between the corresponding state nodes to obtain a factor graph and an initial state estimation vector. Then, based on the factor graph, the initial state estimation vector is iteratively updated to obtain a window-optimized state vector, and the current frame state is extracted from the window-optimized state vector. At the same time, the removed old state nodes are marginalized to obtain the optimized six-degree-of-freedom pose. This constructs a unified probabilistic graphical model that can simultaneously accommodate relative dynamic constraints, visual geometric constraints, and absolute spatial priors, transforming the multi-source fusion problem into a nonlinear least squares optimal estimation problem. Through the above processing, the cumulative drift of the visual-inertial system can be forcibly pulled back using the absolute geometric constraints of the tower, and the sliding window marginalization mechanism maintains a constant computational load while preserving historical prior information. This effectively outputs centimeter-level high-precision poses with both high-frequency dynamic response and long-term global consistency, providing a solid spatial benchmark for the accurate back-calculation of potential hazards.

[0051] More specifically, in a specific example of this application, the pose optimization solution module 150 includes: a factor graph loading unit, used to construct a sliding window containing state nodes, and load the inertial navigation pre-integration factor, visual anchor point factor, and tower geometric constraints between the corresponding state nodes to obtain a factor graph and an initial state estimation vector; an iterative update unit, used to iteratively update the initial state estimation vector based on the factor graph to obtain a window-optimized state vector; and an edge-processing unit, used to extract the current frame state from the window-optimized state vector and perform edge-processing on the removed old state nodes to obtain the optimized six-degree-of-freedom pose.

[0052] Accordingly, the factor graph loading unit is used to construct a sliding window containing state nodes and load the inertial navigation pre-integration factor, visual anchor factor, and tower geometric constraints between the corresponding state nodes to obtain the factor graph and the initial state estimation vector. It should be noted that due to the limited computing resources of the UAV's onboard edge computing platform, performing global nonlinear optimization on the entire trajectory state would lead to computational lag, and a simple linear filtering method is difficult to effectively fuse the tower's absolute geometric constraints with time delay characteristics and the highly nonlinear visual observation information. Based on this, the technical solution of this application further constructs a sliding window containing state nodes and loads the inertial navigation pre-integration factor, visual anchor factor, and tower geometric constraints between the corresponding state nodes to obtain the factor graph and the initial state estimation vector. This models the multi-source sensor fusion problem as a fixed-dimensional probabilistic graphical model, achieving joint optimization of multiple constraints while maintaining a constant computational scale. Through the above processing, the real-time performance and accuracy of pose calculation can be effectively balanced, ensuring that the real-time acquisition of front-end data is not blocked while introducing strong geometric constraints to correct drift.

[0053] More specifically, in a concrete example of this application, the factor graph construction process follows an implementation logic from topological definition to constraint equation loading. First, a fixed-length time sliding window is established, maintaining a certain number of keyframe state nodes. Each state node contains the UAV's position, velocity, attitude quaternions, and zero-bias variables from the accelerometer and gyroscope at the corresponding time. Then, factor edge loading is performed. For two adjacent state nodes within the window, a binary constraint edge is constructed using the inertial navigation pre-integration factor calculated in the previous steps, constraining the relative motion changes between them. For observed rigid feature points within the window, a multi-constraint edge connecting the inverse depth of the feature point and the camera pose is constructed using the visual anchor point factor. In particular, when effective tower geometric constraints exist, they are converted into unary prior factors and directly applied to the corresponding pose state node, introducing absolute spatial position correction. Finally, initial value prediction is performed. The mechanical arrangement results of the inertial navigation data are used as initial value guesses for all state variables to be optimized, combined to form an initial state estimation vector, providing a convergence starting point for subsequent nonlinear iterative optimization.

[0054] Accordingly, the iterative update unit is used to iteratively update the initial state estimation vector based on the factor graph to obtain the window-optimized state vector. It should be noted that since the initial state estimation vector mainly relies on the mechanical arrangement and prediction of inertial navigation data, it is limited by the inherent noise and zero-bias random walk characteristics of the inertial sensor. It does not fully integrate the constraints of visual observation and tower geometric priors, resulting in dynamic and geometric inconsistencies among the state variables. Direct output would contain a large cumulative error. Therefore, the technical solution of this application further iteratively updates the initial state estimation vector based on the factor graph to obtain the window-optimized state vector, thereby constructing and solving a nonlinear least squares problem. A numerical optimization method is used to find the optimal state point that minimizes the sum of all factor residuals. Through the above processing, the measurement errors of multiple sensors can be redistributed and suppressed within a probabilistic framework, thereby effectively obtaining the optimal window state solution that satisfies global geometric consistency and has higher accuracy than single-sensor observations.

[0055] More specifically, in a concrete example of this application, the implementation process employs the Levenberg-Marquardt algorithm to perform nonlinear optimization on the constructed factor graph model. First, a total residual objective function is constructed, which consists of the sum of squared Mahalanobis distances weighted by the covariance matrices of the dynamic residuals of all inertial navigation pre-integration factors, the reprojection residuals of visual anchor factors, and the prior residuals of tower geometric constraint factors within the sliding window. Then, linearization and iterative solving are performed. In each iteration step, the residual vectors corresponding to the current state estimate and the Jacobian matrix with respect to the state variables are calculated, thereby constructing a normal equation system containing the Hessian matrix and gradient vectors. Next, the incremental equations are solved, and the state increment vector is calculated using linear algebra methods such as sparse Cholesky decomposition. This vector is then superimposed on the current initial state estimate vector to update the state points. Finally, convergence is determined, and the linearization and update process is repeated until the magnitude of the state increment is less than a preset small threshold or the decrease in the objective function tends to stabilize. At this point, the output state vector is the window-optimized state vector incorporating multi-source constraints.

[0056] Accordingly, the edge-processing unit is used to extract the current frame state from the window-optimized state vector and perform edge-processing on the removed old state nodes to obtain the optimized six-DOF pose. It should be noted that due to the limited storage space and computing power of the UAV's onboard computing platform, the number of state nodes in the factor graph will increase linearly with the flight time. If all historical states are retained without restriction, the computational load will quickly exceed the hardware's capacity limit, leading to system crash. Conversely, directly discarding old state nodes means losing the constraint information accumulated from historical observations, thus disrupting the global consistency of the trajectory. Based on this, the technical solution of this application further extracts the current frame state from the window-optimized state vector and performs edge-processing on the removed old state nodes to obtain the optimized six-DOF pose. This maintains a constant sliding window dimension while converting the constraint information carried by the historical states that slide out of the window into a priori information matrix and retaining it in the system. Through the above processing, a dynamic balance between computational consumption and positioning accuracy can be effectively achieved, ensuring that the system maintains real-time and high-precision pose output capability during long-endurance inspection missions.

[0057] More specifically, in a concrete example of this application, the execution flow encompasses two key stages: result output and information compression. First, the optimal state is extracted and output. The state variables corresponding to the latest image frame are indexed from the converged window optimization state vector. The rotation matrix and translation vector are parsed and combined into high-confidence six-DOF pose data, which is then transmitted to the hazard location module for 3D inverse calculation as the final navigation result. Next, the old state is marginalized. When a new keyframe is added to the sliding window, the oldest state node with the earliest timestamp within the window is identified as the object to be removed. A locally linearized system containing this node and its associated factors is constructed. The Hessian matrix of this local system is decomposed using the Schur complement algebra method, and the information contribution of the node to be removed to the remaining associated nodes is projected into a priori information matrix and a priori residual vector. Finally, this prior term is constructed as a marginalization probability factor and connected to the oldest frame state node retained within the sliding window. This completes the mathematical compression and transmission of historical trajectory constraint information, providing boundary conditions for the next round of window optimization.

[0058] In the aforementioned power transmission channel hazard inspection system 100 based on computer vision and UAV collaboration, the hazard location output module 160 is used to locate the hazard target pixel from the synchronized multi-source data frames, and to calculate the true three-dimensional coordinates of the hazard point by combining the optimized six-degree-of-freedom pose. It should be noted that since the images acquired by the monocular vision sensor only contain semantic information of the two-dimensional pixel plane and lack direct spatial depth measurement, although the simple semantic segmentation result can identify the presence of vegetation encroachment hazards in the image, it cannot know the actual physical distance and absolute geographical location of the hazard point relative to the power transmission line. Moreover, if the unoptimized original sensor pose is directly used for inverse calculation, inertial navigation drift and extrinsic parameter errors will cause the calculated three-dimensional coordinates to deviate from the true position, which cannot meet the power operation and maintenance requirements for accurate hazard point location. Based on this, the technical solution of this application further locates the target pixels of potential hazards from synchronized multi-source data frames using a semantically weighted mask, and calculates the true three-dimensional coordinates of the hazard points by combining the optimized six-degree-of-freedom pose to obtain the three-dimensional coordinates of the hazard. This allows the rigorously optimized camera pose to serve as a high-precision spatial reference. The depth information of the two-dimensional hazard pixels is recovered using multi-view geometry principles, and this depth information is projected onto a unified world coordinate system. Through the above processing, a precise mapping from two-dimensional image semantics to three-dimensional spatial coordinates can be achieved, ensuring that the output hazard coordinates have centimeter-level absolute positioning accuracy, providing reliable navigation data for subsequent automatic defect removal by UAVs or manual intervention.

[0059] More specifically, in a specific example of this application, the hazard location output module 160 includes: a target location unit, used to extract points of interest and perform epipolar matching on the image data in the synchronized multi-source data frames based on a semantically weighted mask to obtain the hazard target pixel; a depth determination unit, used to determine the estimated depth of the target based on camera intrinsic parameters and the optimized six-degree-of-freedom pose; and a reverse projection unit, used to perform reverse projection of the absolute position of the hazard target pixel in the world coordinate system using the estimated depth of the target and the inverse matrix of the camera intrinsic parameter matrix to obtain the three-dimensional coordinates of the hazard.

[0060] Furthermore, the process first extracts points of interest and performs epipolar matching in the hazard area, traverses the semantically weighted mask, extracts connected regions marked as dynamic interference or vegetation, and calculates the centroid of the connected region or the boundary point closest to the power line as the target pixel of the hazard. The fundamental matrix is ​​calculated using the optimized six-degree-of-freedom pose. In the previous keyframe Chinese search based on epipolar constraints matching points with the same name If a match is successful, the observation point pair between the two frames is output, calculated as follows: in, For the current frame Homogeneous coordinates of pixels in the image. Reference frame The homogeneous coordinates of the matching pixels in the image. To describe the frame With frames The fundamental matrix of the epipolar geometric relationship between them is then determined. Subsequently, depth triangulation based on the optimal pose is performed. The projection matrices of the current frame and the reference frame are obtained using the optimized six-DOF pose. A system of linear triangulation equations is constructed and solved using the least-squares solution to obtain the scene depth value in the camera coordinate system. The calculation formula is as follows: in, The calculated straight-line distance between the potential hazard point and the current camera optical center. The baseline length between the optical centers of the current frame and the reference frame. The angle between the observed ray and the baseline is determined for the feature point. Finally, an inverse projection of the absolute position in the world coordinate system is performed, utilizing the camera intrinsic parameter matrix. The inverse matrix backprojects the target pixels of the potential hazard onto the normalized camera plane, combines the target depth estimation to expand it into a 3D vector in the camera coordinate system, and utilizes the rotation matrix in the optimized six-degree-of-freedom pose. Translation vector Transform it to the world coordinate system; the calculation formula is as follows: in, The absolute three-dimensional coordinates of the potential hazard point in the world coordinate system. This is the optimized rotation matrix from the camera coordinate system to the world coordinate system. This is the optimized translation vector from the camera coordinate system to the world coordinate system. The x and y coordinates of the potential hazard on the image plane. This is the inverse of the camera intrinsic parameter matrix. Through this process, the system ultimately outputs high-precision 3D coordinates of the hidden dangers after geometric correction and depth recovery.

[0061] In summary, the power transmission channel hazard inspection system based on computer vision and UAV collaboration, according to embodiments of this application, is explained. First, it performs clock synchronization and distortion correction on multi-source heterogeneous sensor data, and monitors magnetic field anomalies in real time to dynamically adjust navigation weights, blocking magnetic interference from misleading the course at its source. Based on this, it utilizes semantic masks to shield dynamic background noise such as vegetation, focusing on extracting the rigid structural features of the transmission towers, and uses truss topology analysis to lock visual features onto physical nodes. By spatially registering these high-confidence structural features with a standard tower model, absolute geometric constraints are introduced to correct the cumulative drift of the visual odometry. Finally, within a factor graph optimization framework, multi-source information is fused to calculate a highly robust six-degree-of-freedom pose, and based on this, a three-dimensional inverse projection is performed on the hazard target, thereby achieving centimeter-level precise positioning of power transmission channel hazards.

Claims

1. A power transmission channel hidden danger inspection system based on computer vision and UAV collaboration, characterized in that, include: The data acquisition and synchronization module is used to perform spatiotemporal alignment and magnetic field anomaly detection on the raw visual stream, raw inertial navigation readings, raw magnetometer readings and raw satellite observations collected by the UAV in order to obtain synchronized multi-source data frames and electromagnetic interference status markers. The semantic feature extraction module is used to perform semantic perception and dynamic feature mask construction on the image data in the synchronized multi-source data frames to obtain a set of rigid feature points containing only the tower structure. The robust pre-integration module performs robust integration on the inertial navigation data in the synchronized multi-source data frames based on electromagnetic interference state markers to obtain the inertial navigation pre-integration factor containing the relative pose increment; the geometric constraint construction module performs spatial registration and triangulation on the rigid feature point set and the pre-set standard tower model to obtain the tower geometric constraints and visual anchor point factors; the pose optimization solution module performs factor graph optimization and robust pose solution on the inertial navigation pre-integration factor, visual anchor point factor and tower geometric constraints to obtain the optimized six-degree-of-freedom pose; the hazard location output module locates the hazard target pixel from the synchronized multi-source data frames and, combined with the optimized six-degree-of-freedom pose, calculates the true three-dimensional coordinates of the hazard point to obtain the hazard's three-dimensional coordinates.

2. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 1, characterized in that, The data acquisition and synchronization module includes: an image correction unit, used to perform distortion correction processing on the original visual stream based on the camera intrinsic parameter matrix and distortion model parameters, and extract the exposure timestamp to obtain distorted image frames and image timestamps; a time alignment unit, used to perform linear interpolation and time alignment on the original inertial navigation readings, original magnetometer readings, and original satellite observations with the image timestamps as the main clock reference to obtain a synchronized measurement dataset containing distorted image frames; and an anomaly detection unit, used to perform electromagnetic environment anomaly detection and status marking based on the magnetometer data in the synchronized measurement dataset to output synchronized multi-source data frames and electromagnetic interference status markings.

3. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 1, characterized in that, The semantic feature extraction module includes: a semantic reasoning unit, used to perform image tensor preprocessing and pixel-level semantic reasoning on image data in synchronized multi-source data frames to obtain a semantic label map; a mask generation unit, used to map the pixel categories in the semantic label map to differentiated numerical weights according to a preset category stability mapping table to obtain a semantically weighted mask; and a feature filtering unit, used to perform feature extraction and filtering on image data in synchronized multi-source data frames under rigid region constraints based on the semantically weighted mask to obtain a rigid feature point set.

4. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 1, characterized in that, The robust pre-integration module includes: a zero-bias compensation unit, used to perform zero-bias compensation and measurement sequence correction on the inertial navigation measurement sequence in the synchronized multi-source data frame using the zero-bias estimate optimized at the previous moment to obtain the corrected inertial navigation measurement sequence; a state recursion unit, used to perform pre-integration state recursion in manifold space on the corrected inertial navigation measurement sequence to obtain the relative motion increment including displacement, velocity and rotation changes; and a factor construction unit, used to adjust the heading angle variance term in the noise covariance matrix according to the electromagnetic interference state mark, and combine it with the relative motion increment to construct the error state equation to output the inertial navigation pre-integration factor.

5. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 1, characterized in that, The geometric constraint construction module includes: a visual tracking unit, used to perform feature temporal tracking and triangulation reprojection on a rigid feature point set to obtain visual anchor factors and local sparse point clouds; an iterative registration unit, used to perform rigid transformation iterative solution on the local sparse point cloud and the standard tower model to obtain the registration transformation matrix and registration score; and a geometric constraint construction unit, used to construct tower geometric constraints based on the registration transformation matrix when the registration score meets the confidence threshold.

6. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 1, characterized in that, The pose optimization solution module includes: a factor graph loading unit, used to construct a sliding window containing state nodes, and load the inertial navigation pre-integration factor, visual anchor point factor and tower geometric constraints between the corresponding state nodes to obtain a factor graph and an initial state estimation vector; an iterative update unit, used to iteratively update the initial state estimation vector based on the factor graph to obtain the window-optimized state vector; and an edge-processing unit, used to extract the current frame state from the window-optimized state vector and perform edge-processing on the removed old state nodes to obtain the optimized six-degree-of-freedom pose.

7. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 1, characterized in that, The hazard location output module includes: a target localization unit, used to extract points of interest and perform epipolar matching on the image data in the synchronized multi-source data frames based on a semantically weighted mask to obtain the target pixel of the hazard; a depth determination unit, used to determine the estimated depth of the target based on camera intrinsic parameters and optimized six-degree-of-freedom pose; and a reverse projection unit, used to perform reverse projection of the absolute position of the target pixel of the hazard in the world coordinate system using the estimated depth of the target and the inverse matrix of the camera intrinsic parameter matrix to obtain the three-dimensional coordinates of the hazard.

8. The power transmission channel hidden danger inspection system based on computer vision and UAV collaboration according to claim 3, characterized in that, The feature selection unit includes: a line segment detection subunit, used to perform line segment detection and noise removal on image data in synchronized multi-source data frames based on the binarized region constraints in the semantic weighted mask to obtain a set of truss line segments of the transmission tower; a topology construction subunit, used to take the intersection points of line segments in the truss line segment set as candidate physical nodes and construct a topology weighted response map based on the candidate physical nodes; and a feature locking subunit, used to perform structure-aware nonmaximum suppression and feature locking on the topology weighted response map to obtain a set of rigid feature points.

Citation Information

Cited By

  • An unmanned aerial vehicle real-time target tracking method based on asynchronous semantic correction

    CN122176012A

  • An unmanned aerial vehicle real-time target tracking method based on asynchronous semantic correction

    CN122176012B