Unmanned aerial vehicle motion inversion and bridge cable stable tracking method based on video panoramic segmentation
By constructing a component-level panoramic segmentation model for bridge videos, the problems of multi-category background segmentation and UAV motion inversion in bridge cable inspection scenarios were solved. Stable tracking of bridge cable instances and accurate filtering of UAV motion were achieved, improving the accuracy and reliability of bridge structural health monitoring.
Patent Information
- Application Number
- CN202611043083.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-25
AI Technical Summary
Existing technologies cannot achieve multi-category background segmentation in bridge cable inspection scenarios, make it difficult to maintain the consistency of the segmentation mask within the video time frame, and cannot accurately invert the motion of drones, resulting in difficulties in bridge cable instance segmentation and targetless drone motion filtering.
A component-level panoramic segmentation model for bridge videos is constructed, including a semantic segmentation branch for bridge video background, a branch for segmentation and tracking of bridge cable instances, and a panoramic information fusion module. The model achieves UAV motion inversion and filtering through multi-geometric feature matching.
Stable tracking of bridge cable instances and accurate inversion and filtering of UAV motion were achieved, improving the accuracy and reliability of bridge structural health monitoring.
Smart Images

Figure CN122636652A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary fields of structural health monitoring, intelligent bridge operation and maintenance, computer vision, vibration measurement, and video processing, and particularly to a method for UAV motion inversion and bridge cable stability tracking based on video panoramic segmentation. This method is widely applied in engineering practices such as UAV visual measurement, non-contact vibration identification, structural health monitoring, and intelligent bridge operation and maintenance for long-span cable-stayed bridge systems. Background Technology
[0002] As a critical load-bearing component of cable-stayed bridges, the structural integrity of bridge cables directly affects the overall performance and operational safety of the bridge. Throughout the bridge's lifespan, bridge cables not only endure repeated external forces such as wind loads, vehicle dynamic loads, and impacts, but also face the long-term effects of environmental factors such as corrosion and material aging, making them susceptible to varying degrees of damage. Therefore, conducting real-time monitoring of bridge cables and ensuring the safety and reliability of key components is of great significance for guaranteeing bridge operational safety and extending its service life.
[0003] Vision-based structural vibration measurement systems, with their advantages of being non-contact, high-precision, and capable of multi-point simultaneous monitoring, have been widely used in the health monitoring of large bridge structures. For core components like bridge cables, which have a wide spatial distribution, mounting cameras on mobile measurement platforms such as drones enables rapid surveying. However, the drones themselves move while collecting video of bridge cable vibrations, causing aliasing between the drone's movement and the actual vibration signals of the cables, affecting measurement accuracy. Furthermore, existing vision-based structural vibration measurement methods generally face technical challenges such as structural target identification and stable tracking.
[0004] Existing research typically employs two strategies for inverting and filtering UAV motion. One is to install an inertial measurement unit (IMU) to obtain the displacement matrix by integrating acceleration; however, this integration process introduces accumulated errors, which is particularly detrimental to monitoring minute vibrations. The second method involves manually placing targets on stationary structures and estimating UAV displacement by tracking their motion. However, bridge cable inspections require regular checks, and the targets are not only difficult to place but also susceptible to environmental corrosion, requiring frequent maintenance. It is worth noting that neither of these methods fully utilizes the rich information inherent in the images themselves.
[0005] Segmentation models in computer vision have laid the foundation for the accurate identification of bridge structural components. Segmentation masks based on structural components can select regions for structural feature extraction, and then use feature extraction, tracking, or matching techniques to measure structural motion. However, existing segmentation models are mostly geared towards general image processing tasks, essentially representing binary segmentation of a single foreground and background. They struggle to achieve multi-instance segmentation of the foreground and multi-class segmentation of the background, thus failing to invert and filter out targetless UAV motion based on a stationary marker in the background. Furthermore, image-based segmentation methods struggle to maintain semantic consistency among multiple masks across video frames. In bridge cable inspection scenarios, bridge cables often have similar geometric features, and some videos contain more than one cable instance simultaneously, further increasing the difficulty of achieving automated and stable tracking of bridge cable instances using image segmentation methods.
[0006] In summary, to achieve stable tracking of bridge cable instances and motion inversion and filtering of targetless UAVs, the following challenges urgently need to be overcome: (1) Existing models lack the ability to segment backgrounds in the scenario of bridge cable UAV inspection. At the same time, it is difficult to segment multiple instances of bridge cables with similar geometric features, and it is impossible to maintain the consistency of each mask instance within the video time period. As a result, static calibration structures and bridge cable instances cannot be accurately and stably segmented and aligned.
[0007] (2) Due to the limitations of existing segmentation models, it is difficult to accurately extract information of stationary calibration objects in the background, and it is impossible to realize the inversion and filtering of UAV motion based on image information without a target.
[0008] To address the aforementioned issues, this invention proposes a method for UAV motion inversion and bridge cable stabilization tracking based on video panoramic segmentation. This method achieves semantic segmentation of bridge background information and instance segmentation of bridge foreground information, while maintaining the consistency of relevant information throughout the video timeline. Based on the segmentation results of stationary calibration objects, it enables the inversion and filtering of UAV motion. Summary of the Invention
[0009] The purpose of this invention is to solve the problems in the prior art and to propose a method for UAV motion inversion and bridge cable stability tracking based on video panoramic segmentation.
[0010] This invention is achieved through the following technical solution: This invention proposes a method for UAV motion inversion and bridge cable stability tracking based on video panoramic segmentation, the method comprising: Step 1: Build a component-level panoramic video segmentation model of the bridge; Step 2: Construct a large-scale cable-stayed bridge panoramic segmentation dataset incorporating bridge structural information; Step 3: Design a training strategy for a component-level bridge video panoramic segmentation model; Step 4: Construct a UAV motion inversion and filtering method based on point-line-surface multi-geometric feature matching.
[0011] Furthermore, in step one, the overall architecture of the component-level bridge video panoramic segmentation model is designed; the designed component-level bridge video panoramic segmentation model consists of three parts: the bridge video background semantic segmentation branch, the bridge cable instance segmentation and tracking branch, and the panoramic information fusion module guided by the bridge cable segmentation. The bridge video background semantic segmentation branch is responsible for semantically segmenting the background region of the bridge cable inspection drone video frames and maintaining the semantic consistency of each type of mask between frames. The bridge cable instance segmentation and tracking branch performs instance-level identification and classification of bridge cables, and ensures that the same bridge cable has a consistent tracking identity throughout the entire video sequence by matching the similarity of instance masks between frames. Finally, the panoramic information fusion module guided by bridge cable segmentation merges the video masks output by the two branches to generate a complete panoramic segmentation video.
[0012] Furthermore, the constructed bridge video background semantic segmentation branch architecture consists of six parts: an image encoder, an auto-completion encoder, a memory system, a mask decoder, a classification head, and a mask evaluation and filtering module. The bridge cable instance segmentation and tracking branch architecture consists of two parts: a bridge cable instance segmentation module and a bridge cable instance mask post-processing module. The design of the panoramic information integration module guided by bridge cable segmentation: Based on the segmentation results of bridge cable instances, a panoramic information integration module is designed to generate panoramic segmentation results of bridge cable inspection videos.
[0013] Furthermore, step two specifically includes: Step 21: Select the parameters for UAV video acquisition of bridge cable vibration; Step 22: Annotate the drone video frames in the bridge cable inspection scenario, integrate bridge cable instance information and background semantic information into them, and expand the data scale using data augmentation.
[0014] Furthermore, step three specifically includes: Step 31: Design a training strategy for the semantic segmentation branch of the bridge video background; Step 32: Design a training strategy for the bridge cable instance segmentation module; extract bridge cable instance mask annotations from the bridge panoramic segmentation dataset, and introduce image instance segmentation training loss to enable the model to have the ability to perceive bridge cable instances.
[0015] Furthermore, in step 31, a loss function is designed to fine-tune its parameters. The designed fine-tuning loss function consists of image segmentation loss. With mask classification loss It consists of two parts; The total loss function used for parameter fine-tuning in the image background semantic segmentation part is:
[0016] In the formula, Loss The total training loss for the semantic segmentation branch of the bridge video background; These are the weighting coefficients for the corresponding loss terms; For image basic segmentation loss, To segment auxiliary loss.
[0017] Furthermore, in step 32, the total loss used to train the bridge cable instance segmentation module is:
[0018] In the formula, The total loss for training the bridge cable instance segmentation module; These are the weight coefficients for bounding box regression loss, instance segmentation loss, classification loss, distribution focus loss, and angle loss, respectively.
[0019] Furthermore, step four specifically includes: Step 41: Construct a UAV motion inversion method based on multi-dimensional geometric features. First, separate the background semantic mask from the panoramic segmentation results. Then, through cross-graph feature interaction and iterative optimization, gradually enhance the distinguishability of the features and simultaneously predict the confidence of each point having a corresponding matching point. Finally, use the enhanced feature vectors to calculate the inner product similarity and multiply it with the matching confidence of the corresponding point to obtain the matching score matrix. Finally, through mutual nearest neighbor filtering, only retain point pairs that are bidirectionally consistent and have a score greater than zero as the output matching results. Step 42: Establish and verify the motion filtering method for unmanned aerial vehicles (UAVs).
[0020] Further, step four two specifically involves: applying the estimated UAV translational motion back to subsequent frame images to compensate for UAV motion and achieve motion correction relative to the first frame; to quantitatively evaluate the filtering effect of the correction on UAV motion, the same translational transformation is applied to the multi-category background semantic masks of subsequent frames to obtain the corrected masks, and then the intersection-union ratio (IU / I) scores of each category mask before and after correction and the corresponding mask of the first frame are calculated respectively.
[0021] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the UAV motion inversion and bridge cable stabilization tracking method based on video panoramic segmentation.
[0022] The beneficial effects of this invention are: This invention addresses the problems in existing UAV video segmentation methods for bridge cable vibration inspection, such as the inability to achieve instance-level segmentation of bridge cables, lack of multi-category semantic understanding of background information, difficulty in maintaining temporal consistency of segmentation masks across video frames, difficulty in utilizing image information for UAV motion inversion, and difficulty in achieving stable tracking of bridge cable targets. It proposes a UAV motion inversion and stable tracking method for bridge cable suspenders based on video panoramic segmentation. The method described in this invention has the following improvements and advantages: (1) Construct the overall architecture of the component-level bridge video panoramic segmentation model, and form a unified processing framework for foreground bridge cable instance segmentation and background information semantic segmentation for bridge cable vibration UAV inspection videos, effectively integrating instance-level and semantic-level segmentation capabilities. (2) Design a semantic segmentation branch architecture for bridge video background, introduce a memory mechanism and configure a classification head to achieve multi-category semantic segmentation of the background while enhancing the semantic consistency of each segmentation mask in the video temporal domain; (3) Design a bridge cable instance segmentation module and add an instance mask rotation angle prediction branch in its output head to provide support for subsequent downstream tasks such as mask shape optimization based on angle information and extraction of bridge cable vibration displacement time history; (4) Design a post-processing module for the bridge cable instance mask, optimize and correct the mask shape based on the predicted rotation angle information, and use the mask geometric features to maintain the instance consistency of the same bridge cable target between adjacent frames. (5) Construct a large-scale cable system bridge panoramic segmentation dataset that incorporates bridge structural information, providing a high-quality data foundation covering diverse background semantics and bridge cable instance annotations for subsequent model training, and supporting the joint embedding of semantic and instance information in the panoramic segmentation model; (6) Design a training strategy for the semantic segmentation branch of bridge video background, and use targeted fine-tuning methods to enable the model to fully learn the multi-class background information in the inspection scene, thereby improving the accuracy and robustness of semantic segmentation; (7) Design a training strategy for the bridge cable instance segmentation module, and enhance the model’s ability to identify bridge cable targets through a dedicated fine-tuning method to ensure the accuracy of instance segmentation and inter-frame stability; (8) Construct a UAV motion inversion and filtering method based on point-line-plane multi-geometric feature matching, use multi-dimensional geometric feature matching to achieve high-precision inversion of UAV motion parameters, and design corresponding motion filtering evaluation indexes to quantitatively evaluate the motion compensation effect. Attached Figure Description
[0023] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.
[0024] Figure 1 This is a flowchart of a method for UAV motion inversion and bridge cable stability tracking based on video panoramic segmentation.
[0025] Figure 2 This is a component-level bridge video panoramic segmentation model architecture diagram.
[0026] Figure 3 This is a diagram of the semantic segmentation branch architecture for bridge video backgrounds.
[0027] Figure 4 This is a diagram of the video segmentation mask classification header architecture.
[0028] Figure 5 This is a diagram of the bridge cable instance segmentation and tracking branch architecture.
[0029] Figure 6 This is a diagram of the mask model output header architecture for enhancing bridge cable angle information.
[0030] Figure 7 This is a schematic diagram of a portion of the bridge cable UAV video panoramic segmentation dataset.
[0031] Figure 8 This is a training strategy diagram for the semantic segmentation branch of the bridge video background.
[0032] Figure 9 This is an architecture diagram of a UAV motion inversion method based on multidimensional geometric features. Detailed Implementation
[0033] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0034] Specifically, in combination Figures 1-9 This invention proposes a method for UAV motion inversion and bridge cable stability tracking based on video panoramic segmentation, the method comprising: Step 1: Build a component-level panoramic video segmentation model of the bridge; Step one specifically includes the following steps: Step 11: Design the overall architecture of the component-level bridge video panoramic segmentation model The designed component-level bridge video panoramic segmentation model consists of three parts: a bridge video background semantic segmentation branch, a bridge cable instance segmentation and tracking branch, and a panoramic information fusion module guided by bridge cable segmentation. The model architecture is as follows: Figure 2 As shown in the diagram, the bridge video background semantic segmentation branch is responsible for semantically segmenting the background region of the bridge cable inspection UAV video frames and maintaining semantic consistency of each category of mask across frames. The bridge cable instance segmentation and tracking branch performs instance-level identification and classification of bridge cables, and ensures that the same bridge cable has a consistent tracking identity throughout the entire video sequence through inter-frame instance mask similarity matching. Finally, the panoramic information fusion module guided by bridge cable segmentation merges the video masks output from the two branches to generate a complete panoramic segmented video.
[0035] Steps 1 and 2: Constructing the semantic segmentation branch architecture for bridge video background The semantic segmentation branch for bridge video backgrounds consists of six parts: an image encoder, an auto-completion encoder, a memory system, a mask decoder, a classification head, and a mask evaluation and filtering module. The branch architecture is as follows: Figure 3 As shown.
[0036] The image encoder employs a four-stage hierarchical pyramid structure, extracting multi-scale features from local details to global semantics through progressive downsampling. Specifically, Stage 1 generates a feature map with a resolution of H / 4×W / 4 of the input image, with a relatively small number of channels (e.g., 144 dimensions), focusing on capturing fine-grained information such as edges and textures. Stage 2 halves the resolution to H / 8×W / 8, increasing the number of channels to 288, and begins to perceive intermediate semantics such as object parts and local structures. Stage 3 further downsamples to H / 16×W / 16, reaching 576 channels, upgrading the feature representation to high-level semantics, capable of reflecting complete object or scene categories. Stage 4 finally reduces the resolution to H / 32×W / 32, expanding the number of channels to 1152, focusing on global context and abstract semantics. At the application level, the high-resolution features from the first two stages are fed to the mask decoder to support the restoration of fine details; the low-resolution features from the latter two stages are handed over to the memory system for modeling contextual relationships.
[0037] The auto-completion encoder can encode user-inputted prompts such as points, bounding boxes, or masks into prompt embeddings that the model can understand. When user input is lacking, an automatic mask generator can be used to automatically generate regular grid points across the entire image as prompts, which can then guide the model to perform instance segmentation.
[0038] The memory system is the core module that maintains consistency among mask instances throughout the video timeline. It consists of three parts: a memory attention module, a memory encoder, and a memory bank.
[0039] The memory attention module is a temporal feature enhancement unit based on the Transformer decoder architecture. Its core objective is to dynamically enhance the visual representation of the current frame by incorporating memory information from historical frames. This design enables the model to maintain stable and continuous segmentation output even in challenging scenarios such as severe target occlusion, rapid movement, or significant deformation. Its working mechanism can be summarized as follows: projecting the latent features of the current frame as a query, projecting historical memory features as keys and values respectively, and then injecting the weighted aggregated memory information into the current frame features by calculating the similarity distribution between the query and all memory keys, thereby achieving the temporal reasoning ability of "looking back to understand the present." The specific calculation process is as follows:
[0040] In the formula, Q is the query matrix, projected from the current frame; K is the key matrix, projected from the memory features; and V is the value matrix, projected from the memory features. M represents the current frame features, the output of the image encoder for the current video frame; M represents the memory features, the set of stored historical frame features. It is a learnable linear projection weight matrix used to map input features to the attention space; This is a scaling factor used to stabilize the gradient; The weighted aggregated memory features are the information most relevant to the current frame extracted from the memory bank; softmax is the normalized exponential function that transforms the original scores into a probability distribution; Proj is the output projection layer. The final enhanced features, after memory enhancement, will be used as input to the mask decoder.
[0041] The memory encoder is responsible for transforming the visual information and segmentation results of the current frame into a compact, storable, and reusable memory representation. This memory encoder receives the current frame's visual features from the image encoder and the predicted mask logits from the mask decoder. First, it performs dimensionality reduction on the high-resolution mask to improve computational efficiency. Then, it converts the mask logits into a probability distribution. Next, it performs deep fusion of the downsampled mask and image features through a Fuser module composed of multiple CXBlock layers, allowing the spatial information of the mask to fully interact with the semantic features of the image. The fused features are further augmented with positional encoding to preserve spatial context information. The final output is a fixed-dimensional memory feature vector, which is stored in a memory bank for use by the memory attention module in subsequent frames. The calculation method is as follows:
[0042] In the formula, This represents the final generated memory features; PE stands for positional encoding. Inject spatial location information into the features; This is the predicted mask output by the mask decoder; The Sigmoid function converts the mask into a probability graph. For the mask downsampling module, multi-layer stride convolution is used to gradually reduce the spatial resolution and adjust the number of channels, and finally outputs a feature map with the same dimension as the image features; 1 1. Convolution, used to align the channels of image features; It is a fusion unit composed of multiple CXBlocks.
[0043] The memory bank is used to store the memory features of historical frames generated by the memory encoder, forming a structured temporal information cache. It provides historical context information for the memory attention module, enabling the model to refer to the target state of past frames when processing the current frame. This maintains the temporal consistency of segmentation in complex scenarios such as occlusion and deformation, and supports interactive error correction and long video streaming.
[0044] The mask decoder is responsible for deeply fusing image features, cue encoding information, and temporal enhancement features output by the memory attention module. While completing image-level segmentation, it enables the model to effectively refer to the target state of historical frames, thereby achieving stable and continuous segmentation and tracking in complex scenarios such as occlusion and deformation, and finally generating accurate segmentation masks for each instance of the image background.
[0045] The video segmentation mask classification head is a key module that imbues different background instances in an image with semantic information. It is the core component for achieving background semantic segmentation, and its specific architecture is as follows: Figure 4 As shown.
[0046] This module uses a multilayer perceptron (MLP) to map the mask tokens output from the mask decoder to category logits, and then converts them into the final predicted category labels via an Argmax operation, thereby giving the model the ability to distinguish between multi-class backgrounds with fine detail. The calculation process is as follows:
[0047] In the formula, The class labels predicted by the model; To return the logits of the first c The index with the largest component c This is used to obtain the final category; C Total number of categories; , , These are the weight matrices for the fully connected layer and the output layer, respectively. , , , , are the bias vectors of the fully connected layer and the output layer, respectively; is the activation function; x is the input feature vector.
[0048] By introducing a classification head, the model can accurately identify different background categories such as bridge towers, other bridge components, and distant environmental features in the bridge structure, providing reliable support for achieving accurate semantic segmentation of stationary objects.
[0049] The mask evaluation and filtering module is responsible for comprehensively evaluating the quality and spatial rationality of the generated candidate masks to ensure the accuracy and reliability of the final output. This module consists of three cooperating sub-components: the Intersection over Union (IoU) score prediction head, the object score head, and the non-overlapping constraints module. The IoU prediction head maps mask features to predicted IoU scores using a lightweight multilayer perceptron, estimating the overlap between the predicted mask and the real target, providing a quantitative basis for ranking the confidence of candidate masks. The object score head outputs a scalar score to indicate whether the mask region contains a valid object, effectively suppressing false positive masks caused by misdetection. The non-overlapping constraints module introduces spatial consistency penalties or post-processing logic to force the predicted masks of different instances within the same image to not overlap, thereby avoiding logical conflicts caused by multiple candidate masks vying for the same pixel region.
[0050] Step 13: Build the bridge cable instance segmentation and tracking branch architecture The bridge cable instance segmentation and tracing branch architecture consists of two parts: a bridge cable instance segmentation module and a bridge cable instance masking post-processing module. The branch architecture is as follows: Figure 5 As shown.
[0051] The bridge cable instance segmentation module employs an image instance segmentation model to achieve pixel-level accurate segmentation of the bridge cables. This module first utilizes a backbone network to extract multi-level features from the input image, and then a neck network fuses and enhances these multi-scale features to improve the model's semantic understanding and spatial sensitivity to the target structure. Furthermore, the model incorporates a bridge cable angle prediction branch in its output header. The output of this branch guides the mask post-processing module in optimizing based on angle information and provides support for the accurate extraction of the main vibration direction of the bridge cables in downstream tasks. The output header architecture of the mask model with enhanced bridge cable angle information is as follows: Figure 6 As shown.
[0052] In the segmentation head, all prediction branches employ a three-layer convolutional cascade design. They receive multi-scale feature maps from the Neck, undergo continuous convolution operations, and simultaneously output bounding box regression values, class scores, mask coefficients, and initial inactive values for the rotation angle. For the rotation angle, its initial inactive values are first activated by the Sigmoid function, and then linearly scaled to obtain the actual angle value. The specific transformation process is as follows:
[0053] In the formula, The input feature maps are from multiple levels of Neck. B These are the bounding box regression parameters; S For categorized scores; M These are the mask coefficients; The original logits represent the rotation angle; This is the final output angle of the model.
[0054] The bridge cable instance mask post-processing module aims to refine and optimize the mask to enhance its continuity and maintain consistency across frames throughout the video timeline. First, based on the output mask and its corresponding rotation angle, the module utilizes the mask's maximum width and leverages the structural characteristic that slender bridge cables can typically penetrate any two sides of an image to extend the mask's long side along the rotation direction to both ends of the image, thus completing the mask. This operation fills small holes within the mask and corrects defects where some masks fail to completely cover the bridge cable. Next, the module sorts the masks in the first frame according to their center point coordinates and assigns a unique trajectory number to each mask. Based on this, the module extracts the optimized mask's center point coordinates, aspect ratio, area, and other geometric features, and, in conjunction with the rotation angle output by the model, calculates the geometric similarity between masks frame by frame. This yields a matching score between masks in two consecutive frames. The mask with the highest matching score inherits the trajectory number of the corresponding mask from the previous frame, ensuring a stable identity association for the bridge cable target across frames and enabling long-term reliable tracking.
[0055] Step 14: Design the panoramic information integration module for bridge cable segmentation guidance Given the high confidence level of the bridge cable instance segmentation results, a panoramic information integration module is designed based on these results to generate panoramic segmentation results for bridge cable inspection videos. During mask fusion, when the same pixel contains both foreground bridge cable instance information and background semantic information, the background semantic information is discarded, retaining only the bridge cable instance mask. This ensures the integrity and coherence of the segmented bridge cable region, and the segmentation boundary of the background semantic mask is optimized accordingly. Through this strategy, a panoramic segmentation result for UAV bridge cable inspection videos that combines detailed foreground instance segmentation with complete background semantic information can be obtained.
[0056] Step 2: Construct a large-scale cable-stayed bridge panoramic segmentation dataset incorporating bridge structural information; Step two specifically includes: Step 21: Select the video acquisition parameters for bridge cable vibration drone. According to the sampling theorem, the video sampling frame rate must be at least twice the highest frequency component of the bridge cable vibration signal under test. Based on the inspection report of large cable-stayed bridge systems, the fundamental frequency of the bridge cables in such systems can reach 2.0Hz. According to the harmonic relationship, to achieve the extraction of higher-order frequencies of cable vibration in downstream tasks, the video acquisition frame rate should be no less than 30Hz. In actual acquisition, to avoid affecting the daily operation of the bridge, the distance between the drone and the bridge should be no less than 5m. Since the movement of the bridge cables is a micro-vibration, to ensure accurate extraction of the cable movement, the highest possible resolution should be used. The drone video acquisition parameters for bridge cable vibration inspection scenarios are shown in Table 1.
[0057] Table 1. Parameters for UAV Video Acquisition of Bridge Cable Vibration
[0058] Step 22: Annotate the drone video frames in the bridge cable inspection scenario, integrate bridge cable instance information and background semantic information into them, and expand the data scale using data augmentation.
[0059] Video frames were extracted from UAV videos of bridge cable vibration inspection scenarios. To comprehensively cover various scenarios, fully expand the original data, and address potential unforeseen circumstances during video acquisition, a frame was acquired every 300 or 600 frames, based on the video field of view and frame rate, constructing an original image dataset for bridge cable vibration inspection scenarios. Annotation tools were used to annotate the original images frame by frame, integrating bridge cable instance information and background semantic information into the dataset. This provides multi-class information on bridge structure and complex backgrounds in the inspection scenario for component-level bridge video panoramic models. Background multi-class segmentation is fundamental to achieving UAV motion inversion based on stationary calibration objects. Based on the pinhole camera model, the method for converting real physical coordinates to image coordinates is as follows:
[0060] In the formula, The coordinates of a three-dimensional point in the world coordinate system; Let R be the extrinsic parameter matrix of the camera, t be the rotation matrix, and K be the intrinsic parameter matrix of the camera. Focal length, expressed in pixels; The coordinates of the principal point (the intersection of the optical axis and the imaging plane) in the pixel coordinate system are usually located near the center of the image. s Scale factor; These are the pixel coordinates of a 3D point projected onto the image.
[0061] When a 3D object is projected onto a 2D plane, compression becomes more pronounced with increasing distance. Therefore, only stationary objects approximately on the same plane as the bridge cables will reflect camera motion in the image that aligns with the cable movement. For the cables in the bridge tower area, the towers can serve as clear stationary reference points. The beam structure, bridge surface, and bridge railings can be approximated as being on the same plane and labeled as the same category. Distant features such as buildings and trees are far from the cables and do not require separate classification, so they are grouped into one category. Areas without obvious features, such as the sky and sea, are treated as background information. Based on this, pixel-by-pixel multi-category labeling is performed on the bridge video background and cable instances.
[0062] To expand the dataset size and enhance data diversity, the labeled images and masks were augmented using data enhancement techniques such as translation, scaling, brightness adjustment, contrast adjustment, Gaussian noise addition, and Gaussian blurring. The composite transformation calculation formula is as follows:
[0063] In the formula, I Label the original input image with its corresponding mask; The output image and mask are labeled after composite data augmentation; For geometric transformation functions, first adjust by scaling factor. s Scale it, then translate it by the vector. t Perform translation; This is the contrast adjustment coefficient; This is the brightness offset, a constant directly added to the pixel value to adjust the overall brightness; It is additive Gaussian noise with a mean of 0 and a variance of . The normal distribution; The kernel is a Gaussian blurred convolution with a standard deviation of . .
[0064] The data augmentation process is performed simultaneously on the image and the mask to ensure spatial consistency. Part of the bridge cable UAV video panoramic segmentation dataset is shown below. Figure 7 As shown.
[0065] Step 3: Design a training strategy for a component-level bridge video panoramic segmentation model; Step three specifically includes: Step 31: Design a training strategy for the semantic segmentation branch of the bridge video background. The designed training strategy for semantic segmentation of bridge video background is as follows: Figure 8 As shown.
[0066] The memory system is a fundamental module for maintaining the semantic consistency of video background masks. After pre-training on a large video segmentation dataset, this module ensures the temporal consistency of targets such as vehicles and humans, thus allowing direct transfer to scenarios with small bridge cable movements. However, the image background semantic segmentation branch lacks prior knowledge of the bridge structure, necessitating the design of a corresponding loss function to fine-tune its parameters. The designed fine-tuning loss consists of two parts: image segmentation loss and mask classification loss.
[0067] Image segmentation loss It consists of two parts: the fundamental image segmentation loss. With segmentation auxiliary loss , A binary cross-entropy loss is used to measure the difference between the predicted mask and the true mask:
[0068] In the formula, N Total number of pixels; i For pixel index; The actual pixel label; For the model to the first i The logits output for each pixel correspond to the pixel value of the predicted mask.
[0069] This is a variant of cross-entropy loss, which forces the model to focus on hard-to-classify pixels in edge regions, effectively alleviating the problem of positive and negative sample imbalance. The calculation method is as follows:
[0070] In the formula, For the model to the first i The probability of predicting the "correct category" for each pixel; As an auxiliary loss value; This is a balancing factor used to adjust the relative contribution of positive and negative samples to the loss; To focus on parameters, control the decay rate of easily classified samples.
[0071] Masking Classification Loss Standard multi-class cross-entropy is used to supervise the semantic classification of instances, which is then used to train the classification head. The calculation method is as follows:
[0072] In the formula, B This is the batch size, i.e., the number of masks currently involved in the loss calculation; j Use the sample index to iterate through each mask in the current batch; C Total number of categories; c For the category index, iterate through all possible categories; For the model to the first j The output of the first instance c Logit value; For the first j The actual category label of each instance; For the model to the first j The logit value corresponding to the actual category output by each instance.
[0073] The total loss function used for parameter fine-tuning in the image background semantic segmentation part is:
[0074] In the formula, Loss The total training loss for the semantic segmentation branch of the bridge video background; These are the weighting coefficients for the corresponding loss terms.
[0075] Step 32: Design a training strategy for the bridge cable instance segmentation module Bridge cable instance mask annotations are extracted from the bridge panoramic segmentation dataset, and image instance segmentation training loss is introduced to enable the model to have the ability to perceive bridge cable instances.
[0076] Bounding box regression loss is used to achieve accurate localization of slender bridge cables. The calculation method is as follows:
[0077] In the formula, B is the predicted bounding box; The true bounding box; w , h These are the width and height of the prediction box, respectively; , These are the width and height of the ground truth bounding box, respectively; v is the aspect ratio consistency penalty factor, used to measure the difference in aspect ratio between the predicted bounding box and the ground truth bounding box. This is a dynamic tradeoff coefficient used to balance the aspect ratio penalty term. v The impact of IoU loss; This is the Euclidean distance function, which calculates the straight-line distance between two points; c The diagonal length of the smallest bounding rectangle that covers both the predicted and ground truth boxes; For bounding box regression loss; For the first i The target score for each prospective anchor point; i The index number for the foreground anchor point; The set of anchor points that are identified as positive samples by the label assigner; For the first i The coordinate vector of the prediction box corresponding to each anchor point; For the first i The coordinate vector of the ground truth bounding box corresponding to each anchor point.
[0078] The instance segmentation loss uses a hybrid of binary cross-entropy and Dice loss to optimize bridge cable instance mask prediction. The calculation method is as follows:
[0079] In the formula, This is the binary cross-entropy loss, used to measure the difference between the predicted probability of a single pixel and the true label; m For predicting the logit of the mask; These are the pixel values of the actual mask; The Dice loss is used to measure the overall overlap between the predicted mask and the real mask. It is a very small positive number, used to avoid the denominator being zero; For instance segmentation loss; N The total number of foreground instances; A set of indices for foreground instances; k The index number of the foreground instance; The area of the bounding box corresponding to the k-th instance; u , v () represents pixel space coordinates; For the first k The set of pixel positions inside the bounding box of an instance; The prediction mask logit for the k-th instance at position ( u , v The value at ); The true mask label for the k-th instance at position ( u , v The value at ().
[0080] The classification loss is used to effectively distinguish between the foreground and background of bridge cable instances. The calculation method is as follows:
[0081] In the formula, For classification loss; s The target score vector is the set of target scores for all anchor points. j The index number of the anchor point; For the first j The category logit corresponding to each anchor point; For the first j The target score for each anchor point, represented by a soft label generated by the label assigner, measures how well the anchor point aligns with the real target.
[0082] Distribution Focal Loss (DFL) improves regression accuracy and further enhances target localization accuracy by modeling bounding box coordinates as a discrete probability distribution, guiding the network to learn the uncertainty of boundary positions. The calculation method is as follows:
[0083] In the formula, The distribution focus loss is used to optimize the distribution prediction of bounding box regression; CE( d , y () is the multi-class cross-entropy loss, used to measure the predicted distribution. d With target index y Differences; For the first i Discrete distribution of anchor point predictions; , The first i The left and right boundary indices of the true bounding box coordinates corresponding to each anchor point in the discretized interval.
[0084] The original method indirectly applies angle constraints only through bounding box IoU calculation, which is an implicit optimization; however, the angle information of the bridge cables is crucial for accurately extracting their principal direction vibration characteristics. Therefore, based on the original bounding box regression, classification, and distribution focus loss, an independent angle loss term is explicitly introduced. Explicit supervision is added to the implicit constraints to enhance the model's angle learning ability and improve angle prediction accuracy. The calculation method is as follows:
[0085] In the formula, The angle loss is used to optimize the regression accuracy of the rotation angle of the bridge cable instance; The total number of foreground instances; i The index number of the foreground instance; For smoothing L1 loss function; , These are the predicted rotation angle and the actual rotation angle of the i-th foreground instance, respectively.
[0086] The total loss used to train the bridge cable instance segmentation module is:
[0087] In the formula, The total loss for training the bridge cable instance segmentation module; The weight coefficients for bounding box regression loss, instance segmentation loss, classification loss, distribution focus loss, and angle loss are respectively defined.
[0088] Step 4: Construct a UAV motion inversion and filtering method based on point-line-surface multi-geometric feature matching.
[0089] Step four specifically includes: Step 41: Constructing a UAV motion inversion method based on multi-dimensional geometric features The architecture of the UAV motion inversion method based on multidimensional geometric features is as follows: Figure 9 As shown.
[0090] First, the background semantic mask is extracted from the panoramic segmentation results. Based on the distance relationship between the foreground bridge cables and various background elements, a semantic mask priority strategy is designed: the semantic priority is ranked from high to low as follows: bridge towers, other bridge components, distant features, and background elements such as the sky and sea (river) surface. Based on the semantic masks actually appearing in the image background segmentation results, the highest priority mask region in the current frame is selected as the motion estimation region. Within this selected region, point features, line features, and mask geometric features are extracted using a point feature extractor, a line feature extractor, and a surface feature converter, respectively. All features are then fed into a feature matching network to achieve robust matching between each frame and the first frame, thereby estimating the UAV translational motion of subsequent frames relative to the first frame.
[0091] The matching network first fuses the keypoint coordinates and descriptors of two images to form the initial feature representation of each point. Then, through cross-image feature interaction and iterative optimization, the discriminability of the features is gradually enhanced, and the confidence level of each point having a corresponding matching point is predicted simultaneously. Low-confidence points are pruned immediately and no longer participate in subsequent calculations; the iteration process terminates early when the matching state of the two points tends to stabilize. Finally, the inner product similarity is calculated using the enhanced feature vectors and multiplied by the matching confidence level of the corresponding point to obtain the matching score matrix. Finally, through nearest neighbor filtering, only point pairs that are bidirectionally consistent and have a score greater than zero are retained as the output matching result. For the matching results, the median is taken to filter out outliers, resulting in the estimated UAV motion based on stationary calibration objects in the image background. The calculation method is as follows:
[0092] In the formula, P is a soft-assignment matrix, and its elements are... Point i With point j The matching probability; For row normalization operators; For column normalization operators; To proceed continuously T Alternating column normalization is performed until convergence. To expand the similarity matrix, a complete scoring table containing all possibilities of "match" and "mismatch" is included; Temperature coefficient; This is the final set of reliable matching point pairs for the output. These are the probability values in the soft assignment matrix; For the points in the first frame i The point with the highest probability in the second frame is selected as its candidate match. j ; For the points in the second frame j The point with the highest probability in the first frame is selected as its candidate match. i ; This is the matching probability threshold; The estimated translational displacement of the drone relative to the first frame is given by: [variable name], ... x Components and y The median of the components yields a robust translation estimate; For the j-th matching point in the current frame x , y coordinate; For the j-th matching point in the first frame x , y coordinate.
[0093] Step 42: Establish and verify the motion filtering method for unmanned aerial vehicles (UAVs). The estimated UAV translational motion is then applied to subsequent frames to compensate for the UAV's motion, achieving motion correction relative to the first frame. To quantitatively evaluate the filtering effect of this correction on UAV motion, the same translational transform is applied to multi-class background semantic masks in subsequent frames to obtain corrected masks. Then, the intersection-union (IUU) scores of each class of masks before and after correction are calculated with the corresponding mask in the first frame. The calculation method is as follows:
[0094] In the formula, The original crossover ratio (CROR) is used to measure the performance of the first unmanned aerial vehicle (UAV) without motion correction. t The first frame k The mask and the first one in frame 1 k The degree of overlap between corresponding targets; The corrected crossover ratio (CVR) is used to measure the performance of the UAV after motion correction. t The first frame k The mask and the first one in frame 1 k The degree of overlap between corresponding targets; For the first t The first frame k The segmentation mask for each target; For the first frame k The segmentation mask for each target; After drone motion correction, the first tThe first frame k A segmentation mask for each target.
[0095] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the UAV motion inversion and bridge cable stabilization tracking method based on video panoramic segmentation.
[0096] The memory in this application embodiment can be volatile memory or non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the methods described in this invention is intended to include, but is not limited to, these and any other suitable types of memory.
[0097] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).
[0098] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.
[0099] It should be noted that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuitry in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above methods.
[0100] The above provides a detailed description of the UAV motion inversion and bridge cable stabilization tracking method based on video panoramic segmentation proposed in this invention. Specific examples are used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.
Claims
1. A method for UAV motion inversion and bridge cable stability tracking based on video panoramic segmentation, characterized in that, The method includes: Step 1: Build a component-level panoramic video segmentation model of the bridge; Step 2: Construct a large-scale cable-stayed bridge panoramic segmentation dataset incorporating bridge structural information; Step 3: Design a training strategy for a component-level bridge video panoramic segmentation model; Step 4: Construct a UAV motion inversion and filtering method based on point-line-surface multi-geometric feature matching.
2. The method according to claim 1, characterized in that, In step one, the overall architecture of the component-level bridge video panoramic segmentation model is designed. The designed component-level bridge video panoramic segmentation model consists of three parts: the bridge video background semantic segmentation branch, the bridge cable instance segmentation and tracking branch, and the panoramic information fusion module guided by the bridge cable segmentation. The bridge video background semantic segmentation branch is responsible for semantically segmenting the background region of the bridge cable inspection drone video frames and maintaining the semantic consistency of each type of mask between frames. The bridge cable instance segmentation and tracking branch performs instance-level identification and classification of bridge cables, and ensures that the same bridge cable has a consistent tracking identity throughout the entire video sequence by matching the similarity of instance masks between frames. Finally, the panoramic information fusion module guided by bridge cable segmentation merges the video masks output by the two branches to generate a complete panoramic segmentation video.
3. The method according to claim 2, characterized in that, The constructed bridge video background semantic segmentation branch architecture consists of six parts: image encoder, auto-completion encoder, memory system, mask decoder, classification head, and mask evaluation and filtering module. The bridge cable instance segmentation and tracking branch architecture consists of two parts: a bridge cable instance segmentation module and a bridge cable instance mask post-processing module. The design of the panoramic information integration module guided by bridge cable segmentation: Based on the segmentation results of bridge cable instances, a panoramic information integration module is designed to generate panoramic segmentation results of bridge cable inspection videos.
4. The method according to claim 1, characterized in that, Step two specifically includes: Step 21: Select the parameters for UAV video acquisition of bridge cable vibration; Step 22: Annotate the drone video frames in the bridge cable inspection scenario, integrate bridge cable instance information and background semantic information into them, and expand the data scale using data augmentation.
5. The method according to claim 1, characterized in that, Step three specifically includes: Step 31: Design a training strategy for the semantic segmentation branch of the bridge video background; Step 32: Design a training strategy for the bridge cable instance segmentation module; extract bridge cable instance mask annotations from the bridge panoramic segmentation dataset, and introduce image instance segmentation training loss to enable the model to have the ability to perceive bridge cable instances.
6. The method according to claim 5, characterized in that, In step 31, a loss function is designed and its parameters are fine-tuned. The designed fine-tuning loss function consists of image segmentation loss. With mask classification loss It consists of two parts; The total loss function used for parameter fine-tuning in the image background semantic segmentation part is: In the formula, Loss The total training loss for the semantic segmentation branch of the bridge video background; These are the weighting coefficients for the corresponding loss terms; For image basic segmentation loss, This is used to segment auxiliary losses.
7. The method according to claim 5, characterized in that, In step 3.2, the total loss used to train the bridge cable instance segmentation module is: In the formula, The total loss for training the bridge cable instance segmentation module; These are the weight coefficients for bounding box regression loss, instance segmentation loss, classification loss, distribution focus loss, and angle loss, respectively.
8. The method according to claim 1, characterized in that, Step four specifically includes: Step 41: Construct a UAV motion inversion method based on multi-dimensional geometric features. First, separate the background semantic mask from the panoramic segmentation results. Then, through cross-graph feature interaction and iterative optimization, gradually enhance the distinguishability of the features and simultaneously predict the confidence of each point having a corresponding matching point. Finally, use the enhanced feature vectors to calculate the inner product similarity and multiply it with the matching confidence of the corresponding point to obtain the matching score matrix. Finally, through mutual nearest neighbor filtering, only retain point pairs that are bidirectionally consistent and have a score greater than zero as the output matching results. Step 42: Establish and verify the motion filtering method for unmanned aerial vehicles (UAVs).
9. The method according to claim 8, characterized in that, Step 42 specifically involves: applying the estimated UAV translational motion back to subsequent frame images to compensate for UAV motion and achieve motion correction relative to the first frame; to quantitatively evaluate the filtering effect of this correction on UAV motion, the same translational transformation is applied to the multi-category background semantic masks of subsequent frames to obtain the corrected masks, and then the intersection-union ratio (IU / I) scores of each category mask before and after correction and the corresponding mask of the first frame are calculated respectively.
10. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-9.