Unmanned aerial vehicle based cable-stayed bridge stay cable video multi-target recognition tracking and vibration extraction method

By improving the YOLOv11-obb model and StrongSORT algorithm, and combining SIFT/ORB feature point matching and variational mode decomposition technology, the problems of accuracy identification and multi-target tracking of cable-stayed bridge vibration in UAV videos were solved. High-precision extraction of cable-stayed bridge vibration signals and modal parameter identification were achieved, supporting health monitoring of long-span cable-stayed bridges.

CN120931896BActive Publication Date: 2026-06-19HARBIN INST OF TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HARBIN INST OF TECH
Filing Date
2025-07-23
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing technologies for UAV video-driven cable vibration recognition suffer from problems such as insufficient target detection accuracy, easy failure of multi-target tracking algorithms, and low accuracy of modal parameter recognition. In particular, it is difficult to achieve high-precision cable vibration signal extraction and modal parameter recognition in complex backgrounds and multi-cable intersection scenarios.

Method used

An improved model for detecting slender, tilted targets based on the YOLOv11-obb model is adopted, and the StrongSORT algorithm is integrated for multi-target tracking. The displacement is extracted by combining SIFT/ORB feature point matching and sub-pixel thinning technology. Furthermore, a UAV motion correction algorithm based on variational mode decomposition and joint time-frequency domain screening is designed, and a joint working mode analysis algorithm is constructed.

Benefits of technology

It achieves high-precision detection and multi-target stable tracking of cable stays under complex backgrounds, improves the extraction accuracy of cable stay vibration signals and the identification accuracy of modal parameters, breaks through the limitations of traditional fixed-point visual geography, and provides high-precision technical support for health monitoring of long-span cable stay bridges.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120931896B_ABST
    Figure CN120931896B_ABST
Patent Text Reader

Abstract

This invention proposes a method for multi-target recognition, tracking, and vibration extraction of bridge cable-stayed bridge videos based on unmanned aerial vehicles (UAVs). The method includes: Step 1: Constructing a refined tilted and slender target detection model for bridge cable-stayed bridges based on the YOLOv11 model; Step 2: Proposing a multi-target tracking algorithm that integrates the tilted and slender target detection model and the StrongSORT algorithm; Step 3: Improving the displacement extraction method by combining SIFT / ORB feature point matching and sub-pixel refinement techniques; Step 4: Designing a UAV motion correction algorithm based on variational mode decomposition and time-frequency domain joint screening; Step 5: Constructing a joint working mode analysis algorithm that combines natural excitation technology and random subspace recognition algorithm. This method achieves high-precision extraction of cable-stayed bridge vibration signals and identification of cable-stayed bridge modal parameters, providing technical support for health monitoring of long-span cable-stayed bridges.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of computer vision, target detection, multi-target tracking, vibration signal processing, and bridge structural health monitoring, particularly to a method for multi-target recognition, tracking, and vibration extraction based on UAV-based video of bridge stay cables. The method focuses on solving the challenges of detecting, tracking, and extracting vibration modes of slender, inclined targets such as bridge stay cables in UAV video scenarios. It can be widely applied in engineering fields such as structural health monitoring of long-span bridges, non-contact vibration detection in intelligent construction, and bridge operation and maintenance and disaster prevention and mitigation. Background Technology

[0002] As a critical infrastructure in transportation networks, the structural health monitoring of long-span cable-stayed bridges is crucial for ensuring traffic safety. The stay cables, as the core load-bearing components of cable-stayed bridges, directly reflect the overall performance of the bridge through their vibration. Traditional vibration identification technologies based on fixed-point computer vision are significantly limited by geographical conditions. In complex scenarios such as crossing rivers and valleys, wiring is difficult and the monitoring range is limited, making comprehensive coverage challenging. Drones, with their high mobility, flexible take-off and landing capabilities, and non-contact measurement, can easily reach areas inaccessible to traditional equipment, providing a new path for bridge vibration monitoring that is not constrained by terrain. However, extracting the vibration modes of stay cables from drone videos faces several technical challenges:

[0003] (1) Problem of complex background interference and target characteristic adaptation: The cable-stayed bridge presents a slender and inclined feature in the image. Traditional target detection models are not good at capturing the features of such slender targets, and the complex bridge background is prone to missed detection or positioning deviation. At the same time, the change of perspective during the flight of the UAV will cause dynamic changes in the spatial attitude of the cable-stayed bridge, which further increases the difficulty of detection.

[0004] (2) Problems with UAV motion interference and displacement measurement accuracy: The swaying of the UAV itself during flight introduces low-frequency vibration noise, the amplitude of which is often greater than the actual vibration displacement of the cable. Existing vibration correction methods rely on fixed background reference points or inertial measurement unit data, but cable monitoring scenarios often face situations such as a clean background with no reference points and the UAV not being equipped with an inertial unit, which limits the applicability of traditional correction methods.

[0005] (3) Efficiency issues of multi-target tracking and modal parameter identification: Cable-stayed bridges typically contain dozens to hundreds of cables. Traditional single-target tracking algorithms struggle to achieve stable tracking of multiple cables across frames. Meanwhile, existing modal decomposition methods (such as VMD) do not fully incorporate the differences in vibration characteristics between the UAV and the cables, resulting in insufficient accuracy in modal parameter identification, which fails to meet the needs of engineering applications.

[0006] In recent years, advancements in computer vision and deep learning technologies have provided new solutions to these problems. Deep learning-based target detection algorithms, by introducing adaptive anchor box mechanisms and multi-scale feature fusion, have demonstrated excellent performance in detecting slender targets, effectively improving the positioning accuracy of stay-at-home cables in complex backgrounds. Multi-target tracking algorithms, combined with appearance feature extraction techniques, can construct more robust target motion models. By introducing a spatial attitude transformation matrix, cross-frame correlation of stay-at-home cables under changing viewpoints can be achieved. Signal processing methods such as Variational Mode Decomposition (VMD), through adaptive frequency band segmentation, can separate signals based on the time-frequency characteristics of UAV jitter and stay-at-home cable vibration, exhibiting good noise suppression capabilities in vibration signal decomposition.

[0007] However, existing technologies still have significant shortcomings in UAV video-driven cable-stayed bridge vibration recognition: First, target detection models lack the ability to extract features from slender, inclined targets. Mainstream convolutional neural networks, due to their anchor frame design bias towards conventional rectangular targets, struggle to capture the slender geometric features of cable-stayed bridges. In complex bridge backgrounds (such as water reflections, vegetation obstructions, and interference from bridge tower structures), they often miss detections or have positioning errors, making it difficult to meet the required detection accuracy. Second, multi-target tracking algorithms do not consider the spatial attitude changes of cable-stayed bridges. When the UAV's flight perspective changes, the projection angle and length of the cable-stayed bridge in the video will dynamically change. Algorithms such as StrongSORT, lacking attitude compensation mechanisms, are prone to track trajectory breakage due to feature matching failures, especially in multi-cable crossing scenarios where tracking errors accumulate significantly. Third, mode decomposition methods do not effectively utilize the differences in vibration characteristics between UAVs and cable-stayed bridges, resulting in incomplete noise removal and low accuracy in mode parameter recognition.

[0008] To address the aforementioned issues, this invention proposes a method for multi-target identification, tracking, and vibration extraction of bridge cable-stayed bridge videos based on unmanned aerial vehicles (UAVs). This method improves the detection model for slender, inclined targets based on the YOLOv11-obb model, constructs a multi-target tracking algorithm based on the StrongSORT algorithm, and optimizes the mode decomposition method by updating the search algorithm and screening methods. This achieves high-precision extraction of cable-stayed bridge vibration signals and identification of cable-stayed bridge modal parameters, providing technical support for the health monitoring of long-span cable-stayed bridges. Summary of the Invention

[0009] The purpose of this invention is to solve the problems in the prior art and to propose a method for multi-target recognition, tracking and vibration extraction of bridge cable-stayed bridge videos based on UAVs.

[0010] This invention is achieved through the following technical solution: This invention proposes a method for multi-target recognition, tracking, and vibration extraction based on UAV-based bridge cable-stayed bridge video, the method comprising:

[0011] Step 1: Construct a refined model for detecting inclined and slender targets in bridge stay cables based on the YOLOv11 model;

[0012] Step 2: Propose a multi-target tracking algorithm that integrates the tilted slender target detection model and the StrongSORT algorithm;

[0013] Step 3: Improve the displacement extraction method by combining SIFT / ORB feature point matching and sub-pixel thinning techniques;

[0014] Step 4: Design a UAV motion correction algorithm based on variational mode decomposition and joint time-frequency domain screening;

[0015] Step 5: Construct a joint working mode analysis algorithm that combines natural excitation techniques with random subspace identification algorithms.

[0016] Furthermore, step one specifically includes:

[0017] Step 11: Collect cable stay image data and perform high-precision annotation. Expand the dataset through data augmentation and construct a YOLO format dataset for detecting inclined and slender targets of bridge cable stays.

[0018] Steps 1 and 2: Design the C3K2_DSCA module, introducing dynamic serpentine convolution and axial cross attention to enhance the ability to capture cable features;

[0019] Step 13: Use the Angleloss loss function to optimize the angle prediction accuracy and establish a refined detection model for inclined and slender targets of bridge stay cables.

[0020] Furthermore, in steps one and three, when calculating the Angleloss function, the positive sample angles are first predicted, then normalized, and then the angle distribution focusing loss is calculated based on the selected positive sample angle predictions and the normalized target angles. The specific implementation of this loss function in angle regression is shown below:

[0021] (2)

[0022] In the formula, N The number of positive samples; S This represents the sum of the weights of the positive samples. For sample weights; and Weights for the left and right intervals of the target angle; Predict the distribution based on the angle; and The left and right interval indices for the target angle; This is for calculating cross-entropy.

[0023] Furthermore, step two specifically includes:

[0024] Step 21: Connect the YOLOv11-DSCA target detection model and add angle parameters;

[0025] Step 22: Introduce the OSNet network for appearance extraction to enhance appearance extraction capabilities;

[0026] Steps 2 and 3: Introduce the ProbIOU metric to adapt to matching in rotation scenarios.

[0027] Furthermore, step three specifically includes:

[0028] Step 31: Construct a stable feature point tracking baseline region and divide the ROI region for feature point extraction in each frame of video;

[0029] Step 32: Introduce a sub-pixel corner refinement function and extract sub-pixel level feature point positions through SIFT / ORB feature point matching;

[0030] Step 33: Optimize the RANSAC algorithm to achieve adaptive scaling and hierarchical screening.

[0031] Furthermore, during the execution of the RANSAC algorithm, the dynamic threshold generation mechanism calculates the reprojection error threshold in real time based on the current frame feature distribution characteristics and historical displacement statistics; the dynamic threshold generation formula is as follows:

[0032] (16)

[0033] In the formula, The dynamic threshold for the current frame. The scaling factor. The number of feature points, For the first Displacement of each feature point This is the historical average displacement. Based on the threshold bias.

[0034] Furthermore, step four specifically includes:

[0035] Step 41: Based on the vibration characteristics of UAVs and cable-stayed bridges, construct a joint time-frequency domain screening method;

[0036] Step 42: Introduce the triangular topology optimization search algorithm to optimize the key parameters of the variational mode decomposition algorithm.

[0037] Furthermore, step five specifically includes:

[0038] Step 51: A two-stage joint analysis of vibration data is conducted using natural excitation technology and a random subspace identification algorithm.

[0039] Step 52: Filter and merge similar modalities through hierarchical clustering.

[0040] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the UAV-based method for multi-target recognition, tracking and vibration extraction of bridge cable-stayed bridge video.

[0041] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the UAV-based method for multi-target recognition, tracking, and vibration extraction of bridge cable-stayed bridge videos.

[0042] The beneficial effects of this invention are:

[0043] (1) The present invention optimizes the target detection module. Based on the slender and inclined characteristics of the cable-stayed cable, dynamic snake convolution and axial cross attention mechanism are introduced based on the YOLOv11s framework. The C3K2_DSCA module is designed and the Angleloss loss function is used to optimize the angle prediction accuracy, thus solving the problems of missed detection and positioning deviation in complex backgrounds.

[0044] (2) This invention proposes a multi-target tracking algorithm that integrates tilted target detection. Based on StrongSORT, it adopts the YOLOv11-DSCA detection model, introduces OSNet to extract external features, uses ProbIOU to match the detection box, and improves the angle parameterized Kalman filter to achieve stable tracking of multiple targets across frames.

[0045] (3) Based on multi-target tracking, this invention combines SIFT / ORB feature point matching and sub-pixel displacement tracking technology to provide high-precision trajectory data for vibration signal time-domain analysis and improve displacement calculation accuracy.

[0046] (4) The present invention designs an improved VMD method based on the vibration characteristics of UAVs, introduces a triangular topology optimization search algorithm, and uses a joint mechanism of time-domain exponential envelope fitting and frequency-domain fundamental frequency screening to screen the intrinsic mode functions that conform to the exponential decay law;

[0047] (5) This invention combines natural excitation technology with covariance-driven random subspace recognition algorithm to identify working modes, realize high-precision extraction of cable-stayed bridge mode parameters, and solve the problem of mode aliasing;

[0048] (6) This invention realizes the extraction of cable vibration signals and identification of modal parameters based on UAV video, breaking through the traditional fixed-point visual geographical limitations, and providing a high-precision technical solution for the health monitoring of cable-stayed bridges that is applicable to engineering scenarios. Attached Figure Description

[0049] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0050] Figure 1 This is a flowchart of a method for multi-target recognition, tracking, and vibration extraction of bridge cable stays based on UAV video.

[0051] Figure 2 This is a flowchart of the construction process for a refined detection model of slender, inclined targets for bridge stay cables.

[0052] Figure 3 This is a diagram of the internal architecture of the DSCA module, a key module for feature extraction of cable-stayed bridges, and the dynamic serpentine convolutional block.

[0053] Figure 4 This is a schematic diagram showing the changes made by the dynamic serpentine convolution kernel in the dynamic serpentine convolution block to the characteristics of the cable-stayed bridge.

[0054] Figure 5 This is a diagram of the internal architecture of the YOLOv11-DSCA cable-stayed bridge-specific target detection model.

[0055] Figure 6 This is a flowchart of the multi-target tracking algorithm for cable-stayed bridges.

[0056] Figure 7 This is a diagram of the OSNet network's full-scale feature learning module and multi-scale fusion architecture for extracting the appearance of cable-stayed bridges.

[0057] Figure 8 This is a flowchart of the cable-stayed bridge displacement extraction process, which combines feature point matching algorithms with sub-pixel thinning techniques.

[0058] Figure 9 This is a flowchart of a variational mode decomposition algorithm optimized in combination with a triangular topology optimization search algorithm.

[0059] Figure 10 This is a schematic diagram of the construction of interior points in a triangular topology optimization algorithm used for searching variational mode decomposition.

[0060] Figure 11 This is a flowchart of the joint working modal analysis algorithm applied to the vibration modal analysis of cable-stayed bridges. Detailed Implementation

[0061] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0062] Combination Figures 1-11 This invention proposes a method for multi-target recognition, tracking, and vibration extraction based on UAV-based bridge cable-stayed cable video, the method comprising:

[0063] Step 1: Construct a refined model for detecting inclined and slender targets in bridge stay cables based on the YOLOv11 model;

[0064] Step 1 implementation process is as follows Figure 2 As shown, step one specifically includes:

[0065] Step 11: Collect cable stay image data and perform high-precision annotation. Expand the dataset through data augmentation and construct a YOLO format dataset for detecting inclined and slender targets of bridge cable stays.

[0066] The first step involves collecting and accurately labeling cable-stayed bridge images through on-site photography and online data collection. Data augmentation is then used to expand the dataset, constructing a YOLO-formatted dataset for detecting slender, inclined targets in bridge cable-stayed bridges. The specific steps are as follows:

[0067] The dataset primarily originates from a collaborative integration of on-site photography and online resources. On-site photography needed to cover the actual environment of the bridge, taking photos from multiple angles (eye-level view of the bridge deck structure, upward view of the cable-stayed towers, and downward view of the river surface, etc.) and distances (close-up focusing on cable details, mid-range shots showing partial bridge structures, and distant views capturing the overall outline). To address the limitations of on-site photography in terms of visual completeness and structural diversity, online images were collected to supplement the dataset with panoramic views and images of special cable-stayed cable types.

[0068] All captured images were labeled using roLabelImg. The original XML label file generated by roLabelImg is in VOC format, with the main format being (x0, y0, x1, y1, x2, y2, x3, y3). After obtaining the XML label file, it needs to be converted into a DOTA format label file (TXT), and then the DOTA format label file (TXT) is converted into a YOLO format label file (TXT) using the official code.

[0069] Given the limited number of images collected, data augmentation techniques can be used to expand the dataset. Specifically, this can be achieved through image rotation, affine transformation, mirroring, brightness adjustment, and noise addition, effectively increasing the dataset size. The specific steps are as follows:

[0070] (1) Rotation: Rotate the image at a random angle, and fill the blank areas generated after rotation with black pixels.

[0071] (2) Mirroring: A symmetrical version of an image is generated by flipping it horizontally or vertically.

[0072] (3) Affine: After rotating the image at a random angle, scaling is performed and translation is carried out.

[0073] (4) Brightness adjustment: By randomly adjusting the image brightness parameters, the visual effect of the cable-stayed bridge under different lighting conditions is simulated.

[0074] (5) Add noise: Randomly inject Gaussian noise into the image to enhance the diversity of the data.

[0075] Steps 1 and 2: Design the C3K2_DSCA module, introducing dynamic serpentine convolution and axial cross attention to enhance the ability to capture cable features;

[0076] Stay cables are a type of slender target that often spans the entire interface. For this type of target, a targeted improvement to the C3K2_DSCA module is proposed. The improvements made to the C3K2_DSCA module mainly consist of two parts: introducing dynamic serpentine convolution and adding axial cross attention.

[0077] The introduction of dynamic serpentine convolution and the addition of axial cross attention are both reflected in the basic DSCA module. The internal structure of the DSCA module is as follows: Figure 3 As shown on the left, after image input, it passes through three core branches: a regular convolutional block, a dynamic serpentine convolutional block for the X-axis, and a dynamic serpentine convolutional block for the Y-axis. The regular convolutional block branch extracts general feature maps with basic semantics. The dynamic serpentine convolutional blocks for the X-axis and Y-axis focus on the horizontal edge line structure features along the x-axis and the vertical contour gradient features along the y-axis through depthwise separable convolution, respectively, to specifically capture the axial geometric features of slender targets. The extracted x-axis and y-axis features are then processed by axial feature recombination and a cross-attention module, and concatenated with the basic features along the channel dimension using the channel concatenation function cat to form a composite feature map that integrates multi-directional spatial information and basic semantics. This retains general visual features while enhancing the axial structural information of the target, improving the feature representation ability of slender and tilted targets.

[0078] The internal structure of the dynamic serpentine convolution block module is as follows: Figure 3As shown on the right. The input feature map first enters a 3×3 bias convolutional layer, followed by a batch normalization layer to normalize the result. Then, the offset is limited to between -1 and 1 by the tanh activation function, thus mimicking the range of a snake's sway. The formula for the tanh activation function is as follows:

[0079] (1)

[0080] Next, the value of `morph` (0 or 1) determines whether the next convolutional kernel specializes in the x-direction or the y-direction. When `morph` is 0, the kernel expands along the x-direction; when `morph` is 1, the kernel expands along the y-direction. Taking the x-direction as an example, assuming a convolutional kernel K of size 9×1, the specific position of each grid in K is represented as follows: ,in , representing the horizontal distance from the center grid, such as Figure 4 As shown on the left. Convolution kernel. K Each grid position The selection process is cumulative. Starting from the central position... Initially, the position away from the center grid depends on the position of the previous grid: Compared to Increase offset Therefore, the offsets need to be accumulated to ensure that the convolution kernel conforms to a linear morphological structure. Since the offsets are usually decimals, while the coordinates are usually integers, bilinear interpolation is used to calculate the weighted sum of the feature values ​​corresponding to these 8 positions, thus obtaining the deformed convolution kernel. Figure 4 As shown on the right. Through the above operations, key features can be better perceived. Finally, the deformed convolution kernel is used to perform a convolution operation. The convolution result is normalized by group normalization (GN) and then passed through the SiLU activation function to obtain the output of this convolutional layer.

[0081] The outputs of the dynamic serpentine convolutional blocks in the X and Y directions are not directly concatenated with the output of Conv_0. Instead, feature reorganization is performed first, followed by concatenation in the cross-attention module. Finally, the output of the cross-attention module is concatenated with the base features using the cat function. At the beginning of this stage, two multi-head attention modules, self.attn_x and self.attn_y, are defined to process features along the X and Y axes, respectively.

[0082] The outputs of the dynamic serpentine convolutional blocks in the X and Y directions are the input features feat_x and feat_y, respectively, both with shapes [B, C, H, W], where B is the batch size, C is the number of channels, H is the height, and W is the width. To enable computation using a multi-head attention module, the input features need to be reorganized. For feat_x, the channel dimension is first moved to the end, and then the batch size and height dimensions are merged to obtain a sequence with shape [B*H, W, C]. A similar operation is performed on feat_y to obtain a sequence with shape [B*W, H, C].

[0083] Then, a multi-head attention module is used to perform cross-attention calculation. `self.attn_x(x_seq, y_seq, y_seq)` means that the feature sequence `x_seq` of branch X is used as the query, and the feature sequence `y_seq` of branch Y is used as both the key and the value, calculating the output `attn_x` after branch X has focused on information from branch Y. Similarly, `self.attn_y(y_seq, x_seq, x_seq)` means that the feature sequence `y_seq` of branch Y is used as the query, and the feature sequence `x_seq` of branch X is used as both the key and the value, calculating the output `attn_y` after branch Y has focused on information from branch X. Finally, `attn_x` and `attn_y` are reconstructed through feature recombination to restore their original shape [B, C, H, W], and then concatenated with the results of the basic features.

[0084] Step 13: Use the Angleloss loss function to optimize the angle prediction accuracy and establish a refined detection model for inclined and slender targets of bridge stay cables.

[0085] In steps one and three, the loss function used in conventional object detection is based on horizontal bounding box detection and does not consider the impact of angle on training. Therefore, the Angleloss loss function is introduced. The Angleloss loss function is a relatively effective loss function in rotated object detection. It solves the limitations of traditional point estimation methods in angle prediction by transforming angle prediction into distribution learning. Through steps such as positive sample selection, target angle normalization, loss calculation, and special case handling, the Angleloss loss function can guide the model to learn more accurate angle information.

[0086] The Angleloss function is an improved form of the focusing loss function (dfl_loss). When calculating the Angleloss function, the positive sample angles are first predicted, then normalized. Next, the angle distribution focusing loss is calculated based on the selected positive sample angle predictions and the normalized target angles. Internally, this function models and learns the angle distribution according to the principles of the focusing loss function, ensuring that the model's predicted angle distribution is as close as possible to the true angle distribution. The specific implementation of this loss function in angle regression is shown below:

[0087] (2)

[0088] In the formula, N The number of positive samples; S This represents the sum of the weights of the positive samples. For sample weights; and Weights for the left and right intervals of the target angle; Predict the distribution based on the angle; and The left and right interval indices for the target angle; This is for calculating cross-entropy.

[0089] The calculated loss value is multiplied by a weight, which is calculated from the sum of the target scores and is used to weight the loss of different samples to highlight the influence of important samples.

[0090] By treating angle prediction as distribution learning, Angleloss can better capture the uncertainty and range of change in angles, thereby improving the model's accuracy in angle prediction. Even with slight changes in the target angle, the model can respond more sensitively. Because it considers the distribution information of angles, the model is more robust to targets at different angles and can better handle this periodicity, avoiding the problems that traditional point estimation methods may encounter when dealing with angle periodicity.

[0091] After the above steps, the improved YOLOv11-DSCA model structure is as follows: Figure 5 As shown.

[0092] Step 2: Propose a multi-target tracking algorithm that integrates the tilted slender target detection model and the StrongSORT algorithm;

[0093] Step Two Implementation Process as follows Figure 6 As shown, step two specifically includes:

[0094] Step 21: Connect the YOLOv11-DSCA target detection model and add angle parameters;

[0095] In the technical framework of object detection and tracking, the parameter design of the detection box directly affects the compatibility and performance of subsequent tracking algorithms. The YOLOv11-DSCA model, designed for complex scene target characteristics, outputs an Oriented Bounding Box (OBB), with parameters in the form of (x, y, w, h, r), containing 5 degrees of freedom: center coordinates (x, y), width w, height h, and rotation angle r (in radians, describing the direction of the detection box's deflection relative to the horizontal axis). In contrast, the standard YOLOv11 model outputs an Axis-Aligned Bounding Box (AABB) with only 4 degrees of freedom (x, y, w, h), and by default, the detection box boundary is parallel to the image coordinate axes. This difference is crucial in scenarios where the target exhibits significant rotation or pose changes—the rotated detection box more accurately fits the target contour and reduces redundant background interference, while the horizontal detection box may contain a large number of invalid regions due to target rotation, affecting detection accuracy.

[0096] However, the native design of the mainstream multi-object tracking algorithm StrongSORT is based on the geometric characteristics of horizontal bounding boxes. Its core modules (such as data association, state prediction, and trajectory management) all assume that the bounding boxes are axis-aligned, resulting in structural incompatibility with rotating bounding boxes. For example, StrongSORT's Kalman filter model only maintains the 4-dimensional state of the horizontal bounding box, such as its position and size, and cannot model the dynamic changes in the angle of the rotating bounding box, leading to a significant increase in trajectory prediction error when the target rotates.

[0097] To achieve the fusion of the improved YOLOv11 model with StrongSORT, this invention primarily modifies the Kalman filtering and appearance feature extraction parts of the StrongSORT algorithm. The main modification involves adding a new degree of freedom, a rotation angle *r*, to these two parts. Specifically, in the Kalman filtering part, corresponding state transition and observation matrices are designed. The observation equation directly updates the angle state using the rotation parameters output by YOLOv11, enabling the tracker to predict the target's rotation trajectory. In the appearance feature extraction part, the ROI region for feature extraction is rotated by an angle, allowing the extracted feature region to be combined with the target detection bounding box.

[0098] Step 22: Introduce the OSNet network for appearance extraction to enhance appearance extraction capabilities;

[0099] An OSNet network pre-trained on the Market1501 dataset was used to replace the original feature extraction module of StrongSORT. The core design principle of OSNet is a multi-scale feature fusion mechanism, such as... Figure 7As shown, by integrating feature representations from different receptive fields, the problem of appearance differences of targets under varying sizes, viewpoints, and poses can be effectively addressed. This multi-scale feature processing capability is particularly suitable for small target tracking tasks and can significantly improve feature robustness during cross-frame association.

[0100] Furthermore, OSNet balances lightweight design with computational efficiency in its network architecture design. By optimizing channel dimensions and convolution operation configurations, it significantly reduces the time overhead of feature extraction while maintaining high-performance feature representation, providing a better engineering solution for real-time target tracking systems.

[0101] Steps 2 and 3: Introduce the ProbIOU metric to adapt to matching in rotation scenarios.

[0102] StrongSORT's original IoU calculation formula is based on area. Assuming there are detection boxes A and B, the IoU is calculated using the following formula:

[0103] (3)

[0104] In the formula, This represents the intersection region of detection box A and detection box B. This is the union region of detection box A and detection box B.

[0105] It is evident that the detection boxes for the cable-stayed bridge are tilted, and the angle and aspect ratio of these tilted boxes change drastically between frames. Traditional IoU only calculates the geometric overlap area, failing to capture the directional differences and shape changes of the tilted detection boxes. When the cable-stayed bridge detection boxes rotate, even if the target is actually tracked continuously, the sudden reduction in overlap area may lead to trajectory breakage or incorrect association. Since traditional IoU calculations are inapplicable to rotated boxes, ProbIoU (Probabilistic IoU) is adopted. ProbIoU is a metric for measuring the similarity of rotated bounding boxes, taking into account the uncertainty of the bounding boxes.

[0106] Traditional IoU (Axis-aligned bounding boxes) only calculates geometric overlap, while ProbIoU, for rotated bounding boxes (formatted as (x,y,w,h,r), where r is the rotation angle), models each bounding box as a two-dimensional Gaussian distribution (assuming the width and height are uniformly distributed, with rotation introducing covariance). It defines similarity by measuring the similarity between two Gaussian distributions (Bhattacharyya distance), thus more accurately describing the degree of matching of rotated boxes.

[0107] Assume two rotated bounding boxes, OBB1 and OBB2, and that the width w and height h of the bounding boxes follow a uniform distribution with variance . and Rotation angle r This introduces a covariance term, calculated as follows:

[0108] (4)

[0109] The Bhattacharyya distance is used to measure the similarity of Gaussian distributions. The formula for calculating the Bhattacharyya distance is as follows:

[0110] (5)

[0111] In the formula, and As the mean parameter of the Gaussian distribution, it comes from the center coordinates of OBB1. The center coordinates of ) and OBB2 ).

[0112] After obtaining the Bhattacharyya distance, a range constraint needs to be added to prevent numerical overflow before calculating ProbIoU. When the two distributions completely overlap... When the correlation coefficient is 0, ProbIoU = 1; when completely uncorrelated, ProbIoU = 0. The formula for calculating ProbIoU is as follows:

[0113] (6)

[0114] ProbIoU models rotating bounding boxes using a Gaussian distribution and leverages Bhattacharyya distance to quantify distribution similarity, addressing the shortcomings of traditional IoU in rotating scenarios. Its core principle is to transform a geometric problem into a probabilistic metric, simultaneously considering differences in position, size, and orientation, making it suitable for tilted target matching tasks in complex scenes.

[0115] Step 3: Improve the displacement extraction method by combining SIFT / ORB feature point matching and sub-pixel thinning techniques;

[0116] Step 3 implementation process is as follows Figure 8 As shown, step three specifically includes:

[0117] Step 31: Construct a stable feature point tracking baseline region and divide the ROI region for feature point extraction in each frame of video;

[0118] In the processing of cable-stayed bridge image sequences by target detection models, significant positional shifts and scale fluctuations in the detection bounding boxes of a single cable in consecutive frames occur due to unavoidable attitude jitter during UAV flight, dynamic vibration displacement of the cable itself, and the inherent randomness of the detection algorithm. The pitch and roll motions of the UAV during flight, influenced by airflow, cause continuous changes in the shooting angle, resulting in alterations in the imaging angle and size of the cable. Furthermore, the vibrations of the cable under wind and vehicle loads cause dynamic drift in its position within the image. Additionally, the detection algorithms struggle to perfectly overlap the output detection bounding boxes due to differences in image illumination and background complexity. These combined factors pose a significant challenge to subsequent feature point tracking.

[0119] To construct a stable reference region for feature point tracking, a systematic processing workflow can be adopted. First, the coordinates of the detection boxes in N consecutive frames are aggregated, summarizing the position and size information of the detected cable-stayed bridge detection boxes in each frame. By calculating the spatial union of all detection boxes, the maximum coverage area for feature point extraction can be determined, ensuring that all possible positions of the cable-stayed bridge in consecutive frames are included in the processing scope. Subsequently, the minimum bounding tilted rectangle of this union region is solved using the rotating caliper algorithm. This algorithm can efficiently traverse boundary points to find the tilted rectangle that can contain all detection boxes with the minimum area. This rectangle not only accurately defines the target region but also preserves the spatial orientation characteristics of the cable-stayed bridge, providing precise geometric constraints for subsequent analysis.

[0120] Using the obtained tilted rectangle as the input feature point matching and tracking algorithm for the ROI region effectively overcomes the feature point tracking mismatch problem caused by detection box drift. This method, through geometric constraints, limits the tracking range to a minimum area encompassing the entire temporal trajectory of the cable-stayed bridge, significantly reducing the impact of background environmental interference factors, such as other bridge components, sky, and water surfaces, on feature point extraction and matching. This provides a stable computational domain for the feature point matching algorithm.

[0121] Step 32: Introduce a sub-pixel corner refinement function and extract sub-pixel level feature point positions through SIFT / ORB feature point matching;

[0122] In the field of computer vision, SIFT (Scale Invariant Feature Transform) is a widely used feature extraction and matching algorithm with advantages such as scale invariance, rotation invariance, and illumination invariance.

[0123] This method constructs a Gaussian difference-of-scale (DoG) space, and the calculation formula is as follows:

[0124] (7)

[0125] (8)

[0126] In the formula, ( x,y () represents the image coordinates. The standard deviation of the Gaussian kernel. k This is the scaling factor between adjacent scales.

[0127] Extrema points in the image are identified at different scales; these extrema points are potential feature points. Then, a three-dimensional quadratic function is fitted to precisely determine the location and scale of the feature points. Low-contrast and edge-response feature points are removed through comparison. Finally, the principal direction is determined by calculating the gradient direction histogram within the neighborhood of the selected feature points. The formula for calculating the principal direction is as follows:

[0128] (9)

[0129] (10)

[0130] In the formula, For gradient magnitude, For the gradient direction, For the convolution of the image with a Gaussian kernel, and The interval between adjacent pixels.

[0131] After obtaining the principal direction, the neighborhood of a feature point is divided into multiple sub-regions based on the principal direction. Then, the required gradient direction histogram is calculated for each sub-region. Finally, all histograms for a feature point are combined into a 128-dimensional feature descriptor. Matching feature points are found by comparing the Euclidean distances between feature descriptors, using nearest neighbor distance matching to improve matching accuracy.

[0132] ORB (Oriented FAST and Rotated BRIEF) is a fast and efficient feature extraction and matching algorithm that combines FAST (Features from Accelerated Segment Test) feature point detection with BRIEF (BinaryRobust Independent Elementary Features) feature descriptors.

[0133] The ORB algorithm first uses the FAST algorithm to quickly detect corner points in the image. These corner points are potential feature points, and the calculation formula is as follows:

[0134] (11)

[0135] In the formula, p k For the first circle k 1 pixel,t This is the grayscale threshold.

[0136] Then, the principal direction is determined by calculating the centroid direction within the neighborhood of the feature point, and a principal direction is assigned to each feature point. The calculation formula is as follows:

[0137] (12)

[0138] (13)

[0139] In the formula, The vector direction from the feature point to the centroid serves as a rotation-invariant reference.

[0140] Next, the BRIEF algorithm is used to generate a binary feature descriptor in the neighborhood of the feature point. This descriptor consists of a series of binary bits. Feature point matching is performed by comparing the Hamming distance of these binary bits. The smaller the Hamming distance, the higher the matching degree.

[0141] Both the SIFT matching tracking algorithm and the ORB matching tracking algorithm have a major drawback in displacement tracking: they can only achieve pixel-level displacement tracking, which is far from sufficient. To obtain sub-pixel-level displacement, it is often necessary to use other functions from the OpenCV library.

[0142] This invention implements this functionality using the `cv2.cornerSubPix` function. The process begins by determining the termination condition for sub-pixel thinning, which is fundamental to the entire process. The termination condition is that the thinning process stops when the number of iterations reaches 30 or the precision reaches 0.001. This setting ensures that the thinning process neither wastes computational resources due to excessive iterations nor falls into an infinite loop due to excessively high precision requirements. The error function calculation formula is as follows:

[0143] (14)

[0144] In the formula, W It is the window area surrounding the corner point; These are window functions (such as Gaussian weights) used to reduce the impact of edge pixels; These are the pixel coordinates within the window. It corresponds to the image intensity; It is a prediction function based on the corner position p.

[0145] Next comes the extraction of feature point coordinates. First, the SIFT or ORB algorithm is used to detect feature points and calculate descriptors. Then, feature point matching is performed to filter out high-quality matching points. The coordinates of these high-quality matching points are then extracted and stored in two separate sequences. These coordinates are the foundational data for subsequent sub-pixel thinning.

[0146] The crucial sub-pixel refinement step then begins. To meet the input requirements of the `cv2.cornerSubPix` function, the extracted feature point coordinates need to be reshaped. Then, the `cv2.cornerSubPix` function is called. This function continuously optimizes the feature point coordinates using an iterative algorithm based on the image's grayscale information within the search window, achieving sub-pixel level refinement. The position update calculation formula is as follows:

[0147] (15)

[0148] In the formula, J It is the Jacobian matrix, containing the partial derivatives of the error function with respect to p; W It is a diagonal weight matrix; r It is the residual vector.

[0149] Finally, the refined coordinates are reshaped into a two-dimensional array. The resulting coordinates are the feature point coordinates after sub-pixel refinement. By calculating the differences between these coordinates in adjacent frames, a more accurate sub-pixel displacement can be obtained. This sub-pixel displacement tracking method based on the cv2.cornerSubPix function utilizes local grayscale information of the image and improves the accuracy of feature point localization through iterative optimization, thus making displacement calculation more accurate.

[0150] Step 33: Optimize the RANSAC algorithm to achieve adaptive scaling and hierarchical screening.

[0151] During the execution of the RANSAC algorithm, the dynamic threshold generation mechanism calculates the reprojection error threshold in real time based on the current frame feature distribution characteristics and historical displacement statistics. This threshold calculation method combines the statistical characteristics of feature point displacement with prior knowledge of historical data, combining the displacement variance with the basic threshold through a formula, enabling the threshold to adapt to the feature matching difficulty under different scenarios. When the drone's shooting perspective changes significantly or the cable-stayed bridge vibrates violently, the threshold automatically increases to accommodate a wider range of displacement changes; while in stable scenarios, the threshold decreases accordingly to improve the sensitivity of outlier identification. The dynamic threshold generation formula is as follows:

[0152] (16)

[0153] In the formula, The dynamic threshold for the current frame. This is the scaling factor (calibrated using historical data). The number of feature points, For the first Displacement of each feature point This is the historical average displacement. Based on the threshold bias.

[0154] The adaptive iteration count calculation module dynamically adjusts the RANSAC iteration count based on the initial matching quality. Traditional RANSAC algorithms typically use a fixed number of iterations, which may lead to insufficient iterations in complex scenarios or computational redundancy in simple scenarios. The method described in this invention estimates the outlier rate of the current frame by analyzing the inlier ratio of the previous frame and dynamically calculates the required number of iterations based on the confidence level. When the matching quality is good, the algorithm automatically reduces the number of iterations to improve computational efficiency; while when the matching difficulty is high, the number of iterations is increased to ensure that the optimal model is found. The adaptive iteration count calculation formula is as follows:

[0155] (17)

[0156] In the formula, This represents the iteration number of the current frame. For confidence level parameters, This is the current out-of-frame rate estimate. Minimum number of samples required for model estimation (8 for basic matrix estimation).

[0157] To further improve the algorithm's adaptability to the geometric characteristics of the stay cables, a geometrically constrained enhanced RANSAC model is introduced. This model first assigns weights to feature points based on their distance from the cable's central axis, ensuring that points closer to the axis have a higher selection probability. In each iteration of RANSAC, random sampling is performed according to these weighted probabilities, prioritizing feature points that conform to the cable's geometric distribution. This sampling strategy effectively improves the model's ability to fit the cable features and reduces the impact of background interference points on the model's estimation.

[0158] (18)

[0159] In the formula, For the first The weights of each feature point For the first The distance from each feature point to the central axis of the cable-stayed bridge. The average distance from all feature points to the central axis. This represents the standard deviation of the distance.

[0160] After multiple iterations of the RANSAC algorithm, a dynamic model evaluation and result fusion mechanism is introduced. This mechanism comprehensively scores the RANSAC results under different parameter configurations, with scoring indicators including multiple dimensions such as inlier ratio, reprojection error, and geometric consistency. A weighted summation method is used to calculate the comprehensive score of each model, and the model parameter combination with the highest score and its corresponding inlier set are selected as the optimal result. This multi-dimensional evaluation method can comprehensively measure the model's performance, avoiding the bias that may arise from a single indicator.

[0161] Based on the filtered set of interior points, the vibration displacement vector of the stay cable is calculated. The overall displacement vector of the stay cable is obtained by averaging the coordinate changes of all interior points between the current frame and the reference frame. To further smooth the displacement trajectory and reduce the influence of measurement noise, Kalman filtering is used to optimize the displacement sequence. Kalman filtering utilizes the system's dynamic model and measurement model to recursively estimate the displacement, effectively suppressing random noise and improving the stability and accuracy of displacement measurement. In this way, a complete process from initial feature point matching to high-precision vibration displacement calculation is realized.

[0162] Step 4: Design a UAV motion correction algorithm based on variational mode decomposition and joint time-frequency domain screening;

[0163] Step four implementation process is as follows Figure 9 As shown, step four specifically includes:

[0164] Step 41: Based on the vibration characteristics of UAVs and cable-stayed bridges, construct a joint time-frequency domain screening method;

[0165] Drone motion primarily occurs at frequencies below 0.5 Hz. Disturbances to the onboard camera while the drone is hovering mainly originate from two sources: fuselage vibration caused by the motors and swaying caused by airflow. These disturbances make the video recorded by the onboard camera unstable. An onboard gimbal camera platform can effectively isolate interference and ensure the stability of the camera's optical axis, resulting in stable video. Considering the damping performance of the gimbal, the impact of motor vibration is negligible and can be considered as ordinary noise. However, the impact of airflow cannot be ignored. When the drone changes position due to wind, the position of the target object in the imaging plane also changes, potentially affecting vibration measurements.

[0166] Like most vibration systems, stay cables produce damped vibrations under external excitation. The decay trend of damped vibrations closely approximates that of an exponential function. Therefore, the ideal damped vibration signal can be represented by a basis function consisting of an exponential and a cosine function, as shown in the following formula:

[0167] (19)

[0168] In time-domain analysis, the most significant characteristic of damped vibration is its amplitude decaying exponentially over time, forming a smooth exponential envelope. This characteristic is determined by the energy dissipation nature of the vibration system. Taking engineering structures such as cable-stayed bridges as an example, their vibration signals continuously lose energy under damping, causing the amplitude envelope to decay exponentially. This decay pattern is not only a direct reflection of damped vibration in theoretical mechanics but also a key indicator of structural dynamics in actual engineering. However, UAV swaying interference is affected by random factors such as airflow and control, resulting in irregular signal amplitude fluctuations. The envelope exhibits non-stationary characteristics and no fixed decay pattern, potentially containing sudden spikes or random oscillations, lacking a clear exponential decay trend. To quantify this difference, each IMF component obtained from the decomposition can be subjected to exponential fitting. By analyzing the tightness of the fitting curve and its decay characteristics, effective signals can be screened—the fitting result of the target vibration signal should closely resemble an exponential curve, while the fitting effect of UAV interference is poor, with an irregular envelope. This allows us to eliminate interference components that do not conform to the exponential characteristics.

[0169] In frequency domain analysis, the fundamental frequency of vibration of long-span cable-stayed bridges is influenced by structural parameters and is typically between 0.5 Hz and 5.0 Hz. In contrast, the low-frequency swaying of UAVs is usually below 0.1 Hz. Even if the energy diffuses to higher frequencies as the swaying amplitude increases, the dominant frequency is still mostly concentrated in the range of 0.1 Hz to 1.0 Hz, and the amplitude distribution is unstable.

[0170] A joint screening mechanism can be constructed by combining time-domain and frequency-domain characteristics. First, spectral analysis is used to completely eliminate IMFs with a dominant frequency below 0.1 Hz, reducing interference from this component, which is clearly low-frequency interference from UAVs. For signals between 0.1 Hz and 0.5 Hz, their energy distribution is further observed, eliminating interference components with dispersed energy and no clear fundamental frequency peak. For the retained IMFs in the 0.1 Hz to 5.0 Hz frequency band, their envelope characteristics are analyzed through exponential fitting, screening out components whose envelopes conform to an exponential decay law, eliminating irregular vibrations caused by noise or nonlinear factors. Finally, the modal characteristics of cable-stayed bridge vibration in engineering practice are used to verify whether the screened IMFs conform to the physically defined damped vibration characteristics, avoiding the misselection of interference components.

[0171] Step 42: Introduce the triangular topology optimization search algorithm to optimize the key parameters of the variational mode decomposition algorithm.

[0172] Variational Mode Decomposition (VMD) is a class of adaptive signal decomposition methods. Its core idea is to decompose complex signals into several intrinsic mode functions (IMFs) with finite bandwidth by constructing a variational model and solving for the optimal solution. This algorithm has advantages such as strong noise resistance, small boundary effects, and the ability to preset the number of decomposed modes, and it is widely used in non-stationary signal processing fields (such as mechanical fault diagnosis and biomedical signal analysis). The expression for IMF is as follows:

[0173] (20)

[0174] (twenty one)

[0175] In the formula, yes amplitude, yes The frequency.

[0176] (1) Construction of variational problems

[0177] The goal of VMD is to transform the original signal It is decomposed into K modal components. Each IMF is then subjected to a Hilbert transform to convert it into an analytic signal, calculated using the following formula:

[0178] (twenty two)

[0179] The sum of all modes equals the original signal. The calculation formula is as follows:

[0180] (twenty three)

[0181] The one-sided spectrum of the analytic signal is ,when When the bandwidth is such that: This formula characterizes the mode. At the center frequency Frequency clustering in the vicinity.

[0182] The above conditions can be used to construct a model for a constrained variational problem:

[0183] (twenty four)

[0184] (2) Solving variational problems

[0185] To handle the constraints, the Lagrange multiplier method and a quadratic penalty term are introduced, transforming the constrained optimization problem into a solution using the unconstrained augmented Lagrange function:

[0186] (25)

[0187] in This is a penalty parameter (controlling the strictness of bandwidth constraints). It is a Lagrange multiplier.

[0188] The solution to the variational problem is now transformed into finding the saddle point in the above formula, achieved by using the alternating direction multiplier algorithm. Therefore, the calculation is simplified by frequency domain transformation (Fourier transform), and then... , and (by respectively) , and By iteratively updating the saddle point obtained through Fourier transform, and by adaptive frequency band decomposition, and by iteratively updating the center frequency and bandwidth of the IMF, the optimal solution to the variational problem is obtained after multiple screenings.

[0189] The TTAO algorithm (Triangulation Topology Aggregation Optimizer) is a metaheuristic optimization algorithm based on the topological structure of mathematically similar triangles. It aims to balance global exploration and local exploitation during the optimization process through geometric properties. Its core principle is to divide the population into multiple triangular topological units, each containing three vertices and one internal random vertex. The search range is adaptively shrunk by dynamically adjusting the triangle side lengths (which decay exponentially with the number of iterations). Figure 10 As shown, the algorithm employs two key strategies: Generic Aggregation and Local Aggregation. The former generates new solutions by exchanging optimal individual information among different triangular units, thus enhancing global exploration; the latter achieves refined local search based on the positional perturbations of the optimal and suboptimal individuals within a unit, effectively avoiding getting trapped in local maxima.

[0190] The TTAO algorithm, after randomly generating the initial population, groups three individuals into a triangular unit. It then generates equilateral triangles using polar coordinate transformation, with the internal vertices obtained by linearly weighting the three vertices to ensure diversity in the initial distribution. The process is as follows:

[0191] (26)

[0192] (27)

[0193] (28)

[0194] (29)

[0195] In the formula, , , and It has four vertices. A random number between [0,1] 1, and For the upper and lower bounds of the variable, t For the number of iterations, T This represents the maximum number of iterations.

[0196] Then, drawing on the crossover concept of genetic algorithms, new solutions are generated by linearly combining the best individuals from different units. The fitness of the new solutions is compared with that of the original best / second-best individuals, and the best solution within the unit is updated to expand the search space coverage. If the new solution is better than the current best or second-best solution, the corresponding vertex is updated to improve population diversity. The search calculation formula is as follows:

[0197] (30)

[0198] In the formula, It is a random number. This is the optimal solution for the current cell. This is the optimal solution for the random unit.

[0199] Then, a perturbation vector is constructed using the position difference between the optimal and second-best individuals. The search step size is dynamically adjusted by the decay parameter α to perform a fine-grained search within the local region, thereby improving the convergence accuracy. The calculation formula is as follows:

[0200] (31)

[0201] (32)

[0202] In the formula, This represents the optimal solution for the current cell, where T is the maximum number of iterations.

[0203] The algorithm iteratively optimizes the position and structure of the triangular units, repeating the above aggregation process until the maximum number of iterations T is reached, and returns the global optimum and its fitness value, gradually approaching the global optimum. It is suitable for high-dimensional continuous optimization problems.

[0204] TTAO's general optimization framework can be used to solve for the penalty factor in VMD. When decomposing the mode number K, both can be used as decision variables to form an optimization vector. Aiming at minimizing the envelope entropy, the envelope signal is extracted by performing a Hilbert transform on the IMF components of the signal, and its Shannon entropy is calculated to measure the sparsity and periodicity of the signal. The smaller the entropy value, the more prominent the signal characteristics (such as fault impact signals) and the weaker the noise interference, thus quantifying the VMD decomposition quality and driving the algorithm towards a "low-entropy, optimal decomposition" direction. Through iterative updates, TTAO can efficiently search... The optimal combination of K and V balances the decomposition accuracy and computational efficiency of VMD, making it suitable for engineering scenarios such as non-stationary signal processing.

[0205] Step 5: Construct a joint working mode analysis algorithm that combines natural excitation techniques with random subspace identification algorithms.

[0206] Step four implementation process is as follows Figure 11 As shown, step five specifically includes:

[0207] Step 51: A two-stage joint analysis of vibration data is conducted using natural excitation technology and a random subspace identification algorithm.

[0208] The Natural Excitation Technique (NEXT) is a method for extracting the impulse response function (IRF) of a system based on environmental vibration response data. This method approximates the environmental excitation as white noise, assuming that its power spectral density is approximately constant within the frequency band of interest.

[0209] Stochastic Subspace Identification - Covariance-Driven (SSI-COV) is a system identification method based on subspace projection. Its core idea is to separate the system's state-space model parameters by constructing and decomposing the covariance matrix of the response data. Compared to traditional methods, the SSI-COV algorithm directly processes the statistical properties of the response data without explicitly calculating the input-output relationship, making it particularly suitable for environmental vibration data.

[0210] This invention designs a classic OMA analysis hybrid framework based on the NExT method and the SSI-COV algorithm for analyzing the vibration modes of stay cables from the IMF.

[0211] First, environmental vibration response data is read, and the frequency domain cross-correlation density between signals is calculated using the NExT method. (This is for a single channel.) ( t It is equivalent to the autocorrelation function, i.e. the impulse response function.

[0212] Next, the impulse responses are arranged into block Hankel matrices and singular value decomposition (SVD) is performed to extract the signal subspace. The calculation formula is as follows:

[0213] (33)

[0214] In the formula, To extract the first part of the Hankel matrix constructed based on IRF. r Columns as extracted signal subspace .

[0215] The structure of the Hankel matrix is ​​as follows:

[0216] (34)

[0217] In the formula, h(k) Lagging k The IRF value at time 1.

[0218] Then, the Feature System Implementation Algorithm (ERA) is used to construct the state matrix from the subspace, calculated as follows:

[0219] (35)

[0220] In the formula, and To extract the signal subspace The upper and lower blocks are separated to obtain the inverse; + represents the pseudo-inverse. A This is the state matrix.

[0221] The state matrix is ​​decomposed into eigenvalues ​​to obtain preliminary modal parameters (frequency, damping ratio, mode shape). The calculation formulas are as follows:

[0222] (36)

[0223] The eigenvalues ​​obtained from formula (36) ,in The frequency can be obtained according to formula (37), and the damping ratio can be obtained according to formula (38). The calculation formulas are as follows:

[0224] (37)

[0225] (38)

[0226] Step 52: Filter and merge similar modalities through hierarchical clustering.

[0227] A multi-level modal parameter screening system based on the physical characteristics of vibration signals and prior engineering knowledge can accurately extract the vibration characteristics of cable-stayed bridges. This mechanism effectively eliminates noise interference modes and retains the true structural vibration modes through physical constraints on damping ratios, frequency grouping and clustering, cross-group feature fusion, and modal correlation assessment, significantly improving the reliability and engineering applicability of modal identification.

[0228] Before performing modal identification, the original modal parameters are first screened based on the dynamic characteristics of the cable-stayed structure. Since the damping ratio of actual cable-stayed structures is usually in the range of 0.001 to 0.15, this step eliminates two types of abnormal modes by setting a damping ratio threshold: measurement noise or UAV jitter interference with too small a damping ratio, and non-structural vibrations with too large a damping ratio that exceed the physical properties of the cable-stayed material.

[0229] To accommodate the frequency domain distribution characteristics of the multi-mode communication in cable-stayed bridges, a strategy of grouping frequencies at 0.3 Hz intervals is adopted, grouping modes with frequency errors within ±0.15 Hz into the same group. This frequency domain clustering method provides a foundation for subsequent damping ratio merging, reducing mode splitting caused by subtle frequency differences. Modes within the same group are then weighted and merged based on damping ratio similarity. When the damping ratio difference is less than a set threshold, modes are merged by a weighted average of frequency and damping ratio, with the weight inversely proportional to the difference, thus effectively eliminating parameter fluctuations caused by random noise.

[0230] To capture potential similar vibrations, cross-group damping ratio merging is implemented for modes in different frequency groups. This mechanism allows for the merging of modes with frequency differences within 0.5 Hz and damping ratio differences less than 3%, to accommodate the coupling characteristics of multi-mode cable-stayed bridges. A weighted average method based on frequency and damping ratio is used during merging to further improve the stability of modal parameters.

[0231] Finally, modal correlation assessment is used to cluster the original modes. When the frequency difference and damping ratio difference are less than their respective thresholds, the modal correlation is further determined. For single-channel data, the correlation between amplitude and phase is calculated; for multi-channel data, modal calibrator (MAC) assessment is used, and the filtered modal parameters are finally sorted in ascending order of frequency. The MAC calculation formula is as follows:

[0232] (39)

[0233] In the formula, This is a threshold condition.

[0234] The present invention also proposes an electronic device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps of the UAV-based method for multi-target recognition, tracking and vibration extraction of bridge cable-stayed bridge video.

[0235] The present invention also proposes a computer-readable storage medium for storing computer instructions, which, when executed by a processor, implement the steps of the UAV-based method for multi-target recognition, tracking, and vibration extraction of bridge cable-stayed bridge videos.

[0236] The memory in this application embodiment can be volatile memory or non-volatile memory, or it can include both volatile and non-volatile memory. The non-volatile memory can be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory can be random access memory (RAM), which is used as an external cache. By way of example, but not limitation, many forms of RAM are available, such as static random access memory (SRAM), dynamic random access memory (DRAM), synchronous dynamic random access memory (SDRAM), double data rate synchronous dynamic random access memory (DDRSDRAM), enhanced synchronous dynamic random access memory (ESDRAM), synchronous linked dynamic random access memory (SLDRAM), and direct rambus RAM (DR RAM). It should be noted that the memory used in the methods described in this invention is intended to include, but is not limited to, these and any other suitable types of memory.

[0237] In the above embodiments, implementation can be achieved, in whole or in part, through software, hardware, firmware, or any combination thereof. When implemented in software, it can be implemented, in whole or in part, as a computer program product. The computer program product includes one or more computer instructions. When the computer instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium accessible to a computer or a data storage device such as a server or data center that integrates one or more available media. The available media may be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., high-density digital video discs (DVDs)), or semiconductor media (e.g., solid-state disks (SSDs)).

[0238] In implementation, each step of the above method can be completed by integrated logic circuits in the processor's hardware or by instructions in software. The steps of the method disclosed in the embodiments of this application can be directly implemented by a hardware processor, or by a combination of hardware and software modules in the processor. The software modules can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are omitted here.

[0239] It should be noted that the processor in the embodiments of this application can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method embodiments can be completed by the integrated logic circuitry in the processor's hardware or by instructions in software form. The processor can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. It can implement or execute the methods, steps, and logic block diagrams disclosed in the embodiments of this application. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the methods disclosed in the embodiments of this application can be directly embodied as execution by a hardware decoding processor, or as a combination of hardware and software modules in the decoding processor. The software modules can be located in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above methods.

[0240] The above provides a detailed description of the UAV-based method for multi-target recognition, tracking, and vibration extraction of bridge cable-stayed bridge videos. Specific examples have been used to illustrate the principles and implementation methods of this invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this invention. Therefore, the content of this specification should not be construed as a limitation of this invention.

Claims

1. A method for multi-target recognition, tracking, and vibration extraction of bridge cable-stayed bridge videos based on UAVs, characterized in that, The method includes: Step 1: Construct a refined model for detecting inclined and slender targets in bridge stay cables based on the YOLOv11 model; Step one specifically includes: Step 11: Collect cable stay image data and perform high-precision annotation. Expand the dataset through data augmentation and construct a YOLO format dataset for detecting inclined and slender targets of bridge cable stays. Steps 1 and 2: Design the C3K2_DSCA module, introducing dynamic serpentine convolution and axial cross attention to enhance the ability to capture cable features; Step 13: Optimize the angle prediction accuracy using the Angleloss loss function and establish a refined detection model for inclined and slender targets in bridge stay cables; Step 2: Propose a multi-target tracking algorithm that integrates the tilted slender target detection model and the StrongSORT algorithm; Step 3: Improve the displacement extraction method by combining SIFT / ORB feature point matching and sub-pixel thinning techniques; Step three specifically includes: Step 31: Construct a stable feature point tracking baseline region and divide the ROI region for feature point extraction in each frame of video; Step 32: Introduce a sub-pixel corner refinement function and extract sub-pixel level feature point positions through SIFT / ORB feature point matching; Step 33: Optimize the RANSAC algorithm to achieve adaptive scaling and hierarchical screening; Step 4: Design a UAV motion correction algorithm based on variational mode decomposition and joint time-frequency domain screening; Step four specifically includes: Step 41: Based on the vibration characteristics of UAVs and cable-stayed bridges, construct a joint time-frequency domain screening method; Step 42: Introduce the triangular topology optimization search algorithm to optimize the key parameters of the variational mode decomposition algorithm; Step 5: Construct a joint working mode analysis algorithm that combines natural excitation techniques with random subspace identification algorithms; Step five specifically includes: Step 51: A two-stage joint analysis of vibration data is conducted using natural excitation technology and a random subspace identification algorithm. Step 52: Filter and merge similar modalities through hierarchical clustering.

2. The method of claim 1, wherein, In steps one and three, when calculating the Angleloss function, the positive sample angles are first predicted, then normalized, and finally the angle distribution focusing loss is calculated based on the selected positive sample angle predictions and the normalized target angles. The specific implementation of this loss function in angle regression is shown below: (2) In the formula, N The number of positive samples; S This represents the sum of the weights of the positive samples. For sample weights; and Weights for the left and right intervals of the target angle; Predict the distribution based on the angle; and The left and right interval indices for the target angle; This is for calculating cross-entropy.

3. The method of claim 2, wherein, Step two specifically includes: Step 21: Connect the YOLOv11-DSCA target detection model and add angle parameters; Step 22: Introduce the OSNet network for appearance extraction to enhance appearance extraction capabilities; Steps 2 and 3: Introduce the ProbIOU metric to adapt to matching in rotation scenarios.

4. The method according to claim 3, characterized in that, During the execution of the RANSAC algorithm, the dynamic threshold generation mechanism calculates the reprojection error threshold in real time based on the current frame feature distribution characteristics and historical displacement statistics; the dynamic threshold generation formula is as follows: (16) In the formula, The dynamic threshold for the current frame. The scaling factor. The number of feature points, For the first Displacement of each feature point This is the historical average displacement. Based on the threshold bias.

5. An electronic device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the method according to any one of claims 1-4.

6. A computer-readable storage medium for storing computer instructions, characterized in that, When the computer instructions are executed by the processor, they implement the steps of the method according to any one of claims 1-4.

Citation Information

Patent Citations

  • Bridge cable sheath surface defect detection method

    CN117670825A

  • Real-time cable force identification method and device for inhaul cable based on deep learning

    CN119942322A