An unmanned aerial vehicle structure displacement measurement method based on visual compound motion decomposition

CN122835252APending Publication Date: 2026-09-29HARBIN INST OF TECH +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611347371.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-09-02
Publication Date
2026-09-29

AI Technical Summary

Technical Problem

这些方法能够在一定程度上改善测量结果,但在复杂室外环境中仍可能存在特征匹配稳定性不足、处理流程较长、参数求解受噪声影响、残余误差难以评估以及测量精度受拍摄条件制约等问题

Benefits of technology

1、本发明通过对无人机连续拍摄的图像序列进行视觉跟踪,获得目标点和两个静止参考点的坐标序列,并采用整像素匹配与亚像素细化相结合的方式提取点位轨迹,有利于降低像素量化误差和弱纹理区域跟踪误差对后续速度求取的影响,为结构位移重建提供稳定的坐标输入;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122835252A_ABST
    Figure CN122835252A_ABST
Patent Text Reader

Abstract

The application discloses a kind of unmanned aerial vehicle structure displacement measurement methods based on visual compound motion decomposition, it is related to unmanned aerial vehicle visual displacement monitoring technical field, specifically includes: the image sequence that unmanned aerial vehicle is shot is carried out visual tracking, obtains the coordinate sequence of target point and two stationary reference points;According to the observation speed of coordinate sequence respectively calculated target point and two stationary reference points, based on the observation speed of stationary reference point determines the translation dependent component caused by unmanned aerial vehicle itself movement and the rotation dependent component around optical axis;According to translation dependent component and the rotation dependent component around optical axis and the position relationship of target point relative to stationary reference point, the dependent speed at target point is calculated;According to the observation speed of target point and the dependent speed at target point, the structural response speed in the measurement plane of target point is calculated, numerical integration is carried out to structural response speed, obtains the structural displacement of target point.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of UAV visual displacement monitoring technology, specifically to a method for measuring UAV structural displacement based on visual composite motion decomposition. Background Technology

[0002] Structural dynamic displacement reflects the dynamic response of engineering structures such as bridges, tall towers, cable-stayed structures, and frame structures under vehicle loads, wind loads, seismic action, or artificial excitation. It is an important observation for structural health monitoring, seismic performance assessment, wind-induced vibration analysis, and harmful vibration identification. Accurately obtaining structural displacement response is of great significance for judging structural service performance, identifying abnormal vibrations, and assessing structural safety status.

[0003] With the development of image sensors and computer vision technology, vision-based non-contact displacement measurement methods are increasingly being applied to structural health monitoring. These methods acquire images of structural motion using cameras and combine them with image processing algorithms such as feature point tracking, digital image correlation, and optical flow to obtain the displacement of structural measurement points. They offer advantages such as being non-contact, flexible in deployment, low in equipment cost, and capable of multi-point simultaneous observation. However, fixed cameras typically require stable installation positions and are limited by shooting distance, field of view, obstruction conditions, and on-site safety requirements, making it difficult to simultaneously achieve both measurement accuracy and observation range in outdoor monitoring of large civil structures.

[0004] Unmanned aerial vehicles (UAVs) equipped with cameras can take close-up photos in locations inaccessible to personnel, providing a new platform for measuring the displacement of large structures. However, compared to fixed cameras, UAVs are more susceptible to airflow, control errors, and body vibrations during hovering or low-speed flight. This causes the apparent motion of structural measurement points in the image sequence to be affected by platform drift, attitude fluctuations, and changes in imaging scale. For structures with small vibrations, these disturbances may be on the same order of magnitude as the actual structural displacement, thus reducing the reliability of the displacement measurement results.

[0005] Current UAV visual displacement measurements typically employ methods such as image registration, background feature correction, geometric transformation compensation, frequency domain filtering, or inertial measurement assistance to mitigate the impact of platform motion. While these methods can improve measurement results to some extent, they may still suffer from problems in complex outdoor environments, including insufficient stability of feature matching, lengthy processing procedures, parameter solving being affected by noise, difficulty in assessing residual errors, and measurement accuracy being constrained by shooting conditions.

[0006] Therefore, a method for measuring the structural displacement of UAVs based on visual composite motion decomposition is needed. Summary of the Invention

[0007] To address the shortcomings of existing technologies, the present invention aims to provide a method for measuring the structural displacement of unmanned aerial vehicles (UAVs) based on visual composite motion decomposition.

[0008] To achieve the above objectives, the present invention provides the following technical solution: A method for measuring UAV structural displacement based on visual composite motion decomposition includes: A sequence of images continuously captured by a camera mounted on a drone is obtained. The image frames in the image sequence include a target point on the structure under test and two stationary reference points. Visual tracking is performed on the image sequence to obtain the coordinate sequence of the target point and the two stationary reference points. The observation velocities of the target point and two stationary reference points are calculated based on the coordinate sequence. The translational entrainment component and the rotational entrainment component around the optical axis caused by the motion of the UAV are determined based on the observation velocity of the stationary reference points. The entrainment velocity at the target point is calculated based on the translational entrainment component, the rotational entrainment component around the optical axis, and the positional relationship of the target point relative to two stationary reference points. Based on the observed velocity and the entrainment velocity at the target point, the structural response velocity of the target point in the measurement plane is calculated. The structural displacement of the target point is obtained by numerical integration of the structural response velocity.

[0009] Furthermore, the visual tracking includes two stages: integer pixel matching and subpixel thinning. In the integer pixel matching stage, zero-mean normalized cross-correlation is used to perform integer pixel matching between the reference frame and the current frame to obtain integer pixel displacement. In the subpixel thinning stage, subpixel thinning is performed in the local neighborhood of the position corresponding to the integer pixel displacement to obtain subpixel residual displacement. The integer pixel displacement and the subpixel residual displacement are added together to obtain the total displacement.

[0010] Furthermore, the sub-pixel thinning includes: calculating the sub-pixel displacement component of a single pixel along the local grayscale gradient direction under the assumption of brightness conservation; when the pixel is in direction or When the magnitude of the gray-level gradient in a direction exceeds the corresponding gradient threshold, the pixel is added to the set of valid pixels. An observation matrix is ​​constructed based on the set of valid pixels, and an adaptive weighted average is performed on the observation matrix to obtain the two-dimensional sub-pixel residual displacement.

[0011] Furthermore, the observation velocities of the target point and the two stationary reference points are calculated based on the coordinate sequences, including: calculating the dynamic scale factor based on the pixel distance change between the two stationary reference points; centering and scaling the coordinate sequences corresponding to the target point and the stationary reference points using the midpoint of the reference baseline as the origin; and taking the time derivative of the scale-corrected coordinate sequences to obtain the observation velocities of the target point and the stationary reference points.

[0012] Furthermore, the time derivative includes: smoothing the scale-corrected coordinate sequence using a Savitzky-Gorye differential filter; determining the observation velocity based on the first derivative output by the filter; and interpolating or removing invalid coordinates before performing the time derivative when invalid coordinates exist in the coordinate sequence.

[0013] Furthermore, the translational and rotational components caused by the UAV's own motion are determined based on the observation velocity of the stationary reference points, including: forming a reference baseline based on two stationary reference points; determining the translational component at the midpoint of the reference baseline based on the average component of the observation velocity of the stationary reference points after scale correction; and determining the rotational component around the optical axis based on the relative observation velocity between the two stationary reference points.

[0014] Further, the calculation of the entrainment velocity at the target point includes: representing the position of the target point with the midpoint of the reference baseline as the origin, and obtaining the positional relationship between the target point and the midpoint of the reference baseline; and calculating the entrainment velocity at the target point based on the translational entrainment component, the rotational entrainment component around the optical axis, and the positional relationship between the target point and the midpoint of the reference baseline.

[0015] Furthermore, after obtaining the structural displacement of the target point, the process also includes signal enhancement of the structural displacement: variational mode decomposition is used to perform baseline cleaning of the structural displacement, removing the modal components that characterize low-frequency drift, and obtaining the displacement signal after baseline removal; singular spectrum analysis is used to perform noise reduction processing on the displacement signal after baseline removal, and obtain the enhanced structural displacement.

[0016] Furthermore, it also includes: determining the marker space layout for visual tracking based on the number of effective pixels participating in subpixel refinement and the geometric relationship between the target point and the stationary reference point; the geometric relationship includes the baseline distance between the stationary reference points and the distance from the target point to the stationary reference point.

[0017] Compared with the prior art, the present invention has the following beneficial effects: 1. This invention obtains the coordinate sequence of the target point and two stationary reference points by visually tracking the image sequence captured by the UAV in a continuous manner, and extracts the point trajectory by combining integer pixel matching and sub-pixel thinning. This helps to reduce the impact of pixel quantization error and tracking error in weak texture areas on subsequent velocity calculation, and provides stable coordinate input for structural displacement reconstruction. 2. This invention calculates the dynamic scale factor based on the pixel distance change between two stationary reference points, and performs centering and scale correction on the coordinate sequence. This can compensate for the imaging scale change caused by the movement of the UAV along the optical axis and reduce the displacement conversion error caused by imaging scale drift. In the velocity domain, the observation velocity of the UAV itself is decomposed using the observation velocity of the two stationary reference points. The influence of the UAV platform's translational drift and rotation around the optical axis on the observation motion of the target point is deducted from the observation velocity of the target point, so that the separation process between the actual structural response and the camera's own disturbance has a clear kinematic correspondence. 3. This invention obtains the structural response velocity by subtracting the entanglement velocity at the target point from the observed velocity at the target point, and calculates the structural displacement at the target point through numerical integration. Compared with the method of directly calculating the displacement based on changes in image coordinates, this invention can reduce the superimposed influence of UAV platform disturbances and pixel scale changes on the structural displacement results, improve the reliability of UAV visual structural displacement measurement results, and further use variational mode decomposition to remove modal components that characterize low-frequency drift, and use singular spectrum analysis to denoise the displacement signal after baseline removal, thereby reducing the influence of residual low-frequency drift and high-frequency random noise on the displacement curve and improving the stability of displacement reconstruction results under complex UAV shooting conditions. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating the method for measuring UAV structural displacement based on visual composite motion decomposition. Figure 2 This is a schematic diagram of the visual tracking process; Figure 3 This is a schematic diagram of the spatial layout of the static reference point and the structural target point; Figure 4 This is a comparison chart of the displacement measurement results of the free vibration experiment of the three-story frame structure in Embodiment 2 of the present invention. Detailed Implementation

[0019] Example 1: Refer to Figure 1 This embodiment provides a method for measuring the structural displacement of a UAV based on visual composite motion decomposition, including: In this embodiment, the target point is an artificial marker or a stable, trackable natural feature point set in the main response region of the measured structure to characterize the dynamic displacement response of the structure. The stationary reference point is a reference marker, a stationary background feature point, a marker on a fixed component, or a light spot projected onto a stationary area, which remains stationary relative to the world coordinate system during the measurement period to provide reference for camera-induced motion observation. To facilitate velocity-domain composite motion decomposition, at least two stationary reference points are selected, denoted as stationary reference point 1 and stationary reference point 2. The target point is denoted as target point o. The line connecting the two stationary reference points forms a reference baseline, which is used to determine the camera translational and rotational components around the optical axis.

[0020] S1, Visual trajectory extraction.

[0021] The drone, equipped with a camera, continuously captures images of the structure under test, resulting in an image sequence arranged according to sampling time. Each frame includes at least a target point o, a stationary reference point 1, and a stationary reference point 2. The target point o is an artificially marked point or a stably trackable natural feature point within the main response area of ​​the structure under test; the stationary reference points 1 and 2 are reference markers, stationary background feature points, markers on fixed components, or light spots projected onto the stationary area that remain stationary relative to the world coordinate system during the measurement period.

[0022] Each frame of data in the image sequence includes the image frame number, sampling time, image grayscale or color matrix, and initial regions of interest (ROIs) for the target point and stationary reference point. The initial ROI is a local image patch surrounding the corresponding point, used to define the search range for visual tracking. During the initial processing, the initial ROI can be manually selected within the reference frame, or it can be automatically detected based on target color, target shape, corner features, speckle features, or grayscale features. In subsequent frame processing, the ROI can be updated using the coordinates of that point output from the previous frame as the center.

[0023] Before the measurement begins, the target point o is positioned within the main response plane of the structure, and stationary reference points 1 and 2 are positioned within the area that remains stationary during the measurement period. Preferably, the two stationary reference points are located within the same measurement plane as the target point. When a depth difference exists between the target point and the stationary reference points, if the scale error caused by the depth difference is less than a preset allowable error, the method of this embodiment is still used for measurement; if the scale error caused by the depth difference exceeds the preset allowable error, the scale is corrected using camera intrinsic parameters, point depth, or plane calibration relationships, or a new stationary reference point that is approximately coplanar with the target point is selected.

[0024] Visual tracking is performed on the target point o, stationary reference point 1, and stationary reference point 2 in the image sequence. (Refer to...) Figure 2 The visual tracking process includes two stages: integer pixel matching and subpixel thinning. The following description uses a currently tracked point as an example. The currently tracked point can be the target point o, a stationary reference point 1, or a stationary reference point 2.

[0025] During the integer pixel matching stage, the reference frame image T(x,y) and the current frame image I(x,y) are read. The reference frame image is the image frame used to establish the matching template and is set as the first frame of the image sequence; when the measurement time is long or the illumination changes significantly, a reliable historical frame is used as the updated reference frame. For candidate integer shifts (m,n), the zero-mean normalized cross-correlation function value is calculated. : in, This represents the pixel coordinates of the current point within the region of interest in the reference frame image. grayscale value at that location This represents the average grayscale value of the reference frame region. Indicates the candidate displacement in the current frame The average gray value of the corresponding area This indicates the degree of similarity between the reference frame and the corresponding region of interest in the current frame.

[0026] Search within the preset integer search range The maximum value is used to determine the integer displacement corresponding to the position of the maximum value as the integer pixel displacement. .

[0027] when If the maximum value is lower than the preset matching threshold, or the difference between the maximum and second-largest values ​​is lower than the preset discrimination threshold, it indicates that the current frame has the risk of occlusion, motion blur, sudden illumination changes, or weak texture mismatch. The search range is expanded for rematching. If the threshold requirement is still not met after the new match, the coordinates of that point in the current frame are marked as invalid, and the coordinates are interpolated using adjacent valid frames or the frame is discarded in subsequent coordinate sequences. If the number of consecutive invalid frames exceeds the preset number, a tracking failure flag is output, and a prompt is given to reselect the region of interest or reacquire the image sequence.

[0028] In the subpixel refinement stage, the residual displacement after integer pixel matching is typically within one pixel. The residual displacement is determined after the integer pixel displacement has been set. Afterwards, the current point still has small displacements not represented by integer pixel displacements. The purpose of subpixel thinning is to estimate the residual displacements and add them to the integer pixel displacements to obtain the total subpixel displacement of the current point. A local neighborhood is established near the integer pixel matching position, and the subpixel residual displacements are estimated based on the brightness conservation assumption and the local gray-level gradient. For pixels within the local neighborhood, the subpixel displacement component along the local gray-level gradient direction can be expressed as: in, This represents the grayscale value of the reference image at pixel (x, y). This represents the grayscale value of the image at the corresponding pixel at the current time t. s(x,y,t) represents the gray-level gradient of the reference image at this pixel, and s(x,y,t) represents the sub-pixel displacement component of this pixel along the gray-level gradient direction.

[0029] This formula is used to extract sub-pixel displacement observations from a single effective pixel. To reduce the impact of weak texture regions and grayscale quantization errors, pixels whose gradient magnitudes meet a threshold condition are selected as effective pixels. When the grayscale gradient magnitude of a pixel in the x or y direction exceeds the corresponding gradient threshold, that pixel is added to the set of effective pixels. The gradient threshold can be determined based on the noise level of the reference frame image, the target texture contrast, or calibration experiments; when the number of effective pixels... When the value is below the preset minimum, the texture of the region of interest is insufficient. The region of interest is then expanded, the template position is adjusted, or the whole pixel displacement is retained as the low confidence result for that frame.

[0030] Constructing the observation matrix from the set of effective pixels ,in The number of effective pixels, Indicates the first Each effective pixel provides sub-pixel displacement observations and their corresponding gradient direction information. Adaptive weights are determined based on the grayscale gradient magnitude, the distance from the pixel to the center of the reference frame image, and the matching residual. The observation matrix M is then weighted by the adaptive weight matrix to obtain... .based on Calculate the two-dimensional subpixel residual displacement If the weighted calculation result exceeds one pixel, the subpixel result of that frame is judged as abnormal, and local matching is re-executed or an integer pixel displacement result is used. The total displacement of this point in the image plane is... .

[0031] In this embodiment of the invention, the image plane is the two-dimensional image coordinate plane corresponding to the camera imaging plane, with coordinates in pixels; the measurement plane is the physical plane where the target point of the measured structure exhibits the main displacement response, with coordinates or displacement in physical length. The two are correlated through camera imaging relationship, dynamic scale factor, and physical scale factor.

[0032] Following the above process, the coordinate sequences of the target point o, stationary reference point 1, and stationary reference point 2 on the image plane at each sampling time are obtained. The coordinate sequence of the target point o is denoted as... The coordinate sequence of the stationary reference point 1 is denoted as The coordinate sequence of the stationary reference point 2 is denoted as Each coordinate record includes a point number, frame number, sampling time, x-coordinate, y-coordinate, matching confidence score, and valid identifier.

[0033] S2, velocity domain composite motion decomposition.

[0034] Specifically, the image plane coordinate sequence of the target point o, stationary reference point 1, and stationary reference point 2 output by S1 is read, and the observed motion of the target point is decomposed into the entrainment motion caused by the camera's own motion and the actual motion of the structure in the velocity domain.

[0035] In this embodiment of the invention, a camera coordinate system and a world coordinate system are established. The camera coordinate system moves with the camera mounted on the UAV, while the world coordinate system is fixed relative to the structure under test and a stationary reference point. For any observation point in the image... The true velocity of the observation point in the world coordinate system Observation speed in camera coordinate system And the speed of motion caused by the camera itself satisfy: in, Point Relative to the true velocity in the world coordinate system, Point The observation speed as reflected in the camera coordinate system or image plane observation. This represents the drag velocity caused by the camera's motion relative to the world coordinate system.

[0036] The observed velocity is obtained directly from the image, but neither the entrainment velocity nor the true velocity can be directly measured, and the equations are underdetermined. Therefore, a stationary reference point is introduced into the measurement scenario to provide direct observation of the entrainment velocity. For the stationary reference point... Since it remains stationary in the world coordinate system, its actual velocity is zero, therefore it satisfies: in, Represents a stationary reference point Observation speed in the camera coordinate system Represents a stationary reference point The speed of the pull caused by the camera's movement.

[0037] The observation velocity of a stationary reference point can be used to provide a direct observation of the entrainment velocity. The observation velocity of the stationary reference point in the camera coordinate system directly corresponds to the negative of the entrainment velocity at that point. Under the assumption of local rigid body motion of the camera, the entrainment velocity field in the entire field of view can be characterized by the observation velocities of a finite number of stationary reference points, and substituted into the basic decomposition relation to solve for the true motion of the target point.

[0038] This embodiment establishes a pixel-domain implementation method for in-plane displacement measurement within the principal response plane of the structure. When the UAV gimbal stabilizes the image and captures near-vertical images, the residual camera motion in the local measurement area mainly manifests as in-plane translation, small-angle rotation around the optical axis, and imaging scale changes caused by motion along the optical axis. The scale drift caused by motion along the optical axis is addressed through a dynamic scale factor. Correction is performed; under the assumption of local rigid body motion of the camera, the velocity difference between any two points satisfies The angular velocity around the optical axis is determined by the difference in observed velocities between two stationary reference points. After the entrainment velocity at the stationary reference points is transferred to the target point, the actual velocity of the target point in the world coordinate system is obtained by subtracting this entrainment velocity from the observed velocity at the target point.

[0039] First, a dynamic scaling factor is defined based on the change in pixel distance between two stationary reference points. The pixel distance metric between two stationary reference points at the reference time is denoted as... The pixel distance between two stationary reference points at the current time t is denoted as , For reference time The pixel distance metric between two stationary reference points calculated in the same manner. Dynamic scale factor. for: in, Used to compensate for changes in imaging scale caused by the movement of the UAV along the camera's optical axis. and Using the same pixel distance metric, ensure It can reflect the scale change relationship between the current frame and the reference frame.

[0040] when Zero Less than the preset minimum baseline distance metric or When the reference baseline of the current frame is outside the preset reasonable range, the calculation of the reference baseline is unreliable, and the baseline of the adjacent valid frame is used. Perform interpolation, or mark the current frame as a scale anomalous frame.

[0041] Before scaling, the pixel coordinates are represented by the midpoint of two stationary reference points at the reference time to avoid artificial translation introduced by scaling operations. The midpoint of the two stationary reference points at the reference time is denoted as... For any point At any moment The original pixel coordinates are Subtract it The centered coordinates are then obtained and are still denoted as... Based on centralization, points The scale-corrected pixel coordinates are as follows ,satisfy .in, Used to represent image plane coordinates under the same scale reference. For target point o, stationary reference point 1, and stationary reference point 2, the following are obtained respectively: , ,and .

[0042] Differentiating the centered pixel coordinates and scale factor yields the scale-corrected velocity components. To reduce the amplification effect of direct finite difference on visual tracking noise, this embodiment employs a Savitsky-Gorye differential filter. The Savitsky-Gorye differential filter smooths and differentiates the coordinate sequence within a local time window, outputting the first-order time derivative corresponding to the sampling time. This first-order time derivative is used as the observed velocity of the target point or stationary reference point. The filter window length can be set to 11 frames, and the polynomial order can be set to 5; in some embodiments, the window length and polynomial order can also be adjusted according to the image frame rate, structural vibration frequency, and noise level.

[0043] By the rule of differential products, point The scale-corrected velocity components are: in, Centered pixel coordinates The time derivative obtained by the Savitzky-Gorye differential filter, Dynamic scaling factor The instantaneous rate of change, This represents the pixel velocity component after scale correction.

[0044] If the coordinates at a certain moment are invalid, the coordinate sequence is interpolated and then differentiated. If the number of consecutive invalid frames exceeds a preset number, the velocity component of that time period is marked as invalid and will not participate in subsequent integration.

[0045] After obtaining the scale-corrected velocity components, the camera translational and rotational entrainment components are extracted based on the velocities of the two stationary reference points. The velocity component common to the two stationary reference points is used to characterize the camera translational entrainment component at the midpoint of the reference baseline. Specifically, the camera translational entrainment components in the x and y directions are as follows: in, and As a stationary reference point 1 at time 1 The scale-corrected velocity components, and As a stationary reference point 2 at time... The scale-corrected velocity components, and This represents the camera translational drag velocity component at the midpoint of the reference baseline.

[0046] The relative velocity component perpendicular to the reference baseline between two stationary reference points is used to characterize the rotational entrainment component about the optical axis. Scalar rotation coefficient. Estimated by the following formula: in, Indicates time The scalar rotation coefficient about the optical axis.

[0047] The target point's location is represented with the midpoint of the reference baseline as the origin. The coordinates of target point o relative to the midpoint of the reference baseline are: in, and This represents the scale-corrected coordinate components of the target point o relative to the midpoint of the reference baseline, used to transfer the rotational entrainment components around the optical axis to the target point.

[0048] Furthermore, read the physical scale factor. . It can be determined by the ratio of the known marker size in the reference frame to its pixel size, or by the actual distance between two stationary reference points and the distance between the corresponding pixels in the reference frame. The unit of k is millimeters per pixel.

[0049] Taking into account the physical scale factor k, the velocity components of the target point o in the measurement plane are: in, and These represent the target point o at time [time]. The observed velocity components after scale correction; and These represent the camera translational entrainment velocity components; and These correspond to the velocity corrections in the x and y directions, respectively, after the rotational motion around the optical axis is transmitted to the target point. and These represent the true velocity components in the x and y directions of the target point o in the corresponding measurement plane of the world coordinate system, respectively.

[0050] The total observation speed of the target point, the camera translation drag speed, and the camera rotation drag speed are all expressed as an explicit combination of pixel observations.

[0051] S3, displacement reconstruction and optional signal enhancement.

[0052] Specifically, the initial displacement of the target point is obtained by numerically integrating the actual velocity component of the target point, i.e., the structural response velocity. The physical displacements of the target point o in the x and y directions are as follows: in, This represents the structural displacement of target point o in the x-direction. This represents the structural displacement of target point o in the y-direction. It is the integral variable.

[0053] Under ideal measurement conditions, the displacement obtained by the above integration can be directly used as the reconstruction result.

[0054] In typical UAV measurement scenarios, small depth differences between markers, slight initial misalignment of camera attitude, subpixel noise in visual tracking, and numerical differentiation and integration errors can introduce low-frequency drift or high-frequency random noise into the initial displacement. To further improve the engineering robustness of the method under these conditions, this method performs signal enhancement processing on the initial displacement signal, including variational mode decomposition (VMD) and singular spectrum analysis (SSA).

[0055] Let the initial displacement signal obtained by numerical integration reconstruction be... .when When low-frequency trend terms exist, variational mode decomposition (VMD) is used for baseline cleaning. VMD will... The modal components are decomposed into several modal components, each with a corresponding center frequency. The decomposition objective is to find a sum of modal components equal to... Under the constraints, the sum of the squares of the bandwidths of each modal analytic signal is minimized.

[0056] In this embodiment, the number of modes in variational mode decomposition and penalty parameters Adaptive search is performed using the whale optimization algorithm. The whale optimization algorithm uses envelope entropy as the fitness objective, which evaluates the complexity of the decomposed modal components. In one implementation, the whale optimization algorithm uses a population size of 10, a maximum number of iterations of 100, a penalty parameter search range of 100 to 3500, and a number of modal decompositions. The search range is 3 to 10. Based on the structural fundamental frequency, spectral peak, or pre-calibrated structural dynamic characteristics, modal components with a center frequency lower than the structural fundamental frequency and exhibiting a slow-changing trend are identified as low-frequency trend modes. These low-frequency trend modes are then removed to obtain the baseline-removed displacement signal. .

[0057] If the fundamental frequency of the structure is unknown, the initial displacement signal can be used. The fundamental frequency of the structure is estimated by the main peak of the spectrum; if multiple modes may contain the true response of the structure, these modes are retained to avoid erroneous deletion of the structural response. If the displacement amplitude decreases abnormally after removing low-frequency trend modes, the number of removed modes is reduced or variational mode decomposition is skipped.

[0058] Furthermore, on Singular spectrum analysis was performed to suppress high-frequency random noise introduced by visual measurements and numerical computations. A Hankel matrix is ​​constructed according to a preset window length. Singular value decomposition is performed on the Hankel matrix, and the number of principal components to be retained is determined based on the singular value energy ratio, the location of singular value abrupt changes, or the range of structural response frequencies. The final reconstructed displacement is expressed as... .in, This represents the final reconstructed displacement after signal enhancement. This represents the singular spectrum analysis reconstruction process. This indicates the number of principal components retained. If the structural dominant frequency is weakened or the displacement waveform becomes excessively smoothed after singular spectrum analysis, then adjust... The value can be set to either `value` or singular spectrum analysis can be disabled.

[0059] Output the final reconstructed displacement. The final reconstructed displacement can be used for structural dynamic response analysis, vibration amplitude identification, frequency identification, or for comparison with measurement results from external reference sensors.

[0060] S4. Uncertainty Propagation and Spatial Layout Criteria.

[0061] In detail, this step is used for optimizing the marker layout before measurement or for evaluating accuracy after measurement, and is not a necessary step for reconstructing the target point displacement from the image sequence. If this step is not performed, the target point displacement can still be obtained through S1 to S3; if this step is performed, the measurement uncertainty is evaluated based on the visual tracking error and the geometric relationship of the point, and the layout of the target point and the stationary reference point is determined.

[0062] Reference Figure 3 Let the baseline distance between two stationary reference points be D, and the distances from the target point o to stationary reference point 1 and stationary reference point 2 be respectively... and The visual tracking errors of feature points are independent and have the same standard deviation. D. and The physical length unit can be used, or the image plane length unit unified by the physical scale factor can be used, and all three use the same unit. It can be obtained from repeated tracking experiments at stationary points, calibration experiments, or statistical analysis of subpixel reconstruction residuals.

[0063] Under the conditions that the target point and the stationary reference point are located on the same measurement plane, the visual tracking errors are independent, and local linearization holds, the measurement uncertainty of the target point displacement after composite motion decomposition is... It can be represented as: in, This indicates the uncertainty in the measurement of the target point displacement. This represents the standard deviation of single-point visual tracking error. This represents the baseline distance between two stationary reference points. and These represent the distances from the target point to the two stationary reference points, respectively.

[0064] Specifically, the triangle formed by the two stationary reference points and the target point satisfies the geometric relationship. , Given a certainty, increase and reduce and This can reduce the geometric magnification factor. This applies when the target point is collinear with two stationary reference points, and the target point is located at the midpoint of the line connecting the two stationary reference points. According to the extreme value relationship of the geometric magnification factor, its square is not less than 1.5. Therefore, the displacement measurement uncertainty satisfies... Its minimum value In practical deployment, it is preferable to place the target point close to the midpoint of the line connecting two stationary reference points, with the two stationary reference points positioned on either side of the target point. When site conditions do not allow for the three points to be collinear, it is preferable to select a longer reference baseline and place the target point as close to the reference baseline as possible to reduce measurement uncertainty.

[0065] Furthermore, combined with the number of effective pixels Based on the uncertainty model described above, an error prediction model is constructed to predict the relationship between the single-point tracking error and the uncertainty: in, This represents the root mean square error of the prediction. This represents the fitting coefficient obtained through calibration experiments. This indicates the number of effective pixels participating in weighted subpixel reconstruction.

[0066] A larger value indicates more texture information available for sub-pixel reconstruction within the region of interest, which helps reduce single-point visual tracking errors. Fit coefficient The model was obtained from calibration experiments using the same camera, the same markers, and under similar imaging conditions. When calibration is not complete, the error prediction model can be used for relative comparisons between different layout schemes, but not as a basis for absolute error output.

[0067] like When the number of effective pixels is below a preset threshold, suggestions may include increasing the marker size, improving marker texture contrast, adjusting the camera shooting distance, or reselecting the region of interest. or If the error exceeds the preset allowable error, an option will be provided to adjust the spacing between stationary reference points. Alternatively, the relative positions between the target point and the stationary reference point can be adjusted, allowing for spatial layout optimization based on the target point, two stationary reference points, and the number of effective pixels before formal image acquisition.

[0068] Example 2: In one specific embodiment, a target point and two stationary reference points are arranged on a three-layer frame structure, and a drone is used to collect video of the structure's motion. The three-layer frame structure consists of aluminum columns and a concentrated mass plate, with the top layer exhibiting a free vibration response after being subjected to horizontal excitation. The target point is arranged near the concentrated mass plate on the top layer to measure the horizontal displacement of the top layer; the two stationary reference points are arranged in areas that remain stationary during the measurement period and are simultaneously located within the field of view of the camera mounted on the drone, along with the target point.

[0069] The drone collects video from a hovering position relative to the target observation area. The video resolution is 3840×2160 pixels, and the frame rate is 50 frames per second. The physical scale factor k is determined by the known marker size in the first frame and is defined as the ratio of the physical size of the structure to the number of pixels occupied by the structure image, in millimeters per pixel. A laser displacement sensor synchronously records the horizontal displacement of the top layer as a reference.

[0070] The acquired video is processed according to steps S1-S3 to reconstruct the displacement of the structural measuring points. Finally, step S4 is executed. , , , as well as Evaluate measurement uncertainty and the rationality of spatial layout.

[0071] Reference Figure 4 In this embodiment, the target point displacement curve reconstructed by the method of the present invention is consistent with the curve measured by the laser displacement sensor. Under free vibration conditions, the root mean square error of the displacement reconstructed by this method is 0.0304 mm. This experimental result demonstrates that the above-described visual composite motion decomposition process can reconstruct the displacement of the structural target point under the disturbance conditions of the UAV platform.

[0072] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method for measuring the structural displacement of an unmanned aerial vehicle (UAV) based on visual composite motion decomposition, characterized in that, include: A sequence of images continuously captured by a camera mounted on a drone is obtained. The image frames in the image sequence include a target point on the structure under test and two stationary reference points. Visual tracking is performed on the image sequence to obtain the coordinate sequence of the target point and the two stationary reference points. The observation velocities of the target point and two stationary reference points are calculated based on the coordinate sequence. The translational entrainment component and the rotational entrainment component around the optical axis caused by the motion of the UAV are determined based on the observation velocity of the stationary reference points. The entrainment velocity at the target point is calculated based on the translational entrainment component, the rotational entrainment component around the optical axis, and the positional relationship of the target point relative to two stationary reference points. Based on the observed velocity and the entrainment velocity at the target point, the structural response velocity of the target point in the measurement plane is calculated. The structural displacement of the target point is obtained by numerical integration of the structural response velocity.

2. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 1, characterized in that, The visual tracking includes two stages: integer pixel matching and sub-pixel thinning. In the integer pixel matching stage, zero-mean normalized cross-correlation is used to perform integer pixel matching between the reference frame and the current frame to obtain integer pixel displacement; In the subpixel thinning stage, subpixel thinning is performed in the local neighborhood of the position corresponding to the integer pixel displacement to obtain the subpixel residual displacement; the integer pixel displacement and the subpixel residual displacement are added together to obtain the total displacement.

3. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 2, characterized in that, The subpixel thinning includes: Under the assumption of brightness conservation, the sub-pixel displacement component of a single pixel is calculated along the local gray-level gradient direction; when the pixel is in direction or When the magnitude of the gray-level gradient in a direction exceeds the corresponding gradient threshold, the pixel is added to the set of valid pixels. An observation matrix is ​​constructed based on the set of valid pixels, and an adaptive weighted average is performed on the observation matrix to obtain the two-dimensional sub-pixel residual displacement.

4. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 1, characterized in that, The observed velocities of the target point and the two stationary reference points are calculated based on the coordinate sequence, including: The dynamic scale factor is calculated based on the pixel distance change between two stationary reference points; the coordinate sequences corresponding to the target point and the stationary reference points are centered and scaled using the midpoint of the reference baseline as the origin; the time derivative of the scaled coordinate sequences is calculated to obtain the observation velocities of the target point and the stationary reference points.

5. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 4, characterized in that, The time derivative includes: The scale-corrected coordinate sequence is smoothed using a Savitsky-Gorye differential filter; the observation velocity is determined based on the first derivative of the filter output; when invalid coordinates exist in the coordinate sequence, the invalid coordinates are interpolated or removed before time differentiation is performed.

6. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 1, characterized in that, The observation velocity based on the stationary reference point determines the translational entrainment component and the rotational entrainment component around the optical axis caused by the UAV's own motion, including: A reference baseline is formed based on two stationary reference points; the translational entrainment component at the midpoint of the reference baseline is determined based on the average component of the observation velocity of the stationary reference points after scale correction; and the rotational entrainment component around the optical axis is determined based on the relative observation velocity between the two stationary reference points.

7. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 6, characterized in that, Calculate the entrainment velocity at the target point, including: The position of the target point is represented by the midpoint of the reference baseline, and the positional relationship between the target point and the midpoint of the reference baseline is obtained. The entrainment velocity at the target point is calculated based on the translational entrainment component, the rotational entrainment component around the optical axis, and the positional relationship between the target point and the midpoint of the reference baseline.

8. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 1, characterized in that, After obtaining the structural displacement at the target point, the process also includes signal amplification of the structural displacement: Variational mode decomposition is used to perform baseline cleaning of the structural displacement, removing modal components that characterize low-frequency drift, and obtaining the displacement signal after baseline removal; Singular spectrum analysis was used to denoise the displacement signal after baseline removal, resulting in the enhanced structural displacement.

9. The method for measuring UAV structural displacement based on visual composite motion decomposition according to claim 2, characterized in that, Also includes: The marker space layout for visual tracking is determined based on the number of effective pixels participating in subpixel refinement and the geometric relationship between the target point and the stationary reference point. The geometric relationships include the baseline distance between stationary reference points and the distance from the target point to the stationary reference points.