A monocular high-speed video measurement method and system based on a moving platform

By using the DALC method and the SWDC camera attitude calibration method on the vibration table, the problems of uneven light measurement and camera attitude changes in dark light conditions are solved, and high-precision monocular high-speed video measurement is achieved.

CN119359761BActive Publication Date: 2025-05-09BEIJING UNIV OF CIVIL ENG & ARCHITECTURE +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411342779.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-25
Publication Date
2025-05-09
Estimated Expiration
2044-09-25

AI Technical Summary

Technical Problem

When monocular measurements are performed under dark light conditions of the vibration platform, it is difficult to achieve uniform light, resulting in uneven imaging quality, and the movement of the mobile platform will cause camera posture changes, affecting the extraction of displacement information.

Method used

The DALC method is used to perform subpixel center extraction, combined with the SWDC camera attitude calibration method, and accurate camera attitude calibration is performed using inter-frame geometric constraints, and 3D information is reconstructed through the camera projection model to obtain displacement.

Benefits of technology

Robust and accurate subpixel center extraction and camera attitude calibration under dark light conditions improve measurement accuracy and accuracy, reduce resource consumption and improve work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119359761B_ABST
    Figure CN119359761B_ABST
Patent Text Reader

Abstract

The present invention discloses a monocular high-speed video measurement method and system based on a moving platform, including collecting a monocular high-frame-rate image sequence of a moving platform, preprocessing the image sequence; performing dynamic image block data processing on the image sequence based on adaptive brightness compensation and LSF to obtain an adjusted image; performing camera posture calibration on the adjusted image based on a sliding window dynamic constraint to obtain a calibrated image; obtaining the three-dimensional coordinates and retrieval displacement of the tracking point according to the calibrated image, and outputting the measurement result according to the three-dimensional coordinates and the retrieval displacement. The present invention supports real-time adaptive compensation, dynamic image block processing and camera posture calibration by optimizing the measurement process, and has important application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of measurement, and in particular to a monocular high-speed video measurement method and system based on a moving platform. Background Art

[0002] With the continuous advancement of global urbanization, public infrastructure has occupied an indispensable position in people's daily life. However, daily use and sudden disasters may cause cracks and fall off in these structures, seriously affecting the operational safety of buildings. Therefore, monitoring the seismic performance of public infrastructure structures can avoid unnecessary casualties and property losses. In the initial design stage of infrastructure, scaled model vibration tests are performed using a shaking table to study the dynamic characteristics and response of the structure under earthquakes and other dynamic loads.

[0003] When performing monocular measurements in low light conditions on a vibration table, compensating for illumination is critical to achieving accurate measurements in the large field of view of a high-speed camera. However, it is difficult to achieve uniform illumination across the entire field of view. Using fill lights to partially illuminate artificial targets can result in uneven illumination and poor imaging of certain targets. Dark light areas and limited fill light range can create areas of insufficient illumination, while sunlight or strong light can result in overexposed areas. These issues can blur the edges of target points and reduce the accuracy of geometric center fitting.

[0004] In addition, affected by wind, gravity and engine power, the mobile platform remains in motion during the measurement process, which will cause the attitude of the high-speed camera to change during the measurement process. When the camera moves, the target displacement includes the false displacement caused by the camera movement, which significantly affects the extraction of displacement information. In order to eliminate the influence of the mobile platform movement, accurate camera attitude calibration is required for each frame.

[0005] In order to achieve more flexible and high-precision dynamic monitoring of the vibration state of large-span structures, a monocular high-speed video measurement method and system based on a moving platform are proposed. Firstly, the DALC method is proposed to achieve robust and accurate sub-pixel center extraction for dark light and high exposure targets. Secondly, a SWDC camera posture calibration method is proposed to make full use of inter-frame geometric constraints to obtain the accurate camera posture of each frame. Finally, the 3D information is reconstructed based on the camera projection model and the displacement is obtained. Summary of the invention

[0006] The purpose of the present invention is to provide a monocular high-speed video measurement method and system based on a moving platform.

[0007] To achieve the above object, the present invention is implemented according to the following technical solutions:

[0008] The present invention comprises the following steps:

[0009] Collecting a high-frame-rate image sequence of a moving platform monocular, and preprocessing the image sequence;

[0010] The image sequence is subjected to dynamic image block data processing based on adaptive brightness compensation and LSF to obtain an adjusted image; the dynamic image block data processing sequentially comprises four steps: adaptive compensation of target points in the image block, Gaussian blur and adaptive threshold binarization brightness compensation post-processing, sub-pixel center positioning of the target point based on LSF, and updating the image block to obtain the target point center of the image sequence;

[0011] The camera attitude calibration is performed on the adjusted image based on the sliding window dynamic constraint to obtain a calibrated image; the camera attitude calibration is sequentially composed of three steps: initial camera attitude estimation based on EPnP, minimization of reprojection error optimization, and adjustment between sliding window sequences;

[0012] The three-dimensional coordinates and the retrieval displacement of the tracking point are acquired according to the calibration image, and the measurement result is output according to the three-dimensional coordinates and the retrieval displacement.

[0013] Furthermore, the method for adaptively compensating a target point in an image block includes:

[0014] The approximate center pixel coordinates of the tracking point in the first frame of the image sequence are I(x,y), and the extended range pixel n is used to determine the range of the image block [(x-(n+m),x+(n+m)),(y-(n+m),y+(n+m))];

[0015] Set the search radius to m, and set the range of the image block as the search area for the subsequent frame matching image block;

[0016] According to the characteristics of the target in the image block, the target points can be divided into standard targets, dark targets and overexposed targets, where standard targets are marks collected under normal lighting conditions, dark targets are marks collected under dark conditions, and overexposed targets are marks collected under sufficient lighting conditions;

[0017] Based on comprehensive analysis, the empirical grayscale thresholds of three types of target points were determined: the average grayscale of standard targets was between 80 and 200, the average grayscale of dark targets was lower than 80, and the average grayscale of overexposed targets was higher than 200;

[0018] Based on single-scale Retinex, the brightness and contrast of dark targets are enhanced, and the original image is decomposed into an illumination image and a reflection image. The expression is:

[0019] S(x,y)=R(x,y)·L(x,y)

[0020] The pixel value of the original image is S(x,y), the reflectivity is R(x,y), and the illumination is L(x,y). The enhanced reflectivity is calculated as:

[0021] R enhanced (x,y)=log(1+R(x,y))

[0022] The enhanced reflectivity is R enhanced , the enhanced image is obtained according to the enhanced reflectivity and illumination:

[0023] S enhanced (x,y)=R enhanced (x,y)·L(x,y)

[0024] The enhanced image is S enhanced (x,y);

[0025] For overexposed targets, the input image is divided into M×N sub-images of equal size, and the height by width of the entire image is H×W, where the number of sub-images is

[0026] Calculate the grayscale histogram of the sub-image, use grayscale stretching to obtain the contrast limit threshold, if the frequency of the grayscale level exceeds, then cut the grayscale level and redistribute the excess part evenly to other grayscale levels, the expression is:

[0027]

[0028] The contrast limit threshold is λ, the number of pixels at the kth gray level is H(k), and the number of updated pixels at the kth gray level is

[0029] Redistribute the extra part, the expression is:

[0030]

[0031] The additional part ΔH is used to calculate the additional frequency:

[0032]

[0033] The additional frequency of the kth gray level is H * (k), the upper limit of the grayscale range is L, and the cumulative distribution of the truncated histogram is calculated:

[0034]

[0035] The cumulative distribution of the kth gray level is CDF(k), and the additional frequency of the ith gray level is H * (i), the upper limit of the gray level is k;

[0036] Normalized cumulative distribution, expressed as:

[0037]

[0038] The smallest non-zero value of the CDF is min , the height of the sub-image is M, the width of the sub-image is N, and the grayscale value of the original sub-image is mapped to the enhanced grayscale value according to the normalized cumulative distribution:

[0039]

[0040] The pixel value of the original sub-image is S tile (x,y), the pixel value of the enhanced sub-image is

[0041] For the boundary area between adjacent sub-images, bilinear interpolation is used to achieve smooth transition and obtain the initial compensated image.

[0042] Furthermore, the Gaussian blur and adaptive threshold binarization brightness compensation post-processing method includes:

[0043] Use the Gaussian function as the convolution kernel, perform weighted averaging on the pixels in the initial compensated image, and calculate the two-dimensional Gaussian value:

[0044]

[0045] The coordinates relative to the center of the convolution kernel are (x, y), the two-dimensional Gaussian function of the coordinates is G(x, y), and the standard deviation of the Gaussian distribution is σ;

[0046] The convolution kernel is obtained by discretizing the Gaussian function, and the image is convolved with the convolution kernel;

[0047] Adopting the adaptive threshold method based on neighborhood mean, the threshold of each pixel is determined by averaging the grayscale values ​​in the local area of ​​the image block, and the compensated image is output.

[0048] Furthermore, the method for locating the sub-pixel center of a target point based on LSF and updating an image block and obtaining the target point of an image sequence includes:

[0049] The Canny operator is used to identify the edge points of the circle in the compensated image, and the center of the ellipse is fitted using LSF to achieve sub-pixel positioning. The expression of the ellipse equation is:

[0050]

[0051] The center coordinates of the ellipse are (x c ,y c ), the major axis of the ellipse is a, and the minor axis of the ellipse is b;

[0052] The pixel set on the edge of the ellipse is M = [(x1, y1), (x2, y2), ..., (xn ,y n )], calculate the root mean square error:

[0053]

[0054] The initial center coordinate of the ellipse is x st,c and st,c , the initial length of the major axis is The initial length of the minor axis is

[0055] The Levenberg-Marquardt method is used for nonlinear optimization based on the root mean square error to calculate the ellipse center I with sub-pixel accuracy. c (x c ,y c );

[0056] Will I c (x c ,y c ) as the initial center approximate pixel coordinates, repeat the operation, calculate the precise center approximate pixel coordinates of the tracking target in the next frame, and output the compensated image after the precise center approximate pixel coordinates as the adjusted image.

[0057] Furthermore, the method for estimating the initial camera posture based on EPnP includes:

[0058] The EPnP algorithm is used to estimate the initial camera pose of the frame. The projection model expression of the 3D point is:

[0059] p i =K[R / T]P i

[0060] Where K is the camera internal parameter matrix, R is the rotation matrix, T is the translation vector, and the two-dimensional coordinates of the control point in the i-th frame are p i ;

[0061] EPnP uses four virtual control points to represent all 3D points:

[0062]

[0063] in All 3D points of the i-th control point are P i , the weight coefficient of the i-th control point in the j-th frame is a ij , the jth virtual control point is C j , the projection of the control point on the adjusted image plane can be expressed as:

[0064] c i =K[R / T]C i

[0065] The projection of the control point i on the adjusted image plane is c i , the linear equation is obtained by expressing the linear combination of the 3D points as control points, and the expression is:

[0066]

[0067] Convert the linear equation into a linear least squares problem and solve for R and T.

[0068] Further, the method of minimizing the reprojection error optimization and the sliding window sequence adjustment includes:

[0069] Combine the EPnP algorithm with the Levenberg-Marquardt nonlinear optimization algorithm, use the result of the EPnP algorithm as the initial value, and calculate the minimum reprojection error:

[0070]

[0071] The actual two-dimensional coordinates of the i-th control point are p i , the i-th two-dimensional coordinate of the current estimated external parameter projected from the three-dimensional point is

[0072] Gradually adjust the camera pose until the minimization of the reprojection error converges to a minimum;

[0073] Use sliding window adjustment to optimize the camera pose using inter-frame information, and select the sliding window size N based on accuracy and efficiency requirements;

[0074] Within the window, the camera posture information in each pair of frames is used to triangulate each control point in the field of view and calculate its 3D coordinates, and the error function is adjusted. The expression is:

[0075]

[0076] The true 3D coordinates of the i-th control point in the j-th frame are X i,j , the three-dimensional coordinates estimated by triangulation are

[0077] The above error function is minimized using the Levenberg-Marquardt algorithm, the window slides forward one frame, and the beam adjustment optimization is repeated until all frames of the adjustment image are traversed and the calibration image is output.

[0078] Furthermore, the method of obtaining the three-dimensional coordinates of the tracking point and retrieving the displacement according to the calibration image includes:

[0079] The spatial position of the tracking point in the calibration image within the camera field of view is reconstructed based on the camera projection model, and the two-dimensional coordinates (u, v) are converted into the standardized camera coordinates (x c,y c ,z c ), the expression is:

[0080] (x c ,y c ,z c )=s·K -1 (u,v,1)

[0081] The intrinsic parameter matrix of the camera is K, and the relationship between pixel coordinates and real-world coordinates is established based on the scale factor. The expression is:

[0082]

[0083] The scale factor is s, and the actual physical size of the target surface is d. w , the corresponding pixel size on the image plane is d p ;

[0084] Convert normalized camera coordinates to world coordinates:

[0085] (X,Y,Z)=R·(x c ,y c ,z c )+T

[0086] Where R is the rotation matrix and T is the translation vector, then the displacement calculation of the corresponding tracking point in three directions is:

[0087]

[0088] The coordinates of the tracking point in the first frame in the X, Y and Z directions are X1, Y1 and Z1 respectively, and the coordinates of the corresponding point in the nth frame in the X, Y and Z directions are X n , Y n and Z n ;

[0089] The vibration state and maximum amplitude of the target structure under load are determined according to the dynamic displacement, and the three-dimensional coordinates of the tracking point and the retrieved displacement are output.

[0090] In the second aspect, a monocular high-speed video measurement system based on a moving platform comprises:

[0091] Data acquisition module: used to collect high frame rate image sequences of a moving platform monocular and pre-process the image sequences;

[0092] Image block processing module: used for performing dynamic image block data processing on the image sequence based on adaptive brightness compensation and LSF to obtain an adjusted image; the dynamic image block data processing sequentially comprises four steps: adaptive compensation of target points in the image block, Gaussian blur and adaptive threshold binarization brightness compensation post-processing, sub-pixel center positioning of target points based on LSF, and updating the image block and obtaining the target point center of the image sequence;

[0093] The posture calibration module is used to perform camera posture calibration on the adjusted image based on sliding window dynamic constraints to obtain a calibration image; the camera posture calibration is sequentially composed of three steps: initial camera posture estimation based on EPnP, minimization of reprojection error optimization, and adjustment between sliding window sequences;

[0094] The measurement output module is used to obtain the three-dimensional coordinates and the retrieval displacement of the tracking point according to the calibration image, and output the measurement result according to the three-dimensional coordinates and the retrieval position.

[0095] The beneficial effects of the present invention are:

[0096] The present invention is a monocular high-speed video measurement method and system based on a moving platform. Compared with the prior art, the present invention has the following technical effects:

[0097] The present invention significantly improves the precision and accuracy of moving platform monocular high-speed video measurement through a series of key steps, including preprocessing, adaptive brightness compensation, dynamic image block processing, camera posture calibration, three-dimensional coordinate acquisition and displacement retrieval. By optimizing the measurement process, not only resources are greatly saved, but also the overall work efficiency is improved. The method can realize the automated measurement of moving platform monocular high-speed video, support real-time adaptive compensation, dynamic image block processing and camera posture calibration, and has important application value. Its wide applicability enables the technology to meet the needs of moving platform monocular high-speed video measurement of different types and standards, showing strong versatility and practicality. BRIEF DESCRIPTION OF THE DRAWINGS

[0098] Figure 1 A flow chart of the steps of a monocular high-speed video measurement method and system based on a moving platform of the present invention;

[0099] Figure 2 Lay out schematic diagrams for bridge structures and tracking points;

[0100] Figure 3 Schematic diagram for control point layout;

[0101] Figure 4 This is a before and after comparison of the adaptive brightness compensation results;

[0102] Figure 5 It is a schematic diagram of a sliding window;

[0103] Figure 6 Schematic diagram of camera position;

[0104] Figure 7 It is the displacement curve of tracking point 1;

[0105] Figure 8 It is the displacement curve of tracking point 2;

[0106] Fig. 9 It is the displacement curve diagram of tracking point 3;

[0107] Fig.10 It is the displacement curve diagram of tracking point 4;

[0108] Fig.11 This is a comparison chart between the moving platform monocular high-speed video measurement results and the displacement meter results. DETAILED DESCRIPTION

[0109] The present invention is further described below by means of specific embodiments. The illustrative embodiments and descriptions of the present invention are used to explain the present invention but are not intended to limit the present invention.

[0110] The present invention provides a monocular high-speed video measurement method and system based on a moving platform, comprising the following steps:

[0111] like Figure 1 As shown, in this embodiment, the following steps are included:

[0112] Collecting a high-frame-rate image sequence of a moving platform monocular, and preprocessing the image sequence;

[0113] In the actual evaluation, the second section of the No. 2 bridge of a certain interchange ramp bridge was selected as the engineering background; the curvature radius of urban ramp bridges is usually 40m to 60m. The design curve of this bridge adopts a circular curve with a curvature radius of 50m, a span arrangement of 4×20m, a single-box single-chamber section for the main beam, and a solid reinforced concrete cylindrical pier.

[0114] Artificial target points, as target observation points, can greatly improve the accuracy of video measurement and the speed of target tracking. Therefore, they are often attached to the surface of the object being measured. In this experiment, artificial target points are divided into control points and tracking points. Control points are used to determine the exterior orientation parameters of the two cameras, and tracking points are used to measure the dynamic changes of the observation position on the target. The tracking point consists of a white circle with a black border. There are crosshairs and inner circles above the control point, which are used to measure the three-dimensional coordinates of the control point using a total station. According to the experimental scene and the size of the structural model, the diameter of the artificial circular target is 7cm, and the control point uses a high-precision total station to obtain the three-dimensional coordinates of its center; the schematic diagram of the bridge structure and tracking point layout is as follows Figure 2 The schematic diagram of control point layout is shown in Figure 3As shown; This embodiment aims to explore the dynamic characteristics of the ramp bridge under the action of seismic forces. The key nodes of the ramp bridge are monitored to determine the impact of seismic waves on the structural model. Special attention is paid to the deformation of the connection between the pier and the bridge deck during seismic activity. Target points are placed at the key nodes of the structural model, and the control network is evenly distributed around the shaking table to calibrate the mobile platform of the high-speed video measurement system. In addition, the near-field natural Chi-Chi wave is used as the input waveform, and a specific peak value and duration are set on the shaking table.

[0115] The image sequence is subjected to dynamic image block data processing based on adaptive brightness compensation and LSF to obtain an adjusted image; the dynamic image block data processing sequentially comprises four steps: adaptive compensation of target points in the image block, Gaussian blur and adaptive threshold binarization brightness compensation post-processing, sub-pixel center positioning of the target point based on LSF, and updating the image block to obtain the target point center of the image sequence;

[0116] The camera attitude calibration is performed on the adjusted image based on the sliding window dynamic constraint to obtain a calibrated image; the camera attitude calibration is sequentially composed of three steps: initial camera attitude estimation based on EPnP, minimization of reprojection error optimization, and adjustment between sliding window sequences;

[0117] The three-dimensional coordinates and the retrieval displacement of the tracking point are acquired according to the calibration image, and the measurement result is output according to the three-dimensional coordinates and the retrieval displacement.

[0118] In this embodiment, the method for adaptively compensating a target point in an image block includes:

[0119] The approximate center pixel coordinates of the tracking point in the first frame of the image sequence are I(x,y), and the extended range pixel n is used to determine the range of the image block [(x-(n+m),x+(n+m)),(y-(n+m),y+(n+m))];

[0120] Set the search radius to m, and set the range of the image block as the search area for the subsequent frame matching image block;

[0121] According to the characteristics of the target in the image block, the target points can be divided into standard targets, dark targets and overexposed targets, where standard targets are marks collected under normal lighting conditions, dark targets are marks collected under dark conditions, and overexposed targets are marks collected under sufficient lighting conditions;

[0122] Based on comprehensive analysis, the empirical grayscale thresholds of three types of target points were determined: the average grayscale of standard targets was between 80 and 200, the average grayscale of dark targets was lower than 80, and the average grayscale of overexposed targets was higher than 200;

[0123] Based on single-scale Retinex, the brightness and contrast of dark targets are enhanced, and the original image is decomposed into an illumination image and a reflection image. The expression is:

[0124] S(x,y)=R(x,y)·L(x,y)

[0125] The pixel value of the original image is S(x,y), the reflectivity is R(x,y), and the illumination is L(x,y). The enhanced reflectivity is calculated as:

[0126] R enhanced (x,y)=log(1+R(x,y))

[0127] The enhanced reflectivity is R enhanced , the enhanced image is obtained according to the enhanced reflectivity and illumination:

[0128] S enhanced (x,y)=R enhanced (x,y)·L(x,y)

[0129] The enhanced image is S enhanced (x,y);

[0130] For overexposed targets, the input image is divided into M×N sub-images of equal size, and the height by width of the entire image is H×W, where the number of sub-images is

[0131] Calculate the grayscale histogram of the sub-image, use grayscale stretching to obtain the contrast limit threshold, if the frequency of the grayscale level exceeds, then cut the grayscale level and redistribute the excess part evenly to other grayscale levels, the expression is:

[0132]

[0133] The contrast limit threshold is λ, the number of pixels at the kth gray level is H(k), and the number of updated pixels at the kth gray level is

[0134] Redistribute the extra part, the expression is:

[0135]

[0136] The additional part ΔH is used to calculate the additional frequency:

[0137]

[0138] The additional frequency of the kth gray level is H * (k), the upper limit of the grayscale range is L, and the cumulative distribution of the truncated histogram is calculated:

[0139]

[0140] The cumulative distribution of the kth gray level is CDF(k), and the additional frequency of the ith gray level is H * (i), the upper limit of the gray level is k;

[0141] Normalized cumulative distribution, expressed as:

[0142]

[0143] The smallest non-zero value of the CDF is min , the height of the sub-image is M, the width of the sub-image is N, and the grayscale value of the original sub-image is mapped to the enhanced grayscale value according to the normalized cumulative distribution:

[0144]

[0145] The pixel value of the original sub-image is S tile (x,y), the pixel value of the enhanced sub-image is

[0146] For the boundary area between adjacent sub-images, bilinear interpolation is used to achieve smooth transition and obtain the initial compensated image.

[0147] Figure 4 The adaptive brightness compensation results of dark target points and overexposed target points in the image block in this embodiment are shown.

[0148] In this embodiment, the Gaussian blur and adaptive threshold binarization brightness compensation post-processing method includes:

[0149] Use the Gaussian function as the convolution kernel, perform weighted averaging on the pixels in the initial compensated image, and calculate the two-dimensional Gaussian value:

[0150]

[0151] The coordinates relative to the center of the convolution kernel are (x, y), the two-dimensional Gaussian function of the coordinates is G(x, y), and the standard deviation of the Gaussian distribution is σ;

[0152] The convolution kernel is obtained by discretizing the Gaussian function, and then the convolution kernel is used to convolve the image. Then, an adaptive threshold calculation method based on the neighborhood mean is used to determine the threshold of each pixel by averaging the grayscale values ​​in the local area of ​​the image block to achieve local feature binarization processing;

[0153] Adopting the adaptive threshold method based on neighborhood mean, the threshold of each pixel is determined by averaging the grayscale values ​​in the local area of ​​the image block, and the compensated image is output.

[0154] In this embodiment, the method for locating the sub-pixel center of a target point based on LSF, updating an image block, and obtaining the target point of an image sequence includes:

[0155] The Canny operator is used to identify the edge points of the circle in the compensated image, and the center of the ellipse is fitted using LSF to achieve sub-pixel positioning. The expression of the ellipse equation is:

[0156]

[0157] The center coordinates of the ellipse are (x c ,y c ), the major axis of the ellipse is a, and the minor axis of the ellipse is b;

[0158] The pixel set on the edge of the ellipse is M = [(x1, y1), (x2, y2), ..., (x n ,y n )], calculate the root mean square error:

[0159]

[0160] The initial center coordinate of the ellipse is x st,c and st,c , the initial length of the major axis is The initial length of the minor axis is

[0161] The Levenberg-Marquardt method is used for nonlinear optimization based on the root mean square error to calculate the ellipse center I with sub-pixel accuracy. c (x c ,y c );

[0162] Will I c (x c ,y c ) as the initial center approximate pixel coordinates, repeat the operation, calculate the precise center approximate pixel coordinates of the tracking target in the next frame, and output the compensated image after the precise center approximate pixel coordinates as the adjusted image.

[0163] In this embodiment, the method for estimating the initial camera posture based on EPnP includes:

[0164] The EPnP algorithm is used to estimate the initial camera pose of the frame. The projection model expression of the 3D point is:

[0165] p i =K[R / T]P i

[0166] Where K is the camera internal parameter matrix, R is the rotation matrix, T is the translation vector, and the two-dimensional coordinates of the control point in the i-th frame are pi ;

[0167] EPnP uses four virtual control points to represent all 3D points:

[0168]

[0169] in All 3D points of the i-th control point are P i , the weight coefficient of the i-th control point in the j-th frame is a ij , the jth virtual control point is C j , the projection of the control point on the adjusted image plane can be expressed as:

[0170] c i =K[R / T]C i

[0171] The projection of the control point i on the adjusted image plane is c i , the linear equation is obtained by expressing the linear combination of the 3D points as control points, and the expression is:

[0172]

[0173] Convert the linear equation into a linear least squares problem and solve for R and T.

[0174] In this embodiment, the method of minimizing reprojection error optimization and sliding window sequence adjustment includes:

[0175] Combine the EPnP algorithm with the Levenberg-Marquardt nonlinear optimization algorithm, use the result of the EPnP algorithm as the initial value, and calculate the minimum reprojection error:

[0176]

[0177] The actual two-dimensional coordinates of the i-th control point are p i , the i-th two-dimensional coordinate of the current estimated external parameter projected from the three-dimensional point is

[0178] Gradually adjust the camera pose until the minimization of the reprojection error converges to a minimum;

[0179] like Figure 5 As shown, the sliding window adjustment is used to optimize the camera posture using the inter-frame information, and the size of the sliding window N is selected based on the accuracy and efficiency requirements;

[0180] Within the window, the camera posture information in each pair of frames is used to triangulate each control point in the field of view and calculate its 3D coordinates, and the error function is adjusted. The expression is:

[0181]

[0182] The true 3D coordinates of the i-th control point in the j-th frame are X i,j , the three-dimensional coordinates estimated by triangulation are

[0183] The above error function is minimized using the Levenberg-Marquardt algorithm, the window slides forward one frame, and the beam adjustment optimization is repeated until all frames of the adjusted image are traversed, and the calibrated camera posture information is output. In this embodiment, the generated camera position information is as follows: Figure 6 shown.

[0184] In this embodiment, the method of acquiring the three-dimensional coordinates of the tracking point and retrieving the displacement according to the calibration image includes:

[0185] The spatial position of the tracking point in the calibration image within the camera field of view is reconstructed based on the camera projection model, and the two-dimensional coordinates (u, v) are converted into the standardized camera coordinates (x c ,y c ,z c ), the expression is:

[0186] (x c ,y c ,z c )=s·K -1 (u,v,1)

[0187] The intrinsic parameter matrix of the camera is K, and the relationship between pixel coordinates and real-world coordinates is established based on the scale factor. The expression is:

[0188]

[0189] The scale factor is s, and the actual physical size of the target surface is d. w , the corresponding pixel size on the image plane is d p ;

[0190] Convert normalized camera coordinates to world coordinates:

[0191] (X,Y,Z)=R·(x c ,y c ,z c )+T

[0192] Where R is the rotation matrix and T is the translation vector, then the displacement calculation of the corresponding tracking point in three directions is:

[0193]

[0194] The coordinates of the tracking point in the first frame in the X, Y and Z directions are X1, Y1 and Z1 respectively, and the coordinates of the corresponding point in the nth frame in the X, Y and Z directions are X n , Y n and Z n ;

[0195] The vibration state and maximum amplitude of the target structure under load are determined according to the dynamic displacement, and the three-dimensional coordinates of the tracking point and the retrieved displacement are output.

[0196] In this embodiment, the displacement of tracking points 1-4 in the direction of seismic wave input is as follows: Figure 7-Figure 10 At the same time, the displacement meter is arranged at the tracking point 2 and compared with the measurement results of this embodiment, the measurement accuracy of the displacement meter used can reach 0.1mm, which can be used as a reference for the displacement value. The comparison results are shown in Fig.11 shown.

[0197] In the second aspect, a monocular high-speed video measurement system based on a moving platform comprises:

[0198] Data acquisition module: used to collect high frame rate image sequences of a moving platform monocular and pre-process the image sequences;

[0199] Image block processing module: used for performing dynamic image block data processing on the image sequence based on adaptive brightness compensation and LSF to obtain an adjusted image; the dynamic image block data processing sequentially comprises four steps: adaptive compensation of target points in the image block, Gaussian blur and adaptive threshold binarization brightness compensation post-processing, sub-pixel center positioning of target points based on LSF, and updating the image block and obtaining the target point center of the image sequence;

[0200] The posture calibration module is used to perform camera posture calibration on the adjusted image based on sliding window dynamic constraints to obtain a calibration image; the camera posture calibration is sequentially composed of three steps: initial camera posture estimation based on EPnP, minimization of reprojection error optimization, and adjustment between sliding window sequences;

[0201] The measurement output module is used to obtain the three-dimensional coordinates and the retrieval displacement of the tracking point according to the calibration image, and output the measurement result according to the three-dimensional coordinates and the retrieval position.

[0202] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention should be included in the protection scope of the present invention.

Claims

1. A monocular high-speed video measurement method based on a moving platform, characterized in that: The following steps are involved: Collecting a high-frame-rate image sequence of a monocular on a moving platform, and preprocessing the image sequence; Performing dynamic image block data processing on the image sequence based on adaptive brightness compensation and LSF to obtain an adjusted image; The dynamic image block data processing includes four steps: adaptive compensation of target points in the image block, Gaussian blur and adaptive threshold binarization brightness compensation post-processing, sub-pixel center positioning of the target point based on LSF, and updating the image block and obtaining the target point center of the image sequence; The camera attitude calibration is performed on the adjusted image based on the sliding window dynamic constraint to obtain a calibrated image; the camera attitude calibration is sequentially composed of three steps: initial camera attitude estimation based on EPnP, minimization of reprojection error optimization, and adjustment between sliding window sequences; The three-dimensional coordinates and the retrieval displacement of the tracking point are acquired according to the calibration image, and the measurement result is output according to the three-dimensional coordinates and the retrieval displacement.

2. According to the monocular high-speed video measurement method based on a moving platform as described in claim 1, it is characterized in that: The method for adaptively compensating a target point in an image block comprises: The approximate center pixel coordinates of the tracking point in the first frame of the image sequence are I(x,y), and the extended range pixel n is used to determine the range of the image block [(x-(n+m),x+(n+m)),(y-(n+m),y+(n+m))]; Set the search radius to m, and set the range of the image block as the search area for the subsequent frame matching image block; According to the characteristics of the target in the image block, the target points can be divided into standard targets, dark targets and overexposed targets, where standard targets are marks collected under normal lighting conditions, dark targets are marks collected under dark conditions, and overexposed targets are marks collected under sufficient lighting conditions; Based on comprehensive analysis, the empirical grayscale thresholds of three types of target points were determined: the average grayscale of standard targets was between 80 and 200, the average grayscale of dark targets was lower than 80, and the average grayscale of overexposed targets was higher than 200; Based on single-scale Retinex, the brightness and contrast of dark targets are enhanced, and the original image is decomposed into an illumination image and a reflection image. The expression is: S(x, y) = R(x, y) L(x, y) The pixel value of the original image is S(x, y), the reflectivity is R(x, y), and the illumination is L(x, y). The enhanced reflectivity is calculated as: R enhanced (x,y)=log(1+R(x,y)) The enhanced reflectivity is R enhanced , the enhanced image is obtained according to the enhanced reflectivity and illumination: S enhanced (x,y)=R enhanced (x,y)·L(x,y) The enhanced image is S enhanced (x, y); For overexposed targets, the input image is divided into M×N sub-images of equal size, and the height by width of the entire image is H×W, where the number of sub-images is Calculate the grayscale histogram of the sub-image, use grayscale stretching to obtain the contrast limit threshold, if the frequency of the grayscale level exceeds, then cut the grayscale level and redistribute the excess part evenly to other grayscale levels, the expression is: The contrast limit threshold is λ, the number of pixels at the kth gray level is H(k), and the number of updated pixels at the kth gray level is Redistribute the extra part, the expression is: The additional part ΔH is used to calculate the additional frequency: The additional frequency of the kth gray level is H*(k), the upper limit of the gray range is L, and the cumulative distribution of the truncated histogram is calculated: The cumulative distribution of the kth gray level is CDF(k), and the additional frequency of the ith gray level is H * (i), the upper limit of the gray level is k; Normalized cumulative distribution, expressed as: The smallest non-zero value of the CDF is min , the height of the sub-image is M, the width of the sub-image is N, and the grayscale value of the original sub-image is mapped to the enhanced grayscale value according to the normalized cumulative distribution: The pixel value of the original sub-image is S tile (x,y), the pixel value of the enhanced sub-image is For the boundary area between adjacent sub-images, bilinear interpolation is used to achieve smooth transition and obtain the initial compensated image.

3. According to the monocular high-speed video measurement method based on a moving platform as described in claim 1, it is characterized in that: The Gaussian blur and adaptive threshold binarization brightness compensation post-processing method comprises: Use the Gaussian function as the convolution kernel, perform weighted averaging on the pixels in the initial compensated image, and calculate the two-dimensional Gaussian value: The coordinates relative to the center of the convolution kernel are (x, y), the two-dimensional Gaussian function of the coordinates is G(x, y), and the standard deviation of the Gaussian distribution is σ; The convolution kernel is obtained by discretizing the Gaussian function, and the image is convolved with the convolution kernel; Adopting the adaptive threshold method based on neighborhood mean, the threshold of each pixel is determined by averaging the grayscale values ​​in the local area of ​​the image block, and the compensated image is output.

4. According to the monocular high-speed video measurement method based on a moving platform as claimed in claim 1, it is characterized in that: The method for locating the sub-pixel center of a target point based on LSF, updating an image block, and obtaining the target point of an image sequence comprises: The Canny operator is used to identify the edge points of the circle in the compensated image, and the center of the ellipse is fitted using LSF to achieve sub-pixel positioning. The expression of the ellipse equation is: The center coordinates of the ellipse are (x c ,y c ), the major axis of the ellipse is a, and the minor axis of the ellipse is b; The pixel set on the edge of the ellipse is M = [(x1, y1), (x2, y2), ..., (x n ,y n )], calculate the root mean square error: The initial center coordinate of the ellipse is x st,c and st,c , the initial length of the major axis is The initial length of the minor axis is The Levenberg-Marquardt method is used for nonlinear optimization based on the root mean square error to calculate the ellipse center I with sub-pixel accuracy. c (x c ,y c ); Will I c (x c ,y c ) as the initial center approximate pixel coordinates, repeat the operation, calculate the precise center approximate pixel coordinates of the tracking target in the next frame, and output the compensated image after the precise center approximate pixel coordinates as the adjusted image.

5. The monocular high-speed video measurement method based on a moving platform according to claim 1, characterized in that: The method for estimating an initial camera posture based on EPnP comprises: The EPnP algorithm is used to estimate the initial camera pose of the frame. The projection model expression of the 3D point is: p i =K[R / T]P i Where K is the camera internal parameter matrix, R is the rotation matrix, T is the translation vector, and the two-dimensional coordinates of the control point in the i-th frame are p i ; EPnP uses four virtual control points to represent all 3D points: in All 3D points of the i-th control point are Pi, and the weight coefficient of the i-th control point in the j-th frame is a ij , the jth virtual control point is C j , the projection of the control point on the adjusted image plane can be expressed as: c i =K[R / T]C i The projection of the control point i on the adjusted image plane is c i , the linear equation is obtained by expressing the linear combination of the 3D points as control points, and the expression is: Convert the linear equation into a linear least squares problem and solve for R and T.

6. The monocular high-speed video measurement method based on a moving platform according to claim 1, characterized in that: Methods for minimizing reprojection error optimization and sliding window inter-sequence adjustment include: Combine the EPnP algorithm with the Levenberg-Marquardt nonlinear optimization algorithm, use the result of the EPnP algorithm as the initial value, and calculate the minimum reprojection error: The actual two-dimensional coordinates of the i-th control point are p i , the i-th two-dimensional coordinate of the current estimated external parameter projected from the three-dimensional point is Gradually adjust the camera pose until the minimization of the reprojection error converges to a minimum; Use sliding window adjustment to optimize the camera pose using inter-frame information, and select the sliding window size N based on accuracy and efficiency requirements; Within the window, the camera posture information in each pair of frames is used to triangulate each control point in the field of view and calculate its 3D coordinates, and the error function is adjusted. The expression is: The true 3D coordinates of the i-th control point in the j-th frame are X i,j , the three-dimensional coordinates estimated by triangulation are The above error function is minimized using the Levenberg-Marquardt algorithm, the window slides forward one frame, and the beam adjustment optimization is repeated until all frames of the adjustment image are traversed and the calibration image is output.

7. The monocular high-speed video measurement method based on a moving platform according to claim 1, characterized in that: The method of obtaining the three-dimensional coordinates of the tracking point and retrieving the displacement according to the calibration image comprises: The spatial position of the tracking point in the calibration image within the camera field of view is reconstructed based on the camera projection model, and the two-dimensional coordinates (u, v) are converted to the standardized camera coordinates (x c ,y c , z c ), the expression is: (x c ,y c ,z c )=s·K- 1 ·(u, v, 1) The intrinsic parameter matrix of the camera is K, and the relationship between pixel coordinates and real-world coordinates is established based on the scale factor. The expression is: The scale factor is s, and the actual physical size of the target surface is d. w , the corresponding pixel size on the image plane is dp; Convert normalized camera coordinates to world coordinates: (X,Y, Z)=R·(x c ,y c ,z c )+T Where R is the rotation matrix and T is the translation vector, then the displacement calculation of the corresponding tracking point in three directions is: The coordinates of the tracking point in the first frame in the X, Y and Z directions are X1, Y1 and Z1 respectively, and the coordinates of the corresponding point in the nth frame in the X, Y and Z directions are X n , Y n and Z n ; The vibration state and maximum amplitude of the target structure under load are determined according to the dynamic displacement, and the three-dimensional coordinates of the tracking point and the retrieved displacement are output.

8. A monocular high-speed video measurement system based on a moving platform, characterized in that: include: Data acquisition module: used to collect high frame rate image sequences of a moving platform monocular and pre-process the image sequences; An image block processing module: used for performing dynamic image block data processing on the image sequence based on adaptive brightness compensation and LSF to obtain an adjusted image; The dynamic image block data processing includes four steps: adaptive compensation of target points in the image block, Gaussian blur and adaptive threshold binarization brightness compensation post-processing, sub-pixel center positioning of the target point based on LSF, and updating the image block and obtaining the target point center of the image sequence; The posture calibration module is used to perform camera posture calibration on the adjusted image based on sliding window dynamic constraints to obtain a calibration image; the camera posture calibration is sequentially composed of three steps: initial camera posture estimation based on EPnP, minimization of reprojection error optimization, and adjustment between sliding window sequences; The measurement output module is used to obtain the three-dimensional coordinates and the retrieval displacement of the tracking point according to the calibration image, and output the measurement result according to the three-dimensional coordinates and the retrieval position.

Citation Information

Patent Citations

  • High-speed video measurement method for progressive collapse of latticed shell structure

    CN107421509A

  • Non-contact RTK (Real-Time Kinematic) acquisition and measurement method, system and equipment

    CN116755123A