A phase correlation matching method for inter-frame micro-shift high-speed video measurement
By employing adaptive window neighborhood sub-image extraction, gradient real and virtual dual-channel Gaussian filtering, and robust estimation, the problem of high-precision tracking of minute inter-frame offsets in high-speed video measurement is solved, thereby improving target recognition and tracking accuracy in complex environments.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-21
- Publication Date
- 2026-03-31
AI Technical Summary
Existing high-speed video measurement methods lack effective measures to deal with aliasing noise and other interference, resulting in low tracking accuracy for small inter-frame offsets. In particular, it is difficult to achieve high-precision target recognition and tracking in situations with large changes in lighting and complex backgrounds.
An adaptive window neighborhood sub-image extraction, real and imaginary dual-channel Gaussian filtering of neighborhood sub-image gradients, and robust estimation and iterative optimization method are adopted. The image sequence is dynamically adjusted by adaptive window to obtain neighborhood sub-images, and real and imaginary dual-channel Gaussian filtering and robust estimation are performed. The final sub-pixel offset value is obtained by iterative optimization.
It achieves accurate identification and tracking of minute inter-frame shifts of target points on the surface of high-speed moving objects under varying lighting conditions and complex backgrounds, improving the detection capability of minute inter-frame shifts and enhancing robustness to noise and lighting changes.
Smart Images

Figure CN121074094B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of high-speed video measurement matching technology, and in particular to a phase correlation matching method for high-speed video measurement with small inter-frame offsets. Background Technology
[0002] High-speed video measurement instantaneously records the spatial position and state of moving objects in the form of video or image sequences. Through photogrammetric analysis, it obtains high-precision three-dimensional spatial coordinates of feature points of the moving object, enabling qualitative and quantitative analysis of the object's characteristics and describing its trajectory and motion parameters. Unlike video measurement using ordinary cameras, high-speed video measurement can achieve frame rates of hundreds or even thousands, thus meeting the demands for video measurement of high-speed moving objects. Sequence image target tracking is one of the core technologies of high-precision high-speed video measurement. High-speed video measurement can acquire massive amounts of image sequence data in a short time, ensuring that the displacement of target points between image sequences does not exceed the pixel level or even reaches the sub-pixel level. Extracting motion information of high-speed moving bodies from this data presents the critical challenge of high-precision tracking of minute shifts in high-frame-rate image sequences.
[0003] Image sequence micro-shift tracking techniques are primarily used to detect and track minute displacement changes of targets within image sequences. These techniques mainly include optical flow, template matching, deep learning, and phase correlation methods. Optical flow tracking estimates the continuous motion of a target by solving optical flow equations (such as the Horn-Schunck or Lucas-Kanade equations) to calculate the motion vectors of pixels in the image sequence. Combining this with deep learning algorithms like FlowNet and PWC-Net can further improve tracking accuracy in complex backgrounds. However, this method is sensitive to image noise and illumination variations, and is only suitable for tracking minute shifts in simple scenes with minimal illumination changes. Template matching tracks targets by matching target templates with regions in the image sequence. This requires careful selection of the search method, range, and threshold. It is suitable for scenes with minimal changes in target shape and texture, but is sensitive to target deformation and rotation, limiting its ability to capture minute shifts. Deep learning utilizes convolutional neural networks (CNNs) or recurrent neural networks (RNNs) to learn the motion characteristics of targets, enabling high-precision tracking of nonlinear motion in complex backgrounds. However, it relies on large amounts of training data, and differences between the training data and the actual scene can lead to insufficient ability to recognize minute shifts. Phase correlation is a frequency-domain-based image registration technique. By introducing peak estimation, registration accuracy can be improved to the sub-pixel level (e.g., 1 / 100 pixel level), significantly enhancing the detection capability of minute shifts. Furthermore, it is insensitive to changes in image grayscale, noise, and illumination, making it suitable for high-speed video measurement on large-field-of-view airborne moving platforms with significant illumination variations. However, existing methods lack effective measures to handle aliasing noise and other interference, relying too heavily on parameter tuning to ensure robustness. Therefore, a phase correlation matching method for high-speed video measurement of minute inter-frame shifts is needed. Summary of the Invention
[0004] The purpose of this invention is to provide a phase correlation matching method for high-speed video measurement with small inter-frame offsets, so as to solve the problems existing in the prior art.
[0005] To achieve the above objectives, the present invention is implemented according to the following technical solution:
[0006] On one hand, the present invention includes the following steps:
[0007] Acquire high frame rate image sequences;
[0008] Neighborhood sub-images are obtained by dynamically adjusting the image sequence using an adaptive window.
[0009] Gradient real and imaginary dual-channel Gaussian filtering is performed on the neighborhood sub-image to obtain the normalized cross power spectrum matrix;
[0010] The normalized cross power spectrum matrix is fitted to a plane, and the initial offset value is obtained by iterative optimization after dynamically adjusting the sampling subset through noise estimation index and interior point probability.
[0011] Based on the initial offset value, the normalized cross-power spectrum is updated to calculate the residual offset value. The process is iterated until the residual offset value is lower than a preset threshold. All iterated offset values are then summed to obtain the final sub-pixel offset value.
[0012] Furthermore, the step of dynamically adjusting the image sequence to obtain neighborhood sub-images through an adaptive window specifically includes: using a SuperPoint feature network to extract the first frame feature point set of the target region of a high-speed moving object; using the feature points as the center, extracting the neighborhood of the feature points based on a reference window; combining Sobel gradient quantization of local features; and using the standard deviation of the gradient values in the target region as a threshold to dynamically adjust the window size. Specifically, high gradient regions with gradient values exceeding the threshold use small windows, while smooth regions with gradient values below the threshold use large windows. This process segments the area into several neighborhood sub-images centered on the feature points, which serve as matching templates.
[0013] Furthermore, to preserve the frequency response of salient features in neighboring sub-images, the image representation uses the gradient complex number G(x). m,n ,y m,n Replace the original image intensity function g(x) m,n ,y m,n ):
[0014] G(x m,n ,y m,n ) = G x (x m,n ,y m,n )+jG y (x m,n ,y m,n (1)
[0015] Among them, G x (x m,n ,y m,n ) and G y (x m,n ,y m,n Let x and y represent the gradients in the x and y directions, respectively. m,n and y m,n Indicates image coordinates, m represents the frame number, and n represents the point number;
[0016] Construct a two-dimensional Gaussian filter Q with a normalized cross power spectral matrix for both real and imaginary parts. filtered (u m,n ,v m,n The specific formula is:
[0017]
[0018] Q filtered (u m,n ,v m,n )=Re(Q filtered (u m,n ,v m,n ))+i·Im(Q filtered (u m,n ,v m,n ))(3)
[0019] Where, Q(u) m,n ,v m,n ) represents the normalized cross-power spectrum matrix, F denotes the Fourier transform, * denotes complex conjugate, and u m,n ,v m,n Represents frequency domain coordinates, and The horizontal and vertical offsets to be calculated, Re(Q) filtered (u m,n ,v m,n )) represents filtering the real part, Im(Q filtered (u m,n ,v m,n )) indicates filtering of the imaginary part.
[0020] Furthermore, the robust estimation of the initial offset value specifically includes plane fitting based on the filtered normalized cross power spectrum matrix using an improved HMSS robust estimation algorithm. The improved HMSS robust estimation algorithm dynamically adjusts the size of the sampling subset through noise estimation index and interior point probability, and then obtains the initial offset value through iterative optimization using the LKOS cost function.
[0021] Furthermore, to ensure that each sampled large subset M has a sufficient number of inliers, the size of the subset is dynamically adjusted using noise estimation σ and inlier probability p:
[0022] M=m+[λ·f(σ,p)] (5)
[0023] Where m is the minimum subset size, f(σ,p) is an adjustment function designed based on the noise estimation and the proportion p of interior points, and λ is the adjustment coefficient. If the noise level σ is low and p is high, then f(σ,p) takes a smaller value; if the noise is high, a larger M is chosen to compensate for the error risk. The parameters of the plane model are finally accurately estimated through iterative optimization using the LKOS cost function.
[0024]
[0025] in, This represents the estimated initial offset value, r. (t) This represents the residual iterative model, where t represents the number of iterations. This indicates the termination condition: the average residual of the data sample set from the first two iterations is less than the residual of the k-th sorting. This indicates the phase angle.
[0026] Furthermore, in the iterative offset value summation, a stop condition of 0.01 pixels is used as the preset threshold, and the final offset value is the sum of all iterative offset values:
[0027]
[0028] in, is the final offset value between frames of the image sequence, and s is the number of robust iterations.
[0029] The beneficial effects of this invention are:
[0030] This invention is a phase correlation matching method for high-speed video measurement with small inter-frame offsets. Compared with the prior art, this invention has the following technical advantages:
[0031] This invention leverages the robustness of frequency domain phase correlation matching to changes in illumination and noise. Through three key technologies—adaptive window neighborhood sub-image extraction, real and virtual dual-channel Gaussian filtering of neighborhood sub-image gradients, and robust estimation and iteration—it achieves accurate identification and tracking of minute inter-frame offset sequences of target points on the surface of high-speed moving objects under the influence of illumination, angle, and other factors, thereby ensuring the accuracy of dynamic monitoring. Attached Figure Description
[0032] Figure 1 This is a schematic diagram of the system framework for a high-speed video measurement phase correlation matching method with small inter-frame offsets according to the present invention. Detailed Implementation
[0033] The present invention will be further described below through specific embodiments. The illustrative embodiments and descriptions herein are used to explain the present invention, but are not intended to limit the present invention.
[0034] like Figure 1 As shown, a phase correlation matching method for high-speed video measurement with small inter-frame offsets includes the following steps:
[0035] Acquire high frame rate image sequences;
[0036] Neighborhood sub-images are obtained by dynamically adjusting the image sequence using an adaptive window.
[0037] Gradient real and imaginary dual-channel Gaussian filtering is performed on the neighborhood sub-image to obtain the normalized cross power spectrum matrix;
[0038] The normalized cross power spectrum matrix is fitted to a plane, and the initial offset value is obtained by iterative optimization after dynamically adjusting the sampling subset through noise estimation index and interior point probability.
[0039] Based on the initial offset value, the normalized cross-power spectrum is updated to calculate the residual offset value. The process is iterated until the residual offset value is lower than a preset threshold. All iterated offset values are then summed to obtain the final sub-pixel offset value.
[0040] (1) Phase correlation matching method for high-speed video measurement of small inter-frame offset
[0041] High-speed video measurement can acquire massive amounts of image sequence data in a short time, but to extract motion information of high-speed moving objects, precise and stable tracking of minute offset sequences is required. This stage is a critical step in the high-speed video measurement system of a large field-of-view airborne moving platform, significantly impacting the accuracy and efficiency of subsequent camera pose calibration and 3D coordinate calculation. In actual monitoring, the target on the surface of a high-speed moving object varies in image contrast and texture due to factors such as illumination and angle, leading to low tracking accuracy of minute inter-frame offsets. To address this issue, a phase correlation matching method for high-speed video measurement of minute inter-frame offsets is proposed. This method overcomes key technologies such as adaptive window neighborhood sub-image extraction, real and virtual dual-channel Gaussian filtering of neighborhood sub-image gradients, and robust estimation and iteration of minute offsets, achieving accurate identification and tracking of minute inter-frame offsets in the image sequence. The specific technical methods are as follows:
[0042] 1) Adaptive window neighborhood sub-image extraction: Define P1 = (x 1,1 ,y 1,1 ),...,(x 1,n ,y 1,n Let P1 = (x) be the set of feature points in the target region of a high-speed moving object extracted using the SuperPoint feature network, where n represents the point number. 1,1 ,y 1,1 ),...,(x 1,n ,y 1,n Centered on a target region, an adaptive window adjustment strategy is established to segment several sub-images as matching templates. The neighborhood of feature points is extracted through a baseline window (e.g., 32×32). Local features are quantized using Sobel gradients. The window size is dynamically adjusted using the standard deviation of gradient values in the target region as a threshold: regions exceeding the threshold are high-gradient regions, and a small window (16×16) is used to improve computation speed and suppress interference; regions below the threshold are smooth regions, and a large window (64×64) is used to enhance robustness.
[0043] Adaptive window neighborhood sub-image extraction enhances the robustness of phase correlation matching to noise and illumination changes by establishing an adaptive window adjustment strategy that dynamically adjusts the template size based on the local texture, contrast, or gradient distribution around feature points. It extracts only the necessary high-information regions around the target point as templates, reducing redundant pixels involved in correlation calculations. This maintains matching accuracy while avoiding the computational burden of traditional global search. This adaptability enables the system to adapt to various complex monitoring environments and requirements.
[0044] 2) Real and Imaginary Dual-Channel Gaussian Filtering of Neighborhood Sub-Image Gradients: To preserve the frequency response of significant features of neighborhood sub-images, the image representation adopts the gradient complex number G(x) m,n ,y m,n Replace the original image intensity function g(x) m,n ,y m,n ):
[0045] G(x m,n ,y m,n ) = G x (x m,n ,y m,n )+jG y (x m,n ,y m,n (1)
[0046] Among them, G x (x m,n ,y m,n ) and G y (x m,n ,y m,n Let x and y represent the gradients in the x and y directions, respectively. m,n and y m,n The coordinates are denoted by m, where m represents the frame number and n represents the point number.
[0047] Construct a two-dimensional Gaussian filter Q with a normalized cross power spectral matrix for both real and imaginary parts. filtered (u m,n ,v m,n This addresses the issue of phase entanglement and discontinuity caused by aliasing noise in high-speed camera high-frame-rate shooting, which results in phase angle data smoothing being directly interrupted.
[0048]
[0049] Q filtered (u m,n ,v m,n )=Re(Q filtered (u m,n ,v m,n ))+i·Im(Q filtered (u m,n ,v m,n(3)
[0050] Where, Q(u) m,n ,v m,n ) represents the normalized cross-power spectrum matrix, F denotes the Fourier transform, * denotes complex conjugate, and u m,n ,v m,n Represents frequency domain coordinates, and The horizontal and vertical offsets to be calculated, Re(Q) filtered (u m,n ,v m,n )) represents filtering the real part, Im(Q filtered (u m,n ,v m,n )) indicates filtering of the imaginary part.
[0051] 3) Robust estimation and robust iteration of small offsets: The offset between neighboring sub-images of a sequence is represented in the frequency domain as a linear phase difference in the normalized cross power spectrum.
[0052]
[0053] in, This represents the phase angle. Theoretically, the phase angle... Represents the offset parameter in the Cartesian coordinate system uv. and Defined 2D plane.
[0054] The neighborhood sub-image gradient dual-channel Gaussian filtering method, along with a robust estimation and iterative method for small offsets, preserves the frequency response of significant features in the neighborhood sub-image by using complex gradients for image representation. Furthermore, it constructs a two-dimensional Gaussian filter with a normalized cross-power spectrum matrix for both real and imaginary parts, effectively addressing the phase entanglement discontinuity caused by aliasing noise from high-speed camera high-frame-rate shooting, thus ensuring the accuracy of small offset recognition. Simultaneously, the robust estimation and iterative method for small offsets uses an improved HMSS robust estimation to calculate the initial offset value and a robust iteration to calculate the residual offset value by continuously updating the normalized cross-power spectrum, obtaining the final sub-pixel offset value and improving tracking accuracy.
[0055] ① Robust estimation of initial offset values: An improved HMSS robust estimation algorithm is used for plane fitting, and an adaptive sampling subset size is used instead of a fixed subset size larger than the minimum subset for plane fitting.
[0056] The noise estimation index σ is obtained by calculating the local residual distribution of N randomly selected points. Let p be the probability that each data point is an inlier. This p value represents the probability that most points in the data are inliers, and is usually determined by the data quality and noise level. If there is less noise in the data, the probability p of inliers will be higher; if there is more noise, the probability p of inliers will be lower. To ensure that each large subset M sampled has a sufficient number of inliers, an adjustment mechanism is designed to dynamically adjust the size of the subset using the noise estimation σ and the inlier probability p:
[0057] M=m+[λ·f(σ,p)] (5)
[0058] Where m is the minimum subset size, f(σ,p) is an adjustment function designed based on the noise estimate and the proportion p of interior points, and λ is the adjustment coefficient. If the noise level σ is low and p is high, then f(σ,p) takes a smaller value; if the noise is high, a larger M is chosen to compensate for the error risk. Finally, the parameters of the plane model are accurately estimated through iterative optimization using the LKOS cost function.
[0059]
[0060] in, This represents the estimated initial offset value, r. (t) This represents the residual iterative model, where t represents the number of iterations. This indicates the termination condition, which is that the average residual of the data sample set in the first two iterations is less than the k-th sorted residual.
[0061] ② Sub-pixel phase correlation robust iteration of the final offset value: After obtaining the initial offset value for the first time, update the normalized cross-power spectrum to calculate the residual offset value, iterate until the residual offset value is lower than a threshold (e.g., 0.01 pixels), and the final offset value is the sum of all iterated offset values:
[0062]
[0063] in, is the final offset value between frames of the image sequence, and s is the number of robust iterations.
[0064] This invention achieves precise acquisition of several neighborhood sub-images centered on a target point. It uses the complex gradient of the image instead of the original image intensity function as the image representation method, preserving the frequency response of significant features of the neighborhood sub-images. It constructs a two-dimensional Gaussian filter with a normalized cross-power spectrum matrix of real and imaginary parts to solve the problem of phase entanglement discontinuity caused by aliasing noise from high-speed camera high-frame-rate shooting. By updating the normalized cross-power spectrum to calculate the residual offset value through robust iteration, the final sub-pixel offset value is obtained, thereby achieving accurate identification and tracking of small target shifts in the sequence image.
[0065] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the protection scope of the present invention.
Claims
1. A frame-to-frame micro-shift high-speed video measurement phase correlation matching method, characterized in that, The method comprises the following steps: obtaining a high-frame-rate image sequence; dynamically adjusting the image sequence through an adaptive window to obtain a neighborhood sub-image; performing gradient real and imaginary double-channel Gaussian filtering on the neighborhood sub-image to obtain a normalized cross power spectrum matrix; performing plane fitting on the normalized cross power spectrum matrix, dynamically adjusting a sub-set size through a noise estimation index and an inlier probability, and then iteratively optimizing to obtain an initial offset value; based on the initial offset value, updating the normalized cross power spectrum to calculate a residual offset value, and iteratively optimizing until the residual offset value is lower than a preset threshold value, and then summing all the iterative offset values to obtain a final sub-pixel offset value; To preserve the frequency response of the neighborhood sub-image salient features, the image representation form employs gradient complex substitute original image intensity function : wherein, and respectively represent and gradients in the directions, and represent image coordinates, represent frame numbers, represent point numbers; Constructing real and imaginary dual-channel normalized cross-power spectrum matrix two-dimensional gaussian filter The specific formula is: wherein is a normalized cross-power spectrum matrix, denotes a Fourier transform, denotes a complex conjugate, denotes a frequency domain coordinate, and are the horizontal and vertical direction offsets to be solved, denotes a filtering on the real part, denotes a filtering on the imaginary part; To ensure a large subset of samples There can be enough inliers to estimate noise And inlier probability To dynamically adjust the size of the subset: where, is the minimum subset size, is the ratio of inliers to noise estimate is the adjustment function designed, is the adjustment coefficient, if the noise level is low and is high, then takes a smaller value; if the noise is high, a larger is selected to compensate for the risk of error, and the parameters of the final accurate plane model are iteratively optimized by the LKOS cost function. wherein, represents an estimated initial offset value, represents a residual iterative model, represents the number of iterations, represents a termination condition, i.e. the average residual of the data sample set of the last two iterations is less than the th ordered residual, represents a phase angle.
2. The frame-to-frame fractional shift high speed video measurement phase correlation matching method of claim 1, wherein, the dynamically adjusting the image sequence through an adaptive window to obtain a neighborhood sub-image specifically comprises: extracting a first-frame feature point set of a high-speed moving object target region by using a SuperPoint feature network, extracting a feature point neighborhood based on a reference window with the feature point as the center, quantifying local features by combining Sobel gradient, and dynamically adjusting a window size by taking a standard deviation of gradient values in the target region as a threshold value, wherein a high-gradient region with a gradient value exceeding the threshold value uses a small window, and a smooth region with a gradient value lower than the threshold value uses a large window, and a plurality of neighborhood sub-images centered on the feature points are segmented as matching templates.
3. The frame-to-frame fractional shift high speed video measurement phase correlation matching method of claim 1, wherein, the robust estimation of the initial offset value specifically comprises: based on the filtered normalized cross power spectrum matrix, performing plane fitting by using an improved HMSS robust estimation algorithm, the improved HMSS robust estimation algorithm dynamically adjusts a sub-set size through a noise estimation index and an inlier probability, and then iteratively optimizes to obtain an initial offset value through an LKOS cost function.
4. The frame-to-frame fractional shift high speed video measurement phase correlation matching method of claim 1, wherein, in the summing of the iterative offset values, 0.01 pixels is taken as a preset threshold value stop condition, and the final offset value is the sum of all the iterative offset values: wherein, is the final offset value between the sequence video frames, is the number of robust iterations.
Citation Information
Patent Citations
Frequency domain fast sub picture element global motion estimating method for image stability
CN101068357A
Robust plane fitting-based phase correlation sub-pixel matching method
CN103824287A