Video Global Motion Compensation Method Based on Affine Inverse Transformation Model

By combining SURF and MSAC methods, an affine transformation model is established and inverse transformation is performed, the problem of global motion compensation in the dynamic context is solved, more accurate motion object detection and reduced false alarms are achieved, and the degree of image distortion is minimized.

CN114519832BActive Publication Date: 2025-07-29HANGZHOU DIANZI UNIV
4 Cites 0 Cited by

Patent Information

Application Number
CN202210145390.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-02-17
Publication Date
2025-07-29
Estimated Expiration
2042-02-17

AI Technical Summary

Technical Problem

In video sequences under dynamic backgrounds, traditional global motion compensation methods are difficult to effectively eliminate the impact of global motion, resulting in an increase in false alarms for motion target detection, and the motion target is easily compensated as a background and the extraction is incomplete.

Method used

Combining SURF and MSAC methods, an affine transformation model is established and inverse transformation is performed, global motion is estimated and compensated, and the video sequence is output through a specific view to realize the conversion from dynamic background to static background.

Benefits of technology

It improves the extraction integrity and detection accuracy of the moving target, reduces false alarms, minimizes image distortion, and has better robustness and accuracy.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114519832B_ABST
    Figure CN114519832B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for global motion compensation of video based on an affine inverse transformation model. The present invention combines the SURF and MSAC methods for processing to obtain an affine transformation model of the global motion of a video sequence, estimates the global motion of the video sequence, proposes an inverse transformation model of the affine transformation for compensating the global motion, and realizes the conversion of the dynamic background of the video sequence to a static background. First, feature point matching pairs are extracted between adjacent frames in the video sequence under the dynamic background, and then the feature point matching pairs that do not satisfy the motion transformation are removed. An affine transformation matrix describing the global motion change of the video sequence is obtained by fitting. Finally, an inverse transformation is performed and applied to each adjacent frame in turn to compensate the global motion. An output view is created with the first frame image as the standard, and the video sequence after global motion compensation is output for target detection. The method of the present invention has a small degree of distortion of the image frames after global motion compensation and has better robustness and accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of digital image processing, especially in the fields of digital image processing and target detection, and specifically relates to a method for video global motion compensation based on an inverse affine transformation model. Background Art

[0002] Intelligent analysis technologies such as detecting, recognizing, tracking moving targets and analyzing human behaviors using videos have been widely applied in fields such as military, intelligent transportation, and medicine. Among them, the detection of moving targets is the most fundamental and crucial link. In a video sequence, according to whether the video recording device itself moves, it is mainly divided into target detection under a static background and target detection under a dynamic background. The movement of the background is usually caused by the change in the position of the video recording device, which is called global motion; while the movement of the foreground is the movement of the moving object relative to the video recording device, which is local motion. The main methods for target detection under a static background include frame difference method, background difference method, etc. These methods have achieved obvious effects under a static background, and the methods are mature and accurate. However, compared with target detection under a static background, due to the existence of global motion, the false alarms of target detection for video sequences under a dynamic background increase significantly. Therefore, before detecting the targets under a dynamic background, it is necessary to estimate and compensate the global motion to eliminate the influence brought by the global motion.

[0003] Traditional global motion compensation methods obtain the motion model between frames of video sequence images through feature point matching methods between images such as SIFT (Scale-invariant feature transform) and SURF (Speed Up Robust Features), estimate the global motion, and perform model transformation on the image to predict the next frame image, obtaining a compensated frame image. When the moving target moves slowly, the difference in motion between the moving target and the background in two consecutive frames is small, and the moving target is easily compensated into the background, resulting in incomplete extraction of the moving target. In order to better extract the moving target, the present invention focuses on the conversion from a dynamic background to a static background of the video sequence, establishes an inverse transformation model of affine transformation to act on each frame image in turn to compensate the global motion, and outputs the video sequence after global motion compensation with a specific view. This video sequence can directly use the frame difference method to obtain a relatively complete moving target. Summary of the Invention

[0004] The purpose of the present invention is to provide a method for video global motion compensation based on an inverse affine transformation model.

[0005] The method of the present invention combines the SURF method with the MSAC (M-Estimate Sample Consensus) method to obtain an affine transformation model that can more accurately describe the global motion of a video sequence under a dynamic background. On this basis, an inverse transformation model of the affine transformation is established and applied to each frame image in turn to compensate for the global motion, and a video sequence after global motion compensation is output with a specific view to achieve the conversion of the dynamic background of the video sequence to a static background. The specific steps of the method of the present invention are as follows:

[0006] Step (1): The video sequence V under a dynamic background is an ordered set V N (Θ N ):

[0007] V N (Θ N ) = {I1(θ1), I2(θ2), …, I N (θ N )}; The nth frame image I n (θ n ) = I nb (θ nb ) ∪ I no (θ no ), n = 1, 2, …, N, I nb (θ nb ) represents the set of pixel points forming the background in the nth frame image, I no (θ no ) represents the set of pixel points of all moving objects in the nth frame image, and there are a total of M moving objects,

[0008] The motion state θ nb of the background pixel points in the nth frame image n , b n , t n ) T , where the motion parameters a n , b n , t n respectively describe the scaling, rotation, and translation states of the background, and T represents the transpose; The motion state of the foreground moving object with the centroid of the mth moving object in the nth frame image as the standard This motion state superimposes the background motion state θ nb at this time and its own motion state

[0009] Step (2): Combine the SURF and MSAC methods to generate an affine transformation model describing the global motion:

[0010] (2.1) Establish a set of feature point matching pairs between adjacent frame images in the video sequence V under the dynamic background as follows:

[0011] (2.1.1) Generate all the interest points in the image using the Hessian matrix:

[0012] Before constructing the Hessian matrix, it is necessary to perform Gaussian filtering on the image. At the pixel point P(u, v) in the grayscale image corresponding to the nth frame image, establish the Hessian matrix H:

[0013] L(P, σ) is the convolution of the image I at pixel point P n (P) and the second-order Gaussian template, that is: σ is the Gaussian blur coefficient, and the subscripts uu, uv, vu, vv are second derivative markers; the intermediate parameter

[0014] The determinant of the Hessian matrix for each pixel point is: det(H) = L uu L vv -(0.9L uv ) 2 , if the value of the determinant is not 0, then determine that this pixel point is a possible extreme point, that is, an interest point.

[0015] (2.1.2) Interest point screening:

[0016] At the neighborhood of each interest point with a size of d×d×d, use a filter with a size of d×d to perform non-maximum suppression: Compare each interest point with d 3 -1 pixel points in its scale space and two-dimensional image space neighborhood. If it is not the maximum or minimum value, then eliminate it, and save the remaining key points as feature points.

[0017] (2.1.3) Feature point matching:

[0018] Establish a block of g×g centered on each feature point, and each block contains h×h pixels, and then rotate the block to the feature direction. Calculate the response value for each small block using the Haar wavelet, and then use the feature vector D = [∑d x , ∑|d x |, ∑d y , ∑|d y |] to represent the feature of this small block. ∑d x and ∑d y respectively represent the horizontal and vertical Haar wavelet response values relative to the feature direction, ∑|d x | and ∑|d y| represents the sum of the absolute values of the Haar wavelet responses in the horizontal and vertical directions relative to the feature direction, respectively. Combining the feature vectors of the g×g blocks yields the Z-dimensional feature descriptor for the feature point, where Z = 4×g×g.

[0019] The above method is used to respectively process two adjacent frames of image I in the video sequence V under the dynamic background. k (θ k ), I k+1 (θ k+1 ) to establish a feature point set: E for I k (θ k )’s number of feature points, F is I k+1 (θ k+1 ) is the number of feature points.

[0020] With the feature point set α k The corresponding description subset is With the feature point set α k+1 The corresponding description subset is And R(α k ) and R(α k+1 ) all have Z-dimensional features.

[0021] By comparing the descriptors in the two descriptor sets, the image I is k (θ k ) and I k+1 (θ k+1 ) Feature point matching: Euclidean distance is used to measure the similarity. The shorter the Euclidean distance, the better the matching degree of the two feature points. Finally, the best feature point matching point is selected:

[0022] in, Represents image I k (θ k ) feature point description subset R(α k ) in the descriptor of the e-th feature point, Represents the z-th dimension feature of the descriptor, z = 1, 2, ..., Z; Represents image I k+1 (θ k+1 ) feature point description subset The descriptor of the f-th feature point in, Represents the z-th dimension feature of the descriptor; Representation descriptor and Similarity measurement, when the value of the measurement is less than the set threshold λ, the corresponding feature point and as feature point matching pairs.

[0023] All the feature point matching pairs that meet the above conditions form image I k (θ k ) and the set of feature point matching pairs between I k+1 (θ k+1 ):

[0024] (2.2) Global motion estimation:

[0025] Most of the feature point matching pairs in the set S of feature point matching pairs between adjacent images can be generated by a model, and at least n s point pairs can be used to fit the model parameters, and these parameters are fitted in the following iterative manner, n s ≤ min(E, F):

[0026] (2.2.1) Randomly select n k feature points from the set S of feature point matching pairs, n k ≤ n s , and use these n k feature points to fit a model M k ;

[0027] (2.2.2) For the remaining feature points in S, calculate the transformation error of each point. If the error exceeds the threshold, it is marked as an outlier, otherwise it is recognized as an inlier, and the inliers are augmented and recorded in the current inlier set IS;

[0028] (2.2.3) If the cost function of the current inlier set IS is less than the cost function C best of the optimal inlier set IS best , then update the optimal inlier set to IS best = IS; where W is the error function, is the estimated model parameter, s is a pair of matching points in the set S of feature point matching pairs; L is the loss function, in the formula γ is the error calculated by the error function W, T is the error threshold for distinguishing inliers;

[0029] (2.2.4) Iterative operation. If the iteration number is reached, then end, otherwise, increment the iteration number by one and repeat the above operation. Where w is the probability that n s points are inliers after G iterations, and ω is the proportion of inliers in S.

[0030] The image I k (θ k ) and I k+1 (θk+1 The inlier matching pairs between them form a new set of matching pairs S i , and at the same time, the image I is obtained by the iterative process k (θ k ) and I k+1 (θ k+1 ) - the fitting model of the global motion between them - the affine transformation matrix: In the formula, represents the translation amount of the coordinates of two feature points, reflects the corresponding rotation change between the two feature points, reflects the corresponding scaling change between the two feature points. These six parameters together serve as the estimators of the global motion state of the image I k+1 (θ k+1 ) parameters.

[0031] In the video sequence V under the dynamic background obtained by the above method, for all the affine transformation matrices describing the global motion estimation between adjacent frame images, the change process of the background motion of the entire video sequence is estimated by the change of the motion parameters (a, b, t) among them.

[0032] Step (3) Global motion compensation:

[0033] (3.1) Create an output view:

[0034] Define a rectangular coordinate system u1 - v1 with the upper left corner of the first frame image I1(θ1) in the video sequence as the origin. The coordinates of each pixel (u, v) in the image represent the column number and row number of the pixel in the image array respectively, and its value is the gray level of the pixel point. Establish an imaging plane coordinate system x1 - y1 in centimeters, with the center of the image as the origin. Take the imaging plane coordinate system x1 - y1 as the imaging plane coordinate system of the output view, and take the pixel coordinate range of I1(θ1) on the imaging plane as the pixel range of the output view to create the output view.

[0035] (3.2) Compensate for the global motion:

[0036] Compensate for the global motion of the (k + 1)-th frame image I k+1 (θ k+1 ) as an example of its compensation process: First, perform an inverse transformation on the affine transformation matrix {M1, M2,..., M k} to obtain the inverse transformation matrix Then, apply the inverse transformation matrix of the affine transformation to each pixel of I k+1 (θ k+1 )

[0037] to obtain the new (k + 1)-th frame image (u k+1 , v k+1 ) represents the pixel coordinates of I k+1 (θ k+1 ), and represents the pixel coordinates of the image . The image after image is output through the output view established with the first frame image as the standard, and the (k + 1)-th frame image after global motion compensation is obtained Among them, due to the existence of global motion in the original video sequence V, the scenes captured by different frame images are not completely consistent, that is, some pixel points that appear in the previous frame image may not appear in the subsequent frame image. Therefore, when the pixel coordinates on the coordinate system x1 - y1 exceed the pixel coordinate range of the output view image (such as Figure 2 the non-shadow part of the image in), it means that the scene composed of this pixel point does not appear in the first frame image, so this pixel point will not be included in the output view. As shown in Figure 2 , the pixel value of the shadow part in the output view is the pixel value at the corresponding position in the image , and the pixel values of the remaining parts (the non-shadow part of the output view) are assigned 0

[0038] So far, for the (k + 1)-th frame image I′ k+1 (θ′ k+1 ), the motion state of the moving target marked as i is:

[0039] Among them, θ (k+1)b is the state quantity of the global motion in the image I k+1 (θ k+1 ), is the motion state of the moving target marked as i itself in the image I k+1 (θ k+1 ), is the estimated quantity of the global motion state in the image I k+1 (θ k+1 )

[0040] Perform global motion compensation on {I2(θ2), I3(θ3), …, I N (θ N )} in turn according to the inverse transformation model of the above affine transformation, and connect the compensated images in sequence to form a new video sequence V′:

[0041] V N ′(Θ N ) = {I1(θ1), I′2(θ′2), …, I′ k (θ′ k ), I′k+1 (θ′ k+1 ), …, I′ N (θ′ N )};

[0042] Among them, the image coordinate system and the imaging plane coordinate system of {I1(θ1), I′2(θ′2), …, I′ k (θ′ k ), I′ k+1 (θ′ k+1 ), …, I′ N (θ′ N )} are the same, both established with the image I1(θ1) as the standard, and there is only the movement of the foreground moving object in the video sequence V′.

[0043] For two adjacent grayscale images of the video sequence V′, the difference D k (x, y) = |I′ k+1 (x, y) - I′ k (x, y)| is calculated by the frame difference method. The pixel points where the difference D k (x, y) is greater than or equal to the set threshold are classified as the pixel points of the moving object, while the pixel points where the interpolation D k (x, y) is less than the set threshold are classified as the background pixel points, where I′ k (x, y) and I′ k+1 (x, y) are the grayscale values of the pixel points in two adjacent frames of images I′ k (θ′ k ) and I′ k+1 (θ′ k+1 ) in the video sequence V′.

[0044] The present invention adopts an objective standard for evaluating images - Peak Signal to Noise Ratio (PSNR) to prove the effectiveness of the method in the text. The larger the PSNR value, the less the image distortion. Calculate the PSNR value for the grayscale images of two adjacent frames among all the images that make up the video sequence V in the dynamic background; calculate the PSNR value for the compensated frame image obtained by the traditional global motion estimation algorithm and the original grayscale image. Calculate the PSNR value for two adjacent grayscale images after global motion compensation obtained by the method of the present invention. By comparing the PSNR values obtained by the above three methods, it can be found that the PSNR values obtained by the method of the present invention are the largest, indicating that the image frames after global motion compensation obtained by the method of the present invention have the least distortion degree, that is, the method of the present invention has better robustness and accuracy. Brief Description of the Drawings

[0045] Figure 1 It is the flowchart of the method of the present invention;

[0046] Figure 2 Schematic diagram of the output view of the video sequence. Detailed implementation manner

[0047] The present invention will be further described below with reference to the accompanying drawings.

[0048] The video global motion compensation method based on the affine inverse transformation model has the following specific steps Figure 1 as shown:

[0049] Step (1) The video sequence V under the dynamic background is an ordered set V N (Θ N ), where N = 280 in this embodiment:

[0050] V N (Θ N ) = {I1(θ1), I2(θ2), …, I N (θ N )}; The nth frame image I n (θ n ) = I nb (θ nb ) ∪ I no (θ no ), n = 1, 2, …, N, I nb (θ nb ) represents the set of pixel points forming the background in the nth frame image, and I no (θ no ) represents the set of pixel points of all moving targets in the nth frame image, and there are two moving targets in total.

[0051] The motion state θ of the background pixel points in the nth frame image nb = [a n , b n , t n T , where the motion parameters a n , b n , t n respectively describe the scaling, rotation, and translation states of the background, and T represents the transpose; the motion state of the foreground moving target with the centroids of the two moving targets in the nth frame image as the standard This motion state superimposes the background motion state θ nb at this time and its own motion state

[0052] Step (2) Combine the SURF and MSAC methods to generate an affine transformation model describing the global motion:

[0053] (2.1) Establish the feature point matching pairs between adjacent frame images in the video sequence V under the dynamic background:​

[0054] SURF is a commonly used feature point matching algorithm. When matching two images in the same scene, first, the Hessian matrix is used to generate all the interest points in the image. Then, after removing unreasonable interest points, the feature points are extracted, their descriptors are generated, and the feature point matching pairs are obtained through the comparison of the descriptors.

[0055] (2.1.1) Generating all the interest points in the image using the Hessian matrix:

[0056] The Hessian matrix is a square matrix composed of the second-order partial derivatives of a multivariate function, which describes the gray-scale gradient changes in all directions, obtains all the "suspicious" extreme points, and generates interest points. Before constructing the Hessian matrix, Gaussian filtering needs to be performed on the image. At the pixel point P(u, v) in the grayscale image corresponding to the nth frame image, the Hessian matrix H is established:

[0057] L(P, σ) is the convolution of the image I n (P) and the second-order Gaussian template, that is: σ is the Gaussian blur coefficient, and the subscripts uu, uv, vu, vv are the second-order derivative marks; the intermediate parameter

[0058] The determinant of the Hessian matrix for each pixel point is: det(H) = L uu L vv -(0.9L uv ) 2 , and this formula is also the discriminant of the Hessian matrix. Among them, 0.9 is the weight coefficient. If the value of the determinant is not 0, then this pixel point is determined to be a possible extreme point, that is, an interest point.

[0059] (2.1.2) Screening of interest points:

[0060] In order to obtain the feature points in I n (θ n ) that can truly be used for feature matching, the interest points need to be screened. At the 3×3×3 neighborhood of each interest point, a filter with a size of 3×3 is used for non-maximum suppression: each interest point is compared with 26 pixel points in its scale space and two-dimensional image space neighborhood. If it is not the maximum or minimum value, it is removed, and the remaining key points are saved as feature points.

[0061] (2.1.3) Feature point matching:

[0062] To obtain the feature point matching pairs, it is necessary to generate descriptors for the feature points. A 4×4 block is established centered on each feature point, and each block contains 5×5 pixels. Then, the block is rotated to the feature direction. The response value is calculated for each small block using the Haar wavelet, and then the feature vector D = [∑d x , ∑|d x |, ∑d y , ∑|d y |] represents the feature of the small block. ∑d x and ∑d y respectively represent the Haar wavelet response values in the horizontal and vertical directions relative to the feature direction. ∑|d x | and ∑|d y | respectively represent the sum of the absolute values of the Haar wavelet responses in the horizontal and vertical directions relative to the feature direction. The feature vectors of the 4×4 blocks are combined to obtain the 64-dimensional feature descriptor of the feature point, where Z = 64.

[0063] Using the above method, the feature point sets are established for the adjacent two-frame images I k (θ k ) and I k+1 (θ k+1 ) in the video sequence V under the dynamic background: E is the number of feature points of I k (θ k ), and F is the number of feature points of I k+1 (θ k+1 );

[0064] The descriptor subsets corresponding to the feature point sets α k are The descriptor subsets corresponding to the feature point sets α k+1 are and the descriptors in R(α k ) and R(α k+1 ) all have 64-dimensional features;

[0065] By comparing the descriptors in the two descriptor subsets, the matching of the feature points of the images I k (θ k ) and I k+1 (θ k+1 ) is completed: The Euclidean distance is used to measure their similarity. The shorter the Euclidean distance, the better the matching degree of the two feature points. Finally, the best feature point matching points are selected:

[0066] Among them, represents the descriptor subset R(α k ) of the feature points of the image I k (θk The descriptor of the e-th feature point in denotes the z-th dimensional feature of this descriptor, where z = 1, 2, …, Z; denotes the image I k+1 (θ k+1 ) the set of feature point descriptors of the f-th feature point in denotes the z-th dimensional feature of this descriptor; denotes the descriptor and the measure of similarity. When the value of the measure is less than the set threshold λ, the corresponding feature points and are taken as a feature point matching pair.

[0067] All the feature point matching pairs that meet the above conditions are grouped to form the set of feature point matching pairs between the image I k (θ k ) and I k+1 (θ k+1 ):

[0068] (2.2) Global motion estimation:

[0069] To obtain the accurate motion parameters between adjacent frame images I k (θ k ) and I k+1 (θ k+1 ) in the video sequence V under dynamic background and to estimate the background motion, it is necessary to remove as many feature point matching pairs that do not satisfy the motion transformation in the set S of feature point matching pairs as possible. The present invention combines the MSAC algorithm to remove the outliers in the set S of feature point matching pairs obtained by SURF as much as possible.

[0070] Most of the feature point matching pairs in the set S of feature point matching pairs between images can be generated by a model, and at least n s point pairs can be used to fit the model parameters, and these parameters are fitted in the following iterative manner, n s ≤ min(E, F):

[0071] (2.2.1) Randomly select n k feature points from the set S of feature point matching pairs, n k ≤ n s , and use these n k feature points to fit a model M k ;

[0072] (2.2.2) For the remaining feature points in S, calculate the transformation error for each point. If the error exceeds the threshold, mark it as an outlier; otherwise, identify it as an inlier and record the inliers after expansion into the current inlier set IS.

[0073] (2.2.3) If the cost function of the current inlier set IS best is less than the cost function C best of the optimal inlier set IS best = IS; where W is the error function, is the estimated model parameter, s is a pair of matching points in the feature point matching pair set S; L is the loss function, in the formula γ is the error calculated by the error function W, T is the error threshold for distinguishing inliers;

[0074] (2.2.4) Perform iterative operations. If the iteration count is reached, end; otherwise, increment the iteration count by one and repeat the above operations. Where w is the probability that n s points are inliers after G iterations, usually taking the value of 0.99; ω is the proportion of inliers in S.

[0075] Since the moving speeds of the foreground object and the background are inconsistent, and compared with the background, the number of feature points on the foreground object is much less, so in the process of repeatedly iterating to finally obtain the motion model, the transformation error between the feature points of the foreground object will exceed the threshold and be excluded as an outlier.

[0076] Form a new matching pair set S k (θ k ) and I k+1 (θ k+1 ) from the inlier matching pairs between the images I i , and at the same time obtain the global motion fitting model - the affine transformation matrix between the images I k (θ k ) and I k+1 (θ k+1 ) through the iteration process: In the formula, represents the translation amount of the two feature point coordinates, reflects the corresponding rotation change between the two feature points, reflects the corresponding scaling change between the two feature points. These six parameters together serve as the parameter of the global motion state estimator k+1 (θ k+1 ) of the image I .

[0077] In the video sequence V under the dynamic background obtained by the above method, all the affine transformation matrices describing the global motion estimation between adjacent frame images After that, the change process of the background motion of the entire video sequence can be estimated from the changes in the motion parameters (a, b, t) among them.

[0078] Step (3) Global motion compensation:

[0079] (3.1) Create an output view:

[0080] Define a rectangular coordinate system u1-v1 with the upper left corner of the first frame image I1(θ1) in the video sequence as the origin. The coordinates of each pixel (u, v) in the figure represent the column number and row number of the pixel in the image array respectively, and its value is the gray level of the pixel point. And u1-v1 is the image coordinate system in pixel units. Since the image coordinate system cannot represent the physical position of the pixel in the image, establish an imaging plane coordinate system x1-y1 in centimeters, with the center of the image as the origin. Take the imaging plane coordinate system x1-y1 as the imaging plane coordinate system of the output view, and take the pixel coordinate range of I1(θ1) on the imaging plane as the pixel range of the output view to create the output view, as Figure 2 shown.

[0081] (3.2) Compensate for the global motion:

[0082] Based on the motion parameters of adjacent frame images obtained, the present invention proposes a global motion compensation algorithm based on the affine inverse transformation model to perform motion compensation frame by frame on the video sequence V under the dynamic background. For the convenience of description, perform global motion compensation on the (k + 1)-th frame image I k+1 (θ k+1 ), and take its compensation process as an example: First, perform an inverse transformation on the affine transformation matrix {M1, M2,..., M k} to obtain the inverse transformation matrix Then, apply the inverse transformation matrix of the affine transformation to each pixel of I k+1 (θ k+1 ):

[0083] Obtain the new (k + 1)-th frame image (u k+1 , v k+1 ) represents the pixel coordinates of I k+1 (θ k+1 ), represents the image The pixel coordinates. Since they are obtained by calculating the pixel coordinates through parameters, the position of the image in its imaging plane is changed, and each frame of the image has a different imaging plane coordinate system, resulting in different imaging standards for the images. To solve this problem, the present invention passes the image Output through the output view established with the first frame of the image as the standard to obtain the (k + 1)-th frame of image I′ after global motion compensation k+1 (θ′ k+1 ). Among them, due to the existence of global motion in the original video sequence V, the scenes captured by different frames of images are not completely consistent, that is, some pixel points that appear in the previous frame of the image may not appear in the subsequent frame of the image. So when The pixel coordinates on the coordinate system x1 - y1 exceed the pixel coordinate range of the output view image (such as Figure 2 The non-shadow part of the image in), it means that the scene composed of this pixel point does not appear in the first frame of the image, so this pixel point will not be included in the output view. As shown in Figure 2 , the pixel value of the shadow part in the output view is the pixel value at the corresponding position in the image , and the pixel values of the remaining part (the non-shadow part of the output view) are assigned 0.

[0084] So far, in the (k + 1)-th frame of image I′ k+1 (θ′ k+1 ), the motion state of the moving target marked as i is:

[0085] Among them, θ (k+1)b Is the state quantity of the global motion in the image I k+1 (θ k+1 ), Is the motion state of the moving target marked as i itself in the image I k+1 (θ k+1 ), Is the estimated quantity of the global motion state in the image I k+1 (θ k+1 ).

[0086] Perform global motion compensation on {I2(θ2), I3(θ3), …, I N (θ N )} in turn according to the inverse transformation model of the above affine transformation, and connect the compensated images in sequence to form a new video sequence V′:

[0087] V′ N (Θ N ) = {I1(θ1), I′2(θ′2), …, I′ k (θ′ k ), I′ k+1 (θ′k+1 ), …, I′ N (θ′ N )};

[0088] Among them, the image coordinate system and the imaging plane coordinate system of {I1(θ1), I′2(θ′2), …, I′ k (θ′ k ), I′ k+1 (θ′ k+1 ), …, I′ N (θ′ N )} are the same, both established with the image I1(θ1) as the standard, and there is only the movement of the foreground moving object in the video sequence V′.

[0089] Subtract the adjacent two-frame grayscale images of the video sequence V′ by the frame difference method to get D k (x, y) = |I′ k+1 (x, y) - I′ k (x, y)|. Classify the pixel points where the difference D k (x, y) is greater than or equal to the set threshold as the pixel points of the moving object, while classify the pixel points where the interpolation D k (x, y) is less than the set threshold as the background pixel points. Thus, it is determined whether there is a moving object in the image sequence. Among them, I′ k (x, y) and I′ k+1 (x, y) are the grayscale values of the pixel points in the adjacent two-frame images I′ k (θ′ k ) and I′ k+1 (θ′ k+1 ) in the video sequence V′. In this way, it is determined that there are always two moving objects in the image sequence, and the pixel values belonging to the pixel point set of the moving object are not 0, while the pixel point values belonging to the background are all 0. Thus, the moving object can be clearly extracted from the image.

Claims

1. A video global motion compensation method based on an affine inverse transformation model, characterized in that, The specific method is as follows: Step (1): The video sequence V under dynamic background is an ordered set V composed of N frame images N (Θ N ): V N (Θ N ) = {I1(θ1), I2(θ2), …, I N (θ N )}; The nth frame image I n (θ n ) = I nb (θ nb ) ∪ I no (θ no ), n = 1, 2, …, N, I nb (θ nb ) represents the set of pixel points that make up the background in the nth frame of the image, I no (θ no ) represents the set of pixel points of all moving targets in the nth frame of the image, and there are M moving targets in total, The motion state θ of the background pixel points of the nth frame image nb = [a n , b n , t n T , where the motion parameters a n , b n , t n describe the scaling, rotation, and translation states of the background respectively, and T represents the transpose;​ Motion state of foreground moving object with the centroid of the m-th moving object in the n-th frame image as the standard This motion state superimposes the background motion state θ at this time nb and its own motion state In step (2), the SURF and MSAC methods are combined to generate an affine transformation model for describing global motion: (2.1) Establish a set S of feature point matching pairs between adjacent frame images in the video sequence V under a dynamic background: (2.2) Randomly select point pairs from the set S of feature point matching pairs between adjacent frame images to fit the model parameters. For all affine transformation matrices describing the global motion estimation between adjacent frame images Estimate the change process of the background motion of the entire video sequence based on the changes in the motion parameters (a, b, t). Step (3) Global motion compensation; (3.1) Create an output view: Define a rectangular coordinate system u1-v1 with the upper left corner of the first frame image I1(θ1) in the video sequence as the origin. The coordinates of each pixel (u, v) in the figure represent the column number and row number of the pixel in the image array respectively, and its value is the gray level of the pixel point. Establish an imaging plane coordinate system x1-y1 in centimeters, with the center of the image as the origin of this coordinate system; use the imaging plane coordinate system x1-y1 as the imaging plane coordinate system of the output view, and use the pixel coordinate range of I1(θ1) on the imaging plane as the pixel range of the output view to create the output view; (3.2) Compensate for global motion: For the (k + 1)-th frame image I k+1 (θ k+1 ) perform global motion compensation. Taking its compensation process as an example: First, perform an inverse transformation on the affine transformation matrix {M1, M2, …, M k} to obtain the inverse transformation matrix Then apply the inverse transformation matrix of the affine transformation to each pixel of I k+1 (θ k+1 ): Obtain the new (k + 1)-th frame image (u k+1 , v k+1 ) represents the pixel coordinates of I k+1 (θ k+1 ); represents the pixel coordinates of the image ; Output the image through the output view established with the first frame image as the standard to obtain the (k + 1)-th frame image I′ k+1 (θ′ k+1 ); When the pixel coordinates on the coordinate system x1 - y1 exceed the pixel coordinate range of the output view, it means that the scene composed of this pixel point does not appear in the first frame image, and this pixel point will not be included in the output view, and the output pixel value is the pixel value at the corresponding position in the image , and the pixel values of the remaining parts are assigned 0; The (k + 1)-th frame image I′ k+1 (θ′ k+1 ) in which the motion state of the moving target marked as i is: Among them, θ (k+1)b is the state quantity of the global motion in the image I k+1 (θ k+1 ), is the motion state of the moving target marked as i itself in the image I k+1 (θ k+1 ); is the estimated quantity of the global motion state in the image I k+1 (θ k+1 ). Perform global motion compensation on {I2(θ2), I3(θ3), …, I N (θ N )} in sequence according to the inverse transformation model of the above affine transformation, and connect the compensated images in sequence to form a new video sequence V′: V′ N (Θ N ) = {I1(θ1), I′2(θ2′), …, I′ k (θ′ k ), I′ k+1 (θ′ k+1 ), …, I′ N (θ′ N )}; Among them, the image coordinate systems and imaging plane coordinate systems of {I1(θ1), I′2(θ′2), …, I′ k (θ′ k ), I′ k+1 (θ′ k+1 ), …, I′ N (θ′ N )} are all consistent, established with the image I1(θ1) as the standard, and there is only the movement of foreground moving objects in the video sequence V′; Subtract the difference D between two adjacent grayscale images of the video sequence V' using the frame difference method k (x, y) = |I' k+1 (x, y) - I' k (x, y)|. Classify the pixel points where the difference D k (x, y) is greater than or equal to the set threshold as the pixel points of the moving object, while the pixel points where the difference D k (x, y) is less than the set threshold are classified as background pixel points, where I' k (x, y) and I' k+1 (x, y) are the grayscale values of the pixel points in two adjacent images I' k (θ' k ) and I' k+1 (θ' k+1 ) in the video sequence V'.

2. The video global motion compensation method based on the affine inverse transformation model according to claim 1, characterized in that, The specific method for establishing the set S of feature point matching pairs in step (2.1) is as follows: (2.1.1) Use the Hessian matrix to generate all the interest points in the image: Before constructing the Hessian matrix, Gaussian filtering needs to be performed on the image. Establish a Hessian matrix H at the pixel point P(u, v) in the grayscale image corresponding to the nth frame image: L(P, σ) is the image I at pixel point P n (P) convolved with a second-order Gaussian template, i.e.: σ is the Gaussian blur coefficient, and the subscripts uu, uv, vu, vv are second derivative notations; the intermediate parameter The determinant of the Hessian matrix for each pixel is: det(H) = L uu L vv -(0.9L uv ) 2 , if the value of the determinant is not 0, then the pixel is determined to be a possible extreme point, that is, an interest point; (2.1.2) Interest point screening: At the neighborhood of each interest point with size d×d×d, non-maximum suppression is performed using a filter with size d×d: Each interest point is compared with d 3 -1 pixel points in its scale space and two-dimensional image space neighborhood. If it is not a maximum or minimum value, it is removed, and the remaining key points are saved as feature points; (2.1.3) Feature point matching: Establish a g×g block centered at each feature point, and each block contains h×h pixels. Then rotate the block to the feature direction; calculate the response value for each small block using Haar wavelets, and then use the feature vector D = [∑d x , ∑|d x |, ∑d y , ∑|d y |] to represent the feature of the small block. ∑d x and ∑d y respectively represent the Haar wavelet response values in the horizontal and vertical directions relative to the feature direction. ∑|d x | and ∑|d y | respectively represent the sum of the absolute values of the Haar wavelet responses in the horizontal and vertical directions relative to the feature direction; combine the feature vectors of the g×g blocks to obtain the Z-dimensional feature descriptor of the feature point, Z = 4×g×g; For two adjacent frames of images I k (θ k ) and I k+1 (θ k+1 ) in the video sequence V under a dynamic background, the feature point sets are established by the above method: Let E be the number of feature points of I k (θ k ), and F be the number of feature points of I k+1 (θ k+1 ); The description subset corresponding to the feature point set α k is The description subset corresponding to the feature point set α k+1 is And the descriptors in R(α k ) and R(α k+1 ) both have Z-dimensional features; By comparing the descriptors in the two descriptor sets, the matching of the image I k (θ k ) and I k+1 (θ k+1 ) feature points: The Euclidean distance is used to measure their similarity. The shorter the Euclidean distance, the better the matching degree of the two feature points. Finally, the best feature point matching points are selected: Among them, represents the image I k (θ k ) is the descriptor of the e-th feature point in the feature point description subset R(α k ), represents the z-th dimension feature of this descriptor, z = 1, 2, …, Z; represents the image I k+1 (θ k+1 ) is the descriptor of the f-th feature point in the feature point description subset , represents the z-th dimension feature of this descriptor; represents the measure of the similarity between the descriptors and . When the value of the measure is less than the set threshold λ, the corresponding feature points and are used as feature point matching pairs; All the feature point matching pairs that meet the above conditions form image I k (θ k ) and I k+1 (θ k+1 ) between the set of feature point matching pairs:

3. The video global motion compensation method based on the affine inverse transformation model according to claim 1, characterized in that, In step (2.2), at least n point pairs in the set S of feature point matching pairs between images can be used to fit the model parameters, and these parameters are fitted in the following iterative manner, where n s ≤ min(E, F): s ​ (2.2.1) Randomly select n feature points from the set S of feature point matching pairs k feature points, where n k ≤ n s , and use these n k feature points to fit a model M k ; (2.2.2) For the remaining feature points in S, calculate the transformation error of each point. If the error exceeds the threshold, it is marked as an outlier, otherwise it is recognized as an inlier. The inliers are expanded and recorded in the current inlier set IS; (2.2.3) If the cost function of the current inlier set IS is less than the cost function C best of the optimal inlier set IS best , then update the optimal inlier set to IS best = IS; where W is the error function, is the estimated model parameter, s is a pair of matching points in the feature point matching pair set S; L is the loss function, in the formula, γ is the error calculated by the error function W, and T is the error threshold used to distinguish inliers; (2.2.4) Iterative operation. If the number of iterations is reached then end. Otherwise, increment the number of iterations by one and repeat the above operation. Where w is the probability that n s points are inliers after G iterations, and ω is the proportion of inliers in S; The image I obtained by the MSAC method k (θ k ) and I k+1 (θ k+1 ) form a new set of feature point matching pairs for inliers, and at the same time, the fitting model of the global motion between the image I k (θ k ) and I k+1 (θ k+1 ) is obtained through the iterative process - the affine transformation matrix: In the formula, represents the translation amount of the coordinates of two feature points, reflects the corresponding rotation change between two feature points, reflects the corresponding scaling change between two feature points. These six parameters together serve as the estimation quantity of the global motion state of the image I k+1 (θ k+1 ) parameters.

Citation Information

Patent Citations

  • Full view stabilizing method based on global characteristic point iteration

    CN101316368A

  • Self-adaptive electronic image stabilization method based on long-range view and close-range view switching

    CN103079037A

  • Improved mean shift target tracking method based on surf features

    CN104036523A

  • Electronic image stabilization algorithm based on vehicle-mounted rearview mirror system

    CN113256679A