Dynamic object suppression video image stabilization method based on depth information and global consistency
By introducing depth information and global consistency modeling into video image stabilization technology, combining the Transformer Encoder model and motion propagation module, the problems of dynamic object interference and local receptive field limitation are solved, and the effect and robustness of video image stabilization are significantly improved.
Patent Information
- Application Number
- CN202510093968.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-21
- Publication Date
- 2025-06-03
AI Technical Summary
The existing video image stabilization technology performs poorly when dealing with complex dynamic scenes, especially when a large number of dynamic objects appear in the video, and it is difficult to effectively suppress the interference of dynamic objects to global motion. The local receptive field of traditional methods limits the ability of global consistency modeling.
A dynamic object suppression video image stabilization method based on depth information and global consistency is adopted. Key points in the video are extracted and optical flow estimated, and static and dynamic weights are calculated based on depth information. The Transformer Encoder model is used for global consistency modeling, and dense residual motion is predicted through the motion propagation module, and a stable video sequence is finally output.
It effectively reduces the interference of dynamic objects on video stable results, improves the estimation accuracy of camera motion, improves the effect and robustness of video image stabilization, and achieves better global consistency modeling.
Smart Images

Figure CN120088167A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision and video processing, and particularly relates to a dynamic object suppression video stabilization method based on depth information and global consistency. Background Art
[0002] With the popularization of smartphones, drones, and other portable devices, video shooting has become an indispensable part of daily life and work. However, due to factors such as the shaking of the shooting device and camera movement, handheld videos often exhibit obvious jitter or blur during recording, which has a greater impact on video quality and viewing experience. Especially in dynamic scenes, frequent camera movement and the presence of dynamic objects make video stabilization technology face more complex challenges. To address this issue, video stabilization technology has emerged. It restores a stable video image by removing camera jitter and other unnecessary movements, greatly improving the viewing quality and smoothness of the video.
[0003] Existing video stabilization technologies can be roughly divided into two categories: pixel-level motion estimation-based methods and grid-level motion estimation-based methods. Pixel-level motion estimation-based methods mainly rely on analyzing the motion of each pixel in the video to infer the global motion of the camera, and then correct the image. However, when dealing with complex dynamic scenes, especially when there are a large number of dynamic objects in the video, these methods often perform poorly. Because the motion patterns of dynamic objects and the background are different, the local motion of dynamic objects will interfere with the estimation of the camera's global motion, thus affecting the video stabilization effect.
[0004] To overcome the limitations of pixel-level methods in complex dynamic scenes, researchers have proposed grid-level motion estimation methods. This method divides video frames into several grid blocks to reduce the interference of the local motion of dynamic objects on the global motion estimation. However, grid-level methods also have some problems. First, how to accurately distinguish the motion of dynamic objects from that of the static background remains an unsolved problem. Especially in dynamic scenes, the motion difference between foreground dynamic objects and the background is large, and traditional grid-level estimation-based methods are difficult to effectively suppress the interference of dynamic objects on the global motion. Second, the receptive fields of traditional methods such as convolutional neural networks (CNNs) are small and cannot effectively capture the dependencies between distant grid vertices. Especially in complex dynamic scenes, the limitation of local receptive fields makes it impossible for CNNs to accurately model global consistency.
[0005] On the other hand, with the continuous development of deep learning and computer vision technologies, depth information (such as depth maps, depth estimation, etc.) has become an important research direction in video stabilization technology. Depth information can provide accurate geometric position information for each pixel or grid in the video, thereby helping to distinguish camera motion from the local motion of dynamic objects. In complex dynamic scenes, depth information can effectively reduce the impact of dynamic objects on video stabilization, improve the estimation accuracy of camera motion, and thus enhance the stability of the video. Existing methods often do not make full use of depth information. Therefore, when existing video stabilization methods face dynamic scenes, especially when there is strong interference between camera motion and the motion of dynamic objects, the stabilization effect and robustness are insufficient.
[0006] Therefore, although existing methods have made some progress in video stabilization, they still face problems such as interference from dynamic objects, limitations of local receptive fields, and insufficient global consistency modeling. Summary of the Invention
[0007] In order to overcome the above-mentioned technical deficiencies, the present invention proposes a dynamic object suppression video stabilization method based on depth information and global consistency. By introducing depth information, this method can effectively distinguish camera motion from the local motion of dynamic objects, reduce the interference of dynamic objects on the video stabilization effect. At the same time, the Transformer Encoder model is used to perform global consistency modeling on motion propagation, thereby enhancing the video stabilization effect and robustness.
[0008] The technical solution adopted by the present invention to overcome its technical problems is as follows: A dynamic object suppression video stabilization method based on depth information and global consistency proposed by the present invention includes extracting key points in the video, estimating the optical flow of the key points to obtain the motion information of the key points, and obtaining the depth information of the key points; respectively calculating the static weight and the motion vector residual based on the motion information, depth information, and global homography matrix of each key point, and obtaining the dynamic weight based on the attention mechanism, fusing the dynamic weight with the static weight to obtain the comprehensive weight of the key point, and calculating the weighted motion residual based on the comprehensive weight and the motion vector residual; propagating the motion information of the key points to the initialized grid vertices, predicting the dense residual motion based on the weighted motion residual, and accumulating the dense residual motion to obtain the trajectory; smoothing and motion compensating the trajectory to output a stable video sequence.
[0009] Further, extracting key points in the video, estimating the optical flow of the key points to obtain the motion information of the key points, and obtaining the depth information of the key points specifically include: detecting key points in video frames to obtain the positions of key points in each video frame; estimating the optical flow of the detected key points to obtain the displacement of each key point between consecutive frames as the motion information of the key points; and obtaining the depth information of each key point based on the depth map of each frame.
[0010] Further, calculating the static weight and the motion vector residual respectively corresponding to the motion information, depth information, and global homography matrix of each key point, and obtaining the dynamic weight based on the attention mechanism, fusing the dynamic weight with the static weight to obtain the comprehensive weight of the key point, and calculating the weighted motion residual based on the comprehensive weight and the motion vector residual specifically include: calculating the static weight based on the motion information and depth information of each key point; obtaining the dynamic weight based on the depth information and motion information in combination with the attention mechanism; fusing the dynamic weight with the static weight to obtain the comprehensive weight of the key point; calculating the motion vector residual of the key point based on the original position of the key point, motion information, and global homography matrix; and obtaining the weighted motion residual based on the motion vector residual and the comprehensive weight.
[0011] Further, propagating the motion information of the key points to the initialized grid vertices, and predicting the dense residual motion based on the weighted motion residual specifically include: dividing the grid of the image and calculating the grid vertex positions; and associating the weighted motion residual of the key points to each grid vertex to obtain the dense residual motion.
[0012] Further, associating the weighted motion residual to each grid vertex to obtain the dense residual motion specifically includes: calculating the distance vector between each initial grid vertex and the key point; predicting the motion vector of each grid vertex through the motion propagation module based on the distance vector, weighted motion residual, and depth information of the key point; performing position encoding on the grid vertices; obtaining the motion association vector between the position-encoded grid vertices and the initial grid vertices through the attention mechanism; calculating the dense residual motion based on the motion association vector and the global homography matrix; and accumulating the dense residual motion to obtain the trajectory.
[0013] Reconstruct the motion on the detected key points based on the initialized grid vertex motion, that is, propagate the sparse key point motion information to the grid vertices, thereby improving the accuracy of estimating the camera motion.
[0014] Further, extracting the key points in the video is as shown in formula (1), and estimating the optical flow of the key points is as shown in formula (2).
[0015]
[0016] ui =(u i , v i ) = PWC(I t , I t+1 , p i ) (2)
[0017] Among them, RFNet is a key point detection method for convolutional neural networks, p i is the position of the i-th detected key point, L is the total number of key points detected in each frame, u i is the displacement between the t-th frame and the (t + 1)-th frame of the key point p i , I i and I i+1 are two adjacent images in the video respectively, u i and v i are the motion components in the horizontal and vertical directions respectively, and p i is the position of the i-th key point.
[0018] Furthermore, the static weight is calculated based on the motion information and depth information of each key point as shown in formula (3). The dynamic weight is obtained based on the attention mechanism, that is, the depth information and motion information of each key point are input into the MLP network to obtain the fused dynamic weight, as shown in formula (4). The dynamic weight and the comprehensive weight are fused to obtain the comprehensive weight as shown in formula (5).
[0019]
[0020] A i = MLP(depth i , m i ) (4)
[0021]
[0022] Among them, W i is the static weight, α and β are hyperparameters that control the influence intensity of depth information and motion information respectively, A i is the dynamic weight, m i is the motion information, is the comprehensive weight.
[0023] Based on depth and motion information, a dynamic object suppression weight is designed to distinguish the local motion of dynamic objects from camera motion.
[0024] Furthermore, based on the original position of the key point, motion information and global homography matrix, the motion vector residual of the key point is calculated, as shown in formula (6).
[0025] Δm i = p i + mi -H i (p i ) (6)
[0026] Among them, p i is the original position of the key point, m i is the motion vector of the key point, the global homography matrix H i is estimated by the RANSAC method, and Δm i is the motion vector residual of the key point.
[0027] Furthermore, the motion vector of each vertex is predicted by the motion propagation module based on the distance vector, weighted motion residual, and depth information of the key point, as shown in formula (7),
[0028] Δn k = MP({Δ ^ m i , d ik , depth i}) (7)
[0029] Among them, MP(.) represents the motion propagation module composed of several 1D convolutions, d ik is the distance vector, Δ ^ m i is the weighted motion residual, depth i is the depth information, and Δn k is the motion vector of each vertex v k of the grid.
[0030] The motion correlation vector between the grid vertex after obtaining the position encoding through the attention mechanism and the initial grid vertex adopts formulas (8)-(10),
[0031] Δn' k = ConsistencyTransform(Δn k , PE(p n,m )) (8)
[0032] {p n,m = (r, c)|r ∈ {0, 1,.., N - 1}, c ∈ {0, 1,.., M - 1}} (9)
[0033]
[0034] Among them, p n,m represents the position index of the grid, where r and c are the column numbers of the grid; PE(p n,m ) represents encoding the grid position, obtained by performing sine and cosine transforms on r and c, where sin and cos are periodic functions, and σ is the scaling parameter, Δn'k is an associated motion vector.
[0035] The beneficial effects of the present invention are as follows:
[0036] 1. Using depth information for motion propagation and dynamic object suppression to reduce the interference of dynamic objects on the video stabilization result.
[0037] 2. Reconstructing motion on the detected key points based on the initialized grid vertex motion, that is, propagating the sparse key point motion information to the grid vertices, thereby improving the accuracy of estimating camera motion;
[0038] 3. Providing spatial position information for the Transformer model, adopting two-dimensional position encoding, and mapping the two-dimensional coordinates of each grid vertex to a higher-dimensional space, thereby effectively modeling the global dependencies between distant pixels;
[0039] 4. On the basis of global motion propagation, performing trajectory smoothing, compensation, and frame cropping to eliminate jitter during motion. Description of the Drawings
[0040] Figure 1 is a schematic flowchart of a dynamic object suppression video stabilization method based on depth information and global consistency according to an embodiment of the present invention;
[0041] Figure 2 is a schematic diagram of the motion propagation network structure according to an embodiment of the present invention. Detailed Embodiments
[0042] To facilitate better understanding of the present invention by those skilled in the art, the present invention will be further described in detail below with reference to the drawings and specific embodiments. The following is only exemplary and does not limit the protection scope of the present invention.
[0043] As Figure 1 shown, a schematic flowchart of a dynamic object suppression video stabilization method based on depth information and global consistency according to this embodiment includes the following steps.
[0044] S1. Extract key points in the video frame, estimate the optical flow of the key points to obtain the motion information of the key points, and obtain the depth information of the key points.
[0045] S11. Detect key points in the video frame to obtain the key point positions of each video frame.
[0046] In an embodiment of the present invention, a key point detection method RFNet based on a convolutional neural network is used to detect key points and extract the key point positions in each frame. The key point position p i =(x i , y i ) is the image It A set of points with significant features. The key point detection formula is shown in Formula (1).
[0047]
[0048] Among them, the position of the i-th detected key point is p i , and L is the total number of key points detected in each frame.
[0049] In one embodiment of the present invention, L is set to 512.
[0050] S12. Perform optical flow estimation on the detected key points to obtain the displacement of each key point between consecutive frames as the motion information of the key points.
[0051] To improve the accuracy and efficiency of key point motion estimation, calculate the optical flow of the key points to obtain the displacement of each key point between consecutive frames.
[0052] In one embodiment of the present invention, the PWC-Net network model is used for optical flow estimation, and the optical flow estimation adopts a formula as shown in Formula (2).
[0053] u i =(u i ,v i )=PWC(I t ,I t+1 ,p i ) (2)
[0054] Among them, u i is the displacement between the t-th frame and the (t + 1)-th frame of the key point p i , which will be used as the motion information m i of the subsequent key points. I i and I i+1 are two adjacent images in the video respectively. u i and v i are the horizontal and vertical motion components respectively. The optical flow of the key points describes the motion of pixel points in the image between two adjacent frames. It is a two-dimensional vector representing its displacement in the horizontal (x direction) and vertical (y direction), and is used for subsequent weight calculation.
[0055] S13. Obtain the depth information of each key point based on the depth map of each frame.
[0056] The depth value of the key points in each frame can be obtained from the depth map of each frame. Specifically, we know the coordinate positions of the key points in each frame, and the depth value can be taken from the corresponding position in the depth map as the depth information. The depth information is used for subsequent motion propagation and dynamic object suppression to reduce the interference of dynamic objects on the video stabilization result.
[0057] In one embodiment of the present invention, the Monodepth2 monocular depth estimation algorithm is used to predict a depth map from a monocular image to obtain the depth information of each frame. The formula for depth estimation is shown in Formula (3).
[0058] depth i = Monodepth(I t , p i ) (3)
[0059] where depth i is the depth value corresponding to the key point p i .
[0060] S2. Calculate the static weight and the motion vector residual respectively based on the motion information, depth information, and global homography matrix of each key point, and obtain the dynamic weight based on the attention mechanism. Then fuse the dynamic weight with the static weight to obtain the comprehensive weight of the key point, and calculate the weighted motion residual based on the comprehensive weight and the motion vector residual.
[0061] S21. Calculate the static weight based on the motion information and depth information of each key point.
[0062] Based on the depth and motion information, design a dynamic object suppression weight to distinguish the local motion of dynamic objects from the camera motion.
[0063] For each key point, calculate the comprehensive weight according to its depth value and the amount of motion. If the depth value corresponding to the key point is small and the amount of motion is large, it is regarded as a dynamic object, and its influence on the global motion estimation is reduced.
[0064] The comprehensive weight corresponds to the weight of the global pixel points. For some regions, if the corresponding depth value is small and the amount of motion is large, the calculated weight is large. For the background, through the calculation of the comprehensive weight, a small or almost zero weight will be obtained.
[0065] Specifically, the depth of each key point is depth i , and its motion vector, that is, the motion information, is m i =(u i , v i ), where u i and v i are the horizontal and vertical motion components respectively. To determine whether it is a dynamic object, we combine the depth difference and the motion vector to calculate a comprehensive weight W i , which is used to adjust the influence of dynamic objects on the global motion propagation, as shown in Formula (4).
[0066]
[0067] Among them, α and β are hyperparameters that control the influence intensities of depth information and motion information. If the depth value depth i of a key point p i is small and the amount of motion is large, then its corresponding weight will be small, indicating that this key point may be taken from a dynamic object area. Conversely, if the depth value is large or the amount of motion is small, then the weight of this point will be large, indicating that the area of this point is more likely to be a static background.
[0068] S22, obtaining the dynamic weight based on the attention mechanism.
[0069] To better combine depth information and motion information, an attention mechanism based on a multi-layer perceptron is used to fuse these two features. For each key point p i , we input its depth information depth i and motion information m i into the MLP network to obtain a fused dynamic weight A i , as shown in formula (5):
[0070] A i = MLP(depth i , m i ) (5)
[0071] S23, fusing the dynamic weight and the static weight to obtain the comprehensive weight of the key point.
[0072] Combining the dynamic weight with the calculated weighting coefficient W i to obtain the final comprehensive weight indicating the influence of the dynamic object on the motion propagation of this mesh vertex.
[0073] S24, based on the original position p i of the key point, the motion vector m i and the global homography matrix H i , calculating the motion vector residual Δm i
[0074] Due to the influence of camera motion, all points in each frame of image will move on the same plane, and this motion relationship follows a homography transformation. Based on the original position p i of the key point and its corresponding motion vector m i , we can use the RANSAC method to estimate the global homography matrix H i to describe the plane transformation relationship from the current frame to the target frame, and calculate the motion vector residual Δm i , that is, as shown in formula (6).
[0075] Δmi = p i + m i - H i (p i ) (6)
[0076] S25. Obtain the weighted motion residual based on the motion vector residual and the comprehensive weight.
[0077] Use Δm i to obtain the weighted motion residual, denoted as Δ ^ m i , so as to generate a more accurate dynamic object motion propagation model, thereby enhancing the robustness of the mesh vertex motion estimation.
[0078] S3. Propagate the motion information of the key points to the initialized mesh vertices, predict the dense residual motion based on the weighted motion residual, and accumulate the dense residual motion to obtain the trajectory.
[0079] In order to improve the accuracy of estimating the camera motion, we reconstruct the motion on the detected key points based on the initialized mesh vertex motion, that is, propagate the sparse key point motion information to the mesh vertices. The specific process includes the following:
[0080] S31. Divide the mesh and calculate the mesh vertex positions.
[0081] Assume the image size is H×W. We divide the image into an N×M mesh, and the motion information within each mesh is propagated through the displacements of the sparse key points.
[0082] S32. Associate the weighted motion residual of the key points to each mesh vertex to obtain the dense parametric motion.
[0083] In an embodiment of the present invention, 1D convolution is used to associate the residual motion information of the key points to each mesh vertex.
[0084] Regard the sparse key points and the dense mesh vertices as two point clouds with motion and 2D position distributions, and associate the weighted motion residual of the key points to each mesh vertex based on several 1D convolutions to fine-tune the motion of the mesh vertices. The specific process includes the following.
[0085] S321. Calculate the distance vector d k between each initial mesh vertex v i and the key point p ik , as shown in formula (7).
[0086] d ik = p i - v k (7)
[0087] where \(k = 1,\cdots,N\times M\).
[0088] S322, based on the distance vector \(d\) ik , the weighted motion residual \(\Delta\) ^ m i and the depth information of the key points are predicted by the motion propagation module to obtain the motion vector \(\Delta n\) k for each vertex \(v\) k .
[0089] It should be noted that the motion vector \(\Delta n\) k corresponds to the motion residual of the mesh vertices, and the mesh vertices are generally dense. Similarly, \(\Delta m\) i represents the motion residual of the key points, and the key points are generally sparse. Through the MP network, \(\Delta m\) i is propagated and solved to obtain \(\Delta n\) k .
[0090] For \(\Delta n\) k as a dense residual motion vector of \(M\times N\) dimensions, corresponding one-to-one with the mesh vertices, the motion vector \(\Delta n\) k for each vertex \(v\) k is as shown in formula (8).
[0091] \(\Delta n\) k = MP(\(\{\Delta^m\) i , \(d\) ik , depth i \}) (8)
[0092] where MP(.) represents the motion propagation module composed of several 1D convolutions, and its structure is as Figure 2 shown, \(p\) i and \(v\) k respectively represent the \(i\)-th key point and the \(k\)-th mesh vertex in each frame. The distance vector \(d\) ik , the weighted motion residual \(\Delta^m\) i has a size of \([M\times N,L,2]\), and the depth information has a size of \([L,1]\), and they are input into three independent encoders.
[0093] The encoded features are further embedded through several 1D convolutions. Then, a weight vector of size \([M\times N,L,1]\) is obtained based on the distance embedding, and it is used to aggregate the motion features, distance features, and depth features to generate weighted features of size \([M\times N,1,2]\), and these features are subsequently used to predict the motion information \(\Delta n\) k of the mesh vertices through the motion decoder.
[0094] The overall framework of the motion propagation network structure adopts the Baseline model, improves the pre - processing and post - processing of the model, and introduces a branch for depth information encoding.
[0095] S323, perform position encoding on the grid vertices.
[0096] Since most traditional video stabilization methods are based on convolutional neural networks, their receptive fields are limited to local areas and it is difficult to effectively model the global dependence relationships between pixels at long distances. To overcome this problem and provide spatial position information for the Transformer model, a two - dimensional position encoding method is adopted to map the two - dimensional coordinates of each grid vertex to a higher - dimensional space, thereby providing unique position information for each grid vertex. The input grid size is N×M, and the calculation process of the position encoding is shown in formulas (9) and (10).
[0097] {p n,m =(r,c)|r∈{0,1,..,N - 1},c∈{0,1,..,M - 1}} (9)
[0098]
[0099] where p n,m represents the position index of the grid, where r and c are the column numbers of the grid; PE(p n,m ) represents encoding the grid position, which is obtained by performing sine and cosine transforms on r and c. Here, sin and cos are periodic functions, and σ is a scaling parameter used to adjust the encoding range.
[0100] S324, obtain the motion association vector between the position - encoded grid vertices and the initial grid vertices through the attention mechanism.
[0101] In an embodiment of the present invention, a Transformer Encode module is introduced to capture the dependence relationship of the motion vectors between the grid vertices after the motion propagation through the self - attention mechanism. Specifically, the motion information Δn k of the grid vertices after the propagation, after being processed by the self - attention of the 2D position - encoded Transformer, obtains the cross - region motion association motion vector Δn' k , as shown in formula (11).
[0102] Δn' k =ConsistencyTransform(Δn k ,PE(p n,m )) (11)
[0103] S325, based on the global homography matrix H iDetermine the final dense residual motion Δn k ’.
[0104] It uses Δn k v as the motion parameter of the grid vertices, and obtains H based on the global homography matrix i (v k ) to determine the final dense residual motion Δn k ’.
[0105] Δn k ’ = H i (v k ) + Δn k ’ (12)
[0106] S326. Accumulate the dense residual motion to obtain the trajectory.
[0107] The Δn' value of each frame k is accumulated to obtain a relatively sharp curve.
[0108] S4. Smooth the trajectory and perform motion compensation to output a stable video sequence.
[0109] The motion propagation module estimates the final position of each grid vertex, which is manifested as the trajectory of each grid in a set of videos. We smooth this trajectory using a general method and finally compensate it to the entire image. Geometric transformation is performed on the video frames through image transformation to output a stable video sequence.
[0110] Based on the global motion propagation of S1 - S3, trajectory smoothing, compensation, and frame cropping are performed to eliminate jitter during the motion. The trajectory is smoothed by weighted average or filtering method, and the video frames are compensated according to the smoothed trajectory to restore the video stabilization effect. Specifically, the difference between the smoothed trajectory and the original trajectory is calculated and compensated to optimize the final video stabilization effect.
[0111] A video stabilization method for suppressing dynamic objects based on depth information and global consistency proposed by the present invention is compared with traditional methods such as the MeshFlow method and deep learning-based methods such as the GlobalFlow method, DIFRINT method, and DUT method on the NUS dataset. Using Stability, Distortion, and Cropping as evaluation metrics respectively, the comparison results are shown in Table 1. It can be seen that the video stabilization method for suppressing dynamic objects based on depth information and global consistency proposed by the present invention achieves the best results in the Stability metric and is slightly inferior to the best method in the Distortion and Cropping metrics. The comparison results on the BIT dataset are shown in Table 2, achieving the best results in the Stability and Distortion metrics and being slightly inferior to the best method in Cropping. The video stabilization method for suppressing dynamic objects based on depth information and global consistency proposed by the present invention basically achieves the best results on the tested datasets.
[0112] Table 1
[0113]
[0114] Table 2
[0115]
[0116] In summary, by introducing depth information and global consistency modeling, the present invention can effectively suppress the interference of dynamic objects and improve the stability and visual quality of video stabilization.
[0117] It should be noted that: in other embodiments, the steps of the corresponding methods are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
Claims
1. A video stabilization method for dynamic object suppression based on depth information and global consistency, characterized in that: include: Extract key points in the video, perform optical flow estimation on the key points to obtain the running information of the key points, and obtain the depth information of the key points; Based on the motion information, depth information and global homography matrix of each key point, the static weight and motion vector residual are calculated respectively, and the dynamic weight is obtained based on the attention mechanism. The dynamic weight is fused with the static weight to obtain the comprehensive weight of the key point, and the weighted motion residual is calculated based on the comprehensive weight and motion vector residual. The motion information of the key points is propagated to the initialized mesh vertices, and dense residual motion is obtained based on weighted motion residual prediction, and the dense residual motion is accumulated to obtain the trajectory; The trajectory is smoothed and motion compensated to output a stable video sequence.
2. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 1, characterized in that: The extracting of key points in the video, performing optical flow estimation on the key points to obtain running information of the key points, and obtaining depth information of the key points specifically include: Perform key point detection on the video frame to obtain the key point position of each video frame; The optical flow of the detected key points is estimated to obtain the displacement of each key point between consecutive frames as the running information of the key point; The depth information of each key point is obtained based on the depth map of each frame.
3. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 1, characterized in that: The method includes calculating the static weight and the motion vector residual based on the motion information, depth information and global homography matrix of each key point, obtaining the dynamic weight based on the attention mechanism, fusing the dynamic weight with the static weight to obtain the comprehensive weight of the key point, and calculating the weighted motion residual based on the comprehensive weight and the motion vector residual. Specifically, the method includes: Calculate static weights based on the motion information and depth information of each key point; Dynamic weights are obtained based on depth information and operation information combined with attention mechanism; The dynamic weight and the static weight are combined to obtain the comprehensive weight of the key point; Based on the original position, motion information and global homography matrix of the key points, the motion vector residual of the key points is calculated; A weighted motion residual is obtained based on the motion vector residual and the comprehensive weight.
4. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 1, characterized in that: The method of propagating the motion information of the key points to the initialized mesh vertices and obtaining dense residual motion based on weighted motion residual prediction specifically includes: Divide the image into grids and calculate the grid point positions; The weighted motion residual of the key points is associated to each mesh vertex to obtain the dense residual motion.
5. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 4, characterized in that: The step of associating the weighted motion residual to each mesh vertex to obtain a dense residual motion specifically includes: Calculate the distance vector between each initial mesh vertex and the keypoint; The motion vector of each mesh vertex is predicted through the motion propagation module based on the distance vector, weighted motion residual and depth information of key points; Encode the positions of mesh vertices; The motion correlation vector between the position-encoded mesh vertices and the initial mesh vertices is obtained through the attention mechanism; The dense residual motion is calculated based on the motion correlation vector and the global homography matrix; The dense residual motion is accumulated to obtain the trajectory.
6. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 1, characterized in that: The key points extracted from the video are shown in formula (1), and the optical flow estimation of the key points is shown in formula (2). u i =(u i ,v i )=PWC(I t ,I t+1 ,p i ) (2) Among them, RFNet is a key point detection method of convolutional neural network, p i is the position of the i-th key point detected, L is the total number of key points detected per frame, and u i is the key point p i The displacement between the tth frame and the t+1th frame, I i and I i+1 are two adjacent images in the video, u i and v i are the horizontal and vertical motion components, respectively, i is the position of the i-th key point.
7. The method for video stabilization based on depth information and global consistency for dynamic object suppression according to claim 3, characterized in that: The static weight is calculated based on the motion information and depth information of each key point as shown in formula (3). The dynamic weight is obtained based on the attention mechanism, that is, the depth information and motion information of each key point are input into the MLP network to obtain the fused dynamic weight, as shown in formula (4). The dynamic weight is fused with the comprehensive weight to obtain the comprehensive weight as shown in formula (5). A i =MLP(depth i ,m i ) (4) Among them, W i is a static weight, α and β are hyperparameters that control the influence of depth information and motion information, respectively. i is the dynamic weight, m i For sports information, is the comprehensive weight.
8. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 3, characterized in that: Based on the original position of the key point, the motion information and the global homography matrix, the motion vector residual of the key point is calculated, as shown in formula (6): Δm i =p i +m i -H i (p i ) (6) Among them, p i is the original position of the key point, m i is the motion vector of the key point, the global homography matrix H i It is estimated by RANSAC method, Δm i is the motion vector residual of the key point.
9. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 5, characterized in that: The motion vector of each vertex is predicted by the motion propagation module based on the distance vector, weighted motion residual and depth information of the key point, as shown in formula (7): Δn k =MP({Δ ^ m i ,d ik ,depth i }) (7) Among them, MP(.) represents a motion propagation module composed of several 1D convolutions, d ik is the distance vector, Δ ^ m i is the weighted motion residual, depth i is the depth information, Δn k v for each vertex of the mesh k motion vector.
10. The method for dynamic object suppression video stabilization based on depth information and global consistency according to claim 9, characterized in that: The motion association vector between the mesh vertex and the initial mesh vertex obtained by the attention mechanism after position encoding is obtained using formulas (8)-(10), Δn' k =ConsistencyTransform(Δn k ,PE(p n,m )) (8) {p n,m =(r,c)|r∈{0,1,..,N-1},c∈{0,1,..,M-1}} (9) Among them, p n,m Represents the position index of the grid, where r and c are the column numbers of the grid; PE(p n,m ) represents the encoding of the grid position, which is obtained by performing sine and cosine transforms on r and c, where sin and cos are periodic functions, σ is the scaling parameter, and Δn' k is the associated motion vector.
Citation Information
Cited By
Camera motion estimation method and system under foreground mask based on optical flow guidance
CN120602799A
A camera motion estimation method and system based on optical flow-guided foreground masking
CN120602799B
Multi-camera cooperative anti-shake control system and method
CN121000969A
Multi-camera cooperative anti-shake control system and method
CN121000969B
Dynamic scene adaptive image stabilization processing method and device, equipment and medium
CN121438167A