Video image stabilization frame acquisition method and device based on Kalman filtering

Optimizing the video image stabilization process through Kalman filtering and RANSAC algorithm reduces the computational complexity, real-time video image stabilization, and generates more stable video frames.

CN120128800APending Publication Date: 2025-06-10SHANGHAI WINGTECH INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510509824.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-22
Publication Date
2025-06-10

AI Technical Summary

Technical Problem

The existing video image stabilization algorithm based on Kalman filtering has high computational complexity and is difficult to achieve real-time processing, especially when processing complex motions is huge.

Method used

By calculating the initial homography matrix and performing Kalman filtering, the predicted homography matrix is obtained, and the weighted calculation is combined with the standard unit matrix to obtain the stable homography matrix, the geometric transformation is performed using the prediction and stable homography matrix, and combined with the iterative calculation of RANSAC, the target stable image frame is generated.

Benefits of technology

This reduces the complexity of the algorithm, improves the efficiency of video stabilization, realizes real-time processing, reduces the impact of noise and error, and generates more stable video frames.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120128800A_ABST
    Figure CN120128800A_ABST
Patent Text Reader

Abstract

The embodiment of the invention provides a video image stabilization frame obtaining method and device based on Kalman filtering, and relates to the technical field of data processing, and the method specifically comprises the steps: obtaining a first initial feature point set and a second initial feature point set corresponding to a first target gray level image and a second target gray level image respectively; calculating a corresponding initial homography matrix; performing Kalman filtering on the initial homography matrix to obtain a predicted homography matrix; performing geometric transformation on the first initial feature point set by adopting a prediction homography matrix and a dimensional stability homography matrix to obtain a first target feature point set and a second target feature point set; performing RANSAC iterative calculation based on the first target feature point set and the second target feature point set to obtain a target image stabilization homography matrix; and performing transmission transformation on the first original image based on the target image stabilization homography matrix to generate a target image stabilization frame. According to the invention, the video image stabilization efficiency can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of data processing, and particularly to a method and device for obtaining video stabilization frames based on Kalman filtering. Background Art

[0002] With the increasing popularity of social media and short videos in modern society, videos play an increasingly important role in our daily lives. However, since videos are usually shot by amateur devices such as mobile phones and drones in daily life and do not have image stabilization functions, the shooting effects are not good. If professional hardware devices are used for image stabilization, there are also disadvantages such as inconvenient carrying and high costs. Video stabilization algorithms have become an essential solution to improve video jitter effects.

[0003] Currently, if pure image stabilization algorithms are to achieve real-time processing, Kalman filtering is generally used. However, traditional video stabilization algorithms based on Kalman filtering usually separate the distortion amounts of the homography transformation generated in consecutive frames, and then use the main motion amounts after separating the distortion for filtering to obtain the image stabilization effect. This method requires decomposing the motion vectors, and the calculation process is complex. For example, when dealing with complex motions between video frames, it is necessary to accurately analyze the influence of each motion component on the homography matrix, and the amount of calculation is huge.

[0004] In summary, how to improve the video stabilization efficiency based on Kalman filtering, reduce the algorithm complexity, and achieve real-time processing has become an urgent problem to be solved. Summary of the Invention

[0005] In view of this, embodiments of the present application provide a method and device for obtaining video stabilization frames based on Kalman filtering, which can improve the efficiency of video stabilization.

[0006] In a first aspect, embodiments of the present application provide a method for obtaining video stabilization frames based on Kalman filtering, including:

[0007] Calculating a corresponding initial homography matrix according to a first initial feature point set and a second initial feature point set corresponding to a first target grayscale image and a second target grayscale image respectively; the first target grayscale image and the second target grayscale image are obtained based on a first original image and a second original image corresponding to two adjacent video frames in the video to be stabilized;

[0008] Performing Kalman filtering on the initial homography matrix to obtain a predicted homography matrix, and performing weighted calculation in combination with an identity matrix to obtain a stable homography matrix;

[0009] Performing geometric transformation on the first initial feature point set by using the predicted homography matrix and the stable homography matrix respectively to obtain a first target feature point set and a second target feature point set;

[0010] Perform RANSAC iterative calculation based on the first set of target feature points and the second set of target feature points to obtain the target video stabilization homography matrix;

[0011] Perform perspective transformation on the first original image based on the target video stabilization homography matrix to generate the target video stabilization frame.

[0012] As an optional implementation manner of an embodiment of the present application, before calculating the corresponding initial homography matrix according to the first set of initial feature points and the second set of initial feature points corresponding to the first target grayscale image and the second target grayscale image respectively, the method further includes:

[0013] Parse the video to be stabilized to obtain the first original image and the second original image corresponding to two adjacent video frames;

[0014] Convert the image formats of the first original image and the second original image into grayscale images to generate the first target grayscale image and the second target grayscale image;

[0015] Use the ORB feature point extraction algorithm to extract feature points from the first original image and the second original image to generate the first set of initial feature points and the second set of initial feature points.

[0016] As an optional implementation manner of an embodiment of the present application, the performing RANSAC iterative calculation based on the first set of target feature points and the second set of target feature points to obtain the target video stabilization homography matrix includes:

[0017] Perform feature point matching on the first set of target feature points and the second set of target feature points, and perform iterative calculation through the RANSAC algorithm to obtain the initial homography matrix corresponding to the single-frame images of two adjacent frames.

[0018] As an optional implementation manner of an embodiment of the present application, before calculating the corresponding initial homography matrix according to the first set of initial feature points and the second set of initial feature points corresponding to the first target grayscale image and the second target grayscale image respectively, the method further includes:

[0019] Convert the image formats of the first original image and the second original image into grayscale images to obtain the first initial grayscale image and the second initial grayscale image;

[0020] Reduce the first initial grayscale image and the second initial grayscale image based on a preset reduction factor to generate the first target grayscale image and the second target grayscale image.

[0021] As an alternative implementation manner of an embodiment of the present application, after calculating the corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image respectively, the method further includes:

[0022] Based on a preset amplification factor, perform an amplification operation on the initial homography matrix.

[0023] As an alternative implementation manner of an embodiment of the present application, the performing Kalman filtering on the initial homography matrix to obtain a predicted homography matrix includes:

[0024] Obtain an estimated value of the predicted homography matrix through an estimated value of a historical predicted homography matrix and a target state transition matrix;

[0025] Perform verification and update according to the estimated value of the predicted homography matrix and the initial homography matrix to obtain the predicted homography matrix.

[0026] As an alternative implementation manner of an embodiment of the present application, the performing perspective transformation on the first original image based on the target video stabilization homography matrix to generate a target video stabilization frame includes:

[0027] Convert the image format of the first original image to the YUV format to generate a first YUV image;

[0028] Perform perspective transformation on the first YUV image based on the target video stabilization homography matrix to obtain a target video stabilization frame.

[0029] In a second aspect, an embodiment of the present application provides a video stabilization frame acquisition device based on Kalman filtering, including:

[0030] A calculation unit, configured to calculate a corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image respectively; the first target grayscale image and the second target grayscale image are obtained based on the first original image and the second original image corresponding to two adjacent video frames in a video to be stabilized;

[0031] A filtering unit, configured to perform Kalman filtering on the initial homography matrix to obtain a predicted homography matrix, and perform weighted calculation in combination with an identity matrix to obtain a stability maintenance homography matrix;

[0032] A transformation unit, configured to perform geometric transformation on the first initial feature point set by using the predicted homography matrix and the stability maintenance homography matrix respectively to obtain a first target feature point set and a second target feature point set;

[0033] An acquisition unit, configured to perform RANSAC iterative calculation based on the first set of target feature points and the second set of target feature points, and acquire a target image stabilization homography matrix;

[0034] A generation unit, configured to perform perspective transformation on the first original image based on the target image stabilization homography matrix to generate a target image stabilization frame.

[0035] As an optional implementation manner of an embodiment of the present application, the calculation unit is specifically configured to parse the video to be stabilized, and acquire the first original image and the second original image corresponding to two adjacent video frames; convert the image formats of the first original image and the second original image into grayscale images to generate the first target grayscale image and the second target grayscale image; use the ORB feature point extraction algorithm to extract feature points from the first original image and the second original image to generate the first initial set of feature points and the second initial set of feature points.

[0036] As an optional implementation manner of an embodiment of the present application, the calculation unit is specifically configured to perform feature point matching on the first set of target feature points and the second set of target feature points, and perform iterative calculation through the RANSAC algorithm to acquire an initial homography matrix corresponding to the single-frame images of two adjacent frames.

[0037] As an optional implementation manner of an embodiment of the present application, the calculation unit is specifically configured to convert the image formats of the first original image and the second original image into grayscale images to acquire a first initial grayscale image and a second initial grayscale image; perform downsampling on the first initial grayscale image and the second initial grayscale image based on a preset downsampling factor to generate the first target grayscale image and the second target grayscale image.

[0038] As an optional implementation manner of an embodiment of the present application, the calculation unit is further configured to perform an amplification operation on the initial homography matrix based on a preset amplification factor.

[0039] As an optional implementation manner of an embodiment of the present application, the filtering unit is specifically configured to acquire an estimated value of the predicted homography matrix through an estimated value of the historical predicted homography matrix and a target state transition matrix; perform verification and update based on the estimated value of the predicted homography matrix and the initial homography matrix to acquire the predicted homography matrix.

[0040] As an optional implementation manner of an embodiment of the present application, the generation unit is specifically configured to convert the image format of the first original image into the YUV format to generate a first YUV image; perform perspective transformation on the first YUV image based on the target image stabilization homography matrix to acquire a target image stabilization frame.

[0041] In a third aspect, an embodiment of the present application provides an electronic device, including: a memory and a processor, where the memory is used to store a computer program; the processor is used to cause the electronic device to implement the method for obtaining a video stabilized frame based on Kalman filtering described in any one of the above embodiments when executing the computer program.

[0042] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored. When the computer program is executed by a computing device, the computing device is caused to implement the method for obtaining a video stabilized frame based on Kalman filtering described in any one of the above embodiments.

[0043] The method for obtaining a video stabilized frame based on Kalman filtering provided by the embodiment of the present application: calculating a corresponding initial homography matrix according to a first initial feature point set and a second initial feature point set respectively corresponding to a first target grayscale image and a second target grayscale image; the first target grayscale image and the second target grayscale image are obtained based on a first original image and a second original image corresponding to two adjacent video frames in a video to be stabilized; performing Kalman filtering on the initial homography matrix to obtain a predicted homography matrix, and performing weighted calculation in combination with an identity matrix to obtain a stabilized homography matrix; respectively performing geometric transformation on the first initial feature point set by using the predicted homography matrix and the stabilized homography matrix to obtain a first target feature point set and a second target feature point set; performing RANSAC iterative calculation based on the first target feature point set and the second target feature point set to obtain a target stabilized homography matrix; performing projective transformation on the first original image based on the target stabilized homography matrix to generate a target stabilized frame.

[0044] The present application processes the initial homography matrix through Kalman filtering. Considering the continuity of the motion between video frames, it directly predicts the possible motion transformation of the next frame image relative to the current frame image, obtains a more accurate predicted homography matrix, provides a reliable basis for the stabilized image processing, and then performs weighted calculation in combination with an identity matrix to obtain a stabilized homography matrix. When there is a large prediction error, the identity matrix makes the stabilized homography matrix closer to maintaining the original state of the image, suppressing the influence of noise and errors, and making the transformation of the video frame more stable. Compared with the traditional Kalman filtering that directly separates the distortion amount of the distortion generated by the homography transformation in consecutive frames, and then uses the main motion amount after separating the distortion for filtering to obtain the stabilized image effect, the present application avoids the complex motion vector decomposition process, reduces the algorithm complexity, and improves the efficiency of obtaining the stabilized frame through Kalman filtering. BRIEF DESCRIPTION OF THE DRAWINGS

[0045] The accompanying drawings here are incorporated into the description and form a part of this description, showing embodiments consistent with the present application, and are used together with the description to explain the principles of the present application.

[0046] To more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the accompanying drawings required for the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0047] Figure 1 One of the flowcharts of the steps of the method for obtaining video stabilized frames based on Kalman filtering provided by the embodiments of the present application;

[0048] Figure 2 Another flowchart of the steps of the method for obtaining video stabilized frames based on Kalman filtering provided by the embodiments of the present application;

[0049] Figure 3 A schematic structural diagram for obtaining stabilized images of images in a video provided by the embodiments of the present application;

[0050] Figure 4 A schematic hardware structure diagram of an electronic device provided by the embodiments of the present application. Detailed implementation manners

[0051] In order to be able to more clearly understand the above-mentioned objects, features, and advantages of the present disclosure, the solutions of the present disclosure will be further described below. It should be noted that, without conflict, the embodiments of the present disclosure and the features in the embodiments can be combined with each other.

[0052] Many specific details are set forth in the following description in order to fully understand the present disclosure, but the present disclosure can also be implemented in other ways different from those described herein; obviously, the embodiments in the description are only part of the embodiments of the present disclosure, rather than all of the embodiments.

[0053] In the embodiments of the present application, words such as "exemplary" or "for example" are used to indicate examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the embodiments of the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Exactly speaking, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner. In addition, in the description of the embodiments of the present application, unless otherwise specified, the meaning of "a plurality of" refers to two or more.

[0054] It should be noted that in this article, the term "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not explicitly listed, or elements inherent to such a process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the said element.

[0055] With the development of technology, the applications of social media and short videos in modern society have become increasingly popular. Videos play an increasingly important role in our daily lives. However, since videos are usually shot with amateur devices such as mobile phones and drones in daily life, which do not have the function of image stabilization, the shooting effect is not good. If professional hardware devices are used for image stabilization, there are also disadvantages such as inconvenient carrying and high costs. Therefore, after obtaining the videos shot by various devices, it is necessary to perform image stabilization processing on the videos to improve video jitter. To solve the above problems, this application further proposes a method for obtaining video stabilization frames based on Kalman filtering, which is specifically described as follows:

[0056] An embodiment of this application provides a method for obtaining video stabilization frames based on Kalman filtering. Referring to Figure 1 as shown, the method for obtaining video stabilization frames based on Kalman filtering includes the following steps S101 - S105:

[0057] S101. Calculate the corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image respectively.

[0058] Among them, the first target grayscale image and the second target grayscale image are obtained based on the first original image and the second original image corresponding to two adjacent video frames in the video to be stabilized.

[0059] It should be noted that in the embodiment of this application, two adjacent video frames in the video to be stabilized are used to obtain one stabilized frame, and then the stabilized frames corresponding to all adjacent two - frame video frames of the video to be stabilized are obtained. Then, video encoding is performed through multiple stabilized frames to obtain the stabilized video corresponding to the video to be stabilized.

[0060] Among them, the homography matrix is usually a 3*3 matrix, which describes the perspective transformation relationship between two planes. In computer vision, the homography of a plane is defined as the projection mapping from one plane to another plane. If the homogeneous coordinates are used to map a point P on the calibration board to the point m on the imager, this mapping relationship matrix is the homography matrix between the two planes. In this application, by obtaining the initial homography matrix of two adjacent frames in the video to be stabilized, the coordinate transformation relationship between two adjacent video frames can be comprehensively obtained. In order to accurately analyze the relationship between two adjacent frames of images, it is necessary to extract features from each frame of image that are representative and can describe the local or overall characteristics of the image. These features should be to a certain extent unaffected by common transformations such as translation, rotation, and scaling of the image, so that accurate matching and analysis can be performed subsequently.

[0061] Specifically, since the video frame images obtained by decoding the video to be stabilized are usually RGBA format images or other format images in non-gray scale format, in order to facilitate subsequent feature point extraction and processing, the embodiments of this application convert the format of the first original image and the second original image to obtain the first target gray image and the second target gray image. This is because a gray image only contains luminance information and no color information. Therefore, the data structure of a gray image is relatively simple. Compared with a color image, it is more efficient in storage and processing, and at the same time retains the main structure and features of the image, which is more conducive to the accurate extraction of feature points.

[0062] Because there are often some geometric transformation relationships, changes in features, etc. between two adjacent frames due to factors such as the movement of objects, the movement of the camera, or subtle changes in the scene. By analyzing two adjacent frames of images, these information can be captured for subsequent image stabilization. Then, after extracting the corresponding feature points from the first target gray image and the second target gray image respectively, the first initial feature point set and the second initial feature point set can be generated. Then, by calculating with these two initial feature point sets, using the first initial feature point set and the second initial feature point set, first find the corresponding relationship between the two sets of feature points through a matching algorithm. Then, use an algorithm such as RANSAC (Random Sample Consensus) to calculate the initial homography matrix; among them, the RANSAC algorithm can screen out the correct matching points from a large number of feature point matching pairs and remove the mismatched points, so as to more accurately estimate the two-dimensional projection transformation relationship from the first target gray image to the second target gray image, that is, the initial homography matrix.

[0063] Exemplarily, the video to be stabilized can be composed of 50 video frames. Then, two adjacent video frames can be the second video frame and the third video frame. Thus, the second video frame and the third video frame are the first original image and the second original image respectively. Furthermore, based on the first original image and the second original image, the first target grayscale image and the second target grayscale image are obtained.

[0064] S102. Perform Kalman filtering on the initial homography matrix to obtain a predicted homography matrix, and perform weighted calculation in combination with the standard identity matrix to obtain a stabilized homography matrix.

[0065] It should be noted that Kalman filtering is based on Bayesian estimation theory, which combines the state space model and the observation model of the system. It continuously optimizes the estimation of the system state through two steps: prediction and update. In the prediction step, according to the dynamic model of the system, the state estimation value at the previous moment is used to predict the state at the current moment. In the update step, the predicted value is fused with the actual observation value, and the Kalman gain is used to adjust the predicted value to obtain a more accurate current state estimation value. In this way, Kalman filtering can estimate the true state of the system in real time in the presence of noise and uncertainty.

[0066] Furthermore, in the embodiment of the present application, Kalman filtering is used in video stabilization. By performing Kalman filtering on the initial homography matrix, according to the continuity of the motion between video frames, the possible motion transformation of the next frame image relative to the current frame image can be predicted. Since the camera jitter usually has a certain regularity and continuity, Kalman filtering can utilize this characteristic to accurately predict the motion trend of the image, obtain a predicted homography matrix, and provide a basis for subsequent stabilization processing.

[0067] At the same time, since various noises are inevitably introduced during the video shooting process, and there may also be certain errors in the processes such as feature point extraction and initial homography matrix calculation. Kalman filtering can also effectively suppress the influence of these noises and errors through its optimal estimation characteristic. During the prediction and update processes, it will automatically adjust the weights of the predicted value and the observation value according to the statistical characteristics of the noise and the reliability of the observation value, so that the estimation result is more stable and accurate. The homography matrix processed by Kalman filtering can better reflect the true motion of the image and reduce the unstable factors caused by noise and errors.

[0068] The initial homography matrix is used as the input of the Kalman filter. The Kalman filter is a recursive optimal estimation method that can predict and update the state of the system according to the state equation and observation equation of the system, combining historical information and current observations. In the present invention, the initial homography matrix is processed by the Kalman filter, taking into account the continuity of motion between video frames and the influence of noise, to obtain a predicted homography matrix. The predicted homography matrix can more accurately reflect the true motion trend between video frames.

[0069] Meanwhile, weighted calculation is performed based on the standard identity matrix (the standard identity matrix is a matrix with 1 on the main diagonal and 0 for the remaining elements). By setting appropriate weighting coefficients, a stabilized homography matrix is obtained to suppress possible noise or inaccurate estimations. It should be noted that the weighting coefficients can be adjusted according to the actual video scene and jitter situation to achieve the best image stabilization effect.

[0070] S103: Geometric transformations are respectively performed on the first initial feature point set using the predicted homography matrix and the stabilized homography matrix to obtain a first target feature point set and a second target feature point set.

[0071] In the embodiment of the present application, after obtaining the predicted homography matrix and the stabilized homography matrix, it is necessary to perform geometric transformations on the first initial feature point set using the predicted homography matrix and the stabilized homography matrix.

[0072] Among them, the geometric transformation is a process of coordinate transformation of the feature points extracted from the first target grayscale image according to the predicted homography matrix and the stabilized homography matrix, and then operations such as translation, rotation, scaling, and perspective distortion are performed on the first target grayscale image. Furthermore, the first initial feature point set is transformed by the predicted homography matrix to obtain a first target feature point set, so that the first target feature point set can reflect the positions of each point in the first initial feature point set after transformation predicted by the Kalman filter; the first initial feature point set is transformed by the stabilized homography matrix to obtain a second target feature point set, so as to control the stability degree of the geometric transformation through the second target feature points.

[0073] Performing geometric transformation on the first initial feature point set using the predicted homography matrix is based on the prediction result of the Kalman filter for image motion. The predicted homography matrix predicts the possible positions of the first initial feature points on the second frame image (i.e., the second target grayscale image) according to the motion trend from the previous frame to the current frame, thereby obtaining a first target feature point set. This process takes into account the continuity of motion between video frames and the law of camera jitter, and maps each point in the first initial feature point set to a new position through matrix multiplication to form a first target feature point set.

[0074] Transform using the stable homography matrix. The stable homography matrix is obtained by weighted calculation in combination with the standard unit matrix. To a certain extent, it synthesizes the predicted information and the inherent characteristics of the current frame (the standard unit matrix represents the initial state without any transformation). Use the stable homography matrix to perform a geometric transformation on the first set of initial feature points to obtain the second set of target feature points. The purpose of doing this is to take into account the predicted motion while considering the characteristics of the current frame itself, making the transformation of the feature points more stable and accurate, and avoiding deviations that may occur due to solely relying on prediction. By performing geometric transformations on the first set of initial feature points using these two homography matrices, the first set of target feature points and the second set of target feature points obtained can more accurately reflect the feature correspondence relationship between two adjacent frames of images, providing a reliable basis for subsequent calculation of the target stable image homography matrix and generation of the target stable image frame.

[0075] S104. Perform RANSAC iterative calculation based on the first set of target feature points and the second set of target feature points to obtain the target stable image homography matrix.

[0076] Specifically, the target stable image homography matrix can be understood as a matrix that can describe the best transformation relationship between the first original image and the second original image of two adjacent video frames in the current frame. Through it, the image can be transformed to achieve the stable image effect.

[0077] Among them, the RANSAC (Random Sample Consensus) algorithm is a model fitting method used to process data containing noise and errors. In this case, it is necessary to fit a target stable image homography matrix that can accurately describe the relationship between two frames of images through the first set of target feature points and the second set of target feature points, which is used to describe the perspective transformation relationship between two frames of images.

[0078] The first set of target feature points and the second set of target feature points contain feature point information after different transformations. By further calculating these two sets of feature points, such as calculating the differences and correlations between them, and comprehensively considering the factors of prediction and stabilization, the target stable image homography matrix is finally obtained. The target stable image homography matrix will be used to accurately transform the original image to achieve the purpose of video stabilization.

[0079] Specifically, the specific implementation steps of the above step S104 (perform RANSAC iterative calculation based on the first set of target feature points and the second set of target feature points to obtain the target stable image homography matrix) are as follows:

[0080] Perform feature point matching on the first set of target feature points and the second set of target feature points, and perform iterative calculation through the RANSAC algorithm to obtain the initial homography matrix corresponding to the single-frame images of the two adjacent frames.

[0081] It should be noted that when matching the feature points of the first target feature point set and the second target feature point set, the feature points in the first target feature point set and the feature points in the second target feature point set can be matched according to the distance between the positions of the first target feature point set and the second target feature points in the image.

[0082] Then, iterative calculations are performed through the RANSAC algorithm; in each iteration, a subset is randomly selected from the first target feature point set. For calculating the homography matrix between two frames of images, usually 4 pairs of feature points are randomly selected (because the homography matrix has 8 degrees of freedom, and 4 pairs of feature points can uniquely determine a homography matrix).

[0083] Then, using the extracted subset of feature points, a hypothetical image stabilization homography matrix is calculated according to the corresponding mathematical model, and then this hypothetical image stabilization homography matrix is used to verify all the extracted feature points. The specific verification method is to transform the feature points in the first frame of image to the corresponding positions in the second frame of image according to the model, and then calculate the distance (i.e., the reprojection error) between the transformed feature points and the corresponding feature points in the actual second frame of image.

[0084] Repeat the above iterative process several times (for example, set the number of iterations to N), and record the reprojection error of the current hypothetical image stabilization homography matrix after each iteration. Finally, select the hypothetical homography matrix with the smallest reprojection error as the final result, that is, as the target image stabilization homography matrix. In this way, through multiple iterations, in the presence of noise, a relatively accurate model can be found to describe the relationship between two adjacent frames of images.

[0085] Through the above steps, for each pair of adjacent frames of images in the target image sequence, features are first extracted, and then iterative calculations are performed through the RANSAC algorithm, and a homography matrix that can describe the relationship between the two frames of images can be obtained relatively accurately, providing an important basis for subsequent image processing tasks. Furthermore, the calculated target image stabilization homography matrix can accurately describe the transformation relationship between the two frames of images, providing key parameters for subsequent perspective transformation of the original image to generate a stabilized image frame.

[0086] S105. Perform a perspective transformation on the first original image based on the target image stabilization homography matrix to generate a target stabilized image frame.

[0087] Specifically, in the context of video stabilization, the first original image and the target stabilized image frame can be regarded as two planar images in different states, and the target image stabilization homography matrix characterizes how to map the points in the first original image to the corresponding positions in the target stabilized image frame to eliminate the instability caused by factors such as camera jitter in the image.

[0088] Among them, warping is a transformation method that projects a planar image onto another plane, and it can simulate the perspective effect when a human eye or a camera observes an object. In a two-dimensional image, the warping transformation can be described by a 3x3 homography matrix. In the embodiments of the present application, the target image stabilization homography matrix can define the mapping relationship between the points on the first original image plane and the points on the target image stabilization frame plane.

[0089] After obtaining the target homography matrix, the above-mentioned perspective transformation can be performed on the intermediate frames in the current target image sequence through this target homography matrix to obtain the corresponding target image stabilization frames. Through the above process of performing perspective transformation on the first original image based on the target image stabilization homography matrix, the motion jitter between video frames is effectively compensated, and stable target image stabilization frames are generated, thereby achieving the purpose of video image stabilization and improving the viewing effect and quality of the video.

[0090] In the present application, the initial homography matrix is processed through Kalman filtering. Considering the continuity of the motion between video frames, the possible motion transformation of the next frame image relative to the current frame image is directly predicted to obtain a more accurate predicted homography matrix, which provides a reliable basis for image stabilization processing. Then, combined with the standard identity matrix for weighted calculation to obtain the stable homography matrix. When there is a large prediction error, the standard identity matrix makes the stable homography matrix closer to maintaining the original state of the image, suppressing the influence of noise and errors, and making the transformation of video frames smoother. Compared with the traditional Kalman filtering that directly separates the distortion amount of the distortion generated by the homography transformation in consecutive frames, and then uses the main motion amount after separating the distortion for filtering to obtain the image stabilization effect, the present application avoids the complex process of motion vector decomposition, reduces the algorithm complexity, and improves the efficiency of obtaining image stabilization frames through Kalman filtering.

[0091] As an extension and refinement of the above embodiments, referring to Figure 2 As shown, a method for obtaining a video image stabilization frame based on Kalman filtering provided by the present application further includes the following steps S201 to S209:

[0092] S201. Analyze the video to be image stabilized to obtain the first original image and the second original image corresponding to two adjacent video frames.

[0093] In some embodiments, the video to be image stabilized is formed by sequentially playing a series of consecutive image frames at a certain frame rate. Video analysis is to decode the video file according to its coding format and restore the compressed and stored video data to the original image frame sequence. Common video coding formats include H.264, H.265, MPEG, etc. Different coding formats have different compression algorithms and decoding methods.

[0094] After decoding the video to be image-stabilized, the obtained video frame images are usually images in RGBA format or other non-grayscale format images. To facilitate subsequent feature point extraction and processing, grayscale images remove color information, reduce the data volume, and at the same time retain the main structure and features of the images, which is more conducive to the accurate extraction of feature points. Therefore, it is necessary to convert the format of the first original image and the second original image to obtain the first target grayscale image and the second target grayscale image.

[0095] S202. Convert the image formats of the first original image and the second original image into grayscale images to generate the first target grayscale image and the second target grayscale image.

[0096] Specifically, the above step S202 (converting the image formats of the first original image and the second original image into grayscale images to generate the first target grayscale image and the second target grayscale image) can be refined into the following steps 1 and 2:

[0097] Step 1: Convert the image formats of the first original image and the second original image into grayscale images to obtain a first initial grayscale image and a second initial grayscale image.

[0098] In some embodiments, since the first original image and the second original image obtained by decoding the video to be image-stabilized are not grayscale images, then for the convenience of subsequent feature extraction, after converting the formats of the first original image and the second original image, the first initial grayscale image and the second initial grayscale image corresponding to the first original image and the second original image can be obtained respectively.

[0099] Step 2: Reduce the first initial grayscale image and the second initial grayscale image based on a preset reduction factor to generate the first target grayscale image and the second target grayscale image.

[0100] After obtaining the first initial grayscale image and the second initial grayscale image in the above step 1, in order to greatly improve the single-frame processing speed when extracting feature points of the image in the subsequent process, it is also necessary to reduce the first initial grayscale image and the second initial grayscale image; among them, the reduction factor can be represented in the form of a matrix, referring to the following expression:

[0101]

[0102] Among them, scaleH represents the reduction factor, and s can be understood as the reduction multiple.

[0103] It should be noted that when reducing the image, a reasonable scaling factor is set according to the actual situation. When reducing the image, the number of pixels in the image will decrease. For example, if the side length of the image is reduced to half of the original, then the number of pixels will become one-fourth of the original. When extracting feature points, it is usually necessary to calculate for each pixel or a local area of the pixel, and the reduction in the number of pixels will reduce the computational amount.

[0104] S203. Use the ORB feature point extraction algorithm to extract feature points from the first original image and the second original image, and generate the first initial feature point set and the second initial feature point set.

[0105] Specifically, the ORB (Oriented FAST and Rotated BRIEF) feature point extraction algorithm is a fast and efficient feature point extraction and description algorithm. It combines the speed advantage of the FAST (Features from Accelerated Segment Test) corner detection algorithm and the efficiency of the BRIEF (Binary Robust Independent Elementary Features) descriptor, and adds the calculation of the feature point direction on this basis, making the feature points have rotational invariance.

[0106] In the embodiment of the present application, by using the ORB feature point extraction algorithm to extract corresponding feature points respectively, feature points can be obtained efficiently and quickly, thereby improving the efficiency of obtaining the target stable image frame.

[0107] S204. Calculate the corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image respectively.

[0108] It should be noted that since the first initial grayscale image and the second initial grayscale image are reduced in step 2 of S202, therefore, after obtaining the initial homography matrix, it is necessary to perform an amplification process on the initial homography matrix based on the corresponding amplification factor so that the initial homography matrix is applicable to the size of the original image. Therefore, after obtaining the initial homography matrix in the embodiment of the present application, the following steps also need to be performed:

[0109] Perform an amplification operation on the initial homography matrix based on a preset amplification factor.

[0110] It should be noted that when performing the magnification process, the magnification factor corresponding to the reduction factor in step 2 under S202 needs to be used. That is, when each of the single-frame images in the initial image sequence is reduced to 1 / 4 of the original single-frame image through the reduction factor, at this time, the initial first homography matrix needs to be magnified to 4 times the original initial first homography matrix through the magnification factor.

[0111] Specifically, similarly, the magnification factor can be represented in matrix form and can be obtained by finding the inverse matrix of the reduction matrix. The obtained magnification factor refers to the following expression:

[0112]

[0113] where scaleH_inv represents the inverse matrix of the reduction matrix, that is, the magnification factor corresponding to the reduction factor, and s can be understood as the magnification multiple.

[0114] S205. Perform Kalman filtering on the initial homography matrix to obtain a predicted homography matrix.

[0115] Since Kalman filtering is generally divided into two main processes, including prediction (Prediction) and correction (Update), also known as time update and measurement update, to estimate the state of the system. In video stabilization, the homography matrix is regarded as the state of the system, and the state transition matrix and the preset process noise covariance matrix are used to predict the state of the next frame, that is, to predict the homography matrix corresponding to the next frame, that is, the predicted homography matrix.

[0116] Therefore, the above step S205 (perform Kalman filtering on the initial homography matrix to obtain a predicted homography matrix) can be refined into the following steps S2051 and S2052:

[0117] S2051. Obtain the estimated value of the predicted homography matrix through the estimated value of the historical predicted homography matrix and the target state transition matrix.

[0118] Specifically, when applying Kalman filtering to the estimation of the homography matrix, the homography matrix is regarded as the state variable in Kalman filtering. According to the target state transition matrix A, combined with the estimated value of the predicted homography matrix at the previous moment, the estimated value of the predicted homography matrix at the current moment is estimated; furthermore, the target state transition matrix can describe the transition relationship of the system state from the previous moment k - 1 to the next moment k. For example, if the target moves at a certain speed in the image, the state transition matrix will take this motion pattern into account to predict the position and shape changes of the next frame.

[0119] When predicting the estimated value of the predicted homography matrix of the next frame, it can be obtained by referring to the following formula:

[0120] x′ k = AX k-1

[0121] where x′ k represents the estimated value of the predicted homography matrix at time k; x k-1 represents the estimated value of the predicted homography matrix at time k - 1. It should be noted that the target state transition matrix A can be obtained according to the following state transition equation:

[0122] x k = A * x k-1 + W k

[0123] where W k represents system noise, which can be understood as various interference factors and uncertainties that cannot be accurately predicted and controlled during the video stabilization process. For example, the interference effects brought by the shooting device, shooting environment, etc.

[0124] It should be noted that when calculating the predicted homography matrix between the initial two adjacent frames of a video to be stabilized, since the estimated value of the predicted homography matrix calculated last time cannot be obtained, and thus the specific value of the target state transition matrix A cannot be solved according to the above state transition equation, the target state transition matrix can be set as the identity matrix during the initial calculation. Then, after the clear target state transition matrix is calculated subsequently, the target state transition matrix is replaced by the clear state transition matrix from the identity matrix.

[0125] In the prediction part, in addition to predicting the estimated value of the predicted homography matrix, the estimated value of the covariance of the predicted homography matrix also needs to be predicted; among them, the covariance corresponding to the predicted homography matrix is also represented by a matrix, that is, the covariance matrix P, which needs to participate in the calculation of the Kalman gain during the correction process at the next moment and affects the correction of the estimated value of the predicted homography matrix. Specifically, the estimated value of the covariance matrix corresponding to the predicted homography matrix is calculated with reference to the following formula:

[0126] p′ k = Ap k-1 A T + Q

[0127] where p′ k represents the estimated value of the covariance matrix at time k, p k-1 represents the covariance matrix at time k - 1, A TThe transpose matrix of the matrix representing the target transfer state. Q represents the covariance matrix of the process noise, which is a quantitative representation of the process noise; it refers to the uncertainty existing during the internal state transfer of the current system. For example, in video stabilization, the random jitter of the camera, small rotational deviations, etc. These factors will make the change of the homography matrix from the current frame to the next frame uncertain. It should be noted that the value of Q can be set according to the actual situation, and this application does not make specific limitations.

[0128] After obtaining the observed state Z of the system k is the measurement or observation result of the system state at time k. The observation matrix H is used to describe how to obtain the observed state from the system state. However, there will also be errors or uncertainties during the observation process, which are called observation noise and are represented by V K .

[0129] S2052. Perform calibration update according to the estimated value of the predicted homography matrix and the initial homography matrix to obtain the predicted homography matrix.

[0130] Specifically, in the above prediction part, it does not depend on new measurement values (that is, the homography matrix calculated by denoising between the current adjacent two frames), and only predicts the future state, and the result will be used as the input of the calibration part.

[0131] Then, in the above steps, the prediction of the predicted homography matrix and the covariance matrix corresponding to the predicted homography matrix is completed. In this step, it is necessary to perform calibration update on them. Before the calibration update, it is necessary to first calculate the Kalman gain according to the following formula:

[0132] K g = p k-1 (Hp k-1 H T + R) -1

[0133] where K g is the Kalman gain, H is the initial homography matrix between the adjacent two frames at the current moment; H t is the transpose matrix of the initial homography matrix; R is the covariance matrix of the observation noise, which can be understood as the influence caused by the uncertain factors that appear when calculating the initial homography matrix. It should be noted that when calculating the Kalman gain, the estimated value p k-1 of the covariance matrix corresponding to the predicted homography matrix at the previous moment is required. Therefore, it is necessary to predict and correct the covariance matrix corresponding to the predicted homography matrix in the prediction and calibration parts to ensure the accuracy of the data.

[0134] Furthermore, the estimated value of the predicted homography matrix and the estimated value of the covariance matrix corresponding to the predicted homography matrix are updated and corrected according to the following formula:

[0135] X k = x′ k + K g (Z k - Hx′ k )

[0136] P k = (I - K g H)p′ k

[0137] where X k represents the predicted homography matrix updated after correction; Z k represents the initial homography matrix corresponding to the current moment after denoising; P k represents the covariance matrix corresponding to the predicted homography matrix updated after correction; I is the identity matrix.

[0138] It should be noted that before processing the video to be stabilized, since the initial information cannot be obtained when calculating the predicted homography matrix of the previous few frames, the estimated value of the predicted homography matrix of the previous moment corresponding to the initial moment and the initial value of its corresponding covariance can be set by the staff themselves.

[0139] S206. Perform weighted calculation in combination with the standard identity matrix to obtain the stabilized homography matrix.

[0140] In the embodiment of the present application, if the predicted homography matrix is directly used to transform the video frames, the video may appear unnatural jitter or distortion. By performing weighted calculation with the standard identity matrix, when there are large errors in the prediction, since the standard identity matrix represents no transformation, the stabilized homography matrix will be closer to maintaining the original state of the image, thereby suppressing the influence of these noises and errors and making the transformation of the video frames smoother.

[0141] Furthermore, the stabilized homography matrix obtained by the above weighted calculation method in the present application can, to a certain extent, suppress the noise or inaccurate estimation in the predicted homography matrix, make the transformation of the video frames more stable and natural, and thus achieve a better video stabilization effect.

[0142] Specifically, the stabilized homography matrix can be obtained by weighted combination of the standard identity matrix with reference to the following formula:

[0143] Hpre = alpha * Hpre + (1 - alpha) * Heye

[0144] Among them, Hpre represents the stability homography matrix; alpha is a weighting coefficient with a value range of [0, 1]. The closer the value is to 1, the better the image stabilization effect; Heye is the identity matrix.

[0145] S207: Geometrically transform the first initial feature point set by using the predicted homography matrix and the stability homography matrix respectively to obtain a first target feature point set and a second target feature point set.

[0146] S208: Perform RANSAC iterative calculation based on the first target feature point set and the second target feature point set to obtain the target image stabilization homography matrix.

[0147] S209: Perform perspective transformation on the first original image based on the target image stabilization homography matrix to generate a target image stabilization frame.

[0148] Specifically, the specific implementation steps of the above step S209 (performing perspective transformation on the first original image based on the target image stabilization homography matrix to generate a target image stabilization frame) include the following S2091 and S2092:

[0149] S2091: Convert the image format of the first original image to the YUV format to generate a first YUV image.

[0150] Specifically, the reason for converting the format of the middle frame in the target image sequence to the YUV image format and then performing perspective transformation is that for the RGBA image obtained from the video to be stabilized, when converted to a YUV image, the image size will be reduced to 60% of the original single-frame image. Therefore, using the YUV image for perspective transformation can save memory space and improve processing speed compared to using the original RGBA image for perspective transformation.

[0151] It should be noted that the YUV image format is mainly used for the transmission and storage of video signals. It divides the color information into a luminance component Y and two chrominance components U and V. The Y component represents the luminance information of the image, which determines the black and white degree of the image. For example, in a black and white image, only the Y component is needed to completely represent the image. The U and V components represent the chrominance information of the color. The U component is usually related to blue, and the V component is usually related to red. They describe the hue and saturation of the color.

[0152] S2092: Perform perspective transformation on the first YUV image based on the target image stabilization homography matrix to obtain a target image stabilization frame.

[0153] It should be noted that when performing perspective transformation on the to-be-stabilized image frame using the target homography matrix, the to-be-stabilized image frame is an image in YUV format. When performing perspective transformation using an image in YUV format, the image size will be reduced to 60% of the original. Therefore, performing perspective transformation using an image in YUV format can save memory space and improve processing speed compared to performing perspective transformation using an RGBA image.

[0154] Based on the same inventive concept, as an implementation of the above method, an embodiment of the present application further provides a video stabilized image frame acquisition device based on Kalman filtering. This embodiment corresponds to the foregoing method embodiment. For the convenience of reading, details of the foregoing method embodiment will not be elaborated one by one in this embodiment. However, it should be clear that a video stabilized image frame acquisition device based on Kalman filtering in this embodiment can correspondingly implement all the content in the foregoing method embodiment.

[0155] An embodiment of the present application provides a video stabilized image frame acquisition device based on Kalman filtering, Figure 3 which is a schematic structural diagram for acquiring a stabilized image of an image in a video, as Figure 3 shown. The stabilized image acquisition device 300 for an image in a video includes:

[0156] A calculation unit 301, configured to calculate a corresponding initial homography matrix according to a first initial feature point set and a second initial feature point set respectively corresponding to a first target grayscale image and a second target grayscale image; the first target grayscale image and the second target grayscale image are obtained based on a first original image and a second original image corresponding to two adjacent video frames in a to-be-stabilized video;

[0157] A filtering unit 302, configured to perform Kalman filtering on the initial homography matrix, obtain a predicted homography matrix, and perform weighted calculation in combination with an identity matrix to obtain a stable homography matrix;

[0158] A transformation unit 303, configured to perform geometric transformation on the first initial feature point set by using the predicted homography matrix and the stable homography matrix respectively to obtain a first target feature point set and a second target feature point set;

[0159] An acquisition unit 304, configured to perform RANSAC iterative calculation based on the first target feature point set and the second target feature point set to obtain a target stable homography matrix;

[0160] A generation unit 305, configured to perform perspective transformation on the first original image based on the target stable homography matrix to generate a target stabilized image frame.

[0161] As an alternative implementation manner of the embodiment of the present application, the calculation unit 301 is specifically configured to analyze the video to be stabilized, and obtain the first original image and the second original image corresponding to two adjacent video frames; convert the image formats of the first original image and the second original image into grayscale images to generate the first target grayscale image and the second target grayscale image; extract feature points from the first original image and the second original image by using the ORB feature point extraction algorithm to generate the first initial feature point set and the second initial feature point set.

[0162] As an alternative implementation manner of the embodiment of the present application, the calculation unit is specifically configured to perform feature point matching on the first target feature point set and the second target feature point set, and perform iterative calculation through the RANSAC algorithm to obtain the initial homography matrix corresponding to the single-frame images of two adjacent frames.

[0163] As an alternative implementation manner of the embodiment of the present application, the calculation unit 301 is specifically configured to convert the image formats of the first original image and the second original image into grayscale images to obtain a first initial grayscale image and a second initial grayscale image; perform a shrinking process on the first initial grayscale image and the second initial grayscale image based on a preset shrinking coefficient to generate the first target grayscale image and the second target grayscale image.

[0164] As an alternative implementation manner of the embodiment of the present application, the calculation unit 301 is further configured to perform an enlarging operation on the initial homography matrix based on a preset enlarging coefficient.

[0165] As an alternative implementation manner of the embodiment of the present application, the filtering unit 302 is specifically configured to obtain an estimated value of the predicted homography matrix by using an estimated value of the historical predicted homography matrix and a target state transition matrix; perform verification and update according to the estimated value of the predicted homography matrix and the initial homography matrix to obtain the predicted homography matrix.

[0166] As an alternative implementation manner of the embodiment of the present application, the generating unit 305 is specifically configured to convert the image format of the first original image into the YUV format to generate a first YUV image; perform a perspective transformation on the first YUV image based on the target image stabilization homography matrix to obtain a target image stabilization frame.

[0167] Based on the same inventive concept, the embodiments of the present disclosure further provide an electronic device. Figure 4 The structural schematic diagram of the electronic device provided by the embodiment of the present disclosure is as Figure 4As shown in the figure, the electronic device provided in this embodiment includes: a memory 401 and a processor 402. The memory 401 is used to store a computer program. The processor 402 is used to execute the audio data processing method provided in the above embodiment when executing the computer program.

[0168] Based on the same inventive concept, an embodiment of the present application also provides a computer-readable storage medium. A computer program is stored on the computer-readable storage medium. When the computer program is executed by a processor, the computing device is enabled to implement the video stabilization frame acquisition method based on Kalman filtering provided in the above embodiment.

[0169] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media that contain computer-usable program code.

[0170] The processor can be a central processing unit (CPU), or can also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc.

[0171] The memory may include non-permanent memory in a computer-readable medium, random access memory (RAM) and / or non-volatile memory in the form of, for example, read-only memory (ROM) or flash RAM. The memory is an example of a computer-readable medium.

[0172] Computer-readable media include both permanent and non-permanent, removable and non-removable storage media. The storage media can implement information storage by any method or technology, and the information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data signals and carrier waves.

[0173] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present application.

Claims

1. A video stabilization frame acquisition method based on Kalman filtering, characterized in that: include: Calculate a corresponding initial homography matrix according to a first initial feature point set and a second initial feature point set corresponding to the first target grayscale image and the second target grayscale image respectively; The first target grayscale image and the second target grayscale image are obtained based on a first original image and a second original image corresponding to two adjacent video frames in the video to be stabilized; Performing Kalman filtering on the initial homography matrix to obtain a predicted homography matrix, and performing weighted calculation in combination with a standard unit matrix to obtain a stabilization homography matrix; Using the prediction homography matrix and the stabilization homography matrix to perform geometric transformation on the first initial feature point set, respectively, to obtain a first target feature point set and a second target feature point set; Perform RANSAC iterative calculation based on the first target feature point set and the second target feature point set to obtain a target image stabilization homography matrix; A transmission transformation is performed on the first original image based on the target image stabilization homography matrix to generate a target image stabilization frame.

2. The method according to claim 1, characterized in that: Before calculating the corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image, the method further includes: Parsing the video to be stabilized to obtain the first original image and the second original image corresponding to two adjacent video frames; Converting the image formats of the first original image and the second original image into grayscale images to generate the first target grayscale image and the second target grayscale image; The ORB feature point extraction algorithm is used to extract feature points from the first original image and the second original image to generate the first initial feature point set and the second initial feature point set.

3. The method according to claim 1, characterized in that The performing RANSAC iterative calculation based on the first target feature point set and the second target feature point set to obtain a target image stabilization homography matrix includes: Feature point matching is performed on the first target feature point set and the second target feature point set, and an iterative calculation is performed using a RANSAC algorithm to obtain an initial homography matrix corresponding to the single-frame images of the two adjacent frames.

4. The method according to claim 1, characterized in that Before calculating the corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image, the method further includes: Convert the image formats of the first original image and the second original image into grayscale images to obtain a first initial grayscale image and a second initial grayscale image; The first initial grayscale image and the second initial grayscale image are reduced based on a preset reduction coefficient to generate the first target grayscale image and the second target grayscale image.

5. The method according to any one of claims 1 to 4, characterized in that: After calculating the corresponding initial homography matrix according to the first initial feature point set and the second initial feature point set corresponding to the first target grayscale image and the second target grayscale image, the method further includes: Based on a preset magnification factor, the initial homography matrix is ​​magnified.

6. The method according to claim 1, characterized in that The performing Kalman filtering on the initial homography matrix to obtain a predicted homography matrix includes: Obtaining an estimated value of the predicted homography matrix through an estimated value of the historical predicted homography matrix and a target state transfer matrix; A verification update is performed according to the estimated value of the predicted homography matrix and the initial homography matrix to obtain the predicted homography matrix.

7. The method according to claim 1, characterized in that The performing a transmission transformation on the first original image based on the target image stabilization homography to generate a target image stabilization frame includes: Converting the image format of the first original image into a YUV format to generate a first YUV image; A transmission transformation is performed on the first YUV image based on the target image stabilization homography matrix to obtain a target image stabilization frame.

8. A video stabilization frame acquisition device based on Kalman filtering, characterized in that: include: A calculation unit, configured to calculate a corresponding initial homography matrix according to a first initial feature point set and a second initial feature point set respectively corresponding to a first target grayscale image and a second target grayscale image; the first target grayscale image and the second target grayscale image are obtained based on a first original image and a second original image corresponding to two adjacent frames of the video to be stabilized; A filtering unit, configured to perform Kalman filtering on the initial homography matrix to obtain a predicted homography matrix, and perform weighted calculation in combination with a standard unit matrix to obtain a stabilization homography matrix; A transformation unit, configured to respectively use the prediction homography matrix and the stabilization homography matrix to perform geometric transformation on the first initial feature point set to obtain a first target feature point set and a second target feature point set; an acquisition unit, configured to perform RANSAC iterative calculation based on the first target feature point set and the second target feature point set to acquire a target image stabilization homography matrix; A generating unit is used to perform a transmission transformation on the first original image based on the target image stabilization homography matrix to generate a target image stabilization frame.

9. An electronic device, characterized in that: include: A memory and a processor, wherein the memory is used to store a computer program; and the processor is used to enable the electronic device to implement the video stabilization frame acquisition method based on Kalman filtering as described in any one of claims 1 to 7 when executing the computer program.

10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a computing device, the computing device implements the video stabilization frame acquisition method based on Kalman filtering as described in any one of claims 1-7.