A video anti-shake method, device, storage medium and processor

By dynamically adjusting the crop rate and correcting posture information of video frames, the insufficient field of view caused by video anti-shake is solved, and the video stability and picture quality are improved, which significantly improves the user experience.

CN119946414BActive Publication Date: 2025-06-17ANHUI LISTENAI CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510425550.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-07
Publication Date
2025-06-17
Estimated Expiration
2045-04-07

AI Technical Summary

Technical Problem

Existing video anti-shake methods lead to insufficient field of vision and affect user experience.

Method used

By obtaining the initial correction attitude information of the current video frame, calculating the initial crop rate, and combining the historical frame information for envelope tracking, dynamically adjusting the target crop rate, adjusting the initial correction attitude information to determine the target correction attitude information, and updating the target device data to achieve dynamic adjustment of the field of view.

Benefits of technology

Effectively reduce jitter in video due to handheld shooting and other reasons, enhance the robustness and adaptability of anti-shake processing, improve the utilization rate of field of vision in anti-shake algorithms, and improve user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119946414B_ABST
    Figure CN119946414B_ABST
Patent Text Reader

Abstract

The present application discloses a video anti-shake method, apparatus, storage medium and processor. In this solution, initial correction attitude information corresponding to the current video frame is obtained; based on the initial correction attitude information, an initial cropping rate of the current video frame is calculated; envelope tracking is performed based on the initial cropping rate and the initial cropping rates corresponding to a plurality of consecutive historical video frames to obtain a target cropping rate of the current video frame; based on the target cropping rate, the initial correction attitude information is adjusted to determine target correction attitude information; based on the target cropping rate and the target correction attitude information, target device data corresponding to the current video frame is updated to obtain an adjusted current video frame. Compared with the prior art, there are problems such as low utilization rate of the field of view of the video anti-shake algorithm, poor portability, cumbersome processes, and poor device compatibility. The present application has obvious advantages.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a video stabilization method, device, storage medium and processor. Background Art

[0002] With the rapid development of information technology and artificial intelligence, video stabilization is being used more and more widely in various fields. Traditional video stabilization methods use a fixed field of view to ensure shooting stability. For example, when a smartphone turns on video shooting mode, it often uses optical image stabilization (OIS) + electronic image stabilization (EIS) to enhance the stability of the image. In order to ensure a stable shooting effect, the crop ratio is constant in any scene, but it will cause some unnecessary cropping, resulting in a smaller field of view for preview and recorded videos.

[0003] At present, in order to solve the problem of insufficient field of view caused by video stabilization, a variety of solutions have been proposed and applied, mainly including real-time processing and offline processing. Real-time processing generally uses external physical hardware to expand the field of view or enhance stability, for example, using an augmented lens to increase the field of view, or using an external stabilizer for video stabilization, but this method increases cost and complexity, reduces portability, and is limited by the field of view of the device itself. Offline processing is to turn off stabilization and use a stabilizer during shooting, while recording posture sensor data for later adjustment of the field of view. This method relies on the device's ability to support specific functions and export data, has limited applicability, and complex post-processing. It can be seen that the shortcomings of these two methods greatly affect the user experience in various video application scenarios.

[0004] Therefore, facing the problem of insufficient field of view caused by video stabilization, how to improve user experience is a technical problem that needs to be solved urgently. Summary of the invention

[0005] Based on the above problems, the present application provides a video stabilization method, device, storage medium and processor, the purpose of which is to improve the utilization rate of the field of view in the stabilization algorithm and enhance the user experience by dynamically adjusting the field of view.

[0006] The embodiments of the present application disclose the following technical solutions:

[0007] The first aspect of the present application provides a video anti-shake method, the method comprising:

[0008] Get the initial correction posture information corresponding to the current video frame;

[0009] Based on the initial corrected posture information, calculating an initial cropping rate of the current video frame;

[0010] Perform envelope tracking based on the initial cropping rate corresponding one-to-one to the initial cropping rate of multiple consecutive historical video frames, to obtain the target cropping rate of the current video frame;

[0011] Based on the target cropping rate, adjust the initial corrected pose information to determine the target corrected pose information;

[0012] Based on the target cropping rate and the target corrected pose information, update the target device data corresponding to the current video frame to obtain the adjusted current video frame.

[0013] Optionally, the calculating the initial cropping rate of the current video frame based on the initial corrected pose information includes:

[0014] Obtain the image corresponding to the current video frame;

[0015] Perform equidistant sampling on the edge pixel points of the image at a preset interval to obtain an initial set composed of N sampling points;

[0016] Perform two-dimensional transformation on the initial corrected pose information to obtain a perspective transformation matrix;

[0017] Let the value of k be 1; use the preset magnification factor as the magnification factor for the k-th iteration;

[0018] Magnify the sampling points included in the initial set according to the magnification factor for the k-th iteration to obtain the first magnified point set for the k-th iteration;

[0019] Perform perspective transformation on the points included in the first magnified point set for the k-th iteration based on the perspective transformation matrix to obtain the first transformed point set for the k-th iteration;

[0020] Judge whether the first transformed point set for the k-th iteration satisfies the first iteration end condition to obtain the first judgment result;

[0021] If the first judgment result is negative, adjust the magnification factor for the k-th iteration to obtain the magnification factor for the (k + 1)-th iteration, let the value of k increase by 1, and return to the step of magnifying the sampling points included in the initial set according to the magnification factor for the k-th iteration to obtain the first magnified point set for the k-th iteration;

[0022] If the first judgment result is positive, end the iteration and output the magnification factor for the k-th iteration as the initial magnification factor;

[0023] Based on the initial magnification factor, calculate the initial cropping rate of the current video frame.

[0024] Optionally, performing envelope tracking based on the initial cropping rate and the initial cropping rates corresponding to multiple consecutive historical video frames one by one to obtain the target cropping rate of the current video frame, including:

[0025] Based on the initial cropping rate and the initial cropping rates corresponding to multiple consecutive historical video frames one by one, construct a cropping rate set;

[0026] Using the window sliding technique, calculate the envelope of the cropping rate set to obtain the envelope tracking result of the cropping rate set;

[0027] Analyze the envelope tracking result to determine the target cropping rate of the current video frame.

[0028] Optionally, based on the target cropping rate, adjusting the initial correction attitude information to determine the target correction attitude information, including:

[0029] Based on the target cropping rate, calculate the target magnification ratio;

[0030] Magnify the sampling points included in the initial set according to the target magnification ratio to obtain a second magnified point set;

[0031] Let the value of k be 1; use the initial correction attitude information as the initial data for the k-th iteration;

[0032] Perform a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration;

[0033] Based on the perspective transformation matrix of the k-th iteration, perform a perspective transformation on the points in the second magnified point set to obtain the second transformed point set of the k-th iteration;

[0034] Judge whether the second transformed point set of the k-th iteration satisfies the second iteration end condition to obtain a second judgment result;

[0035] If the second judgment result is no, adjust the initial data of the k-th iteration to obtain the initial data of the (k + 1)-th iteration, let the value of k increase by 1, and return to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration;

[0036] If the second judgment result is yes, the iteration ends, and output the initial data of the k-th iteration as the target correction attitude information.

[0037] Optionally, if the second judgment result is negative, adjust the initial data of the k-th iteration to obtain the initial data of the (k + 1)-th iteration, increment the value of k by 1, and return to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration, including:

[0038] If the second judgment result is negative, adjust the smoothing coefficient in the initial data of the k-th iteration, and perform smoothing processing on the original pose of the target device according to the adjusted smoothing coefficient to obtain the smoothed pose of the target device; the original pose of the target device is calculated based on the target device data corresponding to the acquired current video frame;

[0039] Based on the original pose and the smoothed pose of the adjusted target device, calculate the initial data of the (k + 1)-th iteration;

[0040] Increment the value of k by 1, and return to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration.

[0041] Optionally, the obtaining the initial correction pose information corresponding to the current video frame includes:

[0042] Obtain the target device data corresponding to the current video frame;

[0043] Based on the target device data, calculate the original pose of the target device;

[0044] Perform smoothing processing on the original pose according to the smoothing coefficient to obtain the smoothed pose of the target device;

[0045] Based on the original pose and the smoothed pose, calculate the initial correction pose information corresponding to the current video frame.

[0046] The second aspect of the present application provides a video anti-shake device, and the device includes:

[0047] An information acquisition module, configured to acquire initial correction pose information corresponding to a current video frame;

[0048] A cropping rate determination module, configured to calculate an initial cropping rate of the current video frame based on the initial correction pose information; perform envelope tracking on the initial cropping rate and the initial cropping rates corresponding to a plurality of consecutive historical video frames to obtain a target cropping rate of the current video frame;

[0049] A correction pose information determination module, configured to adjust the initial correction pose information based on the target cropping rate to determine target correction pose information;

[0050] An adjustment module, configured to update the target device data corresponding to the current video frame based on the target cropping rate and the target correction pose information, so as to obtain an adjusted current video frame.

[0051] Optionally, the cropping rate determination module is specifically configured to:

[0052] Obtain an image corresponding to the current video frame;

[0053] Perform equidistant sampling on the edge pixel points of the image at a preset interval to obtain an initial set composed of N sampling points;

[0054] Perform a two-dimensional transformation on the initial correction pose information to obtain a perspective transformation matrix;

[0055] Let the value of k be 1; use the preset magnification factor as the magnification factor for the first iteration;

[0056] Magnify the sampling points included in the initial set according to the magnification factor of the k-th iteration to obtain a first magnified point set of the k-th iteration;

[0057] Perform perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration;

[0058] Determine whether the first transformed point set of the k-th iteration satisfies the first iteration end condition to obtain a first judgment result;

[0059] If the first judgment result is negative, adjust the magnification factor of the k-th iteration to obtain the magnification factor of the (k + 1)-th iteration, let the value of k increase by 1, and return to the step of magnifying the sampling points included in the initial set according to the magnification factor of the k-th iteration to obtain the first magnified point set of the k-th iteration;

[0060] If the first judgment result is positive, the iteration ends, and output the magnification factor of the k-th iteration as the initial magnification factor;

[0061] Based on the initial magnification factor, calculate the initial cropping rate of the current video frame.

[0062] A third aspect of the present application provides a computer-readable storage medium, in which a computer program is stored, and when the program is run by a processor, it implements the video anti-shake method provided in any implementation manner of the first aspect.

[0063] A fourth aspect of the present application provides a processor, which is used to run a computer program, and when the program runs, it executes the video anti-shake method provided in any implementation manner of the first aspect.

[0064] Compared with the prior art, the present application has the following beneficial effects:

[0065] The video anti-shake method provided by the present application calculates the initial correction attitude information and the initial cropping ratio, and adjusts the target cropping ratio of the current frame through envelope tracking in combination with historical frame information. This method can effectively reduce the jitter caused by hand-held shooting in the video, enabling the algorithm to dynamically adapt to different shooting conditions and scene changes, and enhancing the robustness and adaptability of the anti-shake processing. After determining the target cropping ratio, the initial correction attitude information is adjusted according to this cropping ratio to determine the target correction attitude information, which can retain the content of the original picture as much as possible while ensuring the stability of the video, and avoid the problem of picture loss caused by excessive cropping. By updating the target device data corresponding to the current video frame, the adjusted current video frame is obtained to achieve the purpose of dynamic field of view adjustment.

[0066] It can be seen that this method can be integrated into an embedded device, and adaptively adjust the range of the shooting field of view according to the amplitude of the movement and the frequency of the jitter, achieving the purpose of dynamically adjusting the field of view. This method improves the utilization rate of the field of view in the anti-shake algorithm, not only ensures the stability of the video and the picture quality, but also greatly enhances the user experience in various video application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0067] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the following drawings are only some embodiments of the present application. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0068] Figure 1 It is a flowchart of a video anti-shake method provided by an embodiment of the present application;

[0069] Figure 2 It is an effect diagram when there is no black edge in the black edge detection provided by an embodiment of the present application;

[0070] Figure 3 It is an effect diagram when there is a black edge in the black edge detection provided by an embodiment of the present application;

[0071] Figure 4 It is a scaling coefficient and envelope prediction diagram provided by an embodiment of the present application;

[0072] Figure 5 It is a structural schematic diagram of a video anti-shake device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0073] As described above, in order to solve the problem of insufficient field of view caused by video anti-shake, various solutions have been proposed and applied, mainly including two categories: real-time processing and offline processing. Real-time processing generally uses external physical hardware to expand the field of view or enhance stability. For example, an augmented lens is used to increase the field of view, or an external stabilizer is used for video anti-shake. However, this method increases the cost and complexity, reduces portability, and is limited by the field of view of the device itself. Offline processing is to turn off the anti-shake and use a stabilizer during shooting, while recording the attitude sensor data for later field of view adjustment. This method depends on the device's support for specific functions and the ability to export data, with limited applicability and complex post-processing. It can be seen that the deficiencies in these two types of methods greatly affect the user experience in various video application scenarios.

[0074] In view of the above problems, through research, the inventor has proposed a video anti-shake method, device, storage medium, and processor, which obtain the initial correction attitude information corresponding to the current video frame; calculate the initial cropping rate of the current video frame based on the initial correction attitude information; perform envelope tracking based on the initial cropping rate and the initial cropping rates corresponding to multiple consecutive historical video frames to obtain the target cropping rate of the current video frame; adjust the initial correction attitude information based on the target cropping rate to determine the target correction attitude information; and update the target device data corresponding to the current video frame based on the target cropping rate and the target correction attitude information to obtain the adjusted current video frame.

[0075] In order to enable those skilled in the art to better understand the solution of this application, the following will clearly and completely describe the technical solutions in the embodiments of this application with reference to the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. Based on the embodiments in this application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the scope of protection of this application.

[0076] See Figure 1 , which is a flowchart of a video anti-shake method provided by an embodiment of this application. As Figure 1 shown, the method includes the following steps:

[0077] S101. Obtain the initial correction attitude information corresponding to the current video frame.

[0078] In an implementable embodiment, obtaining the initial correction attitude information corresponding to the current video frame includes the following steps:

[0079] S1011. Obtain the target device data corresponding to the current video frame.

[0080] Obtain the data of the current video frame from a target device (such as a camera or a mobile phone). This data usually includes the image itself and the metadata associated with this frame, such as timestamps, sensor data (such as gyroscope and accelerometer data), which are crucial for subsequent calculations of the camera's pose information.

[0081] Obtain motion data and camera data through the Inertial Measurement Unit (IMU) and the camera of the target device respectively. These data are used to calculate the attitude change of the target device's motion. The specific process is as follows: (1) IMU noise reduction, using filtering techniques to reduce the impact of IMU noise on the data, thereby improving the accuracy of attitude estimation; (2) Timestamp synchronization. To ensure an accurate correspondence between the data obtained from the IMU (such as acceleration and angular velocity) and the image frames captured by the camera, timestamp synchronization is required. The synchronization operation usually involves adjusting the clocks of the IMU and the camera, or using software algorithms to compensate for the delay between the clocks of the IMU and the camera. This operation ensures that the IMU data can correctly reflect the motion state at the moment of each frame capture; (3) Calibrate the coordinate systems of the IMU and the camera to obtain the rotation relationship R IC and the translation relationship t IC to align the coordinate systems of the IMU and the camera.

[0082] Through the high-frequency motion data (such as acceleration and angular velocity) provided by the IMU and the visual information of the camera, more accurate attitude estimation can be performed. This method not only improves the performance limitations of individual sensors but also effectively deals with various complex motion scenarios, providing strong support for subsequent applications (such as video stabilization, 3D reconstruction, etc.).

[0083] S1012. Calculate the original pose of the target device based on the data of the target device.

[0084] Based on the obtained data of the target device, especially sensor data, the position and orientation of the camera at the time of each frame capture, that is, the pose information, can be calculated. This process usually involves complex mathematical operations, including but not limited to using IMU data for attitude estimation. Among them, the obtained pose information is divided into two types: original pose and smoothed pose. The original pose is directly calculated from sensor data and may contain some high-frequency noise.

[0085] After obtaining the IMU and coordinate system calibration parameters R IC , t IC , through the formula P c = R IC P I + t IC , convert it into the original pose of the camera, where P c is the attitude of the camera, PI It is the attitude of the IMU, and then the original pose of the camera is obtained through integration and calibration. The integration methods include but are not limited to direct integration method, quaternion integration method, direction cosine matrix integration method, etc. For the convenience of subsequent smoothing operations, the attitude here is represented by quaternions or Euler angles.

[0086] S1013. Smooth the original pose according to the smoothing coefficient to obtain the smoothed pose of the target device.

[0087] After obtaining the original pose of the camera, since it represents the original motion attitude of the camera and records the jitter information, in order to obtain stable data, it is necessary to smooth the original attitude to obtain the smoothed pose. Due to different integration methods, the smoothing methods will be somewhat different, including but not limited to Kalman filtering, complementary filtering, mean filtering, low-pass filtering, etc. It should be noted that since the smoothing coefficient needs to be adjusted frequently in the subsequent steps, in this embodiment, the adopted integration method supports real-time dynamic configuration of the smoothing coefficient.

[0088] S1014. Based on the original pose and the smoothed pose, calculate the initial correction attitude information corresponding to the current video frame.

[0089] By calculating the difference between the original pose and the smoothed pose, the initial correction attitude information corresponding to the current video frame is obtained. The correction attitude information usually includes rotation, scaling, and translation operations. The representation methods of the correction attitude include but are not limited to the forms of quaternions and Euler angles. Obtaining the correction attitude information is a key step, aiming to compensate for the jitter or instability caused by the camera movement during video shooting, so as to ensure that the video output is more stable and smooth.

[0090] S102. Based on the initial correction attitude information, calculate the initial cropping ratio of the current video frame.

[0091] In an implementable embodiment:

[0092] S1021. Obtain the image corresponding to the current video frame.

[0093] S1022. Sample the edge pixel points of the image at equal intervals according to a preset spacing to obtain an initial set composed of N sampling points.

[0094] S1023. Perform a two-dimensional transformation on the initial correction attitude information to obtain a perspective transformation matrix.

[0095] S1024. Let the value of k be 1; use the preset magnification factor as the magnification factor for the k-th iteration.

[0096] Enlarge the sampling points included in the initial set according to the magnification factor of the k-th iteration to obtain the first enlarged point set for the k-th iteration.

[0097] Perform perspective transformation on the points included in the first set of magnified points in the k-th iteration based on the perspective transformation matrix to obtain the first set of transformed points in the k-th iteration.

[0098] Determine whether the first set of transformed points in the k-th iteration satisfies the first iteration end condition to obtain a first judgment result.

[0099] If the first judgment result is no, adjust the magnification factor in the k-th iteration to obtain the magnification factor in the (k + 1)-th iteration, increment the value of k by 1, and return to the step of magnifying the sampling points included in the initial set according to the magnification factor in the k-th iteration to obtain the first set of magnified points in the k-th iteration.

[0100] If the first judgment result is yes, the iteration ends, and the magnification factor in the k-th iteration is output as the initial magnification factor.

[0101] When determining whether the first set of transformed points in the k-th iteration satisfies the first iteration end condition, specifically: determine whether the pixel points included in the current video frame are all within the range framed by the points in the first set of transformed points. If so, there is no black edge, as Figure 2 shown, Figure 2 in which the pixel points included in the current video frame are all within the range framed by the points Pap1, Pap2, Pap3, and Pap4 in the set of transformed points; if not, there is a black edge, as Figure 3 shown, the pixel points included in the current video frame are not all within the range framed by the points Pap1, Pap2, Pap3, and Pap4 in the set of transformed points. Among them, Pap1, Pap2, Pap3, and Pap4 are all points in the set of transformed points.

[0102] S1025. Calculate the initial cropping rate of the current video frame based on the initial magnification factor.

[0103] Assume the magnification factor is b, the cropping rate is (b - 1) / b, and the retention rate is 1 - the cropping rate.

[0104] The value range of the cropping rate is (0, 1), and in the embodiments of the present application, the cropping rate is accurate to two decimal places.

[0105] After performing transformation using the initial correction pose information, it may occur that the image edge exceeds the original field of view, resulting in the generation of black edges. To solve this problem, for the transformed image, identify the black edge area caused by cropping, and based on the black edge detection result, dynamically adjust the cropping rate (for example, using the dichotomy method) to find the minimum cropping rate without black edges, ensuring that the finally output video not only eliminates unnecessary black edges but also retains the original content as much as possible.

[0106] S103. Perform envelope tracking based on the initial cropping rate and the initial cropping rates corresponding to multiple consecutive historical video frames one by one, to obtain the target cropping rate of the current video frame.

[0107] An envelope is a function that describes the outer enclosing contour or boundary of a signal. The envelope itself extracts the trend or contour of the signal in a certain way, and is used to describe the amplitude change, oscillation, fluctuation, etc. of the signal.

[0108] In an implementable embodiment:

[0109] S1031. Construct a cropping rate set based on the initial cropping rate and the initial cropping rates corresponding to multiple consecutive historical video frames one by one.

[0110] S1032. Use the window sliding technique to calculate the envelope of the cropping rate set, to obtain the envelope tracking result of the cropping rate set.

[0111] Among them, the calculation method of the real-time envelope is to calculate the maximum value of a window with a window length of W by using the sliding window technique. Let X be the set composed of cropping rates, and i be the subscript of the set point. Then the envelope point E at the corresponding i + W - 1 position i+W-1 has the following calculation formula:

[0112] ;

[0113] When the envelope point cannot correspond to the timestamp of X, the interpolation method is used to obtain the corresponding envelope point, to complete the corresponding points of the entire sequence, so that it is consistent with the original set in length. The interpolation method is not limited here.

[0114] An important role of envelope tracking is to smooth the change of the cropping rate. By identifying and following the main trend in the cropping rate set, small fluctuations caused by noise or other non-critical factors can be filtered out.

[0115] S1033. Analyze the envelope tracking result to determine the target cropping rate of the current video frame.

[0116] According to the envelope tracking results of multiple consecutive historical video frames, predict the envelope information of the current video frame, and determine the target cropping rate of the current video frame, which also plays a role in smoothing, making the change of the picture field of view natural and smooth. Taking the Kalman prediction method as an example, let the true value of the envelope be E true (t), and the estimated value be , and the update equation is:

[0117] ;

[0118] In the above formula, is the predicted envelope value. Here, it is assumed that the estimated value at the current time t is equal to the true value at the previous time t-1 multiplied by a coefficient a, simulating the physical model of envelope prediction. There will be certain differences in actual use. The update equation is as follows:

[0119] ;

[0120] In the above formula, K(T) is the Kalman gain:

[0121] ;

[0122] In the above formula, R is the observation noise covariance, is the prediction error covariance. By continuously updating the estimated value, the Kalman filter can obtain a smooth and accurate envelope during real-time processing, as shown in Figure 4 shown. Figure 4 The title of Figure 4 is Scaling Coefficient and Envelope Prediction. The horizontal axis represents time, and the vertical axis represents the cropping rate. The light-colored line represents the original data, showing the original fluctuations of the cropping rate over time. It can be seen that there are obvious high-frequency jitters in the original data. The dark-colored line represents the data after envelope smoothing, showing the smooth trend of the cropping rate over time. This line is smoother, removing the high-frequency jitters and retaining the main trend changes.

[0123] In an implementable embodiment, after obtaining the predicted envelope, it further includes:

[0124] Performing truncation processing or smoothing processing on the predicted envelope value.

[0125] According to the actual situation, truncation or smoothing processing can be used alone, or both can be used in combination. The preset range is usually between 0 and the maximum allowable cropping rate. This method can help limit the predicted value within a reasonable range and reduce the instability caused by data fluctuations.

[0126] Envelope tracking is a filtering technique that can help eliminate unnecessary jitters, making the final obtained target cropping rate more stable. In this way, a cropping rate that can adapt to the maximum motion range in the entire video sequence is found while minimizing picture loss.

[0127] S104. Based on the target cropping rate, adjust the initial correction attitude information to determine the target correction attitude information.

[0128] In an optional embodiment:

[0129] S1041. Based on the target cropping rate, calculate the target magnification factor.

[0130] S1042. Magnify the sampling points included in the initial set according to the target magnification ratio to obtain a second magnified point set.

[0131] Among them, the initial set and the initial set in S1022 can be the same set. The second magnified point set is obtained by magnifying the points in the initial set according to the target magnification ratio, while the first magnified point set in S1024 changes with the change of the magnification ratio during the iteration process.

[0132] S1043. Let the value of k be 1; use the initial correction attitude information as the initial data for the k-th iteration.

[0133] Perform a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration.

[0134] Perform a perspective transformation on the points in the second magnified point set based on the perspective transformation matrix of the k-th iteration to obtain the second transformed point set of the k-th iteration.

[0135] Judge whether the second transformed point set of the k-th iteration meets the second iteration end condition to obtain a second judgment result.

[0136] If the second judgment result is no, adjust the initial data of the k-th iteration to obtain the initial data of the (k + 1)-th iteration, increase the value of k by 1, and return to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration.

[0137] If the second judgment result is yes, the iteration ends, and output the initial data of the k-th iteration as the target correction attitude information.

[0138] Among them, the second transformed point set changes with the change of the correction attitude information during the iteration process, while the first transformed point set in S1024 changes with the change of the first magnified point set during the iteration process. The specific process of judging whether the second transformed point set of the k-th iteration meets the second iteration end condition is the same as that of judging whether the first transformed point set of the k-th iteration meets the first iteration end condition in S1024.

[0139] Apply the target cropping ratio to the initial attitude transformation and perform black edge detection. If there are black edges, adjust the initial correction attitude until no black edges appear to obtain the target correction attitude information. This method ensures that the video still looks natural and coherent after cropping the video based on the target cropping ratio.

[0140] S105. Update the target device data corresponding to the current video frame based on the target cropping ratio and the target correction attitude information to obtain the adjusted current video frame.

[0141] After obtaining the target cropping rate and the target correction pose information, the cropping rate and the correction pose are transformed to meet the requirements of the image correction parameters. Image correction requires three parameters: the camera internal parameter, the rotation matrix, and the new camera internal parameter. Among them, the main operation is to fuse the target cropping rate with the camera internal parameter to obtain the new camera internal parameter. Exemplarily, the fusion process is as follows. Let the cropping rate be m, the original camera internal parameter be K, and the new camera internal parameter be K new , where:

[0142] ;

[0143] In the above formula, f x is the focal length of the camera in the x-axis direction of the image coordinate system, which reflects the scaling of the focal length in the horizontal pixel direction. f y is the focal length of the camera in the y-axis direction of the image coordinate system, which reflects the scaling of the focal length in the vertical pixel direction. c x is the coordinate of the principal point on the x axis of the image coordinate system, that is, the horizontal pixel position of the intersection point of the camera optical axis and the image plane. Theoretically, it should be the center of the image width. c y is the coordinate of the principal point on the y-axis of the image coordinate system, that is, the vertical pixel position of the intersection point of the camera optical axis and the image plane. Theoretically, it should be the center of the image height.

[0144] The new camera internal parameter K new is:

[0145] ;

[0146] Use K, K new , and R to complete the stabilization operation of the image. Among them, the new camera internal parameter and the rotation matrix R are both dynamic. The new internal parameter matrix realizes the dynamic field of view change, and the rotation matrix realizes the anti-shake.

[0147] Based on the target cropping rate and the target correction pose information, the above-mentioned corresponding geometric transformations (such as rotation, scaling, and translation) are performed on the current video frame. The processed video frame effectively utilizes the target correction pose information to improve the stability of the video, reduces the visual jitter caused by the movement of the target device, and obtains a more stable video frame with better visual effects, providing a smoother and more comfortable experience for users.

[0148] The video anti-shake method provided by the embodiments of the present application calculates the initial correction attitude information and the initial cropping ratio, and adjusts the target cropping ratio of the current frame through envelope tracking in combination with historical frame information. This method can effectively reduce the jitter caused by handheld shooting and other reasons in the video, enabling the algorithm to dynamically adapt to different shooting conditions and scene changes, and enhancing the robustness and adaptability of the anti-shake processing. After determining the target cropping ratio, the initial correction attitude information is adjusted according to this cropping ratio to determine the target correction attitude information, so as to retain the content of the original picture as much as possible while ensuring the stability of the video, and avoid the problem of picture loss caused by excessive cropping. By updating the target device data corresponding to the current video frame, the adjusted current video frame is obtained to achieve the purpose of dynamic field of view adjustment.

[0149] It can be seen from this that this method can be integrated into an embedded device, and adaptively adjust the range of the shooting field of view according to the amplitude of the movement and the frequency of the jitter, achieving the purpose of dynamically adjusting the field of view. This method improves the utilization rate of the field of view in the anti-shake algorithm, not only ensuring the stability and picture quality of the video, but also greatly enhancing the user experience in various video application scenarios.

[0150] Based on the video anti-shake method introduced in the foregoing embodiments, correspondingly, the present application also provides a video anti-shake device. Figure 5 The following is a schematic structural diagram of the device. As Figure 5 shown, the video anti-shake device includes:

[0151] An information acquisition module 501, configured to acquire the initial correction attitude information corresponding to the current video frame.

[0152] A cropping ratio determination module 502, configured to calculate the initial cropping ratio of the current video frame based on the initial correction attitude information; perform envelope tracking based on the initial cropping ratio and the initial cropping ratios corresponding to a plurality of consecutive historical video frames to obtain the target cropping ratio of the current video frame.

[0153] A correction attitude information determination module 503, configured to adjust the initial correction attitude information based on the target cropping ratio to determine the target correction attitude information.

[0154] An adjustment module 504, configured to update the target device data corresponding to the current video frame based on the target cropping ratio and the target correction attitude information to obtain the adjusted current video frame.

[0155] Optionally, the cropping ratio determination module is specifically configured to:

[0156] Acquire the image corresponding to the current video frame;

[0157] Sample the edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set composed of N sampling points;

[0158] Perform a two-dimensional transformation on the initial correction attitude information to obtain a perspective transformation matrix;

[0159] Let the value of k be 1; use the preset magnification factor as the magnification factor for the k-th iteration;

[0160] Magnify the sampling points included in the initial set according to the magnification factor of the k-th iteration to obtain a first magnified point set for the k-th iteration;

[0161] Perform a perspective transformation on the points included in the first magnified point set for the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set for the k-th iteration;

[0162] Determine whether the first transformed point set for the k-th iteration satisfies the first iteration end condition to obtain a first judgment result;

[0163] If the first judgment result is no, adjust the magnification factor of the k-th iteration to obtain the magnification factor for the (k + 1)-th iteration, let the value of k increase by 1, and return to the step of magnifying the sampling points included in the initial set according to the magnification factor of the k-th iteration to obtain a first magnified point set for the k-th iteration;

[0164] If the first judgment result is yes, the iteration ends, and output the magnification factor of the k-th iteration as the initial magnification factor;

[0165] Based on the initial magnification factor, calculate the initial cropping rate of the current video frame.

[0166] Optionally, the cropping determination module is specifically configured to:

[0167] Construct a cropping rate set based on the initial cropping rate and the initial cropping rates corresponding to multiple consecutive historical video frames;

[0168] Use the window sliding technique to calculate the envelope of the cropping rate set to obtain the envelope tracking result of the cropping rate set;

[0169] Analyze the envelope tracking result to determine the target cropping rate of the current video frame.

[0170] Optionally, the correction attitude information determination module is specifically configured to:

[0171] Calculate a target magnification factor based on the target cropping rate;

[0172] Magnify the sampling points included in the initial set according to the target magnification factor to obtain a second magnified point set;

[0173] Let the value of k be 1; use the initial correction attitude information as the initial data for the k-th iteration;

[0174] Perform a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration;

[0175] Perform a perspective transformation on the points in the second enlarged point set based on the perspective transformation matrix of the k-th iteration to obtain the second transformed point set of the k-th iteration;

[0176] Determine whether the second transformed point set of the k-th iteration satisfies the second iteration end condition to obtain a second judgment result;

[0177] If the second judgment result is no, adjust the initial data of the k-th iteration to obtain the initial data of the (k + 1)-th iteration, increment the value of k by 1, and return to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration;

[0178] If the second judgment result is yes, end the iteration and output the initial data of the k-th iteration as the target correction attitude information.

[0179] Optionally, the correction attitude information determination module is specifically configured to:

[0180] If the second judgment result is no, adjust the smoothing coefficient in the initial data of the k-th iteration, and perform smoothing processing on the original pose of the target device according to the adjusted smoothing coefficient to obtain the smoothed pose of the adjusted target device; the original pose of the target device is calculated based on the target device data corresponding to the acquired current video frame;

[0181] Calculate the initial data of the (k + 1)-th iteration based on the original pose and the smoothed pose of the adjusted target device;

[0182] Increment the value of k by 1, and return to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration.

[0183] Optionally, the information acquisition module is specifically configured to:

[0184] Acquire the target device data corresponding to the current video frame;

[0185] Calculate the original pose of the target device based on the target device data;

[0186] Perform smoothing processing on the original pose according to the smoothing coefficient to obtain the smoothed pose of the target device;

[0187] Calculate the initial corrected pose information corresponding to the current video frame based on the original pose and the smoothed pose.

[0188] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the program is run by a processor, the video anti-shake method introduced in any manner of the method embodiment is implemented.

[0189] In addition, an embodiment of the present application further provides a processor, which is used to run a computer program. When the program runs, it executes the video anti-shake method introduced in any implementation manner of the foregoing method embodiment.

[0190] It should be noted that the various embodiments in this specification are described in a progressive manner. The same or similar parts between the various embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, refer to the partial description of the method embodiment. The device embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components referred to as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0191] As described above, this is only a specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present application should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A video stabilization method, characterized in that: include: Get the initial correction posture information corresponding to the current video frame; Based on the initial correction posture information, calculating an initial cropping rate of the current video frame; Performing envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence to obtain a target cropping rate of the current video frame; Based on the target cropping rate, adjusting the initial correction posture information to determine the target correction posture information; Based on the target cropping rate and the target correction posture information, updating the target device data corresponding to the current video frame to obtain an adjusted current video frame; The calculating the initial cropping rate of the current video frame based on the initial correction posture information includes: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; use the preset magnification as the magnification of the kth iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.

2. The method according to claim 1, characterized in that The step of performing envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of continuous historical video frames in one-to-one correspondence to obtain a target cropping rate of the current video frame includes: Constructing a cropping rate set based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence; Calculating the envelope of the cropping rate set by using a window sliding technique to obtain an envelope tracking result of the cropping rate set; The envelope tracking result is analyzed to determine a target cropping rate of the current video frame.

3. The method according to claim 1, characterized in that The adjusting the initial correction posture information based on the target cropping rate to determine the target correction posture information includes: Calculating a target magnification based on the target cropping ratio; Amplify the sampling points included in the initial set according to the target magnification to obtain a second amplified point set; Let the value of k be 1; use the initial corrected posture information as the initial data of the kth iteration; Performing a two-dimensional transformation on the initial data of the k-th iteration to obtain a perspective transformation matrix of the k-th iteration; Performing perspective transformation on the points in the second magnified point set based on the perspective transformation matrix of the k-th iteration to obtain a second transformed point set of the k-th iteration; Determine whether the second transformation point set of the k-th iteration meets the second iteration end condition, and obtain a second determination result; If the second judgment result is no, adjusting the initial data of the k-th iteration to obtain the initial data of the k+1-th iteration, adding 1 to the value of k, and returning to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration; If the second judgment result is yes, the iteration ends, and the initial data of the kth iteration is output as the target correction posture information.

4. The method according to claim 3, characterized in that: If the second judgment result is no, the initial data of the kth iteration is adjusted to obtain the initial data of the k+1th iteration, the value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the kth iteration to obtain the perspective transformation matrix of the kth iteration is returned, including: If the second judgment result is no, adjusting the smoothing coefficient in the initial data of the kth iteration, smoothing the original posture of the target device according to the adjusted smoothing coefficient, and obtaining an adjusted smoothed posture of the target device; the original posture of the target device is calculated based on the target device data corresponding to the current video frame; Calculating initial data for the k+1th iteration based on the original pose and the adjusted smooth pose of the target device; The value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the k-th iteration is returned to obtain a perspective transformation matrix of the k-th iteration.

5. The method according to claim 1, characterized in that The obtaining of the initial correction posture information corresponding to the current video frame includes: Get the target device data corresponding to the current video frame; Based on the target device data, calculating the original position and posture of the target device; Smoothing the original posture according to a smoothing coefficient to obtain a smoothed posture of the target device; Based on the original posture and the smoothed posture, initial corrected posture information corresponding to the current video frame is calculated.

6. A video anti-shake device, characterized in that: include: An information acquisition module, used to obtain the initial correction posture information corresponding to the current video frame; A cropping rate determination module is used to calculate the initial cropping rate of the current video frame based on the initial correction posture information; perform envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames corresponding to each other, to obtain a target cropping rate of the current video frame; A correction posture information determination module, used to adjust the initial correction posture information based on the target cropping rate and determine the target correction posture information; An adjustment module, configured to update target device data corresponding to a current video frame based on the target cropping rate and the target correction posture information to obtain an adjusted current video frame; The calculating the initial cropping rate of the current video frame based on the initial correction posture information includes: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; use the preset magnification as the magnification of the kth iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.

7. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the program is executed by a processor, the video stabilization method according to any one of claims 1 to 5 is implemented.

8. A processor, characterized in that: Used to run a computer program, which executes the video stabilization method according to any one of claims 1 to 5 when running.

Citation Information

Patent Citations

  • Video stability augmentation method and device, computer equipment and storage medium

    CN110740247A

  • Video jitter processing method and device, electronic equipment and computer readable storage medium

    CN118646953A