Video anti-shake method and device, storage medium and processor
By dynamically adjusting the crop rate and correcting posture information of video frames, the insufficient field of view caused by video anti-shake is solved, and the video stability and picture quality are improved, greatly improving the user experience.
Patent Information
- Application Number
- CN202510425550.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-07
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2045-04-07
AI Technical Summary
Existing video anti-shake methods lead to insufficient field of vision and affect user experience.
By obtaining the initial correction attitude information of the current video frame, calculating the initial crop rate, and combining the historical frame information for envelope tracking, dynamically adjusting the target crop rate, thereby adjusting the initial correction attitude information, updating the target device data, and realizing dynamic adjustment of the field of view.
Effectively reduce jitter in video due to handheld shooting and other reasons, enhance the robustness and adaptability of anti-shake processing, improve video stability and picture quality, and improve user experience.
Smart Images

Figure CN119946414A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to a video stabilization method, device, storage medium and processor. Background Art
[0002] With the rapid development of information technology and artificial intelligence, video stabilization is being used more and more widely in various fields. Traditional video stabilization methods use a fixed field of view to ensure shooting stability. For example, when a smartphone turns on video shooting mode, it often uses optical image stabilization (OIS) + electronic image stabilization (EIS) to enhance the stability of the image. In order to ensure a stable shooting effect, the crop ratio is constant in any scene, but it will cause some unnecessary cropping, resulting in a smaller field of view for preview and recorded videos.
[0003] At present, in order to solve the problem of insufficient field of view caused by video stabilization, a variety of solutions have been proposed and applied, mainly including real-time processing and offline processing. Real-time processing generally uses external physical hardware to expand the field of view or enhance stability, for example, using an augmented lens to increase the field of view, or using an external stabilizer for video stabilization, but this method increases cost and complexity, reduces portability, and is limited by the field of view of the device itself. Offline processing is to turn off stabilization and use a stabilizer during shooting, while recording posture sensor data for later adjustment of the field of view. This method relies on the device's ability to support specific functions and export data, has limited applicability, and complex post-processing. It can be seen that the shortcomings of these two methods greatly affect the user experience in various video application scenarios.
[0004] Therefore, facing the problem of insufficient field of view caused by video stabilization, how to improve user experience is a technical problem that needs to be solved urgently. Summary of the invention
[0005] Based on the above problems, the present application provides a video stabilization method, device, storage medium and processor, the purpose of which is to improve the utilization rate of the field of view in the stabilization algorithm and enhance the user experience by dynamically adjusting the field of view.
[0006] The embodiments of the present application disclose the following technical solutions: The first aspect of the present application provides a video anti-shake method, the method comprising: Get the initial correction posture information corresponding to the current video frame; Based on the initial corrected posture information, calculating an initial cropping rate of the current video frame; Performing envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence to obtain a target cropping rate of the current video frame; Based on the target cropping rate, adjusting the initial correction posture information to determine the target correction posture information; Based on the target cropping rate and the target correction posture information, the target device data corresponding to the current video frame is updated to obtain an adjusted current video frame.
[0007] Optionally, the calculating an initial cropping rate of the current video frame based on the initial corrected posture information includes: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; set the preset magnification as the magnification of the kth iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.
[0008] Optionally, performing envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence to obtain a target cropping rate of the current video frame includes: Constructing a cropping rate set based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence; Calculating the envelope of the cropping rate set by using a window sliding technique to obtain an envelope tracking result of the cropping rate set; The envelope tracking result is analyzed to determine a target cropping rate of the current video frame.
[0009] Optionally, adjusting the initial correction posture information based on the target cropping rate to determine the target correction posture information includes: Calculating a target magnification based on the target cropping ratio; Amplify the sampling points included in the initial set according to the target magnification to obtain a second amplified point set; Let the value of k be 1; use the initial corrected posture information as the initial data of the kth iteration; Performing a two-dimensional transformation on the initial data of the k-th iteration to obtain a perspective transformation matrix of the k-th iteration; Performing perspective transformation on the points in the second magnified point set based on the perspective transformation matrix of the k-th iteration to obtain a second transformed point set of the k-th iteration; Determine whether the second transformation point set of the k-th iteration meets the second iteration end condition, and obtain a second determination result; If the second judgment result is no, adjusting the initial data of the k-th iteration to obtain the initial data of the k+1-th iteration, adding 1 to the value of k, and returning to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration; If the second judgment result is yes, the iteration ends, and the initial data of the kth iteration is output as the target correction posture information.
[0010] Optionally, if the second judgment result is no, adjusting the initial data of the k-th iteration to obtain the initial data of the k+1-th iteration, adding 1 to the value of k, and returning to perform a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration, comprises: If the second judgment result is no, adjusting the smoothing coefficient in the initial data of the kth iteration, smoothing the original posture of the target device according to the adjusted smoothing coefficient, and obtaining an adjusted smoothed posture of the target device; the original posture of the target device is calculated based on the target device data corresponding to the current video frame; Calculating initial data for the k+1th iteration based on the original pose and the adjusted smooth pose of the target device; The value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the k-th iteration is returned to obtain a perspective transformation matrix of the k-th iteration.
[0011] Optionally, the obtaining of initial correction posture information corresponding to the current video frame includes: Get the target device data corresponding to the current video frame; Based on the target device data, calculating the original position and posture of the target device; Smoothing the original posture according to a smoothing coefficient to obtain a smoothed posture of the target device; Based on the original posture and the smoothed posture, initial corrected posture information corresponding to the current video frame is calculated.
[0012] A second aspect of the present application provides a video anti-shake device, the device comprising: An information acquisition module, used to obtain the initial correction posture information corresponding to the current video frame; A cropping rate determination module is used to calculate the initial cropping rate of the current video frame based on the initial correction posture information; perform envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames corresponding to each other, to obtain a target cropping rate of the current video frame; A correction posture information determination module, used to adjust the initial correction posture information based on the target cropping rate and determine the target correction posture information; The adjustment module is used to update the target device data corresponding to the current video frame based on the target cropping rate and the target correction posture information to obtain an adjusted current video frame.
[0013] Optionally, the cropping rate determination module is specifically used to: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; set the preset magnification as the magnification of the first iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.
[0014] A third aspect of the present application provides a computer-readable storage medium, in which a computer program is stored. When the program is executed by a processor, the video stabilization method provided in any implementation of the first aspect is implemented.
[0015] A fourth aspect of the present application provides a processor, which is used to run a computer program, and when the program is run, the video stabilization method provided in any implementation of the first aspect is executed.
[0016] Compared with the prior art, this application has the following beneficial effects: The video stabilization method provided by the present application calculates the initial correction posture information and the initial cropping rate, and performs envelope tracking in combination with the historical frame information to adjust the target cropping rate of the current frame. This method can effectively reduce the jitter in the video caused by handheld shooting and other reasons, so that the algorithm can dynamically adapt to different shooting conditions and scene changes, and enhance the robustness and adaptability of the anti-shake processing. After determining the target cropping rate, the initial correction posture information is adjusted according to this cropping rate to determine the target correction posture information, while ensuring the stability of the video, the content of the original picture is retained as much as possible, and the problem of picture loss caused by excessive cropping is avoided. By updating the target device data corresponding to the current video frame, the adjusted current video frame is obtained to achieve the purpose of dynamic field of view adjustment.
[0017] It can be seen that this method can be integrated into embedded devices, and the range of the shooting field of view can be adaptively adjusted according to the amplitude of movement and the frequency of shaking, thus achieving the purpose of dynamically adjusting the field of view. This method improves the utilization rate of the field of view in the anti-shake algorithm, not only ensuring the stability and picture quality of the video, but also greatly improving the user experience in various video application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0019] Figure 1 A flowchart of a video stabilization method provided in an embodiment of the present application; Figure 2This is a rendering of the effect when the black border is not included in the black border detection provided in the embodiment of the present application; Figure 3 This is a rendering of the effect when black edges are included in the black edge detection provided in the embodiment of the present application; Figure 4 The scaling factor and envelope prediction diagram provided in the embodiment of the present application; Figure 5 A schematic diagram of the structure of a video stabilization device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0020] As described above, in order to solve the problem of insufficient field of view caused by video stabilization, many solutions have been proposed and applied, mainly including real-time processing and offline processing. Real-time processing generally uses external physical hardware to expand the field of view or enhance stability, for example, using an augmented lens to increase the field of view, or using an external stabilizer for video stabilization, but this method increases cost and complexity, reduces portability, and is limited by the field of view of the device itself. Offline processing is to turn off stabilization and use a stabilizer during shooting, while recording posture sensor data for later adjustment of the field of view. This method relies on the device's ability to support specific functions and export data, has limited applicability, and complex post-processing. It can be seen that the shortcomings of these two methods greatly affect the user experience in various video application scenarios.
[0021] In view of the above problems, the inventors have proposed a video stabilization method, device, storage medium and processor after research, which obtain the initial correction posture information corresponding to the current video frame; calculate the initial cropping rate of the current video frame based on the initial correction posture information; perform envelope tracking based on the initial cropping rate and the initial cropping rates of multiple consecutive historical video frames corresponding to each other, and obtain the target cropping rate of the current video frame; adjust the initial correction posture information based on the target cropping rate, and determine the target correction posture information; based on the target cropping rate and the target correction posture information, update the target device data corresponding to the current video frame to obtain the adjusted current video frame.
[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of this application.
[0023] See also Figure 1 , which is a flow chart of a video anti-shake method provided by an embodiment of the present application. Figure 1 As shown, the method comprises the following steps: S101. Obtain initial correction posture information corresponding to the current video frame.
[0024] In an achievable implementation, obtaining the initial correction posture information corresponding to the current video frame includes the following steps: S1011. Obtain target device data corresponding to the current video frame.
[0025] Get the data of the current video frame from the target device (such as a camera or mobile phone). This data usually includes the image itself and metadata related to the frame, such as timestamp and sensor data (such as gyroscope and accelerometer data), which is crucial for the subsequent calculation of the camera's pose information.
[0026] The motion data and camera data are obtained through the inertial measurement unit (IMU) and camera of the target device respectively. These data are used to calculate the posture change of the target device. The specific process is as follows: (1) IMU noise reduction, using filtering technology to reduce the impact of IMU noise on the data, thereby improving the accuracy of posture estimation; (2) Timestamp synchronization, in order to ensure that there is an accurate correspondence between the data obtained from the IMU (such as acceleration and angular velocity) and the image frame captured by the camera, timestamp synchronization is required. The synchronization operation usually involves adjusting the clocks of the IMU and the camera, or using software algorithms to compensate for the delay between the clocks of the IMU and the camera. This operation ensures that the IMU data can correctly reflect the motion state at the moment of capturing each frame; (3) Calibrate the coordinate systems of the IMU and the camera to obtain the rotation relationship R IC and translation relationship IC , so that the coordinate systems of the IMU and the camera remain aligned.
[0027] The high-frequency motion data (such as acceleration and angular velocity) provided by the IMU and the visual information of the camera can be used to make more accurate attitude estimation. This approach not only improves the performance limit of a single sensor, but also effectively copes with various complex motion scenes, providing strong support for subsequent applications (such as video stabilization, 3D reconstruction, etc.).
[0028] S1012: Calculate the original position and posture of the target device based on the target device data.
[0029] Based on the acquired target device data, especially the sensor data, the position and orientation of the camera when taking each frame, i.e., the pose information, can be calculated. This process usually involves complex mathematical operations, including but not limited to using IMU data for pose estimation. Among them, the pose information obtained is divided into two types: raw pose and smoothed pose. The raw pose is calculated directly from the sensor data and may contain some high-frequency noise.
[0030] After the previous step, the IMU and coordinate system calibration parameters R are obtained. IC ,t IC Then, through the formula P c =R IC P I +t IC , converted into the original pose of the camera, where P c is the camera's posture, P I It is the attitude of the IMU, and then the original attitude of the camera is obtained by integration and calibration. The integration methods include but are not limited to direct integration, quaternion integration, direction cosine matrix integration, etc. In order to facilitate subsequent smooth operations, the attitude here is expressed by quaternion or Euler angle.
[0031] S1013. Smoothing the original posture according to a smoothing coefficient to obtain a smoothed posture of the target device.
[0032] After obtaining the original pose of the camera, since it expresses the original motion pose of the camera and records the jitter information, in order to obtain stable data, it is necessary to smooth the original pose to obtain a smooth pose. Due to different integration methods, there will be certain differences in the smoothing methods, including but not limited to Kalman filtering, complementary filtering, mean filtering, low-pass filtering, etc. It should be noted that since the subsequent steps require frequent adjustment of the smoothing coefficient, in this embodiment, the integration method adopted supports real-time dynamic configuration of the smoothing coefficient.
[0033] S1014. Calculate initial corrected posture information corresponding to the current video frame based on the original posture and the smoothed posture.
[0034] By calculating the difference between the original pose and the smoothed pose, the initial corrected pose information corresponding to the current video frame is obtained. The corrected pose information usually includes rotation, scaling and translation operations, and the representation of the corrected pose includes but is not limited to quaternions and Euler angles. The acquisition of corrected pose information is a key step, the purpose of which is to compensate for the jitter or instability caused by camera movement during video shooting, thereby ensuring a more stable and smooth video output.
[0035] S102: Calculate an initial cropping rate of the current video frame based on the initial corrected posture information.
[0036] In one possible implementation: S1021. Obtain an image corresponding to the current video frame.
[0037] S1022: Perform equal-interval sampling on edge pixel points of the image according to a preset interval to obtain an initial set consisting of N sampling points.
[0038] S1023: Perform a two-dimensional transformation on the initial corrected posture information to obtain a perspective transformation matrix.
[0039] S1024, setting the value of k to 1; and using the preset magnification as the magnification of the kth iteration.
[0040] The sampling points included in the initial set are enlarged according to the magnification of the k-th iteration to obtain a first enlarged point set of the k-th iteration.
[0041] Perspective transformation is performed on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration.
[0042] It is determined whether the first transformation point set of the kth iteration satisfies the first iteration end condition to obtain a first determination result.
[0043] If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned.
[0044] If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification.
[0045] When determining whether the first transformation point set of the kth iteration satisfies the first iteration end condition, specifically: determining whether the pixel points included in the current video frame are all within the range selected by the points in the first transformation point set. If so, there is no black edge, such as Figure 2 As shown, Figure 2 The pixels contained in the current video frame are all within the range selected by points Pap1, Pap2, Pap3 and Pap4 in the transformation point set; if not, there will be black edges, such as Figure 3 As shown, the pixel points included in the current video frame are not all within the range selected by the points Pap1, Pap2, Pap3 and Pap4 in the transformation point set. Among them, Pap1, Pap2, Pap3 and Pap4 are all points in the transformation point set.
[0046] S1025. Calculate an initial cropping ratio of the current video frame based on the initial magnification ratio.
[0047] Assume that the magnification is b, the cropping ratio is (b-1) / b, and the retention rate is 1-cropping ratio.
[0048] The cropping rate has a value range of (0, 1), and in the embodiment of the present application, the cropping rate is accurate to two decimal places.
[0049] After applying the initial correction posture information for transformation, the edge of the image may exceed the original field of view, resulting in black edges. To solve this problem, the black edge area caused by cropping is identified for the transformed image, and the cropping ratio is dynamically adjusted (for example, using the binary method) based on the black edge detection result to find the minimum cropping ratio without black edges, ensuring that the final output video eliminates unnecessary black edges while retaining the original content as much as possible.
[0050] S103 , performing envelope tracking based on the initial cropping rate and initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence to obtain a target cropping rate of a current video frame.
[0051] The envelope is a function that describes the outer contour or boundary of a signal. The envelope itself extracts the changing trend or contour of the signal in some way, and is used to describe the amplitude change, oscillation, fluctuation, etc. of the signal.
[0052] In one possible implementation: S1031: construct a cropping rate set based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence.
[0053] S1032: Calculate the envelope of the cropping rate set by using a window sliding technique to obtain an envelope tracking result of the cropping rate set.
[0054] The real-time envelope is calculated by using the sliding window technique to calculate the maximum value of a window with a window length of W. Let X be the set of clipping rates, i is the subscript of the set point, then the envelope point E at the corresponding i+W-1 position is i+W-1 The calculation formula is: ; When the envelope point does not correspond to the timestamp of X, an interpolation method is used to obtain the corresponding envelope point and complete the corresponding points of the entire sequence so that the length is consistent with the original set. The interpolation method is not limited here.
[0055] An important role of envelope tracking is to smooth out changes in cropping rates by identifying and following major trends in the cropping rate set, filtering out small fluctuations that may be caused by noise or other non-critical factors.
[0056] S1033: Analyze the envelope tracking result to determine a target cropping rate of the current video frame.
[0057] Based on the envelope tracking results of multiple consecutive historical video frames, the envelope information of the current video frame is predicted and the target cropping rate of the current video frame is determined, which also plays a smoothing role, making the change of the picture field of view natural and smooth. Taking the Kalman prediction method as an example, let the true value of the envelope be Etrue (t), estimated value is , the update equation is: ; In the above formula, is the predicted envelope value. Here, it is assumed that the estimated value at the current time t is equal to the true value at the previous time t-1 multiplied by a coefficient a. The physical model of simulated envelope prediction will be somewhat different in actual use. The update equation is: ; In the above formula, K(T) is the Kalman gain: ; In the above formula, R is the observation noise covariance, To predict the error covariance, by continuously updating the estimated value, the Kalman filter can obtain a smooth and accurate envelope in real-time processing, such as Figure 4 shown. Figure 4 The title is Scaling Factor and Envelope Prediction. The horizontal axis represents time, and the vertical axis represents cropping rate. The light-colored line represents the original data, showing the original fluctuation of cropping rate over time. It can be seen that the original data has obvious high-frequency jitter. The dark line represents the data after envelope smoothing, showing the smooth trend of cropping rate over time. This line is smoother, removes high-frequency jitter, and retains the main trend changes.
[0058] In an achievable implementation manner, after obtaining the predicted envelope, the method further includes: The predicted envelope value is truncated or smoothed.
[0059] Depending on the actual situation, truncation or smoothing can be used alone, or they can be combined. The preset range is usually between 0 and the maximum allowed clipping rate. This method can help limit the predicted value to a reasonable range and reduce the instability caused by data fluctuations.
[0060] Envelope tracking is a filtering technique that helps remove unwanted jitter and makes the final target crop rate smoother. In this way, a crop rate is found that can accommodate the maximum range of motion in the entire video sequence while minimizing image loss.
[0061] S104: Based on the target cropping rate, adjust the initial correction posture information to determine target correction posture information.
[0062] In an optional embodiment: S1041. Calculate a target magnification based on the target cropping ratio.
[0063] S1042: Enlarge the sampling points included in the initial set according to the target magnification to obtain a second enlarged point set.
[0064] The initial set and the initial set in S1022 may be the same set, the second enlarged point set is obtained by enlarging the points in the initial set according to the target magnification, and the first enlarged point set in S1024 changes with the magnification during the iteration process.
[0065] S1043, setting the value of k to 1; using the initial corrected posture information as initial data for the kth iteration.
[0066] Perform a two-dimensional transformation on the initial data of the k-th iteration to obtain a perspective transformation matrix of the k-th iteration.
[0067] Performing perspective transformation on the points in the second magnified point set based on the perspective transformation matrix of the k-th iteration to obtain a second transformed point set of the k-th iteration.
[0068] It is determined whether the second transformation point set of the kth iteration satisfies the second iteration end condition to obtain a second determination result.
[0069] If the second judgment result is no, the initial data of the kth iteration is adjusted to obtain the initial data of the k+1th iteration, the value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the kth iteration is returned to obtain the perspective transformation matrix of the kth iteration.
[0070] If the second judgment result is yes, the iteration ends, and the initial data of the kth iteration is output as the target correction posture information.
[0071] The second transformation point set changes with the correction posture information during the iteration process, and the first transformation point set in S1024 changes with the first magnification point set during the iteration process. The specific process of determining whether the second transformation point set of the k-th iteration satisfies the second iteration end condition is the same as the specific process of determining whether the first transformation point set of the k-th iteration satisfies the first iteration end condition in S1024.
[0072] Apply the target cropping rate to the initial posture transformation and perform black edge detection. If there is a black edge, adjust the initial correction posture until no black edge appears and obtain the target correction posture information. This method ensures that after cropping the video based on the target cropping rate, the video still looks natural and coherent.
[0073] S105 . Based on the target cropping rate and the target correction posture information, update the target device data corresponding to the current video frame to obtain an adjusted current video frame.
[0074] After obtaining the target cropping rate and the target correction posture information, the cropping rate and the correction posture are converted to meet the image correction parameter requirements. Image correction requires three parameters, camera intrinsic parameters, rotation matrix, and new camera intrinsic parameters. The main operation is to fuse the target cropping rate with the camera intrinsic parameters to obtain the new camera intrinsic parameters. For example, the fusion process is as follows: assuming that the cropping rate is m, the original camera intrinsic parameters are K, and the new camera intrinsic parameters are K new ,in: ;
[0075] In the above formula, f x is the focal length of the camera in the x-axis direction of the image coordinate system, reflecting the scaling of the focal length in the horizontal pixel direction, f y is the focal length of the camera in the y-axis direction of the image coordinate system, reflecting the scaling of the focal length in the vertical pixel direction, c x The principal point is in the image coordinate system x The coordinate on the axis, that is, the horizontal pixel position of the intersection of the camera optical axis and the image plane, should theoretically be the center of the image width, c y It is the coordinate of the principal point on the y-axis of the image coordinate system, that is, the vertical pixel position of the intersection of the camera optical axis and the image plane, which should theoretically be the center of the image height.
[0076] New camera intrinsic parameter K new for: ;
[0077] Use K, K new , R completes the stabilization operation of the image, where the new camera intrinsic parameters and rotation matrix R are both dynamic. The new intrinsic parameter matrix realizes dynamic field of view changes, and the rotation matrix realizes anti-shake.
[0078] Based on the target cropping rate and the target correction posture information, the current video frame is subjected to corresponding geometric transformation (such as rotation, scaling and translation). The processed video frame effectively utilizes the target correction posture information to improve the stability of the video, reduces the visual jitter caused by the movement of the target device, and obtains a more visually stable video frame, providing users with a smoother and more comfortable experience.
[0079] The video stabilization method provided in the embodiment of the present application calculates the initial correction posture information and the initial cropping rate, and performs envelope tracking in combination with the historical frame information to adjust the target cropping rate of the current frame. This method can effectively reduce the jitter in the video caused by handheld shooting and other reasons, so that the algorithm can dynamically adapt to different shooting conditions and scene changes, and enhance the robustness and adaptability of the stabilization process. After determining the target cropping rate, the initial correction posture information is adjusted according to this cropping rate to determine the target correction posture information, while ensuring the stability of the video, the content of the original picture is retained as much as possible, and the problem of picture loss caused by excessive cropping is avoided. By updating the target device data corresponding to the current video frame, the adjusted current video frame is obtained to achieve the purpose of dynamic field of view adjustment.
[0080] It can be seen that this method can be integrated into embedded devices, and the range of the shooting field of view can be adaptively adjusted according to the amplitude of movement and the frequency of shaking, thus achieving the purpose of dynamically adjusting the field of view. This method improves the utilization rate of the field of view in the anti-shake algorithm, not only ensuring the stability and picture quality of the video, but also greatly improving the user experience in various video application scenarios.
[0081] Based on the video stabilization method introduced in the above embodiments, the present application also provides a video stabilization device accordingly. Figure 5 Figure 1 is a schematic diagram of the structure of the device. Figure 5 As shown, the video anti-shake device includes: The information acquisition module 501 is used to acquire the initial correction posture information corresponding to the current video frame.
[0082] The cropping rate determination module 502 is used to calculate the initial cropping rate of the current video frame based on the initial correction posture information; perform envelope tracking based on the initial cropping rate and the initial cropping rates of multiple consecutive historical video frames corresponding to each other to obtain the target cropping rate of the current video frame.
[0083] The correction posture information determination module 503 is used to adjust the initial correction posture information based on the target cropping rate and determine the target correction posture information.
[0084] The adjustment module 504 is used to update the target device data corresponding to the current video frame based on the target cropping rate and the target correction posture information to obtain an adjusted current video frame.
[0085] Optionally, the cropping rate determination module is specifically used to: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; set the preset magnification as the magnification of the kth iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.
[0086] Optionally, the clipping determination module is specifically used to: Constructing a cropping rate set based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence; Calculating the envelope of the cropping rate set by using a window sliding technique to obtain an envelope tracking result of the cropping rate set; The envelope tracking result is analyzed to determine a target cropping rate of the current video frame.
[0087] Optionally, the correction posture information determination module is specifically used for: Calculating a target magnification based on the target cropping ratio; Amplify the sampling points included in the initial set according to the target magnification to obtain a second amplified point set; Let the value of k be 1; use the initial corrected posture information as the initial data of the kth iteration; Perform a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration; Performing perspective transformation on the points in the second magnified point set based on the perspective transformation matrix of the k-th iteration to obtain a second transformed point set of the k-th iteration; Determine whether the second transformation point set of the k-th iteration meets the second iteration end condition, and obtain a second determination result; If the second judgment result is no, adjusting the initial data of the k-th iteration to obtain the initial data of the k+1-th iteration, adding 1 to the value of k, and returning to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration; If the second judgment result is yes, the iteration ends, and the initial data of the kth iteration is output as the target correction posture information.
[0088] Optionally, the correction posture information determination module is specifically used for: If the second judgment result is no, adjusting the smoothing coefficient in the initial data of the kth iteration, smoothing the original posture of the target device according to the adjusted smoothing coefficient, and obtaining an adjusted smoothed posture of the target device; the original posture of the target device is calculated based on the target device data corresponding to the current video frame; Calculating initial data for the k+1th iteration based on the original pose and the adjusted smooth pose of the target device; The value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the k-th iteration is returned to obtain a perspective transformation matrix of the k-th iteration.
[0089] Optionally, the information acquisition module is specifically used to: Get the target device data corresponding to the current video frame; Based on the target device data, calculating the original position and posture of the target device; Smoothing the original posture according to a smoothing coefficient to obtain a smoothed posture of the target device; Based on the original posture and the smoothed posture, initial corrected posture information corresponding to the current video frame is calculated.
[0090] In addition, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program is stored. When the program is executed by a processor, a video stabilization method as described in any of the method embodiments is implemented.
[0091] In addition, an embodiment of the present application further provides a processor, which is used to run a computer program, and when the program is run, the video stabilization method introduced in any implementation manner of the aforementioned method embodiment is executed.
[0092] It should be noted that each embodiment in this specification is described in a progressive manner, and the same or similar parts between the embodiments can be referred to each other, and each embodiment focuses on the differences from other embodiments. In particular, for the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The device embodiment described above is merely schematic, in which the unit described as a separate component may or may not be physically separated, and the component prompted as a unit may or may not be a physical unit, that is, it may be located in one place, or it may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the scheme of this embodiment. Ordinary technicians in this field can understand and implement it without paying creative work.
[0093] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any changes or substitutions that can be easily thought of by a person skilled in the art within the technical scope disclosed in the present application should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.
Claims
1. A video stabilization method, characterized in that: include: Get the initial correction posture information corresponding to the current video frame; Based on the initial correction posture information, calculating an initial cropping rate of the current video frame; Performing envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence to obtain a target cropping rate of the current video frame; Based on the target cropping rate, adjusting the initial correction posture information to determine the target correction posture information; Based on the target cropping rate and the target correction posture information, the target device data corresponding to the current video frame is updated to obtain an adjusted current video frame.
2. The method according to claim 1, characterized in that: The calculating the initial cropping rate of the current video frame based on the initial correction posture information includes: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; use the preset magnification as the magnification of the kth iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.
3. The method according to claim 1, characterized in that The step of performing envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of continuous historical video frames in one-to-one correspondence to obtain a target cropping rate of the current video frame includes: Constructing a cropping rate set based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames in one-to-one correspondence; Calculating the envelope of the cropping rate set by using a window sliding technique to obtain an envelope tracking result of the cropping rate set; The envelope tracking result is analyzed to determine a target cropping rate of the current video frame.
4. The method according to claim 1, characterized in that: The adjusting the initial correction posture information based on the target cropping rate to determine the target correction posture information includes: Calculating a target magnification based on the target cropping ratio; Amplify the sampling points included in the initial set according to the target magnification to obtain a second amplified point set; Let the value of k be 1; use the initial corrected posture information as the initial data of the kth iteration; Performing a two-dimensional transformation on the initial data of the k-th iteration to obtain a perspective transformation matrix of the k-th iteration; Performing perspective transformation on the points in the second magnified point set based on the perspective transformation matrix of the k-th iteration to obtain a second transformed point set of the k-th iteration; Determine whether the second transformation point set of the k-th iteration meets the second iteration end condition, and obtain a second determination result; If the second judgment result is no, adjusting the initial data of the k-th iteration to obtain the initial data of the k+1-th iteration, adding 1 to the value of k, and returning to the step of performing a two-dimensional transformation on the initial data of the k-th iteration to obtain the perspective transformation matrix of the k-th iteration; If the second judgment result is yes, the iteration ends, and the initial data of the kth iteration is output as the target correction posture information.
5. The method according to claim 4, characterized in that If the second judgment result is no, the initial data of the kth iteration is adjusted to obtain the initial data of the k+1th iteration, the value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the kth iteration to obtain the perspective transformation matrix of the kth iteration is returned, including: If the second judgment result is no, adjusting the smoothing coefficient in the initial data of the kth iteration, smoothing the original posture of the target device according to the adjusted smoothing coefficient, and obtaining an adjusted smoothed posture of the target device; the original posture of the target device is calculated based on the target device data corresponding to the current video frame; Calculating initial data for the k+1th iteration based on the original pose and the adjusted smooth pose of the target device; The value of k is increased by 1, and the step of performing a two-dimensional transformation on the initial data of the k-th iteration is returned to obtain a perspective transformation matrix of the k-th iteration.
6. The method according to claim 1, characterized in that The obtaining of the initial correction posture information corresponding to the current video frame includes: Get the target device data corresponding to the current video frame; Based on the target device data, calculating the original position and posture of the target device; Smoothing the original posture according to a smoothing coefficient to obtain a smoothed posture of the target device; Based on the original posture and the smoothed posture, initial corrected posture information corresponding to the current video frame is calculated.
7. A video anti-shake device, characterized in that: include: An information acquisition module, used to obtain the initial correction posture information corresponding to the current video frame; A cropping rate determination module is used to calculate the initial cropping rate of the current video frame based on the initial correction posture information; perform envelope tracking based on the initial cropping rate and the initial cropping rates of a plurality of consecutive historical video frames corresponding to each other, to obtain a target cropping rate of the current video frame; A correction posture information determination module, used to adjust the initial correction posture information based on the target cropping rate and determine the target correction posture information; The adjustment module is used to update the target device data corresponding to the current video frame based on the target cropping rate and the target correction posture information to obtain an adjusted current video frame.
8. The device according to claim 7, characterized in that The cropping rate determination module is specifically used for: Get the image corresponding to the current video frame; Sampling edge pixel points of the image at equal intervals according to a preset interval to obtain an initial set consisting of N sampling points; Performing a two-dimensional transformation on the initial correction posture information to obtain a perspective transformation matrix; Let the value of k be 1; Use the preset magnification as the magnification of the first iteration; Amplify the sampling points included in the initial set according to the amplification factor of the k-th iteration to obtain a first amplified point set of the k-th iteration; Performing perspective transformation on the points included in the first magnified point set of the k-th iteration based on the perspective transformation matrix to obtain a first transformed point set of the k-th iteration; Determine whether the first transformation point set of the k-th iteration meets the first iteration end condition, and obtain a first determination result; If the first judgment result is no, the magnification of the k-th iteration is adjusted to obtain the magnification of the k+1-th iteration, the value of k is increased by 1, and the step of magnifying the sampling points included in the initial set according to the magnification of the k-th iteration to obtain the first magnified point set of the k-th iteration is returned; If the first judgment result is yes, the iteration ends, and the magnification of the kth iteration is output as the initial magnification; Based on the initial magnification, an initial cropping ratio of the current video frame is calculated.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the program is executed by a processor, the video stabilization method according to any one of claims 1 to 6 is implemented.
10. A processor, characterized in that: Used to run a computer program, which executes the video stabilization method according to any one of claims 1 to 6 when running.
Citation Information
Patent Citations
Camera lens smoothing method and device and portable terminal
CN110519507A
Video stability augmentation method and device, computer equipment and storage medium
CN110740247A
Video recording method and equipment
CN114339101A
Real-time video anti-shake method and video anti-shake device
CN114697467A
Image processing method and related equipment
CN115131222A