Aerial video jitter removal processing method and device, electronic equipment and storage medium

By smoothing the holographic matrix in aerial videos and perspective transformation, the error accumulation and environmental interference problems of the aerial video debounce processing method in the prior art are solved, and a more accurate and stable video debounce effect is achieved.

CN120050521APending Publication Date: 2025-05-27HANVON CORP
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510104775.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-22
Publication Date
2025-05-27

AI Technical Summary

Technical Problem

In the prior art, the aerial video debounce processing method based on gyroscopes has problems of error accumulation and external environment interference, resulting in unstable video debounce effect.

Method used

By estimating pixel points motion based on adjacent video image frames in the video to be processed, a homography matrix is ​​obtained, and the homography matrix is ​​smoothed, and finally perspective transformation is performed based on the smooth homography matrix to generate a debounced image frame.

Benefits of technology

It improves the accuracy and stability of video debounce processing, reduces error accumulation compared with traditional methods, and enhances the visual stability of video.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120050521A_ABST
    Figure CN120050521A_ABST
Patent Text Reader

Abstract

The invention discloses an aerial video jitter removal processing method and device and electronic equipment, and belongs to the technical field of video image processing. The method comprises the steps of performing pixel point motion estimation on a to-be-processed video based on adjacent video image frames in the to-be-processed video, and obtaining homography matrixes respectively corresponding to specified video image frames in the to-be-processed video; performing smooth processing on the homography matrix corresponding to the current video image frame to obtain a smooth homography matrix; performing perspective transformation on the current video image frame based on the smooth homography matrix to obtain a shake-free image frame of the current video image frame; and based on each de-jitter image frame, generating the to-be-processed video after de-jitter processing. According to the method, the homography matrix describing the inter-frame motion of the feature points is generated and smoothed, and then the motion restoration of the video image frames is performed based on the smoothed homography matrix, so that compared with the traditional video jitter removal processing based on parameters acquired by a gyroscope arranged on an aircraft, the error is smaller.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of video image processing, and particularly to a method, an apparatus, an electronic device, and a computer-readable storage medium for de-shaking an aerial video of a flapping-wing aircraft. Background Art

[0002] As is well known, in the field of aerial video processing, the aerial video of a flapping-wing aircraft has more severe jitter. In the prior art, the methods for anti-shake processing of aerial videos mainly include: rigidly connecting a gyroscope to the video acquisition device of the aircraft, collecting jitter parameters during the flight of the aircraft through the gyroscope, where the jitter parameters are the angular displacements calculated by the gyroscope period, calculating correction parameters according to the output angular displacement parameters, and then correcting the video frame images collected during the flight of the aircraft through the correction parameters, so as to keep the picture of the video frame images at a predetermined position to dynamically eliminate part of the video jitter generated during the flight. Since the gyroscope outputs angular velocity data in a period, by integrating the angular velocity data, the angular change of the flight motion can be calculated, and the attitude angle of the aircraft can be deduced therefrom. Then, the position and motion trajectory of the aircraft are calculated using the attitude angle, and the compensation displacement is determined according to the motion trajectory. Among them, the compensation displacement is the correction parameter of the video.

[0003] However, due to the error accumulation nature of attitude angle calculation, the gyroscope is easily affected by external environmental factors such as vibration, temperature change, and geomagnetic field change, resulting in continuous increase in the calculated position and trajectory deviation. In addition, the accuracy of the position and trajectory calculated by the gyroscope is also affected by the integration accuracy, resulting in the calculated correction parameters becoming more and more inaccurate over time.

[0004] It can be seen that the method for de-shaking the video collected by traditional aircrafts, especially flapping-wing aircrafts, in the prior art still needs to be improved. Summary of the Invention

[0005] The embodiments of the present application provide a method and an apparatus for de-shaking an aerial video, which helps to improve the video de-shaking effect and enhance the visual stability of the video after de-shaking processing.

[0006] In a first aspect, the embodiments of the present application provide a method for de-shaking an aerial video, including:

[0007] Performing pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtaining a homography matrix corresponding to each specified video image frame in the video to be processed;

[0008] Taking each video image frame in the video to be processed as the current video image frame in turn, and performing running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix;

[0009] Perform a perspective transformation on the current video image frame based on the smooth homography matrix to obtain a de-shaken image frame of the current video image frame;

[0010] Generate the processed video after de-shaking based on the de-shaken image frames of the video image frames in the video to be processed.

[0011] In a second aspect, an embodiment of the present application provides an aerial video de-shaking processing device, including:

[0012] A homography matrix acquisition module, configured to perform pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and acquire the homography matrices respectively corresponding to the specified video image frames in the video to be processed;

[0013] A matrix smoothing processing module, configured to sequentially use each video image frame in the video to be processed as the current video image frame, and perform running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smooth homography matrix;

[0014] An image frame de-shaking processing module, configured to perform a perspective transformation on the current video image frame based on the smooth homography matrix to obtain a de-shaken image frame of the current video image frame;

[0015] A de-shaken video generation module, configured to generate the processed video after de-shaking based on the de-shaken image frames of the video image frames in the video to be processed.

[0016] In a third aspect, an embodiment of the present application further discloses an electronic device, including a memory, a processor, and a computer program stored on the memory and executable on the processor, where the processor implements the aerial video de-shaking processing method described in the embodiment of the present application when executing the computer program.

[0017] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the aerial video de-shaking processing method disclosed in the embodiment of the present application are implemented.

[0018] The aerial video anti-shake processing method disclosed in the embodiments of the present application performs pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtains the homography matrices corresponding to the specified video image frames in the video to be processed; performs running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix; performs perspective transformation on the current video image frame based on the smoothed homography matrix to obtain the anti-shake image frame of the current video image frame; generates the video to be processed after anti-shake processing based on the anti-shake image frames of each video image frame in the video to be processed. By generating a homography matrix that describes the inter-frame motion of feature points, performing smoothing processing on the homography matrix, and then performing motion repair on the video image frames based on the smoothed homography matrix to output the video image frames after anti-shake processing, this method has a smaller error and a more accurate and stable video anti-shake processing effect compared with the traditional video anti-shake processing based on the parameters collected by the gyroscope set on the aircraft.

[0019] The above description is only an overview of the technical solution of the present application. In order to be able to understand the technical means of the present application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features and advantages of the present application more obvious and understandable, the following specifically illustrates the specific embodiments of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] To make the objectives, technical solutions, and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Apparently, the described embodiments are some but not all of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts shall fall within the protection scope of the present application.

[0021] Figure 1 is the flowchart of the aerial video anti-shake processing method disclosed in the embodiments of the present application;

[0022] Figure 2 is the schematic diagram of the current video image frame and the previous video image frame in the aerial video anti-shake processing method disclosed in the embodiments of the present application;

[0023] Figure 3 is the aerial video anti-shake processing method disclosed in the embodiments of the present application for Figure 2 the anti-shake processing effect schematic diagram of the current video image frame in;

[0024] Figure 4 is the schematic diagram of the structure of the aerial video anti-shake processing device disclosed in the embodiments of the present application;

[0025] Figure 5A block diagram of an electronic device for performing the method according to the present application is schematically shown; and

[0026] Figure 6 A storage unit for holding or carrying program code for implementing the method according to the present application is schematically shown. Detailed implementation manners

[0027] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are some, rather than all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the protection scope of the present application.

[0028] Referring to Figure 1 , a method for processing aerial video anti-shake disclosed in the embodiments of the present application includes: step 110 to step 140.

[0029] Step 110, perform pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtain the homography matrices respectively corresponding to the specified video image frames in the video to be processed.

[0030] The method for processing aerial video anti-shake disclosed in the embodiments of the present application can be applied to various scenarios. For example, the video to be processed can be a video collected by a flying vehicle such as a drone, or a video collected by other intelligent terminals.

[0031] Taking the video collected by a flapping-wing aircraft during flight as an example for anti-shake processing, it is first necessary to accurately evaluate the motion that causes video shake, and then it is possible to better perform video shake repair on the video collected during the motion for the corresponding motion. In some embodiments of the present application, adjacent video image frames of the video to be processed can be captured by traversing all video image frames of the video to be processed to estimate the inter-frame motion, so as to perform motion correction.

[0032] In some alternative embodiments, the performing pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtaining the homography matrices respectively corresponding to the specified video image frames in the video to be processed includes: for the current video image frame of adjacent video image frames in the video to be processed, determine the corresponding feature points in the current video image frame and the previous video image frame of the current video image frame by performing feature point detection on the current video image frame and the previous video image frame of the current video image frame; use the sparse optical flow method to perform motion estimation on the corresponding feature points to obtain the homography matrix corresponding to the current video image frame.

[0033] For example, for the current video image frame, on the premise of knowing the previous video image frame, by detecting the corresponding pixel points in the previous video image frame and the current video image frame, and based on the position changes of the detected pixel points in these two adjacent video image frames, motion estimation is performed to obtain the video jitter corresponding to the current video image frame, which is recorded as a homography matrix. Among them, the pixel points used for motion estimation can be pixel points that represent obvious features in the video image frame, which are denoted as "feature points" in the embodiments of the present application.

[0034] Optionally, performing feature point detection on the current video image frame and the previous video image frame of the current video image frame to determine the corresponding feature points in the current video image frame and the previous video image frame includes: performing corner detection on the current video image frame and the previous video image frame of the current video image frame to determine the corresponding feature points in the current video image frame and the previous video image frame. The corner detection method can be used to extract obvious features of the video image frame for estimating the translational motion and rotational motion corresponding to the video image frame. Among them, the translational motion part includes: the displacement in the X-axis direction and the displacement in the Y-axis direction in the image coordinate system of the video image frame, and the rotation is the image rotation angle.

[0035] The homography transformation can be simply understood as being used to describe the position mapping relationship between an object in the world coordinate system and the pixel coordinate system, and the corresponding transformation matrix is called the homography matrix. In the embodiments of the present application, after detecting the feature points in the current video image frame and the previous video image frame, the corresponding relationship between the feature points can be found according to a certain matching rule, and then according to the corresponding relationship, a motion model representing the motion transformation of the pixels (i.e., feature points) in the video image frame is estimated. This motion model can be expressed as a homography matrix. Among them, the values in the homography matrix include components such as displacements and rotation angles representing motion.

[0036] When performing motion estimation, it is not necessary to obtain the motion of each pixel point in the video image frame. Therefore, in the embodiments of the present application, the Lucas-Kanade optical flow method (an optical flow estimation method based on feature point matching) is used for feature tracking. Optical flow is the instantaneous velocity of the pixels of a moving object in space on the observation imaging plane. It is a method that uses the change of pixels in the time domain in the image sequence and the correlation between adjacent video image frames to find the corresponding relationship between the previous video image frame and the current video image frame, so as to calculate the motion information of the object between adjacent video image frames. In the implementation of the present application, the sparse optical flow method used does not need to detect and extract all the feature points in the video image frame, which can appropriately improve the speed of video image anti-shake processing.

[0037] The basic principle of the sparse optical flow method is described below.

[0038] Assume that at time t, a pixel on the image is located at (x, y) and its brightness is I(x, y, t). At time t+dt, the pixel moves to a new position (x+dx, y+dy). Based on the assumption of constant brightness, we can conclude that:

[0039] I(x,y,t)=I(x+dx,y+dy,t+dt); (Equation 1)

[0040] Since the motion of the pixels is assumed to be small, the Taylor series expansion can be used to approximate the expression on the right side of the equation:

[0041]

[0042] Combining equations 1 and 2 above and ignoring higher-order infinitesimal terms, we can obtain:

[0043]

[0044] Transforming equation 3, we can further obtain equation 4:

[0045]

[0046] In equation 4, dx / dt and dy / dt are the speed of the pixel in the x and y directions, i.e., the optical flow. By solving this equation, we can get the movement of the pixel between two adjacent video image frames.

[0047] By comparing the position changes of feature points in the two previous and next video image frames, the optical flow, that is, the motion vector of the feature point, that is, the coordinate change, can be calculated.

[0048] In an embodiment of the present application, the optical flow method interface in the prior art can be called to obtain the coordinates of the feature points in the current video image frame. For example, the coordinates of the feature points extracted from the previous video image frame, the current video image frame, and the previous video image frame are used as input, the optical flow method interface is called, and the previous frame image, the next frame image, and the coordinates of the feature points extracted from the previous frame image in the current video image frame returned by the interface are obtained, and the coordinates of the corresponding feature points in the next frame image and some state numbers are returned. If a feature point corresponding to the feature point in the previous video image frame is found in the current video image frame, the state corresponding to the feature point is 1, otherwise it is 0. In the next step of iteratively detecting feature points, the currently detected feature points are passed as feature points in the previous frame of the video image. According to this method, the positions of the feature points in any two adjacent video image frames can be detected.

[0049] After obtaining the positions of feature points in adjacent video image frames, the coordinate changes of the feature points in the adjacent video image frames, that is, the optical flow, which is also the motion vector of the feature points, can be further obtained. In some embodiments of the present application, when determining the transformation from the image coordinate system in the previous video image frame to the image coordinate system in the current video image frame according to the optical flow (i.e., coordinate changes) of the adjacent video image frames, a perspective transformation with a relatively high degree of freedom is adopted to obtain a homography matrix.

[0050] Performing a perspective transformation on an image projects the image from one geometric plane to another. The perspective transformation ensures that points on the same straight line remain on the same straight line, but no longer guarantees parallelism. The reason is that after a two-dimensional image undergoes a three-dimensional transformation and then is mapped to another two-dimensional space, the original two-dimensional space of the two-dimensional image is different from the mapped two-dimensional space. Using a perspective transformation to determine the coordinate changes of feature points in video image frames helps to improve the video anti-shake effect. The perspective transformation has a higher degree of freedom than the affine transformation and contains more motion parameters, describing the motion more accurately.

[0051] In practical applications, through video stability analysis of the videos obtained by performing coordinate changes and video anti-shake processing using the affine transformation and the perspective transformation respectively, the video stability scores corresponding to the two transformation methods are obtained. The average video stability score obtained using the perspective transformation is 15 higher than that obtained using the affine transformation.

[0052] The perspective transformation is based on matrix operations. Through matrix operations, the pixel correspondence between the old and new images can be quickly found. The perspective transformation is described using the following homogeneous coordinate formula:

[0053] Among them, on the left side of the equation is the transformed homogeneous coordinate, that is, X and Y are the abscissa and ordinate of the pixel points on the transformed video image. The first n - 1 terms are equal to the two-dimensional plane coordinates, and x and y are the abscissa and ordinate of the pixel points on the video image before transformation. On the right side of the equation, H is a 3x3 square matrix, which is the perspective transformation matrix, that is, the homography matrix described in the embodiments of the present application. From the above formula, it can be seen that after determining the homography matrix, the mapping relationship between the video image frames before and after the transformation can be determined. Further, according to the coordinates of the feature points in the video image frames before and after the transformation, the motion between adjacent image frames can be obtained. Corner detection can extract the coordinates of the image points in the previous frame, and the optical flow method can output the coordinates of the image in the next frame. According to the above formula, the homography matrix H can be obtained.

[0054] In some alternative embodiments, to improve the speed of motion estimation, the video image frames are compressed, and feature point detection and motion estimation are performed based on the compressed video image frames. Among them, performing pixel point motion estimation on the to-be-processed video based on adjacent video image frames in the to-be-processed video to obtain the homography matrices respectively corresponding to the specified video image frames in the to-be-processed video includes: performing compression processing on adjacent video image frames in the to-be-processed video to obtain adjacent compressed video image frames; performing pixel point motion estimation based on the adjacent compressed video image frames to obtain the homography matrices respectively corresponding to the compressed video image frames; performing scale restoration processing on the homography matrices respectively corresponding to the compressed video image frames to obtain homography matrices with the same size as the original video image frames in the to-be-processed video, as the homography matrices respectively corresponding to the corresponding video image frames in the to-be-processed video.

[0055] For example, the height and width of the current video image frame and the previous video image frame of the current video image frame can be compressed to 1 / 4 of the original size to obtain two compressed video image frames. Then, pixel point motion estimation is performed based on these two compressed video image frames to obtain the homography matrix corresponding to the compressed video image frame of the current video image frame. Then, scale restoration processing is performed on the homography matrix corresponding to the compressed video image frame to obtain the homography matrix corresponding to the original size, as the homography matrix corresponding to the current video image frame, which is used for coordinate mapping of the original image.

[0056] Specifically, for example, the formula S = R -1 S_tR can be used for scale restoration of the homography matrix, where R is the ratio matrix and S_t is the homography matrix obtained after compression of the video image frame. By performing scale compression on the video image frame and obtaining the homography matrix based on the compressed video image frame for video de-shaking processing, the speed of video de-shaking processing can be significantly improved.

[0057] Optionally, performing pixel point motion estimation based on the adjacent compressed video image frames to obtain the homography matrices respectively corresponding to the compressed video image frames includes: for the current compressed video image frame in the adjacent compressed video image frames, determining the corresponding feature points in the current compressed video image frame and the previous compressed video image frame by performing feature point detection on the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame; performing motion estimation on the corresponding feature points using the sparse optical flow method to obtain the homography matrix corresponding to the current compressed video image frame.

[0058] Optionally, determining corresponding feature points in the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame by performing feature point detection on the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame includes: performing corner detection on the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame to determine corresponding feature points in the current compressed video image frame and the previous compressed video image frame.

[0059] For the specific implementation of performing corner detection on the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame to determine corresponding feature points in the current compressed video image frame and the previous compressed video image frame, refer to the implementation of determining corresponding feature points in adjacent video image frames in the foregoing embodiments, which will not be elaborated here.

[0060] For the specific implementation of performing motion estimation on the corresponding feature points by using the sparse optical flow method to obtain the homography matrix corresponding to the current compressed video image frame, refer to the specific implementation of performing motion estimation on the corresponding feature points by using the sparse optical flow method to obtain the homography matrix corresponding to the current video image frame described in the foregoing embodiments, which will not be elaborated here.

[0061] Step 120: Sequentially use each video image frame in the video to be processed as the current video image frame, and perform running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix.

[0062] Next, perform smoothing filtering on the estimated motion information, that is, the homography matrix, to obtain a new smoothed motion trajectory. That is, perform smoothing processing on the homography matrix to filter out the unnecessary inter-frame motion represented in the homography matrix, where the unnecessary inter-frame motion is the jitter of the video frames.

[0063] Taking the homography matrix as the following 3rd-order matrix as an example:

[0064]

[0065] Among them, H is the homography matrix, A represents the linear transformation parameter, T represents the translation transformation parameter, = [v 1 , v 2 contains the relationship of the intersection points of the edges after transformation, and s is a scaling factor related to V T = [v 1 , v 2 . Generally, s = 1 is obtained through normalization. The purpose of motion smoothing is to make the errors of the parameters A and T in the homography matrix H smaller, that is, the motion jitter amplitude is smaller, that is, the unnecessary inter-frame motion is less.

[0066] When making the video picture more stable by removing unwanted extra movements, traditional and relatively common motion smoothing methods will smooth the entire motion trajectory or the complete cumulative transformation chain. In the embodiments of the present application, only local displacements are smoothed to achieve motion smoothing. When applying smoothing to the complete transformation chain, the cascading of the original complete transformation chain and the smoothed transformation chain often generates cumulative errors, resulting in inaccurate smoothing results. On the contrary, in the present application, only the homography matrix of the current video image frame is smoothed. This method of locally smoothing the displacement from the current video image frame to adjacent video image frames directly calculates a new transformation matrix from the original video image frame to the corresponding motion compensation frame (i.e., the transformed video image frame) only using the transformation matrices of adjacent video image frames, that is, directly smoothing the transformation matrices of adjacent video image frames.

[0067] Optionally, the performing motion smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix includes: performing motion smoothing processing on the homography matrix corresponding to the current video image frame according to the homography matrices corresponding to each video image frame within the image frame window corresponding to the current video image frame to obtain a smoothed homography matrix. That is, the motion information (i.e., the homography matrix) corresponding to the current video image frame can be smoothed by combining the inter-frame motion information (i.e., the homography matrices corresponding to each video image frame) of multiple video image frames before and after the current video image frame.

[0068] Taking the current video image frame represented as I t as an example, the index N t of the video image frames within the image frame window corresponding to the current video image frame t is represented as N t ={j|t - k ≤ j ≤ t + k}, and the local displacement corresponding to the current video image frame I s can be used to calculate the position of each adjacent video image frame I t relative to the current video image frame I t to the motion compensation frame I′ t The perspective transformation S t is calculated according to the following formula:

[0069] where i represents the index number of the video image frame, G(k) is the Gaussian kernel function, k represents the range of adjacent video image frames, and the value of k can be determined according to experimental results. For example, the value of k can be 15; the "*" operator represents the convolution operation. Among them, the local displacement It is calculated based on the coordinates output by the optical flow method and the coordinates obtained by corner detection. For example, the coordinates of the feature points calculated by the optical flow method are subtracted from the coordinates of the feature points detected by corner detection, and the obtained coordinate difference is used as the local displacement.

[0070] Optionally, performing motion smoothing processing on the homography matrix corresponding to the current video image frame according to the homography matrices corresponding to each video image frame within the image frame window corresponding to the current video image frame to obtain a smoothed homography matrix includes: multiplying the homography matrices corresponding to each video image frame within the image frame window corresponding to the current video image frame to obtain a first matrix representing the cumulative motion corresponding to a continuous plurality of video image frames; calculating a second matrix representing the average motion of a single video image frame according to the first matrix; multiplying the second matrix by the inverse matrix of the homography matrix corresponding to the current image frame to obtain a third matrix, and using the third matrix as the smoothed homography matrix corresponding to the current video image frame.

[0071] The following is an example to illustrate the process of calculating a smoothed homography matrix. First, obtain the input matrix array, which includes the homography matrix corresponding to the j-th video image frame, and the homography matrices corresponding to the first K and the last K video image frames before and after the j-th video image frame respectively. Take the average of this matrix array to obtain an average transformation matrix. Then, calculate the inverse matrix of the homography matrix at a specified index position (such as the one corresponding to the j-th video image frame) in the matrix array, and multiply this average transformation matrix by the inverse matrix. Finally, obtain and return a smoothed homography matrix M, and this smoothed homography matrix M is the smoothed homography matrix corresponding to the j-th video image frame. For the smoothing process of the transformation matrix between adjacent video image frames, the corresponding algorithm experiment reproduction method is: after extracting and tracking the optical flow feature coordinates, obtain the 3rd-order homography matrix corresponding to the current video image frame; then, take a sequence of video image frames where the current video image frame is located, calculate the homography matrices of two adjacent video image frames in the sequence to obtain a sequence of homography matrices; then, start performing matrix multiplication from the first homography matrix in the sequence window, and after multiplication, obtain a first matrix. Then, take the average of the first matrix to obtain a second matrix, and multiply the second matrix by the inverse matrix of the homography matrix corresponding to the current video image frame to obtain a third matrix, that is, obtain the smoothed homography matrix.

[0072] As described above, the homography matrix represents the position difference of feature points between adjacent video image frames. The cumulative multiplication of the homography matrices corresponding to adjacent video image frames represents the integration of the position differences of the feature points, and the cumulative multiplication result represents the imaging position of the feature points. Further, by taking the average of the imaging positions, the average imaging positions of the feature points in each video image frame within a period centered on the current video image frame can be obtained. Multiplying the average imaging position by the inverse of the original position is equivalent to subtracting the original position, thus converting from the original position to the smooth position.

[0073] Using the above method, motion estimation is performed on each video image frame in the video image frame to be processed, and the smooth homography matrix corresponding to each video image frame in the video image frame to be processed is obtained respectively.

[0074] After smoothing processing, a new homography transformation matrix, that is, the smooth homography matrix, can be obtained. Based on the smooth homography matrix, the video image frame can be reconstructed through perspective transformation.

[0075] Step 130, perform perspective transformation on the current video image frame based on the smooth homography matrix to obtain the de-shaking image frame of the current video image frame.

[0076] In specific implementation, the current video image frame and the smooth homography matrix can be used as the input of the perspective transformation, and the output image of the perspective transformation is the de-shaking image frame after performing de-shaking processing on the current video image frame. Taking Figure 2 the previous video image frame (a) and the current video image frame (b) shown in as an example, perform perspective transformation on the current video image frame (b) based on the smooth homography matrix calculated in the foregoing steps, and the de-shaking image frame as shown in Figure 3 can be obtained. For the specific implementation manner of performing perspective transformation on the current video image frame based on the smooth homography matrix, reference can be made to the prior art, and it will not be elaborated in the embodiments of the present application.

[0077] Step 140, generate the video to be processed after de-shaking processing based on the de-shaking image frames of each video image frame in the video to be processed.

[0078] For a given video to be processed, perform motion estimation and perspective transformation on each video image frame in the video to be processed in sequence to obtain the de-shaking image frame of each video image frame. Finally, synthesize the de-shaking image frames frame by frame to obtain the stable video corresponding to the video to be processed.

[0079] In summary, the aerial video anti-shake processing method disclosed in the embodiments of the present application estimates the pixel motion of the video to be processed based on adjacent video image frames in the video to be processed, and obtains the homography matrices corresponding to the specified video image frames in the video to be processed; performs running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix; performs perspective transformation on the current video image frame based on the smoothed homography matrix to obtain the anti-shake image frame of the current video image frame; and generates the video to be processed after anti-shake processing based on the anti-shake image frames of the video image frames in the video to be processed. By generating a homography matrix that describes the inter-frame motion of feature points, performing smoothing processing on the homography matrix, and then performing motion restoration on the video image frames based on the smoothed homography matrix to output the video image frames after anti-shake processing, this method has a more accurate and stable video anti-shake processing effect compared with the traditional method of performing video anti-shake processing based on the parameters collected by the gyroscope set on the aircraft.

[0080] Further, the displacement description information of the feature points of adjacent video image frames is used, that is, the homography matrix locally smooths the displacement from the current video image frame to the adjacent video image frame, and directly calculates the motion compensation from the original video image frame to the corresponding one, effectively reducing the error accumulation generated in the motion smoothing process.

[0081] In addition, before tracking features in the current video image frame using the optical flow method, the video image frame is first compressed, which can improve the feature matching speed of the optical flow method.

[0082] The aerial video anti-shake processing method disclosed in the embodiments of the present application restores the motion of the current video image frame with jitter, suppresses the position change between adjacent video image frames, and makes the visual effect between adjacent video image frames smoother and the video image motion smoother during video playback.

[0083] Refer to Figure 4 , the embodiments of the present application also disclose an aerial video anti-shake processing device, and the device includes:

[0084] A homography matrix acquisition module 410, configured to estimate the pixel motion of the video to be processed based on adjacent video image frames in the video to be processed, and obtain the homography matrices corresponding to the specified video image frames in the video to be processed;

[0085] A matrix smoothing processing module 420, configured to sequentially use each video image frame in the video to be processed as the current video image frame, perform running smoothing processing on the homography matrix corresponding to the current video image frame, and obtain a smoothed homography matrix;

[0086] The image frame anti-shake processing module 430 is used to perform perspective transformation on the current video image frame based on the smooth homography matrix to obtain the anti-shake image frame of the current video image frame;

[0087] The anti-shake video generation module 440 is used to generate the processed video to be processed based on the anti-shake image frames of the video image frames in the video to be processed.

[0088] Optionally, the homography matrix acquisition module 410 is further configured to:

[0089] Perform compression processing on adjacent video image frames in the video to be processed to obtain adjacent compressed video image frames;

[0090] Perform pixel point motion estimation based on the adjacent compressed video image frames to obtain the homography matrices corresponding to the compressed video image frames respectively;

[0091] Perform scale reduction processing on the homography matrices corresponding to the compressed video image frames respectively to obtain homography matrices with the same size as the original video image frames in the video to be processed, as the homography matrices corresponding to the corresponding video image frames in the video to be processed.

[0092] Optionally, the performing pixel point motion estimation based on the adjacent compressed video image frames to obtain the homography matrices corresponding to the compressed video image frames respectively includes:

[0093] For the current compressed video image frame in the adjacent compressed video image frames, determine the corresponding feature points in the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame by performing feature point detection on the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame;

[0094] Adopt the sparse optical flow method to perform motion estimation on the corresponding feature points to obtain the homography matrix corresponding to the current compressed video image frame

[0095] Optionally, the performing running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smooth homography matrix further includes:

[0096] Perform motion smoothing processing on the homography matrix corresponding to the current video image frame according to the homography matrices corresponding to the video image frames in the image frame window corresponding to the current video image frame to obtain a smooth homography matrix.

[0097] Optionally, the performing motion smoothing processing on the homography matrix corresponding to the current video image frame according to the homography matrices corresponding to the video image frames in the image frame window corresponding to the current video image frame to obtain a smooth homography matrix includes:

[0098] Multiply the homography matrices corresponding to each video image frame within the image frame window corresponding to the current video image frame to obtain a first matrix representing the cumulative motion corresponding to multiple consecutive video image frames.

[0099] Calculate a second matrix representing the average motion of a single video image frame according to the first matrix.

[0100] Multiply the second matrix by the inverse matrix of the homography matrix corresponding to the current image frame to obtain a third matrix.

[0101] Use the third matrix as the smoothed homography matrix corresponding to the current video image frame.

[0102] The aerial video anti-shake processing device disclosed in the embodiments of the present application is used to implement the aerial video anti-shake processing method described in the embodiments of the present application. The specific implementation manners of the modules of the device will not be elaborated here, and reference may be made to the specific implementation manners of the corresponding steps in the method embodiments.

[0103] The aerial video anti-shake processing device disclosed in the embodiments of the present application estimates the motion of pixel points of the to-be-processed video based on adjacent video image frames in the to-be-processed video, and obtains the homography matrices respectively corresponding to specified video image frames in the to-be-processed video; performs running smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix; performs perspective transformation on the current video image frame based on the smoothed homography matrix to obtain the anti-shake image frame of the current video image frame; and generates the to-be-processed video after anti-shake processing based on the anti-shake image frames of each video image frame in the to-be-processed video. This method generates a homography matrix describing the inter-frame motion of feature points, performs smoothing processing on the homography matrix, and then performs motion repair on the video image frames based on the smoothed homography matrix to output the video image frames after anti-shake processing. Compared with the traditional video anti-shake processing based on the parameters collected by the gyroscope set on the aircraft, it has a more accurate and stable video anti-shake processing effect.

[0104] Further, by using the displacement description information of the feature points of adjacent video image frames, that is, the homography matrix locally smooths the displacement from the current video image frame to the adjacent video image frame, and directly calculates the motion compensation from the original video image frame, effectively reducing the error accumulation generated in the motion smoothing process.

[0105] In addition, before tracking features in the current video image frame using the optical flow method, the video image frame is first compressed, which can improve the feature matching speed of the optical flow method.

[0106] The aerial video anti-shake processing device disclosed in the embodiments of the present application suppresses the position change between adjacent video image frames by performing motion repair on the current video image frame with jitter, so that during the video playback, the visual effect between adjacent video image frames is relatively smooth and the video picture movement is relatively smooth.

[0107] Each embodiment in this specification is described in a progressive manner. The key point of each embodiment is to illustrate the differences from other embodiments. For the same or similar parts among the embodiments, reference can be made to each other. For the device embodiments, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiments.

[0108] The above provides a detailed introduction to an aerial video anti-shake processing method and device of the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method of the present application and its core idea; at the same time, for those of ordinary skill in the art, according to the idea of the present application, there will be changes in the specific implementation manner and application scope. In summary, the content of this specification should not be construed as a limitation to the present application.

[0109] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.

[0110] The various component embodiments of the present application can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. Those skilled in the art should understand that a microprocessor or a digital signal processor (DSP) can be used in practice to implement some or all of the functions of some or all of the components in the electronic device according to the embodiments of the present application. The present application can also be implemented as a device or device program (such as a computer program and a computer program product) for executing part or all of the methods described herein. Such a program for implementing the present application can be stored on a computer-readable medium, or can be in the form of one or more signals. Such signals can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.

[0111] For example, Figure 5An electronic device that can implement the method according to the present application is shown. The electronic device may be a PC, a mobile terminal, a personal digital assistant, a tablet computer, etc. Traditionally, the electronic device includes a processor 510, a memory 520, and program code 530 stored on the memory 520 and executable on the processor 510. When the processor 510 executes the program code 530, the method described in the above embodiments is implemented. The memory 520 may be a computer program product or a computer-readable medium. The memory 520 may be an electronic memory such as a flash memory, an EEPROM (electrically erasable programmable read-only memory), an EPROM, a hard disk, or a ROM. The memory 520 has a storage space 5201 for the program code 530 of a computer program for executing any method step in the above method. For example, the storage space 5201 for the program code 530 may include respective computer programs for implementing various steps in the above method. The program code 530 is computer-readable code. These computer programs may be read from or written into one or more computer program products. These computer program products include program code carriers such as hard disks, compact discs (CDs), memory cards, or floppy disks. The computer program includes computer-readable code, which when running on the electronic device, causes the electronic device to execute the method according to the above embodiments.

[0112] An embodiment of the present application also discloses a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the aerial video anti-shake processing method as described in the embodiments of the present application are implemented.

[0113] Such a computer program product may be a computer-readable storage medium, and the computer-readable storage medium may have storage segments, storage spaces, etc. arranged similarly to the memory 520 in the Figure 5 shown electronic device. The program code may be stored in the computer-readable storage medium in a suitably compressed form, for example. The computer-readable storage medium is generally a portable or fixed storage unit as referred to in Figure 6 the reference. Generally, the storage unit includes computer-readable code 530', which is code read by the processor, and when these codes are executed by the processor, the respective steps in the method described above are implemented.

[0114] As used herein, "one embodiment", "an embodiment", or "one or more embodiments" means that a particular feature, structure, or characteristic described in connection with the embodiment is included in at least one embodiment of the present application. In addition, note that the examples of the phrase "in one embodiment" herein do not necessarily all refer to the same embodiment.

[0115] In the description provided herein, numerous specific details are set forth. However, it will be understood that embodiments of the present application may be practiced without these specific details. In some instances, well-known methods, structures and techniques have not been shown in detail so as not to obscure the understanding of this description.

[0116] In a claim, any reference sign between parentheses shall not be construed as limiting the claim. The word "comprising" does not exclude the presence of elements or steps not listed in a claim. The word "a" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present application may be implemented by means of hardware comprising several distinct elements, and by means of a suitably programmed computer. In a unit claim enumerating several means, several of these means may be embodied by one and the same item of hardware. The use of the words first, second, third and the like do not denote any order. These words may be interpreted as names.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application and are not intended to limit them. Although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some of the technical features. These modifications or replacements do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A method for de-shaking an aerial video, characterized in that: The method comprises: Perform pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtain homography matrices corresponding to designated video image frames in the video to be processed; Taking each video image frame in the video to be processed as the current video image frame in turn, and performing smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix; Performing perspective transformation on the current video image frame based on the smooth homography matrix to obtain a de-jittered image frame of the current video image frame; Based on the de-jittered image frames of each video image frame in the video to be processed, the video to be processed after de-jittering processing is generated.

2. The method according to claim 1, characterized in that The step of performing pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtaining homography matrices corresponding to designated video image frames in the video to be processed, comprises: Compressing adjacent video image frames in the video to be processed to obtain adjacent compressed video image frames; Performing pixel point motion estimation based on the adjacent compressed video image frames to obtain homography matrices corresponding to the compressed video image frames respectively; The homography matrices corresponding to the compressed video image frames are scaled down to obtain homography matrices with the same size as the original video image frames in the video to be processed, as the homography matrices corresponding to the corresponding video image frames in the video to be processed.

3. The method according to claim 2, characterized in that The performing pixel point motion estimation based on the adjacent compressed video image frames to obtain homography matrices corresponding to the compressed video image frames respectively includes: For a current compressed video image frame among adjacent compressed video image frames, determining corresponding feature points in the current compressed video image frame and the previous compressed video image frame by performing feature point detection on the current compressed video image frame and the previous compressed video image frame of the current compressed video image frame; The sparse optical flow method is used to perform motion estimation on the corresponding feature points to obtain a homography matrix corresponding to the current compressed video image frame.

4. The method according to claim 1, characterized in that: The step of performing smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix includes: According to the homography matrices corresponding to each video image frame in the image frame window corresponding to the current video image frame, motion smoothing processing is performed on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix.

5. The method according to claim 4, characterized in that The step of performing motion smoothing processing on the homography matrix corresponding to the current video image frame according to the homography matrix corresponding to each video image frame in the image frame window corresponding to the current video image frame to obtain a smoothed homography matrix includes: Multiplying homography matrices corresponding to the video image frames in the image frame window corresponding to the current video image frame to obtain a first matrix representing accumulated motion corresponding to a plurality of consecutive video image frames; According to the first matrix, a second matrix representing the average motion of a single video image frame is calculated; Multiplying the second matrix by the inverse matrix of the homography matrix corresponding to the current image frame to obtain a third matrix; The third matrix is ​​used as a smooth homography matrix corresponding to the current video image frame.

6. An aerial video de-shaking processing device, characterized in that: The device comprises: A homography matrix acquisition module, used to perform pixel motion estimation on the video to be processed based on adjacent video image frames in the video to be processed, and obtain homography matrices corresponding to designated video image frames in the video to be processed; A matrix smoothing processing module, used to sequentially use each video image frame in the video to be processed as a current video image frame, and perform smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix; An image frame de-jitter processing module, used for performing perspective transformation on the current video image frame based on the smooth homography matrix to obtain a de-jittered image frame of the current video image frame; The de-jittered video generation module is used to generate the video to be processed after de-jittering based on the de-jittered image frames of each video image frame in the video to be processed.

7. The device according to claim 6, characterized in that The homography matrix acquisition module is further used for: Compressing adjacent video image frames in the video to be processed to obtain adjacent compressed video image frames; Performing pixel point motion estimation based on the adjacent compressed video image frames to obtain homography matrices corresponding to the compressed video image frames respectively; The homography matrices corresponding to the compressed video image frames are scaled down to obtain homography matrices with the same size as the original video image frames in the video to be processed, as the homography matrices corresponding to the corresponding video image frames in the video to be processed.

8. The device according to claim 6, characterized in that The step of performing smoothing processing on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix further includes: According to the homography matrices corresponding to each video image frame in the image frame window corresponding to the current video image frame, motion smoothing processing is performed on the homography matrix corresponding to the current video image frame to obtain a smoothed homography matrix.

9. An electronic device comprising a memory, a processor, and a program code stored in the memory and executable on the processor, wherein: When the processor executes the program code, the method according to any one of claims 1 to 5 is implemented.

10. A computer-readable storage medium having program code stored thereon, characterized in that: When the program code is executed by a processor, the steps of the method described in any one of claims 1 to 5 are implemented.