Image stabilization method and device for aerial video of unmanned aerial vehicle, terminal equipment and computer readable storage medium
By performing frame decomposition, anchor frame determination, matching and affine transformation processing on drone aerial video, the jitter and blur problems caused by unstable hovering of drone aerial videos are solved, and the video image stabilization processing is achieved, improving the stability and viewing of the video.
Patent Information
- Application Number
- CN202510143011.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-13
AI Technical Summary
Drone aerial videos are unstable due to external environmental factors and control system errors, resulting in shaking and blurring of the video screen, affecting the clarity and viewing of the video.
By obtaining the pending video of the drone aerial shot, decomposing it into video frames, measuring the difference between adjacent frames, determining the anchor frame, matching it with the anchor frame, deleting the failure point, determining the affine transformation matrix, repairing the failure point, and finally combining the processed video frames to generate the video after stable image.
Reduce feature loss caused by object occlusion or motion blur, improve the stability and clarity of the video, and improve the viewing of drone aerial videos.
Smart Images

Figure CN119996699A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of unmanned aerial vehicle image processing, and in particular to a method, device, terminal equipment and computer-readable storage medium for stabilizing aerial video of unmanned aerial vehicles. Background Art
[0002] Unmanned aircraft, also known as drones, are unmanned aircraft that are controlled by radio remote control equipment and self-contained program control devices, or in some advanced applications, can be operated completely or intermittently autonomously by an onboard computer. With its many advantages such as high flexibility, strong maneuverability and easy operation, drones have shown great application potential in many fields such as aerial photography, environmental monitoring, agricultural plant protection, and logistics distribution.
[0003] When using drones for aerial video shooting, although they have the ability to hover at high altitudes for a long time, due to the influence of external environmental factors (such as wind and airflow) and the slight errors of the drone's own control system, it is often difficult to achieve absolute stability in the hovering state. This unstable state will directly cause the video to be jittery and blurred, greatly affecting the clarity and viewing quality of the video. Therefore, in order to overcome this problem and improve the quality of drone aerial videos, it is urgent to stabilize the aerial videos. Summary of the invention
[0004] The embodiments of the present invention provide a method, apparatus, terminal device and computer-readable storage medium for stabilizing aerial video of a drone, which can reduce feature loss caused by object occlusion or motion blur, thereby improving the stability of the video.
[0005] An embodiment of the present invention provides a method for stabilizing an aerial video of an unmanned aerial vehicle, comprising:
[0006] Obtain the video to be processed taken by the drone, and decompose the video to be processed into several video frames;
[0007] The difference between two adjacent frames is measured to obtain the difference between adjacent frames, and the anchor frame is determined based on the difference between adjacent frames;
[0008] For each video frame, the video frame is matched with the anchor frame. Based on the anchor frame, the failure points where tracking fails or features are lost due to object occlusion are deleted to obtain the anchor pair with successful tracking.
[0009] Determine the affine transformation matrix based on the successfully tracked anchor point pairs;
[0010] Perform Harris corner point detection on the position of the failure point in the anchor frame to determine the pixel of the failure point;
[0011] Mapping the pixels of the failure points to the video frame through the affine transformation matrix to obtain the processed video frame;
[0012] All processed video frames are combined to obtain a stabilized video.
[0013] Furthermore, the difference between two adjacent frames is measured to obtain the difference between adjacent frames, including:
[0014] Measuring the pixel difference between two adjacent frames to obtain the difference between adjacent frames; or
[0015] The similarity between two adjacent frames is measured to obtain the difference between adjacent frames.
[0016] Furthermore, the pixel difference between two adjacent frames is measured to obtain the difference between adjacent frames, including:
[0017] For each pair of pixels between two adjacent frames, the absolute value of the difference between the pixel values is calculated to obtain the pixel difference value;
[0018] The pixel difference values of each pair of pixels between two adjacent frames are summed to obtain the difference between adjacent frames.
[0019] Furthermore, the similarity between two adjacent frames is measured to obtain the difference between adjacent frames, including:
[0020] For two adjacent frames, the brightness mean and variance of each frame and the covariance between the two adjacent frames are calculated, and based on the brightness mean and variance of each frame and the covariance between the two adjacent frames, the brightness similarity, contrast similarity and structure similarity between the two adjacent frames are calculated;
[0021] The difference between adjacent frames is determined based on brightness similarity, contrast similarity and structure similarity.
[0022] Further, the anchor frame is determined according to the difference between adjacent frames, including:
[0023] Comparing each adjacent frame difference with a first preset candidate threshold value;
[0024] When the difference between adjacent frames is greater than a first preset candidate threshold, marking the later video frame in time between the two adjacent frames as a candidate key frame;
[0025] A number of candidate key frames are selected from each candidate key frame as anchor frames at a preset interval; or, each candidate key frame is clustered, and the candidate key frame corresponding to each cluster center obtained by clustering is used as the anchor frame.
[0026] Further, the anchor frame is determined according to the difference between adjacent frames, including:
[0027] Comparing each adjacent frame difference with a second preset candidate threshold value;
[0028] When the difference between adjacent frames is greater than the second preset candidate threshold, the video frame that is later in time between the two adjacent frames is marked as the anchor frame.
[0029] Furthermore, before combining all processed video frames to obtain a stabilized video, the method further includes:
[0030] The difference between two adjacent frames in all processed video frames is measured to obtain the difference between the processed adjacent frames;
[0031] Comparing each processed adjacent frame difference with a preset check threshold, wherein the preset check threshold is less than a preset candidate threshold;
[0032] When the difference between adjacent frames after processing is greater than a preset check threshold, it is determined that there are still failure points in the processed video frame, and a warning signal is output;
[0033] When there is no difference between adjacent frames after processing that is greater than the preset check threshold, the subsequent video stabilization is continued.
[0034] Based on the above method embodiment, the present invention provides a corresponding device embodiment, including: a video frame decomposition module, an anchor frame determination module, an anchor tracking module, an affine transformation matrix confirmation module, a Harris corner point detection module, a pixel mapping module and a video frame combination module;
[0035] The video frame decomposition module is used to obtain the video to be processed taken by the drone and decompose the video to be processed into several video frames;
[0036] An anchor frame determination module is used to measure the difference between two adjacent frames to obtain the difference between adjacent frames, and determine the anchor frame according to the difference between adjacent frames;
[0037] Anchor point tracking module, for each video frame, matches the video frame with the anchor point frame, and uses the anchor point frame as a reference to delete the failed points where tracking fails or features are lost due to object occlusion, and obtain the anchor point pair where tracking is successful;
[0038] An affine transformation matrix confirmation module is used to determine the affine transformation matrix according to the successfully tracked anchor point pairs;
[0039] Harris corner point detection module, used to perform Harris corner point detection on the position corresponding to the failure point in the anchor point frame to determine the pixel of the failure point;
[0040] A pixel mapping module is used to map the pixels of the failure point to the video frame through an affine transformation matrix to obtain a processed video frame;
[0041] The video frame combination module is used to combine all processed video frames to obtain a stabilized video.
[0042] Based on the above-mentioned method embodiment, the present invention provides a corresponding terminal device embodiment, including: a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the steps of the drone aerial video stabilization method as described in the present invention are implemented.
[0043] Based on the above-mentioned method embodiment, the present invention provides a corresponding computer-readable storage medium embodiment, including: a stored computer program, which controls the device where the computer-readable storage medium is located to execute the steps of the drone aerial video stabilization method as described in the present invention when the computer program is running.
[0044] Compared with the prior art, the beneficial effects of the embodiment of this solution are:
[0045] The present invention obtains a video to be processed taken by an unmanned aerial vehicle, then decomposes the video to be processed into a number of video frames, measures the difference between two adjacent frames of the number of video frames, obtains the difference between adjacent frames, quantifies the degree of change between video frames by calculating the difference between adjacent frames, and determines an anchor frame based on the difference between adjacent frames. The anchor frame is a relatively stable frame and serves as a reference for subsequent frames. Determining the anchor frame helps to reduce the amount of calculation and complexity in subsequent processing. Next, for each video frame, the video frame is matched with the anchor frame, and the anchor frame is used as a reference to delete the failure points where tracking fails or features are lost due to object occlusion, reduce false matches and unstable factors, improve the accuracy of image stabilization processing, and obtain the successfully tracked anchor pair; according to the successfully tracked anchor pair, the affine transformation matrix is determined, and the matrix will be used to map the feature points in subsequent frames to stable positions; Harris corner point detection is performed on the position corresponding to the failure point in the anchor frame to obtain the pixels of the failure point, and new and stable feature points are found near the failure point; finally, the pixels of the failure point are mapped to the video frame through the affine transformation matrix to obtain the processed video frame, which can reduce the problem of feature loss caused by object occlusion or motion blur to a certain extent, thereby improving the stability of the video, and all the processed video frames are combined to finally obtain the stabilized video.
[0046] In summary, the present invention confirms relatively stable anchor frames through the difference between adjacent frames, and uses them as a reference, which effectively reduces the amount of calculation and complexity and improves the imaging quality; then, matches each video frame with the anchor frame, deletes the failure points, thereby improving the accuracy of the image stabilization processing; repairs the failure points through affine transformation, thereby improving the accuracy of the image stabilization processing; and finally combines the processed video frames to achieve video generation after image stabilization. BRIEF DESCRIPTION OF THE DRAWINGS
[0047] Figure 1 It is a flowchart of a method for stabilizing an aerial video of a drone provided by an embodiment of the present invention;
[0048] Figure 2 The figure is a schematic diagram of the structure of a drone aerial video stabilization device provided by one embodiment of the present invention. DETAILED DESCRIPTION
[0049] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.
[0050] In the description of the present invention, it should be understood that the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features.
[0051] like Figure 1 As shown, an embodiment of the present invention provides a method for stabilizing an aerial video of a drone, the method comprising at least the following steps:
[0052] Step S1: obtaining a video to be processed taken by a drone, and decomposing the video to be processed into a number of video frames;
[0053] For step S1, in this embodiment, the video taken by the drone can be obtained by physical connection or wireless transmission. Specifically, when the drone performs an aerial photography task, the aerial video is stored in an internal memory or an external storage device (such as an SD card). After the aerial photography is completed, the aerial video file is copied from the internal memory or external storage device of the drone through a physical connection (such as a USB cable) as the video to be processed for subsequent processing. In order to meet the needs of real-time or convenience, the video taken by the drone can also be transmitted wirelessly, such as Wi-Fi, 4G / 5G network or dedicated frequency band, to a ground control station or other designated receiving device in real time to obtain the video to be processed.
[0054] After obtaining the video to be processed, use video processing software, libraries or tools (such as OpenCV, FFmpeg, VLC, etc.) to read the video to be processed, traverse the video to be processed, and split it frame by frame into independent static images. Each video frame represents a picture at a certain moment in the video.
[0055] Optionally, the color of each video frame is converted to a uniform color, preferably gray. It should be noted that since color images contain rich color information, this information may require additional computing resources to process and analyze during processing, and converting video frames to gray can significantly reduce the complexity of the data, because grayscale images only contain brightness information and do not involve color changes. Therefore, this conversion can reduce the computational complexity of image processing tasks and increase processing speed. At the same time, after removing the interference of color information, the algorithm can focus more on the brightness changes and structural features of the image, such as edges, textures, etc.
[0056] Step S2: measuring the difference between two adjacent frames to obtain the difference between adjacent frames, and determining the anchor frame according to the difference between adjacent frames;
[0057] In a preferred embodiment, measuring the difference between two adjacent frames to obtain the difference between adjacent frames includes:
[0058] Measuring the pixel difference between two adjacent frames to obtain the difference between adjacent frames; or
[0059] The similarity between two adjacent frames is measured to obtain the difference between adjacent frames.
[0060] For step S2, after obtaining a number of video frames through step S1, the difference between two adjacent frames is measured to obtain the inter-frame difference. In this embodiment, two methods can be used to measure the difference between two adjacent frames: the first method is to directly measure the pixel difference between two adjacent frames;
[0061] Preferably, measuring the pixel difference between two adjacent frames to obtain the adjacent frame difference includes:
[0062] For each pair of pixels between two adjacent frames, the absolute value of the difference between the pixel values is calculated to obtain the pixel difference value;
[0063] The pixel difference values of each pair of pixels between two adjacent frames are summed to obtain the difference between adjacent frames.
[0064] Specifically, for each pair of pixels in two adjacent frames, the absolute difference between their pixel values is calculated using the following formula to obtain the pixel difference value:
[0065] N(x,y)=|A(x,y)-B(x,y)|
[0066] Among them, N(x,y) represents the pixel difference value of each pair of pixels in two adjacent frames, A(x,y) represents the pixel value of the video frame that is earlier in time between the two adjacent frames, B(x,y) represents the pixel value of the video frame that is later in time between the two adjacent frames, and x and y represent the coordinates of the pixel points.
[0067] Then, all pixel difference values are summed to obtain the overall difference between two adjacent frames, that is, the difference between adjacent frames:
[0068]
[0069] Wherein, IFD represents the difference between adjacent frames, X represents the image length of the video frame, and Y represents the image width of the video frame.
[0070] The second method is to obtain the difference between adjacent frames by measuring the similarity between two adjacent frames;
[0071] Preferably, measuring the similarity between two adjacent frames to obtain the difference between adjacent frames includes:
[0072] For two adjacent frames, the brightness mean and variance of each frame and the covariance between the two adjacent frames are calculated, and based on the brightness mean and variance of each frame and the covariance between the two adjacent frames, the brightness similarity, contrast similarity and structure similarity between the two adjacent frames are calculated;
[0073] The difference between adjacent frames is determined based on brightness similarity, contrast similarity and structure similarity.
[0074] Specifically, the brightness mean and variance of each frame, as well as the covariance between two adjacent frames, are calculated. Then, the brightness similarity, contrast similarity, and structure similarity between two adjacent frames are calculated using the following formula based on the brightness mean, variance, and covariance calculated previously:
[0075]
[0076] Among them, a represents the video frame that comes earlier in time between two adjacent frames, b represents the video frame that comes later in time between two adjacent frames, l(a,b) represents the brightness similarity between two adjacent frames, c(a,b) represents the contrast similarity between two adjacent frames, s(a,b) represents the structural similarity between two adjacent frames, and μ a Indicates the average brightness of the video frame that comes earlier in time between two adjacent frames, μ b Indicates the average brightness of the video frame that is later in time between two adjacent frames, σ a represents the variance of the previous video frame in two adjacent frames, σ b represents the variance of the later video frame in time between two adjacent frames, σ abrepresents the covariance between two adjacent frames, c1, c2 and c3 represent constant terms, which are used to avoid the denominator being zero.
[0077] Combining the calculated brightness similarity, contrast similarity and structural similarity between two adjacent frames, the similarity (SSIM) between two adjacent frames is calculated by the following formula:
[0078] SSIM(a,b)=[l(a,b) α c(a,b) β s(a,b) γ ]
[0079] Among them, SSIM(a,b) represents the similarity between two adjacent frames, α, β and γ represent adjustable constants, and α, β and γ are all greater than 0, which are used to balance the relative importance of brightness, contrast and structure in SSIM calculation. The range of SSIM value is between -1 and 1. The closer the SSIM value is to 1, the more similar the two frames are in brightness, contrast and structure; conversely, the smaller the SSIM value is, the greater the difference between the two frames is.
[0080] In this embodiment, the similarity between two adjacent frames is converted into the dissimilarity between the adjacent frames as the difference between adjacent frames, that is, the difference between adjacent frames is obtained by subtracting the similarity between the two adjacent frames from 1.
[0081] Next, based on the content difference between adjacent frames, anchor frames that can represent key video information are accurately selected through quantitative analysis. In the present invention, any one of the following three methods can be used to determine the anchor frames: fixed interval sampling method, cluster analysis method, and adaptive threshold method.
[0082] Preferably, determining the anchor frame according to the difference between adjacent frames includes:
[0083] Comparing each adjacent frame difference with a first preset candidate threshold value;
[0084] When the difference between adjacent frames is greater than a first preset candidate threshold, marking the later video frame in time between the two adjacent frames as a candidate key frame;
[0085] A number of candidate key frames are selected from each candidate key frame as anchor frames at a preset interval; or, each candidate key frame is clustered, and the candidate key frame corresponding to each cluster center obtained by clustering is used as the anchor frame.
[0086] Specifically, if the fixed interval sampling method is used to determine the anchor frame, first, set the first preset candidate threshold to determine whether the scene has changed. When the difference between adjacent frames is greater than the first preset candidate threshold, it is considered that the scene has changed, and the video frame that is later in time between the two adjacent frames is marked as a candidate key frame, and the frames with larger differences are preliminarily screened out as candidate key frames. Next, set a fixed frame number interval to select anchor frames from the candidate key frames. Starting from the first frame of the candidate key frames, select a frame as the anchor frame at a fixed frame number interval. It should be noted that the "fixed frame number interval" here refers to the calculation according to the order of the candidate key frames in the video, rather than the actual frame number interval.
[0087] The fixed-interval sampling method is simple and efficient, does not require professional machine learning or data analysis background, and is suitable for various video processing tasks, especially for scenarios with high real-time requirements.
[0088] If cluster analysis is used to determine anchor frames, similar to the fixed interval sampling method, first set the first preset candidate threshold to preliminarily screen out frames with large content differences as candidate key frames. Then, select a clustering algorithm to cluster the candidate key frames, such as K-means clustering or hierarchical clustering. The clustering algorithm will calculate the similarity between frames based on the feature vectors of the candidate key frames, and divide the candidate key frames into several clusters. The frames in each cluster have high similarity in content or features, while the frames between different clusters are quite different. For each cluster, select the candidate key frame corresponding to the cluster center as the representative of the cluster, i.e., the anchor frame.
[0089] The clustering analysis method can adaptively adjust the anchor frame selection strategy according to the complexity of the video content. Through feature extraction and clustering calculation, frames with similar content can be grouped into one category, and representative frames can be selected from each category as anchor frames, which can more accurately reflect the structure of the video content.
[0090] It should be noted that, whether it is the fixed interval sampling method or the cluster analysis method, it is necessary to screen out candidate key frames before determining the anchor frame. Through preliminary screening, those frames with high content similarity can be eliminated, thereby reducing the computational burden of subsequent processing steps and improving the efficiency of the overall algorithm. In addition, the screening of candidate key frames helps to retain those frames with significant content changes, which often contain important information or events in the video, thereby ensuring that the selected anchor frames can more accurately represent the video content.
[0091] Preferably, determining the anchor frame according to the difference between adjacent frames includes:
[0092] Comparing each adjacent frame difference with a second preset candidate threshold value;
[0093] When the difference between adjacent frames is greater than the second preset candidate threshold, the video frame that is later in time between the two adjacent frames is marked as the anchor frame.
[0094] Specifically, if the adaptive threshold method is used to determine the anchor frame, the second preset candidate threshold is dynamically adjusted according to the size of the difference between adjacent frames, so as to extract more key frames in places where the scene changes greatly. The second preset candidate threshold is determined in the following way: all the differences between adjacent frames in the video frame are summed and averaged to obtain the average frame difference, which is used to measure the speed of change of the video content. Then, a preset proportional coefficient is set, and the average frame difference is multiplied by the preset proportional coefficient to obtain the second preset candidate threshold. In this way, for each video to be processed, a different second preset candidate threshold can be set according to the different average frame differences of each video to be processed to accurately capture the key frames in the video.
[0095] After determining the second preset candidate threshold, each adjacent frame difference is compared with the second preset candidate threshold. When the adjacent frame difference is greater than the second preset candidate threshold, it is considered that the video content here has changed significantly, and the video frame that is later in time between the two adjacent frames is marked as the anchor frame.
[0096] The adaptive threshold method can dynamically adjust the threshold according to the speed of change of the video content. In places where the scene changes greatly, the threshold will be increased accordingly due to the large difference between frames, so that more key frames can be extracted. In this way, even when the video content changes rapidly, key information can be accurately captured.
[0097] Step S3: For each video frame, the video frame is matched with the anchor point frame, and the failure points where tracking fails or features are lost due to object occlusion are deleted based on the anchor point frame to obtain the anchor point pair where tracking is successful;
[0098] For step S3, after determining the anchor frame through step S2, feature extraction is performed on the anchor frame and each video frame. Preferably, feature extraction can be performed through feature extraction algorithms such as Harris corner detection or Canny edge detection to capture key local features in the video frame; then, the features in the currently processed video frame are matched with the features in the anchor frame, and the similarity or Euclidean distance between the feature points is calculated to find the best matching pair. The Euclidean distance is a commonly used distance measurement method. It calculates the straight-line distance between two points in the feature space and can effectively reflect the similarity between the feature points. In the matching process, if a feature point in the currently processed video frame cannot form an effective correspondence with any feature point in the anchor frame, it is considered that the tracking of the feature point has failed. In addition, even if the feature point finds a match, if the quality of the match is lower than the preset threshold, such as the Euclidean distance is too long, it is considered that the feature is lost due to object occlusion, illumination change or other reasons. For the above two cases, these failed points need to be eliminated, and finally the corresponding feature points that are successfully tracked between the video frame and the anchor frame are obtained, that is, the anchor pair that is successfully tracked.
[0099] Step S4: Determine the affine transformation matrix according to the successfully tracked anchor point pairs;
[0100] For step S4, the corresponding feature points that are successfully tracked between the video frame and the anchor point frame, that is, the successfully tracked anchor point pairs, are obtained through step S3. These anchor point pairs not only record the positions of the feature points in their respective images, but also imply the geometric transformation information between the images. In order to obtain the geometric transformation between the video frame and the anchor point frame, the affine transformation matrix is determined by the least squares method, so that the feature points of the anchor point frame transformed by the affine transformation matrix can be as close as possible to the corresponding feature points in the video frame.
[0101] Step S5: Perform Harris corner point detection on the position of the failure point corresponding to the anchor point frame to determine the pixel of the failure point;
[0102] For step S5, in the feature matching process between the video frame and the anchor frame in step S3, the feature points that cannot be found due to object occlusion, illumination change or other reasons are called failure points. In order to repair these failed feature points, in the anchor frame, for the estimated position of each failure point, the Harris corner detection algorithm is applied. The Harris corner detection algorithm detects corners by calculating the gradient change in the local window of the image. Specifically, the gradient change in the local window of the image is calculated, the autocorrelation matrix M is constructed, and the determinant detM and trace traceM of the autocorrelation matrix M are calculated. Then, the Harris response function R is calculated:
[0103] R = detM-k(traceM) 2
[0104] Among them, k represents an empirical parameter used to adjust the sensitivity of the response function. When the Harris response function R exceeds the preset response threshold, it is considered that there is a corner point near the center of the window. Search in the neighborhood of the estimated position of the failure point to find the corner point with the highest response value R as a replacement for the failure point. If there are multiple corner points, the one with the highest response value is selected.
[0105] Step S6: Mapping the pixels of the failure points to the video frame through an affine transformation matrix to obtain a processed video frame;
[0106] For step S6, after finding the pixel of the failure point in the anchor frame in step S5, the affine transformation matrix calculated in step S4 is used to map the coordinates of the failure point in the anchor frame to new coordinates in the video frame, thereby achieving precise alignment between frames.
[0107] After the above steps, a video frame containing the repaired failure point is obtained, which will be more accurate and stable in tasks such as feature matching, image registration, and video stabilization. At the same time, since the affine transformation retains the collinearity and proportionality of the image, the processed video frame will also be more natural and coherent visually.
[0108] Step S7: combine all processed video frames to obtain a stabilized video.
[0109] In a preferred embodiment, before combining all processed video frames to obtain stabilized video, the method further includes:
[0110] The difference between two adjacent frames in all processed video frames is measured to obtain the difference between the processed adjacent frames;
[0111] Comparing each processed adjacent frame difference with a preset check threshold, wherein the preset check threshold is less than a preset candidate threshold;
[0112] When the difference between adjacent frames after processing is greater than a preset check threshold, it is determined that there are still failure points in the processed video frame, and a warning signal is output;
[0113] When there is no difference between adjacent frames after processing that is greater than the preset check threshold, the subsequent video stabilization is continued.
[0114] For step S7, after the processed video frames are obtained through step S6, that is, before all the processed video frames are combined, the processed video frames are checked to ensure the video quality after stabilization. Specifically, the difference between two adjacent frames in these processed video frames is recalculated to obtain the difference between adjacent frames after processing, the purpose of which is to detect the degree of change between the video frames to determine whether there are unrepaired failure points or anomalies. Then, each difference between adjacent frames after processing is compared with a preset inspection threshold, wherein the preset inspection threshold is less than a preset candidate threshold (the preset candidate threshold includes a first preset candidate threshold and a second preset candidate threshold) to determine whether the difference between adjacent frames is within an acceptable range. If there is a situation where the difference between adjacent frames after processing is greater than the preset inspection threshold, it means that there are still failure points or anomalies in the processed video frame, and an early warning signal is output at this time, indicating that further inspection or repair is required. If there is no situation where the difference between adjacent frames after processing is greater than the preset inspection threshold, it means that the quality of the processed video frame is good, and the subsequent video stabilization processing can continue.
[0115] After checking that all processed video frames have achieved the expected image stabilization quality, all processed video frames are combined according to the time series to obtain the final stabilized video.
[0116] like Figure 2 As shown, based on the above method embodiment, a corresponding device embodiment is provided;
[0117] An embodiment of the present invention provides a drone aerial video stabilization device, comprising: a video frame decomposition module, an anchor frame determination module, an anchor tracking module, an affine transformation matrix confirmation module, a Harris corner point detection module, a pixel mapping module, and a video frame combination module;
[0118] The video frame decomposition module is used to obtain the video to be processed taken by the drone and decompose the video to be processed into several video frames;
[0119] An anchor frame determination module is used to measure the difference between two adjacent frames to obtain the difference between adjacent frames, and determine the anchor frame according to the difference between adjacent frames;
[0120] Anchor point tracking module, for each video frame, matches the video frame with the anchor point frame, and uses the anchor point frame as a reference to delete the failed points where tracking fails or features are lost due to object occlusion, and obtain the anchor point pair where tracking is successful;
[0121] An affine transformation matrix confirmation module is used to determine the affine transformation matrix according to the successfully tracked anchor point pairs;
[0122] Harris corner point detection module, used to perform Harris corner point detection on the position corresponding to the failure point in the anchor point frame to determine the pixel of the failure point;
[0123] A pixel mapping module is used to map the pixels of the failure point to the video frame through an affine transformation matrix to obtain a processed video frame;
[0124] The video frame combination module is used to combine all processed video frames to obtain a stabilized video.
[0125] It can be understood that the above-mentioned device item embodiments correspond to the method item embodiments of the present invention, and can implement the drone aerial video stabilization method provided by any of the above-mentioned method item embodiments of the present invention.
[0126] It should be noted that the device embodiments described above are merely illustrative, and some or all of the modules may be selected according to actual needs to achieve the purpose of the present embodiment. In addition, in the accompanying drawings of the device embodiments provided by the present invention, the connection relationship between the modules indicates that there is a communication connection between them, which may be specifically implemented as one or more communication buses or signal lines. Those of ordinary skill in the art may understand and implement the present invention without creative work.
[0127] Based on the above-mentioned embodiment of the drone aerial video stabilization method, another embodiment of the present invention provides a terminal device, which includes a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor. When the processor executes the computer program, the drone aerial video stabilization method of any embodiment of the present invention is implemented.
[0128] Exemplarily, in this embodiment, the computer program may be divided into one or more modules, which are stored in the memory and executed by the processor to implement the present invention. The one or more module elements may be a series of computer program instruction segments capable of implementing specific functions, which are used to describe the execution process of the computer program in the terminal device.
[0129] The terminal device may be a computing device such as a desktop computer, a notebook, a palmtop computer, a cloud server, etc. The terminal device may include, but is not limited to, a processor and a memory.
[0130] The processor may be a central processing unit (CPU), other general-purpose processors, digital signal processors (DSP), application-specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or any conventional processor, etc. The processor is the control center of the terminal device, and uses various interfaces and lines to connect various parts of the entire terminal device.
[0131] Based on the above method embodiment, another embodiment is provided: another embodiment of the present invention provides a computer-readable storage medium, including a stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the drone aerial video stabilization method described in any one of the above method embodiments of the present invention.
[0132] Wherein, the module / unit integrated in the drone aerial video stabilization device / terminal device, if implemented in the form of a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such an understanding, the present invention implements all or part of the processes in the above-mentioned embodiment method, and can also be completed by instructing the relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and the computer program can implement the steps of the above-mentioned various method embodiments when executed by the processor. Wherein, the computer program includes computer program code, and the computer program code can be in source code form, object code form, executable file or some intermediate form, etc. The computer-readable medium can include: any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, disk, optical disk, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electric carrier signal, telecommunication signal and software distribution medium, etc.
[0133] The above is a preferred embodiment of the present invention. It should be pointed out that a person skilled in the art can make several improvements and modifications without departing from the principle of the present invention. These improvements and modifications are also considered to be within the scope of protection of the present invention.
Claims
1. A method for stabilizing aerial video of a drone, characterized in that: include: Obtain a video to be processed taken by a drone, and decompose the video to be processed into a number of video frames; Measuring the difference between two adjacent frames to obtain an adjacent frame difference, and determining an anchor frame according to the adjacent frame difference; For each video frame, the video frame is matched with the anchor point frame, and the failure points where tracking fails or features are lost due to object occlusion are deleted based on the anchor point frame to obtain the anchor point pair with successful tracking; Determining an affine transformation matrix according to the successfully tracked anchor point pair; Perform Harris corner point detection on the position of the failure point in the anchor frame to determine the pixel of the failure point; Mapping the pixels of the failure point to the video frame through the affine transformation matrix to obtain a processed video frame; All processed video frames are combined to obtain a stabilized video.
2. The method for stabilizing drone aerial video according to claim 1, characterized in that: The difference between two adjacent frames is measured to obtain the difference between adjacent frames, including: Measuring the pixel difference between two adjacent frames to obtain the difference between adjacent frames; or The similarity between two adjacent frames is measured to obtain the difference between adjacent frames.
3. The method for stabilizing drone aerial video according to claim 2, characterized in that: The pixel difference between two adjacent frames is measured to obtain the difference between adjacent frames, including: For each pair of pixels between two adjacent frames, the absolute value of the difference between the pixel values is calculated to obtain the pixel difference value; The pixel difference values of each pair of pixels between two adjacent frames are summed to obtain the difference between adjacent frames.
4. The method for stabilizing drone aerial video according to claim 2, characterized in that: The similarity between two adjacent frames is measured to obtain the difference between adjacent frames, including: For two adjacent frames, the brightness mean and variance of each frame and the covariance between the two adjacent frames are calculated, and based on the brightness mean and variance of each frame and the covariance between the two adjacent frames, the brightness similarity, contrast similarity and structure similarity between the two adjacent frames are calculated; The difference between adjacent frames is determined according to the brightness similarity, contrast similarity and structure similarity.
5. The method for stabilizing drone aerial video according to claim 1, characterized in that: Determining an anchor frame according to the adjacent frame differences includes: Comparing each adjacent frame difference with a first preset candidate threshold value; When the difference between adjacent frames is greater than a first preset candidate threshold, marking the later video frame in time between the two adjacent frames as a candidate key frame; A number of candidate key frames are selected from each candidate key frame as anchor frames at a preset interval; or, each candidate key frame is clustered, and the candidate key frame corresponding to each cluster center obtained by clustering is used as the anchor frame.
6. The method for stabilizing drone aerial video according to claim 1, characterized in that: Determining an anchor frame according to the adjacent frame differences includes: Comparing each adjacent frame difference with a second preset candidate threshold value; When the difference between adjacent frames is greater than the second preset candidate threshold, the video frame that is later in time between the two adjacent frames is marked as the anchor frame.
7. The method for stabilizing drone aerial video according to claim 5, characterized in that: Before combining all processed video frames to obtain stabilized video, the following steps are also included: The difference between two adjacent frames in all processed video frames is measured to obtain the difference between adjacent frames after processing; Comparing each processed adjacent frame difference with a preset check threshold, wherein the preset check threshold is less than a preset candidate threshold; When the difference between adjacent frames after processing is greater than a preset check threshold, it is determined that there are still failure points in the processed video frame, and a warning signal is output; When there is no difference between adjacent frames after processing that is greater than the preset check threshold, the subsequent video stabilization is continued.
8. A drone aerial video stabilization device, characterized in that: include: Video frame decomposition module, anchor frame determination module, anchor tracking module, affine transformation matrix confirmation module, Harris corner detection module, pixel mapping module and video frame combination module; The video frame decomposition module is used to obtain the video to be processed taken by the drone and decompose the video to be processed into a number of video frames; The anchor frame determination module is used to measure the difference between two adjacent frames to obtain the difference between adjacent frames, and determine the anchor frame according to the difference between adjacent frames; The anchor point tracking module is used to match the video frame with the anchor point frame for each video frame, and delete the failed points where the tracking fails or the features are lost due to object occlusion based on the anchor point frame, so as to obtain the anchor point pair with successful tracking; The affine transformation matrix confirmation module is used to determine the affine transformation matrix according to the successfully tracked anchor point pair; The Harris corner point detection module is used to perform Harris corner point detection on the position of the failure point corresponding to the anchor point frame to determine the pixel of the failure point; The pixel mapping module is used to map the pixels of the failure point to the video frame through the affine transformation matrix to obtain a processed video frame; The video frame combination module is used to combine all processed video frames to obtain a stabilized video.
9. A terminal device, characterized in that: include: A processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein when the processor executes the computer program, the method for stabilizing aerial video of a drone as described in any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium, characterized in that: include: A stored computer program, wherein when the computer program is running, the device where the computer-readable storage medium is located is controlled to execute the drone aerial video stabilization method as described in any one of claims 1 to 7.