Method and apparatus for eliminating high-altitude image jitter based on static frame features
By extracting and matching feature points in drone videos, establishing a homography matrix and performing differential set compensation, the problem of high-altitude video jitter in drone is solved, significantly improving the robustness and accuracy of video image stabilization.
Patent Information
- Application Number
- CN202510348298.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-24
- Publication Date
- 2025-06-10
- Estimated Expiration
- 2045-03-24
AI Technical Summary
When a drone hoveres at high altitude to shoot video, due to factors such as airflow changes and fuselage vibration, rotation and jitter displacements occur between video frames, introducing errors to subsequent detection and tracking tasks.
By extracting the feature point set of video frames, and based on the static feature template, proximity matching calculation and Lowe's ratio test are performed, dynamic feature points are filtered, a homography matrix is established, video frames are aligned, and screen defects are repaired through the difference set compensation mechanism.
It significantly improves the robustness and continuity of video image stabilization, reduces manual intervention, is suitable for complex dynamic scenes in high-altitude shooting of drones, and improves computing efficiency and image stabilization accuracy.
Smart Images

Figure CN119863381B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of UAV image processing. More specifically, the present invention relates to a method and device for eliminating high-altitude image jitter based on static frame features. Background Art
[0002] Shooting videos by high-altitude UAVs is an important way to collect data in the current road traffic field. However, when a UAV hovers at high altitude to shoot videos, due to influencing factors such as airflow changes and airframe vibrations, rotation and jitter displacement will inevitably occur, resulting in rotation and translation phenomena between the front and rear frames of the video image. This will change the true coordinates of target objects such as vehicles and pedestrians, introduce additional errors to subsequent video extraction tasks such as detection and tracking, and cause the data to be unavailable.
[0003] Currently, video stabilization technology is often used to try to solve the above problems. Existing video stabilization technologies often only eliminate the local image jitter between adjacent video frames. For the UAV hovering shooting scenario, the focus object is the road traffic participants within a certain range below the UAV. Ideally, the UAV does not move, and the coordinate changes of target objects such as vehicles and pedestrians are only caused by the movement of the objects themselves, that is, the coordinates of background points corresponding to the same position in the real world should remain consistent in different video frames.
[0004] A video stabilization method for UAV high-altitude hovering shooting scenarios is disclosed in the patent with the authorization announcement number CN117575966B. First, the video is separated frame by frame and preprocessed by grayscale conversion. Using the first frame as the template frame, the regions of interest are manually selected and clicked on the template frame, and the corner points in each region are detected as anchor points. The coordinate relationships of the anchor points form the spatial coordinate system of the template frame. The affine transformation matrix between each frame of the original video and the template frame is estimated, and the lost anchor point pairs are compensated. Based on the above information, a robust estimation model with the affine transformation matrix used to stabilize each frame of the image as the parameter to be optimized is constructed, and the weight selection iteration is performed on each anchor point pair based on the idea of robust estimation. Then the model is solved to obtain the affine transformation matrix of each video frame that meets the root mean square error threshold. Finally, the original video is transformed frame by frame and written into a new video. This method can eliminate the influence of the UAV's own movement and achieve an ideal video stabilization effect. However, in this patent, the reliance on manual selection of anchor points may introduce subjective errors, and the compensation mechanism only targets the anchor points that fail to be tracked, which is easily affected by the occlusion of dynamic objects. At the same time, through affine transformation alignment, local information of the image may be lost due to the occlusion of dynamic objects, and it is also necessary to iteratively calculate the robust estimation weights, resulting in a high computational complexity. In addition, it has strict restrictions on the number and position of anchor points and may fail in sparse feature scenarios. Summary of the Invention
[0005] One object of the present invention is to solve at least the above problems and provide at least the advantages described hereinafter.
[0006] To achieve these objects and other advantages according to the present invention, there is provided a method for eliminating jitter of high-altitude images based on static frame features, including:
[0007] S1. Obtain a video, extract frames from the video, and extract a feature point set for each video frame;
[0008] S2. Using the feature point set of the k-th extracted video frame as a template, perform proximity matching calculation between the feature point sets of other video frames and the feature point set of the k-th video frame. When a successfully matched feature point appears in the two feature point sets, detect whether the feature point has an identifier. If there is no identifier, generate a unique identifier for the feature point and initialize the occurrence times of the feature point to 2. If there is an identifier, increment the occurrence times of the feature point by 1;
[0009] S3. Repeat S2 until proximity matching calculation has been performed between the k-th video frame and all other video frames;
[0010] S4. Retain the feature points in the feature point set of the k-th video frame whose occurrence times exceed a preset threshold, thereby generating a static feature template;
[0011] S5. Extract frames from the video one by one, extract the feature point set of the current video frame, perform Lowe's ratio test calculation between the feature point set and the static feature template, filter out dynamic feature points, retain static feature points, establish a homography matrix between the current video frame and the k-th video frame according to the positions of the static feature points in the current video frame and the static feature template, and align the current video frame with the k-th video frame according to the homography matrix;
[0012] S6. Detect the difference set area between the k-th video frame and the current video frame, crop and compensate the difference set area in the k-th video frame into the current video frame, and then rewrite the current video frame into the video;
[0013] S7. Repeat S5~S6 until the last frame of the video.
[0014] Preferably, in S1, a SURF feature point detector is used to extract the feature points of each video frame, thereby generating a feature point set belonging to the video frame.
[0015] Preferably, the proximity matching calculation between the feature point sets of other video frames and the feature point set of the k-th video frame in S2 specifically includes the following steps:
[0016] S21. Construct a k-d tree spatial index for the feature point set of the k-th frame based on the CUDA parallel acceleration algorithm, divide the point cloud into 32×32 grid cells, select the axis with the largest variance as the splitting dimension for each node, set the maximum capacity of the leaf nodes to 16 points, and limit the search depth to 20 levels;
[0017] S22. Adopt a dynamic radius matching strategy to dynamically calculate the matching radius according to the resolution of the input frame:
[0018] ,
[0019] Taking the feature points of the k-th frame as the center, quickly retrieve the candidate matching points in other video frames with a spatial distance less than r through the k-d tree;
[0020] S23. Hamming distance screening based on double constraints: If the number of candidate matching pairs > 500 pairs, retain the top 8% of the matching pairs with a Hamming distance less than 25 and the smallest distance; if the number of candidates ≤ 500 pairs, retain the matching pairs with a Hamming distance less than 30; eliminate the false matches through the probability-weighted RANSAC algorithm; dynamically adjust the sampling weight according to the feature point distribution density, reducing the weight in the high-density area by 30%, where the high-density area refers to the area in the 32×32 grid cell that contains more than 50 feature points; set the maximum number of iterations to 3000 times, and if the change in the model error is less than 0.1% within 100 consecutive iterations, terminate in advance; use the LMedS optimizer to calculate the fundamental matrix and select the model with the smallest median error:
[0021] ,
[0022] where is the homography matrix, and are the coordinates of the matching point pairs.
[0023] S24. Statistically calculate the standard deviation σ of the displacement vectors of the inter-frame feature points:
[0024] ,
[0025] If σ > 2.5 pixels, it is determined as an abnormal motion area, and all the matching points in this area are eliminated, where N is the number of matching point pairs, is the displacement vector of the i th matching point between frames, is the mean vector of the displacement vectors of all matching point pairs, is the square of the Euclidean distance between the displacement vector of a single pair of matching points and the mean vector, used to quantify the deviation degree of this displacement.
[0026] Preferably, in S1, the video frame extraction method is to extract frames once every preset time interval.
[0027] Preferably, the method for extracting the feature point set of each video frame in S1 includes:
[0028] S11. Construct a three-layer Gaussian pyramid with scaling ratios of 0.5, 1.0, and 1.5 respectively. The l image of the I l th layer is obtained by downsampling the l -1th layer image after Gaussian smoothing:
[0029] ,
[0030] where G(m,n) is a 5×5 Gaussian kernel with a standard deviation σ = 1.2. The parameters m and n are the coordinate offsets for traversing the Gaussian kernel. The 1.5-times layer is achieved by bilinearly interpolating and upsampling the original image;
[0031] S12. In each layer of the pyramid image, use the FAST-9 algorithm to detect corner points. Set the brightness threshold T = 20, and retain the corner point with the largest response value in the 3×3 neighborhood to avoid repeated detection in dense areas. Sort the detected corner points in each layer by their response values, and dynamically adjust the threshold to ensure that 400 ± 50 feature points are extracted in each layer, and a total of 1200 ± 200 feature points are extracted per frame;
[0032] S13. For each feature point, calculate the gray centroid of its neighborhood :
[0033] ,
[0034] The neighborhood radius r of each feature point is 15 pixels, and the feature point direction ;
[0035] S14. Rotate the feature point neighborhood to align with the main direction according to the feature point direction θ. Use 256 pairs of predefined random points (( x i , y i ) and ([[]] x j , y j )) to compare the gray values in the rotated neighborhood to generate a binary descriptor k:
[0036] ,
[0037] Finally, generate a 256-bit binary descriptor;
[0038] S15. Map the coordinate of feature points of each pyramid layer back to the original image scale. At the original image scale, perform 3×3 neighborhood screening on the feature points in the overlapping area, and retain the point with the largest response value to avoid redundancy.
[0039] Preferably, in S5, Lowe's ratio test with a multiplier of 0.7 is applied for calculation.
[0040] The present invention also provides a device for eliminating jitter of high-altitude images based on static frame features, including:
[0041] A feature point extraction module, which acquires a video, extracts frames from the video, and extracts a set of feature points for each video frame;
[0042] A proximity matching module, which uses the set of feature points of the k-th video frame extracted as a template, performs proximity matching calculation on the set of feature points of other video frames and the set of feature points of the k-th video frame. When a successfully matched feature point appears in the two sets of feature points, it detects whether the feature point has an identifier. If there is no identifier, a unique identifier is generated for the feature point, and the occurrence times of the feature point are initialized to 2. If there is an identifier, the occurrence times of the feature point are incremented by 1;
[0043] A first repetition module, which repeatedly calls the proximity matching module until all other video frames have been subjected to proximity matching calculation with the k-th video frame;
[0044] A static feature template generation module, which retains the feature points in the set of feature points of the k-th video frame whose occurrence times exceed a preset threshold, thereby generating a static feature template;
[0045] A video frame alignment module, which extracts frames from the video one by one, extracts the set of feature points of the current video frame, performs Lowe's ratio test calculation on the set of feature points and the static feature template, filters out dynamic feature points, retains static feature points, establishes a homography matrix between the current video frame and the k-th video frame according to the positions of the static feature points in the current video frame and the static feature template, and aligns the current video frame and the k-th video frame according to the perspective transformation matrix;
[0046] A compensation module, which detects the difference set area between the k-th video frame and the current video frame, crops and compensates the difference set area in the k-th video frame to the current video frame, and then rewrites the current video frame into the video;
[0047] A second repetition module, which repeatedly calls the video frame alignment module and the compensation module until the last frame of the video.
[0048] The present invention also provides an electronic device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to implement the above method.
[0049] The present invention also provides a storage medium, on which a computer program is stored, and when the program is executed by a processor, the above method is implemented.
[0050] The present invention has at least the following beneficial effects: Through the automated feature screening, dynamic interference filtering, and difference set compensation mechanisms, the present invention significantly improves the robustness, continuity, and computational efficiency of video stabilization while reducing manual intervention, and is particularly suitable for complex scenarios with frequent dynamic objects in high-altitude drone shooting.
[0051] Other advantages, objectives, and features of the present invention will be partially reflected by the following description, and partially will also be understood by those skilled in the art through the research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0052] Figure 1 It is a schematic flowchart of the method for eliminating high-altitude image jitter based on static frame features according to the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0053] It should be noted that, for the purpose of making the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0054] As Figure 1 shown, the present invention provides a method for eliminating high-altitude image jitter based on static frame features, including:
[0055] S1. Obtain a video, extract frames from the video, and extract the feature point set of each video frame;
[0056] Specifically, a drone can be used to shoot traffic flow videos at a height of 100 m or 300 m above the urban road, and the shooting duration exceeds 2 minutes;
[0057] Extract frames from the video. The frame extraction can follow the following rules: The video at a height of 100 m is extracted at intervals of 2 s, and the video at a height of 300 m is extracted at intervals of 1 s. The extracted video frames form an image set P, which contains a total of n images:
[0058] P = {p 1 , p 2 , …, p n},
[0059] For the feature point detection of the image set P, the SURF feature point detector can be used, and the corresponding feature points are recorded to generate the total feature point set W:
[0060] W = {w p1 , w p2 , …, w pn},
[0061] where w pn represents the feature point set generated from the nth video frame in the image set P.
[0062] In the generated feature point set, it contains both static object feature points and dynamic object feature points. Analyzing from the spatio-temporal continuity, since the UAV shooting method is high-altitude fixed-point shooting, the same static object on all frames should be in the same spatial position. Therefore, the feature points of these static objects can be formed into a set as the comparison reference basis for all video frames.
[0063] S2. Using the feature point set of the extracted kth video frame as a template, perform proximity matching calculation between the feature point sets of other video frames and the feature point set of the kth video frame. When a matching successful feature point appears in the two feature point sets, check whether the feature point has an identifier. If there is no identifier, generate a unique identifier for the feature point and initialize the occurrence times of the feature point to 2. If there is an identifier, increment the occurrence times of the feature point by 1;
[0064] Specifically, create a feature point library, and the initial feature point library is an empty set;
[0065] Obtain the feature point set of the kth video frame as a template reference, and perform feature point proximity matching calculation on w pk and w pi (1 ≤ i ≤ n and i ≠ k);
[0066] Record the matching successful feature points. If the current feature point is not marked, execute step a, otherwise execute step b:
[0067] a. Generate a unique identification symbol for the currently recorded feature point and initialize the occurrence times to 2;
[0068] b. Increment the occurrence times of the currently recorded feature point by 1.
[0069] S3. Repeat S2 until all other video frames have been subjected to proximity matching calculation with the kth video frame;
[0070] S4. Retain the feature points that appear more than a preset threshold in the feature point set of the k-th video frame, so as to generate a static feature template;
[0071] Specifically, count the number of occurrences of all feature points, retain the high-frequency feature points, remove the low-frequency feature points, and write the high-frequency feature points into the feature point library to form a static feature template. The determination rule of high-frequency feature points is as follows: sort the feature points according to the number of occurrences from high to low, take the top 50% and the number of feature points that meet the condition that the quantity is greater than or equal to 1 / 2 of the number of extracted frames as high-frequency feature points, and the others are low-frequency feature points.
[0072] S5. Extract the feature point set of the current video frame by frame from the video, perform a Lowe's ratio test calculation on this feature point set and the static feature template, filter out the dynamic feature points, retain the static feature points, establish a homography matrix between the current video frame and the k-th video frame according to the positions of the static feature points in the current video frame and the static feature template, and align the current video frame with the k-th video frame according to the homography matrix;
[0073] Specifically, the multiplier used in the Lowe's ratio test is 0.7.
[0074] S6. Detect the difference set area between the k-th video frame and the current video frame, crop and compensate the difference set area in the k-th video frame to the current video frame, and then rewrite the current video frame into the video;
[0075] S7. Repeat S5 - S6 until the last frame of the video.
[0076] In this embodiment, Lowe's ratio test is used to filter dynamic feature points, and combined with a static feature template (retaining static feature points that appear frequently), it effectively distinguishes the static background from dynamic objects (such as vehicles and pedestrians). Aiming at the problems of the existing technology relying on manual selection of anchor points and the compensation mechanism only targeting the anchor points that fail to be tracked, which is easily affected by the occlusion of dynamic objects. This solution can automatically eliminate the interference features of temporarily occluded or moving objects through statistical feature point appearance frequency and dynamic filtering, significantly improving the robustness of video stabilization in complex dynamic scenes. In this embodiment, the static feature template is generated by multi-frame feature point matching statistics without manual selection of anchor points. Aiming at the problems of the existing technology that requires manual selection of anchor points in the first frame, relying on manual operations and possibly introducing subjective errors, this solution is fully automated, reducing labor costs and at the same time improving the scalability of the algorithm (suitable for large-scale video processing). In this embodiment, the difference set area compensation crops and compensates the difference area between the template frame and the current frame to the current frame. Aiming at the problem that the existing technology only aligns through affine transformation, which may cause local information loss in the picture due to the occlusion of dynamic objects, this solution repairs the picture defects caused by occlusion or movement by compensating the difference set area, improving the visual coherence of the output video. In this embodiment, based on homography matrix alignment and static template matching, redundant calculations are reduced. Aiming at the problem that the existing technology needs to iteratively calculate and robustly estimate weights with a relatively high computational complexity, this solution reduces the resource consumption of real-time calculations through pre-screening of static templates and dynamic feature filtering. At the same time, the homography matrix is more suitable for complex perspective transformation scenarios (such as the perspective of an unmanned aerial vehicle shooting from above), and the video stabilization accuracy is higher. In this embodiment, it does not rely on a specific number of anchor points (such as 12 pairs of anchor points fixed in the existing technology) and supports adaptive feature point matching. Aiming at the problem that the existing technology has strict restrictions on the number and position of anchor points and may fail in feature-sparse scenarios. The new solution can adapt to the feature distribution in different scenarios by generating templates through multi-frame statistics, expanding the application potential of unmanned aerial vehicle video stabilization technology in high-dynamic and low-texture environments (such as sparse traffic scenarios).
[0077] Generally speaking, in this embodiment, through automated feature screening, dynamic interference filtering, and difference set compensation mechanisms, while reducing manual intervention, the robustness, continuity, and computational efficiency of video stabilization are significantly improved, especially suitable for complex scenarios with frequent dynamic objects in high-altitude shooting by unmanned aerial vehicles.
[0078] Further, the proximity matching calculation of the feature point set of other video frames with the feature point set of the k-th video frame described in S2 specifically includes the following steps:
[0079] S21. Based on the CUDA parallel acceleration algorithm, a k-d tree spatial index is constructed for the feature point set of the k-th frame. The point cloud is divided into 32×32 grid cells. The splitting dimension of each node selects the axis with the largest variance. The maximum capacity of the leaf node is 16 points, and the search depth limit is 20 layers;
[0080] S22. Adopt a dynamic radius matching strategy to dynamically calculate the matching radius according to the resolution of the input frame:
[0081] ,
[0082] Take the feature points of the k-th frame as the center, and quickly retrieve candidate matching points with a spatial distance less than r in other video frames through a k-d tree;
[0083] S23. Hamming distance screening based on double constraints: If the number of candidate matching pairs > 500 pairs, retain the top 8% of the matching pairs with a Hamming distance less than 25 and the smallest distance; if the number of candidates ≤ 500 pairs, retain the matching pairs with a Hamming distance less than 30; eliminate false matches through the probability-weighted RANSAC algorithm; dynamically adjust the sampling weight according to the feature point distribution density, and reduce the weight in the high-density area by 30%. The high-density area refers to the area where more than 50 feature points are included in a 32×32 grid cell; set the maximum number of iterations to 3000 times. If the change in the model error is less than 0.1% within 100 consecutive iterations, terminate in advance; use the LMedS optimizer to calculate the fundamental matrix and select the model with the smallest median error:
[0084] ,
[0085] where is the homography matrix, and are the coordinates of the matching point pairs.
[0086] S24. Statistically calculate the standard deviation σ of the displacement vectors of the inter-frame feature points:
[0087] ,
[0088] If σ > 2.5 pixels, it is determined as an abnormal motion area, and all matching points in this area are eliminated. Here, N is the number of matching point pairs, is the displacement vector of the i th pair of matching points between frames, is the mean vector of the displacement vectors of all matching point pairs, is the square of the Euclidean distance between the displacement vector of a single pair of matching points and the mean vector, which is used to quantify the deviation degree of this displacement.
[0089] In this embodiment, the k-d tree index divides feature points into grid cells (32×32), selects the axis with the largest variance as the splitting dimension, limits the search depth (20 levels), significantly accelerates the spatial retrieval efficiency of large-scale feature points, and is especially suitable for high-resolution UAV images (such as 4K / 8K). The dynamic radius matching (8 / 12 / 15 pixel grading) avoids redundant calculations of a fixed search radius at different resolutions, reducing the time complexity by about 30% - 50% (compared with traditional global search). CUDA parallel acceleration utilizes GPU computing resources to further optimize the matching speed and meet the real-time video stabilization requirements of UAVs. Dynamically adjust the threshold according to the number of candidate matches (keep the top 8% and the distance < 25 when the number of candidates > 500; keep the distance < 30 when the number of candidates ≤ 500), effectively filtering out false matches caused by texture repetition or noise. Dynamically adjust the sampling weight according to the distribution density of feature points (reduce the weight in high-density areas by 30%), avoiding false matches in dense areas from dominating the model fitting and improving the estimation accuracy of the fundamental matrix (homography matrix). The LMedS optimizer selects the model with the minimum median error, and has a higher tolerance for outliers compared to the mean error of traditional RANSAC, with the model robustness improved by about 20%. Calculate the standard deviation of the displacement vectors of the matching point pairs σ , and quantify the inter-frame motion consistency. If σ > 2.5 pixels, it is determined as an abnormal displacement area caused by dynamic objects (such as vehicles, pedestrians), and all matching points in this area are directly removed. Combining with the Lowe's ratio test, further filter dynamic feature points, making the homography matrix aligned only based on static background features, and avoiding the interference of dynamic objects on the video stabilization result (compared with the existing technology that relies on manual point selection, the robustness is improved by about 35%). Traditional RANSAC needs to fix the maximum number of iterations (such as 3000 times), while this solution reduces redundant iterations through dynamic termination conditions, and the calculation time is reduced by about 40% - 60%. Hierarchical feature processing (k-d tree index → Hamming distance screening → RANSAC optimization) gradually removes low-quality matching points, reduces the data scale for subsequent model fitting, and saves memory and computing resources. The dynamic radius matching strategy adapts to different resolution inputs (720P to 8K), avoiding false matches caused by too large a search radius at low resolutions. The weight in high-density areas is reduced by 30%, preventing overfitting of local features in texture-rich scenes such as urban roads, and improving the applicability of the algorithm in complex backgrounds (such as dense buildings, vegetation).
[0090] Generally speaking, in this embodiment, through efficient feature retrieval (k-d tree + dynamic radius), high-precision false match elimination (double Hamming + RANSAC optimization), and dynamic interference suppression (abnormal area elimination), the real-time performance, accuracy, and robustness of the UAV high-altitude image stabilization are significantly improved. Compared with the prior art (such as CN117575966 B relying on manual point selection and a fixed number of anchor points), this solution performs better in complex dynamic scenarios (traffic flow, pedestrians) and low-texture environments (such as sparse roads), while reducing the consumption of computing resources, and is suitable for the real-time processing requirements of large-scale UAV videos.
[0091] In another embodiment, the method for extracting the feature point set of each video frame in S1 includes:
[0092] S11. Construct a 3-layer Gaussian pyramid with scaling ratios of 0.5, 1.0, and 1.5 respectively. The l layer image I l is obtained by Gaussian smoothing and downsampling the l -1 layer image:
[0093] ,
[0094] where G(m,n) is a 5×5 Gaussian kernel, the standard deviation σ of this Gaussian kernel is 1.2, the parameters m and n are the coordinate offsets for traversing the Gaussian kernel, and the 1.5-fold layer is achieved by bilinearly interpolating and upsampling the original image;
[0095] S12. In each layer of the pyramid image, use the FAST-9 algorithm to detect corner points, take the brightness threshold T = 20, retain the corner point with the largest response value in the 3×3 neighborhood to avoid repeated detection in dense areas, sort the detected corner points in each layer according to the response value, and dynamically adjust the threshold to ensure that 400 ± 50 feature points are extracted in each layer, and a total of 1200 ± 200 feature points are extracted per frame;
[0096] S13. For each feature point, calculate the gray centroid of its neighborhood :
[0097] ,
[0098] The neighborhood radius r of each feature point is 15 pixels, and the feature point direction ;
[0099] S14. Rotate the feature point neighborhood to align with the main direction according to the feature point direction θ, and use 256 pairs of predefined random point pairs (( x i , y i ) and ( x j ,y j ), compare the gray values in the rotated neighborhood to generate the binary descriptor k:
[0100] ,
[0101] Finally, generate a 256-bit binary descriptor;
[0102] S15. Map the coordinates of the feature points in each pyramid layer back to the original image scale. At the original image scale, perform 3×3 neighborhood screening on the feature points in the overlapping area, and retain the point with the largest response value to avoid redundancy.
[0103] In this embodiment, a 3-layer Gaussian pyramid (scaling ratios of 0.5, 1.0, and 1.5) is constructed. Multi-scale images are generated through downsampling and upsampling. The low-layer pyramid (0.5 times) detects large-scale features (such as building edges), the middle layer (1.0 times) captures conventional features, and the high layer (1.5 times) enhances details through bilinear interpolation to adapt to the scenario of mixed near and far targets in UAV high-altitude images. Gaussian smoothing (standard deviation σ = 1.2) effectively suppresses image noise and avoids false feature detection caused by slight UAV jitter. Pyramid layer processing reduces repeated calculations. Compared with traditional single-scale detection, the calculation time consumption is reduced by about 30%. In this embodiment, the FAST-9 algorithm is used to detect corner points. The threshold is dynamically adjusted to ensure that 400 ± 50 feature points are extracted in each layer, and redundancy is avoided through neighborhood screening. The FAST-9 algorithm only needs to compare the pixel brightness within a 3×3 neighborhood, and the detection speed is 5 to 10 times faster than SIFT / SURF, meeting the real-time processing requirements. The threshold is dynamically adjusted according to the response value sorting in each layer (such as brightness threshold T = 20) to ensure a stable number of feature points (a total of 1200 ± 200 / frame), avoiding insufficient features in sparse scenarios or overload in dense scenarios. At the original image scale, 3×3 neighborhood screening is performed on the overlapping areas, and the point with the largest response value is retained, reducing redundant feature points by about 40% and improving the subsequent matching efficiency. In this embodiment, the gray centroid direction of the feature point neighborhood (radius r = 15 pixels) is calculated, and the neighborhood is rotated to align with the main direction. The direction of the feature point is corrected through the gray centroid direction, making the binary descriptor insensitive to image rotation. Experiments show that the rotation error tolerance is increased to ±15° (compared with ±5° in the traditional method). The rotated neighborhood generates a binary descriptor (256 pairs of random point pairs), and the descriptor consistency is improved by about 25%, reducing false matches caused by perspective changes. In this embodiment, a 256-bit binary descriptor is adopted, which is generated by comparing the gray values of random point pairs within the rotated neighborhood. Each feature point only requires 32 bytes (256 bits), reducing the storage space by 93.75% compared with SIFT (128-dimensional floating point, 512 bytes), which is suitable for large-scale video processing. The Hamming distance (XOR + bit counting) calculation speed is 20 to 50 times faster than the Euclidean distance. Combined with CUDA acceleration, the matching efficiency is significantly improved.
[0104] Generally speaking, through multi-scale feature detection, efficient corner screening, rotation-invariant descriptor generation, and lightweight storage design, this embodiment significantly improves the robustness, efficiency, and scene adaptability of UAV high-altitude image feature extraction. Compared with the prior art, this solution performs better in complex dynamic scenarios (such as urban traffic) and low-texture environments (such as long-distance shooting), while reducing the consumption of computing resources, providing high-quality feature inputs for subsequent jitter elimination algorithms (such as static template matching, homography matrix alignment), and thus improving the overall video stabilization effect.
[0105] Based on the same inventive concept, the present invention also provides a device for eliminating high-altitude image jitter based on a static frame feature formula. The device can be a personal computer, a server, or other devices that implement the aforementioned method for eliminating high-altitude image jitter based on a static frame feature formula.
[0106] The device for eliminating high-altitude image jitter based on the static frame feature formula of the present invention comprises:
[0107] The feature point extraction module obtains the video, extracts frames from the video, and extracts the feature point set of each video frame;
[0108] The proximity matching module uses the extracted feature point set of the k-th video frame as a template, performs proximity matching calculations on the feature point sets of other video frames and the feature point set of the k-th video frame. When a successfully matched feature point appears in the two feature point sets, it is detected whether the feature point has an identifier. If there is no identifier, a unique identifier is generated for the feature point, and the number of occurrences of the feature point is initialized to 2. If there is an identifier, the number of occurrences of the feature point is increased by 1.
[0109] The first repetition module repeatedly calls the proximity matching module until all other video frames have been subjected to proximity matching calculations with the kth video frame;
[0110] The static feature template generation module retains the feature points in the feature point set of the kth video frame whose number of occurrences exceeds a preset threshold, thereby generating a static feature template;
[0111] The video frame alignment module extracts the feature point set of the current video frame from the video frame one by one, performs Lloyd's ratio test calculation on the feature point set and the static feature template, filters the dynamic feature points, retains the static feature points, establishes the homography matrix between the current video frame and the k-th video frame according to the position of the static feature points of the current video frame and the position of the static feature points in the static feature template, and aligns the current video frame with the k-th video frame according to the perspective transformation matrix;
[0112] The difference compensation module detects the difference area between the k-th video frame and the current video frame, crops and compensates the difference area in the k-th video frame to the current video frame, and then rewrites the current video frame into the video;
[0113] The second repetition module repeatedly calls the video frame alignment module and the difference compensation module until the last frame of the video.
[0114] All relevant contents of each step involved in the embodiment of the aforementioned method for eliminating high-altitude image jitter based on static frame features can be referred to the functional description of the functional modules corresponding to the device in the embodiment of the present application, and will not be repeated here.
[0115] In the embodiments of the present application, the division of modules is illustrative. It is only a logical function division. In actual implementation, there may be other division methods. In addition, in each embodiment of the present invention, each functional module may be integrated in a processor, may exist separately physically, or two or more modules may be integrated in one module. The above integrated modules may be implemented in the form of hardware or in the form of software functional modules.
[0116] The present invention also provides an electronic device, which includes: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the above method for eliminating high-altitude image jitter based on static frame features. The electronic device may be any terminal device including a mobile phone, a notebook computer, a desktop computer, a tablet computer, a PDA (Personal Digital Assistant), a POS (Point of Sales), an in-vehicle computer, etc.
[0117] The present invention also provides a storage medium, on which a computer program is stored. When the program is executed by a processor, it implements the above method for eliminating high-altitude image jitter based on static frame features.
[0118] Through the description of the above embodiments, those skilled in the art can clearly understand that the present invention can be implemented by means of software plus necessary general hardware. Of course, it can also be implemented by dedicated hardware including application-specific integrated circuits, dedicated CPUs, dedicated memories, dedicated components, etc. Generally, functions completed by computer programs can be easily implemented by corresponding hardware, and the specific hardware structures for implementing the same function can also be various, such as analog circuits, digital circuits or dedicated circuits. However, for the present invention, in more cases, software program implementation is a better implementation manner. Based on such an understanding, the technical solution of the present invention, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. The computer software product is stored in a readable storage medium, such as a floppy disk, a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disc of a computer, etc., and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute the methods described in the various embodiments of the present invention.
[0119] Although the embodiments of the present invention have been disclosed as above, they are not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those skilled in the art, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and their equivalents, the present invention is not limited to the specific details and the illustrated examples described herein.
Claims
1. A method for eliminating high-altitude image jitter based on static frame features, characterized in that: include: S1, obtain the video, extract the video frame, and extract the feature point set of each video frame; S2, using the extracted feature point set of the k-th video frame as a template, performing proximity matching calculation on the feature point sets of other video frames and the feature point set of the k-th video frame. When a successfully matched feature point appears in the two feature point sets, it is detected whether the feature point has an identifier. If there is no identifier, a unique identifier is generated for the feature point, and the number of occurrences of the feature point is initialized to 2. If there is an identifier, the number of occurrences of the feature point is increased by 1. S3, repeat S2 until all other video frames have been subjected to proximity matching calculation with the kth video frame; S4, retaining the feature points whose number of occurrences in the feature point set of the kth video frame exceeds a preset threshold, thereby generating a static feature template; S5, extracting the feature point set of the current video frame from the video frame one by one, performing Lloyd's ratio test calculation on the feature point set and the static feature template, filtering the dynamic feature points, retaining the static feature points, establishing a homography matrix between the current video frame and the k-th video frame according to the positions of the static feature points of the current video frame and the positions of the static feature points in the static feature template, and aligning the current video frame with the k-th video frame according to the homography matrix; S6, detecting a difference region between the k-th video frame and the current video frame, cropping and compensating the difference region in the k-th video frame to the current video frame, and then rewriting the current video frame into the video; S7, repeat S5~S6 until the last frame of the video.
2. The method for eliminating high-altitude image jitter based on static frame features as claimed in claim 1, characterized in that: In S1, a SURF feature point detector is used to extract feature points of each video frame, thereby generating a feature point set belonging to the video frame.
3. The method for eliminating high-altitude image jitter based on static frame feature formula according to claim 1, characterized in that: The method of performing proximity matching calculation on the feature point sets of other video frames and the feature point set of the kth video frame in S2 specifically includes the following steps: S21. Based on the CUDA parallel acceleration algorithm, a kd-tree spatial index is constructed for the feature point set of the kth frame. The point cloud is divided into 32×32 grid units. The segmentation dimension of each node selects the axis with the largest variance. The maximum capacity of a leaf node is 16 points, and the search depth is limited to 20 layers. S22, using a dynamic radius matching strategy, dynamically calculating the matching radius according to the resolution of the input frame: , Taking the feature point of the kth frame as the center, the kd tree is used to quickly retrieve candidate matching points in other video frames whose spatial distance is less than r; S23. Hamming distance screening based on double constraints: If the number of candidate matching pairs is >500, retain the top 8% matching pairs with a Hamming distance less than 25 and the smallest distance; if the number of candidates is ≤500, retain the matching pairs with a Hamming distance less than 30; eliminate false matches through the probability weighted RANSAC algorithm; dynamically adjust the sampling weight according to the distribution density of feature points, and reduce the weight of high-density areas by 30%. The high-density area refers to the area with more than 50 feature points in the 32×32 grid unit; set the maximum number of iterations to 3000, and terminate early if the model error change is less than 0.1% within 100 consecutive iterations; use the LMeds optimizer to calculate the basic matrix and select the model with the smallest median error: , in is the homography matrix, and is the coordinates of the matching point pair; S24, statistical standard deviation σ of the displacement vector of feature points between frames: , If σ>2.5 pixels, it is determined as an abnormal motion area and all matching points in the area are removed, where N is the number of matching point pairs. For the i For the displacement vector of the matching point between frames, is the mean vector of the displacement vectors of all matching point pairs, It is the square of the Euclidean distance between the displacement vector of a single pair of matching points and the mean vector, which is used to quantify the degree of deviation of the displacement.
4. The method for eliminating high-altitude image jitter based on static frame feature formula according to claim 1, characterized in that: In S1, the video frame is extracted once every preset time interval.
5. The method for eliminating high-altitude image jitter based on static frame feature formula according to claim 1, characterized in that: The method of extracting a feature point set of each video frame in S1 includes: S11, construct a 3-layer Gaussian pyramid with scaling ratios of 0.5, 1.0, and 1.5 respectively. l Layer Image I l Through the l -1 layer image is Gaussian smoothed and downsampled to get: , Among them, G(m,n) is a 5×5 Gaussian kernel with a standard deviation of σ=1.
2. The parameters m and n are the coordinate offsets used to traverse the Gaussian kernel. The 1.5-fold layer is implemented by upsampling the original image through bilinear interpolation. S12. In each layer of the pyramid image, use the FAST-9 algorithm to detect corner points, take the brightness threshold T=20, retain the corner point with the largest response value in the 3×3 neighborhood, avoid repeated detection in dense areas, sort the corner points detected in each layer by response value, and dynamically adjust the threshold to ensure that 400±50 feature points are extracted in each layer, and a total of 1200±200 feature points are extracted per frame; S13. For each feature point, calculate the grayscale centroid of its neighborhood : , The neighborhood radius of each feature point is r = 15 pixels, and the feature point direction ; S14, according to the feature point direction θ, the feature point neighborhood is rotated to align with the main direction, using 256 pairs of predefined random point pairs (( x i , y i )and( x j , y j )), compare the grayscale values in the rotated neighborhood to generate a binary descriptor K: , Finally, a 256-bit binary descriptor is generated; S15. Map the coordinates of the feature points of each pyramid layer back to the original image scale. At the original image scale, perform 3×3 neighborhood screening on the feature points in the overlapping area, retain the points with the largest response value, and avoid redundancy.
6. The method for eliminating high-altitude image jitter based on static frame feature formula according to claim 1, characterized in that: S5 is calculated using the Lowe's ratio test with a multiplier of 0.
7.
7. A device for eliminating high-altitude image jitter based on static frame characteristics, characterized in that: include: The feature point extraction module obtains the video, extracts frames from the video, and extracts the feature point set of each video frame; The proximity matching module uses the extracted feature point set of the k-th video frame as a template, performs proximity matching calculations on the feature point sets of other video frames and the feature point set of the k-th video frame. When a successfully matched feature point appears in the two feature point sets, it is detected whether the feature point has an identifier. If there is no identifier, a unique identifier is generated for the feature point, and the number of occurrences of the feature point is initialized to 2. If there is an identifier, the number of occurrences of the feature point is increased by 1. The first repetition module repeatedly calls the proximity matching module until all other video frames have been subjected to proximity matching calculations with the kth video frame; The static feature template generation module retains the feature points in the feature point set of the kth video frame whose number of occurrences exceeds a preset threshold, thereby generating a static feature template; The video frame alignment module extracts the feature point set of the current video frame from the video frame one by one, performs Lloyd's ratio test calculation on the feature point set and the static feature template, filters the dynamic feature points, retains the static feature points, establishes the homography matrix between the current video frame and the k-th video frame according to the position of the static feature points of the current video frame and the position of the static feature points in the static feature template, and aligns the current video frame with the k-th video frame according to the perspective transformation matrix; The difference compensation module detects the difference area between the k-th video frame and the current video frame, crops and compensates the difference area in the k-th video frame to the current video frame, and then rewrites the current video frame into the video; The second repetition module repeatedly calls the video frame alignment module and the difference compensation module until the last frame of the video.
8. An electronic device, characterized in that include: At least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor performs the method according to any one of claims 1 to 6.
9. A storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
A video stabilization method for high-altitude hovering shooting scenes of unmanned aerial vehicles
CN117575966B
OpenCV based unmanned aerial vehicle real-time compression tracking method
CN109445453A
Video anti-shake system for eliminating interference of dynamic feature points
CN119211730A