Online fusion image stabilization method and system for unmanned aerial vehicle, computer equipment and storage medium

By performing feature matching and trajectory fusion on drone videos, combined with LSTM network and trajectory smoothing modules of residual blocks, the problem of drone video jitter is solved and the generation of stable videos is achieved.

CN120147196APending Publication Date: 2025-06-13HENAN UNIV OF SCI & TECH
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510204280.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-24
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

In the prior art, when processing video shooting of drones, it is difficult to effectively eliminate violent jitter in the video screen in real time, affecting the video quality.

Method used

A drone-oriented online fusion stable image method is adopted. By extracting single-frame images and IMU data from the original video, feature matching and trajectory fusion are performed, and the smooth camera trajectory is predicted using the LSTM network and the trajectory smoothing module of the residual block, and finally a stable video is generated.

Benefits of technology

It realizes effective real-time elimination of violent jitter on the video screen, improves video quality, and enhances the effect and robustness of the algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure BDA0005284074350000021
    Figure BDA0005284074350000021
  • Figure BDA0005284074350000025
    Figure BDA0005284074350000025
  • Figure BDA0005284074350000032
    Figure BDA0005284074350000032
Patent Text Reader

Abstract

An online fusion image stabilization method for an unmanned aerial vehicle comprises the following steps: extracting a single-frame image from an original video shot by the unmanned aerial vehicle, and preprocessing RGB information and gray information of the single-frame image; key feature points are extracted from the single-frame image; performing feature matching between two adjacent single-frame images based on the key feature points by using a feature matching network to generate an unstable image track; iMU data are obtained from the unmanned aerial vehicle, and an IMU motion track is generated; performing track fusion on the unstable image track and the IMU motion track to obtain a fusion track; using a pre-trained trajectory smoothing module to predict a smooth camera trajectory based on the fused trajectory; converting the smooth camera track into grid motion data, and generating a stable image sequence based on the grid motion data and the single-frame image; and synthesizing a stable video based on the stable image sequence. The method can effectively eliminate the violent jitter of the video image in real time.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer vision, and more specifically, to an online fusion image stabilization method, system, computer device, and storage medium for unmanned aerial vehicles (UAVs). Background Art

[0002] With the rapid development of UAVs and target recognition technologies, UAVs equipped with on-board cameras have been widely used in measurement and detection fields such as power line inspection, building safety, environmental monitoring, surveying and mapping, and urban planning management. However, factors such as strong winds, atmospheric turbulence, and the structure of the UAV itself can cause unnecessary jitter in the videos captured by the on-board cameras, resulting in motion blur and artifacts, which affect the video quality. Video image stabilization is a basic task in image processing, and its purpose is to eliminate unwanted jitter and generate a stable video. Therefore, video image stabilization technology is crucial for whether a UAV can complete complex visual tasks.

[0003] Most of the existing video image stabilization algorithms are methods based on deep learning for image processing, and a few are methods based on extracting sensor information as the data source for processing. The motion sensor-based method uses motion sensor data, such as gyroscopes and accelerometers, and can obtain accurate original camera motion. The sensor data is not affected by light and the specific scene content in the video, and can effectively correct image distortion, obtaining a relatively good image stabilization effect in an appropriate scene depth. However, the sensor-based method has poor effects for different scene depths, and the sensor data error gradually increases with time. Especially in the near-scene, more residual parallax will be generated.

[0004] The deep learning-based algorithm automatically learns feature representations and transformation models through an end-to-end training method, rather than manually extracting and matching features. Such methods can better adapt to different types of image jitter and can handle more complex video stabilization tasks. However, due to the complexity of the deep learning model, it requires more computing resources than traditional algorithms. When running on devices with limited computing resources, the deep learning-based algorithm often has a slow computing speed and poor real-time performance. Summary of the Invention

[0005] To solve the deficiencies in the prior art, the present invention provides an online fusion image stabilization method, system, computer device, and storage medium for UAVs, which can effectively and real-time eliminate severe jitter in video frames.

[0006] To achieve the above object, the specific solution adopted by the present invention is as follows: An online fusion image stabilization method for UAVs, comprising the following steps: Extract single-frame images from the original video captured by the UAV, and preprocess the RGB information and grayscale information of the single-frame images; Extract key feature points from a single-frame image; Use a feature matching network to perform feature matching based on key feature points between two adjacent single-frame images to generate an unstable image trajectory; Obtain IMU (Inertial Measurement Unit) data from the drone and generate an IMU motion trajectory; Perform trajectory fusion on the unstable image trajectory and the IMU motion trajectory to obtain a fused trajectory; Use a pre-trained trajectory smoothing module to predict a smooth camera trajectory based on the fused trajectory; Convert the smooth camera trajectory into grid motion data, and generate a stable image sequence based on the grid motion data and the single-frame images; Synthesize a stable video based on the stable image sequence.

[0007] As a further optimization of the above online fusion image stabilization method for drones: The method for extracting key feature points from a single-frame image is: where SuperPoint represents a feature point detector, is the extracted key feature point, is the coordinate of, f i is the i-th single-frame image of the original video, N is the total number of frames of the original video, and L is the total number of detected key feature points.

[0008] As a further optimization of the above online fusion image stabilization method for drones: The method for using a feature matching network to perform feature matching based on key feature points between two adjacent single-frame images includes: Solve the affine transformation matrix of two adjacent single-frame images based on the coordinate data of the key feature points corresponding to the two adjacent single-frame images: where, H i is the affine transformation matrix, are the coordinates of the key feature points corresponding to two adjacent single-frame images respectively, θ i is the rotation angle, dx i and dy i are the x-direction translation distance of the single-frame image along the x direction and the y-direction translation distance along the y direction respectively; Extract the x-direction translation distance dx i and the y-direction translation distance dy i and the rotation angle θ i from the affine transformation matrix as key parameters; Accumulate key parameters frame by frame to obtain an unstable image trajectory

[0009] As a further optimization of the above online fusion image stabilization method for drones: The method of obtaining IMU data from drones includes: Sample the original IMU data of the drone at a frequency of 400 - 600 Hz to obtain sampled data; Extract the three-axis rotational angular velocity data (ω x , ω y , ω z , t) from the sampled data, where ω x is the angular velocity in the x direction, ω y is the angular velocity in the y direction, ω z is the angular velocity in the z direction, and t is the time; Perform rolling shutter correction on the three-axis rotational angular velocity data; Calculate the camera rotation data of the drone based on the three-axis rotational angular velocity data R(t) = Sω(t) * R(t - S), where S is the sampling time interval; Use the spherical linear interpolation algorithm to expand the camera rotation data to obtain IMU data.

[0010] As a further optimization of the above online fusion image stabilization method for drones: The method of using the spherical linear interpolation method to expand the camera rotation data is: R(t f ) = SLERP(R(t a ), R(t b ), (t a - t b ) / (t b - t a ))), where t a and t b are two adjacent camera rotation data, and t a ≤t f ≤t b , and SLERP(·) is the spherical linear interpolation algorithm.

[0011] As a further optimization of the above online fusion image stabilization method for drones: The trajectory smoothing module includes at least two processing units arranged in sequence, and the processing unit includes an LSTM (Long Short-Term Memory) network and a residual block arranged in sequence.

[0012] As a further optimization of the above online fusion video stabilization method for drones: during the process of training the trajectory smoothing module, the overall loss is calculated based on a preset loss estimation method, and the training process is constrained according to the overall loss. The loss estimation method is as follows: where L is the calculated overall loss, and is the smoothing loss, L p is the boundary loss, L d is the distortion loss, N is the number of look-ahead frames, w p , w d are loss weights, w p,i is the normalized Gaussian weight centered on the current frame, α is the tolerable reference prominence value, Ω(R v , R r ) is the spherical angle between the current virtual and real camera poses, β 0 is the preset threshold, β 1 is the control parameter.

[0013] An online fusion video stabilization device for drones, used to implement the above online fusion video stabilization method for drones. The device includes: A data acquisition module, used to extract single-frame images from the original video captured by the drone and obtain IMU data from the drone; a data processing module, used to preprocess the RGB information and grayscale information of the single-frame images, extract key feature points from the single-frame images, and generate an IMU motion trajectory; A feature matching module, used to perform feature matching based on key feature points between two adjacent single-frame images using a feature matching network to generate an unstable image trajectory; A fusion module, used to perform trajectory fusion on the unstable image trajectory and the IMU motion trajectory to obtain a fusion trajectory; A trajectory processing module, used to predict a smooth camera trajectory based on the fusion trajectory, convert the smooth camera trajectory into grid motion data, and generate a stable image sequence based on the grid motion data and the single-frame images; A video regeneration module, used to synthesize a stable video based on the stable image sequence.

[0014] A computer device, including: A memory, used to store a computer program; A processor, used to execute the computer program to implement the above online fusion video stabilization method for drones.

[0015] A storage medium for storing a computer program which, when executed, implements the above-mentioned online fusion video stabilization method for unmanned aerial vehicles.

[0016] Beneficial effects: The present invention first determines an unstable image trajectory based on a single-frame image extracted from an original video, and determines an IMU motion trajectory based on IMU data obtained from a drone. Then, the unstable image trajectory and the IMU motion trajectory are fused, and the fused trajectory is used to predict a smooth trajectory by a trajectory smoothing module constructed based on an LSTM network and a residual block. Finally, a stabilized video is generated based on the smooth trajectory. In the whole processing process, a multi-modal fusion framework that fuses image and IMU data is adopted, which can effectively improve the algorithm effect and robustness, and well solves the problems that the existing video stabilization algorithms overly rely on the quality of feature extraction and lack of rigid constraints in feature extraction, resulting in large-area non-rigid distortion and artifacts. Description of the Drawings

[0017] Figure 1 is a schematic diagram of the overall process of the video stabilization method of the present invention; Figure 2 is a schematic diagram of the structure of the key feature point extraction module; Figure 3 is a schematic diagram of the structure of the feature matching network; Figure 4 is a schematic diagram of the structure of the trajectory optimization module. Detailed Embodiments

[0018] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0019] As Figures 1-4 shown, the present invention first provides an online fusion video stabilization method for unmanned aerial vehicles, including S1 to S8.

[0020] S1. Extract a single-frame image from the original video captured by the drone, and preprocess the RGB information and grayscale information of the single-frame image. The preprocessing methods may include normalization, color space conversion, noise removal, etc., all of which are conventional image processing methods and will not be elaborated here.

[0021] S2. Use the key feature point extraction module to extract key feature points from the single-frame image. The method for extracting key feature points from the single-frame image is: Among them, SuperPoint represents a feature point detector, is the extracted key feature point, is coordinates, f i is the i-th single-frame image of the original video, N is the total number of frames of the original video, and L is the total number of detected key feature points. SuperPoint adopts a lightweight structure and uses pixel-level positive and negative sample annotations, which can quickly generate stable and reliable key points. Moreover, the high-quality descriptors generated by SuperPoint can accurately describe the regional image information around the key points, thus playing an important role in subsequent feature matching.

[0022] S3. Use the feature matching network to perform feature matching based on key feature points between two adjacent single-frame images to generate unstable image trajectories. By performing feature matching between two adjacent single-frame images, the corresponding key feature point of the j-th key feature point in the i-th single-frame image in the (i + 1)-th single-frame image can be obtained, that is, a feature point pair is obtained, denoted as Furthermore, the position of a key feature point in different single-frame images can be obtained. Based on the position change of all key feature points, an unstable image trajectory can be generated. In the present invention, the feature matching network is preferably a LightGlue depth matcher. For a feature point pair, the processing method of the feature matching network is:

[0023] Furthermore, the method of using the feature matching network to perform feature matching based on key feature points between two adjacent single-frame images includes S31 to S33.

[0024] S31. Solve the affine transformation matrix of two adjacent single-frame images based on the coordinate data of the key feature points corresponding to the two adjacent single-frame images: Among them, H i is the affine transformation matrix, are the coordinates of the key feature points corresponding to two adjacent single-frame images respectively, θ i is the rotation angle, dx i and dy i are the x-direction translation distance of the single-frame image along the x direction and the y-direction translation distance along the y direction respectively. More specifically, after obtaining the feature point pair, a linear system is constructed using the feature point pair, and then the constraints of the homography matrix are described. Finally, the direct linear transformation (DLT) algorithm is used to solve the affine transformation matrix H i .

[0025] S32. Extract the x-direction translation distance dx from the affine transformation matrixi , the translation distance dy in the y direction i and the rotation angle θ i are used as key parameters.

[0026] S33. Accumulate the key parameters frame by frame to obtain the unstable image trajectory

[0027] S4. Obtain IMU (Inertial Measurement Unit) data from the drone and generate an IMU motion trajectory. The method for obtaining IMU data from the drone includes S41 to S45.

[0028] S41. Sample the original IMU data of the drone at a frequency of 400 - 600 Hz to obtain sampled data. In a preferred embodiment of the present invention, the sampling is performed at a frequency of 500 Hz.

[0029] S42. Extract the three - axis rotational angular velocity data (ω x , ω y , ω z , t) from the sampled data, where ω x is the angular velocity in the x - direction, ω y is the angular velocity in the y - direction, ω z is the angular velocity in the z - direction, and t is the time instant.

[0030] S43. Perform rolling - shutter correction on the three - axis rotational angular velocity data. At a sampling frequency of 500 Hz, for a video with 60 frames per second, each single - frame image corresponds to 8 sets of three - axis rotational angular velocity data. Also, because the exposure order of the rolling shutter is to scan each row of pixels from top to bottom, each single - frame image is equally divided into 8 pixel blocks of the same size in the column direction, which can be expressed as where represents the k - th set of three - axis rotational angular velocity data corresponding to the i - th image, and represents the k - th pixel block of the i - th image. Further, in the time series, the 8 sets of three - axis rotational angular velocity data respectively correspond to 8 pixel blocks from top to bottom. Use the three - axis rotational angular velocity data as the rotation matrices corresponding to the 8 pixel blocks on each image respectively, and then use this to eliminate the distortion caused by the tiny exposure time difference of the rolling shutter when the drone is in a fast - motion state.

[0031] S44. Calculate the camera rotation data of the drone based on the three - axis rotational angular velocity data R(t)=Sω(t)*R(t - S), where S is the sampling time interval. On the premise of sampling at a frequency of 500 Hz, the value of S is 2 ms. After obtaining the camera rotation data, represent the camera rotation data as a quaternion and store it in a queue.

[0032] S45. Expand the camera rotation data using the spherical linear interpolation algorithm to obtain IMU data. The method for expanding the camera rotation data using the spherical linear interpolation method is as follows: R(t f ) = SLERP(R(t a ), R(t b ), (t a - t b ) / (t b - t a ))), where t a and t b are two adjacent camera rotation data, and t a ≤ t f ≤ t b , and SLERP(·) is the spherical linear interpolation algorithm. Through the spherical linear interpolation algorithm, the camera rotation data at any timestamp t f can be obtained.

[0033] S5. Perform trajectory fusion on the unstable image trajectory and the IMU motion trajectory to obtain a fused trajectory.

[0034] S6. Use the pre-trained trajectory smoothing module to predict a smooth camera trajectory based on the fused trajectory. The trajectory smoothing module includes at least two processing units arranged in sequence. The processing unit includes an LSTM (Long Short-Term Memory) network and a residual block set in sequence. The trajectory smoothing network is constructed by the ability of the LSTM network to capture long-term dependencies in the time series. The LSTM network can also model the historical trajectory and effectively predict future trajectory points according to historical information, so as to adaptively learn and capture complex patterns of trajectory data, including speed changes, trajectory shapes, etc., and finally better eliminate noise and jitter in the trajectory, improving the quality and continuity of the trajectory data. In addition, the introduction of the residual module further enhances the learning ability of the LSTM network and the effect of trajectory smoothing. The residual module directly adds the input signal to the network output through a fully connected layer, which helps the gradient to propagate over multiple time steps and reduces the impact of gradient vanishing.

[0035] To ensure the performance of the trajectory smoothing module, so that the smooth camera trajectory output by the trajectory smoothing module will not lose too much image data due to excessive smoothing resulting in a large cropping rate, while achieving a good video stabilization effect and reducing distortion caused during the image processing process. During the training process of the trajectory smoothing module, the overall loss is calculated based on a pre-set loss estimation method, and the training process is constrained according to the overall loss. The loss estimation method is as follows: Among them, L is the calculated overall loss, and is the smoothing loss, L p is the boundary loss, L d is the distortion loss, N is the number of look-ahead frames, w p ,w d is the loss weight, w p,i is the normalized Gaussian weight centered on the current frame, α is the tolerable reference prominence value, Ω(R v ,R r ) is the spherical angle between the current virtual and real camera poses, β 0 is the preset threshold, β 1 is the control parameter.

[0036] S7. Convert the smoothed camera trajectory into grid motion data, and generate a stable image sequence based on the grid motion data and a single-frame image. Specifically, represent the smoothed camera trajectory as a relative quaternion, and convert the motion data in the form of relative quaternions into grid motion data. Then, through the warp transformation, cooperate the grid motion data with the single-frame image to generate a stabilized image sequence, that is, the stable image sequence.

[0037] S8. Synthesize a stable video based on the stable image sequence. Generating a video based on an image sequence belongs to the prior art in this field and will not be elaborated here.

[0038] The present invention further provides an online fusion image stabilization device for drones, which is used to implement the above-mentioned online fusion image stabilization method for drones. The device includes a data acquisition module, a data processing module, a feature matching module, a fusion module, a trajectory processing module, and a video regeneration module.

[0039] The data acquisition module is used to extract single-frame images from the original video captured by the drone and obtain IMU data from the drone.

[0040] The data processing module is used to preprocess the RGB information and grayscale information of the single-frame image, extract key feature points from the single-frame image, and generate an IMU motion trajectory.

[0041] The feature matching module is used to perform feature matching between two adjacent single-frame images based on key feature points using a feature matching network to generate an unstable image trajectory.

[0042] The fusion module is used to perform trajectory fusion on the unstable image trajectory and the IMU motion trajectory to obtain a fusion trajectory.

[0043] A trajectory processing module, configured to smooth a camera trajectory based on a fused trajectory prediction, convert the smoothed camera trajectory into grid motion data, and generate a stable image sequence based on the grid motion data and a single-frame image.

[0044] A video regeneration module, configured to synthesize a stable video based on the stable image sequence.

[0001] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the above-described systems and devices can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein. In several embodiments provided in the present disclosure, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. The device embodiments described above are merely illustrative. For example, the division of the modules is only a logical function division, and there can be other division methods in actual implementation. For another example, multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some communication interfaces. The indirect coupling or communication connection of the devices or modules can be in an electrical, mechanical, or other form.

[0045] The modules described as separate components may or may not be physically separated. The components displayed as modules may or may not be physical modules, that is, they can be located in one place or distributed to multiple network modules. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0046] The present invention also provides a computer device, including a memory and a processor.

[0047] The memory is used to store a computer program.

[0048] The processor is configured to execute the computer program to implement the above-mentioned online fusion image stabilization method for an unmanned aerial vehicle.

[0049] Finally, the present invention provides a storage medium, configured to store a computer program, and when the computer program is executed, it implements the above-mentioned online fusion image stabilization method for an unmanned aerial vehicle.

[0050] As a carrier for resource storage, the memory can be a read-only memory, a random access memory, a magnetic disk, an optical disc, etc. The resources stored thereon can include an operating system, computer programs, etc. The storage method can be transient storage or permanent storage. The operating system is used to manage and control each hardware device and computer program on the electronic device, and it can be Windows Server, Netware, Unix, Linux, etc. In addition to the computer program that can be used to implement the adaptive emotion regulation method based on personalized reconfigurable music executed by the electronic device disclosed in any of the foregoing embodiments, the computer program can further include computer programs that can be used to complete other specific tasks. The processor can adopt general processor products based on architectures such as X86, IA64, RISC, MIPS, and ARM.

[0051] The foregoing description of the disclosed embodiments enables those skilled in the art to implement or use the present invention. Various modifications to these embodiments will be apparent to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to the embodiments shown herein, but rather to the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An online fusion image stabilization method for unmanned aerial vehicles, characterized in that: The steps include: Extract single-frame images from the original video taken by the drone, and pre-process the RGB information and grayscale information of the single-frame images; Extract key feature points from a single frame image; Use a feature matching network to perform feature matching based on key feature points between two adjacent single-frame images to generate unstable image trajectories; Get IMU (Inertial measurement unit) data from the drone and generate IMU motion trajectories; The unstable image trajectory and the IMU motion trajectory are fused to obtain a fused trajectory; Use the pre-trained trajectory smoothing module to smooth the camera trajectory based on the fused trajectory prediction; The smooth camera trajectory is converted into grid motion data, and a stable image sequence is generated based on the grid motion data and a single frame image; and a stable video is synthesized based on the stable image sequence.

2. The online fusion image stabilization method for UAV according to claim 1, characterized in that: The method to extract key feature points from a single frame image is: Among them, SuperPoint represents a feature point detector, For the extracted key feature points, for The coordinates of i is the i-th single frame image of the original video, N is the total number of frames of the original video, and L is the total number of key feature points detected.

3. The online fusion image stabilization method for drones according to claim 1, characterized in that: Methods for performing feature matching based on key feature points between two adjacent single-frame images using a feature matching network include: Solve the affine transformation matrix of two adjacent single-frame images based on the coordinate data of the key feature points corresponding to the two adjacent single-frame images: Among them, H i is the affine transformation matrix, are the coordinates of the key feature points corresponding to two adjacent single-frame images, θ i is the rotation angle, dx i anddy i are the x-direction translation distance of a single frame image along the x direction and the y-direction translation distance along the y direction respectively; Extract the x-direction translation distance dx from the affine transformation matrix i , y translation distance dy i and the rotation angle θ i As a key parameter; Accumulate key parameters frame by frame to obtain unstable image trajectory 4. The online fusion image stabilization method for unmanned aerial vehicles according to claim 1, characterized in that: Methods for obtaining IMU data from a drone include: The original IMU data of the drone is sampled at a frequency of 400-600 Hz to obtain sampled data; The three-axis rotation angular velocity data (ω x ,ω y ,ω z ,t), where ω x is the angular velocity in the x direction, ω y is the angular velocity in the y direction, ω z is the angular velocity in the z direction, t is the time; Perform rolling shutter correction on the three-axis rotation angular velocity data; The camera rotation data of the drone is calculated based on the three-axis rotation angular velocity data R(t)=Sω(t)*R(tS), where S is the sampling time interval; The camera rotation data is expanded using the spherical linear interpolation algorithm to obtain the IMU data.

5. The online fusion image stabilization method for unmanned aerial vehicles according to claim 4, characterized in that: The method of expanding the camera rotation data using the spherical linear interpolation method is: R(t f )=SLERP(R(t a ),R(t b ),(t a -t b ) / (t b -t a )), where t a and t b Rotate the data for two adjacent cameras, and t a ≤t f ≤t b , SLERP(·) is the spherical linear interpolation algorithm.

6. The online fusion image stabilization method for unmanned aerial vehicles according to claim 1, characterized in that: The trajectory smoothing module includes at least two processing units arranged in sequence, and the processing units include an LSTM (Long Short-Term Memory) network and a residual block arranged in sequence.

7. The online fusion image stabilization method for unmanned aerial vehicles according to claim 6, characterized in that: During the training of the trajectory smoothing module, the overall loss is calculated based on a pre-set loss estimation method, and the training process is constrained according to the overall loss. The loss estimation method is: Among them, L is the calculated overall loss, and is the smoothing loss, L p is the boundary loss, L d is the distortion loss, N is the number of look-ahead frames, w p , w d is the loss weight, w p,i is the normalized Gaussian weight centered on the current frame, α is the tolerable reference salient value, Ω(R v ,R r ) is the spherical angle between the current virtual and real camera postures, β0 is the preset threshold, and β1 is the control parameter.

8. An online fusion image stabilization device for unmanned aerial vehicles, characterized in that: The device is used to implement an online fusion image stabilization method for a drone as described in any one of claims 1 to 7, the device comprising: The data acquisition module is used to extract single-frame images from the original video taken by the drone and obtain IMU data from the drone; the data processing module is used to pre-process the RGB information and grayscale information of the single-frame image, extract key feature points from the single-frame image, and generate the IMU motion trajectory; A feature matching module is used to use a feature matching network to perform feature matching between two adjacent single-frame images based on key feature points to generate unstable image trajectories; A fusion module is used to fuse the unstable image trajectory and the IMU motion trajectory to obtain a fused trajectory; A trajectory processing module, used for predicting a smooth camera trajectory based on the fused trajectory, converting the smooth camera trajectory into grid motion data, and generating a stable image sequence based on the grid motion data and a single frame image; The video reconstruction module is used to synthesize stable videos based on stable image sequences.

9. Computer device, characterized in that include: Memory for storing computer programs; A processor, configured to execute the computer program to implement an online fusion image stabilization method for a drone as described in any one of claims 1 to 7.

10. A storage medium, characterized in that Used to store a computer program, which, when executed, implements an online fusion image stabilization method for drones as described in any one of claims 1 to 7.