Hybrid anti-shake method and system combining optical anti-shake and electronic anti-shake

By synchronously fusing optical image stabilization position sensors with multi-axis IMUs and image data, and combining convolutional neural network analysis to predict shake trends, a close collaboration between optical and electronic image stabilization is achieved. This solves the problems of rough motion perception and low collaboration efficiency in existing technologies, and improves image stability and clarity.

CN121888098APending Publication Date: 2026-04-17YUANYU VISION TECHNOLOGY (SHENZHEN) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610089018.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-01-22
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing hybrid image stabilization technologies rely on the IMU as the sole or primary motion sensing source, resulting in insufficient precision in distinguishing the external environment, a lack of deep closed-loop coupling, low collaborative efficiency, and an inability to dynamically and finely adjust strategies based on the shooting scene.

Method used

By fusion of the actual displacement feedback from the optical image stabilization position sensor with multi-axis IMU and image data at a high precision in hardware, a motion model is constructed. Combined with convolutional neural network analysis of scene features, future shaking trends are predicted, achieving close collaboration between optical and electronic image stabilization.

Benefits of technology

It achieves precise differentiation between intentional and unintentional shaking, dynamically adjusts the stabilization strategy, reduces system latency and compensation errors, improves image stability and clarity, reduces image cropping and rolling shutter effect, and provides an extremely stable and high-definition imaging experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121888098A_ABST
    Figure CN121888098A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of image and video anti-shake, and discloses a hybrid anti-shake method and system combining optical anti-shake and electronic anti-shake. The hybrid anti-shake method is applied to a hybrid anti-shake unit, and specifically comprises the following steps: S11, acquiring original data of angular velocities and linear accelerations of a terminal device in three axial directions acquired by a multi-axis IMU, and acquiring actual displacement vectors of a position sensor and an image sensor in the terminal device during optical anti-shake execution; and establishing a uniform timestamp to carry out precision synchronous alignment on the data of the multi-axis IMU, the optical anti-shake position sensor and the image sensor, and generating a motion track estimation feature. According to the method, the actual displacement feedback of the optical anti-shake position sensor is used as an independent sensing source, hardware-level high-precision synchronous fusion is carried out on the independent sensing source, the multi-axis IMU and the image data to construct the motion model, intentional motion and non-intentional shake are accurately distinguished from the source, and the problems of rough motion sensing and weak cooperative basis are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the technical field of image and video stabilization, and more specifically, to a hybrid stabilization method and system that combines optical stabilization and electronic stabilization. Background Technology

[0002] Current mobile imaging devices generally incorporate image stabilization technology to address camera shake during handheld shooting. The mainstream solutions mainly include Optical Image Stabilization (OIS) and Electronic Image Stabilization (EIS). Optical image stabilization detects device movement using inertial measurement units (IMUs) such as gyroscopes and drives the lens or image sensor to move in the opposite direction of the shake, directly compensating for the shake at the optical path level. This stabilizes the image projected onto the sensor during exposure, effectively improving single-frame image quality, especially beneficial for low-light shooting. However, its compensation range is limited by physical structure and cannot correct rotation and rolling shutter distortion. Electronic image stabilization, on the other hand, operates entirely at the image data processing level. It typically estimates global motion by using IMU data and / or analyzing pixel motion between consecutive frames, and then performs inverse translation, rotation, and cropping transformations on the image to output a stable image. EIS has the advantage of a theoretically large compensation range and the ability to handle complex motion, but it loses effective pixels, may introduce rolling shutter effect, and exacerbates image noise in low light due to insufficient image information and cropping / enlargement.

[0003] To leverage the strengths of both technologies and compensate for their weaknesses, hybrid image stabilization schemes have emerged in the current technology, allowing OIS and EIS to work in "segments": OIS acts as the first line of defense, handling larger amplitude or low-frequency jitter; "residual jitter" that is not completely canceled by OIS is then handled by EIS for secondary processing. In this architecture, OIS is usually independently controlled by IMU data, while EIS independently calculates its required compensation based on IMU and / or video frame analysis. Although the two work together, their data processing paths are relatively independent, and information exchange is limited. Essentially, it is a loosely connected or simple division of labor.

[0004] Although existing hybrid image stabilization technologies have achieved some success, the following technical problems still exist: 1. Existing hybrid anti-shake solutions mainly rely on the IMU as the sole or primary motion sensing source to control OIS. However, the dynamic response characteristics and actual displacement of the OIS actuator itself are not deeply integrated into the global motion estimation as key feedback information, resulting in insufficient accuracy in the system's distinction of the external environment. 2. Most systems use a fixed OIS / EIS working mode or switch based on a simple threshold, and cannot make dynamic and fine-grained strategy adjustments according to the complex content of the shooting scene, the user's creative intention, and the specific spectral characteristics of the shake. 3. The collaboration between OIS and EIS is lagging. EIS usually passively processes the results after OIS compensation. There is a lack of deep closed-loop coupling between the two based on prediction models. The phase delay problem of OIS has not been fundamentally solved. Furthermore, EIS cannot know the precise compensation effect of OIS in advance, resulting in blindness in its processing and low overall collaboration efficiency.

[0005] Based on the above problems, there is an urgent need to design a hybrid image stabilization method that is more closely coordinated and has better image quality in order to break through the technical bottleneck. Summary of the Invention

[0006] The purpose of this invention is to provide a hybrid image stabilization method and system that combines optical image stabilization and electronic image stabilization. By using the actual displacement feedback of the optical image stabilization position sensor as an independent sensing source, and fusion it with multi-axis IMU and image data at the hardware level with high precision to construct a motion model, the invention accurately distinguishes between intentional motion and unintentional jitter from the source, thus solving the problems of coarse motion perception and weak coordination foundation.

[0007] This invention is implemented as follows: a hybrid image stabilization method combining optical and electronic image stabilization, applied to a hybrid image stabilization unit, specifically includes the following steps: S11: Acquire the raw data of angular velocity and linear acceleration of the terminal device in three axes collected by the multi-axis IMU, and the actual displacement vector of the position sensor and image sensor inside the terminal device when performing optical image stabilization. Establish a unified timestamp to synchronize and align the data of the multi-axis IMU, optical image stabilization position sensor and image sensor with high accuracy, and generate motion trajectory estimation features. S12: Analyze the preview image obtained by the image sensor, use the convolutional neural network to identify the feature elements in the scene, combine the motion trajectory estimation features to determine the current type of shaking, and decide the weight allocation ratio and compensation strategy of optical image stabilization and electronic image stabilization in the current image frame. S13: Drive the optical image stabilization unit according to the compensation strategy, feed back the actual displacement vector and input the raw data collected by the multi-axis IMU into the preset advanced motion controller, calculate the amount of optical compensation to offset the current shaking, and predict the shaking trend in the next few milliseconds based on the historical data of the motion trajectory, drive the optical image stabilization unit to enter the expected compensation position in advance, and generate the expected residual motion vector map for electronic image stabilization. S14: The electronic image stabilization unit receives the expected residual motion vector map and image frame, decomposes the residual jitter into global translation, rotation and non-rigid deformation components, applies strong constraint stabilization to the subject in the picture based on scene semantics, applies flexible deformation correction and filling to the background, reduces image cropping, rolling effect and edge deformation while eliminating jitter, and outputs the processed image frame. S15: The processed image frame is fused with the stabilized original image frame in the temporal domain. A machine learning-based image quality evaluation model is introduced to evaluate the sharpness, noise level and motion blur of different regions of each frame. Based on the evaluation results, the fusion weight is dynamically determined to complete the hybrid image stabilization compensation and output a high-resolution and visually stable image or video stream.

[0008] Further, in S11, the raw data of angular velocity and linear acceleration of the terminal device in three axes acquired by the multi-axis IMU, as well as the actual displacement vectors of the position sensor and image sensor inside the terminal device during optical image stabilization, are acquired, including: The terminal device uses a built-in high-frequency multi-axis IMU to acquire raw angular velocity data of the device in the pitch, yaw and roll axes, as well as the corresponding raw acceleration data of the three axes, in real time at a sampling rate of no less than 1000Hz. Simultaneously, a high-precision position sensor integrated within the optical image stabilization unit captures the actual displacement vector of the lens or sensor compensation component in the plane in real time. The actual displacement vector accurately reflects the amount of optical image stabilization performed. At the same time, the raw image data stream generated by the image sensor during exposure is analyzed in real time to extract the global motion vector within the frame as a motion reference at the visual level.

[0009] Furthermore, establishing a unified timestamp allows for precise synchronization and alignment of data from the multi-axis IMU, optical image stabilization position sensor, and image sensor, generating motion trajectory estimation features, including: Initiate a hardware-level synchronization and fusion process, and use a dedicated timing controller to stamp all sensor data with a unified timestamp with microsecond-level precision to ensure that asynchronous data streams from IMU, position sensor and image sensor are strictly aligned on the timeline. The aligned data is preprocessed, with temperature compensation and noise reduction performed on the IMU data, calibration and correction performed on the position sensor data, and the pre-processed information is input into a lightweight filtering algorithm to generate motion trajectory estimation features of the fused motion state snapshot, providing the underlying data foundation for synchronous calibration for subsequent collaborative image stabilization decisions.

[0010] Furthermore, in S12, a convolutional neural network is used to identify feature elements in the scene, and the type of current jitter is determined by combining motion trajectory estimation features, including: A lightweight convolutional neural network deployed at the front end of the image signal processor performs frame-by-frame analysis of the real-time preview frames. The multi-branch architecture of the lightweight convolutional neural network extracts the following key features in parallel: the target detection branch identifies the main subject and outputs its bounding box, position coordinates and size ratio. Optical flow-assisted branch estimation of the independent motion velocity of the main body; The texture analysis branch determines whether a scene contains dense high-frequency details or regular lines through frequency domain transformation; The environmental perception branch assesses the overall illumination intensity and calculates the signal-to-noise ratio of the image; The fusion engine integrates the extracted semantic features with the motion trajectory estimation features from the multi-axis IMU in a multimodal manner. The features are weighted through an attention mechanism and input into the classification model to accurately classify the current jitter into low-frequency oscillation, high-frequency vibration, compound jitter, or regular scanning motion, and output the corresponding quantitative control parameters for each anti-shake module.

[0011] Furthermore, in S13, the actual displacement vector feedback and the raw data acquired by the multi-axis IMU are input into a preset advanced motion controller to calculate the optical compensation amount to counteract the current jitter, including: The actual displacement vector feedback from the optical image stabilization position sensor and the raw data of angular velocity and acceleration collected by the multi-axis IMU are synchronously input to the preset advanced motion controller. The advanced motion controller has a built-in accurate dynamic model of the controlled object. By solving the optimization problem with the goal of minimizing image drift, it calculates in real time the accurate compensation force and displacement command required to drive the optical image stabilization component. The output of the advanced motion controller is modulated in real time by the upper-level strategy engine, which dynamically distributes weight coefficients and constraints based on the joint recognition of scene content and jitter type.

[0012] Furthermore, by predicting the shake trend within the next few milliseconds, the optical image stabilization unit is driven to the expected compensation position in advance, generating the expected residual motion vector map for electronic image stabilization, including: The advanced motion controller constructs a short-time domain motion prediction model based on the synchronized and fused motion data stream. The prediction model analyzes the historical sequence and spectral characteristics of the motion to deduce the most likely angular and linear displacement trends of the terminal device within a preset 2-5 milliseconds. Based on this prediction result, the advanced motion controller outputs drive commands in advance, so that the optical image stabilization unit begins to move towards the predicted compensation position before the actual disturbance arrives. While generating optical image stabilization prediction commands, the advanced motion controller synchronously calculates and outputs the expected residual motion vector map. Each vector in the expected residual motion vector map represents the amount of displacement that is expected to remain at each point on the image plane after the predictive optical compensation is completed. This information is transmitted to the electronic image stabilization pipeline in real time as prior information, so that the electronic image stabilization unit can know the effect and limitations of optical compensation in advance.

[0013] Furthermore, in S14, the residual jitter is kinematically decomposed into global translation, rotation, and non-rigid deformation components. Based on scene semantics, strong constraint stabilization is applied to the main subject in the image, and flexible deformation correction and filling are applied to the background, including: The expected residual motion vector map is received from the advanced motion controller and subjected to high-order motion decomposition. The decomposition separates the global translation and rotation parameters applicable to the entire map. Furthermore, by analyzing the local inconsistencies of the vector field, the non-rigid deformation components caused by off-center rotation or complex vibration of the terminal device are resolved. Differential processing is performed on the foreground subject and background regions identified by the real-time semantic segmentation map. The subject region is stabilized by strong constraint with the target as the center and locked in a specific position in the image. For the background, adaptive mesh deformation correction is performed based on non-rigid deformation components, and context-aware image repair is used to fill in missing pixels caused by deformation and cropping. When fusing multiple frames, weights are dynamically allocated based on the importance of the regions. The clearest frame data is used first for the main region to achieve the optimal balance between stability and image integrity.

[0014] Furthermore, in S15, a machine learning-based image quality evaluation model is introduced to evaluate the sharpness, noise level, and motion blur of different regions in each frame, including: Before temporal fusion, a pre-trained lightweight machine learning image quality evaluation model is called. The evaluation model is guided by the foreground-background segmentation map, divides the image into multiple perceptual unit grids, and independently evaluates the sharpness, noise level and motion blur of each unit to generate the corresponding regionalized multidimensional image quality score map. In the core multi-frame fusion stage, the image quality score map is combined with motion alignment data to dynamically determine the fusion weight of each frame and each pixel region, prioritizing the frame with the highest score in the subject area as the main data source; for high-noise background regions identified in consecutive frames, temporal adaptive weighted averaging is used for noise reduction; for motion-blurred regions, high-frequency details are extracted from adjacent clearer frames, and convolutional neural networks are used for information transmission and reconstructive sharpening to achieve overall image quality improvement.

[0015] Compared with existing technologies, the hybrid image stabilization method and system combining optical and electronic image stabilization provided by this invention have the following beneficial effects: 1. By using the actual displacement feedback of the optical image stabilization (OIS) position sensor as an independent sensing source and fusioning it with multi-axis IMU and image data at a hardware level with high precision to construct a motion model, the system accurately distinguishes between intentional motion and unintentional shaking from the source, solving the problems of coarse motion perception and weak collaborative foundation. 2. By introducing scene semantic analysis and multimodal decision-making based on convolutional neural networks, the stabilization strategy can be dynamically and intelligently adjusted according to the subject's state, ambient light, and shaking spectrum characteristics. This allows for priority use of OIS to preserve image quality in low light and limitation of OIS travel during tracking shots to avoid interference, changing the fixed threshold switching mode and enabling the stabilization system to truly possess scene perception capabilities. 2. By establishing an optical-electronic feedforward control based on a short-term motion prediction model, an expected residual motion vector map is generated, enabling electronic image stabilization to anticipate the precise effect of optical compensation. This achieves deep collaboration between the two at the millisecond level, reducing overall system latency and compensation errors. Furthermore, electronic image stabilization employs refined processing by region and object, combined with non-rigid deformation correction and content-aware fill, to completely eliminate the sense of shake while minimizing image cropping, rolling shutter effect, and edge distortion. This significantly improves the integrity, naturalness, and usable resolution of the image. Finally, a machine learning image quality evaluation model is introduced in the final temporal fusion stage, achieving dynamic fusion based on multi-dimensional evaluations such as sharpness, noise, and motion blur. This ensures that the output image is stable, and its overall image quality even surpasses the limits of single-frame shooting, providing users with an imaging experience that combines ultimate stability, high-definition image quality, and a smooth creative feel.

[0016] A hybrid image stabilization system combining optical and electronic image stabilization, used to perform the above-described hybrid image stabilization method, the hybrid image stabilization system comprising: The multi-source data acquisition module is used to simultaneously acquire and fuse data from the inertial measurement unit, optical image stabilization position sensor, and image sensor to generate a high-precision estimate of the device's motion trajectory. The scene analysis and decision module is used to identify the scene semantic features of the preview screen, and combine the motion trajectory estimation to determine the type of shaking, and dynamically generate the collaborative weight and compensation strategy of optical image stabilization and electronic image stabilization. An optical-electronic control module is used to drive the optical image stabilization unit to perform advance compensation based on a prediction model according to the compensation strategy, and simultaneously generate the expected residual motion vector map. The electronic image processing module is used to perform partitioned motion decomposition and differential stabilization processing on the image frame based on the expected residual motion vector map and scene semantics. The temporal fusion optimization module is used to perform dynamic weighted fusion of processed multi-frame images based on image quality evaluation, and output visually stable and image quality enhanced image or video streams.

[0017] Specifically, the optical-electronic control module includes: Predictive control unit, used to build short-time domain motion prediction model to drive the optical image stabilization unit to the expected compensation position in advance; The residual map generation unit is used to calculate the expected residual displacement of each pixel on the image plane at the same time as the optical compensation command is generated, form the expected residual motion vector map, and pass it to the subsequent electronic image stabilization pipeline. Attached Figure Description

[0018] Figure 1 This is a flowchart illustrating a hybrid image stabilization method combining optical and electronic image stabilization proposed in this invention. Figure 2 This is a schematic diagram illustrating the process of acquiring multi-axis IMU data and the actual displacement vector during optical image stabilization in a hybrid image stabilization method combining optical and electronic image stabilization proposed in this invention. Figure 3 This is a schematic diagram of the structure of a hybrid image stabilization system combining optical and electronic image stabilization proposed in this invention; Figure 4 This is a schematic diagram of the optical-electronic control module in a hybrid image stabilization system that combines optical and electronic image stabilization, as proposed in this invention. Detailed Implementation

[0019] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0020] The implementation of the present invention will be described in detail below with reference to specific embodiments.

[0021] In the accompanying drawings of this embodiment, the same or similar reference numerals correspond to the same or similar components. In the description of this invention, it should be understood that if terms such as "upper," "lower," "left," and "right" indicate the orientation or positional relationship based on the orientation or positional relationship shown in the drawings, they are only for the convenience of describing this invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, the terms used to describe positional relationships in the drawings are only for illustrative purposes and should not be construed as limiting this invention. For those skilled in the art, the specific meaning of the above terms can be understood according to the specific circumstances.

[0022] Reference Figure 1-2 As shown, a hybrid image stabilization method combining optical and electronic image stabilization, applied to a hybrid image stabilization unit, specifically includes the following steps: S11: Acquire the raw data of angular velocity and linear acceleration of the terminal device in three axes collected by the multi-axis IMU, and the actual displacement vector of the position sensor and image sensor inside the terminal device when performing optical image stabilization. Establish a unified timestamp to synchronize and align the data of the multi-axis IMU, optical image stabilization position sensor and image sensor with high accuracy, and generate motion trajectory estimation features. Among them, the high-frequency multi-axis IMU built into the terminal device acquires the raw angular velocity data of the device in the three axes of pitch, yaw and roll in real time, as well as the corresponding raw acceleration data of the three axes, at a sampling rate of no less than 1000Hz. Simultaneously, a high-precision position sensor integrated inside the optical image stabilization unit is used to capture the actual displacement vector of the lens or sensor compensation component in the plane in real time. The actual displacement vector accurately reflects the amount of optical image stabilization performed. At the same time, the raw image data stream generated by the image sensor during exposure is analyzed in real time to extract the global motion vector within the frame as a motion reference at the visual level. S12: Analyze the preview image obtained by the image sensor, use the convolutional neural network to identify the feature elements in the scene, and combine the motion trajectory estimation features to determine whether the current jitter is a low-frequency large amplitude swing, a high-frequency micro vibration or a compound jitter. Based on scene analysis and jitter type identification, decide the weight allocation ratio and compensation strategy of optical image stabilization and electronic image stabilization in the current time image frame. Among these methods, convolutional neural networks are used to identify feature elements in the scene, and motion trajectory estimation features are combined to determine the type of current jitter, including: A lightweight convolutional neural network deployed at the front end of the image signal processor performs frame-by-frame analysis of the real-time preview frames. The multi-branch architecture of the lightweight convolutional neural network extracts the following key features in parallel: the target detection branch identifies the main subject and outputs its bounding box, position coordinates and size ratio. Optical flow-assisted branch estimation of the independent motion velocity of the main body; The texture analysis branch determines whether a scene contains dense high-frequency details or regular lines through frequency domain transformation; The environmental perception branch assesses the overall illumination intensity and calculates the signal-to-noise ratio of the image; The fusion engine uses multimodal fusion to combine the extracted semantic features with the motion trajectory estimation features from the multi-axis IMU. The features are weighted through an attention mechanism and input into the classification model to accurately classify the current jitter into low-frequency swaying, high-frequency vibration, compound jitter or regular scanning motion, and output the corresponding quantitative control parameters of each anti-shake module. The formula for calculating the intention-weighted eigenvector is as follows: ; Z∈R df The weighted context vector is an aggregated representation of all input features, weighted according to their importance, and serves as the core input to the classification model. f i The i-th input feature vector n represents the total number of features, which come from the multimodal fusion engine and include: Scene semantic features derived from convolutional neural networks (CNNs) (such as subject size, texture complexity, and lighting estimation). Motion trajectory characteristics from the inertial measurement unit (IMU) (such as angular velocity spectrum energy, motion periodicity); α i∈[0,1]: Attention weights, which measure the importance of the ii-th feature to the current jitter classification task, and the sum of all weights is 1; S13: Drive the optical image stabilization unit according to the compensation strategy, feed back the actual displacement vector and input the raw data collected by the multi-axis IMU into the preset advanced motion controller, calculate the amount of optical compensation to offset the current shaking, and predict the shaking trend in the next few milliseconds based on the historical data of the motion trajectory, drive the optical image stabilization unit to enter the expected compensation position in advance, and generate the expected residual motion vector map for electronic image stabilization. This includes predicting the shake trend within the next few milliseconds and driving the optical image stabilization unit to the expected compensation position in advance, generating the expected residual motion vector map for electronic image stabilization, including: The advanced motion controller constructs a short-time domain motion prediction model based on the synchronized and fused motion data stream. The prediction model analyzes the historical sequence and spectral characteristics of the motion to deduce the most likely angular and linear displacement trends of the terminal device within a preset 2-5 milliseconds. Based on this prediction result, the advanced motion controller outputs drive commands in advance, so that the optical image stabilization unit begins to move towards the predicted compensation position before the actual disturbance arrives. While generating optical image stabilization prediction commands, the advanced motion controller synchronously calculates and outputs the expected residual motion vector map. Each vector in the expected residual motion vector map represents the amount of displacement that is expected to remain at each point on the image plane after the predictive optical compensation is completed. This information is transmitted to the electronic image stabilization pipeline in real time as prior information, so that the electronic image stabilization unit can know the effect and limitations of optical compensation in advance. The formula for extracting the expected residual motion is: , , ; Where, θ k pred =[θx, θy, θz] k T : The predicted angular displacement vector output by the MPC prediction model at time k; θ k comp The actual compensated angular displacement vector calculated by MPC and executed by the optical image stabilization unit; θ k res : Residual angular displacement vector, i.e., the residual rotation that optical image stabilization cannot fully compensate for; t k pred =[tx, ty, tz] k T Predicted linear displacement vector; t k comp: The actual compensation linear displacement vector of optical image stabilization (translation can be compensated by sensor-shifting OIS); tk k res : Residual linear displacement vector; ω k meas : The angular velocity vector measured in real time by the IMU; ω k comp : The compensated angular velocity vector calculated from the equivalent motion of OIS compensation; ω k res : Residual angular velocity vector, representing high-frequency vibration components; S14: The electronic image stabilization unit receives the expected residual motion vector map and image frame, decomposes the residual jitter into global translation, rotation and non-rigid deformation components, applies strong constraint stabilization to the subject in the picture based on scene semantics to ensure its absolute clarity and stability, applies flexible correction and filling to the background, and minimizes image cropping, rolling effect and edge deformation while eliminating the sense of jitter, and outputs the processed image frame. S15: The processed image frame is fused with the previously stabilized original image frame in the temporal domain. A machine learning-based image quality evaluation model is introduced to evaluate the sharpness, noise level, and motion blur of different regions in each frame. The fusion weight is dynamically determined based on the evaluation results, and a high-resolution and visually stable image or video stream is output.

[0023] In S11 of this embodiment, a unified timestamp is established to synchronize and align the data from the multi-axis IMU, optical image stabilization position sensor, and image sensor with high precision, generating motion trajectory estimation features, including: Initiate a hardware-level synchronization and fusion process, and use a dedicated timing controller to stamp all sensor data with a unified timestamp with microsecond-level precision to ensure that asynchronous data streams from IMU, position sensor and image sensor are strictly aligned on the timeline. The aligned data is preprocessed, with temperature compensation and noise reduction performed on the IMU data, calibration and correction performed on the position sensor data, and the pre-processed information is input into a lightweight filtering algorithm to generate motion trajectory estimation features of the fused motion state snapshot, providing the underlying data foundation for synchronous calibration for subsequent collaborative image stabilization decisions.

[0024] In S13 of this embodiment, the actual displacement vector feedback and the raw data acquired by the multi-axis IMU are input into a preset advanced motion controller to calculate the optical compensation amount to counteract the current jitter, including: The actual displacement vector feedback from the optical image stabilization position sensor and the raw data of angular velocity and acceleration collected by the multi-axis IMU are synchronously input to the preset advanced motion controller. The advanced motion controller has a built-in accurate dynamic model of the controlled object. By solving the optimization problem with the goal of minimizing image drift, it calculates in real time the accurate compensation force and displacement command required to drive the optical image stabilization component. The output of the advanced motion controller is modulated in real time by the upper-level strategy engine. The strategy engine dynamically assigns weight coefficients and constraints based on the joint recognition of scene content and shake type. For example, in low-light static scenes, the controller is instructed to prioritize the use of optical image stabilization for wide-range compensation to protect image quality; when tracking fast-moving objects, its response speed and travel are limited to prevent conflict with the subject's movement, thereby achieving intelligent collaborative image stabilization that conforms to the shooting intent.

[0025] In S14 of this embodiment, the residual jitter is kinematically decomposed into global translation, rotation, and non-rigid deformation components. Based on scene semantics, strong constraint stabilization is applied to the main subject in the image, and flexible deformation correction and filling are applied to the background, including: The expected residual motion vector map is received from the advanced motion controller and subjected to high-order motion decomposition. The decomposition separates the global translation and rotation parameters applicable to the entire map. Furthermore, by analyzing the local inconsistencies of the vector field, the non-rigid deformation components caused by off-center rotation or complex vibration of the terminal device are resolved. Differential processing is performed on the foreground subject and background regions identified by the real-time semantic segmentation map. The subject region is stabilized by strong constraint with the target as the center and locked in a specific position in the image. For the background, adaptive mesh deformation correction is performed based on non-rigid deformation components, and context-aware image repair is used to fill in missing pixels caused by deformation and cropping. When fusing multiple frames, weights are dynamically allocated based on the importance of the regions. The clearest frame data is used first for the main region to achieve the optimal balance between stability and image integrity.

[0026] In S15 of this embodiment, a machine learning-based image quality evaluation model is introduced to evaluate the sharpness, noise level, and motion blur of different regions in each frame, including: Before temporal fusion, a pre-trained lightweight machine learning image quality evaluation model is called. The evaluation model is guided by the foreground-background segmentation map, divides the image into multiple perceptual unit grids, and independently evaluates the sharpness, noise level and motion blur of each unit to generate the corresponding regionalized multidimensional image quality score map. In the core multi-frame fusion stage, the image quality score map is combined with motion alignment data to dynamically determine the fusion weight of each frame and each pixel region, prioritizing the frame with the highest score in the subject area as the main data source; for high-noise background regions identified in consecutive frames, temporal adaptive weighted averaging is used for noise reduction; for motion-blurred regions, high-frequency details are extracted from adjacent clearer frames, and convolutional neural networks are used for information transmission and reconstructive sharpening to achieve overall image quality improvement; Evaluation algorithms for generating corresponding regionalized multidimensional image quality rating maps include: Sharpness assessment: ; Where: I(u, v): grayscale or brightness value of pixel (u, v); ▽I(u, v): The gradient vector of the image at point (u, v); ‖·‖ 2 2: The L2 norm squared of the gradient, i.e., the square of the gradient magnitude, is used to measure the sharpness of edges and textures; Ws(u, v): Spatial weight factor, which has a higher weight in the segmented foreground subject region M(u, v)=1) and a weight of 1.0 in the background region, making the model pay more attention to the sharpness of the subject; |Pi j |:Grid P ij Total number of pixels within; S ij : Sharpness score of grid cell (i, j), the larger the value, the sharper the area; Noise level assessment: ; Where: Ip: An image block of size B×B within Pij; F{·} and F -1 {·}: represent Fourier transform and inverse Fourier transform, respectively; H HPF The transfer function of a high-pass filter is used to suppress low-frequency content in an image and separate out the high-frequency components that mainly contain noise and texture. median(·): The median operation takes the median of the calculated high-frequency component amplitudes to robustly estimate the noise level of the image patch. N ij : Noise level score for grid cell (i, j), the larger the value, the higher the estimated noise in that area; Motion blur assessment:

[0027] Among them: I ij : Grid P ij Image data within; F{Iij}(f): The amplitude of the spectrum of the image after Fourier transform at frequency ff; F low The low-frequency band represents the basic content and gradually changing parts of the image; F mid Mid-frequency bands represent the edge and detail information of an image; motion blur causes a significant attenuation of mid-to-high frequency energy. : Extremely small positive numbers, to prevent the denominator from being zero.

[0028] B ij The motion blur score of grid cell (i, j) indicates that the closer the value is to 1, the greater the loss of mid-frequency energy relative to low-frequency energy, and the more severe the motion blur; the closer the value is to 0, the clearer the image.

[0029] Finally, the model outputs a three-channel score map Q∈R with the same size as the mesh. G×G×3 : ; Among them, S ij norm N ij norm B ij norm These are the S mentioned above. ij N ij B ij The normalized values ​​facilitate unified weighted decision-making in subsequent fusion modules.

[0030] This technical solution uses the actual displacement feedback of the optical image stabilization (OIS) position sensor as an independent sensing source and performs hardware-level high-precision synchronous fusion with multi-axis IMU and image data to construct a motion model. This accurately distinguishes between intentional motion and unintentional shaking from the source, solving the problems of coarse motion perception and weak collaborative foundation. Secondly, it introduces scene semantic analysis and multimodal decision-making based on convolutional neural networks, enabling the image stabilization strategy to dynamically and intelligently adjust according to the subject's state, ambient lighting, and shaking spectrum characteristics. This allows for priority use of OIS to preserve image quality in low light and limitation of OIS travel during tracking shots to avoid interference. This changes the fixed threshold switching mode and enables the image stabilization system to truly possess scene perception capabilities. Reference Figure 3-4 As shown, a hybrid image stabilization system combining optical and electronic image stabilization is used to perform the above-described hybrid image stabilization method. The hybrid image stabilization system includes: The multi-source data acquisition module is used to simultaneously acquire and fuse data from the inertial measurement unit, optical image stabilization position sensor, and image sensor to generate a high-precision estimate of the device's motion trajectory. The scene analysis and decision-making module is used to identify the scene semantic features of the preview image, and combine motion trajectory estimation to determine the type of shaking, and dynamically generate the collaborative weight and compensation strategy of optical image stabilization and electronic image stabilization. The optical-electronic control module is used to drive the optical image stabilization unit to perform advance compensation based on the prediction model according to the compensation strategy, and simultaneously generate the expected residual motion vector map. The electronic image processing module is used to perform partitioned motion decomposition and differential stabilization processing on image frames based on the expected residual motion vector map and scene semantics. The temporal fusion optimization module is used to perform dynamic weighted fusion of processed multi-frame images based on image quality evaluation, outputting visually stable and image-enhanced image or video streams. In the final temporal fusion stage, a machine learning image quality evaluation model is introduced to achieve dynamic fusion based on multi-dimensional evaluation such as sharpness, noise, and motion blur. This makes the output image stable, and its overall image quality even exceeds the limit of single-frame shooting, providing users with an image experience that combines extreme stability, high-definition image quality, and a smooth creative feel.

[0031] In this embodiment, the optical-electronic control module includes: a predictive control unit, used to construct a short-time domain motion prediction model to drive the optical image stabilization unit to the expected compensation position in advance; and a residual map generation unit, used to calculate the expected residual displacement of each pixel on the image plane at the same time as the optical compensation command is generated, forming an expected residual motion vector map, and passing it to the subsequent electronic image stabilization pipeline. By establishing optical-electronic feedforward control based on the short-time domain motion prediction model, the expected residual motion vector map is generated, enabling the electronic image stabilization to know the precise effect of optical compensation in advance, realizing deep collaboration between the two at the millisecond level, reducing the overall system latency and compensation error. In addition, the electronic image stabilization adopts refined processing by region and object, combined with non-rigid deformation correction and content-aware filling, which completely eliminates the sense of shake while minimizing image cropping, rolling shutter effect and edge deformation, and greatly improving the integrity, naturalness and usable resolution of the image.

[0032] In this embodiment, the entire operation process can be automated by computer control. In each operation stage, sensors can be set up to provide signal feedback and ensure that the steps are performed sequentially. These are all conventional knowledge of current automation control, and will not be elaborated on in this embodiment.

[0033] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention should be included within the protection scope of the present invention.

Claims

1. A hybrid image stabilization method combining optical and electronic image stabilization, characterized in that, When applied to a hybrid image stabilization unit, the specific steps include: S11: Acquire the raw data of angular velocity and linear acceleration of the terminal device in three axes collected by the multi-axis IMU, and the actual displacement vector of the position sensor and image sensor inside the terminal device when performing optical image stabilization. Establish a unified timestamp to synchronize and align the data of the multi-axis IMU, optical image stabilization position sensor and image sensor with high accuracy, and generate motion trajectory estimation features. S12: Analyze the preview image obtained by the image sensor, use the convolutional neural network to identify the feature elements in the scene, combine the motion trajectory estimation features to determine the current type of shaking, and decide the weight allocation ratio and compensation strategy of optical image stabilization and electronic image stabilization in the current image frame. S13: Drive the optical image stabilization unit according to the compensation strategy, feed back the actual displacement vector and input the raw data collected by the multi-axis IMU into the preset advanced motion controller, calculate the amount of optical compensation to offset the current shaking, and predict the shaking trend in the next few milliseconds based on the historical data of the motion trajectory, drive the optical image stabilization unit to enter the expected compensation position in advance, and generate the expected residual motion vector map for electronic image stabilization. S14: The electronic image stabilization unit receives the expected residual motion vector map and image frame, decomposes the residual jitter into global translation, rotation and non-rigid deformation components, applies strong constraint stabilization to the subject in the picture based on scene semantics, applies flexible deformation correction and filling to the background, reduces image cropping, rolling effect and edge deformation while eliminating jitter, and outputs the processed image frame. S15: The processed image frame is fused with the stabilized original image frame in the temporal domain. A machine learning-based image quality evaluation model is introduced to evaluate the sharpness, noise level and motion blur of different regions of each frame. Based on the evaluation results, the fusion weight is dynamically determined to complete the hybrid image stabilization compensation and output a high-resolution and visually stable image or video stream.

2. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 1, characterized in that, In S11, the raw data of angular velocity and linear acceleration of the terminal device in three axes acquired by the multi-axis IMU, as well as the actual displacement vectors of the position sensor and image sensor inside the terminal device during optical image stabilization, are acquired, including: The terminal device uses a built-in high-frequency multi-axis IMU to acquire raw angular velocity data of the device in the pitch, yaw and roll axes, as well as the corresponding raw acceleration data of the three axes, in real time at a sampling rate of no less than 1000Hz. Simultaneously, a high-precision position sensor integrated within the optical image stabilization unit captures the actual displacement vector of the lens or sensor compensation component in the plane in real time. The actual displacement vector accurately reflects the amount of optical image stabilization performed. At the same time, the raw image data stream generated by the image sensor during exposure is analyzed in real time to extract the global motion vector within the frame as a motion reference at the visual level.

3. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 2, characterized in that, Establishing a unified timestamp allows for precise synchronization and alignment of data from multi-axis IMUs, optical image stabilization position sensors, and image sensors, generating motion trajectory estimation features, including: Initiate a hardware-level synchronization and fusion process, and use a dedicated timing controller to stamp all sensor data with a unified timestamp with microsecond-level precision to ensure that asynchronous data streams from IMU, position sensor and image sensor are strictly aligned on the timeline. The aligned data is preprocessed, with temperature compensation and noise reduction performed on the IMU data, calibration and correction performed on the position sensor data, and the pre-processed information is input into a lightweight filtering algorithm to generate motion trajectory estimation features of the fused motion state snapshot, providing the underlying data foundation for synchronous calibration for subsequent collaborative image stabilization decisions.

4. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 3, characterized in that, In S12, a convolutional neural network is used to identify feature elements in the scene, and the type of current jitter is determined by combining motion trajectory estimation features, including: A lightweight convolutional neural network deployed at the front end of the image signal processor performs frame-by-frame analysis of the real-time preview frames. The multi-branch architecture of the lightweight convolutional neural network extracts the following key features in parallel: the target detection branch identifies the main subject and outputs its bounding box, position coordinates and size ratio. Optical flow-assisted branch estimation of the independent motion velocity of the main body; The texture analysis branch determines whether a scene contains dense high-frequency details or regular lines through frequency domain transformation; The environmental perception branch assesses the overall illumination intensity and calculates the signal-to-noise ratio of the image; The fusion engine integrates the extracted semantic features with the motion trajectory estimation features from the multi-axis IMU in a multimodal manner. The features are weighted through an attention mechanism and input into the classification model to accurately classify the current jitter into low-frequency oscillation, high-frequency vibration, compound jitter, or regular scanning motion, and output the corresponding quantitative control parameters for each anti-shake module.

5. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 4, characterized in that, In S13, the actual displacement vector feedback and the raw data acquired by the multi-axis IMU are input into a preset advanced motion controller to calculate the optical compensation amount to counteract the current jitter, including: The actual displacement vector feedback from the optical image stabilization position sensor and the raw data of angular velocity and acceleration collected by the multi-axis IMU are synchronously input to the preset advanced motion controller. The advanced motion controller has a built-in accurate dynamic model of the controlled object. By solving the optimization problem with the goal of minimizing image drift, it calculates in real time the accurate compensation force and displacement command required to drive the optical image stabilization component. The output of the advanced motion controller is modulated in real time by the upper-level strategy engine, which dynamically distributes weight coefficients and constraints based on the joint recognition of scene content and jitter type.

6. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 5, characterized in that, Predicting the shake trend within milliseconds, the optical image stabilization unit is driven to the expected compensation position in advance, generating the expected residual motion vector map for electronic image stabilization, including: The advanced motion controller constructs a short-time domain motion prediction model based on the synchronized and fused motion data stream. The prediction model analyzes the historical sequence and spectral characteristics of the motion to deduce the most likely angular and linear displacement trends of the terminal device within a preset 2-5 milliseconds. Based on this prediction result, the advanced motion controller outputs drive commands in advance, so that the optical image stabilization unit begins to move towards the predicted compensation position before the actual disturbance arrives. While generating optical image stabilization prediction commands, the advanced motion controller synchronously calculates and outputs the expected residual motion vector map. Each vector in the expected residual motion vector map represents the amount of displacement that is expected to remain at each point on the image plane after the predictive optical compensation is completed. This information is transmitted to the electronic image stabilization pipeline in real time as prior information, so that the electronic image stabilization unit can know the effect and limitations of optical compensation in advance.

7. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 6, characterized in that, In S14, residual jitter is kinematically decomposed into global translation, rotation, and non-rigid deformation components. Based on scene semantics, strong constraint stabilization is applied to the main subject in the image, and flexible deformation correction and filling are applied to the background, including: The expected residual motion vector map is received from the advanced motion controller and subjected to high-order motion decomposition. The decomposition separates the global translation and rotation parameters applicable to the entire map. Furthermore, by analyzing the local inconsistencies of the vector field, the non-rigid deformation components caused by off-center rotation or complex vibration of the terminal device are resolved. Differential processing is performed on the foreground subject and background regions identified by the real-time semantic segmentation map. The subject region is stabilized by strong constraint with the target as the center and locked in a specific position in the image. For the background, adaptive mesh deformation correction is performed based on non-rigid deformation components, and context-aware image repair is used to fill in missing pixels caused by deformation and cropping. When fusing multiple frames, weights are dynamically allocated based on the importance of the regions. The clearest frame data is used first for the main region to achieve the optimal balance between stability and image integrity.

8. The hybrid image stabilization method combining optical and electronic image stabilization as described in claim 7, characterized in that, In S15, a machine learning-based image quality evaluation model is introduced to evaluate the sharpness, noise level, and motion blur of different regions in each frame, including: Before temporal fusion, a pre-trained lightweight machine learning image quality evaluation model is called. The evaluation model is guided by the foreground-background segmentation map, divides the image into multiple perceptual unit grids, and independently evaluates the sharpness, noise level and motion blur of each unit to generate the corresponding regionalized multidimensional image quality score map. In the core multi-frame fusion stage, the image quality score map is combined with motion alignment data to dynamically determine the fusion weight of each frame and each pixel region, prioritizing the frame with the highest score in the subject area as the main data source; for high-noise background regions identified in consecutive frames, temporal adaptive weighted averaging is used for noise reduction; for motion-blurred regions, high-frequency details are extracted from adjacent clearer frames, and convolutional neural networks are used for information transmission and reconstructive sharpening to achieve overall image quality improvement.

9. A hybrid image stabilization system combining optical and electronic image stabilization, characterized in that, For performing the hybrid image stabilization method according to any one of claims 1-8, the hybrid image stabilization system comprises: The multi-source data acquisition module is used to simultaneously acquire and fuse data from the inertial measurement unit, optical image stabilization position sensor, and image sensor to generate a high-precision estimate of the device's motion trajectory. The scene analysis and decision module is used to identify the scene semantic features of the preview screen, and combine the motion trajectory estimation to determine the type of shaking, and dynamically generate the collaborative weight and compensation strategy of optical image stabilization and electronic image stabilization. An optical-electronic control module is used to drive the optical image stabilization unit to perform advance compensation based on a prediction model according to the compensation strategy, and simultaneously generate the expected residual motion vector map. The electronic image processing module is used to perform partitioned motion decomposition and differential stabilization processing on the image frame based on the expected residual motion vector map and scene semantics. The temporal fusion optimization module is used to perform dynamic weighted fusion of processed multi-frame images based on image quality evaluation, and output visually stable and image quality enhanced image or video streams.

10. A hybrid image stabilization system combining optical and electronic image stabilization as described in claim 9, characterized in that, The optical-electronic control module includes: Predictive control unit, used to build short-time domain motion prediction model to drive the optical image stabilization unit to the expected compensation position in advance; The residual map generation unit is used to calculate the expected residual displacement of each pixel on the image plane at the same time as the optical compensation command is generated, form the expected residual motion vector map, and pass it to the subsequent electronic image stabilization pipeline.