Image anti-shake method, device and storage medium
By calculating the motion data of the image acquisition device and transforming the image pixels using a rotation matrix and translation compensation vector, the high cost and limited effectiveness of existing optical image stabilization and gimbal image stabilization technologies are solved, achieving efficient image stabilization and image quality preservation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHENZHEN HOLLYLAND TECH CO LTD
- Filing Date
- 2026-03-26
- Publication Date
- 2026-06-09
AI Technical Summary
Existing optical image stabilization and gimbal image stabilization technologies suffer from high costs and limited stabilization effects in handheld cameras. Optical image stabilization has a limited compensation range, while gimbal image stabilization mechanisms are large, heavy, and expensive, making them difficult to apply in miniaturized devices.
By calculating the motion data of the image acquisition device, the original coordinates of the image pixels are rotated and translated using a rotation matrix and a translation compensation vector to achieve image stabilization, avoiding reliance on mechanical structures and image cropping.
It reduces the cost of image stabilization, improves the stabilization effect, maintains the original image quality, avoids image shrinkage or resolution reduction, and achieves better image stabilization.
Smart Images

Figure CN122179663A_ABST
Abstract
Description
Technical Field
[0001] This specification relates to the field of image processing technology, and in particular to an image stabilization method, device, and storage medium. Background Technology
[0002] In handheld camera products, image stabilization technology is a crucial element in improving image quality. When shooting handheld, involuntary hand tremors or other external factors (such as ambient vibrations) can easily cause camera lens shift or rotation, resulting in blurry, shaky, or ghosting images, severely impacting image sharpness and stability. To address this issue, various image stabilization technologies have been proposed, primarily optical image stabilization and gimbal-based image stabilization. However, these technologies still have the following drawbacks: Optical image stabilization technology achieves image stabilization by adjusting the position of the lens elements or image sensor through mechanical structures when shaking is detected. However, its image stabilization compensation range is limited by the mobility of the mechanical structure, so it cannot compensate for large-amplitude shaking. Furthermore, the precision and durability of the mechanical components also affect the image stabilization effect.
[0003] Gimbal stabilization technology counteracts camera shake by controlling a multi-axis gimbal mechanism to move the entire lens assembly. However, gimbal mechanisms are large, heavy, and expensive, which limits their application in miniaturized handheld devices. Furthermore, the complexity of the gimbal mechanism's mechanical structure may increase device power consumption and the risk of failure. Summary of the Invention
[0004] In view of this, in order to at least solve the technical problems of high cost and limited stabilization effect in the related technologies for achieving image stabilization, this specification provides an image stabilization method, device and storage medium, and the technical solution adopted is as follows: According to a first aspect of the embodiments of this specification, an image stabilization method is provided, comprising: Based on the motion data of the image acquisition device when acquiring the image to be processed, calculate the rotation angle vector and position translation vector of the image acquisition device relative to the set camera coordinate system; The object distance of the focused object is calculated based on the focal length parameters of the image acquisition device and the image distance of the focused object in the image to be processed. The rotation matrix is calculated based on the X-axis angle in the rotation angle vector. The rotation matrix is used to rotate and transform the original coordinates of the image pixels to counteract the rotational motion of the image acquisition device around the X-axis of the camera coordinate system. The translation compensation vector is calculated based on the field of view, object distance, Y-axis and Z-axis angles in the rotation angle vector of the image acquisition device, and Y-axis and Z-axis translation amounts in the position translation vector. The translation compensation vector is used to offset the rotational motion of the image acquisition device around the Y-axis and Z-axis of the camera coordinate system, as well as the translational motion along the Y-axis and Z-axis. For each pixel in the image to be processed, the original coordinates of the pixel are rotated using a rotation matrix, and then the coordinates after rotation are translated using a translation compensation vector to obtain the target coordinates after pixel jitter removal. The image to be processed is updated based on the target coordinates of each pixel to obtain a stabilized image.
[0005] According to a second aspect of the embodiments of this specification, an electronic device is provided, comprising: processor; Memory used to store processor-executable instructions; The processor implements the image stabilization method provided in the first aspect of the embodiments of this specification by running executable instructions.
[0006] According to a third aspect of the embodiments of this specification, an image acquisition device is provided, comprising: Equipment body; The camera module, configured on the device itself, is used to acquire images to be processed; An inertial measurement unit, configured on the device body, is used to detect the motion data of the device body; The processor is used to process motion data transmitted based on the inertial measurement unit and to process the image to be processed using the image stabilization method provided in the first aspect of the embodiments of this specification to obtain a stabilized image.
[0007] According to a fourth aspect of the embodiments of this specification, a computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the steps of the image stabilization method provided in the first aspect of the embodiments of this specification.
[0008] The technical solutions provided in the embodiments of this specification may include the following beneficial effects: In the embodiments of this specification, for the image to be processed that requires image stabilization, the rotation angle vector and position translation vector of the image acquisition device relative to the camera coordinate system at the time of image acquisition are calculated using the motion data of the image acquisition device when acquiring the image. This converts the actual motion of the image acquisition device at the time of acquisition into rotation angle vectors and position translation vectors. Next, the X-axis angle of rotation around the X-axis of the camera coordinate system is extracted from the rotation angle vector, and a rotation matrix is calculated based on this. Then, the Y-axis and Z-axis angles in the rotation angle vector, and the Y-axis and Z-axis translation amounts in the position translation vector are combined with the field of view of the image acquisition device and the actual object distance of the focused object in the image to be processed to calculate a translation compensation vector. This ensures that the translation compensation vector accurately reflects the coordinate offset of the pixels caused by the motion of the image acquisition device. Subsequently, by first using a rotation matrix to perform a rotation transformation on the original coordinates, and then using a translation compensation vector to perform a translation transformation on the coordinates obtained from the rotation transformation of the mechanism, the actual motion of the image acquisition device—rotational motion on the three axes of the camera coordinate system, and displacement on the Y and Z axes—is transformed into a combination of translation and rotation, and mapped to pixel coordinates. This cancels out the rotational and translational motion of the image acquisition device relative to the camera coordinate system, thereby enabling the coordinates of each pixel in the image to be pulled back to their original position, thus achieving the purpose of image stabilization.
[0009] As can be seen, the image stabilization method provided in the embodiments of this specification is a technical solution that differs from optical image stabilization and gimbal stabilization. It does not rely on mechanical structures to achieve stabilization, but instead uses rotation transformation and translation to compensate for the original coordinates of pixels. It does not involve image content cropping. Therefore, it can not only reduce the cost of stabilization and ensure that the stabilization effect is not limited by mechanical structures, but also avoid the reduction in image size or resolution caused by image cropping. This is beneficial to maintaining the original image quality and avoiding image quality loss, and can better improve the stabilization effect.
[0010] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this specification. Attached Figure Description
[0011] Figure 1 This is a schematic diagram illustrating a camera coordinate system according to an exemplary embodiment of this specification; Figure 2 This is a flowchart illustrating an image stabilization method according to an exemplary embodiment of this specification; Figure 3 This is a flowchart illustrating a calculation scheme for a rotation angle vector and a position translation vector according to an exemplary embodiment of this specification; Figure 4This is a flowchart illustrating a jitter removal scheme according to an exemplary embodiment of this specification; Figure 5 This is a schematic structural diagram of an electronic device provided in this specification according to an exemplary embodiment. Detailed Implementation
[0012] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this specification. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this specification as detailed in the appended claims.
[0013] The following is an explanation of the relevant terms used in the embodiments of this specification: Camera coordinate system: A coordinate system defined by the image acquisition device (such as a camera), which moves with the movement of the image acquisition device. See also... Figure 1 , Figure 1 This is a schematic diagram of a camera coordinate system according to an exemplary embodiment of the present description; the camera coordinate system is a right-handed coordinate system with the intersection of the principal optical axis of the image acquisition device and the CMOS (Complementary Metal Oxide Semiconductor, image sensor) surface 11 as the origin O, the principal optical axis as the X-axis, and the height of the CMOS as the Z-axis. Figure 1 In the middle, 12 represents the lens, and O' represents the optical center.
[0014] Reference coordinate system: refers to the coordinate system formed by rotating the camera coordinate system Z-axis to the direction of gravity at the moment of image stabilization. Its origin coincides with the origin of the camera coordinate system at the moment of image stabilization. Its Z-axis points to the direction of gravity, and its X-axis points to the projection of the camera coordinate system onto the vertical plane of gravity. The reference coordinate system does not move with the image acquisition device.
[0015] Inertial Measurement Unit (IMU): Generally includes a gyroscope and an accelerometer, it can measure the motion data of an object, including angular velocity and acceleration. Its sampling frequency is usually higher than that of the image acquisition device. Therefore, one image sampling interval of the image acquisition device includes multiple data sampling points of the IMU. For example, assuming the sampling frequency of the IMU is 100Hz and the sampling frequency of the image acquisition device is 30Hz, based on this, there are 3 data sampling points in one image sampling interval (such as [t, t+1)). That is, within this time interval, the IMU can collect 3 motion data.
[0016] The embodiments described in this specification will now be described in detail.
[0017] To address the high cost and limited effectiveness of image stabilization in related technologies, this specification provides an image stabilization method. For an image to be processed, the method utilizes motion data from the image acquisition device during image acquisition to calculate the rotation angle vector and translation vector of the image acquisition device relative to the camera coordinate system at the time of image acquisition. This converts the actual motion of the image acquisition device at the acquisition time into rotation angle and translation vectors. Next, the X-axis angle of rotation around the X-axis of the camera coordinate system is extracted from the rotation angle vector, and a rotation matrix is calculated based on this. Furthermore, the Y-axis and Z-axis angles from the rotation angle vector, and the Y-axis and Z-axis translation amounts from the translation vector, are combined with the field of view of the image acquisition device and the actual object distance to the focused object in the image to calculate a translation compensation vector. This ensures that the translation compensation vector accurately reflects the coordinate offset of pixels caused by the motion of the image acquisition device. Subsequently, by first using a rotation matrix to perform a rotation transformation on the original coordinates, and then using a translation compensation vector to perform a translation transformation on the coordinates obtained from the rotation transformation of the mechanism, the actual motion of the image acquisition device—rotational motion on the three axes of the camera coordinate system, and displacement on the Y and Z axes—is transformed into a combination of translation and rotation, and mapped to pixel coordinates. This cancels out the rotational and translational motion of the image acquisition device relative to the camera coordinate system, thereby enabling the coordinates of each pixel in the image to be pulled back to their original position, thus achieving the purpose of image stabilization.
[0018] As can be seen, the image stabilization method provided in the embodiments of this specification is a technical solution that differs from optical image stabilization and gimbal stabilization. It does not rely on mechanical structures to achieve stabilization, but instead uses rotation transformation and translation to compensate for the original coordinates of pixels. It does not involve image content cropping. Therefore, it can not only reduce the cost of stabilization and ensure that the stabilization effect is not limited by mechanical structures, but also avoid the reduction in image size or resolution caused by image cropping. This is beneficial to maintaining the original image quality and avoiding image quality loss, and can better improve the stabilization effect.
[0019] The following combination Figure 2 The image stabilization method provided in the embodiments of this specification will be described. Figure 2 This is a flowchart illustrating an image stabilization method according to an exemplary embodiment, the image stabilization method comprising the following steps: In step S100, based on the motion data of the image acquisition device when acquiring the image to be processed, the rotation angle vector and position translation vector of the image acquisition device relative to the set camera coordinate system are calculated. In step S200, the object distance of the focused object is calculated based on the focal length parameters of the image acquisition device and the image distance of the focused object in the image to be processed. In step S300, a rotation matrix is calculated based on the X-axis angle in the rotation angle vector. The rotation matrix is used to perform rotation transformation on the original coordinates of the image pixels to counteract the rotational motion of the image acquisition device around the X-axis of the camera coordinate system. In step S400, a translation compensation vector is calculated based on the field of view, object distance, Y-axis and Z-axis angles in the rotation angle vector of the image acquisition device, and Y-axis and Z-axis translation amounts in the position translation vector. The translation compensation vector is used to offset the rotational motion of the image acquisition device around the Y-axis and Z-axis of the camera coordinate system, as well as the translational motion along the Y-axis and Z-axis. In step S500, for each pixel in the image to be processed, the original coordinates of the pixel are rotated by a rotation matrix, and then the coordinates after rotation are translated by a translation compensation vector to obtain the target coordinates after pixel jitter removal. In step S600, the image to be processed is updated according to the target coordinates of each pixel to obtain a stabilized image.
[0020] The image stabilization method provided in the embodiments of this specification can be applied to image acquisition devices or other devices with image processing functions, but is not limited thereto. The image stabilization method provided in the embodiments of this specification can be used for real-time stabilization of images to be processed acquired in real time by an image acquisition device, or it can be used for post-processing. Post-processing can be understood as: during the acquisition of images by the image acquisition device, the images to be processed are not processed initially, but can be processed later when the image acquisition device is in standby mode. Alternatively, partial processing of the images to be processed can be performed according to user needs. For example, a user can select the images or videos to be processed from image or video storage space and then trigger the image stabilization function to process the selected images or videos using the image stabilization method provided in the embodiments of this specification.
[0021] During the process of acquiring images to be processed by the image acquisition device, motion data of the image acquisition device during operation can be acquired in real time by a motion detection unit configured in the image acquisition device. The motion detection unit can be an IMU, but is not limited to this.
[0022] Furthermore, the images acquired by the image acquisition device carry acquisition time information, and similarly, the motion data acquired by the IMU also carries acquisition time information. This allows the system to align the image to be processed with the corresponding motion data during subsequent image stabilization processing, based on the acquisition time information of both the image to be processed and the motion data. This ensures the accuracy of the analysis of the motion data acquired by the image acquisition device during image acquisition, thereby guaranteeing the image stabilization effect.
[0023] However, in the above-mentioned case, because the images to be processed and motion data are stored separately, the motion data for the corresponding time needs to be acquired based on the acquisition time information of the images to be processed during the image stabilization process. That is, a time matching operation needs to be performed. Although this operation will not have a significant impact on the computing performance of the image acquisition device, when the thread is busy, such as when a large number of images to be processed need to be stabilized, the same number of time matching operations as the number of images to be processed will need to be performed. Therefore, in some embodiments, in order to ensure the computing performance of the device and avoid increased resource consumption, when storing the images to be processed, the motion data corresponding to the time of the images to be processed can be stored together in the attribute information of the images to be processed. That is, the images to be processed can carry both the acquisition time information and the associated motion data at the same time.
[0024] Based on this, in the process of image stabilization processing of the image to be processed using the image stabilization method provided in the embodiments of this specification, motion data of the image acquisition device when acquiring the image to be processed can be obtained through any of the above-described implementation methods. Then, step S100 is executed, using the motion data corresponding to the image to be processed, to calculate the rotation angle vector and position translation vector of the image acquisition device relative to the camera coordinate system when acquiring the image to be processed. In this regard, the embodiments of this specification provide two technical solutions for calculating the rotation angle vector and position translation vector from different perspectives, as follows: The first method for calculating the rotation angle vector and the position translation vector: This calculation scheme primarily focuses on computational efficiency, aiming to optimize computational efficiency while maintaining a certain level of accuracy. Based on this, in step S100 above, the rotation angle vector and translation vector of the image acquisition device relative to the set camera coordinate system are calculated based on the motion data of the image acquisition device during image acquisition. This includes: In step S111, the rotation angle vector is calculated based on the angular velocity acquired by the inertial measurement unit configured in the image acquisition device at the acquisition time of the image to be processed, and an image sampling interval with the acquisition time as the endpoint. In step S112, the position translation vector is calculated based on the acceleration collected by the inertial measurement unit at the acquisition time and the next consecutive acquisition time.
[0025] The technical principle of the first calculation scheme described in steps S111 and S112 above is explained below: Assuming the image to be processed is acquired at time t, in order to execute steps S111 and S112, the angular velocity at time t can be obtained first. and the acceleration at time t and t+1 and Next, steps S111 and S112 can be executed in parallel or sequentially. When executed sequentially, the order of execution is not limited.
[0026] During step S111, the time interval Δt (i.e., one image sampling interval) between time t-1 and time t is used to calculate the rotation angle vector. The calculation principle can be found in relevant technologies, such as through the formula. The rotation vector in the reference coordinate system is calculated. Then, based on the transformation matrix from the reference coordinate system to the camera coordinate system, the rotation vector is... Transform to the camera coordinate system to obtain the rotation angle vector. ,in, This represents the angle by which the image acquisition device rotates around the X-axis of the camera coordinate system at time t, and will be referred to as the X-axis angle below. This represents the angle by which the image acquisition device rotates around the Y-axis of the camera coordinate system at time t, and will be referred to as the Y-axis angle below. This represents the angle by which the image acquisition device rotates around the Z-axis of the camera coordinate system at time t, and will be referred to as the Z-axis angle below.
[0027] During the execution of step S112, acceleration was used. and To calculate the position translation vector, the calculation principle can be found in relevant techniques. For example, first, the attitude matrix is calculated based on the aforementioned rotation vector using the Rodrigues rotation formula. This attitude matrix refers to the matrix that transforms the camera coordinate system to the reference coordinate system. Then, the attitude matrix and... After multiplying and subtracting the acceleration due to gravity, use Integrating the acceleration obtained after subtracting the gravitational deceleration and adding it to the velocity calculated at time t-1, we can obtain the velocity at time t. Then, the formula can be used. The position of the image acquisition device relative to the reference coordinate system at time t is calculated. ,in, Let t be the position of the image acquisition device relative to the reference coordinate system at time t-1. Similarly, the position of the image acquisition device relative to the reference coordinate system at time t+1 can be calculated. .
[0028] get and Next, the camera displacement change of the image acquisition device relative to the reference coordinate system from time t to time t+1 is calculated. Then, the camera displacement change is transformed back to the camera coordinate system using the aforementioned transformation matrix, yielding the position translation vector of the image acquisition device relative to the camera coordinate system at time t. ,in, The translation of the image acquisition device relative to the X-axis of the camera coordinate system at time t is referred to as the X-axis translation. Let Y be the translation of the image acquisition device relative to the camera coordinate system at time t, hereinafter referred to as the Y-axis translation. The translation of the image acquisition device relative to the Z-axis of the camera coordinate system at time t is referred to as the Z-axis translation.
[0029] In calculating the position translation vector, considering its significant impact on pixel coordinate compensation, using data from time t-1 and time t to calculate the vector might result in motion blur or lag in the final compensated image. However, by using data from time t and time t+1 to calculate the vector, system latency (the delay from image acquisition to display, causing the image to appear after time t) can be offset. This calibrates the image display state from time t to a point after time t (such as time t+1 or a point between time t and t+1), allowing the user's current physical viewpoint to nearly coincide with or overlap with the physical viewpoint displayed in the stabilized image. This enables the image to follow the user's camera rotation synchronously, greatly improving the tracking and smoothness of handheld shooting.
[0030] The second method for calculating the rotation angle vector and the position translation vector: Because the raw data acquired by the IMU contains high-frequency random noise, such as the IMU's own electronic noise or environmental vibration interference, failure to filter out this noise will result in spurious fluctuations in the calculated rotation angle vector and / or position translation vector. Understandably, the image acquisition device may not actually be jittering, but the calculated rotation angle vector and / or position translation vector will indicate that the device is jittering. Therefore, to solve this technical problem, preserve the true motion trend of the image acquisition device, and thus make the subsequent rotation matrix and translation compensation vector more accurate and stable, thereby improving the anti-shake effect and ensuring that the image does not flicker, the second calculation scheme described above can be used to calculate the rotation angle vector and position translation vector.
[0031] Based on this, please refer to Figure 3 , Figure 3This is a flowchart illustrating a calculation scheme for a rotation angle vector and a position translation vector according to an exemplary embodiment of this specification. In step S100, the rotation angle vector and position translation vector of the image acquisition device relative to a set reference coordinate system are calculated based on the motion data of the image acquisition device when acquiring the image to be processed. This includes: In step S121, the angular velocity sequence and acceleration sequence collected by the inertial measurement unit within an image sampling interval starting from the sampling time of the image to be processed are obtained; In step S122, for each angular velocity in the angular velocity sequence, the original rotation angle vector corresponding to the angular velocity is calculated based on the angular velocity and the sampling time interval of the inertial measurement unit; In step S123, all original rotation angle vectors are subjected to low-pass filtering to obtain rotation angle vectors; In step S124, the original position translation vector corresponding to each acceleration in the acceleration sequence is calculated; In step S125, all original position translation vectors are subjected to low-pass filtering to obtain position translation vectors.
[0032] The technical principle of the second calculation scheme described in steps S121 to S125 above is explained below: Since the acquisition frequency of the IMU is higher than that of the image acquisition device, the IMU has multiple data sampling points between time t and time t+1, and an angular velocity and an acceleration are acquired at each data sampling point. The angular velocity and acceleration are stored in the order of acquisition time, so as to obtain the angular velocity sequence and acceleration sequence recorded in step S121.
[0033] Based on this, by executing step S121, the angular velocity sequence and acceleration sequence acquired by the IMU within an image sampling interval starting from the above sampling time can be obtained.
[0034] Subsequently, steps S122 and S124 can be executed in parallel or sequentially. In the case of sequential execution, the order of the two steps is not limited.
[0035] In step S122, the calculation principle for the original rotation angle vector corresponding to each angular velocity is the same as that for step S111, and will not be repeated here. It is worth noting that the difference lies in the time interval used in step S122 for calculating the original rotation angle vector; this time interval is the IMU sampling time interval, not the image sampling time.
[0036] After obtaining the original rotation angle vectors corresponding to all angular velocities in step S122, step S123 is executed to perform low-pass filtering on all the original rotation angle vectors, thereby suppressing high-frequency noise in the IMU data and obtaining a smoothed rotation angle vector, which is beneficial to improving the accuracy and stability of subsequent anti-shake.
[0037] In step S124, the calculation principle for the original position translation vector corresponding to each acceleration is the same as that for step S112, and will not be repeated here. It is worth noting that the difference lies in the acceleration at two different moments used in step S124 for calculating the original position translation vector; these are each pair of adjacent accelerations in the acceleration sequence. For the last acceleration in the acceleration sequence, its corresponding original position translation vector can be calculated using the acceleration at the next moment.
[0038] After obtaining the original position translation vectors corresponding to all accelerations through step S124, step S124 is executed to perform low-pass filtering on all the original position translation vectors, thereby suppressing high-frequency noise in the IMU data and obtaining a smoothed position translation vector, which is beneficial to improving the accuracy and stability of subsequent anti-shake.
[0039] Therefore, by performing low-pass filtering on all original rotation angle vectors and all original position translation vectors, the final rotation angle vectors and position translation vectors can reflect the true motion trend of the image acquisition device. This makes the rotation matrix and translation compensation vectors obtained subsequently based on the rotation angle vectors and position translation vectors more accurate and smoother, thereby achieving a more precise and stable anti-shake effect and improving the reliability of the system in complex environments.
[0040] After obtaining the rotation angle vector through any of the above embodiments, step S300 is executed to calculate the rotation matrix based on the X-axis angle in the rotation angle vector. The rotation matrix can be calculated using the formula... The calculation shows that, for the image to be processed acquired at time t, the X-axis angle at time t can be calculated. Substituting these values into the formula for calculating the rotation matrix above, we can obtain the rotation matrix at time t. .
[0041] The rotation matrix calculation formula only considers the X-axis angle because when the image acquisition device rotates along the X-axis of its own camera coordinate system, it just constitutes the rotation of the camera image (i.e., the image to be processed).
[0042] The execution order of steps S200 and S100 is not sequential. During the execution of step S200, the focus object can first be identified from the image to be processed using an intelligent object recognition algorithm, and the position of the focus object in the image to be processed can be determined based on the recognition result. Then, automatic focusing can be performed based on the position of the focus object. When focusing is successful, the object distance calculation principle in related technologies can be used to calculate the object distance of the focus object based on the focal length of the image acquisition device and the image distance of the focus object in the image to be processed.
[0043] After obtaining the object distance of the focused object in step S200 and the position translation vector in any of the above embodiments, step S400 is executed to calculate the translation compensation vector based on the field of view of the image acquisition device, the object distance, the Y-axis angle and Z-axis angle in the rotation angle vector, and the Y-axis translation amount and Z-axis translation amount in the position translation vector.
[0044] The translation compensation vector was not calculated using the X-axis translation amount because the position change of the image acquisition device along the X-axis of the camera coordinate system only affects the scaling of the image to be processed, and does not affect the original coordinates of the pixels in the image to be processed. Therefore, there is no need to consider the X-axis translation amount.
[0045] For step S400, this embodiment of the specification provides a method for calculating the translation compensation vector, namely: In step S400, based on the field of view, object distance, Y-axis and Z-axis angles in the rotation angle vector of the image acquisition device, and Y-axis and Z-axis translation amounts in the position translation vector, a translation compensation vector is calculated, including: In step S410, the ratio of the Z-axis translation amount to the object distance and the first sum of the Z-axis angle are calculated, and the ratio of the first sum to the tangent of the field of view angle is calculated to obtain the X-axis translation compensation amount. In step S420, the ratio of the Y-axis translation amount to the object distance and the second sum of the Y-axis angle are calculated, and the ratio of the second sum to the tangent of the field of view angle is calculated to obtain the Y-axis translation compensation amount; In step S430, the X-axis translation compensation amount and the Y-axis translation compensation amount are constructed into a 2×1 matrix to obtain the translation compensation vector.
[0046] In practical applications, the technical solutions of steps S410 to S430 above can be integrated into the following calculation formula: In the above formula, This represents the translation compensation vector at time t. This represents the Z-axis angle at time t. This represents the Y-axis angle at time t. This represents the Z-axis translation at time t. This represents the Y-axis translation at time t. This represents the object distance at time t. Indicates the field of view of the image acquisition device. This represents the translational compensation amount along the X-axis. This represents the Y-axis translation compensation amount. Therefore, by integrating the technical solutions of steps S410 to S430 into the above calculation formula, in actual calculations, the translation compensation vector at the corresponding moment can be calculated by substituting the field of view of the image acquisition device, the corresponding Z-axis angle, Y-axis angle, Z-axis translation amount, Y-axis translation amount, and object distance into the above formula, thereby improving computational efficiency.
[0047] After obtaining the rotation matrix and translation compensation vector, step S500 can be executed. This involves first performing a rotation transformation on the original coordinates of each pixel in the image to be processed using the rotation matrix, and then performing a translation transformation using the translation compensation vector, thereby obtaining the target coordinates of each pixel after jitter removal. For different considerations, this specification provides two jitter removal schemes, as follows: The first method for removing jitter: This jitter removal scheme primarily focuses on computational efficiency, aiming to improve processing efficiency while maintaining a certain level of computational accuracy. Based on this, in step S500, after rotating the original coordinates of the pixel using a rotation matrix, a translation compensation vector is used to perform a translation transformation on the rotated coordinates to obtain the target coordinates of the pixel after jitter removal. This includes: In step S510, the original coordinates of the pixel are multiplied by the rotation matrix, and then subtracted from the translation compensation vector to obtain the target coordinates of the pixel.
[0048] Understandably, the original coordinates of each pixel in the image to be processed are processed in step S510 to obtain the target coordinates after jitter removal. The processing procedure of step S510 can be expressed by the following formula: In the above formula, This represents the target coordinates of the i-th pixel in the image to be processed at acquisition time t; The rotation matrix at time t is described above, and its origin will not be repeated here. This represents the translation compensation vector at time t. Its origin can also be found in the relevant records above, and will not be repeated here. This represents the original coordinates of the i-th pixel in the image to be processed at acquisition time t.
[0049] The second method for removing jitter: To achieve better image stabilization, a second shake removal method can be used to process the coordinates of each pixel. Based on this, please refer to [link / reference needed]. Figure 4 , Figure 4 This is a flowchart illustrating a jitter removal scheme according to an exemplary embodiment of this specification. In step S500, after rotating the original coordinates of the pixel using a rotation matrix, the coordinates after rotation are translated using a translation compensation vector to obtain the target coordinates of the pixel after jitter removal, including: In step S521, the original translation compensation vector corresponding to each data sampling point is calculated based on the original rotation angle vector and the original position translation vector corresponding to each data sampling point. In step S522, the angular deviation between the X-axis rotation angle in each original rotation angle vector and the X-axis rotation angle in the rotation angle vector is calculated, and the single-step rotation matrix corresponding to each angular deviation is calculated. In step S523, the displacement deviation between each original translation compensation vector and the translation compensation vector is calculated to obtain the single-step translation compensation vector; In step S524, the original coordinates are subjected to rotation transformation and translation compensation transformation by using all single-step rotation matrices and all single-step translation compensation vectors to obtain the target coordinates.
[0050] Understandably, the purpose of steps S521 to S524 is to integrate the rotation and translation deviations of the image acquisition device at multiple data sampling points of the IMU, and to perform nested rotation and translation compensation transformations on the original coordinates layer by layer through all single-step rotation matrices and all single-step translation compensation vectors. This allows for the accumulation of small rotation increments (i.e., angle deviations in step S522) and small translation increments (i.e., displacement deviations in step S523) of the image acquisition device within an image sampling interval. This makes the accumulation of angle deviations closer to the actual rotational movement of the image acquisition device, and the accumulation of displacement deviations closer to the actual translational movement of the camera. This significantly reduces the probability of image blur or jitter, and can more comprehensively solve various types of jitter. It avoids residual jitter in the image due to using only single-point compensation, thereby better improving the image stabilization effect and enhancing image quality.
[0051] The technical principles of steps S521 to S524 are explained below: The order of execution of steps S521 and S522 is not limited. During the execution of step S521, the original rotation angle vector corresponding to each data sampling point is the original rotation angle vector obtained after processing in step S122 above. Therefore, the calculation principle of the original rotation angle vector can be found in the relevant description above, and will not be repeated here. The original position translation vector is the original position translation vector obtained after processing in step S124 above. The calculation principle of the original translation compensation vector is the same as the calculation principle of the translation compensation vector described in step S400 above, and will not be repeated here. It can be seen that after processing in step S521, multiple original translation compensation vectors with the same number as the number of data sampling points can be obtained, that is, multiple original translation compensation vectors correspond one-to-one with multiple sampling points.
[0052] During step S522, the X-axis rotation angle in each original rotation angle vector is subtracted from the X-axis rotation angle in the rotation angle vector obtained by the above low-pass filtering process to obtain multiple angle deviations corresponding to multiple data sampling points. Then, for each angle deviation, it is substituted into the above rotation matrix calculation formula to obtain multiple single-step rotation matrices corresponding to multiple data sampling points.
[0053] For example, assuming the image to be processed is acquired at time t, and one image sampling interval includes 3 IMU data sampling points, the IMU sampling time interval is... Then, the times of these three data sampling points, in chronological order, can be represented as follows: , , Based on this, within one image sampling interval, three angular deviations and three single-step rotation matrices can be obtained. Let the three angular deviations be: , , Therefore, the three single-step rotation matrices are as follows: , , .
[0054] After obtaining multiple original translation compensation vectors corresponding to multiple data sampling points by executing step S521, step S523 is executed to subtract each original translation compensation vector from the translation compensation vector obtained by the above low-pass filtering process to obtain multiple displacement deviations corresponding to multiple data sampling points. Each displacement deviation is a single-step translation compensation vector.
[0055] Continuing with the previous example, it can be seen that within one image sampling interval, three single-step translation compensation vectors can also be obtained, which are represented as follows: , , .
[0056] After obtaining all single-step rotation matrices and all single-step translation compensation vectors, step S524 is executed to perform rotation and translation compensation transformations on the original coordinates of each pixel using all single-step rotation matrices and all single-step translation compensation vectors, thereby obtaining the target coordinates of each pixel. An implementation example is provided in this specification as follows: In step S524 above, the original coordinates are subjected to rotation transformation and translation compensation transformation using all single-step rotation matrices and all single-step translation compensation vectors to obtain the target coordinates, including: In step S5241, all single-step rotation matrices and all single-step translation compensation vectors are grouped according to time to obtain multiple compensation sets; each compensation set includes a single-step rotation matrix and a single-step translation compensation vector corresponding to the same sampling time. In step S5242, multiple compensation sets are traversed sequentially according to time, and the currently traversed compensation set is taken as the target compensation set. In step S5243, the original coordinates are subjected to rotation and translation compensation transformations through the target compensation set to obtain the coordinates to be iterated. In step S5244, the coordinates to be iterated are used as the new original coordinates, and the above step S5242 is returned to perform iterative compensation on the coordinates to be iterated until the currently obtained coordinates to be iterated are processed by the last compensation set. In step S5245, the final coordinates to be iterated are used as the target coordinates.
[0057] Understandably, steps S5241 to S5245 are a process of stabilizing an original coordinate system. In other words, the original coordinates of each pixel in an image to be processed will be processed by steps S5241 to S5245 to obtain the target coordinates after jitter removal.
[0058] To facilitate understanding and calculation, and to simplify the iterative de-jittering process of the original coordinates, the processing logic described in steps S5241 to S5245 can be integrated into the following formula: The meanings of the symbols in the above formulas can be found in the relevant records above, and will not be repeated here. Among them, This represents the first coordinate to be iterated after the first anti-shake process. This represents the second iterative coordinate obtained after the second anti-shake processing. It can be seen that this second iterative coordinate is obtained by anti-shake processing the first coordinate to be iterated. This process continues until all compensation sets are used to complete the iterative anti-shake processing. The final coordinate to be iterated is the target coordinate after jitter removal.
[0059] Continuing with the example above, where an image sampling interval includes 3 IMU data sampling points, the above formula will be adaptively adjusted to: Similarly, the meanings of the symbols in the above formulas are described in the relevant sections above and will not be repeated here. Therefore, by substituting the corresponding data into the above formulas, the iterative de-jittering process described in steps S5241 to S5245 can be completed, resulting in target coordinates after the original coordinates have been iteratively de-jittered.
[0060] Therefore, by using the shake removal scheme described in steps S5241 to S5245, the high-precision deviations (angle deviations and displacement deviations) accumulated by rotation and translation compensation are integrated and smoothed through affine transformation, which can improve the comprehensiveness, accuracy, real-time performance and robustness of image stabilization, while also ensuring image quality and enhancing the handheld shooting experience.
[0061] After obtaining the target coordinates of all pixels in the image to be processed using any of the above-mentioned jitter removal methods, step S600 is executed to update the image to be processed according to the target coordinates of each pixel, thereby obtaining the jitter-removed and stabilized image. The image update principle can be found in related technologies and will not be detailed here.
[0062] The various technical features in the above embodiments can be combined arbitrarily, as long as there is no conflict or contradiction between the combinations of features. However, due to space limitations, they are not described one by one. Therefore, the arbitrary combination of various technical features in the above embodiments is also within the scope of this specification.
[0063] Corresponding to the foregoing method embodiments, this specification also provides electronic device embodiments. For example... Figure 5 As shown, Figure 5 This is a schematic structural diagram of an electronic device provided in this specification according to an exemplary embodiment. At the hardware level, the electronic device includes a processor 202, an internal bus 204, a network interface 206, memory 208, and non-volatile memory 210, and may also include other hardware required for business operations. One or more embodiments of this specification can be implemented in software, for example, the processor 202 reads the corresponding computer program from the non-volatile memory 210 into memory 208 and then runs it. Of course, besides software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of hardware and software, etc. That is to say, the execution entity of the following processing flow is not limited to individual logic units, but can also be hardware or logic devices.
[0064] Corresponding to the foregoing method embodiments, this specification also provides embodiments of an image acquisition device. The image acquisition device includes: Equipment body; The camera module, configured on the device itself, is used to acquire images to be processed; An inertial measurement unit, configured on the device body, is used to detect the motion data of the device body; The processor is used to process motion data transmitted based on the inertial measurement unit and to process the image to be processed using the image stabilization method described in any of the above embodiments to obtain a stabilized image.
[0065] In addition to the devices mentioned above, the image acquisition device provided in the embodiments of this specification may also include other modules of image acquisition devices in related technologies, such as an AE (auto exposure) module, an AF (auto focus) module, and an AI (Artificial Intelligence) module. Based on this, in some embodiments, to reduce the processor's workload, the motion analysis algorithm portion of the image stabilization method executed by the processor can be assigned to the inertial measurement unit; the object distance calculation portion can be assigned to the AF module; and the object recognition and position determination portion can be assigned to the AI module.
[0066] In some embodiments, to obtain higher quality images, the AE module can adjust the CMOS exposure time according to the motion of the image acquisition device to suppress motion blur and thus obtain better image or video effects. Specifically, when the image acquisition device is moving rapidly, the AE module can reduce the exposure time, and vice versa.
[0067] Corresponding to the foregoing method embodiments, this specification also provides a computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method as described in any of the foregoing embodiments.
[0068] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this manual are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of related data must comply with relevant laws, regulations and standards, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0069] The systems, devices, modules, or units described in the above embodiments can be implemented by a computer processor or entity, or by a product with a certain function. A typical implementation device is a computer, which can be a personal computer, laptop computer, cellular phone, camera phone, smartphone, personal digital assistant, media player, navigation device, email sending and receiving device, game console, tablet computer, wearable device, or any combination of these devices.
[0070] In a typical configuration, a computer may include one or more processors (CPU), input / output interfaces, network interfaces, and memory.
[0071] Memory may include non-persistent storage in computer-readable media, such as random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of computer-readable media.
[0072] Computer-readable media, including both permanent and non-permanent, removable and non-removable media, can store information using any method or technology. Information can be computer-readable instructions, data structures, modules of programs, or other data. Examples of computer storage media include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, CD-ROM, digital versatile optical disc (DVD) or other optical storage, magnetic tape, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transferable medium that can be used to store information accessible by a computing device. As defined herein, computer-readable media does not include transient computer-readable media, such as modulated data signals and carrier waves.
[0073] The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to be limiting of this specification. The singular forms “a,” “the,” and “the” as used in this specification and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term “and / or” as used herein refers to and includes any and all possible combinations of one or more of the associated listed items.
[0074] It should be understood that although the terms first, second, third, etc., may be used in this specification to describe various information, this information should not be limited to these terms. These terms are only used to distinguish information of the same type from one another. For example, without departing from the scope of this specification, first information may also be referred to as second information, and similarly, second information may also be referred to as first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "when".
[0075] It should also be noted that the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes the aforementioned element.
[0076] The foregoing has described specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims may be performed in a different order than that shown in the embodiments and may still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require the specific or sequential order shown to achieve the desired result. In some embodiments, multitasking and parallel processing are possible or may be advantageous.
[0077] The steps in the various embodiments described above are only for clarity. In practice, they can be combined into one step or some steps can be split into multiple steps. As long as they include the same logical relationship, they are all within the scope of protection of this patent. Adding insignificant modifications or introducing insignificant designs to the algorithm or process, but without changing the core design of the algorithm and process, are also within the scope of protection of this application.
[0078] While this specification contains numerous specific implementation details, these should not be construed as limiting the scope of any invention or the scope of the claims, but rather are primarily intended to describe features of specific embodiments of a particular invention. Certain features described in the various embodiments herein may also be implemented in combination in a single embodiment. Conversely, various features described in a single embodiment may also be implemented separately in various embodiments or in any suitable sub-combination. Furthermore, while features may function in certain combinations as described above and even initially claimed in this way, one or more features from a claimed combination may be removed from that combination in some cases, and a claimed combination may refer to a sub-combination or a variation thereof.
[0079] The terms "specific example" or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with embodiments or examples that are included in at least one embodiment or example of this specification. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.
[0080] Other embodiments of this specification will readily occur to those skilled in the art upon consideration of the specification and practice of the invention claimed herein. This specification is intended to cover any variations, uses, or adaptations that follow the general principles of this specification and include common knowledge or customary techniques in the art not claimed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this specification are indicated by the following claims.
[0081] It should be understood that this specification is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this specification is limited only by the appended claims.
[0082] The above description is only a preferred embodiment of this specification and is not intended to limit this specification. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.
Claims
1. An image stabilization method, characterized in that, include: Based on the motion data of the image acquisition device when acquiring the image to be processed, calculate the rotation angle vector and position translation vector of the image acquisition device relative to the set camera coordinate system; The object distance of the focused object is calculated based on the focal length parameters of the image acquisition device and the image distance of the focused object in the image to be processed. Calculate the rotation matrix based on the X-axis angle in the rotation angle vector; The rotation matrix is used to rotate and transform the original coordinates of the image pixels to counteract the rotational motion of the image acquisition device around the X-axis of the camera coordinate system; Based on the field of view of the image acquisition device, the object distance, the Y-axis and Z-axis angles in the rotation angle vector, and the Y-axis and Z-axis translation amounts in the position translation vector, a translation compensation vector is calculated; the translation compensation vector is used to offset the rotational motion of the image acquisition device around the Y-axis and Z-axis of the camera coordinate system, as well as the translational motion along the Y-axis and Z-axis; For each pixel in the image to be processed, the original coordinates of the pixel are rotated using the rotation matrix, and then the coordinates after rotation are translated using the translation compensation vector to obtain the target coordinates of the pixel after jitter removal. The image to be processed is updated based on the target coordinates of each pixel to obtain a stabilized image.
2. The method according to claim 1, characterized in that, The motion data includes angular velocity and acceleration; The step of calculating the rotation angle vector and position translation vector of the image acquisition device relative to the set camera coordinate system based on the motion data of the image acquisition device when acquiring the image to be processed includes: The rotation angle vector is calculated based on the angular velocity acquired by the inertial measurement unit configured in the image acquisition device at the acquisition time of the image to be processed, and an image sampling interval ending at the acquisition time; The position translation vector is calculated based on the acceleration collected by the inertial measurement unit at the acquisition time and at the next consecutive acquisition time.
3. The method according to claim 1, characterized in that, The motion data is acquired by an inertial measurement unit configured in the image acquisition device, and the motion data includes angular velocity and acceleration; an image sampling interval of the image acquisition device includes multiple data sampling points of the inertial measurement unit; The step of calculating the rotation angle vector and position translation vector of the image acquisition device relative to the set camera coordinate system based on the motion data of the image acquisition device when acquiring the image to be processed includes: The inertial measurement unit acquires the angular velocity sequence and acceleration sequence within an image sampling interval starting from the sampling time of the image to be processed; For each angular velocity in the angular velocity sequence, the original rotation angle vector corresponding to the angular velocity is calculated based on the angular velocity and the sampling time interval of the inertial measurement unit; All original rotation angle vectors are low-pass filtered to obtain the rotation angle vectors; Calculate the original position translation vector corresponding to each acceleration in the acceleration sequence; The original position translation vectors are low-pass filtered to obtain the position translation vectors.
4. The method according to any one of claims 1 to 3, characterized in that, The step of calculating the translation compensation vector based on the field of view of the image acquisition device, the object distance, the Y-axis and Z-axis angles in the rotation angle vector, and the Y-axis and Z-axis translation amounts in the position translation vector includes: Calculate the first sum of the ratio of the Z-axis translation to the object distance and the Z-axis angle, and calculate the ratio of the first sum to the tangent of the field of view angle to obtain the X-axis translation compensation. Calculate the ratio of the Y-axis translation to the object distance and the second sum of the Y-axis angle, and calculate the ratio of the second sum to the tangent of the field of view angle to obtain the Y-axis translation compensation amount; The X-axis translation compensation amount and the Y-axis translation compensation amount are constructed into a 2×1 matrix to obtain the translation compensation vector.
5. The method according to any one of claims 1 to 3, characterized in that, The process of rotating the original coordinates of the pixel using the rotation matrix, and then translating the rotated coordinates using the translation compensation vector to obtain the target coordinates of the pixel after jitter removal includes: The target coordinates of the pixel are obtained by performing matrix multiplication with the rotation matrix and then subtraction with the translation compensation vector.
6. The method according to claim 3, characterized in that, The process of rotating the original coordinates of the pixel using the rotation matrix, and then translating the rotated coordinates using the translation compensation vector to obtain the target coordinates of the pixel after jitter removal includes: Calculate the original translation compensation vector for each data sampling point based on the original rotation angle vector and the original position translation vector corresponding to each data sampling point; Calculate the angular deviation between the X-axis rotation angle in each original rotation angle vector and the X-axis rotation angle in the rotation angle vector, and calculate the single-step rotation matrix corresponding to each angular deviation; Calculate the displacement deviation between each original translation compensation vector and the translation compensation vector to obtain the single-step translation compensation vector; The original coordinates are subjected to rotation and translation compensation transformations using all single-step rotation matrices and all single-step translation compensation vectors to obtain the target coordinates.
7. The method according to claim 6, characterized in that, The process of performing rotation and translation compensation transformations on the original coordinates using all single-step rotation matrices and all single-step translation compensation vectors to obtain the target coordinates includes: All single-step rotation matrices and all single-step translation compensation vectors are grouped according to time to obtain multiple compensation sets; each compensation set includes a single-step rotation matrix and a single-step translation compensation vector corresponding to the same sampling time. The multiple compensation sets are traversed sequentially according to time, and the currently traversed compensation set is taken as the target compensation set. The original coordinates are subjected to rotation and translation compensation transformations using the target compensation set to obtain the coordinates to be iterated. The coordinates to be iterated are used as the new original coordinates, and the process of traversing the multiple compensation sets in chronological order and using the currently traversed compensation set as the target compensation set is returned to perform iterative compensation on the coordinates to be iterated until the currently obtained coordinates to be iterated are obtained by the last compensation set. The final coordinates to be iterated are used as the target coordinates.
8. An electronic device, characterized in that, include: processor; Memory used to store processor-executable instructions; The processor implements the method of any one of claims 1 to 7 by executing the executable instructions.
9. An image acquisition device, characterized in that, include: Equipment body; The camera module, configured on the device body, is used to acquire images to be processed; An inertial measurement unit, configured on the device body, is used to detect the motion data of the device body; A processor is configured to process the image to be processed based on the motion data transmitted by the inertial measurement unit, and to obtain a stabilized image by means of any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.