An IMU-based off-chip TDI low-light spatiotemporal enhancement method and system

Through IMU-assisted multiple exposure and image registration technology, the problem of weak image grayscale response in low-illumination environments is solved, and the image brightness resolution and signal-to-noise ratio are improved.

CN114119766BActive Publication Date: 2025-08-15ZHEJIANG DALI TECH
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202111228216.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-10-21
Publication Date
2025-08-15
Estimated Expiration
2041-10-21

AI Technical Summary

Technical Problem

In low-illumination environments, the image grayscale response is weak, resulting in low image contrast, uneven light and darkness, and difficult to distinguish the details of the shadowed area, which makes it difficult to follow-up processing in application fields such as reconnaissance and video surveillance.

Method used

The off-chip TDI low-illumination space-time enhancement method based on IMU is adopted to perform multiple exposures through the motion of the camera on the gimbal, and the multi-frame images are aligned with the IMU inertial navigation information, and the image registration technology is used to superimpose grayscale information to improve the image brightness resolution.

Benefits of technology

Through the alignment of inter-frame images and the superposition of grayscale values, the brightness resolution and signal-to-noise ratio of the image are improved, and the amount of brightness information of the image is enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114119766B_ABST
    Figure CN114119766B_ABST
Patent Text Reader

Abstract

An off-chip TDI low-light spatiotemporal enhancement method based on an IMU includes the following steps: controlling an electric pan-tilt platform to drive camera motion, using the IMU to obtain measurement information while the camera captures images; jointly calibrating the camera and IMU; obtaining the relative position relationship of the camera between the current frame and the reference frame using the real-time measurement information output by the IMU, back-calculating the image data of the current frame to the reference frame, and determining the coarse matching relationship between image pixels between frames; estimating and predicting the motion trajectories of the previous frame and the current frame, and then performing motion compensation; and fusing and outputting multiple frames of continuous images. The present invention fully utilizes the IMU inertial measurement unit to perform coarse camera motion estimation between frames, and performs precise registration through image registration technology, maximizing the use of motion and image information between frames. Through pixel fusion, the brightness resolution and signal-to-noise ratio of the image are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to an IMU-based off-chip TDI low-illumination spatiotemporal enhancement method and system, belonging to the technical field of image processing. Background Art

[0002] In certain usage scenarios, due to insufficient ambient light or uneven illumination, images often exhibit low contrast, uneven brightness, and difficulty distinguishing details in shadow areas. For example, in backlit or nighttime shooting situations, the image sensor can capture less light, resulting in dark, blurred images and difficulty observing details, which greatly complicates subsequent processing in applications such as reconnaissance and video surveillance. Low-light temporal and spatial enhancement technology (LITSET) reconstructs corresponding high-brightness resolution images from multiple frames of observed low-light images. It has important application value in monitoring equipment, satellite imagery, and medical imaging. Summary of the Invention

[0003] The technical problem to be solved by the present invention is: in order to solve the problem that the grayscale response of a single image is weak due to weak ambient light, an off-chip TDI (time delay integration) low-light spatiotemporal enhancement method and system based on IMU (Inertial measurement unit) is provided. The movement of the camera on the gimbal has a measurable sub-pixel level displacement. Within a certain time and displacement range, the data obtained by multiple exposures, and the overlapping parts of the imaging can be regarded as the superposition after the sensor detects and takes the image, thereby obtaining different effective information. After the image information is effectively integrated with the IMU inertial navigation information, the overlapping part of the image can be reconstructed to improve the brightness resolution. That is, multiple low-light images are aligned and the grayscale values of the corresponding pixels are superimposed to enhance the grayscale information and improve the brightness resolution of the image.

[0004] The purpose of the present invention is achieved through the following technical solutions:

[0005] An off-chip TDI low-light spatiotemporal enhancement system based on an IMU includes a system input terminal, a system control terminal, a system processing terminal, and a system output terminal;

[0006] The system input terminal is used to control the camera to perform dynamic measurement with constant rotation to obtain image information and the inter-frame displacement of the camera;

[0007] The system control terminal is used to control the rotation of the system input terminal and the image acquisition of the camera;

[0008] On the system processing side, the continuous frame images captured by the camera are fused. The fusion is carried out in the following way: the current frame in the image is aligned with the reference frame, and then the inter-frame image alignment, inter-frame motion estimation and motion compensation are performed, and finally the continuous frame images are fused.

[0009] The system output end is used to output the fused image.

[0010] Preferably, the system input includes an electric gimbal, an IMU, a camera lens, and a camera sensor;

[0011] Motorized pan / tilt head for controlling the camera to rotate constantly for dynamic measurement;

[0012] IMU is used to measure the inter-frame displacement of the camera;

[0013] A camera lens is used to optically image the surface information of an object onto an image plane;

[0014] The camera sensor converts the light signal returned by the camera lens into an electrical signal.

[0015] Preferably, the system control end uses a synchronization signal to adjust the pace of IMU and camera lens measurement, that is, the IMU inertial navigation information of each frame of image is synchronously returned when the camera lens collects it.

[0016] Preferably, the specific process of fusion at the system processing end is:

[0017] Through the three-axis attitude angle and angular velocity provided by the IMU, the flight information of the aircraft, the flight altitude, etc., the blur and geometric distortion caused by translation and rotation in the current frame of the image relative to the reference frame are aligned and corrected to complete the coarse motion estimation of the inter-frame image; then the image registration algorithm is used to perform fine alignment to complete the alignment of the inter-frame images, and the processed image is used to perform inter-frame motion estimation and motion compensation based on the current image. The time relationship and motion relationship between the current frame and the previous frame are used to complete the estimation and compensation of the moving object, and the continuous frame motion images are fused to increase the information content of the image.

[0018] An off-chip TDI low-light spatiotemporal enhancement method based on IMU includes the following steps:

[0019] Control the electric gimbal to move the camera, and use the IMU to obtain measurement information while the camera captures images;

[0020] Joint calibration of camera and IMU;

[0021] The real-time measurement information output by the IMU is used to obtain the relative position relationship between the camera in the current frame and the reference frame. The image data of the current frame is back-calculated to the reference frame to determine the rough matching relationship between the image pixels between the frames.

[0022] Estimate and predict the motion trajectory of the previous frame and the current frame, and then perform motion compensation;

[0023] Fusion output of multiple frames of continuous images.

[0024] Preferably, the process of joint calibration of the camera and IMU is:

[0025] According to the relative position relationship between IMU, camera and frames, three-dimensional coordinate transformation is used to obtain a three-dimensional homogeneous transformation matrix;

[0026] By modeling the IMU-camera platform as a fixed rigid body, the joint calibration of the camera and IMU is completed by using the pose of the IMU center reference system in the IMU fixed reference system, the pose of the camera frame reference system in the IMU center reference system, the pose of the camera frame reference system in the calibration board reference system, and the pose of the calibration board reference system in the IMU fixed reference system.

[0027] Preferably, in the specific process of coarse matching between image pixels between frames:

[0028] Using the real-time measurement data of IMU, the relative position relationship between the camera of the current frame and the reference frame is obtained, and the position of the reference frame camera relative to the IMU is obtained. The pose of the reference frame IMU relative to the base coordinate system The current frame camera pose relative to the IMU The pose of the current frame IMU relative to the base coordinate system The image data of the current frame is then back-calculated to the reference frame to determine the rough matching relationship between the image pixels between frames. The transformation formula is as follows:

[0029]

[0030] Preferably, the absolute error sum criterion is used for motion estimation and registration, and the calculation formula of the absolute error sum criterion is as follows:

[0031]

[0032] Where D(i,j) is the SAD value of pixel (i,j), Template is the current frame template image, whose size is M×N, Search is the reference frame search image, Template(s,t) is the pixel grayscale value of the template image, and Search(i+s-1,j+t-1) is the pixel grayscale value of the search image.

[0033] Preferably, a diamond search method is used to search for the best registration area; and two search templates of different shapes and sizes are used: a large diamond search template contains 9 candidate positions; and a small diamond search template contains 5 candidate positions.

[0034] Preferably, the DS algorithm search process is as follows: in the initial stage, the large diamond search template is repeatedly used until the best matching block falls on the center of the large diamond; then the small diamond search template is used to achieve accurate positioning of the best matching block.

[0035] Compared with the prior art, the present invention has the following beneficial effects:

[0036] The present invention makes full use of the IMU inertial measurement unit to estimate the rough motion of the camera between frames, and performs precise registration through image registration technology, making maximum use of the motion and image information between frames, and improving the brightness resolution and signal-to-noise ratio of the image through pixel fusion. BRIEF DESCRIPTION OF THE DRAWINGS

[0037] Figure 1 Schematic diagram of the low-light spatiotemporal enhancement system.

[0038] Figure 2 A numbering diagram of the world coordinate system to the camera coordinate system.

[0039] Figure 3 Schematic diagram of the relationship between the IMU fixed reference system, the reference system of the IMU connected to the camera, the reference system of the camera frame, and the reference system of the calibration plate.

[0040] Figure 4 Schematic diagram of the transformation relationship between various reference frames. DETAILED DESCRIPTION

[0041] To make the objectives, technical solutions and advantages of the present invention more clear, the embodiments of the present invention will be described in further detail below with reference to the accompanying drawings.

[0042] Example 1:

[0043] Due to the low ambient light, the grayscale response of a single image is weak. Therefore, by aligning multiple low-light images and superimposing the grayscale values of corresponding pixels, the grayscale information can be enhanced and the brightness resolution of the image can be improved. This system solution is divided into four parts: system input, system control, system processing, and system output. Figure 1 As shown, the transformation from the world coordinate system to the camera coordinate system is as follows Figure 2 shown.

[0044] The system inputs include: a motorized gimbal that controls the camera's constant rotation for dynamic measurement; an IMU (Instrumental Measurement Unit) that measures an object's three-axis attitude angle (or angular rate) and acceleration, and uses the information returned by the IMU to determine the camera's inter-frame displacement; a camera lens that optically captures the object's surface information onto the image plane; a camera sensor that converts the optical signal returned by the lens into an electrical signal; and power and data transmission equipment.

[0045] System control end: controls the pan-tilt rotation and camera image acquisition, and transmits digital video signals in real time; uses synchronization signals to align the IMU and detector measurements, that is, the IMU inertial navigation information of each frame of image is synchronously returned when the camera is acquiring the image.

[0046] System processing: Using the three-axis attitude angles and angular velocities along the axes provided by the IMU, as well as the aircraft's flight information, including linear velocity and altitude, the system aligns and corrects the blur and geometric distortion caused by translation and rotation in the current frame relative to the reference frame, completing coarse motion estimation between frames. Due to the accuracy limitations of the device's measurements, it is impossible to accurately align each frame, so an image registration algorithm is used to precisely align the images between frames. Inter-frame motion estimation and motion compensation are performed on the processed image using the current image as the standard. The temporal and motion relationships between the current and previous frames are used to estimate and compensate for moving objects, and continuous frame motion images are fused to increase the amount of information in the image.

[0047] At the system output end, the output multi-frame low-light enhanced video improves the signal-to-noise ratio of the output image.

[0048] Example 2:

[0049] The method of the present invention comprises the steps of:

[0050] 1. Control the electric gimbal to move the camera, and obtain IMU measurement information while the camera captures images.

[0051] 2. Perform joint calibration of the camera and IMU to determine the intrinsic parameters of the camera and the extrinsic parameters of the camera relative to the IMU.

[0052] 3. The real-time angle data output by the IMU is used to obtain the relative position relationship of the camera between the current frame and the reference frame, so that the image data of the current frame is back-calculated to the reference frame to determine the coarse matching relationship between the image pixels between frames.

[0053] 4. Estimate and predict the possible motion trajectories of the previous frame and the current frame, and determine the moving area in the image. After the matching relationship between the feature points of the moving object is obtained through inter-frame registration, the affine transformation matrix can be calculated to achieve motion compensation.

[0054] 5. After processing and fusing multiple frames of continuous images, the obtained low-light enhanced image is output according to the timing of the output end.

[0055] The specific process of step 2 is as follows:

[0056] According to the relative position relationship between IMU, camera and frame, three-dimensional coordinate transformation is required. Similar to two-dimensional coordinate transformation, the target position can be obtained by translating and rotating the coordinate system. The translation transformation matrix is obtained from the displacement distance in the three-axis direction, and the formula is as follows, where (T x 、T y 、T z ) is the displacement in the three-axis direction, P(P x 、P y 、P z ) is the initial coordinate matrix, T is the translation matrix, and P1 is the transformed coordinate matrix:

[0057]

[0058] The rotation transformation can be achieved by fixing one axis and rotating the plane formed by the other two axes around the fixed axis. The rotation transformation can be combined to obtain the rotation transformation. Taking the z axis as an example, the formula is as follows, where P1 is the coordinate matrix after transformation, R Z is the transformation matrix around the z-axis, P is the original coordinate matrix, and θ is the z-axis rotation angle:

[0059]

[0060] Similarly, by rotating φ and ω around the x-axis and y-axis, we can get the transformation matrix R around the x-axis x , the transformation matrix R around the y axis y :

[0061]

[0062]

[0063] The transformation is performed in the order of rotation around x, y, and z. The rotation matrix R is:

[0064]

[0065]

[0066] Finally, the three-dimensional homogeneous transformation matrix is obtained:

[0067]

[0068] By modeling the IMU-camera platform as a fixed rigid body, at any frame, four independent reference frames are considered, such as Figure 3As shown:

[0069] 1) The IMU fixed reference frame {B} with its origin at the IMU base;

[0070] 2) The (instantaneous) reference frame {E} of the IMU-connected camera;

[0071] 3) The reference frame {C} of the camera frame (instantaneous), with its origin at the optical center of the camera and its z-axis aligned with the optical axis of the lens;

[0072] 4) The calibration plate reference system {K} is defined on the calibration plate plane.

[0073] The transformation relationship between each reference system is determined as follows: Figure 4 As shown:

[0074] A: The position of the IMU center (instantaneous) reference frame in the IMU fixed reference frame can be obtained through the IMU's real-time parameters.

[0075] B: The pose of the camera frame (instantaneous) reference frame in the IMU center (instantaneous) reference frame. This transformation is fixed. Once the transformation is determined, the actual position of the camera can be calculated at any time through the real-time data and time integration of the IMU, and the transformation relationship between the current frame and the reference frame reference frame can be calculated.

[0076] C: The pose of the camera frame (instantaneous) reference system in the calibration plate reference system, that is, solving the extrinsic parameters (instantaneous) of the camera relative to the calibration plate.

[0077] D: The pose of the calibration plate reference frame in the IMU fixed reference frame. During the actual off-chip TDI process, the calibration plate does not exist. This transformation is only used to assist in solving the pose of the camera frame (instantaneous) reference frame in the IMU center (instantaneous) reference frame.

[0078] By moving the IMU to two positions and ensuring that the camera can see the calibration plate at both positions, we can obtain the initial pose A1 of the IMU center (instantaneous) reference system in the IMU fixed reference system, the pose C1 of the camera frame (instantaneous) reference system in the calibration plate reference system, the current pose A2 of the IMU center (instantaneous) reference system in the IMU fixed reference system, and the pose C2 of the camera frame (instantaneous) reference system in the calibration plate reference system. Then, we can construct a spatial transformation loop as follows: Figure 2 As shown, then:

[0079]

[0080] This is converted into the typical problem of solving AX=XB. By definition, X is a 4×4 homogeneous transformation matrix:

[0081]

[0082] Let the transformation relationship between the current frame and the reference frame be T NtoO ,but:

[0083]

[0084] Where T NtoO Also a 4×4 homogeneous transformation matrix: R and T are the rotation matrix and translation matrix respectively.

[0085] In this way, the actual position of the camera in the current frame can be calculated through the real-time data and time integration of the IMU, and the transformation relationship T between the current frame and the reference frame can be calculated. NtoO , and then use T NtoO The image data of the current frame is back-calculated to the reference frame using the intrinsic parameters obtained by camera calibration to determine the rough matching relationship between image pixels between frames. The calculation formula is as follows:

[0086]

[0087] Where (u1 v1) is the coordinate of the pixel point in the reference frame, (u2 v2) is the coordinate of the pixel point in the current frame, is the camera intrinsic parameter matrix, dx and dy are the pixel sizes in the x and y directions, and (uv) is the principal point coordinate of the camera.

[0088] The specific process of step 3 is as follows:

[0089] Using the real-time measurement data of IMU, the relative position relationship between the camera of the current frame and the reference frame is obtained, and the position of the reference frame camera relative to the IMU is obtained. The pose of the reference frame IMU relative to the base coordinate system The current frame camera pose relative to the IMU The pose of the current frame IMU relative to the base coordinate system The image data of the current frame is then back-calculated to the reference frame to determine the rough matching relationship between the image pixels between frames. The transformation formula is as follows:

[0090]

[0091] The absolute error and SAD criteria are used for motion estimation and precise registration. They are most commonly used because they do not require multiplication operations and are simple and convenient to implement. The SAD calculation formula is as follows:

[0092]

[0093] Where D(i,j) is the SAD value of pixel (i,j), Template is the current frame template image, whose size is M×N, Search is the reference frame search image, Template(s,t) is the pixel grayscale value of the template image, and Search(i+s-1,j+t-1) is the pixel grayscale value of the search image.

[0094] The Diamond Search (DS) method is used to search for the optimal registration region. It is simple, robust, and efficient, making it one of the best-performing fast search algorithms currently available. The basic idea behind Diamond Search is to exploit the fact that the shape and size of the search template significantly influence the speed and accuracy of motion estimation algorithms. When searching for the optimal matching point, choosing a small search template can lead to a local optimum, while choosing a large search template may prevent the optimal point from being found. Therefore, the DS algorithm utilizes two search templates of different shapes and sizes, based on the basic patterns of motion vectors in video images. The Large Diamond Search Pattern (LDSP) contains nine candidate positions, while the Small Diamond Search Pattern (SDSP) contains five candidate positions.

[0095] The DS algorithm's search process is as follows: Initially, a large diamond search template is repeatedly used until the optimal match falls within the center of the large diamond. Due to the large step size of the LDSP, the search range is wide, enabling coarse positioning and preventing the search from becoming trapped in a local minimum. Once coarse positioning is complete, the optimal point is assumed to lie within the diamond-shaped area enclosed by the eight points surrounding the LDSP. Smaller diamond search templates are then used to accurately locate the optimal match, minimizing fluctuations and improving motion estimation accuracy.

[0096] After obtaining the matching relationship between the feature points of the moving object through inter-frame registration, the affine transformation matrix can be calculated to achieve motion compensation.

[0097] The affine transformation of the image is as follows:

[0098]

[0099] Among them, (u, v) is the pixel coordinate before affine transformation, (x, y) is the pixel coordinate after affine transformation, is the affine transformation matrix.

[0100] The motion compensation radial transformation model obtained based on the affine transformation model of the image is as follows:

[0101]

[0102] Among them, (x t+1 ,yt+1 ) is the coordinate of the feature pixel point of the current frame, (x t ,y t ) is the coordinate of the feature pixel point of the reference frame, is the affine transformation matrix between the feature points of the current frame and the reference frame.

[0103] The N feature points obtained by registration can be used to obtain N affine transformation models. The energy minimization method is used to find the six parameters, and the following equations are used:

[0104]

[0105] Where x1, x2, ... x n and y1, y2, ...y n are the pixel coordinates of N feature points in the current frame, X1, X2, ... X n and Y1, Y2, ...Y n are the pixel coordinates of N feature points in the reference frame, E x 、E y is the sum of the horizontal and vertical variances.

[0106] To make E x 、E y Minimum, first find its partial derivative, then set the derivative to zero, and finally get the 6 parameters a-f as follows:

[0107]

[0108]

[0109]

[0110] The contents not described in detail in the specification of the present invention belong to the common knowledge of those skilled in the art.

[0111] Although the present invention has been disclosed above in terms of preferred embodiments, it is not intended to limit the present invention. Any person skilled in the art may make possible changes and modifications to the technical solutions of the present invention by using the methods and technical contents disclosed above without departing from the spirit and scope of the present invention. Therefore, any simple modifications, equivalent changes and modifications made to the above embodiments based on the technical essence of the present invention without departing from the content of the technical solutions of the present invention shall fall within the scope of protection of the technical solutions of the present invention.

Claims

1. An off-chip TDI low-light spatiotemporal enhancement system based on IMU, characterized by: Including system input end, system control end, system processing end, and system output end; The system input terminal is used to control the camera to perform dynamic measurement with constant rotation to obtain image information and the inter-frame displacement of the camera; The system control terminal is used to control the rotation of the system input terminal and the image acquisition of the camera; On the system processing side, the continuous frames captured by the camera are fused. This fusion is done by using the three-axis attitude angles and angular velocity provided by the IMU, the aircraft's flight information, and the flight altitude. Blurs and geometric distortions caused by translation and rotation in the current frame relative to the reference frame are corrected to achieve coarse motion estimation between frames. The process of joint calibration of camera and IMU is as follows: According to the relative position relationship between IMU, camera and frames, three-dimensional coordinate transformation is used to obtain a three-dimensional homogeneous transformation matrix; By modeling the IMU-camera platform as a fixed rigid body, the joint calibration of the camera and IMU is completed using the pose of the IMU center reference system in the IMU fixed reference system, the pose of the camera frame reference system in the IMU center reference system, the pose of the camera frame reference system in the calibration board reference system, and the pose of the calibration board reference system in the IMU fixed reference system. Using the real-time measurement data of the IMU, the relative position relationship of the camera between the current frame and the reference frame is obtained, and the image data of the current frame is back-calculated to the reference frame to determine the coarse matching relationship between the image pixels between the frames; Then, an image registration algorithm is used to precisely register the images to complete the alignment between frames. Inter-frame motion estimation and motion compensation are performed on the processed images using the current image as the standard. The time and motion relationship between the current frame and the previous frame is used to complete the estimation and compensation of moving objects. Continuous frame motion images are then fused to increase the information content of the image. The system output end is used to output the fused image.

2. The enhancement system according to claim 1, characterized in that At the system control end, a synchronization signal is used to align the IMU and camera lens measurements, that is, the IMU inertial navigation information of each frame of image is synchronously returned when the camera lens collects the image.

3. An off-chip TDI low-light spatiotemporal enhancement method based on IMU, characterized in that: The steps include: Control the electric gimbal to move the camera, and use the IMU to obtain measurement information while the camera captures images; Joint calibration of the camera and IMU; the process of joint calibration of the camera and IMU is as follows: According to the relative position relationship between IMU, camera and frames, three-dimensional coordinate transformation is used to obtain a three-dimensional homogeneous transformation matrix; By modeling the IMU-camera platform as a fixed rigid body, the joint calibration of the camera and IMU is completed using the pose of the IMU center reference system in the IMU fixed reference system, the pose of the camera frame reference system in the IMU center reference system, the pose of the camera frame reference system in the calibration board reference system, and the pose of the calibration board reference system in the IMU fixed reference system. The real-time measurement information output by the IMU is used to obtain the relative position relationship between the camera in the current frame and the reference frame. The image data of the current frame is back-calculated to the reference frame to determine the rough matching relationship between the image pixels between the frames. Estimate and predict the motion trajectory of the previous frame and the current frame, and then perform motion compensation; Fusion output of multiple frames of continuous images.

4. The enhancement method according to claim 3, characterized in that The absolute error and criterion are used for motion estimation and registration. The calculation formula of the absolute error and criterion is as follows: Where D(i,j) is the SAD value of pixel (i,j), Template is the current frame template image, whose size is M×N, Search is the reference frame search image, Template(s,t) is the pixel grayscale value of the template image, and Search(i+s-1,j+t-1) is the pixel grayscale value of the search image.

5. The enhancement method according to claim 3, characterized in that: The diamond search method is used to search for the best registration area. Two search templates of different shapes and sizes are used: the large diamond search template contains 9 candidate positions; the small diamond search template contains 5 candidate positions.

6. The enhancement method according to claim 5, characterized in that: The DS algorithm search process is as follows: in the initial stage, the large diamond search template is repeatedly used until the best matching block falls in the center of the large diamond; then the small diamond search template is used to accurately locate the best matching block.

Citation Information

Patent Citations

  • Dynamic background video object extraction based on enhancement type diamond search and three-frame background alignment

    CN102917223A

  • Method for searching for motion estimation

    CN103763563A

  • Calibration method for relative attitude of binocular stereo camera and inertial measurement unit

    CN105606127A

  • Dual-eye three-dimensional visual measurement method and system fused with IMU calibration

    CN110296691A