Coaxial motion drift self-correction method, system, equipment and medium
By using a coaxial motion drift self-correction method, and leveraging visible light image sequences to drive precise inter-frame registration and spatiotemporal statistics of infrared images, the problems of large system size and high power consumption in non-uniformity correction in infrared imaging are solved, achieving zero-interruption, stable image correction and high-precision target detection.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-09
- Publication Date
- 2026-03-13
AI Technical Summary
In existing infrared imaging technologies, non-uniformity correction methods rely on mechanical shutters or blackbodies, resulting in large system sizes, high power consumption, and the inability to achieve continuous imaging with zero interruptions. Furthermore, direct fusion of infrared and visible light can easily produce pixel-level misalignment, affecting image quality and target recognition accuracy.
The scene motion prior is obtained by using a visible light image sequence mounted coaxially, which drives the fine registration between infrared image frames. Gain and bias drift are jointly estimated in the spatiotemporal domain to achieve non-uniformity correction without shutter or blackbody. Subsequently, it is fused with visible light at the pixel level for target detection.
It achieves zero-disruption non-uniformity correction of infrared images, eliminates banding and ghosting, and aligns infrared and visible light pixels at the pixel level, thereby improving the recognition confidence and image quality of target detection.
Smart Images

Figure CN121660936A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of PTZ camera imaging technology, specifically to a coaxial motion drift self-correction method, system, device, and medium. Background Technology
[0002] In the field of infrared imaging technology, non-uniformity correction (NUC) is a core component for ensuring image quality. Due to factors such as manufacturing processes, material inhomogeneity, and variations in the working environment, the response gain and bias of each pixel in an infrared focal plane array (IRFPA) drift to varying degrees over time and with temperature. This results in image distortion such as stripes, fixed pattern noise (FPN), and flicker, severely impacting observation, temperature measurement, and subsequent algorithm processing. Therefore, it is essential to periodically perform NUC processing on infrared images.
[0003] Currently, the mainstream NUC methods mainly fall into two categories:
[0004] Reference source-based correction methods, such as shutter correction or blackbody radiation source calibration, obtain real-time response parameters for each pixel by blocking the infrared light path or introducing a uniform blackbody radiation surface, thereby calculating correction coefficients. Although this method offers high correction accuracy, it has significant drawbacks: First, it requires periodic interruptions to imaging, resulting in the loss of keyframes, which cannot meet the requirements of applications with zero tolerance for frame loss, such as continuous monitoring and target tracking. Second, the introduction of a mechanical shutter or blackbody radiation source increases the system's size, weight, power consumption, and cost, making it difficult to integrate, especially on miniaturized platforms such as drones, handheld thermal imagers, and wearable devices with extremely high load and endurance requirements. Third, mechanical shutters have poor reliability in harsh environments such as high and low temperatures, humidity, and dust, and are prone to jamming, failure, and other malfunctions, reducing system lifespan.
[0005] Scene-based statistical correction methods, such as constant gain models, neural network methods, and Kalman filtering, utilize motion information from the scene in infrared image sequences to estimate the gain and bias drift of each pixel through inter-frame grayscale statistical models. Their advantage is that they do not require interrupting imaging, making them suitable for dynamic scenes. However, they heavily rely on sufficiently rich inter-frame image content and accurate registration. Infrared images themselves have low signal-to-noise ratios and weak textures, especially in areas with weak textures such as the sky, water surfaces, and snow. Traditional motion estimation algorithms based on infrared technology (such as block matching and optical flow methods) struggle to obtain reliable motion vectors, leading to large inter-frame registration errors. This, in turn, distorts the correction coefficients estimated by the statistical model, resulting in ghosting or overcorrection, ultimately reducing image quality. Furthermore, as infrared detectors evolve towards smaller pixels and higher frame rates, the drift speed of gain and bias increases significantly. Traditional scene-based statistical methods struggle to keep pace with these drift changes, making it impossible to maintain stable correction results over extended periods.
[0006] In recent years, although some studies have attempted to introduce visible light images to assist infrared correction, these efforts have mostly been limited to visible light guided filtering or image fusion enhancement, and a systematic framework of pre-registration followed by statistical closed-loop correction has not yet been formed. This has failed to fundamentally solve the NUC performance degradation problem caused by infrared inter-frame registration failure. Furthermore, existing infrared-visible light fusion systems typically perform fusion detection directly on the original infrared and visible light images. Due to inconsistencies in the imaging perspective, resolution, and distortion parameters of the two modalities, pixel-level alignment errors are significant, leading to edge ghosting and target splitting in the fused image, severely impacting the accuracy of subsequent target recognition and detection. Adding an additional registration module introduces computational latency and system complexity, making it difficult to meet real-time requirements.
[0007] In summary, existing technologies cannot achieve continuous imaging with zero interruptions while ensuring high image quality, nor can they simultaneously meet the multiple requirements of system miniaturization, low power consumption, and ease of use in subsequent fusion detection. Therefore, there is an urgent need for a novel NUC method that can overcome the bottleneck of poor reliability in motion estimation of infrared images without relying on mechanical shutters or blackbodies, achieve long-term, stable, and uninterrupted non-uniformity correction, and provide high-quality images with pixel-level alignment for subsequent infrared-visible light fusion detection. Summary of the Invention
[0008] The technical problem this invention aims to solve is that infrared correction relies on a shutter blackbody, which leads to registration errors and frame drops due to the weak texture of infrared frames. Furthermore, direct fusion of infrared and visible light can easily result in pixel-level misalignment. The goal is to provide a coaxial motion drift self-correction method, system, device, and medium. This method uses a high-texture visible light sequence to output scene motion priors, driving precise registration between infrared frames to achieve geometric uniformity. Then, through joint statistics in the spatiotemporal domain, gain and bias drift are estimated in real-time without a blackbody or shutter, completing zero-interruption non-uniformity correction and completely eliminating banding and ghosting. After correction, infrared and visible light are naturally aligned at the pixel level, allowing for fusion detection without additional registration. Target edges are clear and ghosting-free, significantly improving recognition confidence.
[0009] This invention is achieved through the following technical solution:
[0010] The first aspect of this invention provides a coaxial motion drift self-correction method, comprising the following specific steps:
[0011] Acquire a sequence of visible light images as a priori output for scene motion;
[0012] By utilizing scene motion priors, inter-frame registration is performed on the infrared image sequence to obtain a spatially aligned infrared image sequence.
[0013] Spatiotemporal statistics are performed on spatially aligned infrared image sequences to estimate the gain drift coefficient and bias drift coefficient.
[0014] The non-uniformity correction of the current infrared image is performed based on the estimated gain drift coefficient and bias drift coefficient, and the corrected infrared image is output.
[0015] The corrected infrared image and the visible light image are fused pixel by pixel for target detection.
[0016] Furthermore, the acquisition of the visible light image sequence as a priori output of scene motion specifically includes:
[0017] Acquire visible light image sequences with coaxial mounting and extract feature points from the visible light image sequences;
[0018] Feature points are extracted from two adjacent frames and feature matching is performed to generate an initial matching set;
[0019] Foreground filtering is performed on the initial matching set to remove mismatches corresponding to moving targets, resulting in a background matching set;
[0020] Based on the background matching set, a robust estimation algorithm is used to solve for the global motion model parameters, and the global motion vector of the camera is obtained as the scene motion prior output.
[0021] Furthermore, the extraction of visible light image sequence feature points specifically includes:
[0022] An integral image is generated pixel-by-pixel from the input visible light image sequence, enabling the grayscale sum of any rectangular region to be obtained in constant time.
[0023] The second-order differential image is obtained by approximating the second-order Gaussian derivative in the x, y, and xy directions using a box filter;
[0024] Calculate the Hessian determinant response value of each pixel. If the response value of a candidate point is greater than that of n neighbors at the same time, the output is the visible light image sequence feature point.
[0025] Furthermore, the step of using scene motion priors to perform inter-frame registration on the infrared image sequence to obtain a spatially aligned infrared image sequence specifically includes:
[0026] The camera's global motion vector is converted into inter-frame motion in the infrared coordinate system through visible light-infrared motion mapping;
[0027] Motion compensation is performed on the original infrared image sequence based on inter-frame motion in the infrared coordinate system to complete inter-frame registration and obtain a spatially aligned infrared image sequence.
[0028] Furthermore, the step of performing spatiotemporal statistics based on the spatially aligned infrared image sequence to estimate the gain drift coefficient and bias drift coefficient specifically includes:
[0029] For a spatially aligned infrared image sequence, a sliding window with a length of N frames is established along the time dimension at each pixel location.
[0030] Calculate the mean and variance of the pixel's grayscale value within the sliding window;
[0031] Using the preset reference standard deviation and reference mean as a benchmark, and combining the mean and variance of the pixel's grayscale value, the gain drift coefficient and bias drift coefficient are obtained respectively.
[0032] Furthermore, the step of performing non-uniformity correction on the current infrared image based on the estimated gain drift coefficient and bias drift coefficient, and outputting the corrected infrared image, specifically includes:
[0033] Based on the estimated gain drift coefficient and offset drift coefficient, a gain drift coefficient matrix and an offset drift coefficient matrix with the same resolution as the original infrared image frame are constructed, wherein the gain drift coefficient matrix and the offset drift coefficient matrix are updated with at least one of the following factors: ambient temperature, working time or scene statistical characteristics.
[0034] Each pixel of the original infrared image is corrected based on the gain drift coefficient matrix and the bias drift coefficient matrix with the same resolution as the original infrared image frame.
[0035] The corrected pixel value array is dynamically range mapped to obtain the corrected infrared image and output it.
[0036] Furthermore, the step of pixel-level fusion of the corrected infrared image and the visible light image for target detection specifically includes:
[0037] Acquire a visible light image at the same time as the corrected infrared image, which has already been pixel-level registered;
[0038] Detail feature maps are extracted from the corrected infrared image and visible light image respectively, and the fusion weight of each pixel is calculated based on the detail feature maps to obtain a weight map;
[0039] The corrected infrared image and the visible light image are fused pixel-level based on the weight map to generate a fused image;
[0040] The fused image is input into a pre-trained target detection network to obtain and output the target category and location information.
[0041] A second aspect of the present invention provides a system for implementing a coaxial motion drift self-correction method, comprising:
[0042] Infrared detector module for continuously outputting infrared image sequences without the need for mechanical baffles;
[0043] The visible light imaging module is used to synchronously acquire visible light image sequences of the same scene with the infrared detector module, and to output the scene motion prior.
[0044] A motion-scene analysis module, coupled to the visible light imaging module and the infrared detector module, is used to perform optical flow or feature point motion detection based on the visible light image sequence and / or the infrared image sequence to generate a scene motion confidence signal;
[0045] The inter-frame registration module, coupled to the motion-scene analysis module, is used to perform inter-frame registration on the infrared image sequence using the scene motion prior provided by the visible light image sequence when the scene motion confidence signal indicates that there is sufficient motion, so as to obtain a spatially aligned infrared image sequence.
[0046] The spatiotemporal statistical estimation module, coupled with the inter-frame registration module and the motion-scene analysis module, is used to perform statistical operations on the spatially aligned infrared image sequence in the spatiotemporal domain under the control of the scene motion confidence signal, so as to estimate the gain drift coefficient and bias drift coefficient that change slowly over time.
[0047] The non-uniformity correction module, coupled to the spatiotemporal statistical estimation module and the infrared detector module, is used to perform pixel-level non-uniformity correction on the current infrared image using the latest estimated gain drift coefficient matrix and the bias drift coefficient matrix, and output the corrected infrared image.
[0048] A pixel-level fusion module, coupled to the non-uniformity correction module and the visible light imaging module, is used to perform pixel-level fusion of the corrected infrared image and the corresponding visible light image to generate a fused image;
[0049] The target detection module, coupled to the pixel-level fusion module, is used to perform target detection on the fused image and output target category and location information.
[0050] A third aspect of the present invention provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement a coaxial motion drift self-correction method.
[0051] A fourth aspect of the present invention provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements a coaxial motion drift self-correction method.
[0052] Compared with the prior art, the present invention has the following advantages and beneficial effects:
[0053] By outputting scene motion priors using high-texture visible light sequences, precise inter-frame registration of infrared images is driven, achieving geometric uniformity. Then, joint statistics in the spatiotemporal domain are performed to estimate gain and bias drift in real time without blackbodies or shutter speeds, completing zero-interruption non-uniformity correction and completely eliminating banding and ghosting. After correction, infrared and visible light images are naturally aligned at the pixel level, allowing for fusion detection without additional registration. Target edges are clear and ghosting-free, significantly improving recognition confidence. Attached Figure Description
[0054] To more clearly illustrate the technical solutions of the exemplary embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as a limitation of the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort. In the drawings:
[0055] Figure 1 This is a flowchart of the image processing in an embodiment of the present invention. Detailed Implementation
[0056] To make the objectives, technical solutions, and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the embodiments and accompanying drawings. The illustrative embodiments and descriptions of the present invention are only used to explain the present invention and are not intended to limit the present invention.
[0057] As one possible implementation method, such as Figure 1 As shown, this embodiment provides a coaxial motion drift self-correction method, including the following specific steps:
[0058] S1. Obtain the visible light image sequence as the prior output of scene motion;
[0059] S2. Use scene motion priors to perform inter-frame registration on the infrared image sequence to obtain a spatially aligned infrared image sequence.
[0060] S3. Perform spatiotemporal statistics on the spatially aligned infrared image sequence to estimate the gain drift coefficient and offset drift coefficient;
[0061] S4. Based on the estimated gain drift coefficient and bias drift coefficient, perform non-uniformity correction on the current infrared image and output the corrected infrared image;
[0062] S5. The corrected infrared image and the visible light image are fused pixel by pixel to perform target detection.
[0063] In this embodiment, 1. Due to the local distortion often caused by platform jitter or non-rigid target motion between infrared frames, traditional registration models based on global affine / projection are prone to ghosting and cumulative errors. This embodiment first extracts scene motion priors using a visible light sequence with high signal-to-noise ratio and high spatial resolution, and then uses this as a constraint to perform inter-frame mapping on the infrared sequence, which can reduce the average pixel registration error and weaken motion artifacts. 2. Since infrared focal plane arrays exhibit slow gain and bias drift with operating temperature, noise, etc., traditional two-point calibration requires interruption of observation. This embodiment uses the scene statistical constancy assumption to perform joint estimation in the spatial-temporal domain on the spatially aligned infrared sequence, which can calculate the gain drift coefficient and bias drift coefficient of each pixel online without the need for a blackbody, and can effectively suppress noise of fixed patterns. 3. Due to the low resolution and weak edges of long-wave infrared, direct target detection has a high false alarm rate. This embodiment performs pixel-level fusion (weighted, wavelet, or sparse representation) between the infrared frame that has undergone non-uniformity correction and the synchronous visible light frame, which preserves the infrared thermal features and injects visible light texture features, thereby improving detection accuracy.
[0064] S1. Obtain the visible light image sequence as the prior output of scene motion.
[0065] S11. Since coaxial mounting can directly map the visible light motion vector to the infrared coordinate system, avoiding external parameter calibration errors, visible light image sequences of coaxial mounting are acquired.
[0066] Then, an integral image is generated pixel by pixel based on the input visible light image sequence. For each frame in the input visible light image sequence, its integral image is calculated pixel by pixel. The value of any point on the integral image is defined as the sum of the gray values of all pixels in the rectangular region formed by the top left corner (0,0) to (x,y) of the original image. Even if the gray sum of any rectangular region can be obtained in constant time, in order to detect blob features at different scales.
[0067] The second-order differential image is obtained by approximating the second-order Gaussian derivative in the x, y, and xy directions using a box filter. In this template, the weight of the white region is +1, the weight of the black region is -1, and the weight of the gray region is 0. For example, the Lxx filter is a structure with a white vertical bar in the middle and black vertical bars on both sides. First, an integral image is generated from the original image, so that the sum of any rectangular region only requires 4 additions and subtractions. Then, box-shaped weight templates of different sizes are used to approximate the second-order partial derivative of Gaussian image. Here, Lxx is positive in the middle column and negative in the left and right columns, Lyy is positive in the middle row and negative in the top and bottom rows, and Lxy is positive at the four corners and negative at the four "sides". The template is split into several rectangles and quickly accumulated using the integral image, that is, to find stable feature points that are extreme in both scale and space. For each position (x,y) and each scale s in scale space, these box-shaped filters are slid on the integral image, and the response value of the filter coverage area is calculated to obtain the second-order differential approximation values Lxx, Lyy, and Lxy at different scales, that is, to obtain three response images. After combining the three images into a Hessian matrix, speckle detection (det(H)=Dxx∙Dyy−0.81∙Dxy²) or any required high-order feature analysis can be performed. The larger the absolute value of the response value, the more obvious the "spot" feature of the point at that scale (it may be a bright point, a dark point, or a corner point).
[0068] In the three-dimensional scale space (x, y, scale s), (x, y) represents the location, and s is the optimal scale. The Hessian response value of each candidate point is compared with its eight spatial neighbors at the same scale, as well as its neighbors at adjacent (upper and lower) scales. A point is considered a local extremum only if its response value is simultaneously greater than (or simultaneously less than, depending on whether a maxima or minima is being sought) all its neighboring points. The (x, y, s) triples selected through this process are then output as the feature points of the visible light image sequence.
[0069] S12. Extract feature points from two adjacent frames and perform feature matching to generate an initial matching set.
[0070] S13. Perform foreground filtering on the initial matching set to remove mismatches corresponding to moving targets, and obtain the background matching set. Since foreground motion is an outlier in the background motion model, it needs to be removed before global estimation: use RANSAC homography matrix inner / outer point partitioning, or use motion vector direction consistency and amplitude threshold for fast filtering, or use visible light inter-frame optical flow to determine the foreground.
[0071] S14. Based on the background matching set, a robust estimation algorithm is used to solve for the global motion model parameters, obtaining the camera's global motion vector as the prior output of scene motion. From the visible light image sequence, the camera motion of each frame relative to the previous frame (or a certain reference frame) is estimated with high precision. Here, global motion mainly refers to the continuous motion of the entire background caused by the camera's own translation, rotation, or jitter, rather than the independent motion of individual objects in the scene. Feature points are extracted and matched on two consecutive visible light images to obtain a set of reliable matching point pairs. Based on the camera motion and the complexity of the application scenario, a suitable 2D motion transformation model is selected. Using the matched feature point pairs, the parameter matrix of the above motion model is solved using a robust estimation algorithm. The robust estimation algorithm can eliminate outlier point matching pairs that belong to moving objects in the scene, thereby ensuring that the estimated parameter matrix truly represents the global background motion caused by the camera.
[0072] S2. Use scene motion priors to perform inter-frame registration on the infrared image sequence to obtain a spatially aligned infrared image sequence.
[0073] S21. Convert the camera's global motion vector into inter-frame motion in the infrared coordinate system through visible light-infrared motion mapping. Under the premise that the two cameras have been spatially calibrated, describe the projection relationship of the same spatial plane on the visible light imaging plane and the infrared imaging plane. Map the motion estimated in the visible light coordinate system to the infrared image coordinate system to obtain a motion transformation suitable for infrared images.
[0074] S22. Motion compensation is performed on the original infrared image sequence based on the inter-frame motion in the infrared coordinate system to complete the inter-frame registration and obtain the spatially aligned infrared image sequence.
[0075] The original infrared image sequence is aligned frame by frame. For the t-th frame of the infrared image, the cumulative motion transformation from it to the reference frame needs to be calculated. This can be obtained using chain multiplication: (Assume that the accumulation starts from frame 0 and ends at frame t).
[0076] For each integer coordinate pixel in the aligned image, its corresponding, usually non-integer, coordinate in the original infrared image is found using the inverse of the cumulative transform. An interpolation algorithm (such as bilinear interpolation) is then used to calculate the pixel value of the aligned image based on the pixel values surrounding (x,y) in IR(t). This process is repeated across all pixels in the aligned image to obtain the spatially aligned infrared image sequence. Direct feature matching on two low-contrast, low-texture infrared images is very difficult and unreliable. Accurate motion calculated from high-texture visible light images fundamentally avoids the inherent limitations of infrared images. By stabilizing the infrared sequence, the visible light and infrared sequences maintain consistency in their motion trajectories.
[0077] S3. Perform spatiotemporal statistics on the spatially aligned infrared image sequence to estimate the gain drift coefficient and bias drift coefficient.
[0078] S31. For an infrared image sequence that has completed inter-frame registration and ensures that the same ground feature always falls at the same pixel location, a fixed-length sliding window is opened at each pixel along the time direction. Multiple frames of grayscale values are continuously stored within the window, forming a pure time signal. The average value and dispersion of this signal are calculated to obtain the grayscale level and fluctuation magnitude of the pixel in the current time period. Since all frames have completed inter-frame registration through visible light motion priors, the grayscale values within the window can be regarded as radiometric sampling of the same ground feature at continuous moments, minimizing spatial misalignment caused by scene motion.
[0079] S32. Within the sliding window, calculate the temporal mean μ and variance σ² of the grayscale sequence for this pixel. Simultaneously, introduce a "scene static" criterion: if the difference between the current window mean and the reference mean is less than a preset threshold (e.g., 5DN), it is considered that the scene radiation has not changed significantly, and the next step can proceed; otherwise, the window is discarded. This can shield grayscale jumps caused by target movement or sudden heat sources, preventing them from being misjudged as detector drift.
[0080] S33, with a preset reference gain G ref Reference bias O ref and reference standard deviation σ ref Based on the mean and variance of the pixel's grayscale value, the gain drift coefficient and bias drift coefficient are obtained respectively:
[0081] Compare the current standard deviation σ with σ ref The ratio is directly used as the gain drift coefficient α, reflecting the change in pixel response sensitivity;
[0082] Compare the current mean μ with the "expected gray level" G ref ·L ref +O refThe difference is used as the offset drift coefficient β, which reflects the offset of the pixel's DC operating point, where L ref This represents the average dark signal level.
[0083] Finally, an exponential low-pass filter (IIR) is applied to both α and β to suppress random noise and obtain smooth and reliable drift coefficients.
[0084] S41. Based on the estimated gain drift coefficient and offset drift coefficient, construct a gain drift coefficient matrix and an offset drift coefficient matrix with the same resolution as the original infrared image frame. The gain drift coefficient matrix and offset drift coefficient matrix are updated according to at least one of the following factors: ambient temperature, operating time, or scene statistical characteristics. Ambient temperature trigger: The internal temperature sensor triggers every 30 seconds.
[0085] S4. Based on the estimated gain drift coefficient and bias drift coefficient, perform non-uniformity correction on the current infrared image and output the corrected infrared image.
[0086] If ΔT exceeds the set threshold after each reading, the estimated gain drift coefficient and bias drift coefficient are immediately updated across the entire map. Working duration trigger: During continuous operation, a forced update is performed once per cycle to prevent long-term noise accumulation. Scene statistics trigger: When the global average scene radiation changes significantly and persists for a certain number of frames, a local incremental update is initiated.
[0087] S42. First, perform a subtraction and a division operation on each pixel of the original infrared frame: subtract the offset drift coefficient at the corresponding position from the original grayscale value of the current pixel, and then divide by the gain drift coefficient at the same position to obtain the linearly corrected grayscale value. To prevent the calculation result from exceeding the valid range, negative or excessively high values are truncated, retaining the valid range and temporarily storing it with a higher bit width for subsequent processing. Then, the dynamic range mapping stage begins. First, the grayscale distribution of the entire corrected image is statistically analyzed, and a very small number of the brightest and darkest pixels are automatically removed. The remaining part is linearly stretched to the standard eight-bit grayscale range. Then, adaptive histogram equalization with limited contrast is performed in a small area to make local details more prominent while suppressing excessive enhancement in large areas. If manual review is required, the grayscale can be further converted into pseudo-color. The entire process simultaneously outputs high-bit-depth original corrected data for algorithm closure and low-bit-depth enhanced images for display or network inference. The two data streams are kept synchronized with minimal latency. After correction, the noise of the fixed pattern in the image is significantly reduced, the edges of moving targets are clean, and there is no visible trailing or ghosting.
[0088] S5. The corrected infrared image and the visible light image are fused pixel by pixel to perform target detection.
[0089] S51. Image acquisition: First, select a visible light frame that is perfectly aligned with the timestamp of the current calibrated infrared frame. Since the two cameras have already completed coaxial installation and pixel-level registration, the visible light image can be directly regarded as corresponding to the infrared image pixel by pixel, without the need for additional registration.
[0090] S52. Detail Feature Map and Weight Map: For infrared images, calculate their local contrast to obtain a detail map that reflects the edges and intensity abrupt changes of the heat source; for visible light images, calculate their Laplacian energy or gradient magnitude to obtain a detail map that reflects texture and structural information. Compare the two detail maps pixel by pixel, and generate a fused weight map with the same resolution as the image, according to the principle of "the clearer the better." High weights are assigned to areas with rich visible light textures and areas with prominent infrared heat sources, while the weights for other areas transition smoothly to avoid boundary jumps.
[0091] S53. Pixel-level weighted fusion: Using the aforementioned weight map, the infrared and visible light images are weighted and added pixel by pixel: channels with larger weights contribute more, and channels with smaller weights contribute less. The fused image retains both the details and textures of the visible light and the thermal characteristics of the infrared target, resulting in a natural overall appearance with continuous edges and no ghosting or halos.
[0092] S54. Object Detection: The fused image is fed into a pre-trained object detection network. Within the same frame, the network utilizes both visible light texture information and high-contrast infrared features to output more reliable object categories and locations. The final result includes the category label and bounding box coordinates for each object.
[0093] Example 2, as one possible implementation, provides a system for implementing a coaxial motion drift self-correction method, comprising:
[0094] Infrared detector module for continuously outputting infrared image sequences without the need for mechanical baffles;
[0095] The visible light imaging module is used to synchronously acquire visible light image sequences of the same scene with the infrared detector module and output them as priors of scene motion.
[0096] The motion-scene analysis module, coupled with the visible light imaging module and the infrared detector module, is used to perform optical flow or feature point motion detection based on visible light image sequences and / or infrared image sequences to generate scene motion confidence signals.
[0097] The inter-frame registration module, coupled with the motion-scene analysis module, is used to perform inter-frame registration of the infrared image sequence using the scene motion prior provided by the visible light image sequence when the scene motion confidence signal indicates that there is sufficient motion, so as to obtain the spatially aligned infrared image sequence.
[0098] The spatiotemporal statistical estimation module, coupled with the inter-frame registration module and the motion-scene analysis module, is used to perform statistical operations on the spatially aligned infrared image sequence in the spatiotemporal domain under the control of the scene motion confidence signal, in order to estimate the gain drift coefficient and bias drift coefficient that change slowly over time.
[0099] The non-uniformity correction module, coupled with the spatiotemporal statistical estimation module and the infrared detector module, is used to perform pixel-level non-uniformity correction on the current infrared image using the latest estimated gain drift coefficient matrix and bias drift coefficient matrix, and output the corrected infrared image.
[0100] The pixel-level fusion module, coupled with the non-uniformity correction module and the visible light imaging module, is used to perform pixel-level fusion of the corrected infrared image and the corresponding visible light image to generate a fused image.
[0101] The target detection module, coupled with the pixel-level fusion module, is used to perform target detection on the fused image and output target category and location information.
[0102] Example 3, as a possible implementation, provides an electronic device including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, it implements a coaxial motion drift self-correction method.
[0103] Example 4, as one possible implementation, provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements a coaxial motion drift self-correction method.
[0104] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for self-correcting coaxial motion drift, characterized in that, The specific steps include the following: Acquire a sequence of visible light images as a priori output for scene motion; By utilizing scene motion priors, inter-frame registration is performed on the infrared image sequence to obtain a spatially aligned infrared image sequence. Spatiotemporal statistics are performed on spatially aligned infrared image sequences to estimate the gain drift coefficient and bias drift coefficient. The non-uniformity correction of the current infrared image is performed based on the estimated gain drift coefficient and bias drift coefficient, and the corrected infrared image is output. The corrected infrared image and the visible light image are fused pixel by pixel for target detection.
2. The coaxial motion drift self-correction method according to claim 1, characterized in that, The acquisition of the visible light image sequence as a priori output for scene motion specifically includes: Acquire visible light image sequences with coaxial mounting and extract feature points from the visible light image sequences; Feature points are extracted from two adjacent frames and feature matching is performed to generate an initial matching set; Foreground filtering is performed on the initial matching set to remove mismatches corresponding to moving targets, resulting in a background matching set; Based on the background matching set, a robust estimation algorithm is used to solve the global motion model parameters, and the global motion vector of the camera is obtained as the scene motion prior output.
3. The coaxial motion drift self-correction method according to claim 2, characterized in that, The extraction of feature points from the visible light image sequence specifically includes: An integral image is generated pixel-by-pixel from the input visible light image sequence, enabling the grayscale sum of any rectangular region to be obtained in constant time. The second-order differential image is obtained by approximating the second-order Gaussian derivative in the x, y, and xy directions using a box filter; Calculate the Hessian determinant response value of each pixel. If the response value of a candidate point is greater than that of n neighbors at the same time, the output is the visible light image sequence feature point.
4. The coaxial motion drift self-correction method according to claim 1, characterized in that, The step of using scene motion priors to perform inter-frame registration of the infrared image sequence to obtain a spatially aligned infrared image sequence specifically includes: The camera's global motion vector is converted into inter-frame motion in the infrared coordinate system through visible light-infrared motion mapping; Motion compensation is performed on the original infrared image sequence based on inter-frame motion in the infrared coordinate system to complete inter-frame registration and obtain a spatially aligned infrared image sequence.
5. The coaxial motion drift self-correction method according to claim 1, characterized in that, The method of performing spatiotemporal statistics on the spatially aligned infrared image sequence to estimate the gain drift coefficient and bias drift coefficient specifically includes: For a spatially aligned infrared image sequence, a sliding window with a length of N frames is established along the time dimension at each pixel location. Calculate the mean and variance of the pixel's grayscale value within the sliding window; Using the preset reference standard deviation and reference mean as a benchmark, and combining the mean and variance of the pixel's grayscale value, the gain drift coefficient and bias drift coefficient are obtained respectively.
6. The coaxial motion drift self-correction method according to claim 1, characterized in that, The process of performing non-uniformity correction on the current infrared image based on the estimated gain drift coefficient and bias drift coefficient, and outputting the corrected infrared image, specifically includes: Based on the estimated gain drift coefficient and offset drift coefficient, a gain drift coefficient matrix and an offset drift coefficient matrix with the same resolution as the original infrared image frame are constructed, wherein the gain drift coefficient matrix and the offset drift coefficient matrix are updated with at least one of the following factors: ambient temperature, working time or scene statistical characteristics. Each pixel of the original infrared image is corrected based on the gain drift coefficient matrix and the bias drift coefficient matrix with the same resolution as the original infrared image frame. The corrected pixel value array is dynamically range mapped to obtain the corrected infrared image and output it.
7. The coaxial motion drift self-correction method according to claim 1, characterized in that, The step of fusing the corrected infrared image and the visible light image at the pixel level for target detection specifically includes: Acquire a visible light image at the same time as the corrected infrared image, which has already been pixel-level registered; Detail feature maps are extracted from the corrected infrared image and visible light image respectively, and the fusion weight of each pixel is calculated based on the detail feature maps to obtain a weight map; The corrected infrared image and the visible light image are fused pixel-level based on the weight map to generate a fused image; The fused image is input into a pre-trained target detection network to obtain and output the target category and location information.
8. A system, characterized in that, The coaxial motion drift self-correction method applied to any one of claims 1-7 includes: Infrared detector module for continuously outputting infrared image sequences without the need for mechanical baffles; The visible light imaging module is used to synchronously acquire visible light image sequences of the same scene with the infrared detector module, and to output the scene motion prior. A motion-scene analysis module, coupled to the visible light imaging module and the infrared detector module, is used to perform optical flow or feature point motion detection based on the visible light image sequence and / or the infrared image sequence to generate a scene motion confidence signal; The inter-frame registration module, coupled to the motion-scene analysis module, is used to perform inter-frame registration on the infrared image sequence using the scene motion prior provided by the visible light image sequence when the scene motion confidence signal indicates that there is sufficient motion, so as to obtain a spatially aligned infrared image sequence. The spatiotemporal statistical estimation module, coupled with the inter-frame registration module and the motion-scene analysis module, is used to perform statistical operations on the spatially aligned infrared image sequence in the spatiotemporal domain under the control of the scene motion confidence signal, so as to estimate the gain drift coefficient and bias drift coefficient that change slowly over time. The non-uniformity correction module, coupled to the spatiotemporal statistical estimation module and the infrared detector module, is used to perform pixel-level non-uniformity correction on the current infrared image using the latest estimated gain drift coefficient matrix and the bias drift coefficient matrix, and output the corrected infrared image. A pixel-level fusion module, coupled to the non-uniformity correction module and the visible light imaging module, is used to perform pixel-level fusion of the corrected infrared image and the corresponding visible light image to generate a fused image; The target detection module, coupled to the pixel-level fusion module, is used to perform target detection on the fused image and output target category and location information.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the program, it implements the coaxial motion drift self-correction method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by the processor, it implements the coaxial motion drift self-correction method as described in any one of claims 1 to 7.
Citation Information
Cited By
Infrared imaging shutter-free temperature drift correction method fusing physical information reinforcement learning
CN122084122A