A general camera near-infrared image generation method
By employing adaptive subpixel image registration, spectral decoupling and dynamic gain compensation, and guided filtering techniques, the problems of spectral crosstalk and nonlinear deformation in near-infrared image generation by general-purpose CMOS cameras are solved, achieving high-precision and low-cost near-infrared image generation suitable for various application scenarios.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-03-06
- Publication Date
- 2026-06-19
Smart Images

Figure CN121810828B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of image processing, and more particularly to a method for generating near-infrared images from a universal camera. Background Technology
[0002] With the rapid development of computer vision technology, artificial intelligence, and security monitoring, camera modules, as the entry point for image information acquisition, are increasingly widely used in various scenarios. Among them, near-infrared imaging technology has become an important component of the current image processing and acquisition field due to its unique advantages in night vision, penetration of fog and haze, and biometric recognition.
[0003] Existing near-infrared image generation methods rely on dedicated near-infrared cameras or complex optical systems, which are expensive, bulky, and difficult to deploy. They also cannot utilize existing general-purpose camera resources, limiting the widespread application of near-infrared imaging technology. At the same time, the RGB filters of general-purpose CMOS cameras have spectral crosstalk in the near-infrared band. Simply subtracting two frames of images taken by switching filters directly will produce a serious color channel imbalance problem, resulting in poor quality of the generated near-infrared images. At the algorithm level, high-temperature air turbulence will cause nonlinear deformation of the camera image. In particular, there is a time difference of 33ms to 100ms between two frames of images under telephoto lenses. The image shift caused by turbulence will produce serious edge artifacts when directly subtracting. Summary of the Invention
[0004] To address the shortcomings of existing technologies, this invention provides a universal near-infrared image generation method for cameras, thus solving the above problems.
[0005] To achieve the above objectives, the present invention provides the following technical solution: a universal near-infrared image generation method for cameras, comprising the following steps:
[0006] S1. Adaptive subpixel image registration: Acquire two images, one under the open filter and the other under the open filter. Images under cutoff filters The two frames of images are aligned at the sub-pixel level based on optical flow or feature point pyramid algorithm, with an alignment accuracy of 0.1 pixels, in order to eliminate nonlinear deformation caused by high-temperature air turbulence.
[0007] S2. Spectral Decoupling and Dynamic Gain Compensation: The following mathematical model is applied to the registered image to calculate the near-infrared image:
[0008] ;in It is the exposure gain compensation factor. It is a spectral correction matrix;
[0009] S3. Edge enhancement based on guided filtering: The visible light image is used as the guiding image, and the near-infrared image is enhanced using joint bilateral filtering or guided filtering. Edge enhancement and noise suppression are performed to transfer the high signal-to-noise ratio structural information of the visible light image to the near-infrared image, while the low sensitivity of the near-infrared band to turbulence is used to suppress image jitter.
[0010] Preferably, the sub-pixel alignment based on optical flow in step S1 includes:
[0011] S11. Construct an energy function that includes data items and a smoothing term:
[0012] ;in For pixels Displacement vector field at point, and These are the displacements in the horizontal and vertical directions, respectively. Represents the image spatial domain. The regularization coefficient is... and They are respectively and Spatial gradient;
[0013] S12. The displacement field is solved by minimizing the energy function E(u) using an iterative optimization algorithm, and the integer pixel displacement is refined to sub-pixel level through Taylor expansion until the displacement increment is less than 0.05 pixels, achieving an alignment accuracy of 0.1 pixels.
[0014] Preferably, the sub-pixel alignment based on the feature point pyramid in step S1 includes:
[0015] S13. Construct a Gaussian pyramid hierarchical structure. The top layer uses feature point matching to estimate the initial transformation matrix, and the correlation coefficient is maximized layer by layer.
[0016] Where w(x,y;P) represents the coordinates of the pixel (x,y) after projection transformation according to parameter P. and Images and The average gray level in the corresponding region Represents the image spatial domain;
[0017] S14. Solve using an iterative optimization algorithm. Maximize the optimal transformation parameters And 0.1 pixel-level alignment accuracy is achieved through bicubic interpolation.
[0018] Preferably, the exposure gain compensation coefficient in step S2 The real-time estimation method is as follows: Construct the background noise energy function:
[0019] ;in Representing the image spatial domain, by minimizing the energy function Solving for the optimal The value is chosen to minimize the background noise energy of the difference image.
[0020] Preferably, in step S2, the spectral correction matrix This is a 3×3 matrix used to correct spectral crosstalk of the RGB filter in the near-infrared band of a CMOS sensor, and its form is as follows:
[0021] ;in , , These correspond to the correction coefficients for the red, green, and blue color channels, respectively.
[0022] Preferably, the joint bilateral filtering in step S3 uses the following filter kernel:
[0023] ;in Image guided by visible light. For pixels The neighborhood, and These are the spatial Gaussian kernel and the range Gaussian kernel, respectively. This is the normalization factor.
[0024] Preferably, the guiding filter in step S3 adopts a local linear model, assuming the guiding image... With output image In local window The inner linear relationship is satisfied:
[0025] ; where linear coefficients and By minimizing the following cost function in the window Internal solution:
[0026] ;in This is the regularization parameter.
[0027] Preferably, the image under the open filter Images under cutoff filters The acquisition time difference is 33ms to 100ms.
[0028] Preferably, the edge enhancement based on guided filtering employs joint bilateral filtering or guided filtering, using visible light images to guide NIR image reconstruction.
[0029] Preferably, the method is used to suppress image jitter caused by high-temperature air turbulence and improve the accuracy of image edge recognition.
[0030] This invention provides a universal method for generating near-infrared images from a camera. Compared with existing technologies, it has the following advantages:
[0031] This invention effectively solves key problems such as image jitter caused by high-temperature air turbulence and edge artifacts, color channel imbalance, and low signal-to-noise ratio in traditional near-infrared image generation through a three-step processing flow: adaptive subpixel image registration, spectral decoupling and dynamic gain compensation, and edge enhancement based on guided filtering. The method employs subpixel alignment technology based on optical flow or feature point pyramids, achieving an accuracy of 0.1 pixels and eliminating nonlinear deformation caused by turbulence. Real-time online estimation of compensation coefficients using spectral decoupling and dynamic gain compensation models solves the problems of exposure differences and uneven spectral response. Edge enhancement and noise suppression are performed using a combined bilateral filter or guided filter guided by visible light images, transferring clear structural information to near-infrared images. Simultaneously, the low sensitivity of the near-infrared band to turbulence is utilized to suppress image jitter. The overall solution requires no dedicated near-infrared imaging hardware, is applicable to general CMOS cameras, processes image frames with a time difference of 33ms to 100ms, and has good hardware compatibility and real-time processing capabilities.
[0032] 2. This invention eliminates the need for dedicated near-infrared imaging hardware; it only requires adding a switchable filter to a general-purpose CMOS camera, reducing hardware costs by over 70%. It can process image frames with time differences ranging from 33ms to 100ms, demonstrating excellent real-time processing capabilities. Through flexible configuration of algorithm parameters, it can adapt to various application scenarios such as real-time video stream monitoring, drone inspection, bridge health monitoring, and building deformation measurement. The real-time processing version can reach 30fps, while the high-precision monitoring version can achieve sub-pixel displacement monitoring accuracy of 0.01 pixels, equivalent to micrometer-level actual displacement resolution in telephoto scenarios. This provides a low-cost, high-precision, and easily deployable technical solution for visual monitoring in high-temperature and turbulent environments. Attached Figure Description
[0033] Figure 1 This is a flowchart of a general near-infrared image generation method for cameras proposed in this invention. Detailed Implementation
[0034] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0035] Please see Figure 1 The present invention provides the following technical solutions, specifically including the following embodiments:
[0036] Example 1:
[0037] A general method for generating near-infrared images from a camera includes the following steps:
[0038] S1. Adaptive subpixel image registration: Acquire two images, one under the open filter and the other under the open filter. Images under cutoff filters The two frames are aligned at the sub-pixel level using optical flow or feature point pyramid algorithms, achieving an alignment accuracy of 0.1 pixels, to eliminate nonlinear deformation caused by high-temperature air turbulence, thus improving the image quality under the open filter. Images under cutoff filters The acquisition time difference is 33ms to 100ms, and the sub-pixel alignment based on optical flow includes:
[0039] S11. Construct an energy function that includes data items and a smoothing term:
[0040] ;in For pixels Displacement vector field at point, and These are the displacements in the horizontal and vertical directions, respectively. Represents the image spatial domain. The regularization coefficient is... and They are respectively and Spatial gradient;
[0041] S12. The displacement field is solved by minimizing the energy function E(u) using an iterative optimization algorithm, and the integer pixel displacement is refined to sub-pixel level through Taylor expansion until the displacement increment is less than 0.05 pixels, achieving an alignment accuracy of 0.1 pixels.
[0042] Sub-pixel alignment based on feature point pyramids includes:
[0043] S13. Construct a Gaussian pyramid hierarchical structure. The top layer uses feature point matching to estimate the initial transformation matrix, and the correlation coefficient is maximized layer by layer.
[0044] Where w(x,y;P) represents the coordinates of the pixel (x,y) after projection transformation according to parameter P. and Images and The average gray level in the corresponding region Represents the image spatial domain;
[0045] S14. Solve using an iterative optimization algorithm. Maximize the optimal transformation parameters And 0.1 pixel-level alignment accuracy is achieved through bicubic interpolation.
[0046] S2. Spectral Decoupling and Dynamic Gain Compensation: The following mathematical model is applied to the registered image to calculate the near-infrared image:
[0047] ;in It is the exposure gain compensation factor. It is the spectral correction matrix and the exposure gain compensation coefficient. The real-time estimation method is as follows: Construct the background noise energy function:
[0048] ;in Representing the image spatial domain, by minimizing the energy function Solving for the optimal The value is chosen to minimize the background noise energy of the difference image, and the spectral correction matrix is used. This is a 3×3 matrix used to correct spectral crosstalk of the RGB filter in the near-infrared band of a CMOS sensor, and its form is as follows:
[0049] ;in , , The correction coefficients correspond to the red, green, and blue color channels, respectively.
[0050] S3. Edge enhancement based on guided filtering: The visible light image is used as the guiding image, and the near-infrared image is enhanced using joint bilateral filtering or guided filtering. Edge enhancement and noise suppression are performed to transfer the high signal-to-noise ratio structural information of the visible light image to the near-infrared image. Simultaneously, the low sensitivity of the near-infrared band to turbulence is utilized to suppress image jitter. Edge enhancement based on guided filtering employs joint bilateral filtering or guided filtering. NIR image reconstruction is guided by the visible light image. The joint bilateral filtering uses the following filter kernel:
[0051] ;in Image guided by visible light. For pixels The neighborhood, and These are the spatial Gaussian kernel and the range Gaussian kernel, respectively. As a normalization factor, the guided filter uses a local linear model, assuming the guided image... With output image In local window The inner linear relationship is satisfied:
[0052] ; where linear coefficients and By minimizing the following cost function in the window Internal solution:
[0053] ;in This is the regularization parameter.
[0054] Example 2:
[0055] Based on Example 1, the difference between this example and Example 1 is that the algorithm is optimized for real-time video stream scenarios to reduce computational complexity and meet real-time requirements;
[0056] S1. Adaptive subpixel image registration: Acquiring two frames of images with a time difference of 33ms (corresponding to a frame rate of 30fps), i.e., the image under the open filter. Images under cutoff filters Subpixel alignment based on fast optical flow includes:
[0057] S2. Spectral Decoupling and Dynamic Gain Compensation: A lookup table method is used to accelerate exposure gain compensation. A lookup table for β values under different lighting conditions is pre-calculated, and the optimal β value is quickly indexed based on the average brightness of the current frame. The spectral correction matrix uses preset calibration values and is not estimated online to reduce computational overhead.
[0058] S3. Edge enhancement based on fast guided filtering: Fast guided filtering is used to replace standard guided filtering. Linear coefficients are approximated by box filtering. The calculation is completed in constant time complexity using integral graph technology, and the overall processing speed is improved to 30fps real-time processing.
[0059] Example 3: The difference between this example and Example 1 is that it is optimized for long-distance monitoring scenarios using telephoto lenses, focusing on solving the severe image shift and blurring problems caused by atmospheric turbulence.
[0060] S1. Adaptive Subpixel Image Registration: Two frames of images with a time difference of 100ms are acquired to fully separate random displacements caused by turbulence. A method combining phase-correlation-based frequency domain registration and feature point pyramids is adopted. First, the pixel-level initial displacement is obtained through the inverse Fourier transform of the cross-power spectrum. Then, the residual nonlinear deformation is estimated using the dense optical flow method. A turbulence constraint term is introduced to suppress high-frequency jitter, and the alternating direction multiplier method is used for iterative solution.
[0061] S2. Spectral Decoupling and Adaptive Dynamic Gain Compensation: Adaptive window estimation is employed to address the drastic lighting variations characteristic of long-focus scenes. Values. Divide the image into multiple sub-blocks and estimate the value independently for each sub-block. The values are analyzed, outliers are removed through a consistency check, a spatially adaptive exposure compensation field is constructed, and then bilinear interpolation is used to obtain a continuous value across the entire image. Distribution. The spectral correction matrix employs an online calibration strategy, using a known neutral gray reference region in the scene to estimate the response of each channel in real time;
[0062] S3. Edge Enhancement Based on Multi-Scale Guided Filtering: Addressing the rich detail but high noise characteristics of telephoto images, multi-scale guided filtering fusion is employed. Guided filtering is performed at three scales: the original image, half-downsampled, and quarter-downsampled. The final image is then reconstructed using a Laplacian pyramid fusion method, with adaptive weights determined based on the local variance of each scale. The final output near-infrared image is used for sub-pixel displacement monitoring, achieving a monitoring accuracy of 0.01 pixels.
[0063] The following are the test results of the above embodiments:
[0064]
[0065] This invention effectively solves the problems of image jitter caused by high-temperature air turbulence and edge artifacts, color channel imbalance, and low signal-to-noise ratio in traditional near-infrared image generation through a three-step processing flow: adaptive sub-pixel image registration, spectral decoupling and dynamic gain compensation, and edge enhancement based on guided filtering. This method requires no dedicated hardware, is applicable to general CMOS cameras, achieves registration accuracy at the 0.1 pixel level, and can process image frames with time differences from 33ms to 100ms, significantly improving near-infrared image quality and stability. It effectively enhances the accuracy of edge recognition and displacement monitoring in telephoto scenarios. Based on Example 1, the algorithm is optimized for real-time video stream scenarios. By using the pyramid Lucas-Kanade optical flow method and lookup table method to accelerate exposure gain compensation and fast guided filtering, the processing speed is increased to 30fps real-time processing, with a single frame processing time of less than 15ms, meeting the processing requirements of real-time video streams. It is suitable for applications with high real-time requirements such as video surveillance and drone inspection, and is deeply optimized for long-distance monitoring scenarios. By combining frequency domain coarse registration with turbulence-constrained optical flow, sub-block adaptive gain compensation, and multi-scale guided filtering fusion, nonlinear deformation caused by atmospheric turbulence is effectively suppressed, registration accuracy is improved to 0.05 pixels, dynamic range is improved by 6dB, signal-to-noise ratio is improved by more than 8dB, and sub-pixel displacement monitoring accuracy of 0.01 pixels can be achieved.
[0066] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for universal camera near-infrared image generation, characterized in that: Includes the following steps: S1, adaptive sub-pixel image registration: collect two images, respectively, the image under the open filter and the image under the cut-off filter , based on the optical flow method or the feature point pyramid algorithm, the two images are aligned at the sub-pixel level, and the alignment accuracy reaches 0.1 pixel level, to eliminate the nonlinear deformation caused by high temperature air turbulence; S2. Spectral Decoupling and Dynamic Gain Compensation: The following mathematical model is applied to the registered image to calculate the near-infrared image: ;in It is the exposure gain compensation factor. It is a spectral correction matrix; Exposure gain compensation coefficient in step S2 The real-time estimation method is as follows: Construct the background noise energy function: ;in Representing the image spatial domain, by minimizing the energy function Solving for the optimal The value is chosen to minimize the background noise energy of the difference image; Spectral correction matrix in step S2 This is a 3×3 matrix used to correct spectral crosstalk of the RGB filter in the near-infrared band of a CMOS sensor, and its form is as follows: ;in , , The correction coefficients correspond to the red, green, and blue color channels, respectively. S3. Edge enhancement based on guided filtering: The visible light image is used as the guiding image, and the near-infrared image is enhanced using joint bilateral filtering or guided filtering. Edge enhancement and noise suppression are performed to transfer the high signal-to-noise ratio structural information of the visible light image to the near-infrared image, while the low sensitivity of the near-infrared band to turbulence is used to suppress image jitter.
2. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: The sub-pixel alignment based on optical flow in step S1 includes: S11. Construct an energy function that includes data items and a smoothing term: ;in For pixels Displacement vector field at point, and These are the displacements in the horizontal and vertical directions, respectively. Represents the image spatial domain. The regularization coefficient is... and They are respectively and Spatial gradient; S12. The displacement field is solved by minimizing the energy function E(u) using an iterative optimization algorithm, and the integer pixel displacement is refined to sub-pixel level using Taylor expansion until the displacement increment is less than 0.05 pixels, achieving an alignment accuracy of 0.1 pixels.
3. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: The sub-pixel alignment based on the feature point pyramid in step S1 includes: S13. Construct a Gaussian pyramid hierarchical structure. The top layer uses feature point matching to estimate the initial transformation matrix, and the correlation coefficient is maximized layer by layer. Where w(x,y;P) represents the coordinates of the pixel (x,y) after projection transformation according to parameter P. and Images and The average gray level in the corresponding region Represents the image spatial domain; S14. Solve using an iterative optimization algorithm. Maximize the optimal transformation parameters And 0.1 pixel-level alignment accuracy is achieved through bicubic interpolation.
4. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: The joint bilateral filtering in step S3 uses the following filtering kernel: ;in Image guided by visible light. For pixels The neighborhood, and These are the spatial Gaussian kernel and the range Gaussian kernel, respectively. This is the normalization factor.
5. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: The guiding filter in step S3 uses a local linear model, assuming the guiding image... With output image In local window The inner linear relationship is satisfied: Where linear coefficients and By minimizing the following cost function in the window Internal solution: ;in This is the regularization parameter.
6. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: Image under the open filter Images under cutoff filters The acquisition time difference is 33ms to 100ms.
7. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: The edge enhancement based on guided filtering employs joint bilateral filtering or guided filtering, using visible light images to guide NIR image reconstruction.
8. The method for generating near-infrared images from a universal camera according to claim 1, characterized in that: The method is used to suppress image jitter caused by high-temperature air turbulence and improve the accuracy of image edge recognition.
Citation Information
Patent Citations
Image fusion model training method, image fusion method, electronic equipment and medium
CN120147150A
Image processing device, control method for the same and program
JP2016066892A