A cross-domain encoding and decoding near-infrared image color generation system and method
Through the cross-domain encoding and decoding near-infrared image color generation system, the problems of noise interference, edge loss and color loss in near-infrared image color generation are solved, and more accurate color mapping and visual effects are achieved, which is applied to night monitoring, unmanned driving, agricultural remote sensing and biomedical imaging.
Patent Information
- Application Number
- CN202411158202.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-08-22
- Publication Date
- 2025-09-23
- Estimated Expiration
- 2044-08-22
AI Technical Summary
Existing near-infrared image color generation technologies fail to effectively retain and utilize the contextual information of the image, resulting in inaccurate and lack of intuitive color generation.
A cross-domain encoding and decoding near-infrared image color generation system is adopted, including a preprocessing module, a color content encoding module and a cross-domain decoding module. Deep learning technology is used to extract and map the features of near-infrared images and visible light images. Through denoising edge preservation, color attention units and feature network pyramid units, image noise removal, edge preservation and color generation are achieved.
The color generation effect and visual quality of near-infrared images have been significantly improved, and the generated images are more detailed and realistic, suitable for night monitoring, unmanned driving, agricultural remote sensing, biomedical imaging and other fields.
Smart Images

Figure CN119107379B_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of infrared image processing and display technology, and specifically relates to a cross-domain encoding and decoding near-infrared image color generation system and method. Background Art
[0002] Due to its unique penetrating properties and sensitivity to material properties, near-infrared light offers significant advantages in a wide range of fields. It not only penetrates obstructions such as fog and smoke, providing clear nighttime vision, but also offers high accuracy in detecting plant moisture content and soil composition. However, near-infrared images are typically presented in grayscale and lack intuitive color information, limiting their intuitiveness and application. Therefore, research on near-infrared image color generation techniques to imbue monochrome near-infrared images with vivid color is crucial.
[0003] Near-infrared image color generation methods are divided into two categories. One is the traditional colorization method, which maps the near-infrared image to the color space through histogram matching, pseudo-color mapping, etc. This method has low computational complexity and is simple and easy to implement, but its effect is limited by the limitations of traditional algorithms and cannot restore the color information of the real scene; the other is the learning-based method, which learns the mapping relationship between near-infrared images and color images through a large amount of training data to realize the color generation process. It has achieved good results in color generation and low information loss. However, this method does not consider the retention and utilization of the contextual information of the near-infrared image, and does not reasonably utilize the scene depth and material property information that are crucial to generating accurate result images. Summary of the Invention
[0004] The purpose of this application is to overcome the defect in the related art that the preservation and utilization of context information in near-infrared images are not considered.
[0005] To achieve the above objectives, the present application proposes a cross-domain encoding and decoding near-infrared image color generation system, the system comprising:
[0006] A preprocessing module is used to preprocess the selected near-infrared image to obtain a denoised near-infrared image;
[0007] a color content encoding module, configured to generate an encoded visible light color encoding map, a visible light content encoding map, a denoised near-infrared color encoding map, and a denoised near-infrared content encoding map from the visible light image and the denoised near-infrared image;
[0008] The cross-domain decoding module is used to decode the visible light color coding image and the denoised near-infrared content coding image to obtain a visible light near-infrared color image; and to decode the denoised near-infrared color coding image and the visible light content coding image to obtain a near-infrared visible light grayscale image.
[0009] As an improvement to the above system, the pre-processing module includes a denoising edge-preserving sub-module;
[0010] The denoising edge preservation submodule is used to 01 The calculation is the denoised near-infrared image I1 with noise removed and edge preserved;
[0011] The denoising edge-preserving submodule includes several convolution layers, dense connection units and contour convolution units connected in series to perform convolution operations; wherein,
[0012] The densely connected unit is used to perform feature extraction and noise suppression on the near-infrared image;
[0013] The contour convolution unit is used to enhance the edge of the image.
[0014] As an improvement of the above system, the color content encoding module includes a first color encoding submodule, a second color encoding submodule, a first content encoding submodule and a second content encoding submodule; wherein,
[0015] The first color coding submodule is used to convert the denoised near-infrared image I1 into the corresponding denoised near-infrared color coding image I 21 ;
[0016] The second color encoding submodule is used to convert the visible light image I 02 Converted into the corresponding denoised visible light color coding image I 23 ;
[0017] The first content encoding submodule is used to convert the denoised near-infrared image I1 into the corresponding denoised near-infrared content encoding image I 22 ;
[0018] The second content encoding submodule is used to convert the visible light image I 02 Converted into the corresponding denoised visible light content coding image I 24 .
[0019] As an improvement to the above system, the first color coding submodule and the second color coding submodule have the same structure, both including a number of convolution units, depthwise separable convolution units and color attention units; wherein,
[0020] The convolution units extract the denoised near-infrared image I1 and the visible light image I1 by stacking multiple convolution layers. 02 characteristics;
[0021] The depthwise separable convolution unit operates independently on each channel of the input image to extract the denoised near-infrared image I1 and the visible light image I 02 local features of
[0022] The color attention unit learns to identify important color features in the input image.
[0023] As an improvement to the above system, the first content encoding submodule and the second content encoding submodule have the same structure; both include a number of convolution units, hole channel attention units and feature network pyramid units; wherein,
[0024] The convolution unit is used to extract the denoised near-infrared image I1 and the visible light image I 02 shallow features of
[0025] The hole channel attention unit is used to expand the receptive field and enhance the feature representation of important channels, so that the model can capture a wider range of contextual information and pay more attention to the features that are beneficial to the color generation task;
[0026] The feature network pyramid unit connects the denoised near-infrared image I1 and the visible light image I1 through a top-down path and lateral connections. 02 Features at different levels are fused to generate a content encoding graph with semantic and detail information.
[0027] As an improvement of the above system, the cross-domain decoding module includes a fully connected network unit, an AdaIn unit, several residual units and an upsampling unit;
[0028] The fully connected network unit achieves fine control of color and content information by learning the mapping relationship between the two domains;
[0029] The AdaIn unit introduces the information of each target domain by adjusting the variance and mean of the input feature map, so that the generated visible near-infrared color image I 31 With the color style and visual effects of the visible light image domain, the generated near-infrared visible light grayscale image I 32 Color style and visual effects with near-infrared image domain;
[0030] The residual unit retains the low-frequency information of the input image by introducing a jump connection, and learns the mapping relationship of the high-frequency information at the same time, decoding the content coding map into a visible light near-infrared color image I 31 and near-infrared visible light grayscale image I 32 ;
[0031] The upsampling unit is used to restore the resolution of the feature map to the same size as the input image.
[0032] As an improvement to the above system, the loss function L total for:
[0033] Ltotal =λ1L con +λ2L col +λ3L recon
[0034] Among them, λ1, λ2, λ3 are hyperparameters;
[0035] L con is the content consistency loss function:
[0036] L con =||E con (I1)-E con (I 31 -I 02 )||1+||E con (I 02 )-E con (I 32 -I1)||1
[0037] Among them, E con is the first content encoding submodule or the second content encoding submodule; I1 is the denoised near-infrared image; I 31 is a visible light near infrared color image; I 02 is a visible light image; I 32 is the near-infrared visible light grayscale image; ||·||1 is the L1 norm;
[0038] L col is the color consistency loss function:
[0039] L col =||E col (I1)-E col (I 31 -I 02 )||1+||E col (I 02 )-E col (I 32 -I1)||1
[0040] Among them, E col is a first color coding submodule or a second color coding submodule;
[0041] L recon To reconstruct the loss function:
[0042] L recon =||I1-I 32 ||1+||I 02 -I 31 ||1.
[0043] This application also provides a cross-domain encoding and decoding near-infrared image color generation method, which is implemented based on the above system, and includes:
[0044] Step S1: Processing the selected near-infrared image through a pre-processing module to obtain a denoised near-infrared image;
[0045] Step S2: Inputting the visible light image and the denoised near-infrared image into a color content encoding module to extract near-infrared key content features, thereby obtaining an encoded visible light color coding map, a visible light content coding map, a denoised near-infrared color coding map, and a denoised near-infrared content coding map;
[0046] Step S3: Input the visible light color coding image and the denoised near-infrared content coding image into the cross-domain decoding module for decoding to obtain a visible light near-infrared color image; input the denoised near-infrared color coding image and the visible light content coding image into the cross-domain decoding module for decoding to obtain a near-infrared visible light grayscale image.
[0047] As an improvement to the above method, step S2 includes:
[0048] The denoised near-infrared image I1 is input into the first color coding submodule and the first content coding submodule to obtain the feature representation of the near-infrared domain color and content information, i.e., the denoised near-infrared color coding image I 21 And denoised near infrared content coding image I 22 ;
[0049] At the same time, the visible light image I 02 Input the second color coding submodule and the second content coding submodule to obtain the feature representation of visible light domain color and content information, that is, visible light color coding image I 23 and visible light content coding diagram I 24 .
[0050] Compared with the prior art, the advantages of this application are:
[0051] 1. The denoising edge-preserving submodule provided in this application can effectively identify and remove the noise in near-infrared images caused by insufficient ambient light and sensor noise, so that it can effectively identify and remove the existing noise, and can retain the edge information of the image to the greatest extent, so that it can more accurately identify different areas in the image and achieve more refined color mapping. The color encoding submodule combines convolution blocks, depthwise separable convolutions and color attention units to extract rich color features from the denoised near-infrared images and visible light images, and converts these features into features that the model can understand and use. It further learns the mapping relationship between the near-infrared image and the visible light image, and pays more attention to the important color features in the image through the color attention module, thereby enhancing the color expression and visual effect of the color generation results. The content encoding submodule combines convolution blocks, void channel attention units and feature network pyramid units to effectively extract deep features and contextual information of denoised near-infrared images and visible light images. By making full use of feature representations and contextual information at different levels, the content encoding submodule can generate visible light near-infrared color images that are delicate and realistic that match the near-infrared images, and generate near-infrared visible light grayscale images that match the visible light images; the decoding submodule combines residual units, fully connected network units, AdaIn units and upsampling units to realize color generation of near-infrared images and near-infrared information supplementation of visible light images. The cross-domain decoding module works together with the color content encoding module to ensure the efficient implementation of the entire cross-domain encoding and decoding near-infrared image color generation task.
[0052] 2. To solve the problems of noise interference and blurred edges in near-infrared images, this application designs a denoising edge-preserving submodule as a preprocessing stage for cross-domain encoding and decoding of near-infrared image color generation. This submodule uses densely connected units to extract features and suppress noise, and then uses contour convolution units to enhance the edges of the image. This not only removes noise from the near-infrared image but also retains edge information, providing preliminary processing for the subsequent color generation step. To solve the problems of missing color information and inaccurate mapping, this application designs a color content encoding module. This module uses deep learning technology to learn a large amount of training data and establish a mapping relationship between near-infrared image features and color information. During the color generation process, the color content encoding module can actively extract features of near-infrared images and visible light images, and assign appropriate color values to each pixel in the image through the learned mapping relationship. To solve the problem of resolution and detail restoration of near-infrared and visible light images, this application designs a cross-domain decoding module. This module gradually improves the resolution and details of the image while maintaining the image color. It introduces residual learning technology to effectively restore the edges, textures and other details of the denoised near-infrared and visible light images.
[0053] 3. The technical solution of the present application can significantly improve the color generation effect and visual quality of near-infrared images, and at the same time obtain visible light images with near-infrared information. By solving the problems of reducing noise interference, color loss, mapping inaccuracy and original resolution recovery of near-infrared images, images with richer colors and more details are provided, and they are applied to night monitoring, unmanned driving, agricultural remote sensing, biomedical imaging and other fields. In night monitoring, visible light near-infrared image color images can provide clearer target contours and color distinctions, and improve the recognition ability and response speed of the monitoring system. In the field of unmanned driving, this technology helps vehicles perceive the surrounding environment more accurately and improve the safety and reliability of autonomous driving. In agricultural remote sensing, visible light near-infrared color images can reveal the growth status and nutritional status of crops, providing strong support for precision agriculture. In the field of biomedical imaging, this technology helps doctors observe and analyze biological tissue structures more intuitively, improving the accuracy and efficiency of diagnosis. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 The figure shows a flow chart of a method for generating color of a near-infrared image by cross-domain encoding and decoding;
[0055] Figure 2 The figure shows the principle diagram of the cross-domain encoding and decoding near-infrared image color generation system;
[0056] Figure 3 The figure shows the structural principle diagram of the denoising edge preservation submodule;
[0057] Figure 4 Shown is a schematic diagram of the structural principle of the first color coding submodule and the second color coding submodule;
[0058] Figure 5 The figure shows the structural principle diagram of the first content encoding submodule and the second content encoding submodule;
[0059] Figure 6 The figure shows a schematic diagram of the structural principle of the cross-domain decoding module. DETAILED DESCRIPTION
[0060] The technical solution of this application is described in detail below with reference to the accompanying drawings.
[0061] This application provides a system and method for generating color for near-infrared images using cross-domain encoding and decoding. Specifically, it relates to a system and method for generating color for near-infrared images that combines color encoding and decoding of content in both the near-infrared and visible light domains. This application aims to address the limitations of current near-infrared image color generation technology.
[0062] In response to the problems of noise interference, edge loss, missing color information, inaccurate mapping and low resolution in the near-infrared image color generation algorithm, this application proposes a cross-domain encoding and decoding near-infrared image color generation system and method consisting of a preprocessing module 1, a color content encoding module 2 and a cross-domain decoding module 3.
[0063] The pre-processing module 1 includes a denoising edge-preserving sub-module 101 , which is intended to denoise the near-infrared image while preserving the edges, making it visually clearer and sharper.
[0064] Color content encoding module 2 includes a first color encoding submodule 201, a first content encoding submodule 202, a second color encoding submodule 203, and a second content encoding submodule 204. The first color encoding submodule 201 and the second color encoding submodule 203 employ the same structure, aiming to leverage deep learning techniques to establish complex mappings between denoised near-infrared images and color information, and between visible light images and near-infrared information. They automatically assign color values to each pixel in the near-infrared and visible light images, transforming the denoised near-infrared image from colorless to rich color information, and the visible light image from rich color information to near-infrared information, ensuring the accuracy and consistency of color information. The second content encoding submodule 202 and the second content encoding submodule 204 employ the same structure, aiming to leverage feature representations and contextual information at different levels to generate detailed, realistic visible light near-infrared color images that match the input near-infrared image, and near-infrared-visible grayscale images that match the input visible light image and contain near-infrared information.
[0065] The decoding module 3 includes a cross-domain decoding submodule 301, which ensures the stability of the network during training and the retention of low-frequency information of the image, and jointly achieves fine control of color and content information, making the resulting image more diverse in color style and visual effects.
[0066] like Figure 1 As shown, the present application provides a cross-domain encoding and decoding near-infrared image color generation method, which adopts a cross-domain near-infrared image color generation network composed of a denoising edge preservation module and a deep learning network. The network includes a preprocessing module 1, a color content encoding module 2 and a cross-domain decoding module 3; the method includes the following steps:
[0067] S1: Process the selected near-infrared image through the preprocessing module 1 to obtain a denoised near-infrared image;
[0068] The near infrared image is input to the pre-processing module 1, which specifically includes a denoising edge preservation submodule 101. 01 Calculated as the denoised near-infrared image I1;
[0069] The denoising edge preservation submodule 101 assists in denoising the near-infrared image while preserving edge information as much as possible, enhancing the detail expression of the image and better representing the texture and other information of the near-infrared image;
[0070] The denoising edge-preserving submodule 101 includes a series of convolution layers, dense connection modules and contour convolution modules that are connected in series to perform convolution operations. The dense connection modules are used to extract near-infrared image features and suppress noise. The contour convolution modules are used to predict and enhance the image edges to generate a denoised near-infrared image I1 with noise removed and edge preserved.
[0071] like Figure 3 As shown, the specific steps of the workflow of the denoising edge preservation submodule 101 are as follows:
[0072] The near-infrared image is convolved through a series of convolutional layers, densely connected units, and contour convolution units. The densely connected units extract near-infrared image features and suppress noise. The contour convolution units predict and enhance image edges. The two outputs are fused to produce a denoised near-infrared image I1 that removes noise and preserves edges. The denoising edge-preserving submodule 101 is configured with two 3×3 convolutions, two ReLU activation functions, three densely connected units, and two contour convolution units.
[0073] S2: Input the visible light image and the denoised near-infrared image into the color content encoding module 2 to extract the near-infrared key content features, and obtain the encoded visible light color coding map, visible light content coding map, denoised near-infrared color coding map, and denoised near-infrared content coding map;
[0074] like Figure 2 As shown, the color content encoding module 2 specifically includes a first color encoding submodule 201 , a first content encoding submodule 202 , a second color encoding submodule 203 and a second content encoding submodule 204 .
[0075] S201: Input the denoised near-infrared image I1 into the first color encoding submodule 201 and the first content encoding submodule 202 to obtain the feature representation of the near-infrared domain color and content information, i.e., the denoised near-infrared color encoding image I1. 21 And denoised near infrared content coding image I 22 ;
[0076] S202: At the same time, the visible light image I 02 The input is sent to the second color encoding submodule 203 and the second content encoding submodule 204 to obtain the characteristic representation of the visible light domain color and content information, that is, the visible light color encoding image I 23 and visible light content coding diagram I 24 .
[0077] It can be understood that in steps S201 and S202, the first color encoding submodule 201 and the second color encoding submodule 203 adopt the same structure, which is responsible for converting the denoised near-infrared image I1 and the visible light image I 02 Converted into the corresponding denoised near-infrared color coding image I 21 and visible light color coding diagram I 23 .
[0078] The first color encoding submodule 201 and the second color encoding submodule 203 each include a series of convolution units, depthwise separable convolution units, and color attention unit components, which extract rich color features from the input image, further learn the mapping relationship between the near-infrared image and the color image, pay more attention to the important color information in the image, and enhance the color expression and visual effect of the result; a series of convolution units extract the denoised near-infrared image I1 and the visible light image I1 by stacking multiple convolution layers. 02 The features are more conducive to the decoding operation; the depth-separable convolution unit operates independently on each channel of the input image, effectively extracting the denoised near-infrared image I1 and the visible light image I 02 The local features of the image are analyzed and the expressiveness of the features is improved through cross-channel information exchange. The color attention unit recognizes the important color features in the input image through learning, making the model more sensitive to important color features and improving the color recognition ability of the model.
[0079] It can be understood that in steps S201 and S202, the first content encoding submodule 202 and the second content encoding submodule 204 adopt the same structure and are responsible for converting the denoised near-infrared image I1 and the visible light image I 02 Converted into the corresponding denoised near infrared content coding image I 22 and visible light content coding diagram I 24 ;
[0080] The first content encoding submodule 202 and the second content encoding submodule 204 both include a series of convolution units, hole channel attention units, and feature network pyramid units for extracting the denoised near infrared image I1 and the visible light image I 02 The deep features and context information of the image are used to improve the color generation ability of the image and enhance the visual effect of the image; a series of convolution units are used to preliminarily extract the denoised near-infrared image I1 and the visible light image I 02 The shallow features of the feature network are obtained by the hole channel attention unit, which is used to expand the receptive field and enhance the feature representation of important channels, so that the model can capture a wider range of contextual information and pay more attention to the favorable features for the color generation task; the feature network pyramid unit connects the denoised near-infrared image I1 and the visible light image I2 through a top-down path and lateral connections. 02Features at different levels are fused to generate a content encoding graph with richer semantic information and detailed information.
[0081] like Figure 4 As shown, the workflow of the first color content encoding submodule 201 and the second color content encoding submodule 203 is as follows:
[0082] The first color content encoding submodule 201 and the second color content encoding submodule 203 use the same structure to encode the denoised near infrared image I1 and the visible light image I 02 The denoised near-infrared color coding image I is generated by inputting the denoised near-infrared color coding image I into the first color coding submodule 201 and the second content coding submodule 203 respectively. 22 , Visible light color coding diagram I 23 ;
[0083] The first color encoding submodule 201 and the second color encoding submodule 203 extract rich color features from the input image through a series of convolution units, depthwise separable convolution units and color attention unit components, further learn the mapping relationship between the near-infrared image and the color image, pay more attention to the important color information in the image, and enhance the color expression and visual effect of the result; a series of convolution units are used to extract the denoised near-infrared image I1 and the visible light image I1 by stacking multiple convolution layers. 02 The features are more conducive to the decoding operation; the depth-separable convolution unit is used to operate each channel of the input image independently, effectively extracting the denoised near-infrared image I1 and the visible light image I 02 The local features of the image are analyzed, and the expressiveness of the features is improved through cross-channel information exchange. The color attention unit is used to identify important color features in the input image through learning, making the model more sensitive to important color features and improving the model's color recognition ability. The first color encoding submodule 201 and the second color encoding submodule 203 are configured with a total of two 3×3 convolutions, four batch normalization layers, two depthwise separable convolutions, three ReLU activation functions, three color attention units, and one 1×1 convolution.
[0084] like Figure 5 FIG. 2 is a schematic diagram showing the structural principles of the first content encoding submodule 202 and the second content encoding submodule 204. The specific steps of the module workflow are as follows:
[0085] The denoised near-infrared image I1 and the visible light image I 02 The data are inputted into the first content encoding submodule 202 and the second content encoding submodule 204 respectively to generate a denoised near infrared content encoding image I. 21 , Visible Light Content Coding Figure I 24 ;
[0086] The first content encoding submodule 202 and the second content encoding submodule 204 are used to extract the denoised near-infrared image I1 and the visible light image I1 through a series of convolution units, hole channel attention units and feature network pyramid units. 02 The deep features and context information of the image are used to improve the color generation ability of the image and enhance the visual effect of the image; a series of convolution units are used to preliminarily extract the denoised near-infrared image I1 and the visible light image I 02 The shallow features of the image are obtained by the method; the hollow channel attention unit is used to expand the receptive field and enhance the feature representation of important channels, so that the model can capture a wider range of contextual information and pay more attention to the favorable features for the color generation task; the feature network pyramid unit is used to connect the denoised near-infrared image I1 and the visible light image I2 through top-down paths and lateral connections. 02 Features at different levels are fused to generate a content encoding graph with richer semantic information and detailed information; the first content encoding submodule 202 and the second content encoding submodule 204 are configured with a total of 2 3×3 convolutions, 2 batch normalization layers, 2 void channel attention units and 2 feature network pyramid units.
[0087] S3: Input the visible light color coding image and the denoised near-infrared content coding image into the cross-domain decoding module 3 for decoding to obtain a visible light near-infrared color image; input the denoised near-infrared color coding image and the visible light content coding image into the cross-domain decoding module 3 for decoding to obtain a near-infrared visible light grayscale image.
[0088] The cross-domain decoding module 3 includes a decoding submodule 301, which converts the visible light color coding image 1 into a 23 and denoised near infrared content coding image I 22 Input to the cross-domain decoding module 3 for decoding to obtain the visible light near infrared color image I 31 ; The denoised near infrared color coding image I 21 Visible light content encoding diagram I 24 Input to the cross-domain decoding module 3 for decoding to obtain the near-infrared visible light grayscale image I 32 ;
[0089] It can be understood that the decoding submodule 301 is composed of a residual unit, an AdaIn unit, a fully connected network unit, and an upsampling unit, which realizes the effective color generation of the near-infrared image and the near-infrared information supplementation of the visible light image. The residual module retains the low-frequency information of the input image by introducing a skip connection, and learns the mapping relationship of the high-frequency information at the same time, decoding it into a more natural visible light near-infrared color image I 31 and near-infrared visible light grayscale image I 32 ; The AdaIn unit introduces the information of each target domain by adjusting the variance and mean of the input feature map, so that the generated visible near-infrared color image I31 It has the same color style and visual effect as the visible light image domain, making the generated near-infrared visible light grayscale image I 32 It has the color style and visual effects of the near-infrared image domain; the fully connected network unit can achieve fine control of color and content information by learning the mapping relationship between the two domains; the upsampling unit is used to gradually restore the resolution of the feature map to the same size as the domain input image.
[0090] like Figure 6 The figure shows the structural principle of the cross-domain decoding module 3. The specific steps of the workflow of this module are as follows:
[0091] The color coding image and the content coding image are input into the decoding submodule 301 respectively to generate a visible light near infrared color image I 31 and near-infrared visible light grayscale image I 32 ;
[0092] The color coding map and content coding map achieve effective color generation of near-infrared images and near-infrared information supplement of visible light images through residual units, AdaIn units, fully connected network units and upsampling units; the residual unit is used to retain the low-frequency information of the input image by introducing jump connections, while learning the mapping relationship of high-frequency information, and decoding it into a more natural visible light near-infrared color image I 31 and near-infrared visible light grayscale image I 32 ; The AdaIn unit introduces the information of each target domain by adjusting the variance and mean of the input feature map, so that the generated visible near-infrared color image I 31 It has the same color style and visual effect as the visible light image domain, making the generated near-infrared visible light grayscale image I 32 It has the color style and visual effects of the near-infrared image domain; the fully connected network unit can achieve fine control of color and content information by learning the mapping relationship between the two domains; the upsampling unit is used to gradually restore the resolution of the feature map to the same size as the input image; the decoding submodule 301 is configured with a fully connected network unit, an AdaIn unit, three residual units and an upsampling unit;
[0093] The loss function of the cross-domain encoding and decoding near-infrared image color generation method of this application is set for near-infrared image color generation to obtain near-infrared images with richer color information and visible light images with richer near-infrared information. The losses include the following:
[0094] Content consistency loss function:
[0095] L con =||E con (I1)-E con (I 31-I 02 )||1+||E con (I 02 )-E con (I 32 -I1)||1 (1)
[0096] Among them, E con is the first content encoding submodule or the second content encoding submodule; I1 is the denoised near-infrared image; I 31 is the visible light near infrared image; I 02 is a visible light image; I 32 is the near-infrared visible light grayscale image; ||·||1 is the L1 norm.
[0097] Color consistency loss function:
[0098] L col =||E col (I1)-E col (I 31 -I 02 )||1+||E col (I 02 )-E col (I 32 -I1)||1 (2)
[0099] Among them, E col is the first color coding submodule or the second color coding submodule; I1 is the denoised near infrared image; I 31 is the visible light near infrared image; I 02 is a visible light image; I 32 is a near-infrared visible light grayscale image.
[0100] Reconstruction loss function:
[0101] L recon =||I1-I 32 ||1+||I 02 -I 31 ||1 (3)
[0102] Among them, I1 is the denoised near-infrared image; I 32 is the near-infrared visible light grayscale image; I 02 is a visible light image; I 31 It is a visible light near-infrared image.
[0103] Total loss function:
[0104] L total =λ1L con +λ2L col +λ3L recon (4)
[0105] Among them, λ1, λ2, and λ3 are hyperparameters in the network, λ1=1, λ2=1, and λ3=10.
[0106] The content consistency loss function is used to quantify the consistency between the result image and the input image, ensuring that key information such as the content and structure of the input image is retained as much as possible, making the generated result more natural and realistic.
[0107] The color consistency loss function calculates the difference between the generated image and the reference image in the color space and minimizes this difference, thereby maintaining the color distribution of the result image and improving the overall quality and visual effect of the image.
[0108] The reconstruction loss function is used to quantify the difference between the data of the input image and the result image, guiding the color generation model to continuously iteratively optimize and improve the reconstruction ability of the model.
[0109] The total loss function assigns weights to each loss, flexibly adjusts the importance of different types of losses in the training process, and balances each loss function to enable the model to achieve better performance on multiple objectives.
[0110] This application also has the following advantages:
[0111] Food safety testing: Visible-light near-infrared color imaging can analyze the color, texture, and other characteristics of the surface of fruits and vegetables to determine their maturity, freshness, and quality, providing a scientific basis for the grading, storage, and transportation of agricultural products.
[0112] Environmental monitoring and protection: Visible-near infrared images can reveal the distribution of pollutants in water bodies, such as algae blooms and oil spills, and visually demonstrate the extent and scope of pollution through color differences.
[0113] Biomedical research: In biomedical imaging, near-infrared light has good penetration into certain biological tissues, and the colored images help observe and analyze the structure and functional changes of tissues.
[0114] Industrial Automation and Quality Control: In the field of industrial automation, near-infrared color images can be used to monitor the appearance quality of products on the production line in real time, such as color consistency and defect detection, to improve production efficiency and product quality.
[0115] In order to further illustrate the relevant effects of the technical solution of this application, the following experiments are conducted: This application selects the public dataset NIR_VIS as the training and test sets, and uses the denoising edge preservation module to assist in near-infrared image denoising and edge preservation, enhance the image detail expression, and better represent the texture and other information of the near-infrared image; the color coding module further extracts the features of the input image, and learns the denoised near-infrared image I1 and the visible light image I 02The mapping relationship between them pays more attention to the important color information between the two, making the model more sensitive to important color features and improving color recognition ability; the content encoding module is used to extract the denoised near-infrared image I1 and the visible light image I 02 The deep features and contextual information of the image are combined to improve the color generation capability of the image and enhance the visual effect of the image; the cross-domain decoding module realizes effective color generation of near-infrared images and supplements the near-infrared information of visible light images; to verify the performance of the present invention, the color generation effect of the present invention is compared with the enhancement algorithm of deep learning in recent years, using PSNR, SSIM, and LPIPS as evaluation indicators, respectively, and using 20 groups of near-infrared images as tests. The comparison results are shown in Table 1. The test results show that the cross-domain encoding and decoding near-infrared image color generation method provided by this application performs well in PSNR, SSIM, and LPIPS.
[0116] Table 1 Algorithm comparison results
[0117] PSNR↑ SSIM↑ LPIPS↓ CycleGAN 19.04 0.71 0.439 CDGAN 25.96 0.88 0.441 I2VGAN 28.59 0.91 0.364 MCL 28.40 0.82 0.277 This application method 30.31 0.93 0.203
[0118] The present application may also provide a computer device comprising: at least one processor, memory, at least one network interface, and a user interface. The various components in the device are coupled together via a bus system. It will be understood that the bus system is used to enable communication between these components. In addition to a data bus, the bus system also includes a power bus, a control bus, and a status signal bus.
[0119] The user interface may include a display, a keyboard, or a pointing device, such as a mouse, a trackball, a touchpad, or a touch screen.
[0120] It is understood that the memory in the embodiments disclosed in the present application may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. Among them, the non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM), which is used as an external cache. By way of example and not limitation, many forms of RAM are available, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate synchronous DRAM (DDRSDRAM), enhanced synchronous DRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DRRAM). The memories described herein are intended to include, but are not limited to, these and any other suitable types of memory.
[0121] In some embodiments, the memory stores the following elements, executable modules or data structures, or a subset or an extension thereof: an operating system and applications.
[0122] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, which are used to implement various basic services and handle hardware-based tasks. Application programs include various application programs, such as media players and browsers, which are used to implement various application services. The program that implements the method of the embodiment of the present disclosure can be included in the application program.
[0123] In the above embodiment, the processor may also call a program or instruction stored in the memory, specifically, a program or instruction stored in the application program, to:
[0124] Perform the steps of the above method.
[0125] The above method can be applied to or implemented by a processor. The processor may be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by hardware integrated logic circuits in the processor or by software instructions. The above processor may be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The above-disclosed methods, steps, and logic block diagrams can be implemented or executed. The general-purpose processor may be a microprocessor or any conventional processor. The steps of the above-disclosed method can be directly implemented and executed by a hardware decoding processor, or by a combination of hardware and software modules in the decoding processor. The software module can be located in a storage medium well-known in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, etc. The storage medium is located in the memory, and the processor reads the information in the memory and, in conjunction with its hardware, completes the steps of the above method.
[0126] It is understood that the embodiments described herein may be implemented using hardware, software, firmware, middleware, microcode, or a combination thereof. For hardware implementation, the processing unit may be implemented in one or more application-specific integrated circuits (ASICs), digital signal processors (DSPs), digital signal processing devices (DSPDs), programmable logic devices (PLDs), field-programmable gate arrays (FPGAs), general-purpose processors, controllers, microcontrollers, microprocessors, or other electronic units or combinations thereof for performing the functions described herein.
[0127] For software implementation, the technology of the present application can be implemented by executing the functional modules (e.g., procedures, functions, etc.) of the present application. The software code can be stored in a memory and executed by a processor. The memory can be implemented in the processor or external to the processor.
[0128] The present application may also provide a non-volatile storage medium for storing a computer program. When the computer program is executed by a processor, each step in the above method embodiment can be implemented.
[0129] Finally, it should be noted that the above embodiments are intended only to illustrate the technical solutions of this application and are not intended to limit the scope of the present invention. Although this application has been described in detail with reference to the embodiments, it should be understood by those skilled in the art that modifications or equivalent substitutions to the technical solutions of this application do not depart from the spirit and scope of the technical solutions of this application and should be encompassed by the claims of this application.
Claims
1. A cross-domain encoding and decoding near-infrared image color generation system, characterized in that: The system comprises: A preprocessing module is used to preprocess the selected near-infrared image to obtain a denoised near-infrared image; a color content coding module, configured to generate a coded visible light color coding map and a coded visible light content coding map from the visible light image, and generate a denoised near infrared color coding map and a denoised near infrared content coding map from the denoised near infrared image; and A cross-domain decoding module is used to decode the visible light color coding image and the denoised near-infrared content coding image to obtain a visible light near-infrared color image; and decode the denoised near-infrared color coding image and the visible light content coding image to obtain a near-infrared visible light grayscale image; The cross-domain decoding module includes a fully connected network unit, an AdaIn unit, several residual units and an upsampling unit; The fully connected network unit achieves fine control of color and content information by learning the mapping relationship between the two domains; The AdaIn unit introduces the information of each target domain by adjusting the variance and mean of the input feature map, so that the generated visible near-infrared color image I 31 With the color style and visual effects of the visible light image domain, the generated near-infrared visible light grayscale image I 32 Color style and visual effects with near-infrared image domain; The residual unit retains the low-frequency information of the input image by introducing a jump connection, and learns the mapping relationship of the high-frequency information at the same time, decoding the content coding map into a visible light near-infrared color image I 31 and near-infrared visible light grayscale image I 32 ; The upsampling unit is used to restore the resolution of the feature map to the same size as the input image.
2. The cross-domain encoding and decoding near-infrared image color generation system according to claim 1, characterized in that: The pre-processing module includes a denoising edge preservation sub-module; The denoising edge preservation submodule is used to 01 The calculation is the denoised near-infrared image I1 with noise removed and edge preserved; The denoising edge-preserving submodule includes several convolution layers, dense connection units and contour convolution units connected in series to perform convolution operations; wherein, The densely connected unit is used to perform feature extraction and noise suppression on the near-infrared image; The contour convolution unit is used to enhance the edge of the image.
3. The cross-domain encoding and decoding near-infrared image color generation system according to claim 1, characterized in that: The color content encoding module includes a first color encoding submodule, a second color encoding submodule, a first content encoding submodule and a second content encoding submodule; wherein, The first color coding submodule is used to convert the denoised near-infrared image I1 into the corresponding denoised near-infrared color coding image I 21 ; The second color encoding submodule is used to convert the visible light image I 02 Converted into the corresponding denoised visible light color coding image I 23 ; The first content encoding submodule is used to convert the denoised near-infrared image I1 into the corresponding denoised near-infrared content encoding image I 22 ; The second content encoding submodule is used to convert the visible light image I 02 Converted into the corresponding denoised visible light content encoding image I 24 .
4. The cross-domain encoding and decoding near-infrared image color generation system according to claim 3, characterized in that: The first color coding submodule and the second color coding submodule have the same structure, both including a number of convolution units, depth-separable convolution units and color attention units; wherein, The convolution units extract the denoised near-infrared image I1 and the visible light image I1 by stacking multiple convolution layers. 02 characteristics; The depthwise separable convolution unit operates independently on each channel of the input image to extract the denoised near-infrared image I1 and the visible light image I 02 local features of The color attention unit learns to identify important color features in the input image.
5. The cross-domain encoding and decoding near-infrared image color generation system according to claim 3, characterized in that: The first content encoding submodule and the second content encoding submodule have the same structure; both include several convolution units, hole channel attention units and feature network pyramid units; wherein, The convolution unit is used to extract the denoised near-infrared image I1 and the visible light image I 02 shallow features of The hole channel attention unit is used to expand the receptive field and enhance the feature representation of important channels, so that the model can capture a wider range of contextual information and pay more attention to the features that are beneficial to the color generation task; The feature network pyramid unit connects the denoised near-infrared image I1 and the visible light image I1 through a top-down path and lateral connections. 02 Features at different levels are fused to generate a content encoding graph with semantic and detail information.
6. The cross-domain encoding and decoding near-infrared image color generation system according to claim 3, characterized in that: The loss function L of the system total for: L total =λ1L con +λ2L col +λ3L recon Among them, λ1, λ2, λ3 are hyperparameters; L con is the content consistency loss function: L con =||E con (I1)-E con (I 31 -I 02 )||1+||E con (I 02 )-E con (I 32 -I1)||1 Among them, E con is the first content encoding submodule or the second content encoding submodule; I1 is the denoised near-infrared image; I 31 is a visible light near infrared color image; I 02 is a visible light image; I 32 is the near-infrared visible light grayscale image; ||·||1 is the L1 norm; L col is the color consistency loss function: L col =||E col (I1)-E col (I 31 -I 02 )||1+||E col (I 02 )-E col (I 32 -I1)||1 Among them, E col is a first color coding submodule or a second color coding submodule; L recon To reconstruct the loss function: L recon =||I1-I 32 ||1+||I 02 -I 31 ||1。 7. A method for generating color from a cross-domain encoded and decoded near-infrared image, implemented based on the system of any one of claims 1 to 6, comprising: Step S1: Processing the selected near-infrared image through a pre-processing module to obtain a denoised near-infrared image; Step S2: Inputting the visible light image and the denoised near-infrared image into a color content encoding module to extract near-infrared key content features, thereby obtaining an encoded visible light color coding map, a visible light content coding map, a denoised near-infrared color coding map, and a denoised near-infrared content coding map; Step S3: Input the visible light color coding image and the denoised near-infrared content coding image into the cross-domain decoding module for decoding to obtain a visible light near-infrared color image; input the denoised near-infrared color coding image and the visible light content coding image into the cross-domain decoding module for decoding to obtain a near-infrared visible light grayscale image.
8. The cross-domain encoding and decoding near-infrared image color generation method according to claim 7, characterized in that: The step S2 comprises: The denoised near-infrared image I1 is input into the first color coding submodule and the first content coding submodule to obtain the feature representation of the near-infrared domain color and content information, i.e., the denoised near-infrared color coding image I 21 And denoised near infrared content coding image I 22 ; At the same time, the visible light image I 02 Input the second color coding submodule and the second content coding submodule to obtain the feature representation of visible light domain color and content information, that is, visible light color coding image I 23 and visible light content coding diagram I 24 .
Citation Information
Patent Citations
Image processing method, computer equipment and computer readable storage medium
CN116980724A
Method and system for near infrared-visible light image conversion
CN117314734A