A method for converting visible light to infrared images
Through a supervised infrared image conversion network, combined with multi-level feature extraction and rotation consistency loss, the problem of inaccurate mapping of unsupervised methods and weak generalization performance of supervised methods is solved, and the high-quality generation and generalization ability of infrared images is achieved.
Patent Information
- Application Number
- CN202211377615.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-11-04
- Publication Date
- 2025-07-18
- Estimated Expiration
- 2042-11-04
AI Technical Summary
The existing unsupervised image conversion methods are difficult to achieve accurate mapping from the source domain to the target domain. The supervised method has weak generalization performance for unfamiliar scenes, and some conversion methods fail to explicitly distinguish target individuals, resulting in target features coupling and poor generation effect.
The supervised infrared image conversion network is adopted to gradually generate infrared images through multi-level feature extraction modules, material feature separation modules, thermal data feature separation modules and thermal physics simulation modules, and the rotation consistency loss constraint feature separation module is used to reduce manual labeling costs and improve the generation effect.
The discovery ability of the same type of target in strange scenarios and the accurate mapping of the source domain to the target domain is achieved, which reduces the cost of manual labeling, and the generated infrared images are closer to the real images, improving the generation effect and generalization capabilities.
Smart Images

Figure CN115713456B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of image processing, and more specifically, relates to a method for converting visible light into infrared images. Background Art
[0002] With the development of computer vision technology, target detection technology based on infrared images is increasingly widely used in military and civilian fields. In the military field, infrared image guidance systems can greatly improve the accuracy and stability of guidance. At the same time, detection and recognition based on infrared images is also one of the important means of battlefield perception and exploration of important military facilities. In the civilian field, infrared remote sensing, drone navigation, unmanned driving and infrared security all use infrared images as one of the important information sources. Detection and recognition technology based on infrared images can greatly improve the stability and applicability of these systems. Therefore, the development of infrared image detection and recognition technology has important research significance, among which target detection methods based on deep learning are the main research direction at present. However, target detection methods based on deep learning require a large amount of training data. Too little data will affect the generalization performance of the network. Therefore, using deep learning technology to generate infrared images has great research value.
[0003] Thanks to the development of GAN, research on image-to-image conversion has also made great progress. Image-to-image conversion is similar to conditional generation. Given a source domain image, its style is transferred to the target domain while keeping the content of the image unchanged, such as selfie2anime, horse2zebraz and other methods. Pix2pix was first proposed to solve the problem of image-to-image conversion. The limitation of this type of method is that it requires a pixel-level paired dataset. However, obtaining a paired dataset requires a lot of effort, and some paired datasets are not available, such as medical images. To solve this problem, methods based on cycle consistency and content similarity constraints have been proposed, among which the most representative method is cycleGAN. In addition, there are some methods for multimodal generation, which convert source domain images to multiple target domains, such as starGAN. Although unsupervised image-to-image conversion methods are more widely used, a large number of studies have shown that supervised methods can learn more accurate mappings between source and target domains. Unsupervised methods may have perturbations during training, resulting in inaccurate mappings. Therefore, supervised image-to-image conversion methods still have good application value because they can directly calculate pixel-level differences. However, for supervised methods, networks trained using paired datasets have poor generalization performance for unfamiliar scenes.
[0004] In summary, existing unsupervised image conversion methods are difficult to achieve accurate mapping from the source domain to the target domain. Supervised methods have weak generalization performance for unfamiliar scenarios. At the same time, some conversion methods perform feature extraction on the entire image to image generation without further explicitly distinguishing target individuals, resulting in feature coupling between targets and poor generation effects. Methods that explicitly distinguish target individuals divide the entire process into multiple stages and involve humans in the intermediate process, making the processing process cumbersome. Summary of the Invention
[0005] In view of the above defects or improvement requirements of the prior art, the present invention provides a visible light to infrared image conversion method, aiming to solve the technical problem of difficult to balance the accurate mapping from the source domain to the target domain and the generalization performance of unfamiliar scenarios in the prior art.
[0006] To achieve the above object, the present invention provides a visible light to infrared image conversion method, including:
[0007] S1. Obtain paired visible light infrared images as a training set;
[0008] S2. Construct an infrared image conversion network; the infrared image conversion network includes: a plurality of feature extraction modules connected hierarchically from bottom to top, a plurality of parallel material feature separation modules corresponding to the levels of the feature extraction modules, a plurality of feature coupling modules connected hierarchically from top to bottom corresponding to the levels of the feature extraction modules, as well as a thermal data feature separation module and a thermophysical simulation module;
[0009] Among them, the hierarchically connected feature extraction modules are used to perform feature extraction on visible light images at different scales; the parallel material feature separation modules are used to respectively divide the material categories of feature maps at different scales to obtain the material features corresponding to the feature maps at different scales; the thermal data feature separation module is used to obtain global thermal data features from the feature map at the largest scale; the cascaded feature coupling modules are used to select the most fully expressed material features from feature maps at different scales, fuse the selected material features with the corresponding thermal data features, and superimpose them with the features of the previous level to gradually generate a heat map; the thermophysical simulation module is used to convert the heat map into an infrared image;
[0010] S3. Perform supervised iterative training on the infrared image conversion network using the training set to obtain a trained infrared conversion model;
[0011] S4. Input the visible light image to be converted into the trained infrared conversion model to obtain the corresponding infrared image.
[0012] Further, during the training process of the infrared image conversion network, the following rotation consistency loss is used to constrain the feature separation module and the thermal data feature separation module:
[0013]
[0014] and represent the thermal data features before and after the image transformation respectively, and represent the material features of the i-th layer before and after the image transformation respectively, and T(·) represents the inverse transformation of the image transformation.
[0015] Furthermore, the material feature separation module includes a convolutional layer, a pooling layer, and a pixel-level classifier connected in sequence.
[0016] Furthermore, the thermal data feature separation module includes a convolutional layer, a global average pooling layer, and a feature mapping layer connected in sequence.
[0017] Furthermore, the feature coupling module includes a material selection unit, a feature mapping unit, an adaptive instance normalization unit, a convolutional layer, and an upsampling layer;
[0018] Among them, the material selection unit selects the features of a specific channel through the following operation;
[0019]
[0020] where is the channel weighting coefficient, represents the material features of the i-th scale, einsum represents Einstein summation, and based on parameters, the channels of are re-weighted. When the parameter is closer to 1, it means that the features of the channel are expressed, otherwise they are suppressed;
[0021] The feature mapping unit further performs feature mapping on the input global material thermal data features to obtain the thermal data features corresponding to the selected material features;
[0022] The adaptive instance normalization unit first performs adaptive feature fusion on the output of the current-level material selection unit and the output of the current-level feature mapping unit; then superimposes the fused feature map with the previous-level feature map and performs instance normalization processing.
[0023] Furthermore, the feature mapping layer and the feature mapping unit are implemented using an EqualizedLinear fully connected layer.
[0024] Furthermore, the thermophysical simulation module is composed of three residual blocks connected in series.
[0025] Generally speaking, compared with the prior art by the above technical solutions conceived by the present invention, the following beneficial effects can be achieved.
[0026] (1) Based on supervised training through a paired dataset, the present invention extracts material features and thermal data features in an image, and gradually embeds the extracted material features and thermal data features into a heat map, progressively generating the heat map. The entire process operates at the feature level for the target individual in the image. Compared with the traditional method of generating the entire image, it can reduce the coupling relationship between target features, achieve a better generation effect, ensure the discovery ability of the same type of target in an unfamiliar scene, and achieve the purpose of balancing the mapping accuracy from the source domain to the target domain and the generalization ability in an unfamiliar scene.
[0027] (2) The present invention uses an unsupervised method to impose certain constraints on the material feature separation module and the thermal data feature separation module, reducing the cost of manual annotation. At the same time, it alleviates the problem that thermal data cannot be directly obtained to a certain extent, achieving the purpose of balancing the explicit distinction of target individuals and the complexity of image processing, being more suitable for the actual application scenario, and having a high practical value. Description of the Drawings
[0028] Figure 1 Schematic diagram of a visible light to infrared image conversion method provided by an embodiment of the present invention;
[0029] Figure 2 Detailed structure diagram of each module in a visible light to infrared image conversion method provided by an embodiment of the present invention;
[0030] Figure 3 Separation result diagram of material features of the input infrared image at different feature layers by the material feature separation module provided by an embodiment of the present invention;
[0031] Figure 4 Comparison diagram of conversion results between the visible light to infrared image conversion algorithm provided by an embodiment of the present invention and other classical algorithms; where (a) is the result diagram on the KAIST dataset, and (b) is the result diagram on the M3FD dataset;
[0032] Figure 5 Experimental result of the visible light image conversion in an unfamiliar scene provided by an embodiment of the present invention. Detailed Embodiments
[0033] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0034] In view of the fact that existing unsupervised image conversion methods are difficult to achieve accurate mapping from the source domain to the target domain, supervised methods have weak generalization performance for unfamiliar scenarios, and at the same time, some conversion methods perform feature extraction on the entire image to image generation without further explicitly distinguishing target individuals, resulting in feature coupling between targets and poor generation effects. Methods for explicitly distinguishing target individuals divide the entire process into multiple stages and involve human participation in the intermediate process, making the processing process cumbersome. The present invention provides a visible light to infrared image conversion method based on feature separation-coupling. Its overall idea is as follows: Extract multi-scale features of the input visible light image through a feature extraction module so that the features of different-sized targets in the scene are fully expressed. Then, extract the material features in the feature maps of different scales through a material feature separation module At the same time, model the relationship between the pixel values of the input image and the material thermal data through a thermal data feature separation module and based on encode to obtain material thermal data features Subsequently, the material features and thermal data features at different scales are jointly input into multiple cascaded feature coupling modules. The selective expression of the material features is carried out through the built-in material selection module and feature mapping module in the module to further improve the network's ability to express features. Then, the material features and thermal data features are fused through the built-in adaptive instance normalization module and superimposed with the features of the previous level to gradually generate a heat map. Finally, the generation of the infrared image is completed through a thermophysical simulation module. At the same time, during the whole process, an unsupervised rotation consistency constraint is adopted to enhance the effectiveness of the material feature separation module and the thermal data feature separation module.
[0035] Compared with general methods, this method pays more attention to the overall object, decouples the material features and thermal data features during the feature extraction stage and embeds them into the latent space, couples the material and thermal data in a progressive generation manner during the decoding stage to obtain a heat map, and finally obtains an infrared image through the thermophysical simulation process. By this means, the integrity of the object in the generated image can be improved, and at the same time, the generalization performance for scene changes can be improved.
[0036] A visible light to infrared image conversion method based on feature separation-coupling provided by the present invention can be applied to the conversion from visible light images to infrared images, and can also be applied to the conversion from grayscale images to infrared images. In the embodiments of the present invention, taking the conversion from visible light images to infrared images as an example, it is described in detail as follows:
[0037] To achieve the above object, the present invention provides a visible light to infrared image conversion method, including:
[0038] S1. Obtain paired visible light infrared images as the training set;
[0039] S2. Build an image conversion network structure based on a deep neural network; the image conversion network structure is as follows Figure 1 As shown, including:
[0040] Material feature separation module, used to separate feature maps from different scales Different material features are segmented, and the material categories of the feature points are divided by pixel-level classifier to obtain material features
[0041] Thermal data feature separation module, which is used to model the relationship between the input visible light image pixel value and the global material thermal data features. This module further mines the input feature map The features are encoded and the global material thermal data features are obtained through the feature mapping layer.
[0042] The feature coupling module generates heat maps progressively from top to bottom in a cascade manner, fusing material features of different scales at different levels. and material thermal data characteristics The material selection module is used to select the materials that need to participate in the fusion calculation at the current scale, and the material thermal data features are further analyzed through the feature mapping module. Mapping is performed to obtain the thermal data corresponding to the selected material, and then both are input into the adaptive instance normalization module to complete the fusion of material features and thermal data features corresponding to the material. Finally, the fused features are superimposed with the features of the previous level and a heat map is progressively generated through the upsampling layer.
[0043] The thermal physics simulation module is used to fit the conversion process from thermal map to infrared image. It is composed of multiple residual blocks superimposed to finally generate the corresponding infrared image.
[0044] refer to Figure 2 In this embodiment, the material feature separation module is composed of a convolution layer, a pooling layer and a pixel-level classifier. The convolution layer is implemented by a convolution operation with a step size of 1 through a convolution kernel of size 3×3. The size of the input and output feature maps is kept unchanged before and after the convolution layer, and the number of output channels is the number of categories of the preset material; the pooling layer uses a maximum pooling with a kernel size of 2×2 to perform a 2-fold downsampling operation on the input feature map; the pixel-level classifier is intended to separate different materials from the multi-scale feature map. Specifically, for the input feature maps of different scales, First, the input features are processed through the convolution layer Extract features to get Where B is the number of images input to the network during a test, and C is the feature The number of channels, H and W are the size of the feature map at the current scale, and D is The number of channels is also equal to the number of material categories. Then, for each feature point x in the H×W feature map w,h (w = 1…W, h = 1…H), it is classified to determine which material category D_i (i = 1…D) it belongs to, thereby obtaining the feature distribution of the material The classifier is implemented using the multinomial classifier normalization exponential function, and its specific expression is:
[0045]
[0046] Then the calculation process of the material feature separation module can be expressed as:
[0047]
[0048]
[0049] In the formula represents a conventional convolution kernel, n×n represents the size of the convolution kernel, and Maxpool represents the max pooling operation.
[0050] To further demonstrate the effect of the material feature separation module, the features of the visible light image extracted based on the material feature separation module were visualized as Figure 3 shown. In Figure 3 , the first column represents the input original visible light image, the second column represents the visualization result of the features extracted by the i-th material feature separation module and the third column represents the visualization result of the features extracted by the j-th material feature separation module . It can be intuitively seen from the visualization results that the material feature separation modules at different levels effectively divide the materials in the original image, and at the same time, different materials are fully expressed at the feature levels suitable for their expression.
[0051] In this embodiment, the thermal data feature separation module consists of a convolutional layer, a global average pooling layer, and a feature mapping layer. Among them, the convolutional layer further extracts features from the input feature map to obtain intermediate layer features where B is the number of images input to the network during a single test process, C is the number of channels of the features , H and W are the sizes of the feature map at the current scale, and D is the number of channels of , which is also the number of material categories; as an alternative embodiment, the convolutional layer is implemented by performing a convolution operation with a 3×3 convolution kernel and a stride of 1; the global average pooling layer performs global average pooling on the intermediate layer features channel by channel to obtain the dimension-reduced feature Fbrightness ∈ {B, D}; The feature mapping layer, for F brightness Performs feature mapping to obtain material thermal data features Completes the mapping from the visible light image brightness value to the global material thermal data features. Optionally, the feature mapping layer is implemented using a common EqualizedLinear fully connected layer.
[0052] The calculation process of the thermal data feature separation module is as follows:
[0053]
[0054]
[0055] Among them, in the formula, kernel 3×3 Represents a conventional convolution kernel with a kernel size of n×n, GAP represents global average pooling, and EqualizedLinear is a fully connected layer.
[0056] In this embodiment, the feature coupling module consists of a material selection unit, a feature mapping unit, an adaptive instance normalization unit, a convolutional layer, and an upsampling layer.
[0057] Among them, the material selection unit is used to selectively express the input features For different materials, the target sizes are different, so the feature expression capabilities under different scale feature maps are different. By participating in the training through the material selection unit, different materials participate in feature fusion at the scale where the feature expression is the most sufficient and are finally superimposed on the main feature map to achieve better generation.
[0058] Different from the conventional convolution operation, the material selection unit essentially selects the features of specific channels for expression. In this embodiment, the method adopted is to learn a learnable parameter (i represents the i-th scale), And the elements in the parameter are all between 0 and 1. The selected features are obtained through the following operations:
[0059]
[0060] Among them Is the channel weighting coefficient, Represents the material feature of the i-th scale, einsum represents Einstein summation, and based on The parameters of The channels of are reweighted. When the parameter is closer to 1, it means that the features of that channel are expressed, otherwise they are suppressed.
[0061] The feature mapping unit, for the input global material thermal data features Perform further feature mapping to obtain the thermal data features corresponding to the selected material features Optionally, the feature mapping unit is implemented using a single layer of EqualizedLinear fully connected layer.
[0062] The adaptive instance normalization module processes the feature map of the previous level of the input Output of the current level material selection module And the output of the current level feature mapping module Perform adaptive feature fusion.
[0063] First, for And Perform feature fusion,
[0064]
[0065]
[0066] where μ(·) represents calculating the mean value of features sample-by-sample and channel-by-channel, and σ(·) represents calculating the variance of features sample-by-sample and channel-by-channel.
[0067] Subsequently, the fused feature map is superimposed with the feature map of the previous level And perform instance normalization processing,
[0068]
[0069] The convolutional layer is implemented by performing a convolution operation with a stride of 1 using a 3×3 convolutional kernel. The size and number of channels of the input and output feature maps remain unchanged before and after the convolutional layer; the upsampling layer is implemented by bilinear interpolation, and the input image is upsampled by a factor of 2 to gradually increase the size of the feature map until it reaches the size of the input image.
[0070] The calculation process of the feature coupling module is as follows:
[0071]
[0072]
[0073]
[0074]
[0075]
[0076] where, Upsample represents upsampling, InstanceNorm represents instance normalization operation, represents a conventional convolutional kernel with a kernel size of n×n.
[0077] In this embodiment, the thermophysical simulation module is composed of three residual blocks connected in series, and is used to fit the conversion process from the thermal map to the infrared image. The residual network effectively avoids the problem of network degradation through skip connections. The thermophysical process can be fitted by connecting multiple residual blocks in series. The calculation formula for a single residual block is:
[0078] x l+1 = x l + F(x l , W l )
[0079] where x l represents the input of the residual block, x l+1 represents the output of the residual block, F(·) represents the residual branch, and W l represents the weight parameter of the residual branch. Then the calculation process of the thermophysical simulation module can be expressed as:
[0080]
[0081] where, represents the output of the last layer feature coupling module, and ResiBlock represents the residual block.
[0082] S3. Use the training set to perform supervised iterative training on the infrared image conversion network to obtain a trained infrared conversion model;
[0083] In this embodiment, an unsupervised method is used to impose a certain degree of constraint on the material feature separation module and the thermal data feature separation module. Specifically, a single training process is split into two stages. In the first stage, for a batch of input visible light data, a complete process of feature extraction to infrared image conversion is performed, and the corresponding thermal data features and material features at each level are recorded. In the second stage, the visible light images in the first batch are randomly rotated by a certain angle, and then a complete process of feature extraction to infrared image conversion is also performed, and the corresponding thermal data features and material features at each level are recorded. During the process of randomly rotating the angle, the pose of the target changes, and the material represents the specific details of the target. Therefore, the distribution of the material features changes and shows the same angle change compared to the original material features ; randomly rotating the angle will not cause the loss of the target. The thermal data features are related to the presence of the target and independent of the pose of the target. Therefore, compared to the original thermal data features Exhibits invariant characteristics, and through this implicit constraint, the effectiveness of the extracted material features and thermal data features can be ensured.
[0084] In this embodiment, a paired visible light-infrared data set is used to train the entire image conversion model, and the loss function of the entire training process is:
[0085] L total = L GAN + L pix + L consitence
[0086] L GAN = E[D w (I infrared )] - E[D w (G θ (I visible ))]
[0087] L pix = E[||I infrared - G θ (I visible )||1]
[0088]
[0089] Among them, L GAN is the adversarial generation adversarial loss, used to constrain the generated infrared image, I infrared and I visible represent the infrared and visible light images respectively, D w is the discriminator, G θ is the generator; L p is the pixel loss, used to achieve content consistency constraints; L consitence是 is the rotation consistency loss, used to ensure the effectiveness of the separation of material features and thermal data features, and represent the thermal data features before and after image transformation respectively, and respectively represent the material features of the i-th layer before and after the image transformation, and T(·) represents the inverse transformation of the image transformation. The loss function includes the generative adversarial loss, the content consistency loss, and the rotation consistency loss. Among them, the content consistency loss is calculated based on the generated infrared image and the corresponding original infrared image according to the above formula, and the rotation consistency loss is calculated based on the intermediate extracted features in the two stages during a single training according to the above formula. Different weights are adopted for each part of the loss to jointly constrain the optimization of the network parameters. First, the network completes the conversion tasks of the two stages, then calculates the corresponding losses based on the intermediate layer features and the generated results and backpropagates to adjust the relevant model parameters of the generator. After that, the parameters of the generator are fixed, and the corresponding losses are calculated based on the generated results and the corresponding original infrared images and backpropagated to adjust the relevant model parameters of the discriminator.
[0090] Optionally, in this embodiment, the number of training epochs is 200, the learning rate is initialized to 0.0001, the Adam optimizer of the adaptive learning rate method is selected, and the input images are processed using conventional data augmentation methods, including operations such as random cropping, flipping, scaling, and normalization, and finally a unified input size of (256, 256) is obtained.
[0091] S4. Input the visible light image to be converted into the trained infrared conversion model to obtain the corresponding infrared image, and realize the conversion from the visible light image to the infrared image end-to-end.
[0092] To verify the effect of the present invention, comparative experiments with other classical algorithms were carried out on two public datasets. The two datasets are mainly street scene data, including objects such as zebra crossings, pedestrians, vehicles, street lights, houses, and trees. The first dataset is from the KAIST dataset, which contains 51 different scenarios with a total of 2800 pairs of paired visible light-infrared images. Among them, 41 scenarios with a total of 2686 pairs of images are used as training samples, and 10 scenarios with a total of 114 pairs of images are used as test samples; the second dataset is from the M3FD dataset, which contains 149 different scenarios with a total of 4200 pairs of paired visible light-infrared images. Among them, 126 scenarios with a total of 3860 pairs of images are used as training samples, and 23 scenarios with a total of 340 pairs of images are used as test samples. In this paper, the structural similarity SSIM and the peak signal-to-noise ratio PSNR are used as the criteria for measuring performance. The calculation method of SSIM is as follows:
[0093]
[0094] where x and y respectively represent the real infrared image and the corresponding generated infrared image, μ x and μ y respectively represent the means of x and y, σ x and σy respectively represent the standard deviations of x and y, σ xy represents the covariance of x and y, and c1 and c2 are constants respectively to avoid a zero denominator.
[0095] The calculation method of PSNR is as follows:
[0096]
[0097]
[0098] Among them, x and y respectively represent the real infrared image and the corresponding generated infrared image, m and n represent the width and height of the image, and x(i,j) and y(i,j) respectively represent the pixel values of the corresponding images at the position (i,j). Table 1 shows the comparison results between the present invention and other classical methods, Table 1
[0099]
[0100]
[0101] As can be seen from Table 1, the SSIM indexes of the generated images and the corresponding real infrared images of the method proposed by the present invention on the two data sets are better than those of other classical methods, and there is a large improvement on the M3FD data set; for the PSNR index, it can be observed that the higher the SSIM index of the generated image, the slightly higher the corresponding PSNR index will be, but the method proposed by the present invention can reduce the increase of the PSNR index on the basis of significantly improving the SSIM index, that is, improve the quality of the generated image.
[0102] Furthermore, in order to visually display the generation effect, different methods are randomly selected to convert multiple same visible light images, and the comparison results are as Figure 4 shown in (a)-(b). As can be seen from the experimental results, the infrared images generated by the present invention are the closest to the real infrared images in terms of target integrity and overall tone, indicating that the algorithm proposed by the present invention has a better generation effect. At the same time, the present invention also further demonstrates the generation effect of the method on visible light images in unknown scenes, as Figure 5 shown, further verifying the generation ability of the present invention for infrared images in different scenarios. In Figure 5 the visible light data comes from randomly downloaded scene pictures on the Internet. Through the experimental results, it can be intuitively seen that the method proposed by the present invention can better generate the corresponding infrared images, verifying the superiority of the method of the present invention.
[0103] Those skilled in the art can easily understand that the above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent replacements, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A method for converting visible light to infrared images, characterized in that, Including: S1. Obtain paired visible light and infrared images as the training set; S2. Construct an infrared image conversion network; The infrared image conversion network includes: a plurality of feature extraction modules connected hierarchically from bottom to top, a plurality of parallel material feature separation modules corresponding to the levels of the feature extraction modules, a plurality of feature coupling modules connected hierarchically from top to bottom corresponding to the levels of the feature extraction modules, as well as a thermal data feature separation module and a thermophysical simulation module; Among them, the hierarchically connected feature extraction modules are used to extract features of different scales from visible light images; the parallel material feature separation modules are used to classify the material categories of feature maps of different scales respectively to obtain the material features corresponding to the feature maps of different scales; the thermal data feature separation module is used to obtain global thermal data features from the feature map of the largest scale; the cascaded feature coupling modules are used to select the most fully expressed material features from feature maps of different scales, fuse the selected material features with the corresponding thermal data features, and superimpose them with the features of the previous level to gradually generate a heat map; the thermophysical simulation module is used to convert the heat map into an infrared image; S3. Perform supervised iterative training on the infrared image conversion network using the training set to obtain a trained infrared conversion model; S4. Input the visible light image to be converted into the trained infrared conversion model to obtain the corresponding infrared image.
2. The visible light to infrared image conversion method according to claim 1, wherein During the training process of the infrared image conversion network, the following rotation consistency loss is used to constrain the feature separation module and the thermal data feature separation module: and represent the thermal data features before and after the image transformation respectively, and represent the material features of the i-th layer before and after the image transformation respectively, and T(·) represents the inverse transformation of the image transformation.
3. The method for converting visible light to infrared images according to claim 2, wherein The material feature separation module includes a convolutional layer, a pooling layer, and a pixel-level classifier connected in sequence.
4. A visible light to infrared image conversion method according to claim 3, characterized in that The thermal data feature separation module includes a convolutional layer, a global average pooling layer, and a feature mapping layer connected in sequence.
5. The visible light to infrared image conversion method according to claim 4, characterized in that The feature coupling module includes a material selection unit, a feature mapping unit, an adaptive instance normalization unit, a convolutional layer, and an upsampling layer; Among them, the material selection unit selects features of specific channels through the following operations; wherein is the channel weighting coefficient, represents the material feature of the i-th scale. Einsum represents Einstein summation. Based on parameters, re-weight the channels of When the parameter is closer to 1, it means that the feature of the channel is expressed, otherwise it is suppressed; The feature mapping unit further performs feature mapping on the input global material thermal data features to obtain the thermal data features corresponding to the selected material features; The adaptive instance normalization unit first performs adaptive feature fusion on the output of the current-level material selection unit and the output of the current-level feature mapping unit; then superimposes the fused feature map with the feature map of the previous level and performs instance normalization processing.
6. A method for converting visible light to infrared images according to claim 5, characterized in that, The feature mapping layer and the feature mapping unit are implemented using an EqualizedLinear fully connected layer.
7. A visible light to infrared image conversion method according to any one of claims 1-6, characterized in that, The thermophysical simulation module is composed of three residual blocks connected in series.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the program is executed by a processor, it implements the visible light to infrared image conversion method described in any one of claims 1-7.
Citation Information
Patent Citations
Complex scene segmentation method fusing visible light and infrared thermal image features
CN112700371A
Typical target material attribute extraction method and device fusing visible light and near infrared information
CN113537233A