Self-adaptive contrast learning infrared image colorization method and system based on human visual features

By constructing an adaptive contrast learning infrared image colorization method based on human eye visual characteristics, the color distortion and detail loss problems in infrared image conversion to RGB images are solved, high-quality colorization effect is achieved, and the robustness and generalization ability of the model are improved.

CN120451314APending Publication Date: 2025-08-08CHANGCHUN UNIV OF SCI & TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510577587.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-08-08

AI Technical Summary

Technical Problem

The prior art is difficult to convert infrared images into RGB images with high quality, which has problems such as color distortion, insufficient contrast and loss of detailed information, and it is difficult to collect paired data sets, which affects the training effect of deep learning models.

Method used

Adaptive contrast learning method based on human eye visual characteristics is adopted to build a generator and discriminator network, and the generator network is optimized, including initial convolution, feature downsampling, global residual and spatial group attention modules, through adaptive chroma proportion perception strategy, new positive and negative sample construction, and chroma similarity loss, the generator network is optimized, including initial convolution, feature downsampling, global residual and spatial group attention modules, to realize infrared image colorization.

Benefits of technology

Generate more realistic colored results that are in line with human vision, retain detailed information, and improve the quality and generalization ability of colored images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120451314A_ABST
    Figure CN120451314A_ABST
Patent Text Reader

Abstract

The invention discloses an adaptive contrast learning infrared image colorization method and system based on human visual features, and belongs to the technical field of image colorization. A self-adaptive contrast learning infrared image colorization method based on human visual features comprises the following steps: step 1, preparing an infrared data set: preprocessing a training set I and a training set II, and preprocessing a training set image with a fixed size; the first training set is a KAI ST data set, the second training set is an FL IR data set, and a fixed-size training set image is output by taking an original data set image as input; 2, constructing a network model, wherein the whole network is composed of an adversarial network, and the adversarial network comprises a generator, a discriminator, a sample marking operation and a generation fitting operation; contrast learning is introduced, brand new positive and negative samples are constructed, and a generator is designed, so that an output infrared image colorization result is more real and accords with local and overall feelings of visual observation of human eyes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image colorization, and in particular to a method and system for colorizing infrared images using adaptive contrast learning based on human visual characteristics. Background Art

[0002] Image colorization technology converts captured monochrome or grayscale images into more visually meaningful color images. It has found significant applications in a variety of fields, including military reconnaissance, medical diagnosis, security monitoring, and industrial inspection, and holds broad development prospects and significant research value. However, converting single-channel infrared images into high-quality three-channel RGB images still faces numerous technical challenges. First, due to fundamental differences in the imaging principles between infrared and visible light images, the quality of the resulting images often falls short of expectations, often exhibiting issues such as color distortion and insufficient contrast. Second, during the conversion process, detailed information in the original infrared image is easily lost, particularly texture features and edge information. Furthermore, the need to simultaneously capture infrared and visible light images of the same scene makes the collection of paired datasets challenging, such as device synchronization and environmental control. This hinders the effectiveness of model training based on deep learning methods. With the continuous advancement of research on human visual characteristics, scholars in various fields have been inspired to propose cutting-edge techniques widely used in deep learning, such as convolutional layers, pooling layers, normalization layers, and attention mechanisms. Continuously exploring the characteristics of human form remains crucial today.

[0003] The Chinese patent, granted with the number CN117876530B and titled "A Method for Colorizing Infrared Images Based on Reference Images," first constructs a generative adversarial network (GAN) consisting of a generator and a discriminator. The generator includes a dual-path feature mining module, an M-pooling layer, an A-pooling layer, a fusion attention module, a multi-information saliency converter, a cross-scale information aggregation module, a contextual information refinement module, a convolutional layer, and a T-type function to perform cross-chromatic colorization on the input infrared image. Next, the input dataset is preprocessed to a fixed size for training the entire convolutional neural network. The network model is then trained and fine-tuned by minimizing the loss value. Finally, the resulting model parameters are solidified to facilitate direct recall of the parameters for image colorization next time. The colorized images obtained by this method lack detailed information and differ significantly from human visual effects. Furthermore, using a reference infrared image dataset, the generalization performance is poor.

[0004] Therefore, we propose an adaptive contrast learning infrared image colorization method and system based on human visual characteristics to solve the above problems. Summary of the Invention

[0005] (1) Technical problems solved

[0006] In view of the deficiencies in the prior art, the present invention provides a method and system for infrared image colorization based on adaptive contrast learning of human visual characteristics, which solves the problems raised in the above background technology.

[0007] (2) Technical solution

[0008] In order to achieve the above-mentioned purpose, the present invention specifically adopts the following technical solutions:

[0009] An adaptive contrast learning infrared image colorization method based on human visual characteristics includes the following steps:

[0010] Step 1: Prepare infrared datasets: Preprocess training set 1 and training set 2, and fix the size of training set images after preprocessing; training set 1 is the KAIST dataset, and training set 2 is the FLIR dataset. Use the original dataset images as input to output fixed-size training set images;

[0011] Step 2: Build the network model: The overall network consists of an adversarial network, including a generator, a discriminator, a sample labeling operation, and a generation and fitting operation;

[0012] Step 3: Input training set images: Input the dataset prepared in step 1 into the model built in step 2 for training;

[0013] Step 4: Obtain the minimized loss function value and optimal evaluation index: Construct a minimized composite loss function to obtain a better evaluation index; the loss functions used by the generator include adversarial loss, feature loss, total variation loss, contrast loss, and chromatic similarity loss;

[0014] Step 5: Fine-tune the model: Input the infrared image into the model for training, compare the obtained results with image evaluation indicators, and then fine-tune the model parameters to obtain more accurate model parameters;

[0015] Step 6, save the model: solidify the finalized model parameters and output them as a model file. Then, in future testing and model hardware deployment, the model can be directly called for colorized output to obtain a colored image.

[0016] Furthermore, the preprocessing in step 1 adopts methods such as image cropping, image flipping and image type conversion to convert the input image size into the required size of the network input.

[0017] Furthermore, the sample labeling operation in step 2 is specifically as follows: using the generated colorized image and the unpaired color image to generate positive and negative samples in contrastive learning; performing a block operation on the generated colorized image, specifically, labeling 8 pixels in the colorized image with a size of 256×256 as a labeling unit; and labeling any target area v consisting of four labeling units with the same common point. i Classify the marked units according to the number of overlapping with the connected area, and classify the marked units with the target area v i The region where two labeled units overlap is named the strongly correlated region The region that overlaps with the target region by one labeled unit is named the weakly correlated region. The region with no overlapping marked units with the target region is named as the no-correlation region Therefore, we can get the value relative to the target area v i There are 4 relatively strong relationships with it 4 with relatively weak relationships 983 unrelated At the same time, the target area relative to the input color image is defined as the corresponding related area therefore There is only 1; the ultimate goal is to move the strongly correlated area Weakly correlated regions Perform feature fitting to fit positive samples Pull into target area v i and fitting positive samples The distance between them, further away from the target area v i No related areas and corresponding related areas the distance between them;

[0018] The generation and fitting operation defines a positive sample generation and fitting method based on the sample labeling operation; the target area v i The labeled units are numbered and the fitting operation is performed on each of the labeled units, so the positive sample of the upper left corner of the target area can be obtained. Fitting positive sample in the upper right corner Fitting positive sample in the lower left corner And the positive sample fitted in the lower right corner The area marked v in the upper left corner of the target area i_1 Generate fitting positive samples The process and v i_2 、v i_3 and v i_4 Generate fitting positive samples and The process is the same;

[0019] The upper left corner fits the positive sample It is expressed as follows:

[0020]

[0021] Among them, v i_1 is the numbered area in the upper left corner of the target area, For v i The overlapping marked unit is the labeled area v i_1 Weakly correlated regions The result obtained after average pooling can be expressed as: Where A(·) represents average pooling. Similarly, and and v i There are strongly correlated regions and one of the overlapping regions is the labeled region v i_1 The result obtained after average pooling can be expressed as and k i_0 、k i_1 、k i_2 and k i_3 They are defined as:

[0022]

[0023] in, For v i The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, so

[0024] Therefore, the final fitting positive sample can be expressed as:

[0025]

[0026] Furthermore, the generator in step 2 includes an initial convolution, a feature downsampling module, a spatial group attention module, a global residual module, a feature upsampling module, and a skip connection;

[0027] The initial convolution has a convolution kernel size of 1×1 and a step size of 1, which is used to perform high-dimensional mapping on the input image channels to extract high-dimensional detail information;

[0028] The feature downsampling module consists of a convolutional layer, a maximum pooling layer, a cross enhancement module, and a splicing operation, which is used to perform feature mining and fusion on multi-level downsampling information to reduce information loss during the downsampling process.

[0029] The spatial group attention module consists of global maximum pooling, Γ-type function, batch normalization layer, S-type function and matrix dot product, where the pooling kernel in global maximum pooling is the entire feature map, and the expression of Γ-type function can be expressed as: where x i is the i-th feature map, X is the input, which satisfies X=x1…n}, highlighting the different weights of the feature map to achieve the extraction and weighted highlighting of key features;

[0030] The global residual module consists of convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, convolution layer 7, convolution layer 8, convolution layer 9, convolution layer 10, channel superposition operation and convolution addition, wherein the convolution kernel size of convolution layer 1 and convolution layer 10 is 1×1, the step size is 1, the convolution kernel size of convolution layer 2 is 1×3, the step size is 1, the convolution kernel size of convolution layer 3 is 3×1, the step size is 1 The convolution kernel size of the fourth convolution layer is 1×5, with a step size of 1, the convolution kernel size of the fifth convolution layer is 5×1, with a step size of 1, the convolution kernel size of the sixth convolution layer is 1×7, with a step size of 1, the convolution kernel size of the seventh convolution layer is 7×1, with a step size of 1, the convolution kernel size of the eighth convolution layer is 1×9, with a step size of 1, and the convolution kernel size of the ninth convolution layer is 9×1, with a step size of 1, which is used to enhance and output the features of high-dimensional feature information;

[0031] The feature upsampling module consists of deconvolution, normalization and L-type function, which expands the size of the input feature map to half of its original size.

[0032] Furthermore, the discriminator in step 2 includes convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, normalization, R-type function, S-type function and spatial group attention module; wherein the convolution blocks in convolution layer 1, convolution layer 2, convolution layer 3 and convolution layer 4 use a convolution kernel size of 4×4 and a step size of 2, the convolution block in convolution layer 5 uses a convolution kernel size of 4×4 and a step size of 1, and the spatial group attention module is the same as the spatial group attention module network in the generator.

[0033] Furthermore, the contrast loss in step 4 realizes self-supervision of the colorized image without true value by constructing appropriate positive and negative samples, thereby generating a more realistic colorization effect;

[0034] The chromaticity similarity loss constructs a chromaticity ratio relationship between the input image and the unpaired color image in the contrast network, thereby building a feature bridge between the unpaired input images and realizing colorization of the unpaired infrared image.

[0035] An adaptive infrared image colorization system based on human visual characteristics, comprising:

[0036] An image acquisition module, used for acquiring an image to be colored;

[0037] The image adaptation module is used to perform adaptive processing on the acquired colorized image and divide it into a training set and a test set;

[0038] The model training module is used to input the processed training set into the designed network for training, obtain the optimal generation network parameters by constructing the optimal loss function, save the parameters, and use the test set to perform model colorization verification and generate a colorized image;

[0039] The contrast optimization module is used to visualize the intermediate feature maps generated during the model generation process and compare the final image with the input infrared image and the reference color image to further visually observe the image colorization effect;

[0040] The quality assessment module is used to evaluate whether the quality of the final colorized image meets the preset quality requirements; if it does, the generated colorized image is used as the final colorization result and the model parameters are solidified; if it does not, the quality improvement module is activated;

[0041] The quality improvement module is used to reacquire the dataset and continue training the model using the new dataset, repeating the model training module, comparison and optimization module, and quality assessment module until an image that meets the preset quality requirements is generated;

[0042] The hardware deployment module is used to deploy the model that meets the quality assessment module on hardware. Deploying the model on suitable hardware can help users complete the infrared image colorization process more conveniently.

[0043] (3) Beneficial effects

[0044] Compared with the existing technology, the present invention provides a method and system for infrared image colorization based on adaptive contrast learning of human visual characteristics, which has the following beneficial effects:

[0045] 1. The present invention proposes an adaptive chromaticity ratio perception strategy to learn more accurate chromaticity ratio information in continuous dynamic adversarial learning, thereby colorizing unpaired infrared images and obtaining more realistic colorization results.

[0046] 2. This invention proposes a new method for constructing positive and negative samples. By dividing the target area into chromaticity levels at different angles, positive and negative samples of different levels are obtained. Feature fitting is then performed on samples of different levels, and the final result is continuously shortened to the positive sample and further shortened to the negative sample.

[0047] 3. The present invention proposes chromaticity similarity loss, which builds a chromaticity ratio relationship between the input image and the unpaired color image in the comparison network, thereby building a feature bridge between the unpaired input images and realizing the colorization of the unpaired infrared image.

[0048] 4. The present invention proposes a generator network, designs a feature downsampling module, a global residual module and a spatial group attention module, and optimizes the generation process through different modules to obtain a colorized image that is more in line with the visual characteristics of the human eye. BRIEF DESCRIPTION OF THE DRAWINGS

[0049] Figure 1 is a flow chart of the method of the present invention;

[0050] Figure 2 It is a working principle diagram of the present invention;

[0051] Figure 3 This is a schematic diagram of the operation principle of marking samples at different levels in the present invention;

[0052] Figure 4 A diagram showing the principle of operation for generating a fitted positive sample in the present invention;

[0053] Figure 5 A generator network diagram for generating an adversarial network in the present invention;

[0054] Figure 6 This is a schematic diagram of the specific composition of all feature downsampling modules of the present invention;

[0055] Figure 7 This is a schematic diagram of the specific composition of all cross-enhancement modules of the present invention;

[0056] Figure 8 This is a schematic diagram of the specific composition of all spatial group attention modules of the present invention;

[0057] Figure 9 Schematic diagram of the specific composition of all global residual modules of the present invention;

[0058] Figure 10 This is a schematic diagram of the specific composition of all feature upsampling modules in the present invention;

[0059] Figure 11 A discriminator network diagram for generating an adversarial network in the present invention;

[0060] Figure 12 This is a schematic diagram comparing relevant indicators of the method proposed in the present invention;

[0061] Figure 13 Schematic diagram of the system of the present invention. DETAILED DESCRIPTION

[0062] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0063] Example

[0064] like Figure 1-13 As shown, an embodiment of the present invention proposes an infrared image colorization method based on adaptive contrast learning of human visual characteristics, which includes the following steps.

[0065] Step 1. Prepare the dataset and adapt it: Both the KAIST dataset and the FLIR dataset are composed of continuous frames of different videos. The adjacent images are not much different. Therefore, a certain degree of cleaning is performed on both datasets as the input of the entire network. Finally, 4755 images are obtained from the KAIST dataset as training dataset 1, and 3918 images are obtained from the FLIR dataset as training dataset 2.

[0066] Step 2: Build a network model, such as Figure 2 As shown in the figure, an adaptive contrast learning infrared image colorization method based on human visual features and a system working principle diagram are shown, which specifically includes a generator, a sample labeling operation, a generation fitting operation and a discriminator.

[0067] like Figure 3 As shown in Figure ①, it shows how to use the generated colorized image and the unpaired color image to generate positive and negative samples in contrastive learning; the generated colorized image is divided into blocks, specifically, 8 pixels in the colorized image with a size of 256×256 are marked as a labeling unit; the target area v i (v i The target area (the target area consisting of any four marker units with the same common point in the colorized image) is graded according to the number of marker units overlapping with the connected area, and the marker units overlapping with the target area v are classified into i The region where two labeled units overlap is named the strongly correlated region The region that overlaps with the target region by one labeled unit is named the weakly correlated region. The region with no overlapping marked units with the target region is named as the no-correlation region Therefore, we can get the value relative to the target area v i There are 4 relatively strong relationships with it 4 with relatively weak relationships 983 unrelated like Figure 3 As shown in Figure ②, the target area defined in the relative position of the input color image is the corresponding related area therefore There is only 1; Figure 3 As shown in Figure ③, the ultimate goal is to separate the strongly correlated areas Weakly correlated regions Perform feature fitting to fit positive samples Pull into target area v i and fitting positive samples The distance between them, further away from the target area v i No related areas and corresponding related areas The distance between them.

[0068] like Figure 4 As shown in the figure, a positive sample generation fitting method based on sample labeling operation is defined; the target area v i The labeled units are numbered and the fitting operation is performed on each of the labeled units, so the positive sample of the upper left corner of the target area can be obtained. (Here only the upper left corner of the target area is shown i_1 Generate fitting positive samples The process of v i_2 、v i_3 and v i_4 The process of the three regions is similar and will not be repeated here):

[0069]

[0070] Among them, v i_1 is the numbered area in the upper left corner of the target area, For v i The overlapping marked unit is the labeled area v i_1 Weakly correlated regions The result obtained after average pooling can be expressed as: Where A(·) represents average pooling. Similarly, and and v i There are strongly correlated regions and one of the overlapping regions is the labeled region v i_1 The result obtained after average pooling can be expressed as and ki_0 、k i_1 、k i_2 and k i_3 They are defined as:

[0071]

[0072] in, For v i The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, so

[0073] Therefore, the final fitting positive sample can be expressed as:

[0074]

[0075] in, To place the four marked areas in the target area v i The relative positions are stitched together.

[0076] like Figure 5 As shown in , the generator consists of initial convolution, feature downsampling module 1, spatial group attention module 1, feature downsampling module 2, spatial group attention module 2, feature downsampling module 3, spatial group attention module 3, global residual module 1, global residual module 2, global residual module 3, global residual module 4, global residual module 5, global residual module 6, global residual module 7, global residual module 8, global residual module 9, feature upsampling module 3, spatial group attention module 4, feature upsampling module 2, spatial group attention module 5, feature upsampling module 1, spatial group attention module 6, and jump connection; the initial convolution maps the input image channels to high dimensions to extract high-dimensional detail information. The convolution kernel size is 1×1, the step size is 1, and the number of image channels is increased in dimension; the feature downsampling module is as shown in Figure 6 As shown, it consists of convolution layer 1, convolution layer 2, maximum pooling, cross enhancement module and splicing operation. The convolution kernel size used in convolution layer 1 is 3×3 and the step size is 2. The convolution kernel size used in convolution layer 2 is 1×1 and the step size is 1. The pooling kernel of maximum pooling is 2×2. The cross enhancement module is as shown in Figure 7As shown in , it consists of convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, convolution layer 7, convolution layer 8, convolution layer 9, convolution layer 10, normalization operation, channel superposition operation and convolution addition. The convolution kernel size of convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 7, convolution layer 8 and convolution layer 9 is 3×3, with a stride of 2. Convolution layer 4, convolution layer 5 and convolution layer 6 are hollow convolutions with a convolution kernel size of 3×3, a stride of 1, an image padding of 2, and a hollow rate of 2. The convolution kernel size of convolution layer 10 is 1×1 and a stride of 1. The spatial group attention module is shown in Figure 8 As shown in Figure 1, it consists of global maximum pooling, Γ-type function, batch normalization layer, S-type function and matrix dot product. The pooling kernel in global maximum pooling is the entire feature map, and the Γ-type function expression can be expressed as: where x i is the i-th feature map, X is the input, which satisfies X={x1…n}; the global residual module is as follows Figure 9 As shown in , it consists of convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, convolution layer 7, convolution layer 8, convolution layer 9, convolution layer 10, channel superposition operation and convolution addition, where the convolution kernel size of convolution layer 1 and convolution layer 10 is 1×1, with a step size of 1, the convolution kernel size of convolution layer 2 is 1×3, with a step size of 1, the convolution kernel size of convolution layer 3 is 3×1, with a step size of 1, the convolution kernel size of convolution layer 4 is 1×5, with a step size of 1, the convolution kernel size of convolution layer 5 is 5×1, with a step size of 1, the convolution kernel size of convolution layer 6 is 1×7, with a step size of 1, the convolution kernel size of convolution layer 7 is 7×1, with a step size of 1, the convolution kernel size of convolution layer 8 is 1×9, with a step size of 1, and the convolution kernel size of convolution layer 9 is 9×1, with a step size of 1; the feature upsampling module is as shown in . Figure 10 As shown in , the convolution layer 1 is deconvolution, the convolution kernel size is 4×4, the step size is 2, and the padding is 1; the discriminator consists of convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, normalization, R-type function, S-type function and spatial group attention module, among which the convolution blocks in convolution layer 1, convolution layer 2, convolution layer 3 and convolution layer 4 use a convolution kernel size of 4×4 and a step size of 2, and the convolution block in convolution layer 5 uses a convolution kernel size of 4×4 and a step size of 1. The spatial group attention module is the same as the spatial group attention module network in the generator. The composition network of the discriminator is shown in Figure 11 shown.

[0077] In general, the infrared image colorization process is to determine the input infrared data set, preprocess the data set, and input the processed data set into the generator to generate a colorized image. Since the input color image is not paired with the input infrared image, the generated colorized image is subjected to a contrastive learning method and an adaptive chromaticity ratio perception strategy is designed to construct positive and negative samples to find the deep color relationship between the unpaired images, thereby reconstructing a more realistic colorized image. The infrared image is used as a reference, and the colorized image and the color image are respectively input into the discriminator to identify the difference between the generated colorized image and the color image, and the difference is quantified in the range of 0-1 to complete the evaluation of the degree of image colorization.

[0078] In order to ensure the robustness of the network and retain more network information, the present invention uses three activation functions, namely L-type function, R-type function and S-type function. The last layer of the generator uses L-type function, and the last layer of the feature downsampling module, cross enhancement module, global residual module and feature upsampling module all use L-type function; the last layer of the discriminator uses S-type function, and the rest are R-type function; the activation function used in the spatial group attention module is all S-type function; the L-type function, R-type function and S-type function are defined as follows

[0079]

[0080] In order to enhance the generalization ability of the network and prevent data overfitting, normalization operations are generally used in this invention. In this invention, all normalization operations are performed using batch normalization. Normalization is used before various activation functions to make the output of the previous layer distributed with a mean of 0 and a variance of 1, that is, to normalize the input of the next layer, so that it can have a certain gradient when passing through the activation function, thereby avoiding the value being too large and entering the saturation area.

[0081] Step 3: Input the training set images: Input the training dataset 1 and training dataset 2 in the dataset obtained in step 1 into the network constructed in step 2 for training.

[0082] Step 4. Obtain the minimized loss function value and optimal evaluation index: By constructing a reasonable loss function, the model can be constrained to optimize in the direction of the designer's attention. Therefore, the smaller the loss function, the better the robustness of the model. Selecting appropriate evaluation indicators can effectively reflect the quality of the model and the degree of image distortion, and measure the role of the colorization network.

[0083] In step 4, the network output and label loss function are calculated to achieve better fusion effect by minimizing the loss function; during the training process, the loss function uses adversarial loss Feature loss Total variational loss Contrastive loss and chromatic similarity loss Combine them optimally with certain weights to minimize the total loss function.

[0084] In order to encourage the network to output color results with more realistic details, an adversarial loss is used; the adversarial loss is used to learn the implicit relationship between the thermal infrared image and the colorized image, which is defined as:

[0085]

[0086] Among them, x is the input infrared image, y is the input color image, G(·) is the output of the entire generation network, and D(·,·) is the input of the two pictures into the discriminator.

[0087] The feature loss can effectively minimize the brightness and contrast differences between the color image and the colorized image. If the adversarial network focuses too much on the adversarial loss, the generated colorized image will over-exploit the implicit information between the images and fail to bring clearer detail information to the colorization. The feature loss can be expressed as:

[0088]

[0089] Among them, φ k (·) represents the feature representation of the kth max pooling layer in the VGG-16 network, C k H k W k represents the size of the image feature representation in the k max pooling layers, represents the generated colorized image, y represents the input color image, and ||·||1 represents the L1 norm.

[0090] In order to make the generated colorized image have a smoother transition, the present invention adopts the total variation loss. The calculation of the total variation loss does not require paired data. Adding the total variation loss to the total loss can make the edge transition of the generated colorized image smoother, which is more in line with the visual characteristics of the human eye. The total variation loss is defined as:

[0091]

[0092] Wherein, β is a constant, and the present invention adopts β=2.

[0093] In order to construct the color information between unpaired images, this paper adopts contrast loss to achieve self-supervision of the colorized image by constructing appropriate positive and negative samples, thereby generating a more realistic colorization effect. The correct construction of positive and negative samples will directly affect the effectiveness of contrast loss. Therefore, the contrast loss is defined as:

[0094]

[0095] in, v i is the i-th target area, is the target area v i The corresponding fitting positive sample is represents the kth non-correlated area sample of the i-th target area, where the total number of non-correlated area samples is K=983. It represents the corresponding relevant area, τ is the temperature parameter, and the value of the present invention is 0.07.

[0096] In order to deeply explore the color relationship between the unpaired colorized image and the color image, this paper proposes a chromaticity similarity loss. By obtaining the chromaticity ratio information of the target area in the color image and applying it to the colorized image, and simultaneously calculating the chromaticity ratio information in the colorized image, the chromaticity similarity loss between the two images is obtained. The chromaticity similarity loss is defined as:

[0097]

[0098] Among them, v i_1 is the numbered area in the upper left corner of the target area, For v i The overlapping marked unit is the labeled area v i_1 Weakly correlated regions The result obtained after average pooling can be expressed as: Where A(·) represents average pooling. Similarly, and and v i There are strongly correlated regions and one of the overlapping regions is the labeled region v i_1 The result obtained after average pooling can be expressed as and For v i The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, so c0, c1, c2, and c3 are proportional weight coefficients. The setting of proportional weights is based on preliminary experiments on the training data set, and the final weights are set to 10, 1, 5, and 5.

[0099] In order to achieve better model performance, the above five losses are combined to design the total loss of the model, constrain the model as a whole, and the total loss function is designed as follows:

[0100]

[0101] Among them, λ adv ,λ ft ,λ tv ,λ self and λ sss They represent the weights that control the complete loss function. The setting of the weights is based on preliminary experiments on the training data set, and the final weights are set to 0.03, 1, 1.25, 1, and 1 respectively.

[0102] In step 4, appropriate evaluation metrics for the network include Peak Signal-to-Noise Ratio (PSNR), Structural Similarity (SSIM), and Perceptual Image Similarity (LPIPS). PSNR is a fundamental metric in objective image quality evaluation. Its value directly reflects the pixel-level mean squared error (MSE) between the generated image and the original image. A larger PSNR value indicates less image distortion, better quality, and a more balanced color distribution. Structural Similarity (SSIM): Perceptual similarity is an objective evaluation metric that measures image similarity in terms of brightness, contrast, and network quality. It reflects the subjective quality of an image. A higher SSIM value indicates a higher degree of image restoration and better visual quality. Perceptual Image Similarity (LPIPS): The LPIPS metric aims to quantify the perceptual similarity between colorized images and colorized images. It is related to human visual perception. By training on manually annotated datasets, it learns feature extraction and distance metrics that better reflect human visual perception and better reflect the subjective perception of the degree of image perceptual information retention. A lower LPIPS value indicates a closer quality between the generated image and the ground-truth image, resulting in a better image quality. The definitions of PSNR, SSIM, and Perceptual Image Similarity are as follows:

[0103]

[0104] Among them, (2 n -1) 2 It represents the square of the maximum signal value, the number of bits of each sample value, and the maximum value of 8-bit representation is 255; MSE represents the mean of the squares of the differences between corresponding pixels of two images, and the expression is as follows: μ x and μ y Represent the mean and variance of images x and y respectively, and Represents the standard deviation of images x and y, σ xy represents the covariance of images x and y, C1 and C2 are constants; d is the distance between x and x0, and w1 is a trainable weight parameter.

[0105] The number of training runs is set to 400, the learning rate for the first 200 training runs is set to 0.0001, and the learning rate for the next 200 training runs is gradually decreased from 0.0001 to 0; the number of input images is fixed at 1 each time; the Adam optimizer is selected as the network parameter optimizer; its advantages mainly lie in its simple implementation, efficient computation, low memory requirements, and parameter updates that are not affected by gradient scaling, making the parameters relatively stable; when the discriminator's ability to determine whether the generated colorized image is a true color image is balanced with the generator's ability to generate colorized images, the network is considered to have been basically trained, and the discriminator output should be close to 0.5 at this time.

[0106] Step 5. Fine-tune the model: Use test dataset 1 and test dataset 2 to test and fine-tune the trained network; use test dataset 1 and test dataset 2 as input, calculate the evaluation index value in step 4, and then compare it with the expected evaluation index value, and then fine-tune the model until the expected evaluation index value is achieved.

[0107] Step 6. Save the model: Encapsulate the model that meets the desired evaluation indicators and save all parameter information in the network; the network can be applied to infrared image input of any size and complete the task of converting the input infrared image into the generated colorized image.

[0108] Among them, the implementation of convolution, activation function, splicing operation, addition operation, multiplication operation and normalization are algorithms well known to those skilled in the art, and the specific processes and methods can be found in corresponding textbooks or technical literature.

[0109] The present invention constructs an infrared image colorization method and system based on adaptive contrast learning of human visual characteristics, which can convert an input infrared image into a corresponding colorized image without going through other intermediate steps, avoiding the need for manual design of relevant colorization rules or other additional annotations; under the same experimental conditions, the feasibility and superiority of the method are further verified by calculating the relevant evaluation indicators of the colorized images obtained with the existing method; the relevant indicators of the existing technology and the method proposed by the present invention are compared. Figure 12 As shown; Figure 12 The figure shows the colorization effect evaluation of this method and the existing three algorithms on the KAIST dataset and FLIR dataset.

[0110] Existing Method 1: A generative adversarial network for colorization of unpaired infrared images consists of two generators and two discriminators, forming a recurrent adversarial network. The generator uses three downsampling modules, nine residual modules, and three upsampling modules; the discriminator uses six convolutional layers and a sigmoid function to distinguish the colorized image from the color image. This method was trained and tested using the training and test sets described in step 1 of the present invention, respectively.

[0111] Existing Method 2: A contrastive learning generative adversarial network for colorization of unpaired infrared images, consisting of a generator and a discriminator. The generator uses three downsampling modules, nine residual modules, and three upsampling modules; the discriminator uses six convolutional layers and a sigmoid function to distinguish the colorized image from the color image. Contrastive learning constructs positive and negative samples by selecting a target region in the image and then rotating and transforming the chromaticity of the target region to construct the positive sample, while the negative sample is the corresponding target region in the color image. This method is trained and tested using the training and test sets described in step 1 of the present invention, respectively.

[0112] Existing Method 3: A contrastive learning generative adversarial network for colorization of unpaired infrared images, consisting of a generator and a discriminator. The generator uses three downsampling modules, nine residual modules, and three upsampling modules; the discriminator uses six convolutional layers and R-type functions to distinguish generated images from real images. Contrastive learning constructs positive and negative samples by selecting a target region in an image. The positive sample is a frequency-domain transformation of the target region, and high-frequency information is used as an auxiliary tool to approach the positive sample, thereby achieving close proximity to the positive sample in the frequency domain. The negative sample is the corresponding target region in the color image. This method is trained and tested using the training set and test set described in step 1 of the present invention, respectively.

[0113] Figure 12 The figure shows the evaluation of the colorization effect of the proposed method and the existing three algorithms on the KAIST dataset and the FLIR dataset. The proposed method has a higher peak signal-to-noise ratio, higher structural similarity and lower perceived image similarity than the existing methods. These indicators also further illustrate that the proposed method has better colorization quality.

[0114] like Figure 13 The present invention also provides an adaptive contrast learning infrared image colorization system based on human visual characteristics. The system mainly includes an image acquisition module, an image adaptation module, a model training module, a contrast optimization module, a quality assessment module, a quality improvement module and a hardware deployment module.

[0115] The image acquisition module is used to acquire the image to be colored.

[0116] The image adaptation module is used to perform adaptive processing on the acquired colorized image and divide it into a training set and a test set.

[0117] The model training module is used to input the processed training set into the designed network for training, obtain the optimal generation network parameters by constructing the optimal loss function, save the parameters, use the test set to perform model colorization verification, and generate a colored image.

[0118] The contrast optimization module is used to visualize the intermediate feature maps generated during the model generation process and compare the final generated image with the input infrared image and the reference color image to further intuitively observe the image colorization effect.

[0119] The quality assessment module is used to evaluate whether the quality of the final colorized image meets the preset quality requirements; if it does, the generated colorized image will be used as the final colorization result, and the model parameters will be solidified; if it does not meet the preset quality requirements, the quality improvement module will be started.

[0120] The quality improvement module is used to reacquire the dataset and continue training the model using the new dataset, repeating the model training module, comparison and optimization module, and quality assessment module until an image that meets the preset quality requirements is generated.

[0121] The hardware deployment module is used to deploy the model that meets the quality assessment module on hardware. Deploying the model on suitable hardware can help users complete the infrared image colorization process more conveniently.

[0122] Furthermore, the image to be colorized in the image acquisition module is an infrared image, or can be a grayscale image, and does not require a true value image.

[0123] Furthermore, the preprocessing in the image processing module includes image cropping and image flipping, and the final image ratio of the training set and the test set is 10:1.

[0124] Furthermore, the visualization image during the training process in the comparison optimization module is the attention heat map in the spatial group attention module. By observing the attention heat map, it is possible to clearly understand whether the focus area of the attention module is the target area; the visualization image at the end of training is saved after the training is completed, and the model is predicted. The predicted image is compared with the input infrared image and the colorization effect can be intuitively felt.

[0125] Furthermore, the evaluation indicators of the quality assessment module are peak signal-to-noise ratio (PSNR), structural similarity (SSIM) and perceptual image similarity (LPIPS); the higher the value of PSNR, the better; the higher the value of structural similarity, the better; and the higher the value of perceptual image similarity, the better.

[0126] Furthermore, the hardware deployment of the hardware deployment module needs to be deployed on suitable hardware that supports the algorithm model. The hardware should at least include memory, processor and communication interface. By storing the input model in the memory, the model can be read at any time. The processor is used to load the model and perform calculations. The communication interface is used to accept the input training data set and test data set, and output the output results for display or transmission to other terminal devices.

[0127] Finally, it should be noted that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the aforementioned embodiments, those skilled in the art will be able to modify the technical solutions described in the aforementioned embodiments or substitute equivalents for some of the technical features. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. An adaptive contrast learning infrared image colorization method based on human visual characteristics, characterized by: The steps include: Step 1: Prepare infrared datasets: Preprocess training set 1 and training set 2, and fix the size of training set images after preprocessing; training set 1 is the KAIST dataset, and training set 2 is the FLIR dataset. Use the original dataset images as input to output fixed-size training set images; Step 2: Build the network model: The overall network consists of an adversarial network, including a generator, a discriminator, a sample labeling operation, and a generation and fitting operation; Step 3: Input training set images: Input the dataset prepared in step 1 into the model built in step 2 for training; Step 4: Obtain the minimized loss function value and optimal evaluation index: Construct a minimized composite loss function to obtain a better evaluation index; the loss functions used by the generator include adversarial loss, feature loss, total variation loss, contrast loss, and chromatic similarity loss; Step 5: Fine-tune the model: Input the infrared image into the model for training, compare the obtained results with image evaluation indicators, and then fine-tune the model parameters to obtain more accurate model parameters; Step 6, save the model: solidify the finalized model parameters and output them as a model file. Then, in future testing and model hardware deployment, the model can be directly called for colorized output to obtain a colored image.

2. The method for infrared image colorization based on adaptive contrast learning of human visual characteristics according to claim 1, characterized in that: The preprocessing in step 1 adopts methods such as image cropping, image flipping and image type conversion to convert the input image size into the required size of the network input.

3. The method for infrared image colorization based on adaptive contrast learning of human visual characteristics according to claim 1, characterized in that: The sample labeling operation in step 2 is specifically as follows: using the generated colorized image and the unpaired color image to generate positive and negative samples in contrastive learning; performing a block operation on the generated colorized image, specifically, labeling 8 pixels in the colorized image of size 256×256 as a labeling unit; and labeling any target area v consisting of four labeling units with the same common point. i The labeled units are classified according to the number of overlapping with the connected regions, and the target region v i The region where two labeled units overlap is named the strongly correlated region The region that overlaps with the target region by one labeled unit is named the weakly correlated region. The region with no overlapping marked units with the target region is named as the no-correlation region Therefore, we can get the value relative to the target area v i There are 4 relatively strong relationships with it 4 with relatively weak relationships 983 unrelated At the same time, the target area relative to the input color image is defined as the corresponding related area therefore There is only one; the ultimate goal is to move the strongly correlated area Weakly correlated regions Perform feature fitting to fit positive samples Pull into target area v i and fitting positive samples The distance between them, further away from the target area v i No related areas and corresponding related areas the distance between them; The generation and fitting operation defines a positive sample generation and fitting method based on the sample labeling operation; the target area v i The labeled units are numbered and the fitting operation is performed on each of the labeled units, so the positive sample of the upper left corner of the target area can be obtained. Fitting positive sample in the upper right corner Fitting positive sample in the lower left corner And the positive sample fitted in the lower right corner The area marked v in the upper left corner of the target area i_1 Generate fitting positive samples The process and v i_2 、v i_3 and v i_4 Generate fitting positive samples and The process is the same; The upper left corner fits the positive sample It is expressed as follows: Among them, v i_1 is the numbered area in the upper left corner of the target area, For v i The overlapping marked unit is the labeled area v i_1 Weakly correlated regions The result obtained after average pooling can be expressed as: Where A(·) represents average pooling. Similarly, and and v i There are strongly correlated regions and one of the overlapping regions is the labeled region v i_1 The result obtained after average pooling can be expressed as and k i_0 、k i_1 、k i_2 and k i_3 They are defined as: in, For v i The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, For The corresponding area on the color image, so Therefore, the final fitting positive sample can be expressed as:

4. The method for infrared image colorization based on adaptive contrast learning of human visual characteristics according to claim 1, characterized in that: The generator in step 2 includes an initial convolution, a feature downsampling module, a spatial group attention module, a global residual module, a feature upsampling module, and a skip connection; The initial convolution has a convolution kernel size of 1×1 and a step size of 1, which is used to perform high-dimensional mapping on the input image channels to extract high-dimensional detail information; The feature downsampling module consists of a convolutional layer, a maximum pooling layer, a cross enhancement module, and a splicing operation, which is used to perform feature mining and fusion on multi-level downsampling information to reduce information loss during the downsampling process. The spatial group attention module consists of global maximum pooling, Γ-type function, batch normalization layer, S-type function and matrix dot product, where the pooling kernel in global maximum pooling is the entire feature map, and the expression of Γ-type function can be expressed as: where x i is the i-th feature map, X is the input, which satisfies X=x1…n}, highlighting the different weights of the feature map to achieve the extraction and weighted highlighting of key features; The global residual module consists of convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, convolution layer 7, convolution layer 8, convolution layer 9, convolution layer 10, channel superposition operation and convolution addition, wherein the convolution kernel size of convolution layer 1 and convolution layer 10 is 1×1, the step size is 1, the convolution kernel size of convolution layer 2 is 1×3, the step size is 1, the convolution kernel size of convolution layer 3 is 3×1, the step size is 1 The convolution kernel size of the fourth convolution layer is 1×5, with a step size of 1, the convolution kernel size of the fifth convolution layer is 5×1, with a step size of 1, the convolution kernel size of the sixth convolution layer is 1×7, with a step size of 1, the convolution kernel size of the seventh convolution layer is 7×1, with a step size of 1, the convolution kernel size of the eighth convolution layer is 1×9, with a step size of 1, and the convolution kernel size of the ninth convolution layer is 9×1, with a step size of 1, which is used to enhance and output the features of high-dimensional feature information; The feature upsampling module consists of deconvolution, normalization and L-type function, which expands the size of the input feature map to half of its original size.

5. The method for infrared image colorization based on adaptive contrast learning of human visual characteristics according to claim 1, characterized in that: The discriminator in step 2 includes convolution layer 1, convolution layer 2, convolution layer 3, convolution layer 4, convolution layer 5, convolution layer 6, normalization, R-type function, S-type function and spatial group attention module; the convolution blocks in convolution layer 1, convolution layer 2, convolution layer 3 and convolution layer 4 use a convolution kernel size of 4×4 and a step size of 2, the convolution block in convolution layer 5 uses a convolution kernel size of 4×4 and a step size of 1, and the spatial group attention module is the same as the spatial group attention module network in the generator.

6. The method for infrared image colorization based on adaptive contrast learning of human visual characteristics according to claim 1, characterized in that: In step 4, the contrast loss constructs appropriate positive and negative samples to achieve self-supervision of the colorized image without true value, thereby generating a more realistic colorization effect; The chromaticity similarity loss constructs a chromaticity ratio relationship between the input image and the unpaired color image in the contrast network, thereby building a feature bridge between the unpaired input images and realizing colorization of the unpaired infrared image.

7. An adaptive infrared image colorization system based on human visual characteristics, used to execute the adaptive infrared image colorization method based on human visual characteristics according to any one of claims 1 to 6, characterized in that: The system comprises: An image acquisition module, used for acquiring an image to be colored; The image adaptation module is used to perform adaptive processing on the acquired colorized image and divide it into a training set and a test set; The model training module is used to input the processed training set into the designed network for training, obtain the optimal generation network parameters by constructing the optimal loss function, save the parameters, and use the test set to perform model colorization verification and generate a colorized image; The contrast optimization module is used to visualize the intermediate feature maps generated during the model generation process and compare the final image with the input infrared image and the reference color image to further visually observe the image colorization effect; The quality assessment module is used to evaluate whether the quality of the final colorized image meets the preset quality requirements; if it does, the generated colorized image is used as the final colorization result and the model parameters are solidified; if it does not, the quality improvement module is activated; The quality improvement module is used to reacquire the dataset and continue training the model using the new dataset, repeating the model training module, comparison and optimization module, and quality assessment module until an image that meets the preset quality requirements is generated; The hardware deployment module is used to deploy the model that meets the quality assessment module on hardware. Deploying the model on suitable hardware can help users complete the infrared image colorization process more conveniently.

Citation Information

Patent Citations

  • Infrared image colorization method and system based on generative adversarial network

    CN116645569A

  • Self-adaptive infrared image colorization method and system based on human visual features

    CN119722843A