Unmanned aerial vehicle infrared image coloring method based on topological semantic structure loss
By introducing topological structure semantic constraint module and discriminator guide attention sampling strategy in infrared image shading method, the problems of image texture distortion and blurred details after infrared image shading in the prior art are solved. The generated color images are more in line with human visual habits and improve the image acquisition effect.
Patent Information
- Application Number
- CN202510542252.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing infrared image shading methods often experience texture distortion and blurred details in shading images, which fail to effectively capture the topologically perceived semantic relationship between patches.
The infrared image shading method based on topological semantic structure loss is adopted. By building a generative network containing encoder and decoder, the topological structure semantic constraint module and discriminator are introduced to guide attention sampling strategy, capture cross-domain topological semantic information, and ensure the semantic consistency of infrared images and color images in the topological structure.
It effectively solves the problems of texture distortion and blurred details when infrared images are mapped to color images. The generated color images are more in line with human observation and improves the effect of drone's night image acquisition.
Smart Images

Figure CN120070644A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and deep learning, and specifically relates to an unmanned aerial vehicle (UAV) infrared image coloring method based on topological semantic structure loss. Background Art
[0002] In a completely dark environment, a color-spectrum imaging system cannot generate clear images, while a thermal infrared imaging system is insensitive to lighting conditions and has strong penetrability. Therefore, it has become a powerful tool for night perception and is suitable for night detection imaging of UAVs. However, due to the low contrast and monochromatic nature of infrared images, it limits human understanding. Infrared image coloring is to convert a single-channel infrared image into a multi-channel color image that conforms to human visual habits while maintaining semantic consistency, so as to improve its interpretability and visualization effect. At present, this technology has many extensive applications. For example, in medical imaging, colored infrared images help to identify diseased tissues or regions. In security monitoring and military reconnaissance, colored infrared images can help personnel identify targets faster and more accurately. Therefore, it is of great significance to convert night infrared images into corresponding daytime color images.
[0003] Most existing infrared image coloring methods are based on supervised learning with paired color and infrared images, such as image fusion networks and image-to-image translation networks that require paired datasets for training. Supervised coloring methods usually require a large amount of labeled data, but it is relatively difficult to obtain a large number of pixel-level aligned cross-domain image pairs in practice. To solve this problem, in recent years, more and more research has been devoted to using unsupervised image-to-image translation methods to achieve infrared image coloring, including two main lines. The first main line is to introduce a generative adversarial network (GAN), and through the adversarial training of the generator and discriminator, the generated image gradually approaches the target image. For example, CycleGAN uses cyclic consistency constraints to make the reconstructed image consistent with the original image after reverse conversion. SN-DCR uses spectral normalization to improve the problem of low image quality caused by mode collapse in the GAN network. The second main line is to enhance the feature alignment between the input image and the output image through contrastive learning methods. For example, CUT and FastCUT achieve effective cross-domain conversion by maximizing the mutual information between the corresponding patches of the input image and the output image. NEGCUT generates negative samples based on the input image through a novel negative sample generator.
[0004] When the above methods are applied to infrared image colorization, the colorization effect is improved to a certain extent, but there are still limitations. For the generative adversarial network, color mismatches may occur during the encoding-decoding process of the generator, and the model is also prone to collapse during the training process, resulting in low training efficiency, which causes texture distortion or loss of specific targets such as vehicles and pedestrians in the colorized images. The contrast learning method only performs matching patch by patch, using patches at the same position in the infrared image and the color image as positive samples, and patches at different positions as negative samples. Therefore, it ignores the topological-aware semantic relationship between patches, resulting in problems such as texture distortion and detail blurring in the colorized images.
[0005] Therefore, to solve these problems, this paper proposes an infrared image colorization method with topological-aware semantic constraints to solve the above problems of texture distortion and detail blurring in the colorized images caused by ignoring the topological-aware semantic relationship between patches, which helps to generate color pictures that are more in line with human observation. Summary of the Invention
[0006] The purpose of the present invention is to provide a method for colorizing UAV infrared images based on topological semantic structure loss, mainly to solve the problems of texture distortion and detail blurring in the colorized images in the prior art.
[0007] To achieve the above object, the technical solution adopted by the present invention is as follows: A method for colorizing UAV infrared images based on topological semantic structure loss includes the following steps: S1, constructing a non-paired infrared image and a multi-channel color image with the same image size as the input data set; building a generative network including an encoder and a decoder, where the encoder and the decoder are respectively composed of L layers of sub-networks, and the infrared image is input into the generative network for feature extraction and preliminary colorization; S2, introducing a topological structure semantic constraint module, sharing the adjacency matrix between specific patches of infrared-domain and color-domain images, and using a graph neural network to capture the node feature representations of graphs in different domains; constraining the mutual information between graph nodes through a contrast loss function to ensure the semantic consistency of the infrared image and the color image in terms of topological structure; S3, guiding the attention sampling strategy to select the most informative patches during the training process through a discriminator, calculating the importance score of the patches using the discriminator, and performing oversampling, sorting, and importance sampling to guide the generator to maintain content consistency during the encoding and decoding stages; S4, constructing a contrast loss function based on the image topological structure, and at the same time introducing an adversarial loss function, a total variation loss function, and a multi-stage loss function to train the generated color image; S5: Input the infrared images in the test set of the dataset into the trained infrared image colorization network to generate multi-channel color images that conform to human visual habits, while maintaining the scene structure information and details.
[0008] Further, in the step S1, the generation network uses a Resnet-based feature extractor, and constrains the similarity of the representations of the encoder and decoder features in the latent space in the generator to ensure content consistency constraints during the generation process. Among them, the content consistency constraints use an asymmetric contrastive learning structure to constrain the features in multiple stages of the network. The specific constraint process is as follows: Since the semantic levels of the encoder and decoder features are the same in the same stage, define and as the features of the l th stage of the encoder and decoder respectively; sample and patches at the same position, and map them to the -dimensional latent space through two projection heads with shared weights K to obtain and . Further, add a prediction head to to obtain . Introduce stop gradient, that is, , so that the encoder feature remains static during propagation , avoiding its being updated by the gradient, thereby providing a stable target representation for contrastive learning. Finally, a multi-stage content loss is defined, using an asymmetric contrastive learning structure to constrain the features in multiple stages of the network, ensuring that the patch features sampled by the encoder and decoder in the same stage are as similar as possible in the latent space representation, thereby achieving content consistency. The expression of the multi-stage content loss function is as follows: In the formula, represents the expectation of the sample X from the dataset x ; represents Euclidean norm normalization; L represents the total number of constraint stages; Q represents the total number of patch samplings.
[0009] Further, in the step S2, the implementation principle of the topological structure semantic constraint module is to utilize the characteristic that the topological structures of infrared images and color images are the same. Among them, the patch feature of the infrared image is used as the node of the graph structure , and the patch feature As a graph structure of the nodes, and calculate the patch features and the patch features to determine the connection relationship by calculating the cosine similarity between them; if the cosine similarity between two patches is greater than or equal to a set threshold t , then there is a relationship considered to exist between these two patches, that is, they are associated nodes in the graph structure; among them, the adjacency matrix and the features , have the following relationship: ; Among them, the nodes interact through the adjacency matrix of the graph to capture the information of local topological perception semantics; v represents the node features of the color image obtained after graph convolution, which represents the updated information containing the aggregated adjacent node features and also reflects the local topological structure; definition: ; In the formula, is the node feature of the infrared image of the i th node, is the node feature of the color image of the i th node; D is the degree matrix of the node, representing the degree of each node, is the shared k hop representation weight matrix; K represents the number of hops used in the graph convolution operation.
[0010] Furthermore, in the step S2, maximizing the mutual information is expressed as: ; In the formula, x represents the infrared image, represents the color image; N represents the number of nodes in the graph neural network, exp represents the exponential calculation, and T represents the transpose operation.
[0011] Furthermore, in the step S3, the specific steps of the discriminator-guided attention sampling strategy are as follows: S31: Use the discriminator to evaluate the generated image, and adjust the output of the discriminator to the same resolution as the color image feature map in the generator Decoder through interpolation to obtain the attention score matrix: Wherein, represents interpolating the discriminator output to the resolution ; S32: Sampling based on importance, uniformly sampling patches from the generated feature map patches, where is the oversampling ratio, > 1, Q is the total number of patch samplings; Let the sampling set represent all oversampled patches, and for each sampled patch corresponding attention score is sorted in ascending order, where is the patch in the feature map pixel position coordinates, select the top patches with the highest attention scores from the sorted patches, is the importance sampling ratio; The selected patch set is represented as ; S33: Perform uniform sampling, uniformly sample patches from the remaining patches to obtain a set of patches as , then the finally selected patch set can be expressed as: .
[0012] Furthermore, in the step S4, the expression of the adversarial loss function is: ; Wherein, x is the original infrared image, y is a random real color image, is the generator, is the discriminator, represents the expectation of the sample X in the noise distribution x , represents the expectation of the sample y in the real data distribution Y.
[0013] Furthermore, in the step S4, the expression of the total variation loss function is: ; The final loss function is: ; Wherein, is the topological perception semantic loss weight.
[0014] Compared with the prior art, the present invention has the following beneficial effects: (1) Through the topological perception semantic constraint module and discriminator-guided attention sampling strategy that share the adjacency matrix, the present invention captures cross-domain topological semantic information and enhances the mutual information between nodes, solves the problems of texture distortion and detail blurring when mapping infrared images to color images in the prior art, and improves the effect of drone night image acquisition.
[0015] (2) The present invention only uses unpaired infrared and color images for training, without relying on infrared data sets. By constructing the graph structures of infrared and color images, the cross-modal topological semantic consistency is ensured. Experimental results show that the present invention can generate color images that conform to human visual habits, while retaining details and scene structures in complex scenes, improving the performance and visual quality of infrared image coloring. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] Figure 1 is the overall flowchart of the infrared image coloring method based on topological semantic structure loss.
[0017] Figure 2 is the architecture diagram of the infrared image coloring method based on topological semantic structure loss.
[0018] Figure 3 is the discriminator-guided attention sampling strategy diagram.
[0019] Figure 4 is the generator architecture diagram.
[0020] Figure 5 are the experimental results of infrared image coloring of the infrared image coloring method based on topological semantic structure loss on the KAIST data set.
[0021] Figure 6 are the experimental results of infrared image coloring of the infrared image coloring method based on topological semantic structure loss on the FLIR data set. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The present invention will be further described below in conjunction with the drawings and embodiments. The implementation manners of the present invention include but are not limited to the following embodiments. Embodiment
[0023] As Figure 1 shown, an infrared image coloring method for drones based on topological semantic structure loss disclosed by the present invention includes the following steps: S1, constructing unpaired infrared images and multi-channel color images with the same image size as the input data set; building a generation network including an encoder and a decoder, where the encoder and the decoder are respectively composed of LIt consists of sub-networks. The infrared image is input into the generation network for feature extraction and preliminary coloring. Among them, the generation network uses a Resnet-based feature extractor and constrains the similarity of the representations of the encoder and decoder features in the latent space in the generator to ensure content consistency constraints during the generation process. Among them, the content consistency constraint uses an asymmetric contrast learning structure to constrain the features in multiple stages of the network. The specific constraint process is as follows: As Figure 4 shown, an asymmetric contrast learning structure as shown in Figure 3 is introduced in the same stage of the encoder and decoder: the decoder, as the online network (the dark part), includes a projection and a prediction stage, while the encoder, as the target network (the light part), only includes a projection stage. The two networks share the weights of the projection layer, and the online network is optimized through multi-stage content losses. The generator is divided into two parts, the encoder and the decoder, and each part consists of L levels of sub-networks. Since the semantic levels of the encoder and decoder features are the same at the same stage, define and as the features of the l th stage of the encoder and decoder respectively; sample and patches at the same position and map them to the -dimensional latent space through two projection heads with shared weights K to obtain and . Further, add a prediction head to to obtain . To prevent the collapse solution, that is, the network learns to output constants to minimize the loss, stop gradients are introduced, that is , so that the encoder feature remains static during propagation and is avoided from being updated by gradients, thus providing a stable target representation for contrast learning. Finally, a multi-stage content loss is defined, which uses an asymmetric contrast learning structure to constrain the features in multiple stages of the network to ensure that the patch features sampled by the encoder and decoder at the same stage are as similar as possible in the latent space, thereby achieving content consistency. Its expression is as follows: In the formula, X represents the expectation of the sample x from the data set ; L represents Euclidean norm normalization; Q represents the total number of constraint stages; represents the total number of patch samplings.
[0024] S2. Introduce a topological structure semantic constraint module, share the adjacency matrix between specific patches of infrared-domain and color-domain images, and use a graph neural network to capture the node feature representations of graphs in different domains; constrain the mutual information between graph nodes through a contrastive loss function to ensure the semantic consistency of infrared images and color images in terms of topological structure. Among them, the implementation principle of the topological structure semantic constraint module is to utilize the same topological structure characteristics of infrared images and color images. Among them, the patch features of the infrared image are used as the nodes of the graph structure , and the patch features of the color image are used as the nodes of the graph structure , and calculate the cosine similarity between the patch features and the patch features to determine the connection relationship ; if the cosine similarity between two patches is greater than or equal to the set threshold t , then there is a relationship considered to exist between these two patches, that is, they are associated nodes in the graph structure. Among them, the adjacency matrix and the features , have the following relationship: ; Among them, nodes interact through the adjacency matrix of the graph to capture information on local topological perception semantics; v represents the node features of the color image obtained after graph convolution, indicating the updated information containing aggregated adjacent node features and also reflecting the local topological structure. Definition: ; In the formula, is the node feature of the infrared image of the i th node, is the node feature of the color image of the i th node; D is the degree matrix of the nodes, indicating the degree of each node, L2Norm is L2 a normalization operation, is the shared k skip representation weight matrix; K represents the number of hops used in the graph convolution operation.
[0025] In this embodiment, maximizing the mutual information is expressed as: ; In the formula, x represents the infrared image, represents the color image; Ndenotes the number of nodes in the graph neural network, exp represents exponential calculation, and T represents the transpose operation.
[0026] S3. Through the discriminator-guided attention sampling strategy, the most informative patches during the training process are selected. The discriminator is used to calculate the importance scores of the patches, and oversampling, sorting, and importance sampling are performed to guide the generator to maintain content consistency during the encoding and decoding phases. As Figure 3 shown, the specific steps of the discriminator-guided attention sampling strategy are as follows: S31: Use the discriminator to evaluate the generated image. To ensure that the color image generated by the generator is consistent with the input infrared image in terms of feature size, the output of the discriminator is adjusted to the same resolution as the color image feature map in the generator Decoder to obtain the attention score matrix: In the formula, represents interpolating the discriminator output to the resolution ; S32: Sampling based on importance. To ensure the sufficiency of information, patches are uniformly sampled from the generated feature map , where is the oversampling ratio, >1, Q is the total number of patch samplings; let the sampling set represent all oversampled patches. For each sampled patch , the corresponding attention score is sorted in ascending order, where is the pixel position coordinate of the patch in the feature map . The top patches with the highest attention scores are selected from the sorted patches, where is the importance sampling ratio; the selected patch set is denoted as ; S33: Perform uniform sampling. Uniformly sample patches from the remaining patches to obtain the patch set . Then the finally selected patch set can be expressed as: .
[0027] S4. Construct a contrastive loss function based on the image topological structure, and at the same time introduce an adversarial loss function, a total variation loss function, and a multi-stage loss function to train the generated color image. The expression of the adversarial loss function is: ; In the formula, x is the original infrared image, y is a random real color image, is the generator, is the discriminator, represents the expectation of the sample X in the noise distribution x , represents the expectation of the sample y in the real data distribution Y.
[0028] The expression of the total variation loss function is: ; The final loss function is: ; In the formula, is the topological perception semantic loss weight; in this embodiment, the topological perception semantic loss weight is set to 0.1.
[0029] S5: Input the test set infrared images in the dataset into the trained infrared image colorization network to generate multi-channel color images that conform to human visual habits, and maintain the scene structure information and details.
[0030] In this embodiment, as Figure 2 shown, the input infrared image x and the generated image G(x) first extract features E and through the feature extractor , and send them into the topological structure semantic constraint module. This module selects patches through the score matrix D generated by the discriminator to improve the training efficiency and generation effect. After sampling the features and through the discriminator-guided attention sampling strategy, they are used as the nodes of the graph structure and input into the GNN to extract the graph structure representations Z and V of the infrared image and the color image respectively. Through the shared graph neural network, the topological structure relationship between patches is constructed to maximize the mutual information between the infrared image and the generated image, that is, to minimize the topological semantic loss to maintain semantic consistency. The generator G generates a visible light image according to the topologically optimized features, and the discriminatorD Perform quality assessment on the generated images, which is used to supervise the realism of the generated images. Finally, the generated color images are further optimized for details through the total variation loss to ensure smooth and realistic textures.
[0031] Figure 5 It is a graph showing the experimental results of infrared image colorization of the infrared image colorization method based on topological semantic structure loss on the KAIST dataset. The first column is the real infrared image, and the ninth column is the true color image generated by the present invention. Except for the FastCUT method, almost all methods have achieved a certain degree of success in the infrared colorization task, but the finally generated color pictures often lack color information and local information, resulting in an unrealistic appearance. For example, in the first row of pictures, except for the Ours method, the traffic signs on the road are distorted to varying degrees. In the second row of pictures, only CycleGAN, IRC, and Ours restored the detailed information of the vehicle, and NEGCUT and SN-DCR misidentified the building as the sky. In contrast, the infrared image colorization method based on topological semantic structure loss shows considerable progress in terms of fidelity and details, while retaining details and maintaining clear contours.
[0032] Figure 6 It is a graph showing the experimental results of infrared image colorization of the infrared image colorization method based on topological semantic structure loss on the FLIR dataset. The first column is the real infrared image, and the ninth column is the true color image generated by the present invention. The infrared image colorization method based on topological semantic structure loss also achieved significantly better results. For example, in the first row of pictures, only the Ours method can colorize the tree crown while retaining the edge features of the vehicle. In the fourth row of pictures, only IRC and Ours retained the detailed features of the traffic lights. This is because the method in this paper uses GNN to perform structured constraints on local detail information, and information can be transmitted between different patches through GNN. The local features of objects such as traffic lights and trees in a patch in the infrared image can be transmitted to adjacent patches, ensuring a more complete presentation of these objects in the converted color image. This cross-patch transmission and fusion of information helps to eliminate artifacts caused by a single patch being unable to capture complete details.
[0033] The above embodiments are only one of the preferred embodiments of the present invention and should not be used to limit the protection scope of the present invention. Any modification or polishing made without substantial meaning in the main design idea and spirit of the present invention, as long as the technical problems solved are still consistent with the present invention, should be included in the protection scope of the present invention.
Claims
1. A method for colorizing UAV infrared images based on topological semantic structure loss, characterized in that: The following steps are involved: S1, construct unpaired infrared images and multi-channel color images with consistent image sizes as input data sets; build a generative network including an encoder and a decoder, where the encoder and the decoder are respectively L The infrared image is input into the generative network for feature extraction and preliminary colorization. S2, introduces a topological structure semantic constraint module, shares the adjacency matrix between specific patches of infrared and color domain images, and uses graph neural networks to capture node feature representations of graphs in different domains; by contrasting the loss function to constrain the mutual information between graph nodes, the semantic consistency of the topological structure of infrared and color images is ensured; S3, uses the discriminator to guide the attention sampling strategy to select the most informative patches during training, uses the discriminator to calculate the importance scores of patches, and performs oversampling, sorting, and importance sampling to guide the generator to maintain content consistency during encoding and decoding stages; S4, constructs a contrast loss function based on the image topology structure, and introduces adversarial loss function, total variation loss function and multi-stage content loss function to train the generated color images; S5: Input the test set infrared images in the dataset into the trained infrared image colorization network to generate multi-channel color images that conform to human visual habits and maintain scene structure information and details.
2. The method for colorizing an infrared image of a drone based on topological semantic structure loss according to claim 1, characterized in that: In step S1, the generative network adopts a feature extractor based on Resnet, and constrains the representation similarity of the encoder and decoder features in the generator in the latent space to ensure content consistency constraints during the generation process; wherein the content consistency constraints use an asymmetric contrast learning structure to constrain the features of multiple stages in the network, and the specific constraint process is as follows: Since the semantic level of encoder and decoder features is the same at the same stage, we define and The encoder and decoder are l Characteristics of each stage; Sampling and Patches at the same position and passed through two projection heads with shared weights Map them to K -dimensional latent space, we get and , further to Add a prediction head ,get , introduce the stopping gradient, that is , so that the encoder features Remain static during propagation , to prevent it from being updated by the gradient, thereby providing a stable target representation for contrastive learning; finally, a multi-stage content loss is defined, and the features of multiple stages in the network are constrained using an asymmetric contrastive learning structure to ensure that the representation of patch features sampled by the encoder and decoder at the same stage in the latent space is as similar as possible, thereby achieving content consistency; the multi-stage content loss function expression is as follows: In the formula, Represents the dataset X The sample x expectations; represents Euclidean norm normalization; L represents the total number of constraint stages; Q Indicates the total number of patch samples.
3. The method for colorizing an infrared image of a drone based on topological semantic structure loss according to claim 2 is characterized in that: In step S2, the topological structure semantic constraint module is implemented by utilizing the same topological structure of the infrared image and the color image; wherein the patch features of the infrared image are As a graph structure Nodes, patch features of color images As a graph structure nodes and calculate patch features and patch features The cosine similarity between them determines the connection relationship ; The cosine similarity between two patches is greater than or equal to the set threshold t , then the two patches are considered to be related, that is, they are associated nodes in the graph structure; where the adjacency matrix and Features , The relationship is as follows: ; Among them, nodes interact with each other through the adjacency matrix of the graph, thereby capturing local topology-aware semantic information; v Represents the color image node features obtained after graph convolution, which includes the updated information after aggregating the features of adjacent nodes and also reflects the local topological structure; Definition: ; In the formula, For the i The node features of the infrared image of the nodes, For the i Node features of color images of nodes; D is the node degree matrix, indicating the degree of each node, It is shared k Jump representation weight matrix; K Represents the number of hops used in the graph convolution operation.
4. The method for colorizing an infrared image of a drone based on topological semantic structure loss according to claim 3 is characterized in that: In step S2, maximizing mutual information is expressed as: ; In the formula, x represents an infrared image, Represents a color image; N represents the number of nodes in the graph neural network, exp represents exponential calculation, and T represents transpose operation.
5. The method for colorizing an infrared image of a drone based on topological semantic structure loss according to claim 4 is characterized in that: In step S3, the specific steps of the discriminator-guided attention sampling strategy are as follows: S31: Use the discriminator to evaluate the generated image and interpolate the output of the discriminator Adjust to the color image feature map in the generator Decoder Same resolution , get the attention score matrix: In the formula, Indicates interpolating the discriminator output to the resolution ; S32: Sampling based on importance, generating feature maps Uniform sampling patches, including is the oversampling ratio, >1, Q is the total number of patch samples; let the sampling set Represents all oversampled patches, for each sampling patch The corresponding attention score Sort in ascending order, where It's a patch In the feature map The pixel position coordinates in the sorted patches are selected with the highest attention score. patches, including is the important sampling ratio; the selected patch set is expressed as ; S33: Perform uniform sampling, from the remaining uniform sampling in the patch patches, and the set of patches is , then the final selected patch set It can be expressed as: 。 6. The method for colorizing an infrared image of a drone based on topological semantic structure loss according to claim 5, characterized in that: In step S4, the expression of the adversarial loss function is: In the formula, x is the original infrared image, y is a random real color image, is a generator, is the discriminator, Represents the noise distribution X Medium Sample x expectations, Represents the sample in the real data distribution Y y expectations.
7. The method for colorizing an infrared image of a drone based on topological semantic structure loss according to claim 6, characterized in that: In step S4, the expression of the total variation loss function is: ; The final loss function is: ; In the formula, is the topology-aware semantic loss weight.
Citation Information
Patent Citations
Infrared image colorization method and system based on large kernel convolution and graph contrast learning
CN118470153A
Computer Vision Systems and Methods for Diverse Image-to-Image Translation Via Disentangled Representations
US20210224947A1
Cited By
Ship anti-corrosion spraying quality detection method and system based on machine vision
CN120976145A