A high-fidelity infrared remote sensing image generation method, device and equipment

By cropping and pre-training visible light remote sensing images, combining radiation spectral characteristic mapping data and a preset loss function, and using the U-Net architecture and feature fusion module of Pix2Pix GAN, the problem of insufficient realism and reliability in the generation of infrared remote sensing images in existing technologies is solved, and high-fidelity infrared remote sensing images are generated.

CN119600134BActive Publication Date: 2025-11-11XIDIAN UNIV
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202411628225.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-11-14
Publication Date
2025-11-11
Estimated Expiration
2044-11-14

AI Technical Summary

Technical Problem

Existing technologies for generating high-fidelity infrared remote sensing images suffer from poor realism and low simulation efficiency due to physical modeling and simulation methods, while deep learning generation methods suffer from poor interpretability and poor generalization, resulting in insufficient reliability and realism of the generated results.

Method used

By acquiring visible light remote sensing images, cropping them, and then inputting them into a pre-trained infrared image generation network, a high-fidelity infrared remote sensing image is obtained by using the pre-trained image generation module and discrimination module, combined with radiation spectral characteristic mapping data and a preset loss function. The network adopts the U-Net architecture of Pix2Pix GAN, deformable convolution, and multi-scale feature fusion discrimination module, and designs a preset loss function that includes adversarial loss, radiation spectral characteristic loss, and gradient penalty loss.

Benefits of technology

It improves the realism and reliability of infrared remote sensing images, enhances the generalization ability and interpretability of deep learning models, and generates infrared remote sensing images that perform excellently in terms of realism and reliability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119600134B_ABST
    Figure CN119600134B_ABST
Patent Text Reader

Abstract

This invention provides a method, apparatus, and device for generating high-fidelity infrared remote sensing images. The method includes: acquiring a visible light remote sensing image to be processed; cropping the visible light remote sensing image to obtain a cropped visible light remote sensing image; and inputting the cropped visible light remote sensing image into a pre-trained infrared image generation network to generate a corresponding high-fidelity infrared remote sensing image. Specifically, during the training process of the pre-trained infrared image generation network, visible light remote sensing image samples are used as input data, while radiation spectral characteristic mapping data is embedded as prior information into the design of the pre-trained infrared image generation network. A specific loss function is designed to guide the learning direction of the pre-trained infrared image generation network. This allows the interpretability of the physical model to enhance the generalization ability and reliability of the deep learning model, thereby improving the fidelity and reliability of the high-fidelity infrared remote sensing image ultimately generated by the pre-trained infrared image generation network.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of infrared remote sensing image generation technology, specifically to a method, apparatus, and device for generating high-fidelity infrared remote sensing images. Background Technology

[0002] Infrared remote sensing imaging technology has demonstrated broad application potential in various fields such as aerospace, environmental monitoring, urban planning, infrared guidance, and disaster management. These fields have a strong demand for high-fidelity infrared remote sensing image data to support the development of technologies such as remote sensing target detection, image interpretation and evaluation, and computer vision. However, compared to the ease of obtaining visible light remote sensing data, the acquisition of infrared remote sensing images is not only fraught with difficulties but also extremely costly, and the amount of measured data is insufficient to fully meet the needs of measuring the infrared radiation characteristics of high-value targets and conducting algorithm research. Therefore, developing a technical method that can accurately generate high-fidelity infrared remote sensing images from visible light remote sensing images is of great significance for promoting the development of related technologies.

[0003] Currently, methods for generating infrared remote sensing images are mainly divided into two categories: physical modeling simulation generation and deep learning intelligent generation. Physical modeling simulation methods, based on radiative transfer theory and physical laws, offer good physical interpretability and reliability. However, when dealing with complex scenes, they struggle to fully capture the nonlinear and heterogeneous characteristics of the radiation field, leading to discrepancies between simulation results and actual infrared images. Furthermore, this method is inefficient and requires a high level of expertise from operators. Deep learning generation methods, especially generative adversarial networks (GANs), with their powerful nonlinear mapping capabilities and automatic feature learning abilities, can significantly improve simulation efficiency when handling complex scenes. However, due to a lack of high-quality paired data, deep learning models are susceptible to noise and overfitting, reducing the reliability of the generated results. Simultaneously, the black-box nature of deep learning methods limits their interpretability, further increasing the difficulty of generating high-quality infrared remote sensing images.

[0004] In summary, while existing technologies have achieved the generation of infrared remote sensing images by addressing both the remote sensing imaging mechanism and deep learning networks, they have not overcome the inherent contradictions of poor realism and low simulation efficiency in physical modeling and simulation methods, and poor interpretability and generalization in deep learning generation methods. Therefore, the generated high-fidelity infrared remote sensing images still require improvement in terms of realism and reliability. Summary of the Invention

[0005] To address the aforementioned problems in the prior art, the present invention provides a method, apparatus, and device for generating high-fidelity infrared remote sensing images.

[0006] The technical problem to be solved by this invention is achieved through the following technical solution:

[0007] In a first aspect, the present invention provides a method for generating high-fidelity infrared remote sensing images, comprising:

[0008] Acquire the visible light remote sensing image to be processed;

[0009] The visible light remote sensing image to be processed is cropped to obtain the cropped visible light remote sensing image;

[0010] A cropped visible light remote sensing image is input into a pre-trained infrared image generation network to generate a corresponding high-fidelity infrared remote sensing image. The pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module. It is trained using visible light remote sensing image samples as input data and radiation spectral characteristic mapping data and the target infrared remote sensing image as prior information, combined with a preset loss function. The radiation spectral characteristic mapping data represents the spectral characteristic mapping relationship between the visible light remote sensing image sample and the corresponding infrared remote sensing image sample. The preset loss function characterizes the matching relationship between the generated infrared remote sensing image data and the target infrared remote sensing image and the radiation spectral characteristic mapping data.

[0011] Optionally, the visible light remote sensing image to be processed is cropped to obtain a cropped visible light remote sensing image, including:

[0012] The visible light remote sensing image to be processed is cropped according to the preset cropping size to obtain the visible light remote sensing cropped image.

[0013] Optionally, the pre-trained image generation module adopts the U-Net architecture in Pix2Pix GAN, where the U-Net architecture in Pix2Pix GAN has fast Fourier transform processing in the feature calculation process of the corresponding channel attention and spatial attention.

[0014] The pre-trained image discrimination module adopts the discrimination module in Pix2Pix GAN. The discrimination module in Pix2Pix GAN is equipped with a deformable convolution module and a multi-scale feature fusion discrimination module. The deformable convolution module is set between the preset convolutional layers of the discrimination module in Pix2Pix GAN.

[0015] The multi-scale feature fusion and discrimination module is used to perform multi-scale discrimination on infrared remote sensing image generation data and target infrared remote sensing images.

[0016] Optionally, the generation process of the pre-trained infrared image generation network includes:

[0017] Obtain the paired dataset; the paired dataset contains multiple sample pairs of visible light remote sensing image samples and their corresponding infrared remote sensing image samples;

[0018] The spectral reflectance of the paired dataset at each pixel is obtained based on inversion calculation;

[0019] Based on a pre-defined land cover classification method, visible light remote sensing image samples are classified by material to obtain infrared remote sensing image labeled samples.

[0020] The infrared remote sensing image labeled samples and spectral reflectance are matched to obtain material spectral reflectance matching data.

[0021] Based on material spectral reflectance matching data and a preset quantitative characterization model, radiation spectral characteristic mapping data are obtained.

[0022] The initial infrared image generation network is trained based on radiation spectral characteristic mapping data and paired datasets to obtain a pre-trained infrared image generation network; the initial infrared image generation network and the pre-trained infrared image generation network have the same structure.

[0023] Optionally, the initial infrared image generation network is trained based on radiation spectral characteristic mapping data and a paired dataset to obtain a pre-trained infrared image generation network, including:

[0024] The initial infrared image generation network is iteratively trained based on radiation spectral characteristic mapping data, paired datasets, and a preset loss function.

[0025] The initial infrared image generation network that meets the preset stopping conditions is used as the pre-trained infrared image generation network.

[0026] The preset stopping conditions include: the number of iterations meets the preset iteration threshold or the value of the preset loss function continues to converge.

[0027] Optionally, the preset loss function is expressed as:

[0028] Loss=c1L GD +c2Loss IR +c3Loss GP +c4Loss FSM ;

[0029] Where Loss represents the preset loss function, L GD Represents the basic loss function, Loss IR Loss represents the loss of spectral characteristics. GP Represents gradient penalty loss, Loss FSM Let c1 represent the feature space matching loss, c2 represent the first weight value, c3 represent the second weight value, c4 represent the third weight value, and c4 represent the fourth weight value. The base loss function is the sum of the adversarial loss and the conditional loss.

[0030] Optionally, the loss of radiation spectral characteristics is expressed as:

[0031]

[0032] Where i represents the i-th paired dataset, N represents the total number of paired datasets, and I mapping Represents the radiation spectral characteristic mapping data, I input G(I) represents a sample of a visible light remote sensing image. input ) represents the data generated by infrared remote sensing image, and |·| represents taking the absolute value.

[0033] Optionally, the preset quantization representation model is expressed as:

[0034]

[0035] Among them, I mapping (k,λ p ,λ q ,ρ k (m,n) represents the spectral reflectance ρ in a paired dataset of pixel size (m,n). k The material k is from the visible light band λ p to the infrared band λ q The radiation spectral characteristics mapping data; m represents the number of horizontal pixels in the paired dataset, and n represents the number of vertical pixels in the paired dataset; This indicates that material k in the infrared band λ q The corresponding material spectral reflectance matching data, For material k in the visible light band λ p Matching data for the spectral reflectance of the corresponding material at that location; This indicates that in a pixel size of (m k ,n k In the paired dataset, the spectral reflectance is ρ k Material k in the infrared band λ q Total radiation at the location; This indicates that in a pixel size of (m k ,n k In the paired dataset, the spectral reflectance is ρ k Material k in the visible light band λ p Total radiation at the location; m k n represents the number of horizontal pixels of material k in the paired dataset. k This represents the vertical pixel count of material k in the paired dataset. The spectral reflectance is represented by ρ. k Material k in the infrared band λ q The total radiation at that location The spectral reflectance is represented by ρ. k Material k in the visible light band λp The total radiation at that location Indicates the infrared band λ q The material's spectral reflectance matching data is as follows: The total radiation corresponding to the k materials, In the visible light band, λ p The material's spectral reflectance matching data is as follows: The total radiation corresponding to each of the k materials.

[0036] In a second aspect, the present invention provides a high-fidelity infrared remote sensing image generation device, which includes: an acquisition unit, a cropping unit, and an image generation unit.

[0037] The acquisition unit is used to: acquire visible light remote sensing images to be processed;

[0038] The cropping unit is used to: perform image cropping processing on the visible light remote sensing image to be processed, and obtain a cropped visible light remote sensing image;

[0039] The image generation unit is used to: input visible light remote sensing cropped images into a pre-trained infrared image generation network to generate corresponding high-fidelity infrared remote sensing images; wherein, the pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module; the pre-trained infrared image generation network is trained by taking visible light remote sensing image samples as input data and radiation spectral characteristic mapping data as prior information, combined with a preset loss function; the radiation spectral characteristic mapping data is the spectral characteristic mapping relationship between visible light remote sensing image samples and corresponding infrared remote sensing image samples; the preset loss function is used to characterize the matching relationship between the infrared remote sensing image generation data and the target infrared remote sensing image and radiation spectral characteristic mapping data.

[0040] Thirdly, the present invention provides a high-fidelity infrared remote sensing image generation device, comprising: a processor, a storage medium and a bus, wherein the storage medium stores machine-readable instructions executable by the processor, and when the high-fidelity infrared remote sensing image generation device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the high-fidelity infrared remote sensing image generation method of the first aspect described above.

[0041] This invention provides a method, apparatus, and device for generating high-fidelity infrared remote sensing images. The method includes: acquiring a visible light remote sensing image to be processed; cropping the visible light remote sensing image to obtain a cropped visible light remote sensing image; and inputting the cropped visible light remote sensing image into a pre-trained infrared image generation network to generate a corresponding high-fidelity infrared remote sensing image. The pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module. The network is trained using visible light remote sensing image samples as input data and radiation spectral characteristic mapping data and a target infrared remote sensing image as prior information, combined with a preset loss function. The radiation spectral characteristic mapping data represents the spectral characteristic mapping relationship between visible light remote sensing image samples and corresponding infrared remote sensing image samples. The preset loss function characterizes the matching relationship between the generated infrared remote sensing image data and the target infrared remote sensing image and the radiation spectral characteristic mapping data. In this invention, during the training process of the pre-trained infrared image generation network, visible light remote sensing image samples are used as input data, while radiation spectral characteristic mapping data and target infrared remote sensing images are used as prior information and embedded into the design of the pre-trained infrared image generation network. A specific loss function is designed to guide the learning direction of the pre-trained infrared image generation network. This allows the interpretability of the physical model to enhance the generalization ability and reliability of the deep learning model, ultimately improving the realism and reliability of the high-fidelity infrared remote sensing images generated by the pre-trained infrared image generation network.

[0042] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description

[0043] Figure 1 A flowchart illustrating a high-fidelity infrared remote sensing image generation method provided in an embodiment of the present invention;

[0044] Figure 2 A schematic diagram of the structure of a pre-trained image generation module is shown as an example;

[0045] Figure 3 An exemplary schematic diagram of the structure of a pre-trained image discrimination module with added deformable convolutional modules is shown;

[0046] Figure 4 An exemplary diagram of the architecture of a pre-trained infrared image generation network after adding a multi-scale feature fusion discrimination module is shown.

[0047] Figure 5 This is a schematic diagram of the structure of a high-fidelity infrared remote sensing image generation device provided in an embodiment of the present invention;

[0048] Figure 6This is a schematic diagram of the structure of a high-fidelity infrared remote sensing image generation device provided in an embodiment of the present invention. Detailed Implementation

[0049] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.

[0050] To improve the realism and reliability of the high-fidelity infrared remote sensing images generated by the pre-trained infrared image generation network, this invention provides a method for generating high-fidelity infrared remote sensing images. Figure 1 This is a flowchart illustrating a high-fidelity infrared remote sensing image generation method provided in an embodiment of the present invention, as shown below. Figure 1 As shown, it includes:

[0051] S101. Acquire the visible light remote sensing image to be processed.

[0052] Visible light remote sensing images to be processed refer to images acquired from ground or aerial platforms (such as satellites, drones, etc.) that reflect the reflection of sunlight on the Earth's surface. These images typically contain rich information about ground features, such as vegetation, water bodies, and buildings. In this embodiment of the invention, such images are crucial for generating high-quality infrared images.

[0053] S102. Perform image cropping processing on the visible light remote sensing image to be processed to obtain a cropped visible light remote sensing image.

[0054] Understandably, the visible light remote sensing images acquired directly typically have very high resolution and wide scene coverage. Processing the entire image directly could require significant computational resources and time. Cropping the image reduces the processing scope, thereby decreasing computational complexity and improving efficiency. Furthermore, cropping ensures that the images input to the pre-trained infrared image generation network conform to the model's input specifications, avoiding size mismatch issues during processing.

[0055] Optionally, S102 may specifically include:

[0056] The visible light remote sensing image to be processed is cropped according to the preset cropping size to obtain the visible light remote sensing cropped image.

[0057] It should be noted that the above-mentioned preset cropping size is the size information that meets the input of the pre-trained infrared image generation network model.

[0058] S103. Input the visible light remote sensing cropped image into the pre-trained infrared image generation network to generate the corresponding high-fidelity infrared remote sensing image.

[0059] The pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module. The network takes visible light remote sensing image samples as input data and uses radiation spectral characteristic mapping data and the target infrared remote sensing image as prior information, and is trained together with a preset loss function. The radiation spectral characteristic mapping data represents the spectral characteristic mapping relationship between visible light remote sensing image samples and their corresponding infrared remote sensing image samples. The preset loss function characterizes the matching relationship between the generated infrared remote sensing image data and the target infrared remote sensing image and the radiation spectral characteristic mapping data.

[0060] This invention provides a method for generating high-fidelity infrared remote sensing images, comprising: during the training process of a pre-trained infrared image generation network, using visible light remote sensing image samples as input data, and simultaneously embedding radiation spectral characteristic mapping data and target infrared remote sensing images as prior information into the design of the pre-trained infrared image generation network, and designing a specific loss function to guide the learning direction of the pre-trained infrared image generation network. This can enhance the generalization ability and reliability of deep learning models by utilizing the interpretability of physical models, ultimately improving the fidelity and reliability of the high-fidelity infrared remote sensing images generated by the pre-trained infrared image generation network.

[0061] Optionally, the pre-trained image generation module adopts the U-Net architecture in Pix2Pix GAN, where the U-Net architecture in Pix2Pix GAN has fast Fourier transform processing in the feature calculation process of the corresponding channel attention and spatial attention.

[0062] The pre-trained image discrimination module adopts the discrimination module in Pix2Pix GAN. The discrimination module in Pix2Pix GAN is equipped with a deformable convolution module and a multi-scale feature fusion discrimination module. The deformable convolution module is set between the preset convolutional layers of the discrimination module in Pix2Pix GAN.

[0063] The multi-scale feature fusion and discrimination module is used to perform multi-scale discrimination on infrared remote sensing image generation data and target infrared remote sensing images.

[0064] Figure 2 An exemplary schematic diagram of the pre-trained image generation module is shown. For example... Figure 2 As shown, the pre-trained image generation module provided in this embodiment of the invention adopts the U-Net architecture, and sets up a Fast Fourier Transform channel spatial attention mechanism on the existing U-Net architecture in Pix2Pix GAN. That is, Fast Fourier Transform processing is set in the feature calculation process of the corresponding channel attention and spatial attention, such as... Figure 2As shown in the FFCSAM module, after processing by the pre-trained image generation module, a 3*256*256 dimensional feature map is finally obtained. Figure 3 An exemplary schematic diagram of the structure of a pre-trained image discrimination module with added deformable convolutional modules is shown, such as... Figure 3 As shown, this embodiment of the invention adds deformable convolution to the discriminant module in the existing Pix2Pix GAN, such as... Figure 3 As shown in the green section, in addition Figure 3 The discrimination module also includes a feature map module, a confidence map module, downsampling, and fusion convolution processing. Figure 3 The placement of deformable convolutions is merely illustrative; the specific placement can be flexibly set according to requirements, and this embodiment of the invention does not limit this. Furthermore, both the feature map module and the confidence map module are implemented using different convolutional layers. Figure 4 An exemplary diagram of the architecture of a pre-trained infrared image generation network with the addition of a multi-scale feature fusion discrimination module is shown. Figure 4 As shown, the pre-trained infrared image generation network performs comparative discrimination on three scales. First, the input visible light remote sensing image samples are downsampled to three dimensions. Then, the generated infrared remote sensing image data and the target infrared remote sensing image (real data) are downsampled to the corresponding three dimensions. Correspondence comparison is performed on each dimension. The comparators on the three dimensions are called discriminator 1, discriminator 2, and discriminator 3, respectively.

[0065] It should be noted that by incorporating Fast Fourier Transform (FFT) processing into the feature calculation of channel attention and spatial attention within the U-Net architecture, the learning ability of the pre-trained infrared image generation network to image features can be enhanced. This helps the pre-trained infrared image generation network extract richer information from the input visible light remote sensing images, providing strong support for generating high-quality infrared remote sensing images.

[0066] The discriminator module of Pix2Pix GAN is itself a powerful image discrimination tool, effectively distinguishing between real and generated images. By implementing deformable convolutional modules, the discriminator module can more flexibly capture features from visible light remote sensing images, thereby improving discrimination accuracy. The introduction of a multi-scale feature fusion discriminator module enables multi-scale discrimination of both infrared remote sensing image generation data and target infrared remote sensing images. This design helps the pre-trained infrared image generation network better adapt to image inputs of different scales, enhancing the model's robustness.

[0067] Furthermore, the deformable convolution module enables the discrimination module to process image data more efficiently, reducing computational redundancy. Simultaneously, the multi-scale feature fusion discrimination module helps the pre-trained infrared image generation network learn key features from visible light remote sensing image samples more quickly, thereby accelerating the training process.

[0068] Optionally, the generation process of the pre-trained infrared image generation network includes:

[0069] Obtain the paired dataset; the paired dataset contains multiple sample pairs of visible light remote sensing image samples and their corresponding infrared remote sensing image samples;

[0070] The spectral reflectance of the paired dataset at each pixel is obtained based on inversion calculation;

[0071] Based on a pre-defined land cover classification method, visible light remote sensing image samples are classified by material to obtain infrared remote sensing image labeled samples.

[0072] The infrared remote sensing image labeled samples and spectral reflectance are matched to obtain material spectral reflectance matching data.

[0073] Based on material spectral reflectance matching data and a preset quantitative characterization model, radiation spectral characteristic mapping data are obtained.

[0074] The initial infrared image generation network is trained based on radiation spectral characteristic mapping data and paired datasets to obtain a pre-trained infrared image generation network; the initial infrared image generation network and the pre-trained infrared image generation network have the same structure.

[0075] In this embodiment of the invention, Landsat8 data can be used for experiments. Additionally, preprocessing operations such as cropping, rotating, and mirroring can be performed on the paired dataset to meet the needs of subsequent network training. The image size in the paired dataset of this invention is generally set to 256*256 pixels, and there are a total of 6000 pairs in the paired dataset.

[0076] The simulation process of infrared remote sensing images requires radiometric calculations or mapping transformations that conform to the characteristics of different landforms. Therefore, the visible light remote sensing image samples can first be classified and numbered. This can be done in the following two ways, with the following preset landform classification methods:

[0077] (1) Manually supervised / semi-supervised classification can be mainly completed using Photoshop and ENVI software;

[0078] (2) Deep learning intelligent classification can be accomplished using classification networks and label data.

[0079] In this embodiment of the invention, reflectance information of infrared remote sensing images in multiple spectral bands can be obtained through inversion calculation; specifically, each type of satellite has its own corresponding inversion calculation formula, which will not be elaborated upon in the comparison of this embodiment of the invention.

[0080] In addition, after obtaining the spectral reflectance, it can be compared with the information in the spectral reflectance database to obtain the spectral reflectance curve with the highest matching degree for each material, which can be used as the spectral reflectance of that material.

[0081] Optionally, the initial infrared image generation network is trained based on radiation spectral characteristic mapping data and a paired dataset to obtain a pre-trained infrared image generation network, including:

[0082] The initial infrared image generation network is iteratively trained based on radiation spectral characteristic mapping data, paired datasets, and a preset loss function.

[0083] The initial infrared image generation network that meets the preset stopping conditions is used as the pre-trained infrared image generation network.

[0084] The preset stopping conditions include: the number of iterations meets the preset iteration threshold or the value of the preset loss function continues to converge.

[0085] In this embodiment of the invention, based on a deep understanding of the radiation mechanism of a scene and the design of the network architecture, a data-driven approach and physical knowledge are effectively combined through radiation spectral characteristic loss. This guides the initial infrared image generation network to follow the real radiative transfer laws during the learning process, helping it better understand the radiative transfer characteristics between visible light and infrared light, thereby enhancing the training effect and generalization ability of the initial infrared image generation network. Simultaneously, leveraging the powerful learning ability of the initial infrared image generation network, it improves the accurate representation of the complex random distribution, multipath scattering, and nonlinear changes of the radiation field between visible light and infrared remote sensing images, as well as the representation of image texture features, thereby improving the accuracy and detail of feature mapping. Specifically, the loss function formed by the mapping relationship of radiation spectral characteristics from visible light to infrared remote sensing (radiation spectral characteristic mapping data) is used as the entry point and quantitatively coupled into the network model to finally construct the preset loss function.

[0086] The default loss function is expressed as:

[0087] Loss=c1L GD +c2Loss IR +c3Loss GP +c4Loss FSM ;

[0088] Where Loss represents the preset loss function, L GD Represents the basic loss function, LossIR Loss represents the loss of spectral characteristics. GP Represents gradient penalty loss, Loss FSM Let c1 represent the feature space matching loss, c2 represent the first weight value, c3 represent the second weight value, c4 represent the third weight value, and c4 represent the fourth weight value. The base loss function is the sum of the adversarial loss and the conditional loss.

[0089] Specifically, in the implementation of this embodiment, a preset loss function, which includes the basic loss function, the radiation spectral characteristic loss, the gradient penalty loss, and the feature space matching loss, can be used as the overall loss function in the training process of the pre-trained infrared image generation network.

[0090] Optionally, the loss of radiation spectral characteristics is expressed as:

[0091]

[0092] Where i represents the i-th paired dataset, N represents the total number of paired datasets, and I mapping Represents the radiation spectral characteristic mapping data, I input G(I) represents a sample of a visible light remote sensing image. input ) represents the data generated by infrared remote sensing image, and |·| represents taking the absolute value.

[0093] Optionally, the preset quantization representation model is expressed as:

[0094]

[0095] Among them, I mapping (k,λ p ,λ q ,ρ k (m,n) represents the spectral reflectance ρ in a paired dataset of pixel size (m,n). k The material k is from the visible light band λ p to the infrared band λ q The radiation spectral characteristics mapping data; m represents the number of horizontal pixels in the paired dataset, and n represents the number of vertical pixels in the paired dataset; This indicates that material k in the infrared band λ q The corresponding material spectral reflectance matching data, For material k in the visible light band λ p Matching data for the spectral reflectance of the corresponding material at that location; This indicates that in a pixel size of (m k ,n k In the paired dataset, the spectral reflectance is ρ k Material k in the infrared band λ q Total radiation at the location; This indicates that in a pixel size of (m k ,n k In the paired dataset, the spectral reflectance is ρ k Material k in the visible light band λ p Total radiation at the location; m k n represents the number of horizontal pixels of material k in the paired dataset. k This represents the vertical pixel count of material k in the paired dataset. The spectral reflectance is represented by ρ. k Material k in the infrared band λ q The total radiation at that location The spectral reflectance is represented by ρ. k Material k in the visible light band λ p The total radiation at that location Indicates the infrared band λ q The material's spectral reflectance matching data is as follows: The total radiation corresponding to the k materials, In the visible light band, λ p The material's spectral reflectance matching data is as follows: The total radiation corresponding to each of the k materials.

[0096] It should be noted that in this embodiment, the calculation of total radiation can be based on the scene's radiation energy transmission law, including reflected radiation, self-radiation, and the reflectivity information of ground material (as described in the aforementioned preset quantitative characterization model). The solar radiation, sky background radiation, atmospheric path radiation, and atmospheric transmittance involved in reflected radiation can be calculated using the Modtran model. Self-radiation can be calculated using Planck's formula.

[0097] To verify the effectiveness of the high-fidelity infrared remote sensing image generation method provided in this embodiment of the invention, simulation experiments were also conducted. The results are as follows:

[0098] Experimental Setup: All experiments in this invention were conducted on a Windows operating system. The CPU used was an Intel(R) Core(TM) i7-10700F, the GPU was an NVIDIA GeForce RTX 4070 with 16GB of VRAM, the programming language was Python, CUDA Version: 11.1, and Cu DNN Version: 8.9.7. The deep learning framework used to build the network was PyTorch Version: 2.3.1, with an L1 distance loss coefficient of 100, a physical loss coefficient of 100, a perceptual loss coefficient of 10, and a gradient penalty loss coefficient of 10. In the image preprocessing stage, the images were first cropped to 256*256 pixels, the batch size was set to 1, and the Adam optimizer was used. The initial learning rate was 0.0002 for the first 100 rounds, and then the learning rate linearly decreased from 0.0002 to 0 for the next 100 rounds.

[0099] Evaluation metrics are determined as follows: This embodiment of the invention uses two metrics, Peak Signal-to-Noise Ratio (PSNR) and Structural Similarity to Image (SSIM), to evaluate and analyze the generated results. PSNR measures the pixel-level difference between the generated image and the real target image; a higher PSNR value indicates that the pixel values ​​of the generated image are closer to the real target image. SSIM, in measuring image quality, is more consistent with human visual perception and can measure the structural or brightness differences between the generated image and the real target image; a higher SSIM value indicates greater structural similarity between the two images.

[0100] Specifically, a series of ablation experiments were conducted using the original Pix2PixGAN with a supervised learning strategy as a benchmark. The experimental results for the shortwave high-fidelity infrared remote sensing image generation task are shown in Table 1. Experiment 1: The supervised image transfer model Pix2PixGAN was trained with default parameters and 1800 sets of images, and the index values ​​PSNR and SSIM were calculated. Experiment 2: Based on Experiment 1, a method based on deformable convolution + multi-scale feature fusion discrimination was adopted, and the evaluation index results were calculated. Experiment 3: Based on Experiment 2, a channel attention + spatial attention mechanism was designed, and the evaluation index results were calculated. Experiment 4: Based on Experiment 3, radiation spectral characteristic mapping data was introduced, and an improved preset loss function was designed, namely the high-fidelity infrared remote sensing image generation method with coupled scene radiation mechanism and generative adversarial network proposed in this invention, and the evaluation index results were calculated.

[0101] Table 1 Comparative experimental results for the shortwave high-fidelity infrared remote sensing image generation task.

[0102]

[0103]

[0104] As can be seen from the experimental results in Table 1 above, the high-fidelity infrared remote sensing images obtained based on the method of this invention achieve the best results in both SSIM and PSNR evaluation metrics. That is, the infrared remote sensing images generated using this method have the highest fidelity and the best texture detail. Furthermore, this method demonstrates high accuracy and strong adaptability in generating high-fidelity infrared remote sensing images in short, medium, and long wavelength bands. This proves that the method of this invention can effectively improve the accuracy, confidence, and generalization ability of the base model.

[0105] In summary, the high-fidelity infrared remote sensing image generation method proposed in this invention has the following technological innovations:

[0106] 1. Based on the scene radiation energy transmission law, and utilizing the material spectral reflectance characteristics and the spectral dependence of radiation field energy, a quantitative characterization model of the mapping relationship between visible light remote sensing and infrared remote sensing radiation spectral characteristics was constructed. This model was then quantitatively coupled with the final preset loss function. By utilizing the interpretability of physical modeling and the powerful feature extraction capability of network model, the accuracy and generalization of the pre-trained infrared image generation network were improved.

[0107] 2. Based on the Pix2PixGAN model, a fast Fourier transform process is designed to be set in the feature calculation process of channel attention and spatial attention. Starting from both channel attention and spatial attention, parameters and computing power are saved. At the same time, the frequency domain features of the learning image can be extracted, which can help the pre-trained image generation module to learn the comprehensive texture features of the image more efficiently, and improve the generator's ability to extract features and transform them.

[0108] 3. Based on the Pix2PixGAN model, by designing a deformable convolution module and a multi-scale feature fusion discrimination module, the adaptability and generalization ability of the pre-trained image discrimination module to unknown changes are improved. In this way, the details and features of visible light remote sensing images at different scales can be captured, thereby improving the accuracy of the pre-trained image discrimination module.

[0109] 4. A pre-defined loss function was designed, which includes adversarial loss, conditional loss, radiation spectral characteristic loss, gradient penalty loss, and feature space matching loss. This effectively promotes the convergence and balance of the pre-trained infrared image generation network and improves its stability and generalization ability.

[0110] The method provided in this embodiment of the invention can be applied to electronic devices. Specifically, the electronic device can be a desktop computer, a portable computer, a smart mobile terminal, a server, etc., and this embodiment of the invention does not limit the application to such devices.

[0111] Based on the same inventive concept, embodiments of the present invention also provide a high-fidelity infrared remote sensing image generation device. Figure 5 This is a schematic diagram of the structure of a high-fidelity infrared remote sensing image generation device provided in an embodiment of the present invention, as shown below. Figure 5 As shown, it includes: an acquisition unit 501, a cropping unit 502, and an image generation unit 503;

[0112] The acquisition unit 501 is used to: acquire a visible light remote sensing image to be processed;

[0113] The cropping unit 502 is used to: perform image cropping processing on the visible light remote sensing image to be processed, and obtain a visible light remote sensing cropped image;

[0114] The image generation unit 503 is used to: input a cropped visible light remote sensing image into a pre-trained infrared image generation network to generate a corresponding high-fidelity infrared remote sensing image; wherein, the pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module; the pre-trained infrared image generation network is trained by taking visible light remote sensing image samples as input data and radiation spectral characteristic mapping data as prior information, combined with a preset loss function; the radiation spectral characteristic mapping data is the spectral characteristic mapping relationship between visible light remote sensing image samples and corresponding infrared remote sensing image samples; the preset loss function is used to characterize the matching relationship between the infrared remote sensing image generation data and the target infrared remote sensing image and the radiation spectral characteristic mapping data.

[0115] Figure 6 A schematic diagram of a high-fidelity infrared remote sensing image generation device provided in an embodiment of the present invention includes: a processor 610, a storage medium 620, and a bus 630. The storage medium 620 stores machine-readable instructions executable by the processor 610. When the high-fidelity infrared remote sensing image generation device is running, the processor 610 and the storage medium 620 communicate via the bus 630. The processor 610 executes the machine-readable instructions to perform the steps of the above-described method embodiment. Specific implementation methods and technical effects are similar and will not be described in detail here.

[0116] The storage medium may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device. Optionally, the storage medium may also be at least one storage device located remotely from the aforementioned processor.

[0117] The processors mentioned above can be general-purpose processors, including central processing units (CPUs), network processors (NPs), etc.; they can also be digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0118] It should be noted that the terms "first," "second," etc., are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of the invention described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of apparatuses and methods consistent with some aspects of the invention.

[0119] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Furthermore, those skilled in the art can combine and integrate the different embodiments or examples described in this specification.

[0120] Although the invention has been described herein in conjunction with various embodiments, those skilled in the art, by reviewing the accompanying drawings and the disclosure, will understand and implement other variations of the disclosed embodiments in carrying out the claimed invention. In the description of the invention, the word "comprising" does not exclude other components or steps, "a" or "an" does not exclude a plurality, and "a plurality" means two or more, unless otherwise explicitly specified. Furthermore, while different embodiments may describe certain measures, this does not mean that these measures cannot be combined to produce good results.

[0121] The above description, in conjunction with specific preferred embodiments, provides a further detailed explanation of the present invention. It should not be construed that the specific implementation of the present invention is limited to these descriptions. For those skilled in the art, various simple deductions or substitutions can be made without departing from the inventive concept, and all such modifications and substitutions should be considered within the scope of protection of the present invention.

Claims

1. A method for generating high-fidelity infrared remote sensing images, characterized in that, include: Acquire the visible light remote sensing image to be processed; The visible light remote sensing image to be processed is cropped to obtain a cropped visible light remote sensing image; The cropped visible light remote sensing image is input into a pre-trained infrared image generation network to generate a corresponding high-fidelity infrared remote sensing image. The pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module. The network is trained using visible light remote sensing image samples as input data and radiation spectral characteristic mapping data and the target infrared remote sensing image as prior information, combined with a preset loss function. The radiation spectral characteristic mapping data represents the spectral characteristic mapping relationship between the visible light remote sensing image samples and the corresponding infrared remote sensing image samples. The preset loss function characterizes the matching relationship between the infrared remote sensing image generation data and the target infrared remote sensing image and the radiation spectral characteristic mapping data. The generation process of the pre-trained infrared image generation network includes: Obtain a paired dataset; the paired dataset contains multiple sample pairs of visible light remote sensing image samples and corresponding infrared remote sensing image samples; The spectral reflectance of the paired dataset at each pixel is obtained based on inversion calculation; The visible light remote sensing image samples are classified by material based on a preset land cover classification method to obtain infrared remote sensing image labeled samples. The infrared remote sensing image labeled samples and the spectral reflectance are matched to obtain material spectral reflectance matching data. Based on the material spectral reflectance matching data and the preset quantization characterization model, the radiation spectral characteristic mapping data is obtained; The initial infrared image generation network is trained based on the radiation spectral characteristic mapping data and the paired dataset to obtain the pre-trained infrared image generation network; the initial infrared image generation network has the same structure as the pre-trained infrared image generation network.

2. The high-fidelity infrared remote sensing image generation method according to claim 1, characterized in that, The step of cropping the visible light remote sensing image to obtain a cropped visible light remote sensing image includes: The visible light remote sensing image to be processed is cropped according to a preset cropping size to obtain the cropped visible light remote sensing image.

3. The high-fidelity infrared remote sensing image generation method according to claim 1, characterized in that, The pre-trained image generation module adopts the U-Net architecture in Pix2Pix GAN, wherein the U-Net architecture in Pix2Pix GAN is equipped with fast Fourier transform processing in the feature calculation process of the corresponding channel attention and spatial attention. The pre-trained image discrimination module adopts the discrimination module in Pix2Pix GAN, which is equipped with a deformable convolution module and a multi-scale feature fusion discrimination module. The deformable convolutional module is disposed between the preset convolutional layers of the discriminant module in the Pix2Pix GAN; The multi-scale feature fusion and discrimination module is used to perform multi-scale discrimination on the infrared remote sensing image generation data and the target infrared remote sensing image.

4. The method for generating high-fidelity infrared remote sensing images according to claim 1, characterized in that, The step of training the initial infrared image generation network based on the radiation spectral characteristic mapping data and the paired dataset to obtain the pre-trained infrared image generation network includes: The initial infrared image generation network is iteratively trained based on the radiation spectral characteristic mapping data, the paired dataset, and the preset loss function. The initial infrared image generation network that meets the preset stopping conditions is used as the pre-trained infrared image generation network; The preset stopping conditions include: the number of iterations meets a preset iteration threshold or the value of the preset loss function continues to converge.

5. The high-fidelity infrared remote sensing image generation method according to claim 4, characterized in that, The preset loss function is expressed as follows: ; in, Indicates the preset loss function. Represents the basic loss function. This indicates the loss of spectral characteristics. This represents the gradient penalty loss. Represents the feature space matching loss. This represents the first weight value. This represents the second weight value. This represents the third weight value. This represents the fourth weight value, where the basic loss function is the sum of the adversarial loss and the conditional loss.

6. The high-fidelity infrared remote sensing image generation method according to claim 5, characterized in that, The loss of the radiation spectral characteristics is expressed as: ; in, Indicates the first A pairing dataset, This represents the total number of the paired datasets. This represents the data mapping the spectral characteristics of radiation. This represents a sample of visible light remote sensing images. This represents the infrared remote sensing image generation data. This indicates taking the absolute value.

7. The method for generating high-fidelity infrared remote sensing images according to claim 1, characterized in that, The preset quantitative representation model is expressed as follows: ; in, Indicates a pixel size of In the paired dataset, the spectral reflectance is material From the visible light band To the infrared band Radiative spectral characteristics mapping data; This represents the number of pixels in the horizontal direction of the paired dataset. This represents the number of vertical pixels in the paired dataset; Indicates material In the infrared band The corresponding material spectral reflectance matching data, For material In the visible light band Matching data for the spectral reflectance of the corresponding material at that location; Indicates a pixel size of In the paired dataset, the spectral reflectance is material In the infrared band Total radiation at the location; Indicates a pixel size of In the paired dataset, the spectral reflectance is material In the visible light band Total radiation at the location; Represents the materials in the paired dataset The number of horizontal pixels, Represents the materials in the paired dataset The number of vertical pixels, Indicates spectral reflectance material In the infrared band The total radiation at that location Indicates spectral reflectance material In the visible light band The total radiation at that location Indicating in the infrared band The material's spectral reflectance matching data is as follows: of The total radiation corresponding to each material. Indicates the visible light band The material's spectral reflectance matching data is as follows: of The total radiation corresponding to each material.

8. A high-fidelity infrared remote sensing image generation device, characterized in that, The high-fidelity infrared remote sensing image generation device includes: an acquisition unit, a cropping unit, an image generation unit, and a model training unit; The acquisition unit is used to: acquire a visible light remote sensing image to be processed; The cropping unit is used to: perform image cropping processing on the visible light remote sensing image to be processed, to obtain a visible light remote sensing cropped image; The image generation unit is used to: input the visible light remote sensing cropped image into a pre-trained infrared image generation network to generate a corresponding high-fidelity infrared remote sensing image; wherein, the pre-trained infrared image generation network consists of a pre-trained image generation module and a pre-trained image discrimination module; the pre-trained infrared image generation network is trained using visible light remote sensing image samples as input data, and using radiation spectral characteristic mapping data and the target infrared remote sensing image as prior information, combined with a preset loss function; the radiation spectral characteristic mapping data is the spectral characteristic mapping relationship between the visible light remote sensing image sample and the corresponding infrared remote sensing image sample; the preset loss function is used to characterize the matching relationship between the infrared remote sensing image generation data and the target infrared remote sensing image and the radiation spectral characteristic mapping data. The model training unit is used for: acquiring a paired dataset; the paired dataset contains multiple sample pairs of visible light remote sensing image samples and corresponding infrared remote sensing image samples; obtaining the spectral reflectance of the paired dataset at each pixel based on inversion calculation; classifying the visible light remote sensing image samples into materials based on a preset land cover classification method to obtain infrared remote sensing image labeled samples; matching the infrared remote sensing image labeled samples and the spectral reflectance to obtain material spectral reflectance matching data; obtaining the radiation spectral characteristic mapping data based on the material spectral reflectance matching data and a preset quantization characterization model; training an initial infrared image generation network based on the radiation spectral characteristic mapping data and the paired dataset to obtain the pre-trained infrared image generation network; the initial infrared image generation network and the pre-trained infrared image generation network have the same structure.

9. A high-fidelity infrared remote sensing image generation device, characterized in that, include: The device includes a processor, a storage medium, and a bus. The storage medium stores machine-readable instructions executable by the processor. When the high-fidelity infrared remote sensing image generation device is running, the processor communicates with the storage medium via the bus, and the processor executes the machine-readable instructions to perform the steps of the high-fidelity infrared remote sensing image generation method as described in any one of claims 1-7.

Citation Information

Cited By

  • Cross-modal infrared remote sensing image intelligent calculation fusion generation method, device and equipment

    CN122199279A