Infrared image confrontation generation method based on physical quantity and semantic mask guidance

By constructing a multi-channel fusion tensor and introducing a boot module into the generator network for feature modulation, the problem of insufficient physical consistency and detail fidelity in the existing infrared image generation methods is solved, and high-quality infrared images that conform to physical laws are generated.

CN120431201APending Publication Date: 2025-08-05XIDIAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510475702.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-16
Publication Date
2025-08-05

AI Technical Summary

Technical Problem

The existing infrared image generation methods based on deep learning have shortcomings in the physical consistency and detail fidelity of generated images, and it is difficult to generate high-quality infrared images that conform to actual physical laws.

Method used

By constructing object categories, material categories, emissivity and temperature matrices, multi-channel fusion tensors are generated, and a guide module is introduced into the generator network of infrared image adversity generation models for layer-by-layer feature modulation, embedding physical quantities and semantic features to ensure that the generated infrared image conforms to the physical radiation laws.

Benefits of technology

The generated infrared images can truly reflect the thermal radiation characteristics of the target, have sufficient details and follow physical laws, improving the quality and accuracy of the generated images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120431201A_ABST
    Figure CN120431201A_ABST
Patent Text Reader

Abstract

The invention particularly relates to an infrared image confrontation generation method based on physical quantity and semantic mask guidance. The method comprises the steps that physical quantity parameters and semantic parameters to be processed are acquired; constructing a corresponding object category matrix, a material category matrix, an emissivity matrix and a temperature matrix; the object category matrix, the material category matrix, the emissivity matrix and the temperature matrix are coded and spliced along the channel dimension, and a fusion tensor is obtained; inputting the fusion tensor into a generator network of a trained infrared image adversarial generative model, and performing layer-by-layer up-sampling processing on the fusion tensor to obtain a target infrared image; wherein the generator network comprises a plurality of consecutive up-sampling layers configured with boot modules; and the guiding module is used for calculating corresponding modulation parameters according to the input characteristics of the corresponding upper sampling layer, and performing characteristic modulation on the input characteristics of the layer by utilizing the modulation parameters so as to embed physical quantity and semantic characteristics. According to the method, the infrared image which has sufficient details and conforms to the physical law can be generated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of computer perspective technology, and in particular to an infrared image adversarial generation method guided by physical quantities and semantic masks, a storage medium, a computer program product, and an electronic device. Background Art

[0002] Infrared images are an important form of visual information in related technologies and are widely used in industry, medicine, military and other fields. Infrared images are generated by detecting the thermal radiation information of objects and can obtain the thermal characteristic information of the target in low-light or completely dark environments. Compared with visible light images, infrared images have unique thermal radiation characteristics and can reflect the temperature distribution and thermal characteristics of the target. The generation and processing of infrared images provide important support for tasks such as target detection, temperature monitoring and anomaly analysis. However, how to generate simulated images that are highly similar to real infrared images remains a key technical challenge. Existing deep learning-based infrared image generation methods usually use generative adversarial network technology, but there is also the problem of low quality of generated images.

[0003] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present invention, and therefore may include information that does not constitute prior art known to ordinary technicians in this field. Summary of the Invention

[0004] The present invention provides an infrared image adversarial generation method and system based on physical quantity and semantic mask guidance, a computer program product, and an electronic device, which can overcome the defects in the prior art to a certain extent.

[0005] Other features and advantages of the present invention will become apparent from the following detailed description, or may be learned in part by practice of the present invention.

[0006] According to a first aspect of the present invention, a method for adversarial generation of infrared images guided by physical quantities and semantic masks is provided, the method comprising:

[0007] Obtaining physical quantity parameters and semantic parameters to be processed; and constructing corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix for the physical quantity parameters and semantic parameters; wherein the physical quantity parameters include: emissivity, temperature; the semantic parameters include: object type, material type;

[0008] Encoding the object category matrix, material category matrix, emissivity matrix, and temperature matrix respectively and concatenating them along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities;

[0009] The fused tensor is input into the generator network of the trained infrared image adversarial generative model, and the fused tensor is upsampled layer by layer to obtain the target infrared image; wherein the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to feature modulate the input features of the layer to embed physical quantities and semantic features.

[0010] In some exemplary embodiments, encoding the object category matrix, the material category matrix, the emissivity matrix, and the temperature matrix separately and concatenating them along the channel dimension includes:

[0011] Add an offset to the material matrix and re-encode it to obtain the re-encoded material matrix; perform one-hot encoding on the re-encoded material matrix to obtain the corresponding material tensor;

[0012] Based on the preset weighting coefficients, weighted processing is performed on the emissivity matrix and the temperature matrix respectively;

[0013] The material tensor, the weighted emissivity matrix, and the temperature matrix are concatenated along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities.

[0014] In some exemplary embodiments, the modulation parameters include a scaling parameter and an offset parameter having spatial dimensions.

[0015] In some exemplary embodiments, inputting the fused tensor into a generator network of a trained infrared image adversarial generative model, and performing layer-by-layer upsampling on the fused tensor to obtain a target infrared image includes:

[0016] In an upsampling layer, upsampling and normalizing the input features of the upsampling layer are performed in sequence to obtain first feature data;

[0017] The guiding module performs convolution processing on the input features of the upsampling layer to obtain corresponding second feature data; and performs two-dimensional convolution processing on the second feature data to obtain corresponding modulation parameters;

[0018] Modulating the first characteristic data using a modulation parameter to obtain a modulation characteristic;

[0019] The modulation feature is fused with the first feature data to obtain the output feature of the upsampling layer; wherein the output feature is configured as the input feature of the next adjacent upsampling layer.

[0020] In some exemplary embodiments, the matrix size of the modulation parameters is the same as the size of the output features of the corresponding upsampling layer.

[0021] In some exemplary embodiments, inputting the fused tensor into a generator of a trained infrared image adversarial generative model and performing layer-by-layer upsampling on the fused tensor to obtain a target infrared image includes:

[0022] In an upsampling layer, feature extraction and normalization are performed on the input features of the upsampling layer in sequence to obtain third feature data;

[0023] The guiding module performs convolution processing on the input features of the upsampling layer to obtain corresponding second feature data; and performs two-dimensional convolution processing on the second feature data to obtain corresponding modulation parameters;

[0024] modulating the third characteristic data using a modulation parameter to obtain a modulation characteristic;

[0025] The modulation feature and the third feature data are subjected to feature fusion processing to obtain fused feature data; and the fused feature data is subjected to upsampling processing to obtain the output feature of the upsampling layer; wherein the output feature is configured as the input feature of the next adjacent upsampling layer.

[0026] In some exemplary embodiments, pre-training an infrared image adversarial generative model includes:

[0027] A sample dataset is constructed based on original infrared images of different object types. The annotation information of each original infrared image includes the temperature, material, emissivity, and radiance of the object in the infrared image.

[0028] Define the object category matrix, material category matrix, emissivity matrix, and temperature matrix corresponding to the original infrared image;

[0029] Add an offset to the material matrix and re-encode it to obtain the re-encoded material matrix; perform one-hot encoding on the re-encoded material matrix to obtain the corresponding material tensor;

[0030] Based on the preset weighting coefficients, emissivity and temperature are weighted respectively;

[0031] The material tensor, the weighted emissivity matrix, and the temperature matrix are concatenated along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities.

[0032] The fused tensor is input into a generator network, and feature extraction is performed on the fused tensor to obtain initial features; upsampling is performed layer by layer using the initial features of multiple consecutive upsampling layers to obtain and generate an infrared image; wherein, each upsampling layer calculates a corresponding modulation parameter based on the input features of the corresponding upsampling layer using a guidance module, and performs feature modulation on the input features of the layer using the modulation parameter to embed physical quantities and semantic features;

[0033] The fused tensor, generated infrared image, and original infrared image are input into the multi-scale discriminator, and multiple sub-discriminators are used to discriminate the input step by step and output the discrimination results. The input of the subsequent sub-discriminator is the down-sampling result of the adjacent previous sub-discriminator.

[0034] The generator and discriminator are trained alternately, and the model hyperparameters are updated until the model converges.

[0035] According to a second aspect of the present invention, there is provided an infrared image adversarial generation system guided by physical quantities and semantic masks, the system comprising:

[0036] The data preprocessing module is used to obtain the physical quantity parameters and semantic parameters to be processed; and construct the corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix for the physical quantity parameters and semantic parameters; wherein the physical quantity parameters include: emissivity, temperature; the semantic parameters include: object type, material type;

[0037] A data splicing module is used to encode the object category matrix, material category matrix, emissivity matrix, and temperature matrix respectively and splice them along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities;

[0038] The infrared image generation module is used to input the fused tensor into the generator network of the trained infrared image adversarial generative model, perform layer-by-layer upsampling on the fused tensor, and obtain the target infrared image; wherein, the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to perform feature modulation on the input features of the layer to embed physical quantities and semantic features.

[0039] According to a third aspect of the present invention, a computer program product is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned infrared image adversarial generation method based on physical quantity and semantic mask guidance is implemented.

[0040] According to a fourth aspect of the present invention, there is provided an electronic device, comprising:

[0041] processor; and

[0042] a memory for storing executable instructions of the processor;

[0043] The processor is configured to execute the above-mentioned infrared image adversarial generation method guided by physical quantities and semantic masks by executing the executable instructions.

[0044] According to a fifth aspect of the present invention, a storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the above-mentioned infrared image adversarial generation method based on physical quantity and semantic mask guidance is implemented.

[0045] The infrared image adversarial generation method based on physical quantity and semantic mask guidance provided by the embodiment of the present invention, by pre-training the infrared image adversarial generation model based on the generative adversarial network, can first construct the corresponding matrix for the physical quantity parameters and semantic parameters currently input by the user, and splice them into a fusion tensor, so as to maintain the information balance between the semantic and physical quantities when the model is input. The fusion tensor is configured as the input parameter of the model generator; by configuring a guidance module in each upsampling layer of the generator, the modulation parameters of this layer can be calculated, and the physical quantity and semantic features of the input features of this layer can be embedded using the modulation parameters; thereby, the generation of infrared images with sufficient details and physical laws is achieved; and the generated infrared images follow the physical radiation laws to ensure that they can truly reflect the thermal radiation characteristics of the target.

[0046] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] The accompanying drawings are incorporated into and constitute a part of this specification, illustrate embodiments consistent with the present invention, and together with the description, serve to explain the principles of the present invention. Obviously, the drawings described below are only some embodiments of the present invention, and it is clear that those skilled in the art can derive other drawings based on these drawings without inventive effort.

[0048] Figure 1 A schematic diagram schematically illustrates an infrared image adversarial generation method based on physical quantity and semantic mask guidance according to an exemplary embodiment of the present invention;

[0049] Figure 2 A schematic diagram schematically illustrates a guide module according to an exemplary embodiment of the present invention;

[0050] Figure 3 A schematic diagram schematically illustrates the structure of a generative adversarial network model based on multi-layer guidance and multi-scale discrimination according to an exemplary embodiment of the present invention;

[0051] Figure 4 A schematic diagram schematically illustrating the comparison of different infrared images according to an exemplary embodiment of the present invention;

[0052] Figure 5 A schematic diagram schematically illustrating comparison of details of a car in infrared images generated by different methods according to an exemplary embodiment of the present invention;

[0053] Figure 6 A schematic diagram schematically illustrating comparison of human body details in infrared images generated by different methods according to an exemplary embodiment of the present invention;

[0054] Figure 7 A schematic diagram schematically illustrating comparison of house details in infrared images generated by different methods according to an exemplary embodiment of the present invention;

[0055] Figure 8 A schematic diagram schematically illustrates an infrared image adversarial generation system guided by physical quantities and semantic masks according to an exemplary embodiment of the present invention;

[0056] Figure 9 The figure schematically shows the composition of an electronic device in an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0057] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be embodied in many forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that this disclosure will be thorough and complete and will fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.

[0058] In addition, the accompanying drawings are merely schematic illustrations of the present invention and are not necessarily drawn to scale. Identical reference numerals in the figures denote identical or similar parts, and thus repetitive descriptions thereof will be omitted. Some of the blocks shown in the accompanying drawings are functional entities that do not necessarily correspond to physically or logically separate entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different networks and / or processor devices and / or microcontroller devices.

[0059] In related technologies, deep learning-based infrared image generation methods typically use generative adversarial networks (GANs). The core of this approach is to use GANs to learn the distributional characteristics of real infrared images to generate simulated images that are highly similar to real infrared images in terms of visual effects and statistical characteristics. Generally, real infrared image data is first cleaned and enhanced using image preprocessing techniques to remove background noise and optimize the contrast of key features, providing high-quality input data for GAN training. The GAN generator is then used to generate an initial infrared image from random noise, and the discriminator is used to evaluate and provide feedback on the authenticity of the generated image. During training, the generator and discriminator are continuously optimized through an adversarial game. The discriminator extracts deep features from the image to improve its ability to distinguish between real and generated images, while the generator gradually learns the distributional characteristics of real infrared images to produce more realistic infrared images. However, there are certain technical defects in the existing technical solutions, including: (1) Insufficient physical consistency of the generated images: The existing GAN-based generation methods mainly focus on the realism of visual effects, but due to the lack of constraints on physical laws, the generated infrared images may not accurately reflect the actual physical properties, limiting their reliability in engineering applications; (2) Insufficient detail fidelity: Due to insufficient or uneven distribution of training data, the infrared images generated by GAN may be blurred or distorted in detail features such as edges and textures. Especially in complex scenes, the generated images are difficult to meet the requirements of high-precision simulation.

[0060] In view of the shortcomings and deficiencies of the existing technology, this example embodiment provides an infrared image adversarial generation method based on physical quantity and semantic mask guidance. Figure 1 As shown, the method may include the following steps:

[0061] Step S11, obtaining physical quantity parameters and semantic parameters to be processed; and constructing corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix for the physical quantity parameters and semantic parameters; wherein the physical quantity parameters include: emissivity, temperature; the semantic parameters include: object type, material type;

[0062] Step S12: Encode the object category matrix, material category matrix, emissivity matrix, and temperature matrix respectively and concatenate them along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities;

[0063] Step S13, inputting the fused tensor into the generator network of the trained infrared image adversarial generative model, performing upsampling processing on the fused tensor layer by layer, and obtaining the target infrared image; wherein, the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to perform feature modulation on the input features of the layer to embed physical quantities and semantic features.

[0064] Hereinafter, each step of the infrared image adversarial generation method guided by physical quantities and semantic masks in this example implementation will be described in more detail with reference to the accompanying drawings and embodiments.

[0065] In this example embodiment, the infrared image adversarial generation method guided by physical quantities and semantic masks can be executed by a smart terminal device, or by a server communicating with the smart terminal device; or, it can be executed collaboratively between the smart terminal device and the server. A trained infrared image adversarial generation model can be pre-deployed on the smart terminal device or the server. The user can input the currently customized physical quantities and semantic parameters on the smart terminal device, create a corresponding infrared image generation task, execute the infrared image generation task locally on the terminal device, and output a target infrared image; or, the infrared image generation task can be sent to the server, which executes the infrared image adversarial generation method guided by physical quantities and semantic masks, generates a target infrared image, and feeds it back to the terminal device.

[0066] In step S11, the physical quantity parameters and semantic parameters to be processed are obtained; and corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix are constructed for the physical quantity parameters and semantic parameters; wherein the physical quantity parameters include: emissivity, temperature; and the semantic parameters include: object type, material type.

[0067] For example, users can customize physical quantity parameters and semantic parameters based on the desired infrared image. Semantic parameters include object type and material type. For semantic parameters, specific objects in infrared images can be defined based on object categories, such as vehicles, buildings, or pedestrians; as well as the specific materials of the objects; for example, the materials of the vehicle body, windows, and tires are light-colored coated steel plates, glass, and rubber, respectively; or the materials of human clothing and shoes are cotton and canvas, respectively; and so on. Physical quantity parameters include emissivity and temperature, which can be used to define the temperature of the object and the temperature of the material. In addition, infrared data such as radiance can also be defined.

[0068] For customized physical quantity parameters and semantic parameters, corresponding matrix data can be constructed. Specifically, the category matrix of the object can be set as Each pixel represents the target category number, ranging from {1, 2, ..., C}; the material category matrix is Each pixel represents the material category number, ranging from {1, 2, ..., K}; the emissivity matrix is The temperature matrix is Represents the emissivity and temperature values of each pixel.

[0069] In step S12, the object category matrix, material category matrix, emissivity matrix, and temperature matrix are encoded respectively and spliced along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities.

[0070] Exemplarily, encoding the object category matrix, the material category matrix, the emissivity matrix, and the temperature matrix respectively and splicing them along the channel dimension includes:

[0071] Step S21, adding an offset to the material matrix and re-encoding it to obtain a re-encoded material matrix; performing one-hot encoding on the re-encoded material matrix to obtain a corresponding material tensor;

[0072] Step S22, performing weighted processing on the emissivity matrix and the temperature matrix based on preset weighting coefficients;

[0073] In step S23 , the material tensor, the emissivity matrix after weighted processing, and the temperature matrix are concatenated along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities.

[0074] Specifically, in order to achieve combined encoding of object categories and material categories, an offset of the material number can be defined for each object category, so that the same material has different numbers in different target categories.

[0075] For object category c∈{1,2,…,C}, define the offset Δ c , represents the starting value of the material number of object category c. The recoded material matrix is Its value is calculated by the following formula:

[0076] M new (i, j) = M (i, j) + Δ O(i,j)

[0077] Among them, M(i,j) is the original material number, O(i,j) is the target category of pixel (i,j), Δ O(i,j) is the offset corresponding to the object category.

[0078] Perform one-hot encoding on the re-encoded material matrix and convert it into a (H, W, K total ) tensor M one-hot, where K total is the total number of material numbers after recoding.

[0079] For each pixel position (i, j), the re-encoded material number M new (i,j) generates a K-length total Vector:

[0080]

[0081] Among them, we get

[0082] At the same time, in order to maintain the balance between semantics and physical quantities in information representation, we set α T is the emissivity weighting coefficient, β ε is the weighted coefficient of temperature, and the weighted method is used to adjust the influence of emissivity ε and temperature T. Finally, M one-hot , ε′ and T′ are concatenated along the channel dimension to obtain the final input tensor X. The formula can be expressed as:

[0083] T′=T·(1-α T )

[0084] ε′=ε·(1+β ε )

[0085] X=Concat(M one-hot ,T′,ε′)

[0086] The fused input tensor can be represented by a multi-channel two-dimensional matrix. The value of each pixel can represent the semantic category of the location and the emissivity and temperature values of the semantically corresponding material. Directly using a semantic mask as input makes it difficult to capture complex semantic information. Based on Planck's formula, we know that the material directly associated with the semantics will directly determine the transmittance and temperature values that match the material. Therefore, it is necessary to use convolution operations to map the physical information represented by semantics and physical quantities into a continuous high-dimensional space to better enable the generator to learn. Although the one-hot encoding representing semantics and material converts the material and semantic information into a binary matrix, the material semantic information is preserved through the distribution of channel positions.

[0087] In step S13, the fused tensor is input into the generator network of the trained infrared image adversarial generative model, and the fused tensor is upsampled layer by layer to obtain the target infrared image; wherein, the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to perform feature modulation on the input features of the layer to embed physical quantities and semantic features.

[0088] Exemplarily, inputting the fused tensor into a generator network of a trained infrared image adversarial generative model, upsampling the fused tensor layer by layer, and obtaining a target infrared image includes:

[0089] Step S31: In an upsampling layer, upsampling and normalizing the input features of the upsampling layer are performed in sequence to obtain first feature data;

[0090] Step S32: The guiding module performs convolution processing on the input features of the upsampling layer to obtain corresponding second feature data; and performs two-dimensional convolution processing on the second feature data to obtain corresponding modulation parameters;

[0091] Step S33, modulating the first characteristic data using a modulation parameter to obtain a modulation characteristic;

[0092] Step S34: performing feature fusion processing on the modulation feature and the first feature data to obtain the output feature of the upsampling layer; wherein the output feature is configured as the input feature of the next adjacent upsampling layer.

[0093] Exemplarily, the modulation parameters include scaling parameters and offset parameters having spatial dimensions.

[0094] Exemplarily, the matrix size of the modulation parameters is the same as the size of the output features of the corresponding upsampling layer.

[0095] Specifically, refer to Figure 3 As shown, the generator includes multiple consecutive upsampling layers, each of which is equipped with a guidance module. The fused tensor serves as the input to the first upsampling layer, and the output of the first upsampling layer serves as the input to the second upsampling layer, and so on. Each upsampling layer can include at least one convolutional layer and a pooling layer, arranged in sequence.

[0096] Take the first upsampling layer as an example, its input is the fused tensor. Figure 2 As shown, the first guiding module in the upsampling layer can first use a convolution layer to extract features from the fused tensor to generate corresponding second feature data; and then use a two-dimensional convolutional network to convolve the second feature data to generate corresponding modulation parameters, namely, the scaling parameter γ and the offset parameter β.

[0097] γ Phy =Conv r (Concat(M one-hot ,T′,ε′))

[0098] β Phy =Conv β (Concat(M one-hot,T′,ε′))

[0099] Among them, M one-hot is the encoded semantic and material information (i.e., fusion tensor); T′, ε′ are the emissivity and temperature matrices modified based on the weighting of temperature and transmittance, respectively.

[0100] Two sets of convolutions are used to extract spatially adaptive modulation parameters from the semantic mask and the physical quantity matrix, while ensuring that the γ and β matrices are consistent in shape with the output feature map of the generator at this layer.

[0101] The scaling parameter γ and the offset parameter β are matrices that can be used to calculate the mean (first-order mean) and standard deviation of the value of each pixel in each channel within the minimum batch for normalization. A denormalized semantic and physical quantity guidance module is introduced in each layer of the generator, so that γ and β are dynamically generated from the matrix after the input semantic mask and physical quantity are fused through two-dimensional convolution. They are no longer simple global learnable parameters guided by the loss function. The mathematical expression of two-dimensional convolution is as follows:

[0102]

[0103] in, is the value of the kth channel at position (i, j) of the output feature map, x i+p,j+q,m is the value of the mth channel of the input feature map at position (i+p,j+q), is the weight of the convolution kernel's kth output channel corresponding to the mth input channel, b k is the bias of the kth output channel; K is the size of the convolution kernel, C in is the number of channels of the input feature map.

[0104] Two-dimensional convolutions are advantageous in extracting spatial information. The local connections and weight sharing characteristics of two-dimensional convolutional neural networks enable them to efficiently capture the spatial layout of input data and excel in tasks related to semantic masking. The translation invariance of two-dimensional convolutional neural networks enables them to accurately extract spatial information and generate stable semantic masks even when the target position changes.

[0105] The two-dimensional convolutional network can identify the semantics of different target materials by channel position and learn the relationship between material, emissivity, and temperature through multi-channel fusion. This change allows the normalization layer to not only utilize the statistical information of the feature map itself, but also combine the input semantic mask with physical quantities to perform spatially adaptive normalization on the feature map.

[0106] Taking the first upsampling layer as an example, its input is the fused tensor. The upsampling layer can include at least one convolutional layer (for feature extraction and upsampling of the fused tensor) and a normalization layer, arranged in sequence, and outputs the first feature data. This achieves the conversion from a low-resolution feature map to a high-resolution feature map. After each upsampling, the resolution of the feature map is doubled, and the number of channels is halved.

[0107] For the first upsampling layer, after obtaining the scaling parameter γ, offset parameter β, and upsampled features of this layer, the modulation parameter matrices γ and β can be applied to the standardized feature map (i.e., the first feature data) to embed semantic and physical information into the feature map, and obtain the feature map x modulated by the Phy-SPADE module. norm .

[0108]

[0109] Then the modulated feature map x is transformed into norm When added to the original feature (first feature data), the generated infrared image is consistent with the input multi-source information in semantics, temperature and transmittance.

[0110] The output of the first upsampling layer can be used as the input of the second upsampling layer and perform the same calculation process as the first upsampling layer.

[0111] For example, the generator can include 3 or 5 consecutive upsampling layers. In addition, the last layer of the generator can be set as a convolution layer to convert the feature map output by the last upsampling layer into an infrared image. The formula can be expressed as:

[0112] I fake =Conv(F final )

[0113] Among them, F final is the last layer feature map of the generator, I fake The generated infrared image.

[0114] For example, in other exemplary embodiments of the present disclosure, inputting the fused tensor into a generator of a trained infrared image adversarial generative model, upsampling the fused tensor layer by layer, and obtaining a target infrared image includes:

[0115] Step S41: In an upsampling layer, feature extraction and normalization are performed on the input features of the upsampling layer in sequence to obtain third feature data;

[0116] Step S42: The guiding module performs convolution processing on the input features of the upsampling layer to obtain corresponding second feature data; and performs two-dimensional convolution processing on the second feature data to obtain corresponding modulation parameters;

[0117] Step S43, modulating the third characteristic data using the modulation parameter to obtain a modulation characteristic;

[0118] Step S44: perform feature fusion processing on the modulation feature and the third feature data to obtain fused feature data; and perform upsampling processing on the fused feature data to obtain the output feature of the upsampling layer; wherein the output feature is configured as the input feature of the next adjacent upsampling layer.

[0119] Specifically, for any upsampling layer in the generator, its main network can include at least one convolutional layer and a normalization layer; the convolutional layer can be used to extract features from the input of this layer, and the normalization layer is used to normalize the feature map output by the convolutional layer and output the third feature data.

[0120] The modulation parameters of the up-sampling layer take the third characteristic data as input and use the calculation process of the above embodiment to calculate the modulation parameters corresponding to the layer.

[0121] The third feature data is then modulated using the adjustment parameters to obtain a modulated feature map (i.e., a modulated feature). A residual network can then be used to perform feature fusion processing between the modulated feature map and the third feature data to embed semantic and physical parameters and obtain fused feature data. This fused feature data is then upsampled to obtain upsampled features with expanded resolution, which serve as the output data of this layer.

[0122] In some example embodiments, the guiding module can be a normalization layer set in each upsampling network structure, which implements convolutional learning in each normalization layer of the generator to generate modulation parameters with spatial adaptability, encode information about the layout of semantics and physical quantities, and enable the generator to retain physical quantity and semantic information while enjoying the benefits of normalization.

[0123] Exemplarily, the method further includes: pre-training an infrared image adversarial generation model, including:

[0124] Step S51, constructing a sample data set based on original infrared images of different object types; wherein the annotation information of each original infrared image includes the temperature, material, emissivity, and radiance of the object in the infrared image;

[0125] Step S52, defining the object category matrix, material category matrix, emissivity matrix, and temperature matrix corresponding to the original infrared image;

[0126] Step S53: adding an offset to the material matrix and re-encoding it to obtain a re-encoded material matrix; performing one-hot encoding on the re-encoded material matrix to obtain a corresponding material tensor;

[0127] Step S54, performing weighted processing on the emissivity and temperature based on the preset weighting coefficients;

[0128] Step S55: concatenate the material tensor, the weighted emissivity matrix, and the temperature matrix along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities;

[0129] Step S56: Input the fused tensor into the generator network, perform feature extraction on the fused tensor to obtain initial features; perform upsampling layer by layer using the initial features of multiple consecutive upsampling layers to obtain and generate an infrared image; wherein, each upsampling layer uses a guidance module to calculate a corresponding modulation parameter based on the input features of the corresponding upsampling layer, and uses the modulation parameter to perform feature modulation on the input features of the layer to embed physical quantities and semantic features;

[0130] Step S57: Input the fused tensor, generated infrared image, and original infrared image into a multi-scale discriminator, use multiple sub-discriminators to discriminate the input step by step, and output the discrimination results; wherein the input of the subsequent sub-discriminator is the down-sampling result of the adjacent previous sub-discriminator;

[0131] Step S58: Alternately train the generator and discriminator, and update the model hyperparameters until the model converges.

[0132] Specifically, a sample dataset can be constructed using the original infrared images. The sample dataset can include multiple different types of objects, such as vehicles, residential buildings, and human bodies. Vehicles include cars, vans, and trucks; buildings include residential houses, mid-rise residences, and high-rise residences; and human bodies are male bodies.

[0133] Spherical sampling points were established for all objects at 10° intervals, with a vertical range of -90° to 90° in elevation and a horizontal range of -180° to 180° in azimuth. Each object had 614 viewpoints and 614 images, all with a resolution of 512 x 512 pixels. Vehicle bodies, windows, and tires were constructed from light-colored, pre-painted steel, glass, and rubber, respectively. Vehicle body temperatures ranged from 30°C to 70°C, with data samples collected every 10°C. Civilian building walls, windows, and roofs were constructed from sandstone plaster, glass, and concrete, respectively. Civilian building wall temperatures ranged from 15°C to 30°C, with data samples collected every 5°C. Human clothing and shoes were constructed from cotton and canvas, respectively. Human clothing temperatures ranged from 15°C to 30°C, with data samples collected every 5°C.

[0134] The sample dataset is divided into training and test sets, each containing images of all object classes at different viewing angles and temperatures. Full-angle images of each object at elevations of [-80°, -60°, -40°, -20°, 0°, 20°, 40°, 60°, and 80°] are selected. Furthermore, images of vehicles at temperatures of 30°C, 50°C, and 60°C, and of civil buildings and human bodies at temperatures of 15°C and 25°C are selected as the training set, totaling 5,508 images. Images of vehicles at 40°C and 70°C, and of buildings and human bodies at other elevations of 20°C and 30°C, totaling 4,060 images, are selected as the test set. Each object has its own annotation file, providing infrared data such as temperature, material emissivity, and radiance.

[0135] For infrared images in the training set, the emissivity, temperature, target semantics, and material semantics can be encoded in multiple channels through the physical quantity and semantic fusion module to achieve the fusion of semantics and physical quantities.

[0136] Specifically, the object category matrix can be set as Each pixel represents the object category number, ranging from {1, 2, ..., C}. The material category matrix is Each pixel represents a material category number, ranging from {1, 2, ..., K}. The emissivity matrix is The temperature matrix is Represents the emissivity value and temperature value of each pixel. In order to achieve the combined encoding of object category and material category, we define an offset of the material number for each object category, so that the same material has different numbers in different object categories. For object category c∈{1,2,…,C}, define the offset Δ c , indicating the starting value of the material number of category c.

[0137] The re-encoded material matrix is Its value is calculated by the following formula:

[0138] M new (i, j) = M (i, j) + Δ O(i,j)

[0139] Among them, M(i,j) is the original material number, O(i,j) is the target category of pixel (i,j), Δ O(i,j) is the offset corresponding to the target category.

[0140] Perform one-hot encoding on the re-encoded material matrix and convert it into a (H, W, K total ) tensor M one-hot Among them, K totalIs the total number of re-encoded texture numbers. For each pixel position (i, j), the re-encoded texture number M new (i,j) generates a K-length total Vector:

[0141]

[0142] Finally got At the same time, in order to maintain the balance between semantics and physical quantities in information representation, we set α T is the emissivity weighting coefficient, β ε is the weighted coefficient of temperature, and the weighted method is used to adjust the influence of emissivity ε and temperature T. Finally, M one-hot , ε′ and T′ are concatenated along the channel dimension to obtain the final fused tensor X, which is used as the input of the generator. The formula can be expressed as:

[0143] T′=T·(1-α T )

[0144] ε′=ε·(1+β ε )

[0145] X=Concat(M one-hot ,T′,ε′)

[0146] In the generator of a generative adversarial model, the denormalized semantic and physical quantity guidance module, Phy-SPADE, can be used to convolve the physical quantity and semantic mask to generate modulation parameters with spatial dimensions. By encoding the layout information of the physical quantity and semantics, the generator can preserve the physical quantity and semantic information while achieving the benefits of normalization.

[0147] Considering the role of the normalization layer in a neural network, its main purpose is to normalize the feature maps, making network training more stable and accelerating convergence. Because when training a neural network, the distribution of feature maps may change dramatically, affecting the speed and effectiveness of the network. The normalization layer standardizes the feature maps and adjusts their distribution to a more stable range, thereby alleviating these problems. The core idea of the normalization layer is to normalize each channel of the feature map. The formula is as follows:

[0148]

[0149] Where x is the pixel value of the feature map of the current layer; μ is the mean of the feature map; σ is the standard deviation of the feature map; is the normalized feature map.

[0150] Considering that the general normalization layer relies on the statistical information of the feature map itself, it is insensitive to the spatial position of the feature map and cannot utilize the spatial layout information of the input semantic mask and physical quantity. In the process of generating infrared images, the problem of semantic and physical quantity information loss will occur.

[0151] Therefore, this method introduces a denormalized semantic and physical quantity guidance module in each layer of the generator, so that the scaling parameter γ and offset parameter β are dynamically generated from the matrix after the input semantic mask and physical quantity are fused through two-dimensional convolution. They are no longer simple global learnable parameters guided by the loss function. The mathematical expression of two-dimensional convolution is as follows:

[0152]

[0153] in, is the value of the kth channel at position (i, j) of the output feature map, x i+p,j+q,m is the value of the mth channel of the input feature map at position (i+p,j+q), is the weight of the convolution kernel's kth output channel corresponding to the mth input channel, b k is the bias of the kth output channel. K is the size of the convolution kernel, C in is the number of channels of the input feature map.

[0154] Two-dimensional convolutions are advantageous in extracting spatial information. The local connections and weight sharing characteristics of two-dimensional convolutional neural networks enable them to efficiently capture the spatial layout of input data and excel in tasks related to semantic masking. The translation invariance of two-dimensional convolutional neural networks enables them to accurately extract spatial information and generate stable semantic masks even when the target position changes.

[0155] The fused tensor obtained by fusing the semantic mask and physical quantities through the aforementioned fusion module can be represented by a multi-channel two-dimensional matrix. The value of each pixel represents the semantic category at that location and the emissivity and temperature values of the semantically corresponding material. Directly using the semantic mask as input can fail to capture complex semantic information. Based on Planck's formula, we know that the material directly associated with the semantics directly determines the transmittance and temperature values that match the material. Therefore, a convolution operation is needed to map the physical information represented by both semantics and physical quantities into a continuous high-dimensional space to facilitate better learning for the generator. While the one-hot encoding representing the target semantics and material converts the material and semantic information into a binary matrix, the material semantic information is preserved through the distribution of channel positions. The two-dimensional convolutional network can identify the semantics of different target materials by channel position and learn the relationship between material, emissivity, and temperature through multi-channel fusion. This change allows the normalization layer to not only utilize the statistical information of the feature map itself, but also combine the input semantic mask and physical quantities to perform spatially adaptive normalization of the feature map.

[0156] γ Phy =Conv r (Concat(M one-hot ,T′,ε′))

[0157] β Phy =Conv β (Concat(M one-hot ,T′,ε′))

[0158]

[0159] Among them, M one-hot is the encoded target semantics and material information, and T′,ε′ is the modified emissivity and temperature matrix based on the weighting of temperature and transmittance. norm is the characteristic map after modulation.

[0160] The spatially adaptive modulation parameters are extracted from the semantic mask and the physical quantity matrix through two sets of convolutions, while ensuring that the γ and β matrices are consistent with the shape of the output feature map of the generator at this layer. The modulation parameter matrices γ and β are then applied to the normalized feature map to embed the semantic and physical information into the feature map, and the feature map x modulated by the Phy-SPADE module is obtained. norm .

[0161] For the generator network, the input fusion tensor X is input into the generator network, and the fusion matrix between semantics, emissivity and temperature is feature mapped to generate the initial feature map. The formula can be expressed as:

[0162] F init =Conv(X)

[0163] Where, the dimension of X is C seg,phy ×H×W,C seg,phy F is the sum of the semantic category plus the emissivity and temperature channels, H and W are the resolutions of the input feature map. init Is the initial feature map, shape C init ×H×W;C init =16×n f , represents the deep feature space. Conv is a 3*3 convolution operation that maps the sparse semantic mask to a compact feature space, providing a basis for subsequent layer-by-layer generation.

[0164] In the generator network, in the first upsampling layer, the multi-layer semantic and physical quantity guidance module (Phy-SPADE) of this layer can be used to perform the initial feature map F init The first guided module generates a feature map modulated by semantic material information and emissivity temperature, which is expressed as:

[0165] F guided =Phy-SPADE(F init ,X)

[0166] Then upsample the modulated feature map to generate a higher resolution feature map. The formula is expressed as:

[0167]

[0168] Among them, c is the channel index; i′∈[0,2H-1], j′∈[0,2W-1], is the output feature map F up Pixel index of s h and s w is the scaling factor of height and width, and each pixel value F of the output feature map up (c,i′,j′) is directly taken from the input feature map F guided The nearest pixel value in the input feature map is determined by scaling down the pixel index i′, j′ of the output feature map to the pixel index range of the input feature map.

[0169] The feature map output by the first upsampling layer can be used as the input of the next adjacent upsampling layer.

[0170] For example, the generator can include six consecutive upsampling layers, each with a guide module. The upsampled feature map is further extracted into deep semantic and physical features in a higher resolution space through two layers of Phy-SPADE modules, while keeping the number of feature channels unchanged. The formula can be expressed as:

[0171] F up1=Phy-SPADE(F up ,X)

[0172] F up2 =Phy-SPADE(F up1 ,X)

[0173] Alternatively, in some exemplary embodiments, the fused tensor may be configured as the input of the guidance module in the first upsampling layer, and the features output by the guidance module in the first upsampling layer may be configured as the input of the guidance module in the second upsampling layer; and so on.

[0174] The generator gradually generates feature maps from low resolution to high resolution by upsampling layer by layer. After each upsampling, the resolution of the feature map is doubled. At the same time, the feature map is modulated by the Phy-SPADE module to ensure that it conforms to the semantic layout and physical laws. The layer-by-layer generation process can be expressed as:

[0175]

[0176] in, is the feature map of the i-th layer, and UpSample is the upsampling operation.

[0177] To gradually restore the spatial details of the feature map, after two layers of Phy-SPADE modules that maintain the same channels, the number of channels in the output feature map of each subsequent Phy-SPADE module is gradually reduced. After each upsampling, the number of channels is halved until the final output channel number meets the requirements for the output infrared image. In the final layer of the generator, a convolutional layer converts the feature map into an infrared image. The formula is expressed as:

[0178] I fake =Conv(F final )

[0179] Among them, F final is the last layer feature map of the generator, I fake The generated infrared image.

[0180] Specifically, in order to achieve the coordinated optimization of the generator and the discriminator, an alternating training strategy is adopted to train the generator and the discriminator. In each training iteration, the generator and the discriminator perform their respective optimization steps. The update frequency of the generator is determined by the hyperparameter D steps_per_G Control, that is, every D steps_per_G After the update, the generator performs an optimization step. In addition, the discriminator is updated in each iteration to ensure that it can timely discriminate and optimize the samples generated by the generator.

[0181] Specifically, the multi-scale discriminator consists of multiple sub-discriminators (D1, D2, ..., Dn ); the structure of each sub-discriminator is similar, but the resolution of their input feature maps is different.

[0182] Specifically, multiple sub-discriminators can be initialized first; the input of the multi-scale discriminator consists of the multi-channel fusion tensor X, the generated image I fake and the real image I real The real image is the original infrared image in the sample dataset.

[0183] The discriminator needs to judge whether the input infrared image meets the semantic conditions and physical laws to determine whether it is true or false. The initial feature map is X1=Concat(X,I fake ,I real ), the multi-scale discriminator traverses each sub-discriminator D k To achieve step-by-step discrimination, the first sub-discriminator D1 receives the input initial feature map X1, and the subsequent sub-discriminators D k The input is the result of downsampling the input feature map of the previous sub-discriminator X k , each sub-discriminator D k Need to input feature map X k Make a judgment.

[0184] Each sub-discriminator in the multi-scale discriminator is an independent PatchGAN discriminator. Assume that the input feature map X k The size is N*N, D k The input feature map X k Divide into multiple local patches, each of which is n*n in size, and the entire input feature map X l It can be divided into M*M small blocks.

[0185] in, Sub-discriminator D k Need to input feature map X k The discrimination process can be expressed by the following formula:

[0186]

[0187] Among them, D i,j (X k ) is the discriminant output of the (i, j)th patch, indicating whether the patch is real or fake.

[0188] Sub-discriminator D k The input layer is used to receive the feature map X k , which is used to pass the feature matrix to the subsequent convolution operation. kThe convolution layer is used to gradually extract the features of the local patch. Each convolution layer is followed by a LeakyReLU activation function and an instance normalization layer. The last layer of PatchGAN is a convolution layer, whose output is a two-dimensional feature map D k (X k ), D k (X k ) is of size M*M, and each element corresponds to the discrimination result of a patch, indicating the true or false probability of each patch.

[0189] Assume that the output of the K-th discriminator is O k , then after traversing all sub-discriminators, the output of each sub-discriminator (O1, O2, ..., O n ), then these outputs are fused to obtain the final output result of the multi-scale discriminator, which can be expressed as:

[0190]

[0191] By utilizing a multi-scale discriminator with a gradually reduced resolution design, the discriminator is able to discriminate the features generated by the multi-layer guided generator from a global to local, multi-scale perspective. The discriminator and generator are repeatedly trained using an alternating training strategy, updating the weights of the adversarial generative network model until the network is able to generate realistic infrared images.

[0192] For example, refer to Figure 4 The figure shows a comparison diagram of the infrared images generated by the present invention and the results of other methods; the first row shows a real infrared simulation image; the second row shows an infrared image generated by the existing technology based on the semantic method; the third row shows an infrared image with controllable temperature and realistic details generated by the present invention. Figure 5-Figure 7 The following is a comparison diagram of infrared image details. The first column shows a realistic infrared simulated image; the second column shows details of an infrared image generated using semantic methods using existing techniques; and the third column shows details of a temperature-controlled, realistic infrared image generated by the present invention. Comparing the figures above shows that compared to existing techniques, the present invention's method is able to generate infrared images with sufficient detail that conforms to physical laws. The infrared images generated by this method exhibit superior overall quality and detail.

[0193] The proposed infrared image adversarial generation method, guided by physical quantities and semantic masks, utilizes a physical quantity and semantic mask fusion module to implement multi-channel encoding of target semantics, material information, emissivity, and temperature, based on semantic one-hot representation and a weighted balancing mechanism, thus achieving physical and semantic coupling. A guidance module is placed in each upsampling layer of the generator network. Modulation parameters with spatial dimensions are generated by convolving the physical quantity with the semantic material information. These spatial modulation parameters enable the generator to learn infrared image generation while preserving both physical and semantic material information.

[0194] In addition, by using a denormalized semantic and physical quantity guidance module, convolutional learning is performed at each normalized layer of the generator to generate spatially adaptive modulation parameters, encoding information about the layout of semantic and physical quantities, allowing the generator to enjoy the benefits of normalization while retaining physical and semantic information. At the same time, because each residual block of the generator operates at a different scale, the semantic and physical quantity features are downsampled to match the multi-scale resolution of the generator, and the generator is trained using a multi-scale discriminator and loss function, allowing the discriminator to judge the features generated by the multi-level guidance generator in a global to local multi-scale manner. Compared with existing technologies, it can generate infrared images with sufficient details and physical laws.

[0195] It should be noted that the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0196] Furthermore, the embodiment of this example also provides an infrared image adversarial generation system 80 guided by physical quantities and semantic masks, including:

[0197] The data preprocessing module 801 is used to obtain physical quantity parameters and semantic parameters to be processed; and construct corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix for the physical quantity parameters and semantic parameters. The physical quantity parameters include emissivity and temperature; the semantic parameters include object type and material type.

[0198] A data splicing module 802 is used to encode the object category matrix, material category matrix, emissivity matrix, and temperature matrix respectively and splice them along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities;

[0199] The infrared image generation module 803 is used to input the fused tensor into the generator network of the trained infrared image adversarial generative model, perform upsampling processing on the fused tensor layer by layer, and obtain the target infrared image; wherein, the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to feature modulate the input features of the layer to embed physical quantities and semantic features.

[0200] The functional implementation of each module in the infrared image countermeasure generation system 80 has been explained in detail in the corresponding method embodiment and will not be repeated here.

[0201] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to an embodiment of the present invention, the features and functions of two or more modules or units described above can be concretized in one module or unit. Conversely, the features and functions of one module or unit described above can be further divided into multiple modules or units to be concretized.

[0202] Figure 9 A schematic diagram of an electronic device suitable for implementing an embodiment of the present invention is shown.

[0203] It should be noted that Figure 9 The electronic device 1000 shown is only an example and should not limit the functions and scope of use of the embodiments of the present invention.

[0204] like Figure 9 As shown, electronic device 1000 includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in read-only memory (ROM) 1002 or the program loaded from storage portion 1008 into random access memory (RAM) 1003. Various programs and data required for system operation are also stored in RAM 1003. CPU 1001, ROM 1002 and RAM 1003 are connected to each other via bus 1004. Input / output (I / O) interface 1005 is also connected to bus 1004.

[0205] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, and the like; an output section 1007 including devices such as a cathode ray tube (CRT), a liquid crystal display (LCD), and a speaker; a storage section 1008 including a hard disk and the like; and a communication section 1009 including a network interface card such as a LAN (Local Area Network) card or a modem. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. Removable media 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, or a semiconductor memory, is installed in the drive 1010 as needed, so that computer programs read therefrom can be installed into the storage section 1008 as needed.

[0206] In particular, according to an embodiment of the present invention, the process described below with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present invention includes a computer program product that includes a computer program carried on a storage medium, the computer program containing program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 1009 and / or installed from a removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the various functions defined in the system of the present application are performed.

[0207] Specifically, the above-mentioned electronic device can be a computer, a tablet computer or a server device.

[0208] It should be noted that the storage medium shown in the embodiments of the present invention can be a computer-readable signal medium or a computer-readable storage medium or any combination of the above. The computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or component, or any combination of the above. More specific examples of computer-readable storage media can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present invention, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, device or device. In the present invention, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, which carries computer-readable program code. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium may also be any storage medium other than a computer-readable storage medium that can transmit, propagate, or transport a program for use by or in connection with an instruction execution system, apparatus, or device. Program code contained on the storage medium may be transmitted using any suitable medium, including but not limited to wireless, wired, or any suitable combination thereof.

[0209] The flowcharts and block diagrams in the accompanying drawings illustrate the possible implementation architecture, functions and operations of the systems, methods and computer program products according to various embodiments of the present invention. In this regard, each box in the flowchart or block diagram can represent a module, program segment, or a part of code, and the above-mentioned module, program segment, or a part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram or flowchart, and the combination of boxes in the block diagram or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0210] The units involved in the embodiments of the present invention may be implemented in software or hardware, and the units described may also be provided in a processor. In some cases, the names of these units do not limit the units themselves.

[0211] It should be noted that, as another aspect, the present application also provides a storage medium, which can be included in an electronic device; or it can exist independently without being installed in the electronic device. The above storage medium carries one or more programs, and when the above one or more programs are executed by an electronic device, the electronic device implements the method described in the following embodiments. For example, the electronic device can implement the following Figure 1 The individual steps of the method are shown.

[0212] In one embodiment, the present application provides a computer program product, including a computer program, which implements the steps in the above-mentioned method embodiments when executed by a processor.

[0213] Furthermore, the above-described figures are merely illustrative of the processes included in the method according to exemplary embodiments of the present invention and are not intended to be limiting. It is readily understood that the processes illustrated in the above-described figures do not indicate or limit the temporal order of these processes. Furthermore, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0214] Other embodiments of the present invention will readily occur to those skilled in the art after considering the specification and practicing the invention herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein. The description and examples are to be considered as exemplary only, with the true scope and spirit of the invention being indicated by the claims.

[0215] It should be understood that the present invention is not limited to the exact construction described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof, which is limited only by the appended claims.

Claims

1. A method for infrared image adversarial generation based on physical quantity and semantic mask guidance, characterized in that: The method comprises: Obtaining physical quantity parameters and semantic parameters to be processed; and constructing corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix for the physical quantity parameters and semantic parameters; wherein the physical quantity parameters include: emissivity, temperature; the semantic parameters include: object type, material type; Encoding the object category matrix, material category matrix, emissivity matrix, and temperature matrix respectively and concatenating them along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities; The fused tensor is input into the generator network of the trained infrared image adversarial generative model, and the fused tensor is upsampled layer by layer to obtain the target infrared image; wherein the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to feature modulate the input features of the layer to embed physical quantities and semantic features.

2. The method according to claim 1, characterized in that The encoding of the object category matrix, the material category matrix, the emissivity matrix, and the temperature matrix respectively and splicing them along the channel dimension includes: Add an offset to the material matrix and re-encode it to obtain the re-encoded material matrix; perform one-hot encoding on the re-encoded material matrix to obtain the corresponding material tensor; Based on the preset weighting coefficients, weighted processing is performed on the emissivity matrix and the temperature matrix respectively; The material tensor, the weighted emissivity matrix, and the temperature matrix are concatenated along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities.

3. The method according to claim 1, characterized in that The modulation parameters include scaling parameters and offset parameters having spatial dimensions.

4. The method according to claim 3, characterized in that The step of inputting the fused tensor into the generator network of the trained infrared image adversarial generative model and upsampling the fused tensor layer by layer to obtain the target infrared image includes: In an upsampling layer, upsampling and normalizing the input features of the upsampling layer are performed in sequence to obtain first feature data; The guiding module performs convolution processing on the input features of the upsampling layer to obtain corresponding second feature data; and performs two-dimensional convolution processing on the second feature data to obtain corresponding modulation parameters; Modulating the first characteristic data using a modulation parameter to obtain a modulation characteristic; The modulation feature is fused with the first feature data to obtain the output feature of the upsampling layer; wherein the output feature is configured as the input feature of the next adjacent upsampling layer.

5. The method according to claim 4, characterized in that The matrix size of the modulation parameters is the same as the size of the output features of the corresponding upsampling layer.

6. The method according to claim 3, characterized in that The step of inputting the fused tensor into the generator of the trained infrared image adversarial generative model and upsampling the fused tensor layer by layer to obtain the target infrared image includes: In an upsampling layer, feature extraction and normalization are performed on the input features of the upsampling layer in sequence to obtain third feature data; The guiding module performs convolution processing on the input features of the upsampling layer to obtain corresponding second feature data; and performs two-dimensional convolution processing on the second feature data to obtain corresponding modulation parameters; modulating the third characteristic data using a modulation parameter to obtain a modulation characteristic; The modulation feature and the third feature data are subjected to feature fusion processing to obtain fused feature data; and the fused feature data is subjected to upsampling processing to obtain the output feature of the upsampling layer; wherein the output feature is configured as the input feature of the next adjacent upsampling layer.

7. The method according to any one of claims 1 to 6, characterized in that The method further includes: pre-training an infrared image adversarial generation model, including: A sample dataset is constructed based on original infrared images of different object types. The annotation information of each original infrared image includes the temperature, material, emissivity, and radiance of the object in the infrared image. Define the object category matrix, material category matrix, emissivity matrix, and temperature matrix corresponding to the original infrared image; Add an offset to the material matrix and re-encode it to obtain the re-encoded material matrix; perform one-hot encoding on the re-encoded material matrix to obtain the corresponding material tensor; Based on the preset weighting coefficients, emissivity and temperature are weighted respectively; The material tensor, the weighted emissivity matrix, and the temperature matrix are concatenated along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities. The fused tensor is input into a generator network, and feature extraction is performed on the fused tensor to obtain initial features; upsampling is performed layer by layer using the initial features of multiple consecutive upsampling layers to obtain and generate an infrared image; wherein, each upsampling layer calculates a corresponding modulation parameter based on the input features of the corresponding upsampling layer using a guidance module, and performs feature modulation on the input features of the layer using the modulation parameter to embed physical quantities and semantic features; The fused tensor, generated infrared image, and original infrared image are input into the multi-scale discriminator, and multiple sub-discriminators are used to discriminate the input step by step and output the discrimination results. The input of the subsequent sub-discriminator is the down-sampling result of the adjacent previous sub-discriminator. The generator and discriminator are trained alternately, and the model hyperparameters are updated until the model converges.

8. An infrared image adversarial generation system based on physical quantity and semantic mask guidance, characterized by: The system comprises: The data preprocessing module is used to obtain the physical quantity parameters and semantic parameters to be processed; and construct the corresponding object category matrix, material category matrix, emissivity matrix, and temperature matrix for the physical quantity parameters and semantic parameters; wherein the physical quantity parameters include: emissivity, temperature; the semantic parameters include: object type, material type; A data splicing module is used to encode the object category matrix, material category matrix, emissivity matrix, and temperature matrix respectively and splice them along the channel dimension to obtain a multi-channel fusion tensor based on semantics and physical quantities; The infrared image generation module is used to input the fused tensor into the generator network of the trained infrared image adversarial generative model, perform layer-by-layer upsampling on the fused tensor, and obtain the target infrared image; wherein, the generator network includes a plurality of consecutive upsampling layers configured with a guidance module; the guidance module is used to calculate the corresponding modulation parameters of the input features of the corresponding upsampling layer, and use the modulation parameters to perform feature modulation on the input features of the layer to embed physical quantities and semantic features.

9. A computer program product comprising a computer program, characterized in that When the computer program is executed by a processor, the infrared image adversarial generation method based on physical quantity and semantic mask guidance according to any one of claims 1 to 7 is implemented.

10. An electronic device, characterized in that: include: processor; as well as a memory for storing executable instructions of the processor; The processor is configured to execute the infrared image adversarial generation method based on physical quantity and semantic mask guidance according to any one of claims 1 to 7 by executing the executable instructions.