Image generation method and device

By preprocessing the reference condition data and inputting it into the image generation model for processing, the problem of high complexity of the image generation method in the prior art is solved, and the effect of simplifying the complexity of the image generation method and reducing the computational complexity is achieved.

CN120032210APending Publication Date: 2025-05-23ROCKET FORCE UNIV OF ENG
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510032505.4
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-08
Publication Date
2025-05-23

AI Technical Summary

Technical Problem

In the prior art, the image generation method has a high complexity when processing long sequences, resulting in an increase in computational complexity. No effective technical solution has been proposed.

Method used

An image generation method is proposed, which preprocesses the reference condition data, improves the data quality, and inputs the preprocessed data into the image generation model, including a generator and a discriminator. The generator processes the target condition data through the downsampling module, the visual attention module and the upsampling module, and the discriminator outputs discriminatory information to indicate the discriminating results of the reference image and the real image.

Benefits of technology

The complexity of the image generation method is simplified, the computational complexity is reduced, and the problem of high complexity of the image generation method in the prior art is solved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032210A_ABST
    Figure CN120032210A_ABST
Patent Text Reader

Abstract

The invention provides an image generation method and device, and relates to the technical field of data processing. The method specifically comprises the steps that reference condition data are obtained, the reference condition data are preprocessed, target condition data are obtained, and the preprocessing operation is used for improving the data quality of the reference condition data; inputting the target condition data into an image generation model, and outputting a target image, the image generation model being a model obtained by training an initial image generation model by using a training sample set in advance, the image generation model comprising a generator and a discriminator, the generator processes the target condition data through the down-sampling module, the visual attention module and the up-sampling module to generate a reference image, the discriminator outputs discrimination information when the reference image serves as input, and the discrimination information is used for indicating a discrimination result of the reference image and the real image. According to the invention, the problem of high complexity of an image generation method in the prior art is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of data processing technology, and in particular to an image generation method and device. Background Art

[0002] The conditional generative adversarial network (CGAN) is an extension of the generative adversarial network (GAN). It introduces conditional information to control the properties of the generated samples. The conditional generative adversarial network consists of a generator and a discriminator. The generator generates data based on the conditional information, while the discriminator determines whether the generated data is real based on the conditional information.

[0003] Conditional generative adversarial networks are usually used in related technologies to generate or convert various images that meet the conditions. However, traditional conditional generative adversarial networks need to consider information on more time steps when processing long sequences. The generator and discriminator need to operate on more time steps, which increases the complexity of the calculation.

[0004] With regard to the problem of high complexity of image generation methods in the prior art, no effective technical solution has been proposed yet. Summary of the invention

[0005] The embodiments of the present application provide an image generation method and device to at least solve the problem of high complexity of the image generation method in the prior art.

[0006] According to one aspect of an embodiment of the present application, there is provided an image generation method, comprising: obtaining reference condition data, and performing a preprocessing operation on the reference condition data to obtain target condition data, wherein the preprocessing operation is used to improve the data quality of the reference condition data; inputting the target condition data into an image generation model, and outputting a target image, wherein the image generation model is a model obtained by pre-training an initial image generation model using a training sample set, the image generation model comprises a generator and a discriminator, the generator processes the target condition data through a downsampling module, a visual attention module, and an upsampling module to generate a reference image, and the discriminator outputs discrimination information when taking the reference image as input, and the discrimination information is used to indicate the discrimination result between the reference image and the real image.

[0007] According to another aspect of an embodiment of the present application, an image generating device is also provided, including: an acquisition unit, used to acquire reference condition data, and perform a preprocessing operation on the reference condition data to obtain target condition data, wherein the preprocessing operation is used to improve the data quality of the reference condition data; an input unit, used to input the target condition data into an image generation model, and output a target image, wherein the image generation model is a model obtained by pre-training an initial image generation model using a training sample set, and the image generation model includes a generator and a discriminator, the generator processes the target condition data through a downsampling module, a visual attention module, and an upsampling module to generate a reference image, and the discriminator outputs discrimination information when taking the reference image as input, and the discrimination information is used to indicate the discrimination results between the reference image and the real image.

[0008] Optionally, the above-mentioned image generation device also includes a training unit, which is used to train the initial image generation model using a training sample set; the training unit includes: a first input subunit, which is used to input the sample condition data in the training sample set into the initial generator in the initial image generation model to generate a sample generated image corresponding to the sample condition data; a second input subunit, which is used to input the sample generated image and the sample real image corresponding to the sample condition data into the initial discriminator in the initial image generation model to obtain a discrimination result; a first determination subunit, which is used to determine the training loss value according to the discrimination result and a preset loss determination method, and alternately correct the initial generator and the initial discriminator based on the loss value; a second determination subunit, which is used to determine the initial generator as a generator and the initial discriminator as a discriminator when the initial generator and the initial discriminator meet the preset conditions to obtain the image generation model.

[0009] Optionally, the above-mentioned initial generator includes an initial downsampling module, which is used to perform downsampling operations on the sample condition data for a preset number of times to obtain downsampled data; an initial visual attention module, which is used to extract data from the downsampled data to obtain data of interest; and an initial upsampling module, which performs upsampling operations on the data of interest for a preset number of times to obtain a sample generated image.

[0010] Optionally, the above-mentioned initial visual attention module includes a pooling submodule, which is used to perform average pooling and maximum pooling on the downsampled data, respectively, to obtain average pooling data and maximum pooling data; a connection submodule, which is used to connect the average pooling data and the maximum pooling data to obtain connection feature data; a convolution module, which is used to perform target convolution operations on the average pooling data, the maximum pooling data and the connection feature data to obtain convolution feature data; an input submodule, which is used to input the convolution feature data into the visual state space block for visual data processing and output data of interest.

[0011] Optionally, the convolution module includes a convolution submodule, which is used to perform convolution operations on average pooling data, maximum pooling data, and connection feature data, respectively, to obtain first convolution data corresponding to the average pooling data, second convolution data corresponding to the maximum pooling data, and third convolution data corresponding to the connection feature data; a first product calculation module, which is used to multiply the first convolution data with the down-sampled data to obtain first product data, and multiply the second convolution data with the down-sampled data to obtain second product data; a second product calculation module, which is used to multiply the first product data with the third convolution data to obtain third product data, and multiply the second product data with the third convolution data to obtain fourth product data, wherein the convolution feature data includes the third product data and the fourth product data.

[0012] Optionally, the above-mentioned visual state space block includes a normalization module, which is used to normalize the convolution feature data to obtain normalized data; a target processing module, which is used to perform a target processing operation on the normalized data to obtain target processed data, and the target processing operation is used to enhance the data characteristics of the normalized data; a reference processing module, which is used to perform a reference processing operation on the normalized data to obtain reference processed data, and the reference processing operation is used to enhance the generalization ability of the model; and a residual connection module, which is used to perform a residual connection operation on the reference processed data, the target processed data and the convolution feature data to obtain data of interest.

[0013] Optionally, the above-mentioned first determination subunit includes: an acquisition module, which is used to perform feature acquisition operations according to the downsampling process of the downsampling module and the upsampling process of the upsampling module to obtain multiple groups of acquired features; a first determination module, which is used to determine the first loss value based on the multiple groups of acquired features and the discrimination results; a second determination module, which is used to determine the second loss value based on the generation process of the sample-generated image of the initial generator, and determine the third loss value based on the generation process of the sample-generated image and the discrimination process of the initial discriminator; a third determination module, which is used to determine the training loss value based on a preset weight set, a first loss value, a second loss value and a third loss value, wherein the preset weight set includes preset weights corresponding to the first loss value, the second loss value and the third loss value, respectively.

[0014] According to another aspect of the embodiments of the present application, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the above image generation method.

[0015] According to another aspect of the embodiments of the present application, an electronic device is also provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the image generation method as described above.

[0016] The above-mentioned image generation method solves the problem of high complexity of the image generation method in the prior art and simplifies the complexity of the image generation method. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] In order to more clearly illustrate the specific implementation methods of the present application or the technical solutions in the prior art, the drawings required for use in the specific implementation methods or the description of the prior art will be briefly introduced below. Obviously, the drawings described below are only some implementation methods of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative work.

[0018] Figure 1 is a schematic diagram of a hardware environment of an optional image generation system according to an embodiment of the present invention;

[0019] Figure 2 is a flow chart of an optional image generation method according to an embodiment of the present invention;

[0020] Figure 3 is a schematic diagram of an optional image generation method according to an embodiment of the present invention;

[0021] Figure 4 is a schematic diagram of another optional image generating method according to an embodiment of the present invention;

[0022] Figure 5 is a schematic diagram of another optional image generating method according to an embodiment of the present invention;

[0023] Figure 6 is a schematic diagram of a network structure of an optional VSS block according to an embodiment of the present invention;

[0024] Figure 7 is an optional SS2D schematic diagram according to an embodiment of the present invention. DETAILED DESCRIPTION

[0025] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present application.

[0026] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and are not necessarily used to describe a specific order or sequence. In addition, the terms "including" and "having" and any of their variations are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices. It should be noted that, in the absence of conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments.

[0027] Before introducing the technical solution of the present application, some background technical knowledge involved in the present application is first introduced and explained. The following related technologies can be arbitrarily combined with the technical solution of the embodiment of the present application as optional solutions, and they all belong to the protection scope of the embodiment of the present application. The embodiment of the present application includes at least part of the following contents.

[0028] GANs: Generative adversarial networks, which consist of a generative model and a discriminative model. The generative model is responsible for capturing the distribution of sample data, while the discriminative model is generally a binary classifier that determines whether the input is real data or a generated sample. The optimization process of this model is a "binary minimax game" problem. During training, one of the parties (the discriminative network or the generative network) is fixed, and the parameters of the other model are updated, alternating and iterating. Finally, the generative model can estimate the distribution of sample data.

[0029] CGAN: Conditional Generative Adversarial Network, is a variant of Generative Adversarial Network, which introduces additional conditional information when generating data. In traditional generative adversarial networks, the adversarial training between the generator and the discriminator is unconditional, that is, the generator simply generates data from random noise, while the discriminator tries to distinguish between the generated data and the real data. In contrast, the conditional generative adversarial network allows the user to specify the specific conditions of the data to be generated. By inputting conditional information into the generator and the discriminator, the conditional generative adversarial network can learn to generate more structured and diverse data under given conditions.

[0030] In the prior art, conditional generative adversarial networks are usually used to generate or convert various images that meet the conditions. However, traditional conditional generative adversarial networks need to consider more information on time steps when processing long sequences, and the generator and discriminator need to operate on more time steps, which increases the complexity of the calculation. In particular, when converting visible light images into infrared images, due to the limitations of application scenarios and guarantee capabilities, the reference images in infrared imaging guidance are usually visible light images, while the real-time images are infrared images. The imaging mechanisms of infrared images and visible light images are different, resulting in significant differences in their features, which in turn increases the difficulty of scene matching in infrared imaging guidance. Using infrared simulation technology to generate infrared features of scenes in the required environment can not only effectively reduce the cost of obtaining infrared data, but also generate infrared data that is difficult to obtain in field tests under a variety of natural environments and scene conditions. The generated infrared data can be applied to aviation, navigation, meteorology, geology, and agriculture, providing basic and reliable data for tasks such as detection, classification, positioning, identification, and tracking.

[0031] At present, the mainstream infrared image simulation technology can be divided into two categories: one is the image simulation technology based on infrared characteristic modeling, and the other is the infrared image simulation technology based on deep learning. The former is based on the establishment of a mathematical model related to the infrared characteristics of the scene, and the infrared simulation effect of the scene is displayed through computer simulation; the latter uses a deep learning algorithm to train a large number of visible light scenes and their infrared characteristic effects to obtain a mapping model between the visible light and infrared image of the scene, and then generates the infrared characteristics of the visible light scene based on this model and displays it in the form of an image.

[0032] Infrared simulation technology based on infrared characteristic modeling can generate better target infrared texture, but its automation level is not high, and it has deficiencies in generating infrared characteristics of target scenes under interference such as smoke and shadows, batch processing of data, etc. Deep learning has shown remarkable performance in feature extraction, spatial transformation, data fitting, etc. in image processing. Deep generative models are an important component of image generation based on deep learning, among which generative adversarial networks (GANs) show superior performance in image transformation than other generative models.

[0033] According to the different types of application data, the conversion algorithms of visible light to infrared images based on the GANs architecture can be divided into paired (i.e., visible light and infrared images of the same scene) and unpaired types. The classic methods of visible light to infrared image conversion based on the Pix2Pix framework include ThermalGAN, LayerGAN, InfraGAN, and IR-GAN, etc. These methods are aimed at paired data. Among them, ThermalGAN and LayerGAN generate infrared images through multimodal data generation. The ThermalGAN algorithm generates infrared images in two steps: first, the average temperature infrared image of the target is generated using the visible light image and the temperature vector, and then the average temperature infrared image of the target and the visible light image are used to generate a more refined infrared image. LayerGAN includes two methods for generating infrared images: one method uses temperature vectors, semantic segmentation images, and thermal segmentation images to generate infrared images; the other method uses visible light images, semantic segmentation images, and thermal segmentation images to generate infrared images. The LayerGAN algorithm is of great significance for the task of pedestrian re-identification based on infrared images, but it has high data requirements and requires multimodal data. InfraGAN and IR-GAN strengthen the constraints of edge information of generated infrared images by constructing new generative networks and loss functions, which effectively alleviates the problem of edge distortion of generated infrared images. However, these two methods do not deeply study the feature mapping relationship between visible light images and infrared images.

[0034] Classic methods of visible-to-infrared image conversion models based on the CycleGAN framework include: SIR-GAN, DAGAN, and CVIIT, etc. This type of method targets unpaired data. SIR-GAN takes traditional infrared simulation data as input to obtain infrared images with rich texture information. DAGAN uses a novel dual attention mechanism to automatically predict flashovers in indoor fires using visible-to-infrared conversion. CVIIT introduces contrast loss to ensure the consistency of content between the generated image and the source image, which can improve the quality of image conversion under reduced light conditions. The edge-guided multi-domain visible-to-infrared image conversion algorithm retains more details of the infrared image by constraining the consistency of edge information between the generated image and the source image. Although these methods relax the requirements on data, they produce inferior image results compared to methods based on paired data.

[0035] Although generative adversarial networks (GANs) have been successful in converting visible light images to infrared images, they still have some shortcomings. The main problems include the different imaging mechanisms of visible light and infrared images, which leads to a decrease in the accuracy of feature mapping of GANs, resulting in less than ideal fitting results. In addition, the generated infrared images also have problems such as poor texture consistency and loss of details.

[0036] In order to solve the above problems, the embodiments of the present application provide an image generation method and device. The image generation method in the embodiments of the present application is an image generation method based on Mamba and multi-scale feature contrast learning. When the conditional data in the reference conditional data is a visible light image, and the target image is an infrared image, the image generation method in the embodiments of the present application can also be understood as an image generation method from a visible light image to an infrared image based on Mamba and multi-scale feature contrast learning. The present application names this method V2I-GAN. Based on the CGAN framework, V2I-GAN introduces the Mamba attention module in the conditional generation network model (i.e., the above-mentioned conditional generation adversarial network). This module constructs a feature mapping relationship between visible light images and infrared images based on the state space, which can enable the image generation model to focus on the key data in the conditional data during the processing of the target conditional data (when the conditional data is the above-mentioned visible light image, it can also be understood as focusing on the key areas of the visible light image). As an optional implementation, the above-mentioned image generation method can be applied to, but is not limited to, Figure 1 The hardware environment diagram of the image generation system composed of the terminal device 102 and the server 104 is shown in FIG. Figure 1 As shown, the terminal device 102 is connected to the server 104 via a network 110 , and the network 110 may include but is not limited to: a wired network and a wireless network.

[0037] The terminal device 102 is also provided with a display 106, a processor 108 and a memory 112. The display 106 can be used to display the target image and, when the reference condition data is an image, to display the reference condition data. The processor 108 can be used to perform data processing on the reference condition data, and the memory 112 can be used to store relevant data information involved in this application.

[0038] The server 104 may be a single server, a server cluster consisting of multiple servers, or a cloud server. The server 104 includes a database 114 and a processing engine 116. The database 114 may be used to store reference condition data, target condition data, target images, etc., and the processing engine 116 may be used to process the reference condition data.

[0039] According to one aspect of the embodiment of the present invention, the image generation system may further perform the following steps: First, the terminal device 102 executes S102 to send an image generation request to the server 104 via the network 110; then, the server 104 executes Figure 1As shown in S104 to S106, reference condition data is obtained, and preprocessing operations are performed on the reference condition data to obtain target condition data, wherein the preprocessing operation is used to improve the data quality of the reference condition data; the target condition data is input into the image generation model, and the target image is output, wherein the image generation model is a model obtained by pre-training the initial image generation model with a training sample set, and the image generation model includes a generator and a discriminator, the generator processes the target condition data through a downsampling module, a visual attention module, and an upsampling module to generate a reference image, and the discriminator outputs discrimination information when the reference image is used as input, and the discrimination information is used to indicate the discrimination result between the reference image and the real image. Finally, S108 is executed to send the target image to the terminal device 102.

[0040] The above-mentioned image generation method simplifies the complexity of image generation, thereby solving the problem of high complexity of image generation methods in the prior art.

[0041] The above is only an example and is not limited in this embodiment.

[0042] As an optional implementation, please refer to Figure 2 The flowchart of the image generation method shown in FIG. The method may include the following steps:

[0043] S202, obtaining reference condition data, and performing a preprocessing operation on the reference condition data to obtain target condition data, wherein the preprocessing operation is used to improve the data quality of the reference condition data;

[0044] S204, input the target condition data into the image generation model, and output the target image, wherein the image generation model is a model obtained by pre-training the initial image generation model with a training sample set, and the image generation model includes a generator and a discriminator, the generator processes the target condition data through a downsampling module, a visual attention module, and an upsampling module to generate a reference image, and the discriminator outputs discrimination information when taking the reference image as input, and the discrimination information is used to indicate the discrimination result between the reference image and the real image.

[0045] It should be noted that the reference condition data in the above S202 includes condition data and noise data (or referred to as noise data), and the condition data includes multiple types, for example, the condition data can be any label information, such as category labels (such as image categories), text descriptions, images (such as visible light images, facial expression images of faces), etc. The above preprocessing operations specifically include: S202-1, standardizing and normalizing the reference condition data to obtain standard data, and the standardization and normalization are used to adjust the data scale of the condition data to a scale that is unified with the sample condition data in the training sample set, which helps the generator to generate the final target image more stably; S202-1, using the condition data in accordance with the condition data. The standard data is encoded using an encoding method corresponding to the data type of the conditional data to obtain encoded data. The encoding process avoids the output of the generator being affected by the data type problem (for example, when the conditional data in the reference conditional data is discrete data such as category labels, one-hot encoding is used to convert the reference conditional data into a data format suitable for input to the generator); S202-3, a data enhancement operation is performed on the encoded data using a data enhancement method corresponding to the data type of the conditional data to obtain target conditional data. The data enhancement operation increases the diversity of the conditional data, which helps the generator learn richer features, thereby generating a target image that better meets user expectations. For example, when the conditional data is of image type, the coded data is rotated (rotated at different angles), translated (moved in the horizontal and vertical directions), scaled (enlarged or reduced coded images, images also belong to data, therefore, in the technical solution of the present application, when the conditional data is of image type, the corresponding data can be called images or regions, for example, the conditional data can be called conditional images, the data of interest can be called regions of interest, and so on), cropped (randomly select sub-regions of the image corresponding to the coded data), etc.; when the conditional data is a category label, the coded data corresponding to the conditional data is data enhanced, and the coded data corresponding to the noise data is data weakened, but the intensity of data enhancement and data weakening is a pre-set appropriate intensity, which will not cause a certain data to be too strong or a certain data to be too weak, thereby avoiding the inaccurate generated target image. Through the above preprocessing operations, the image generation model can better utilize the target conditional data obtained by the preprocessing operation to generate a target image that better meets specific requirements. Then the above preprocessing operations can not only be used to improve the data quality of the reference conditional data, but also can be used to improve the accurate feature extraction of the data by the generator in the subsequent image generation model.

[0046] The image generation model in S204 above includes a generator and a discriminator, and the generator in this solution includes a downsampling module, a visual attention module and an upsampling module. The target condition data is processed by the downsampling module, the visual attention module and the upsampling module to generate a reference image. The discriminator outputs discrimination information indicating the discrimination result between the reference image and the real image when the reference image is used as input; in addition, the reference image is also input with a real image, and the discriminator compares the reference image with the real image to obtain the probability that the reference image is the real image (i.e., the discrimination result). Finally, the image generation model can output the target image according to the discrimination result, and outputting the target image according to the probability specifically includes determining the reference image as the target image for output when the discrimination result indicates that the probability of the reference image being the real image is greater than the preset probability (this operation is to avoid the problem of inaccurate output of the target image due to system or other reasons, so it is also designed to allow the discriminator in the image generation model to generate a discrimination result, and finally output the target image according to the discrimination result, so as to avoid the problem of inaccurate output of the target image due to system reasons). The downsampling module, visual attention module and upsampling module included in the generator of the above image generation model are all obtained by pre-training the initial image generation model with a training sample set, specifically, the initial upsampling module, initial visual attention module and initial downsampling module included in the initial generator of the initial image generation model are trained to be the downsampling module, visual attention module and upsampling module included in the generator of the final image generation model. The target image in the above S204 can be understood as, but not limited to, an image that meets the conditions indicated by the conditional data, such as an infrared image.

[0047] The discriminator of the V2I-GAN algorithm (the image generation method of this application) is constructed by PatchGAN. The Markov discriminator is a discriminative model composed of convolutional layers, whose output is a matrix of size N×N, and the average value of the matrix is ​​used as the true or false output. Since each value in the output matrix corresponds to a receptive field in the original image, that is, to a regional block (patch) in the original image, the generative adversarial network (GANs) with this structure is also called PatchGAN. Taking an image of size 256×256 as an example, traditional GANs map the 256×256 image to a single scalar output, representing "true" or "false", but this is not easy to reflect the local features of the image. In contrast, PatchGAN maps the 256×256 image to a matrix of size N×N, and each N×N represents whether the regional block in the image is true or false. PatchGAN realizes the extraction and representation of local image features, enabling the model to pay more attention to the detailed information of the image, which is conducive to generating higher quality images.

[0048] It should be noted that after obtaining the target image, the present application also uses t-SNE technology to visualize the feature distribution of the image, such as the real infrared image, the generated reference image, and when the conditional data is a visible light image, it also includes the visualized visible light image.

[0049] The following takes the reference condition data as a visible light image and the target image as an infrared image as an example. Figure 3 The process of generating target images by the above image generation model is described as follows:

[0050] like Figure 3 As shown, first, the reference condition data includes a visible light image (i.e., the above-mentioned condition data) and random noise (i.e., the above-mentioned noise data), and the reference condition data is preprocessed to obtain the target condition data; the target condition data is input into the image generation model, and the generator in the image generation model includes a downsampling module, a visual attention module, and an upsampling module to process the target condition data to generate a reference image, and then the real image and the reference image are input into the discriminator, and the discriminator discriminates the real image and the reference image and outputs a discrimination result, and finally the infrared image output according to the discrimination result is the target image.

[0051] The image generation model of the generator in the conditional generative adversarial network obtained by training the initial image generation model of the internal structure of the generator in the conditional generative adversarial network not only outputs a more accurate generated image, but also simplifies the complexity of image generation.

[0052] As an optional implementation, the above-mentioned training of the initial image generation model using a training sample set includes: S1, inputting the sample condition data in the training sample set into the initial generator in the initial image generation model to generate a sample generated image corresponding to the sample condition data; S2, inputting the sample generated image and the sample real image corresponding to the sample condition data into the initial discriminator in the initial image generation model to obtain a discrimination result; S3, determining the training loss value according to the discrimination result and a preset loss determination method, and alternately correcting the initial generator and the initial discriminator based on the loss value; S4, when the initial generator and the initial discriminator meet the preset conditions, determining the initial generator as a generator, and determining the initial discriminator as a discriminator to obtain an image generation model.

[0053] It should be noted that before using the training sample set to train the initial image generation model, it is necessary to construct an initial generator and an initial discriminator, the initial generator includes an initial upsampling module, an initial visual attention module and an initial downsampling module; the initial image generation model is constructed according to the initial generator and the initial discriminator; then the training sample set in S1 above is obtained, the training sample set includes multiple sample condition data, and also includes a sample real image corresponding to each sample condition data. The sample condition data is sample data corresponding to the reference condition data, each piece of sample condition data includes sample conditions corresponding to the condition data in the reference condition data, and also includes sample noise data corresponding to the noise data in the reference condition data, and the data input to the initial discriminator in the initial image generation model is the sample generated image generated by the initial generator and the sample real image corresponding to the sample condition data input to the initial generator; the initial discriminator will obtain a discrimination result based on the sample generated image and the sample real image, and then the training loss value can be determined according to the discrimination result and the preset loss method, and the initial generator and the initial discriminator are alternately corrected based on the loss value, so that the image generated by the initial generator is closer to the real image, and the discrimination result obtained by the initial discriminator is more accurate. When the initial generator and the initial discriminator in the initial image generation model reach a balance (that is, the above-mentioned initial generator and the initial discriminator both meet the preset conditions), the initial image generation model at this time is determined as the image generation model, and at this time, the initial generator in the initial image generation model is determined as the generator in the image generation model, and the discriminator in the initial image generation model is determined as the discriminator in the image generation model.

[0054] The operations in S3 specifically include: performing feature acquisition operations according to the downsampling process of the downsampling module and the upsampling process of the upsampling module to obtain multiple sets of acquired features; determining a first loss value according to the multiple sets of acquired features and the discrimination results; determining a second loss value according to the generation process of the sample-generated image of the initial generator. And determining a third loss value according to the generation process of the sample-generated image and the discrimination process of the initial discriminator; determining a training loss value according to a preset weight set, a first preset loss value, a second preset loss value, and a third loss value, wherein the preset weight set includes preset weights corresponding to the first loss value, the second loss value, and the third loss value, respectively.

[0055] The first loss value mentioned above can be understood as, but not limited to, a multi-scale feature contrast loss, the second loss value can be understood as, but not limited to, an L1 loss, and the third loss can be understood as, but not limited to, a conditional generative adversarial network loss. The following is a detailed description of the three losses:

[0056] The following is an example of the input conditional data being a visible light image and the generated target image being an infrared image to illustrate the method of determining the above training loss value: In order to enhance the detail features of the generated image, this application introduces a method based on a multi-scale feature comparison learning strategy in the process of training the initial image generation model, such as Figure 4 As shown, compared with the existing method for determining the multi-scale feature contrast loss, the multi-scale feature contrast loss method in this application only requires a set of encoders and decoders (such as Figure 4 As shown, there are three arrows from the encoder and decoder pointing to the attention-guided sampling respectively. This part can be understood as each layer in the network structure of the encoder and decoder pointing to the attention-guided sampling respectively, corresponding to the three downsamplings and three upsamplings mentioned above, and "three times" can be understood as the three-layer convolutional network structure in the generator), which simplifies the training process to a certain extent.

[0057] The conditional generative adversarial network consists of an encoder and a decoder, each consisting of L stages. For ease of interpretation, the input x is represented as h 0 , through the encoding part, the output result consists of L stages, including a series of encoding features of different sizes, namely The encoded features output by the encoder are fed into the decoder, which provides a series of decoded features, namely The outputs of the corresponding layers of the encoder and decoder have the same dimension, i.e., h l and h 2L-1 have the same size, denoted as (h l ,h l ), which represents a pair of encoding-decoding features with the same size.

[0058] like Figure 4 As shown, for any pair of encoding-decoding features (h l ,h l ), sample a query vector from the position of the output feature map of the decoder at stage l, sample positive samples from the corresponding position of the encoder, and sample N negative samples from other positions of the encoder. Then, use a two-layer MLP (multi-layer perceptron) with a shared linear layer to transform the query vector, positive samples, and negative samples into a K-dimensional embedding space, and obtain vectors q, respectively, and is a set of real numbers. In order to avoid the collapse of contrastive learning mode, the vector q, and Mapped to the unit sphere through L2 normalization. Contrastive learning focuses on making the query vector closer to the positive sample in the embedding space, but away from the negative sample, which represents a (N+1) class classification problem. Therefore, the cross quotient function is used in this application for calculation, which represents the probability of selecting a negative sample as a positive sample. The function is as follows:

[0059]

[0060] where · represents the dot product of the vector; τ is the temperature coefficient used to scale the distance between the query vector and the positive and negative samples, which can be set to 0.07, for example. By extending its application to multi-scale features, (h l ,h l ) can be expressed as:

[0061] The total loss value (i.e. the training loss value mentioned above) includes L1 loss and LSGAN loss in addition to the above-mentioned multi-scale feature contrast learning loss, which are defined as: L1 loss:

[0062] Conditional Generative Adversarial Network Loss:

[0063]

[0064] The total loss value is:

[0065] in, is the contrastive learning loss; E(·) represents the expected operation; x∈X in the following table represents the data from the visible light image; y∈Y in the following table represents the data from the corresponding real infrared image; y represents the label information from the real infrared image; D(y) is the probability that the discriminator can accurately judge that the real data is real; G(x) represents the target domain image (i.e., infrared image) generated based on the source domain image (i.e., visible light image); D(G(x)) is the probability that the discriminator can accurately judge that the generated data is real data, λ 1 , 2 , 3 They are the weights of LSGAN (least squares loss), L1 and contrastive learning loss respectively.

[0066] By training the initial image generation model in the above-mentioned manner of the present application, a relatively accurate image generation model of the output target image can be obtained. At the same time, by adopting the multi-scale feature contrast learning loss, the generated target image can be constrained at different scales to reduce the problem of detail loss. And the multi-scale feature contrast learning loss in the present application utilizes the attention map of the discriminator to strengthen the learning of difficult samples.

[0067] As an optional implementation, the above-mentioned input of sample condition data in the training sample set into the initial generator in the initial image generation model to generate a sample generated image corresponding to the sample condition data includes: S1, the initial downsampling module in the initial generator performs a downsampling operation on the sample condition data for a preset number of times to obtain the downsampled data; S2, the initial visual attention module in the initial generator performs data extraction on the downsampled data to obtain the data of interest; S3, the initial upsampling module in the initial generator performs an upsampling operation on the region of interest for a preset number of times to obtain the sample generated image.

[0068] It should be noted that the preset times in the above S1 and S3 are determined according to the structure of the initial generator (that is, determined according to the network structure of the initial generator). Assuming that the initial generator adopts the U-NET network structure, then in the process of downsampling the sample conditional data, how many downsampling operations (such as convolution, a kind of downsampling operation) will be performed, and when upsampling the data of interest, how many upsampling operations will be performed, and the specific number of upsampling or downsampling operations depends on the designed U-NET network structure.

[0069] It can be understood that, when the condition data in the reference condition data is the image (e.g., a visible light image), the region of interest in S2 can also be referred to as the region of interest. The sample generated image is, for example, an infrared image, that is, the condition data in the reference condition data can be a visible light image, and the target image is an infrared image, then the image generation method in the present application can also be understood, but not limited to, as an image conversion method for converting a visible light image into an infrared image.

[0070] Through the above implementation of the present application, the sample condition data is processed by using an initial generator constructed differently from the existing generator, thereby generating a sample generated image, further avoiding the problem that due to the different imaging mechanisms of visible light and infrared images, the accuracy of feature mapping is reduced when using the existing conditional generative adversarial network to convert visible light images into infrared images, thereby resulting in an unsatisfactory fitting effect. At the same time, the problem of poor texture consistency and loss of details in the generated infrared image is avoided.

[0071] As an optional implementation, the initial visual attention module in the above-mentioned initial generator extracts data from the downsampled data to obtain data of interest, including: S1, performing average pooling and maximum pooling on the downsampled data respectively to obtain average pooling data and maximum pooling data; S2, connecting the average pooling data and the maximum pooling data to obtain connection feature data; S3, performing target convolution operation on the average pooling data, the maximum pooling data and the connection feature data to obtain convolution feature data; S4, inputting the convolution feature data into the visual state space block for visual data processing, and outputting the data of interest.

[0072] The average pooling in S1 above averages all values ​​in the neighborhood, thereby retaining the background features of the data (for example, retaining the background information of the image) to obtain an average representation of the overall data features (for example, if the above conditional data is an image, then the downsampled data is still an image, then the average pooling can be understood as dividing the downsampled data into multiple sub-images, and then calculating the average of all pixel values ​​in each sub-image, and using the average as the output value of the sub-image, thereby reducing the amount of calculation); the maximum pooling selects the maximum value in the neighborhood as the output, thereby retaining the texture features of the data (for example, if the above conditional data is an image, then the downsampled data is still an image, then the maximum pooling can be understood as dividing the downsampled data into multiple sub-images, and then selecting the maximum pixel value from each sub-image as the output value of the sub-image, thereby sampling and reducing the size and features of the downsampled data, which can also reduce the amount of calculation).

[0073] The connection operation in the above S2 can be understood, but is not limited to, as connecting the average pooling data and the maximum pooling data to obtain the overall connection feature data; then performing a target convolution operation on the average pooling data, the maximum pooling data and the connection feature data to obtain the convolution feature data, and finally performing subsequent visual data processing on the convolution feature data, and the visual data processing here can be understood, but is not limited to, as extracting the data of interest from the convolution feature data.

[0074] Through the above-mentioned implementation of the present application, more accurate data features (ie, the above-mentioned data of interest) can be extracted from the downsampled data, thereby improving the accuracy of the target image generated by the image generation model.

[0075] As an alternative implementation, the above-mentioned target convolution operation on the average pooling data, the max pooling data, and the concatenated feature data to obtain the convolution feature data includes: S1, performing convolution operations on the average pooling data, the max pooling data, and the concatenated feature data respectively to obtain a first convolution operation corresponding to the average pooling data, a second convolution operation corresponding to the max pooling data, and a third convolution operation corresponding to the concatenated feature data; S2, multiplying the first convolution data by the downsampled data to obtain a first product data, and multiplying the second convolution data by the downsampled data to obtain a second product data; S3, multiplying the first product data by the third convolution data to obtain a third product data, and multiplying the second product data by the third convolution data to obtain a fourth product data, where the convolution feature data includes the third product data and the fourth product data.

[0076] The convolution operation in the above S1 can also be referred to as convolution (a mathematical operation). The operations in the above S1 are to perform convolution on the average pooling data to obtain the first convolution data; perform convolution on the max pooling data to obtain the second convolution data; perform convolution on the concatenated feature data to obtain the third convolution data; and then perform S2, which is to respectively calculate the products of the first convolution data and the second convolution data with the downsampled data to obtain the product of the first convolution data and the downsampled data (i.e., the above-mentioned first product data) and the product of the second convolution data and the downsampled data (i.e., the above-mentioned second product data).

[0077] In the above S3, the first product data and the second product data are respectively multiplied by the third convolution data to obtain the third product data and the fourth product data. Then, the third product data and the fourth product data are the finally obtained convolution feature data.

[0078] The following takes the conditional data as a visible light image and the generated target image as an infrared image as an example, and combines Figure 5 (To simplify the drawings, the above-mentioned conditional data and the process of preprocessing the reference conditional data are omitted in this figure. Figure 5 The visible light image shown is the conditional data in the target conditional data obtained after preprocessing) to elaborate on the process of the above-mentioned generator (or initial generator) processing the input target conditional data (or sample conditional data) in detail:

[0079] The Visual State Space (VSS) block integrates the advantages of the Selective State Space Models (SSMs) into visual data processing, including a global perception field, input-dependent weight parameters, and linear computational complexity, and effectively extracts features from visible light images. Drawing on the performance advantages of the VSS block, the present application designs a VSS attention module (i.e., a visual attention module, with a structure as Figure 5As shown in Figure 2.1, the V2I-GAN generator is mainly composed of three parts: downsampling module, VSS attention module and upsampling module. The network structure is shown in Figure 2.1. Figure 5 As shown in the figure, the specific working process of the generator can be described as follows: the input visible light image x first undergoes three downsampling operations (such as Figure 5 The image size is continuously reduced, while the number of channels is continuously increased, thereby obtaining the features of the visible light image. The extracted features are passed through the VSS attention module to obtain the features of the region of interest in the visible light image. Finally, after three upsampling (such as Figure 5 After the operation in the decoder shown in the figure), the image size continues to increase and the number of channels gradually decreases, and it is reconstructed into an infrared image

[0080] The visible light image is first downsampled to obtain f x , and then the input of the VSS attention module is processed by adaptive average pooling and adaptive maximum pooling to obtain Avg(f x ) and Max(f x ), after the convolution process, they are multiplied with the original input to obtain f x *Conv(Avg(f x )) and f x *Conv(Max(f x )). Then the result is combined with the convolutional feature Conv(Cat(Avg(f x ),Max(f x ))) is multiplied and added to the input VSS block. The final output feature of the VSS attention module before entering the VSS block can be specifically expressed as:

[0081] Through the above-mentioned implementation of the present application, the data calculation process is simplified, while data loss is avoided, and the accuracy of the obtained convolution feature data is guaranteed. At the same time, the visual attention module and the multi-scale feature contrast learning loss can promote each other to generate high-quality and reliable infrared images.

[0082] As an optional implementation, the above-mentioned input of convolution feature data into the visual space block for visual data processing and output of data of interest includes: S1, the visual state space block normalizes the convolution feature data to obtain normalized data; S2, performs a target processing operation on the normalized data to obtain target processed data, and the target processing operation is used to enhance the data characteristics of the normalized data; S3, performs a reference processing operation on the normalized data to obtain reference processed data, and the reference processing operation is used to enhance the generalization ability of the model; S4, performs a residual connection operation on the reference processed data, the target processed data and the convolution feature data to obtain the data of interest.

[0083] It should be noted that the above-mentioned visual state space block (hereinafter referred to as VSS block) is one of the most important components of the above-mentioned visual attention module (hereinafter referred to as VSS attention module). The VSS block originates from VMamba and is a key component of the generation network. After the convolution feature data is input into the visual state space block, the visual state space block will perform a series of data processing on the convolution feature data, including the above-mentioned normalization, target processing operation, reference processing operation, residual connection, etc. The above-mentioned target processing operation and reference processing operation can be understood, but not limited to, as processing on two processing paths for normalized data. The specific processing processes on different processing paths are different. The target processing operation is the operation on the main path, and the reference processing operation is the operation on the secondary path. In this application, the operation on the main path includes layer normalization operation, linear transformation operation, depth separable convolution operation, operation corresponding to the activation function, selective scanning operation, layer normalization operation; the operation on the secondary path only includes linear transformation operation and operation corresponding to the activation function. Finally, the data on different processing paths and the initial convolution feature data are residually connected to obtain the final data of interest. It should be noted that, before the above S4, the target processing data and the reference processing data are also multiplied pixel by pixel and linearly transformed, and the operation in the above S4 can be understood, but not limited to, as a residual connection operation between the data obtained after the linear transformation and the normalized data.

[0084] The following combination Figure 6 to Figure 7 The above S1 to S4 (i.e., the processing in the VSS block) are described in detail. Figure 6As shown in the figure, after the convolution feature data is input into the visual state space block, the convolution feature data is firstly subjected to layer normalization (used to standardize the input of each layer so that its mean is 0 and its variance is 1, thereby reducing the input distribution variation between layers and improving the training effect and generalization ability of the model) to obtain normalized data; then it is divided into two paths for processing. In the side path (the path corresponding to the reference processing operation), the normalized data is processed by the linear layer (i.e., linear transformation is performed to map the normalized data to a new space, change the dimension of the data, and facilitate further processing by the subsequent layers) and the block with SiLU activation function. In the main path (the path corresponding to the target processing operation), the normalized data is first processed by the linear layer, the depth-separable convolution block and the SiLU activation function block, and then guided by the 2D-Selective-Scan (SS2D) module to enhance feature extraction. Next, the extracted features are layer normalized and merged with the results of the side path by pixel-by-pixel multiplication. Finally, a linear layer is used to mix the features, and the output features (i.e., the data obtained after pixel-by-pixel multiplication of the reference processed data and the target processed data) are residually connected with the input data (i.e., the normalized data) to form the output of the VSS block.

[0085] Selective Scan 2D (SS2D) consists of three parts: a scan-expand operation (expanding the input data into sequences along four different traversal paths), an S6 block (i.e., selective scanning using an S6 block, processing each sequence separately), and a scan-fusion operation (i.e., reshaping and merging the resulting sequences to form an output map), as shown in Figure 7 As shown (to avoid unclear image numbers, the order of image numbers has been noted after each image). Figure 7 As shown in (a), the scan expansion operation expands the input image into sequences along four different directions from top left to bottom right, from bottom right to top left, from top right to bottom left, and from bottom left to top right. These sequences are processed by the S6 block to extract features, ensuring that a thorough scan is performed from all four directions to capture various features. Figure 7 As shown in (b), the scan fusion operation sums and merges the sequences from the four directions, restoring the output image to the same size as the input image. The specific process can be described as follows: given the input feature s (the data feature obtained after the activation function), it is processed by SS2D until the output feature s 0 It can be described as:

[0086] s d = expand(s,d) Formula 1: Scan expansion

[0087] S6 Block

[0088] Scan Fusion

[0089] Among them, d in formula 1 is the expansion direction, s d is the result obtained by scanning expansion, and The sequences in the four directions obtained by scanning and expansion are respectively scanned and expanded and then selectively scanned by the S6 block to obtain the scanning results.

[0090] Through the above-mentioned embodiments of the present application, not only an image can be generated, but also an image of interest in the input image can be further generated.

[0091] It should be noted that, for the above-mentioned method embodiments, for the sake of simplicity, they are all described as a series of action combinations, but those skilled in the art should know that the present invention is not limited by the described action sequence, because according to the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the present invention.

[0092] According to another aspect of an embodiment of the present invention, an image generating device for implementing the above-mentioned image generating method is also provided.

[0093] The specific manner in which each unit in the above device embodiment performs operations has been described in detail in the embodiment of the method, and will not be elaborated on again here.

[0094] According to another aspect of an embodiment of the present invention, an electronic device for implementing the above-mentioned image generation method is also provided, and the electronic device may be a terminal device or a server. This embodiment is illustrated by taking the electronic device as a terminal device as an example. The electronic device includes: at least one processor; and a memory connected to the at least one processor in communication; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor performs the steps in any one of the above-mentioned method embodiments. The above-mentioned electronic device may be located in at least one network device among a plurality of network devices of a computer network. The above-mentioned processor may be configured to execute the above-mentioned image generation method through a computer program.

[0095] According to one aspect of the present application, a computer-readable storage medium is provided, and a processor of a computer device reads the computer instructions from the computer-readable storage medium, and the processor executes the computer instructions, so that the computer device executes the above-mentioned image generation method.

[0096] Those skilled in the art will appreciate that the implementation of all or part of the processes in the above method embodiments can be accomplished by instructing related hardware through a computer program. The program can be stored in a computer-readable storage medium. When the program is executed, it can include processes such as those in the above method embodiments.

[0097] The above is only a preferred embodiment of the present invention. It should be pointed out that for ordinary technicians in this technical field, several improvements and modifications can be made without departing from the principle of the present invention. These improvements and modifications should also be regarded as the scope of protection of the present invention.

Claims

1. An image generation method, characterized in that: include: Acquiring reference condition data, and performing a preprocessing operation on the reference condition data to obtain target condition data, wherein the preprocessing operation is used to improve the data quality of the reference condition data; The target condition data is input into an image generation model, and a target image is output, wherein the image generation model is a model obtained by pre-training an initial image generation model with a training sample set, and the image generation model includes a generator and a discriminator, and the generator processes the target condition data through a downsampling module, a visual attention module, and an upsampling module to generate a reference image, and the discriminator outputs discrimination information when taking the reference image as input, and the discrimination information is used to indicate the discrimination result between the reference image and the real image.

2. The method according to claim 1, characterized in that: The initial image generation model is trained using a training sample set, including: Inputting the sample condition data in the training sample set into the initial generator in the initial image generation model to generate a sample generated image corresponding to the sample condition data; Inputting the sample generated image and the sample real image corresponding to the sample condition data into the initial discriminator in the initial image generation model to obtain a discrimination result; Determine a training loss value according to the discrimination result and a preset loss determination method, and alternately correct the initial generator and the initial discriminator based on the loss value; When the initial generator and the initial discriminator meet preset conditions, the initial generator is determined as the generator, and the initial discriminator is determined as the discriminator to obtain the image generation model.

3. The method according to claim 2, characterized in that Inputting the sample condition data in the training sample set into the initial generator in the initial image generation model to generate a sample generated image corresponding to the sample condition data includes: The initial downsampling module in the initial generator performs a preset number of downsampling operations on the sample condition data to obtain downsampled data; The initial visual attention module in the initial generator performs data extraction on the downsampled data to obtain data of interest; The initial up-sampling module in the initial generator performs up-sampling operations on the data of interest for a preset number of times to obtain the sample generated image.

4. The method according to claim 3, characterized in that: The initial visual attention module in the initial generator performs data extraction on the downsampled data to obtain data of interest, including: Performing average pooling and maximum pooling on the downsampled data respectively to obtain average pooled data and maximum pooled data; Connecting the average pooling data and the maximum pooling data to obtain connection feature data; Performing a target convolution operation on the average pooled data, the maximum pooled data, and the connection feature data to obtain convolution feature data; The convolution feature data is input into the visual state space block for visual data processing, and the data of interest is output.

5. The method according to claim 4, characterized in that Performing a target convolution operation on the average pooled data, the maximum pooled data, and the connection feature data to obtain convolution feature data, including: Performing convolution operations on the average pooling data, the maximum pooling data, and the connection feature data respectively to obtain first convolution data corresponding to the average pooling data, second convolution data corresponding to the maximum pooling data, and third convolution data corresponding to the connection feature data; Multiplying the first convolution data with the down-sampled data to obtain first product data, and multiplying the second convolution data with the down-sampled data to obtain second product data; The first product data is multiplied by the third convolution data to obtain third product data, and the second product data is multiplied by the third convolution data to obtain fourth product data, wherein the convolution feature data includes the third product data and the fourth product data.

6. The method according to claim 4, characterized in that Inputting the convolution feature data into the visual state space block for visual data processing and outputting the data of interest, including: The visual state space block normalizes the convolution feature data to obtain normalized data; Performing a target processing operation on the normalized data to obtain target processed data, wherein the target processing operation is used to enhance data features of the normalized data; Performing a reference processing operation on the normalized data to obtain reference processed data, wherein the reference processing operation is used to enhance the generalization ability of the model; The reference processed data, the target processed data and the convolution feature data are subjected to a residual connection operation to obtain the data of interest.

7. The method according to claim 3, characterized in that Determining a training loss value according to the discrimination result and a preset loss determination method includes: Performing feature acquisition operations according to the downsampling process of the downsampling module and the upsampling process of the upsampling module to obtain multiple groups of acquisition features; Determine a first loss value according to the plurality of groups of the collected features and the discrimination results; Determine a second loss value according to a generation process of the sample-generated image of the initial generator, and determine a third loss value according to the generation process of the sample-generated image and the discrimination process of the initial discriminator; The training loss value is determined according to a preset weight set, the first loss value, the second loss value and the third loss value, wherein the preset weight set includes preset weights corresponding to the first loss value, the second loss value and the third loss value, respectively.

8. An image generating device, characterized in that: include: an acquisition unit, configured to acquire reference condition data and perform a preprocessing operation on the reference condition data to obtain target condition data, wherein the preprocessing operation is configured to improve the data quality of the reference condition data; An input unit is used to input the target condition data into an image generation model and output a target image, wherein the image generation model is a model obtained by pre-training an initial image generation model with a training sample set, and the image generation model includes a generator and a discriminator. The generator processes the target condition data through a downsampling module, a visual attention module, and an upsampling module to generate a reference image, and the discriminator outputs discrimination information when taking the reference image as input, and the discrimination information is used to indicate the discrimination result between the reference image and the real image.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the image generation method according to any one of claims 1 to 7.

10. An electronic device, characterized in that: The electronic device includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor executes the image generation method described in any one of claims 1-7.

Citation Information

Patent Citations

  • Infrared target image generation method and device, equipment and storage medium

    CN118037873A

  • Training method and device of image generator

    CN118072108A

  • Fine tuning and control diffusion model

    CN118172620A