Conditional multi-target scene image generation method, device and equipment

By drawing geometric figures on real images as a conditional SinGAN, the conditional SinGAN, which is used as a control condition, solves the problem of uncontrollable images generated in traditional methods, and accurately controls the number and position of targets in multi-target scene images, and significantly improves the generated image quality and controllability.

CN116188934BActive Publication Date: 2025-08-26NAT UNIV OF DEFENSE TECH
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211092907.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-09-08
Publication Date
2025-08-26
Estimated Expiration
2042-09-08

AI Technical Summary

Technical Problem

In the traditional multi-objective scene image generation method, the generated image is uncontrollable and the targets are randomly distributed, which cannot meet the needs of precise control.

Method used

Conditional SinGAN is used to draw geometric figures at the set position of the real image as a control condition, and the loss function of the generator and discriminator is trained layer by layer, and the training control image control image is used to control the similarity between the generated image and the control condition and the proximity of the real image to achieve precise control of the generated image.

Benefits of technology

The precise control of the number and position of the generated image under the introduced control conditions is realized. The generated image not only maintains the overall information of the real image, but also meets the target position and quantity specified by the control conditions, improving the controllability and quality of the generated image.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116188934B_ABST
    Figure CN116188934B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, and device for conditional multi-target scene image generation. The method comprises the following steps: obtaining a multi-target scene image and a training control image as a control condition; invoking a conditional single-image generative adversarial network improved based on SinGAN; incorporating control conditions into the loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network; performing layer-by-layer iterative training starting from the bottom layer of the conditional single-image generative adversarial network based on the multi-target scene image and the training control image; in the iterative training of each layer of the network, using the training control image to control the generated image of the generator in the current layer to be similar to the training control image, and controlling the discriminator in the current layer to control the degree of proximity between the generated image and the input image; and outputting the generated pseudo multi-target scene image when the training of the top layer of the conditional single-image generative adversarial network is completed. This achieves precise control over the generated image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing technology, and in particular to a method, device and equipment for generating conditional multi-target scene images. Background Art

[0002] With the development of multimedia and military fields, public opinion warfare and intelligence warfare increasingly rely on technical support, and image generation technology is one of the key supporting technologies. Within image generation technology, the generation of multi-target scene images is a hot topic. Since acquiring real image material in multimedia and military fields is generally difficult, research on solving this problem with limited image data has important practical applications.

[0003] Among the traditional methods for generating multi-target scene images, SinGAN (Single Image Generative Adversarial Network) is a typical example. SinGAN is designed for multi-target scene images and can generate images using a single multi-target image. In SinGAN, the initial input is random noise, and the final generated image is also generated based on random noise, which solves the problem of lack of training materials in the past. However, in the process of realizing the present invention, the inventors found that the aforementioned traditional methods for generating multi-target scene images have the technical problem of uncontrollable generated images. Summary of the Invention

[0004] Based on this, it is necessary to provide a conditional multi-target scene image generation method, a conditional multi-target scene image generation device and a computer device to address the above technical problems, which can achieve precise control of the generated image.

[0005] To achieve the above objectives, the embodiments of the present invention adopt the following technical solutions:

[0006] In one aspect, an embodiment of the present invention provides a conditional multi-target scene image generation method, comprising the steps of:

[0007] Acquire a multi-target scene image and a training control image as a control condition; the training control image is a multi-target scene image with geometric figures drawn at set positions;

[0008] Call the improved conditional single-image generative adversarial network based on SinGAN; control conditions are added to the loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network;

[0009] Based on the multi-target scene images and training control images, we perform iterative training layer by layer starting from the bottom layer of the conditional single-image generative adversarial network.

[0010] In the iterative training of each layer of the network, the training control image is used to control the generated image of the generator in the current layer to be similar to the training control image, and the training control image is used to control the closeness of the generated image to the input image through the discriminator of the current layer; the input image includes multi-target scene images and training control images;

[0011] When the top-level training of the conditional single-image generative adversarial network is completed, a pseudo multi-target scene image generated after learning the multi-target scene image is output.

[0012] On the other hand, a conditional multi-target scene image generation device is also provided, comprising:

[0013] An image acquisition module is used to acquire a multi-target scene image and a training control image as a control condition; the training control image is a multi-target scene image with geometric figures drawn at set positions;

[0014] The network call module is used to call the improved conditional single-image generative adversarial network based on SinGAN. The loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network are controlled by adding control conditions.

[0015] The network training module is used to perform iterative training layer by layer starting from the bottom layer of the conditional single-image generative adversarial network based on multi-target scene images and training control images;

[0016] A conditional control module is used to control the image generated by the generator in the current layer to be similar to the training control image during the iterative training of each layer of the network, and to control the degree of similarity between the generated image and the input image through the discriminator in the current layer using the training control image. The input image includes a multi-target scene image and a training control image.

[0017] The image generation module is used to output a pseudo multi-target scene image generated after learning the multi-target scene image when the top-level training of the conditional single-image generative adversarial network is completed.

[0018] On the other hand, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and when the processor executes the computer program, the steps of any one of the above-mentioned conditional multi-target scene image generation methods are implemented.

[0019] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the computer program implements any one of the steps of the above-mentioned conditional multi-target scene image generation method.

[0020] One of the above technical solutions has the following advantages and beneficial effects:

[0021] The above-mentioned conditional multi-target scene image generation method, device and equipment obtains a real multi-target scene image and then draws a geometric figure at a set position on the multi-target scene image. The multi-target scene image after the figure is drawn is called a training control image, which serves as the control condition of the subsequent network. The improved conditional single-image generative adversarial network based on SinGAN is called, and the bottom layer of the conditional single-image generative adversarial network is iteratively trained layer by layer based on the aforementioned original multi-target scene image and the training control image. In the iterative training of each layer of the conditional single-image generative adversarial network, the training control image is used as a control condition for training control so that the output image of the generator is similar to the control condition in terms of factors such as the number and layout of targets. The discriminator improves the training stability by generating adversarial loss, and balances the closeness between the generated image and the control condition and the real image by reconstruction loss, so as to achieve precise control of the number and position of targets in the generated image under the premise of introducing control conditions, thereby achieving the effect of precise control of the generated image. BRIEF DESCRIPTION OF THE DRAWINGS

[0022] Figure 1 Schematic diagram of the local network structure of Conditional SinGAN in one embodiment;

[0023] Figure 2 1 is a flow chart of a conditional multi-objective scene image generation method according to an embodiment;

[0024] Figure 3 Schematic diagram of the block learning principle in one embodiment;

[0025] Figure 4 A schematic diagram of the internal structure of a generator in one embodiment;

[0026] Figure 5 is a schematic diagram of an original input image and control conditions in one embodiment;

[0027] Figure 6 Schematic diagram of the original input image, control conditions, and output image of Conditional SinGAN in one embodiment;

[0028] Figure 7 A schematic diagram comparing the original image, PS result, SinGAN-generated image, and Conditional SinGAN-generated image in one embodiment;

[0029] Figure 8 Schematic diagram of the module structure of a conditional multi-target scene image generation device in one embodiment. DETAILED DESCRIPTION

[0030] In order to make the purpose, technical solutions and advantages of this application more clear, the following further describes this application in detail with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not intended to limit this application.

[0031] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art to which this application pertains. The terms used herein in the specification of this application are for the purpose of describing specific embodiments only and are not intended to limit this application.

[0032] In addition, the technical solutions between the various embodiments of the present invention can be combined with each other, but it must be based on the fact that ordinary technicians in this field can implement it. When the combination of technical solutions is mutually contradictory or cannot be implemented, it should be deemed that such a combination of technical solutions does not exist and is not within the scope of protection required by the present invention.

[0033] While traditional SinGAN solves the problem of lack of training material, its image generation capabilities are unstable, and the targets in the generated multi-target scene images are randomly and chaotically distributed. Therefore, based on SinGAN, this application proposes Conditional SinGAN, a conditional single-image generative adversarial network, to enhance the controllability of the generated results. The new generative adversarial network Conditional SinGAN can overcome the aforementioned shortcomings of traditional image generation methods.

[0034] Conditional SinGAN is based on the pyramid structure of SinGAN. Because SinGAN is designed for multi-target scene images, it can use a single multi-target image for image generation. In SinGAN, the initial input is random noise, and the final generated image is also generated based on random noise. SinGAN lacks control conditions, which fundamentally leads to the randomness of the generated images, that is, the generated images are uncontrollable and easily generate images that do not conform to human visual cognition. The new network provided in this application will add control conditions so that the network can be guided by the control conditions when generating images.

[0035] In summary, this application addresses the technical problem of uncontrollable generated images in traditional multi-target scene image generation methods and provides a new conditional multi-target scene image generation method. In this new method, the conditional single image generative adversarial network (Conditional SinGAN) adopts a pyramid-type generative adversarial network structure. Control conditions are added to the generator and discriminator of each layer of the pyramid in order to better control the image generation in a specified direction. The network structure after adding the control conditions is as follows: Figure 1 shown.

[0036] Each time the generator outputs a generated image, it is compared against a control condition. The control condition here is not the class label used in traditional CGANs, as such a simple one-hot vector cannot capture the rich information in multi-object scene images. In the new approach, a simple geometric pattern is drawn at a specific location in the real image as a control condition, which is then fed into the generator and discriminator, achieving the goal of using the real image as the control condition. In traditional SinGANs, the discriminator inputs the generated image of the current layer and a real image of the same size. In ConditionalSinGANs, the generator inputs are noise, the original image, and the control condition. The discriminator receives the real image, a fake image generated by the generator, and the control condition. Each discriminator layer determines whether the image output by the generator at that layer is sufficiently realistic. In the Conditional SinGAN architecture, the lower layers generate the overall image structure, while the higher layers gradually refine the image. Unlike traditional techniques, however, each layer incorporates the control condition. The resulting image maintains the overall information of the real image while also meeting the target location and quantity specified by the control condition.

[0037] See also Figure 2 In one embodiment, the present invention provides a conditional multi-target scene image generation method, comprising the following steps S12 to S20:

[0038] S12, acquiring a multi-target scene image and a training control image as a control condition; the training control image is a multi-target scene image with geometric figures drawn at set positions.

[0039] It can be understood that there are many options for the shape of the geometric figure, which can be selected specifically according to the shape outline of the target in the multi-target scene image, such as a geometric color block abstracted from the shape of the target, any geometric shape that is similar or dissimilar to the target shape, etc. The specific selection of the set position can also be flexibly selected according to the position distribution, quantity, etc. of the targets in the multi-target scene image. For example, the set position can be selected as one or more positions on the multi-target scene image where pseudo targets are expected to be generated. For different multi-target scene images, different corresponding training control images can be formed after drawing. By drawing a simple geometric pattern at a set position of the real image (such as the multi-target scene image obtained above) as a control condition and inputting it into the generator and discriminator of each layer, the purpose of using the real image as the control condition is achieved.

[0040] S14, calling the improved conditional single-image generative adversarial network based on SinGAN; the loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network are added with control conditions;

[0041] S16, based on the multi-target scene images and the training control images, iteratively train the conditional single-image generative adversarial network layer by layer starting from the bottom layer;

[0042] S18, in the iterative training of each layer of the network, the training control image is used to control the generated image of the generator in the current layer to be similar to the training control image, and the training control image is used to control the closeness of the generated image to the input image through the discriminator of the current layer; the input image includes a multi-target scene image and a training control image.

[0043] It can be understood that the local structure of the conditional single-graph generative adversarial network can be seen in the above Figure 1 The hint, Figure 1 The structure of two adjacent layers (n layer and n-1 layer) in the network and its data processing flow are illustrated. Regarding the description of the generator and discriminator modules, you can refer to the corresponding description of the generator and discriminator in SinGAN for similar understanding, and will not be elaborated in this specification. Compared to SinGAN, the generator and discriminator of each layer of the Conditional SinGAN network of this application are both controlled by conditions.

[0044] In order to make the image with geometric figures drawn become the control condition that can truly participate in the process execution in the network, the loss functions of the generator and discriminator of SinGAN need to be improved. Conditional SinGAN also adopts a layer-by-layer training method, starting from the Nth layer (the coarsest layer). After the training of each layer, the parameters are fixed, and then the next layer is trained. The generator loss of Conditional SinGAN is composed of content loss. Content loss is used to measure the difference between the control condition and the generated image of each layer. In this embodiment, the position, number and posture of the targets in the multi-target scene image are considered, so the L1 norm is used as the form of the loss function, which is more conducive to the expression of detailed information such as the number, position and posture of the targets in the image.

[0045] In some embodiments, let x n Represents the real image of the nth layer after upsampling, y n Is the control condition of the nth layer, then the conditional single-graph generative adversarial network in the nth layer, the generator's loss function is as follows:

[0046] L rec =||(G n (z n ,x n+1 ↑ r )-y n || (1)

[0047] Among them, Gn represents the generator of the n-th layer network, z n represents the random noise input to the nth layer network, x n+1 Represents the real image of the n+1th layer after upsampling, ↑ represents upsampling, and r represents the upsampling factor.

[0048] The loss function of the generator shown in formula (1) can make the image output by each layer of the generator closer to factors such as the number, position and posture of the targets in the control conditions.

[0049] In some embodiments, the discriminator loss function consists of two parts: adversarial loss and reconstruction loss. The use of generative adversarial loss can improve the stability of training and is the loss function of traditional generative adversarial networks. The purpose of using reconstruction loss is to make the final output image similar to both the real image and the control condition. x can represent the real image, represents the pseudo image generated at the nth scale, which is obtained by upsampling layer by layer. y represents the control condition, z represents the noise, G and D represent the generator and discriminator respectively. The loss function of the discriminator of the nth layer network in the conditional single-image generative adversarial network is:

[0050]

[0051] Among them, the adversarial loss is the same as the traditional adversarial loss in form. After adding the control condition, the adversarial loss L adv (G,D) is:

[0052]

[0053] When constructing the loss function for a conditional single-image generative adversarial network, using a traditional loss function (such as the L2 norm) facilitates the establishment of the objective function, making the network model more adept at generating diverse images. However, in this generation task, while the discriminator's performance remains unchanged, the generator must not only deceive the discriminator but also generate images as realistic as possible. Therefore, in the conditional single-image generative adversarial network, both the generator and the discriminator use the L1 norm to measure the difference between the output image and the control condition, and the L2 norm to measure the difference between the output image and the real image.

[0054] Compared to the L2 norm, the L1 norm focuses more on the inherent information of the image rather than the pixel information. The target position, quantity factors, and posture contained in the inherent information are exactly the parts that need to be controlled when the control conditions need to control the generated image. Therefore, using the L1 norm as the loss between the generated image and the control condition can ensure the controllability of the generated image. The mathematical expression of this process is shown in the above formula (1). The L2 norm between the real image and the output image is used as the other part of the loss function because the L2 norm is the Euclidean distance and it focuses more on the difference between the pixels of the two images. It can ensure that the difference between the output image and the real image at the pixel level is small, so that the blurriness of the output image is reduced. The mathematical expression of this process is shown in the following formula (6). Therefore, the design of this loss function can not only improve the controllability of the generated image, so that the target is generated at the position specified by the control condition, but also reduce the blurriness of the generated image, ensuring the quality of the generated image.

[0055] Let x n represents the real image of the nth layer after sampling, represents the pseudo image generated by the nth layer (i.e., the generated image), φ is a convolutional network with 5 convolutional modules. The specific details of the module structure of the convolutional network are given in Table 1 below. ↑ represents upsampling, r is the upsampling factor, and β is a constant. φ n represents a 5-layer fully convolutional network consisting of 3×3 convolution-BN-LerkyReLU on the nth layer. The derivation process of the reconstruction loss is as follows. Formula (5) describes the internal workings of the generator:

[0056]

[0057]

[0058] Reconstruction loss L rec (G) is:

[0059]

[0060] And when n=N (i.e. the first scale), the initial input of the generator is only random noise, and the reconstruction loss L rec (G) is:

[0061] L rec =||(G N (z * )-x N ||2+β||(G N (z * )-y N || (7)

[0062] Among them, D nrepresents the discriminator of the nth layer network, α represents the hyperparameter, which is the factor for adjusting the loss, E represents the expected value of the distribution function, y n represents the control condition of the nth layer, The generated image generated by the n+1th layer, z * represents a fixed noise spectrum, G N represents the generator of the Nth layer network, x N represents the real image after upsampling the Nth layer network, β represents a constant, y N Indicates the control condition of the Nth layer.

[0063] The reconstruction loss aims to minimize the L2 norm of the difference between the output of the nth layer generator and the original image at that scale, given the noise as input, while also minimizing the L1 norm of the difference between the output of that layer and the control condition. The parameter β is a tuning factor, ranging from 0 to 1. A larger value indicates that the generated image is closer to the control condition, and vice versa. The reconstruction loss is introduced primarily to improve the controllability of the image generated by the network model. It can control the final output image of the network model to retain the basic semantics of the real image and, by adjusting the β parameter, bring the output image closer to the control condition. During training, if the generator parameters are updated, the discriminator parameters are fixed; if the discriminator parameters are updated, the generator parameters are fixed. After training each layer, the generator and discriminator parameters of that layer are fixed, and then training of the next layer is carried out, and training is repeated layer by layer.

[0064] In the Conditional SinGAN network model, the generator loss ensures that the generator's output images resemble the control condition in terms of object number, layout, and pose. The discriminator, through a generative adversarial loss, improves training stability and uses a reconstruction loss to balance the proximity of the generated images to the control condition and the real image. This allows for precise control of the number, position, and pose of objects in the generated images while still introducing the control condition.

[0065] S20, when the top-level training of the conditional single-image generative adversarial network is completed, output a pseudo multi-target scene image generated after learning the multi-target scene image.

[0066] The above-mentioned conditional multi-target scene image generation method obtains a real multi-target scene image and then draws a geometric figure at a set position on the multi-target scene image. The multi-target scene image after the figure is drawn is called a training control image, which serves as the control condition of the subsequent network. The improved conditional single-image generative adversarial network based on SinGAN is called, and the bottom layer of the conditional single-image generative adversarial network is iteratively trained layer by layer based on the original multi-target scene image and the training control image. In the iterative training of each layer of the conditional single-image generative adversarial network, the training control image is used as a control condition for training control so that the output image of the generator is similar to the control condition in terms of factors such as the number and layout of targets. The discriminator improves the training stability through the generation of adversarial loss and balances the closeness between the generated image and the control condition and the real image through the reconstruction loss, so as to achieve precise control of the number and position of targets in the generated image under the premise of introducing the control condition, thereby achieving the effect of precise control of the generated image.

[0067] In one embodiment, the process of using the training control image to control the generator in the current layer to generate an image similar to the training control image includes:

[0068] After the generator of the current layer generates a generated image, it is determined whether the generated image is close to the training control image;

[0069] If so, the generated image is determined to be similar to the training control image and the generated image is transferred to the discriminator of the current layer;

[0070] If not, the generated image is discarded, and the generator of the current layer is instructed to generate a new generated image and perform the next round of proximity judgment until the judgment result is yes.

[0071] It is understandable that Figure 1 As shown, in the network of the current layer n, the generated image generated by the n+1 layer After upsampling and random noise z n , real image x n Input the current layer n together for processing, and the generator generates the generated image of the current layer n Then determine the generated image With the training control image y n Is it close? If so, an image will be generated. Transmitted to the discriminator of the current layer n.

[0072] In one embodiment, the process of using a training control image to control the closeness of a generated image to an input image through a discriminator of a current layer includes:

[0073] After receiving the generated image, if the similarity judgment result between the generated image and the training control image is determined to be similar, and the authenticity judgment result between the generated image and the multi-target scene image is determined to be true, the generated image is upsampled and the sampling result is transmitted to the next layer of network for processing.

[0074] It is understandable that Figure 1 As shown, in the network of the current layer n, the discriminator receives the generated image After that, determine the generated image With the training control image y n Is it similar and determine the generated image With the real image x n Is it true? That is, the two images are judged by the discriminator of the generative adversarial network. When the discriminator cannot distinguish the authenticity of the image, it is judged to be true. If the judgment results are both yes, then the image is generated. Perform upsampling and send to the next layer of network for processing.

[0075] Regarding the model training of the conditional single-image generative adversarial network, specifically: in order to solve the problem that the traditional generative adversarial network model has a large demand for training data and cannot perform single-image learning, the network model of the present invention adopts a single-image training mode. Based on the training of SinGAN block learning and combined with the pyramid-like multi-layer GAN structure, the network model of the present invention also adopts a block learning method to train a single image. A single image is only a sampling point when training the GAN model, which is far from enough to be used for model training. Therefore, a single image needs to be divided into small blocks for training. Assuming that the size of the input original image is 400*400 pixels, the steps of the entire training process are as follows:

[0076] (1) Using low-resolution images to train the model can effectively reduce the training cost. In the low-level pyramid GAN, the original image is first downsampled to a certain size, and then the downsampled image is cut into small blocks. These low-resolution small image blocks are trained, which is conducive to faster training. For example, for a 400*400 pixel image, it is first downsampled to 40*40 to obtain a low-resolution image. This low-resolution image is then divided into blocks, such as 11*11 pixel blocks, and about 800 blocks can be obtained. The size of each block is about 1 / 800 of the entire image. Using these image blocks for training has low training overhead, and the obtained image blocks retain relatively complete overall information of the entire image. Therefore, through training, the low-level pyramid can learn the overall information of the image.

[0077] (2) Using a high-resolution image training model, the local details of the image can be learned. Although the overall information of the image is obtained in step (1), the image details generated by the low-resolution image training low-level model are missing, and the generated image is blurry. Therefore, in the high-level pyramid GAN, the image details need to be improved. The original image is downsampled to 200*200 pixels and then divided into 11*11 image blocks, which can obtain about 30,000 image blocks. The size of each image block is about 1 / 30,000 of the entire image. The 11*11 image block is smaller than the 11*11 image block relative to the 200*200 image than the 40*40 image. Therefore, in the high-level pyramid GAN, the image blocks obtained are relatively fine, and the image blocks obtained by training with these image blocks have detailed information.

[0078] (3) Pyramid GAN completes the acquisition of the overall information of the image at the lower level. The output obtained in the lower level is upsampled and used as one of the inputs of the higher level. It is gradually refined in the upward process, and finally obtains an image that has both overall information and local details.

[0079] The principles of the above three steps are as follows Figure 3 As shown in Figure 1, the structure of the generator G and discriminator D at each level of the pyramid is identical, consisting of five groups of 3x3 convolutional modules with a sliding stride of 1. Therefore, the final receptive field of the generator and discriminator before the image output is calculated to be 11x11. These convolutional modules perform conventional feature extraction. The composition of these five groups of convolutional modules is shown in Table 1. Regardless of the input image size, after the image enters the generator and discriminator, the size of each sliding convolution is 11x11. Therefore, in the lower layers, where the input image is smaller, the image blocks are relatively large and retain more overall information. In the top layers, where the input image is larger, the image is relatively small but retains more details. During the iteration process, the receptive field size remains unchanged, only the image size changes. This also relatively changes the proportion of the obtained image block size in the entire image, with the proportion decreasing as one moves up the pyramid.

[0080] Table 1

[0081] Generator Discriminator Conv(3*3) Conv(3*3) BatchNorm BatchNorm LeakyReLU LeakyReLU

[0082] In order to ensure that the overall information learned by the low-level output is not weakened when it is passed upward as the input of the high-level layer, the generator adopts a residual network structure. The input of this residual network structure is the superposition of a certain size of noise and the upsampling result of the previous layer output. After 5 groups of 3*3 layer convolution processing, it is superimposed with the upsampling result of the previous layer output. This can not only ensure that the randomness introduced by random noise is not weakened as the number of layers increases, thereby making the generated image more diverse, but also refine the details that were not generated in the previous layer through the convolution layer. The internal structure of the generator is as follows Figure 4 As shown in Figure 5, it completes the operation of the above formula (5). The discriminator uses Patch-GAN, which can perform sliding convolution processing on the image output by the generator. The discriminator judges the authenticity of an 11*11 image block in the image each time, and finally outputs an average value as the output of the discriminator.

[0083] Using a pyramid-like, multi-layer GAN allows for block-by-block training of a single image. Introducing a residual network into the GAN generator preserves the overall information as it is passed up the pyramid. In this way, by learning from a single image, the pyramid GAN learns the overall image information at lower levels and refines the image's details at higher levels. Ultimately, the resulting image is similar overall to the original, but with slightly different details. This addresses the issue of scarce training data to a certain extent.

[0084] In summary, by introducing conditional control into the network and then using single-image training, the network model can be trained with a single image. The number, position, and posture of targets in the pseudo-multi-target scene image generated by the trained network model are all controllable.

[0085] In one embodiment, in order to more intuitively and comprehensively illustrate the above-mentioned conditional multi-object scene image generation method, the following is one of the experimental examples of the above-mentioned method:

[0086] Since there is currently no publicly available multi-target scene image dataset, the data used in this experiment were all collected from the Internet. When collecting data, we collected as many multi-target scene images as possible, and most of these collected images met the definition of multi-target scene images after screening. Since the network model mentioned above in this application can be trained with a single image, the dataset is relatively small. The collected data can include 100 multi-target military scene images in four major categories: tanks, ships, aircraft, and missiles. This experiment is based on this self-built dataset.

[0087] The method of making control conditions in this experiment is: first select a multi-target scene image, then use drawing software to draw simple geometric shapes at the set position of the image according to the requirements, such as in the sky, on the sea surface, and in the desert, and use the image with the drawn geometric shapes as the control condition of the network. Figure 5 As shown, the left side shows the original real image, and the right side shows the training control image that has been artificially processed as a control condition, where the boxed part is the artificially added part.

[0088] The inputs of the experiment are random noise, control conditions, and real images. There are no specific requirements for the sizes of real images and control conditions. Each time the network model is trained, it needs to input a real image, a control condition, and random noise (generated by the program). Since the network model is a pyramid network, training needs to be performed layer by layer. The output of the previous layer is upsampled and used as the input of the next layer, and it is iterated all the way to the 0th layer. It should be noted that the highest layer of the pyramid is the 0th layer, and the lowest layer is the Nth layer. The output of the network model is a fake multi-target scene image. The number and layout of the newly added targets in the image are basically the same as those of the control condition. Once a model is trained, it can be used unlimited times to generate fake multi-target scene images under different control conditions.

[0089] To facilitate analysis of experimental results and reduce training overhead, the minimum size of the real image and control condition input during training was downsampled to ε pixels (width or height) and input to the bottom layer of the pyramid. The upsampling factor r = 4 / 3 was set, and the pyramid iteration was terminated by upsampling the real image and control condition to 10*ε pixels. This means that the size of the output image at the top layer is 10 times that of the bottom layer. After each model training, multiple tests can be performed, and the control conditions can be modified for each test. The test results will vary depending on the control conditions.

[0090] Based on the upsampling factor, the number of pyramid layers is 7 to 8, and the resulting image is half the size of the original real image. Because the training images are relatively small, smaller images can be divided into fewer blocks, which saves training time. Therefore, the input image at the bottom layer is a 20-fold downsampling of the real image. During model training, to achieve a balance between the generated fake images under control conditions and the real images, the parameter β is set to 0.5.

[0091] This experiment also included a comparative experiment under the same experimental conditions. Considering the model's need for single-image processing capabilities, the learning models used in this comparative experiment were the traditional SinGAN and the Conditional SinGAN proposed in this application, both trained using the same training set. During the testing phase, using the same input, the existing Ps, SinGAN, and Conditional SinGAN models were tested on the same data, yielding different outputs for comparative analysis.

[0092] Results Analysis: The experiment was conducted according to the above experimental setup. Both SinGAN and Conditional SinGAN were trained using the same training data: a self-built dataset of four categories of multi-object scene images. During PS processing, the images were also downsampled to the same size as the output of SinGAN and Conditional SinGAN.

[0093] The test results of the trained Conditional SinGAN on four categories of multi-target scene images are as follows: Figure 6 As shown in the figure, from left to right are the real image, control condition and output image. The four categories of targets are tanks, aircraft, ships and missiles.

[0094] from Figure 6 As can be seen, the output image contains targets that are not present in the real image. These new targets differ from the real image in terms of number, position, and pose. Furthermore, the number and position of the new targets in the output image are constrained by control conditions and are essentially identical to the number and position of the manually added geometric shapes in the control conditions. The forged images of helicopters and missiles achieve precise control of the positions of the forged targets, while the forged images of ships achieve precise control of the number of forged targets. Careful observation revealed no instances of targets in the output images that are inconsistent with human visual perception. If these forged images were to appear online, they would be extremely confusing to users without relevant background knowledge.

[0095] Because there are not many models capable of generating multi-target scene images, the generation results of SinGAN are selected here for comparison with Conditional SinGAN. At the same time, the traditional method PS is also included in the comparison. When PS is applied to the image, in order to make the comparison of the three methods more convincing, the number and position of the forged targets are kept as close as possible to the control condition. This example will make a qualitative comparison of the generation effects of each method. Figure 7 The following figure shows the results of testing the same image using various methods. From left to right, they are the original image, the image obtained by the Ps method, the image generated by SinGAN, and the image generated by Conditional SinGAN.

[0096] from Figure 7As can be seen in the figure, due to technical limitations, the images processed using PS only undergo simple transformations such as copying, translating, and scaling the original real targets when adding new targets. Redundant background exists around the edges of the targets, and the subtracted target edges are incomplete. This phenomenon is particularly evident in the tank image, where the forged tank is a copy, scaled down, and translated version of the real tank. These factors result in the newly added targets appearing somewhat rigid and stiff, reducing image fidelity. SinGAN can generate multi-target scene images, but due to its lack of control conditions, some of the targets in the generated multi-target scene images exhibit blurring, disintegration, and overlap, which are inconsistent with human visual perception. The generated multi-target scene images appear cluttered and disorganized, and the number of targets generated by SinGAN is also uncontrollable. These phenomena are particularly evident in the fighter jet image, where the forged fighter jet appears disintegrated and severely distorted. Conditional SinGAN, on the other hand, generates relatively realistic multi-target scene images that conform to human visual perception by applying control conditions.

[0097] Because qualitative evaluations are less persuasive, this experiment used an existing AI image quality assessment algorithm, called "Image Quality Assessment," to evaluate the results. Table 2 shows the average scores for each forgery method. Where, a higher score indicates higher image quality.

[0098] Table 2

[0099] Forgery method Assessment score SinGAN 15.56 Conditional SinGAN 19.06

[0100] As can be seen in Table 2, SinGAN achieved the lowest score because the objects in its generated multi-object scene images were random, many were distorted, and the distribution of some objects did not conform to human visual perception. However, SinGAN's generated results also contained a small number of good forged samples, where the objects were randomly distributed. Conditional SinGAN achieved the highest score because the objects in its generated multi-object scene images were controlled by the control conditions, resulting in a reasonable distribution and no distortion. The quantitative evaluation results of this evaluation algorithm are generally consistent with human perceptual judgment.

[0101] Experimental results show that the realism of multi-target scene images forged by Conditional SinGAN is better than that of images generated by Ps-based image generation methods and SinGAN-based image generation methods.

[0102] It should be understood that although Figure 2The steps in the flowchart are shown in sequence as indicated by the arrows, but these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise specified in this document, there is no strict order restriction for the execution of these steps, and these steps can be executed in other orders. Figure 2 At least part of the steps may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily performed at the same time, but can be performed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed in turn or alternately with other steps or at least part of the sub-steps or stages of other steps.

[0103] See also Figure 8 In one embodiment, a conditional multi-target scene image generation device 100 is also provided, comprising an image acquisition module 11, a network calling module 13, a network training module 15, a condition control module 17, and an image generation module 19. The image acquisition module 11 is used to acquire a multi-target scene image and a training control image as a control condition; the training control image is a multi-target scene image with geometric figures drawn at a set position. The network calling module 13 is used to call a conditional single-image generative adversarial network improved based on SinGAN; control conditions are added to the loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network. The network training module 15 is used to perform iterative training layer by layer starting from the bottom layer of the conditional single-image generative adversarial network based on the multi-target scene image and the training control image.

[0104] The conditional control module 17 is used to control the similarity of the generated image of the generator in the current layer to the training control image during the iterative training of each network layer, and to control the proximity of the generated image to the input image through the discriminator in the current layer using the training control image. The input image includes a multi-target scene image and a training control image. The image generation module 19 is used to output a pseudo multi-target scene image generated after learning the multi-target scene image when the training of the top layer of the conditional single-image generative adversarial network is completed.

[0105] The above-mentioned conditional multi-target scene image generation device 100, through the collaboration of various modules, obtains a real multi-target scene image, and then draws a geometric figure at a set position on the multi-target scene image. The multi-target scene image after the figure is drawn is called a training control image, which serves as a control condition for the subsequent network. The improved conditional single-image generative adversarial network based on SinGAN is called, and the bottom layer of the conditional single-image generative adversarial network is iteratively trained layer by layer based on the aforementioned original multi-target scene image and the training control image. In the iterative training of each layer of the conditional single-image generative adversarial network, the training control image is used as a control condition for training control, so that the output image of the generator is similar to the control condition in terms of factors such as the number and layout of targets, while the discriminator improves the training stability by generating adversarial loss, and balances the closeness between the generated image and the control condition and the real image by reconstruction loss, so as to achieve precise control of the number and position of targets in the generated image under the premise of introducing control conditions, thereby achieving the effect of precise control of the generated image.

[0106] In one embodiment, the conditional multi-target scene image generation apparatus 100 may also be used to implement other processing functions of the conditional multi-target scene image generation method.

[0107] Regarding the specific limitations of the conditional multi-target scene image generation device 100, please refer to the corresponding limitations of the conditional multi-target scene image generation method above, which will not be repeated here. The various modules in the above-mentioned conditional multi-target scene image generation device 100 can be implemented in whole or in part by software, hardware, and a combination thereof. The above-mentioned modules can be embedded in or independent of a device with a specific data processing function in the form of hardware, or can be stored in the memory of the aforementioned device in the form of software, so that the processor can call and execute the operations corresponding to the above modules. The aforementioned device can be, but is not limited to, various types of computer devices or graphics processing elements in the field.

[0108] On the other hand, a computer device is provided, comprising a memory and a processor, wherein the memory stores a computer program, and the processor can implement the following steps when executing the computer program: obtaining a multi-target scene image and a training control image as a control condition; the training control image is a multi-target scene image with a geometric figure drawn at a set position; calling a conditional single-image generative adversarial network improved based on SinGAN; the loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network are added with control conditions; based on the multi-target scene image and the training control image, iterative training is performed layer by layer starting from the bottom layer of the conditional single-image generative adversarial network; in the iterative training of each layer of the network, the training control image is used to control the generated image of the generator in the current layer to be similar to the training control image, and the training control image is used to control the degree of closeness between the generated image and the input image through the discriminator of the current layer; the input image includes the multi-target scene image and the training control image; when the training of the top layer of the conditional single-image generative adversarial network is completed, the pseudo multi-target scene image generated after learning the multi-target scene image is output.

[0109] In one embodiment, when the processor executes the computer program, it can also implement the steps or sub-steps added in each embodiment of the above-mentioned conditional multi-target scene image generation method.

[0110] On the other hand, a computer-readable storage medium is provided, on which a computer program is stored, which implements the following steps when executed by a processor: obtaining a multi-target scene image and a training control image as a control condition; the training control image is a multi-target scene image with geometric figures drawn at a set position; calling a conditional single-image generative adversarial network improved based on SinGAN; the loss functions of the generator and discriminator of each layer of the conditional single-image generative adversarial network are added with control conditions; based on the multi-target scene image and the training control image, iterative training is performed layer by layer starting from the bottom layer of the conditional single-image generative adversarial network; in the iterative training of each layer of the network, the training control image is used to control the generated image of the generator in the current layer to be similar to the training control image, and the training control image is used to control the degree of closeness between the generated image and the input image through the discriminator of the current layer; the input image includes the multi-target scene image and the training control image; when the training of the top layer of the conditional single-image generative adversarial network is completed, the pseudo multi-target scene image generated after learning the multi-target scene image is output.

[0111] In one embodiment, when the computer program is executed by a processor, it can also implement the steps or sub-steps added in each embodiment of the above-mentioned conditional multi-target scene image generation method.

[0112] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus dynamic random access memory (Rambus DRAM, abbreviated as RDRAM) and interface dynamic random access memory (DRDRAM).

[0113] The technical features of the above embodiments can be combined arbitrarily. To make the description concise, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0114] The above embodiments merely illustrate several implementation methods of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the present invention. It should be noted that a person skilled in the art could make various modifications and improvements without departing from the spirit of the present application, all of which fall within the scope of protection of the present application. Therefore, the scope of protection of the present patent application shall be determined by the appended claims.

Claims

1. A conditional multi-objective scene image generation method, characterized in that: Including steps: Acquire a multi-target scene image and a training control image as a control condition; the training control image is the multi-target scene image with geometric figures drawn at set positions; Call the improved conditional single-image generative adversarial network based on SinGAN; The control condition is added to the loss function of the generator and discriminator of each layer of the network in the conditional single-image generative adversarial network; Performing layer-by-layer iterative training starting from the bottom layer of the conditional single-image generative adversarial network according to the multi-target scene image and the training control image; In the iterative training of each layer of the network, the training control image is used to control the generated image of the generator in the current layer to be similar to the training control image, and the training control image is used to control the degree of closeness between the generated image and the input image through the discriminator of the current layer; the input image includes the multi-target scene image and the training control image; When the top-level training of the conditional single-image generative adversarial network is completed, outputting a pseudo multi-target scene image generated after learning the multi-target scene image; Among them, the loss function of the generator of the n-th layer network in the conditional single-graph generative adversarial network is L rec : L rec =||(G n (z n ,x n+1 ↑ r )-y n || Among them, G n represents the generator of the n-th layer network, z n Represents the random noise input to the nth layer network, y n represents the control condition of the nth layer, x n+1 represents the real image of the n+1th layer after upsampling, ↑ represents upsampling, and r represents the upsampling factor; The loss function of the discriminator of the n-th layer network in the conditional single-image generative adversarial network is: Among them, the adversarial loss L adv (G,D) is: Reconstruction loss L rec (G) is: And when n=N, the reconstruction loss L rec (G) is: L rec =||(G N (z * )-x N ||2+β||(G N (z * )-y N || Among them, G n Denotes the generator of the n-th layer network, D n represents the discriminator of the nth layer network, α represents the hyperparameter, E represents the expected value of the distribution function, x represents the real image, x n represents the real image of the nth layer network after upsampling, y represents the control condition, y n represents the control condition of the nth layer, z represents random noise, Indicates the generated image, represents the generated image generated by the nth layer, The generated image generated by the n+1th layer, z * represents a fixed noise spectrum, ↑ represents upsampling, r represents the upsampling factor, G N represents the generator of the Nth layer network, x N represents the real image after upsampling the Nth layer network, β represents a constant, y N Indicates the control condition of the Nth layer.

2. The conditional multi-objective scene image generation method according to claim 1, characterized in that: The process of using the training control image to control the generator in the current layer to generate an image similar to the training control image includes: After the generator of the current layer generates the generated image, determining whether the generated image is close to the training control image; If so, determining that the generated image is similar to the training control image and transmitting the generated image to the discriminator of the current layer; If not, the generated image is discarded, and the generator of the current layer is instructed to generate a new generated image and perform the next round of proximity judgment until the judgment result is yes.

3. The conditional multi-objective scene image generation method according to claim 1 or 2, characterized in that: The process of controlling the closeness between the generated image and the input image by the discriminator of the current layer using the training control image includes: After receiving the generated image, if it is determined that the similarity judgment result between the generated image and the training control image is similar, and the authenticity judgment result between the generated image and the multi-target scene image is true, the generated image is upsampled and the sampling result is transmitted to the next layer of network for processing.

4. A conditional multi-target scene image generation device, characterized in that: include: An image acquisition module, used to acquire multi-target scene images and training control images as control conditions; The training control image is the multi-target scene image with geometric figures drawn at set positions; The network calling module is used to call the improved conditional single-image generative adversarial network based on SinGAN; The control condition is added to the loss function of the generator and discriminator of each layer of the network in the conditional single-image generative adversarial network; A network training module, configured to perform layer-by-layer iterative training starting from the bottom layer of the conditional single-image generative adversarial network based on the multi-target scene image and the training control image; a conditional control module, configured to, during iterative training of each layer of the network, use the training control image to control the image generated by the generator in the current layer to be similar to the training control image, and use the training control image to control the degree of proximity between the generated image and the input image through the discriminator in the current layer; the input image includes the multi-target scene image and the training control image; An image generation module is configured to output a pseudo multi-target scene image generated after learning the multi-target scene image when the top-level training of the conditional single image generation adversarial network is completed; Among them, the loss function of the generator of the n-th layer network in the conditional single-graph generative adversarial network is L rec : L rec =||(G n (z n ,x n+1 ↑ r )-y n || Among them, G n represents the generator of the n-th layer network, z n Represents the random noise input to the nth layer network, y n represents the control condition of the nth layer, x n+1 represents the real image of the n+1th layer after upsampling, ↑ represents upsampling, and r represents the upsampling factor; The loss function of the discriminator of the n-th layer network in the conditional single-image generative adversarial network is: Among them, the adversarial loss L adv (G,D) is: Reconstruction loss L rec (G) is: And when n=N, the reconstruction loss L rec (G) is: Among them, G n Denotes the generator of the n-th layer network, D n represents the discriminator of the nth layer network, α represents the hyperparameter, E represents the expected value of the distribution function, x represents the real image, x n represents the real image of the nth layer network after upsampling, y represents the control condition, y n represents the control condition of the nth layer, z represents random noise, Indicates the generated image, represents the generated image generated by the nth layer, The generated image generated by the n+1th layer, z * represents a fixed noise spectrum, ↑ represents upsampling, r represents the upsampling factor, G N represents the generator of the Nth layer network, x N represents the real image after upsampling the Nth layer network, β represents a constant, y N Indicates the control condition of the Nth layer.

5. The conditional multi-objective scene image generation device according to claim 4, characterized in that: The condition control module is used to implement the following functions in the process of using the training control image to control the image generated by the generator in the current layer to be similar to the training control image: After the generator of the current layer generates the generated image, determining whether the generated image is close to the training control image; If so, determining that the generated image is similar to the training control image and transmitting the generated image to the discriminator of the current layer; If not, the generated image is discarded, and the generator of the current layer is instructed to generate a new generated image and perform the next round of proximity judgment until the judgment result is yes.

6. The conditional multi-objective scene image generation device according to claim 5, characterized in that: The condition control module is used to implement the following functions in the process of controlling the closeness between the generated image and the input image through the discriminator of the current layer using the training control image: After receiving the generated image, if it is determined that the similarity judgment result between the generated image and the training control image is similar, and the authenticity judgment result between the generated image and the multi-target scene image is true, the generated image is upsampled and the sampling result is transmitted to the next layer of network for processing.

7. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the conditional multi-target scene image generation method according to any one of claims 1 to 3 are implemented.

8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the conditional multi-target scene image generation method according to any one of claims 1 to 3 are implemented.

Citation Information

Patent Citations

  • Single ship target SAR image generation method based on generative adversarial network

    CN112052899A

  • Segmentation Guided Image Generation With Adversarial Networks

    US20190295302A1