Image generation device, image generation method and non-transitory computer readable storage medium

The image generation device improves authenticity and diversity by using a feature generator with a randomization process and discriminator training, enabling high-quality image generation for neural network training and diverse applications.

JP2025098968APending Publication Date: 2025-07-02NTT DOCOMO INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
JP2024216074
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-20
Filing Date
2024-12-11
Publication Date
2025-07-02

AI Technical Summary

Technical Problem

Current image generation tools lack authenticity and diversity, making them unreliable for training neural network models that require high-quality images.

Method used

An image generation device comprising a feature generator that extracts initial features, applies a randomization process, and uses a discriminator to train the generator, enhancing both authenticity and diversity through a feature extraction unit and random network unit.

Benefits of technology

The device generates images with high diversity and authenticity, suitable for training neural networks, and can be applied in various scenarios including target detection, artistic creation, and dataset expansion.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2025098968000001_ABST
    Figure 2025098968000001_ABST
Patent Text Reader

Abstract

To provide an image generation device, an image generation method and a non-transitory computer readable storage medium.SOLUTION: An image generation device comprises: a feature quantity generator which generates feature quantities on the basis of first input data; an image generator which generates a first image on the basis of second input data and the feature quantities; and a discrimination apparatus which discriminates the generated first image, reversely propagates output of the discrimination apparatus to the feature quantity generator to train the feature quantity generator. The feature quantity generator comprises: a feature extraction unit which extracts an initial feature quantity from the first input data; and a random network unit which performs randomization processing to the initial feature quantity to generate feature quantities at random distance from the initial feature quantity.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence, and particularly to an image generation device, an image generation method, and a non-transitory computer-readable storage medium.

Background Art

[0002] In recent years, with the development of artificial intelligence (AI) technology, various image generation technologies based on AI have emerged one after another. For example, some image generation tools can automatically generate an image intended by a user based on the user's text description or image reference.

[0003] However, current mainstream image generation tools in the industry (such as tools based on simple diffusion models, tools based on ChatGPT, etc.) still have the problem that the authenticity and diversity of the generated images are low, and it is impossible for users to completely rely on such image generation tools to obtain ideal images.

[0004] In some application scenarios (such as target detection), in order to improve the accuracy of a neural network model, many high-authenticity and high-diversity generated images are required to train the neural network model. It is obvious that it is unrealistic to rely only on the above-mentioned image generation models to provide training data.

Summary of the Invention

[0005] This application is made in view of the above problems. In one exemplary aspect, the present disclosure provides an image generation apparatus including a feature quantity generator that generates a feature quantity based on first input data, an image generator that generates a first image based on second input data and the feature quantity, and a discriminator that discriminates the generated first image and backpropagates the output of the discriminator to the feature quantity generator to train the feature quantity generator. The feature quantity generator includes a feature extraction unit that extracts an initial feature quantity from the first input data, and a random network unit that performs a randomization process on the initial feature quantity and generates the feature quantity having a random distance from the initial feature quantity.

[0006] In some embodiments, the first input data includes a category of a target object and at least one of guidewords related to the target object, and the second input data includes a reference image of the target object.

[0007] In some embodiments, the second input data further includes a category of the target object and a bounding box in the reference image surrounding the target object.

[0008] In some embodiments, the image generator is an image generator based on a diffusion model, and the feature quantity generator and the discriminator are a generator and a discriminator based on an adversarial generation network, respectively.

[0009] In some embodiments, the discriminator includes a classification unit that generates a first probability distribution related to the first image, and a loss conversion unit that generates a second probability distribution that is a simulation of the first probability distribution based on the feature quantity, and the loss conversion unit determines a first loss based on the first probability distribution and the second probability distribution and outputs the loss as the output of the discriminator.

[0010] In some embodiments, when the difference between the first probability distribution and the second probability distribution is smaller than a first threshold value, or when the number of training times reaches a second threshold value, the training of the feature generator is stopped.

[0011] In another exemplary aspect, the present disclosure provides a neural network-based system including an image generator as described above and a neural network model. After the feature generator is trained, the image generated by the image generator is used as training data to train the neural network model.

[0012] In some embodiments, after the feature generator is trained, the image generated by the image generator, the real image including the target object, and the enhanced image of the real image are all used as training data to train the neural network model.

[0013] In another exemplary aspect, the present disclosure provides an image generation method including generating a feature amount based on first input data, generating a first image based on second input data and the feature amount, discriminating the generated first image, and backpropagating the output of the discriminator to the feature generator that generates the feature amount to train the feature generator. Generating a feature amount based on the first input data includes extracting an initial feature amount from the first input data and performing a randomization process on the initial feature amount to generate the feature amount having a random distance from the initial feature amount.

[0014] In yet another exemplary aspect, the present disclosure provides a non-transitory computer-readable storage medium storing computer instructions that, when executed by a processor, perform operations including generating a feature amount based on first input data, generating a first image based on second input data and the feature amount, discriminating the generated first image, and backpropagating an output of a discriminator to a feature amount generator that generates the feature amount to train the feature amount generator. Generating a feature amount based on the first input data includes extracting an initial feature amount from the first input data and performing a randomization process on the initial feature amount to generate the feature amount having a random distance from the initial feature amount.

Brief Description of Drawings

[0015]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

DETAILED DESCRIPTION OF THE INVENTION

[0016] Hereinafter, preferred embodiments of the present disclosure will be described in more detail.

[0017] It should be noted that each step described in the embodiment of the method of the present disclosure may be executed in a different order and / or executed in parallel. In addition, the method embodiment may include other steps and / or some steps may be omitted.

[0018] As used herein, the term "comprising" and its variations are inclusive, meaning "including but not limited to". The term "based on" means "at least partially based on". The term "one embodiment" represents "at least one embodiment", the term "another embodiment" represents "at least one another embodiment", and the term "some embodiments" represents "at least some embodiments". Related definitions of other terms are given in the following description.

[0019] It should be noted that concepts such as "first", "second", etc. mentioned in the present disclosure are only used to distinguish different devices, modules, or units, and do not limit the order or interdependence of the functions executed by these devices, modules, or units.

[0020] It should be noted that the modifiers "one" and "a plurality of" mentioned in the present disclosure are not limiting but general. Unless otherwise explicitly indicated in the context, they should be understood as "one or a plurality of".

[0021] In order to improve the authenticity and diversity of the generated images, the present disclosure provides an image generation device. FIG. 1 shows an exemplary functional block diagram of an image generation device 1000 according to an embodiment of the present disclosure.

[0022] As shown in FIG. 1, the apparatus 1000 functionally includes a feature generator 1100, an image generator 1200, and a discriminator 1300.

[0023] The feature generator 1100 generates features based on the first input data. The image generator 1200 generates a first image based on the second input data and the features generated by the feature generator 1100. The discriminator 1300 discriminates the first image generated by the image generator 1200, and backpropagates the output of the discriminator to the feature generator 1100 to train the feature generator 1100.

[0024] In order to facilitate the discrimination process of the discriminator 1300, at least a part of the data (e.g., the actual category of the object) included in the first input data is also provided to the discriminator 1300 as an input.

[0025] In order to improve the diversity of the images generated by the apparatus 1000, the feature generator in the apparatus 1000 according to the present application can generate features having a certain degree of randomness based on the first input data.

[0026] The so-called "diversity" means, for the image generation apparatus 1000, a plurality of images of different styles that the apparatus 1000 can generate for the same or similar inputs. For example, "diversity" can be represented by enabling the generation of different objects, shapes, and structures within the image. This is particularly important for applications such as artistic creation, movie special effects, and the expansion of training datasets. Also, "diversity" can represent different textures, writing effects, rendering effects, etc. of the generated images, and can visually diversify the generated images.

[0027] FIG. 2 shows an exemplary functional block diagram of the feature generator 1100 in an image generation apparatus according to an embodiment of the present disclosure.

[0028] As shown in FIG. 2, functionally, the feature generator 1100 includes a feature extraction unit 1110 and a random network unit 1120. The first input data is provided to the feature extraction unit 1110. The feature extraction unit 1110 performs a feature extraction operation on the first input data to extract initial features from the first input data.

[0029] The extracted initial features are provided to the random network unit 1120, and the random network unit 1120 performs a randomization process on the initial features to generate features having a random distance from the initial features.

[0030] In the present application, the first input data and the second input data may be data directly input by the user, or may be data extracted by the apparatus 1000 based on the user's input.

[0031] Note that in the present application, the first input data may usually be text data. For example, the first input data may include at least one of the category of the target object that the user wants to generate and a guide word related to the target object.

[0032] FIG. 3 shows a schematic diagram of exemplary input data and output data in an image generation apparatus according to an embodiment of the present disclosure.

[0033] As shown in FIG. 3, when the user wants to generate an image in which the target object is a "bird", as the first input data, the category of the target object (for example, bird) and / or a guide word related to the target object (for example, multicolored bird) can be provided to the feature generator 1100.

[0034] The feature generator 1100 generates features with a certain degree of randomness based on the first input data and provides them to the image generator 1200. Here, the so-called "features" generally refer to a feature matrix that characterizes the features of each dimension of the target object and is known to those skilled in the art in the field of neural networks, but will not be elaborated here.

[0035] Also, the second input data is provided to the image generator 1200 together with the features generated by the feature generator 1100.

[0036] In some examples, the second input data may be, for example, a reference image including the target object that the user wants to generate, or it may be image data.

[0037] In other examples, the second input data may include text data such as, for example, the category, size, pose, etc. of the target object.

[0038] Optionally, the second input data may further include a bounding box in the reference image that surrounds the target object, such as the diagonal coordinates (x1, y1), (x2, y2), etc. of the bounding box.

[0039] FIG. 3 shows some examples of the second input data, including, for example, a reference image Pref including a target object (e.g., "bird"), a character indicating the category of the target object (e.g., "bird"), and a bounding box (e.g., the white frame bbox in the image Pref) in the reference image Pref that surrounds the target object.

[0040] Note that the examples of the first input data and the second input data shown in FIG. 3 are merely illustrative and not limiting. For example, in some other embodiments, the second input data of plain text is also possible. For example, when generating a bird, the second input data may be the category name "bird" of the target object. In this case, the image generator can perform image generation based on the text, thereby outputting an image including the target object "bird" (for example, the first image P1 shown in FIG. 3).

[0041] To easily understand the beneficial contribution of the feature generator 1100 in the present specification to the diversity of the generated images, the specific operation mode of the image generation apparatus 1000 according to the present application will be described in detail with reference to FIGS. 3 and 4 in combination.

[0042] FIG. 4 shows an exemplary schematic diagram of the feature generator 1100 in the image generation apparatus 1000 according to an embodiment of the present disclosure generating features.

[0043] Assume that the user aims to generate an image of a bird, and the first input data provided to the feature generator 1100 by the user is a guiding word in the form of text such as "bird". In this case, the feature extraction unit 1110 in the feature generator 1100 shown in FIG. 2 extracts features from the first input data to obtain an initial feature amount O (indicated by a black circle), also called an anchor feature amount, shown in FIG. 4.

[0044] For convenience of explanation, assume that the feature amount extracted by the feature extraction unit 1110 shown in FIG. 2 is a feature amount related to the dimensional data of the color of the bird. That is, if the color of the bird characterized by the initial feature amount O shown in FIG. 4 is "gray", when the initial feature amount O characterizing "gray" is input to the random network unit 1120 shown in FIG. 2, the random network unit 1120 performs a randomization process on the initial feature amount O to generate a random feature amount.

[0045] For example, the random network unit 1120 may include one random number generator that can generate one random offset amount. The randomization process means that the random network unit 1120 adds a random offset amount to the initial feature amount O so that there is a random feature distance (or random offset amount) between the newly generated random feature amount and the initial feature amount O. The gray circles in FIG. 4 show some examples of random feature amounts.

[0046] For example, the random feature amount S1 shown in FIG. 4 may be a feature amount in which the color of the bird subjected to the randomization process is "blue", and the random feature amount S2 shown in FIG. 4 may be a feature amount in which the color of the bird subjected to the randomization process is "yellow".

[0047] Furthermore, when the random feature amount (for example, S1 or S2 shown in FIG. 4) randomized through the random network unit 1120 shown in FIG. 2 is input to the image generator 1200 shown in FIG. 1 or FIG. 3, the image generator 1200 can generate more diverse images based on the random feature amount and the second input data.

[0048] For example, when the randomization process is not performed on the initial feature amount O or when only the second input data is input to the image generator 1200, the image generator 1200 can generate a series of images of various forms of gray birds. When the randomization process is performed on the extracted initial feature amount O and the random feature amount is input to the image generator 1200, the image generator 1200 may generate images of birds of various colors, and the diversity of the generated images is greatly increased.

[0049] Note that the randomization process of the feature amount described above for the color of the bird is merely exemplary. In actual applications, the feature amount extracted from the first input data is a plurality of feature amounts represented by a multi-dimensional matrix and can represent multiple-dimensional features of the target object.

[0050] For example, taking the guideword where the first input data is "bird" as an example, the feature amounts represented by the multi-dimensional matrix may include not only the feature amounts related to the color of the bird, but also the feature amounts related to other dimensions such as the size of the bird, the type of the bird, the posture of the bird, etc. These feature amounts can be randomized through the random network unit 1120 according to the present application so as to generate various combinations of random feature amounts.

[0051] In some embodiments, for example, when the random feature amounts output by the random network unit 1120 shown in FIG. 2 include feature amounts in various dimensions such as the color, size, type, posture, etc. of the bird, the image generator 1200 shown in FIGS. 1 and 3 can advantageously utilize these random feature amounts to generate various bird images with freely combined various colors, sizes, types, and postures.

[0052] It should be noted that the randomization process of the initial feature amount O by the above-mentioned random network unit 1120 does not mean an unrestricted randomization process. The randomization process of the initial feature amount O by the random network unit 1120 is preferably performed within a certain threshold range.

[0053] For example, as shown in FIG. 4, the range of the randomization process can be limited within the range delimited by the dashed circle shown in the figure. That is, the randomization process performed on the initial feature amount O by the random network unit 1120 shown in FIG. 2 needs to be performed within a specific distance radius R so that the generated random feature amount does not deviate greatly from the initial feature amount O.

[0054] In some examples, for example, since the random offset amount generated by the random network unit 1120 is equal to or less than the threshold radius R, and the newly generated random feature amount is limited within the threshold radius R shown in FIG. 4, one threshold radius R may be set for the feature amounts of each dimension.

[0055] For example, with respect to the feature quantity of the size of a bird, in order to avoid the finally generated image lacking authenticity, the size of the bird corresponding to the generated random feature quantity should not exceed the size of a real-world bird, and the corresponding threshold radius R should be reasonably set. For example, with respect to the feature quantity of the color of a bird, in order to avoid the finally generated image lacking authenticity, the color of the bird corresponding to the generated random feature quantity should not exceed the color of a real-world bird, and the corresponding threshold radius R should also be reasonably set.

[0056] The threshold radius R corresponding to the feature quantity of each dimension is usually adjusted step by step and optimized step by step in the process of training the feature quantity generator 1100, so that the diversity of the generated images can be ensured and the authenticity of the generated images can be ensured.

[0057] Evaluating the performance of an image generation device only based on the diversity of the generated images is often insufficient. In actual applications, the image generation device hopes to generate images with high diversity (for example, images of birds with various colors, various categories, various postures, and various sizes), and at the same time hopes that the generated images do not deviate from reality, that is, it is necessary to have a high degree of authenticity such that humans cannot visually distinguish whether the image was created by a human or generated by a machine.

[0058] Examples with high diversity but low authenticity include, for example, when a user attempts to generate an image of a bee. An image generation device can generate a large number of photos of bees in various forms based on the keyword "bee", indicating that the diversity performance of the image generation device is sufficient. However, upon detailed observation of each generated image, it was found that there are many images of "false" bees. For example, images with rabbit ears growing from a bee's head, images with butterfly wings growing on a bee, or images of a bee swimming in water. Such images can be easily discriminated by the human eye as being generated by a machine, so their authenticity is not sufficient.

[0059] Hereinafter, by combining FIGS. 1 to 4 and further combining FIGS. 5 to 7, various embodiments will be further described on how the image generation device according to the present application improves the diversity of the generated images and at the same time ensures the authenticity of the generated images.

[0060] As described above in combination with FIG. 1, the discriminator 1300 discriminates the first image generated by the image generator 1200, and backpropagates the output of the discriminator to the feature quantity generator 1100 to train the feature quantity generator 1100. In the present application, the setting of the discriminator 1300 and the related backpropagation and training process are the main means to improve the authenticity of the generated images.

[0061] In the present application, the main task of the discriminator 1300 is to discriminate whether the first image is a real image or an image generated by a machine. From this, it can be understood that the greater the distribution difference between the first image and the real image (Ground Truth), the easier it is for the discriminator to distinguish whether the first image is real or generated by a machine. The main task of the image generator 1200 is to generate a first image as close as possible to the real image so as to deceive the discriminator 1300 and enable the discriminator 1300 to output a positive result.

[0062] In some embodiments, the image generator 1200 according to the present application may be, for example, an image generator based on a diffusion model, and the feature generator 1100 and the discriminator 1300 may be a generator and a discriminator based on a generative adversarial network (GAN), respectively.

[0063] In a generative adversarial network (GAN), the target of the generator G is to generate real data as much as possible to deceive the discriminator D. The target of the discriminator D is to distinguish the data generated by the generator G from the real data as much as possible. In this way, the generator G and the discriminator D constitute a dynamic "game theory".

[0064] In the present application, the discriminator 1300 is a well-trained discriminator, that is, it is assumed that the discriminator 1300 can sufficiently distinguish the authenticity of the generated first image.

[0065] It should be noted that in the present application, the image generator 1200 is also a well-trained image generator, that is, it is assumed that the image generator 1200 itself can already generate a first image close to a real image based on the input to the maximum extent.

[0066] In this case, the impact on the authenticity performance of the device 1000 focuses on the representation of the feature generator 1100. The representation of the feature generator 1100 depends on the training process of how the discriminator 1300 provides an appropriate output (for example, the loss calculated by the loss function, described later) and backpropagates it to the feature generator 1100. The feature generator 1100 can adjust the network parameters based on the loss function and provide appropriate features so that the image generator 1200 can generate a first image close to a real image.

[0067] Therefore, the training described in this application actually corresponds to fixing the image generator 1200 and the discriminator 1300 and training the feature generator 1100 alone.

[0068] In the model according to the present disclosure, an image generator 1200 is inserted between the feature generator 1100 and the discriminator 1300, and the image generator 1200 cannot perform effective backpropagation (for example, due to an algorithm such as multiple iterations in a diffusion model). Therefore, the discrimination result generated by the discriminator 1300 cannot be directly backpropagated to the feature generator 1100 through the image generator 1200 to train the feature generator 1100. For this reason, in addition to the classifier, this application also proposes a discriminator with a correction ability (which may also be called a loss conversion ability).

[0069] FIG. 5 shows an exemplary functional block diagram of the discriminator 1300 in the image generation apparatus 1000 according to an embodiment of the present disclosure.

[0070] As shown in FIG. 5, the discriminator 1300 may functionally include a classification unit 1310 and a loss conversion unit 1320.

[0071] For example, when a first image is provided to the classification unit 1310, the classification unit 1310 classifies the first image and generates a first probability distribution related to the first image. As described above, when the classification unit 1310 of the discriminator 1300 is sufficiently trained, the first probability distribution generated by the classification unit 1310 may be considered as the actual probability distribution characterizing the first image category.

[0072] Note that inputting the first input data shown in FIG. 5 (for example, the classification "bird" shown in FIG. 3) to the classification unit 1310 is to calculate the loss, which will be described later.

[0073] As described above, the discrimination result generated by the discriminator 1300 cannot be backpropagated to the feature quantity generator 1100 via the image generator 1200 to train the feature quantity generator 1100. Therefore, in order to train the feature quantity generator, the present application proposes to achieve the purpose of training the feature quantity generator 1100 by having the loss conversion unit 1320 in the discriminator 1300 generate a second probability distribution based on the feature quantity received from the feature quantity generator 1100, calculating a loss based on the first probability distribution etc. and the second probability distribution etc., and directly backpropagating the loss to the feature quantity generator 1100.

[0074] FIG. 6 shows a schematic diagram of a discriminator in an image generation apparatus according to an embodiment of the present disclosure simulating a probability distribution.

[0075] For example, as shown in FIG. 6, the classification unit 1310 can classify the generated first image and generate a first probability distribution P1 related to the first image as shown on the left side of FIG. 6. The loss conversion unit 1320 can generate a second probability distribution P2 that can be regarded as a simulation to the first probability distribution P1 based on the input feature quantity. For example, when not sufficiently trained, as shown in the figure, there is an obvious difference between the second probability distribution P2 simulated based on the feature quantity by the loss conversion unit 1320 and the first probability distribution P1.

[0076] Then, the loss conversion unit 1320 can determine a first loss based on the first probability distribution and the second probability distribution and output it as the output of the discriminator 1300.

[0077] In the present application, the overall loss of the image generation apparatus (i.e., the above-described first loss) depends on the corresponding losses of three functional units: the random network unit 1120, the classification unit 1310, and the loss conversion unit 1320.

[0078] FIG. 7 shows a schematic diagram of a discriminator in an image generation apparatus according to an embodiment of the present disclosure determining a first loss.

[0079] As shown in FIG. 7, the loss MSL1 of the random network unit 1120 can be calculated based on the following equation (1). MSL1 = MSE(feature quantity in , feature quantity out ) (1) However, the feature quantity in represents the feature quantity input to the random network unit 1120, for example, the initial feature quantity shown in FIG. 4. The feature quantity out represents the feature quantity output by the random network unit 1120, for example, the random feature quantity shown in FIG. 4. MSE represents the mean squared error loss function that measures the average value of the square of the difference between the feature quantity in and the feature quantity out .

[0080] The loss MSL1 of the random network unit 1120 is used to ensure that the random disturbance added by the random network unit 1120 to the input feature quantity is kept within a reasonable range.

[0081] Note that FIG. 7 also shows the loss MSL2 of the loss conversion unit 1320, which can be calculated based on the following equation (2). MSL2 = MSE(classification unit out , loss conversion unit out ) (2) However, the classification unit out represents the output of the classification unit 1310 (for example, the first probability distribution described above), and the loss conversion unit out represents the output of the loss conversion unit 1320, for example, the second probability distribution generated by simulating the first probability distribution based on the feature quantity out described above. MSE represents the mean squared error loss function that measures the average value of the square of the difference between the classification unit out and the loss conversion unit out .

[0082] The loss MSL2 of the loss conversion unit 1320 is used to ensure that the loss conversion unit 1320 can accurately simulate the output of the classification unit 1310.

[0083] Note that FIG. 7 also shows the loss CEL1 of the classification unit 1310, which can be calculated based on the following equation (3). CEL1 = CEL(GT, loss conversion unit out ) (3) However, GT represents the actual value input from the first input data shown in FIG. 1, such as the actual classification "bird" shown in FIG. 3, into the classification unit 1310, and the loss conversion unit out represents the output of the loss conversion unit 1320, such as the second probability distribution generated by simulating the first probability distribution based on the above-described feature amount out . CEL represents the cross-entropy loss function that measures the similarity between two probability distributions. For example, when GT represents the classification "bird", the user measures the similarity between the generated first image and the bird image.

[0084] The loss CEL1 of the classification unit 1310 is used to calculate the cross-entropy loss based on the output of the loss conversion unit 1320 and the actual value GT.

[0085] Based on the above three loss functions, the loss function L that determines the first loss backpropagated to the feature generator 1100 is calculated by weighting the above three loss functions, and is shown in, for example, equation (4). L = a * MSL1 + b * MSL2 + c * CEL1(4) However, a, b, and c represent the weights of the loss MSL1 of the random network unit 1120, the loss MSL2 of the loss conversion unit 1320, and the loss CEL1 of the classification unit 1310, respectively.

[0086] The first loss L calculated by equation (4) can be backpropagated to the feature generator 1100 to train the feature generator 1100.

[0087] As long as an appropriate first loss is determined, the distance between the first probability distribution of the first image generated thereafter and the generated second probability distribution can be made as close as possible.

[0088] Training the feature generator 1100 includes training the feature extraction unit 1110 shown in FIG. 2, and also includes training the random network unit 1120. As a result, the feature generator 1100 outputs more realistic random feature quantities. Further, the image generator 1200 can generate a first image of a more realistic image based on the second input data and the more realistic random feature quantities.

[0089] In some embodiments, it is possible to determine whether to stop training the feature generator 1100 based on the difference (or distance) between the first probability distribution (for example, the true probability distribution) of the first image and the second probability distribution. For example, when the difference between the first probability distribution and the second probability distribution is smaller than the first threshold, training of the feature generator 1100 can be stopped.

[0090] For example, when the feature generator 1100 and the discriminator 1300 in the present application are designed based on the GAN architecture, and when the image generator 1200 generates an image that is exactly the same as the true data based on the feature quantities output by the feature generator 1100, the discriminator 1300 cannot determine the result, and the probabilities of true and false are both 50%, which is a matter of random guessing. In this case, the feature generator 1100 has been sufficiently trained and will not update its own network parameters.

[0091] Optionally, when the number of training times reaches a second threshold (for example, 100 times), training of the feature generator 1100 is stopped.

[0092] As described above, a specific embodiment in which the discriminator 1300 trains the feature generator 1100 by backpropagating its output to improve the authenticity of the first image generated by the apparatus 1000 has been introduced.

[0093] When the feature generator 1100 is sufficiently trained and can generate images with high diversity and high authenticity, the generated images can be applied to various scenes.

[0094] The most typical application scenario is to train another neural network (for example, a target detection) model as training data.

[0095] The present disclosure further provides a neural network-based system. The system includes the image generation device described above with respect to any one of FIGS. 1 to 7, and a neural network model. When the training of the feature generator 1100 shown in the figure is completed (for example, when the difference between the first probability distribution and the second probability distribution is smaller than the first threshold, or when the number of training times of the feature generator 1100 reaches the second threshold), the image generated by the image generator 1200 is used as training data to train the neural network model.

[0096] For example, when it is detected that the difference between the first probability distribution and the second probability distribution of the first image generated by the image generator 1200 is smaller than the first threshold, or the number of training times reaches the second threshold, the image generated by the image generator 1200 thereafter is used as training data to train the neural network model.

[0097] For example, the neural network model may be a model for applications such as target detection, image beautification, and image restoration.

[0098] In addition, in order to use more diverse training data and improve the accuracy of the neural network model, after the training of the feature generator 1100 is completed, in order to train the neural network model, the image generated by the image generator 1200 thereafter, the real image including the target object, and the enhanced image of the real image may all be used as training data.

[0099] For example, the enhanced image of the real image may include an image generated after performing operations such as rotation, mirroring, scaling, trimming, etc. on the real image.

[0100] Therefore, the image generation device according to the present application is beneficial for providing a large amount of training data with high diversity and high authenticity.

[0101] Of course, the image generation device according to the present application is not limited to the application that provides training data, and the benefits in other related scenarios (such as artistic creation, image beautification, image restoration, image enhancement, etc.) are also obvious.

[0102] Another aspect of the present application provides an image generation method. FIG. 8 shows an exemplary flowchart of an image generation method 8000 according to an embodiment of the present disclosure.

[0103] Specifically, the method 8000 includes a step S8100 of generating a feature amount based on the first input data, a step S8200 of generating a first image based on the second input data and the feature amount, and a step S8300 of discriminating the generated first image and backpropagating the output of the discriminator to the feature amount generator that generates the feature amount to train the feature amount generator.

[0104] In step S8100, generating a feature amount based on the first input data further includes extracting an initial feature amount from the first input data and performing a randomization process on the initial feature amount to generate a feature amount having a random distance from the initial feature amount.

[0105] Similar to the descriptions related to FIGS. 1 to 7, the first input data and the second input data may be directly input by the user or may be extracted based on the user's input.

[0106] In the present application, the first input data may be, for example, text data in general. For example, the first input data may include the category of the target object and at least one of the guidewords related to the target object.

[0107] The second input data may be, for example, a reference image including the target object, or may be image data. In other examples, the second input data may include text data such as, for example, the category of the target object. Optionally, the second input data may further include a bounding box in the reference image surrounding the target object, such as diagonal coordinates (x1, y1), (x2, y2), etc. of the bounding box.

[0108] Note that the various features, various input data and output data, and various operation steps performed by each functional unit described with respect to FIGS. 1 to 5 are similarly applicable to the method embodiment shown in FIG. 8, unless clearly inapplicable.

[0109] In other aspects, the present disclosure further provides a non-transitory computer-readable storage medium for generating an image. FIG. 9 shows a schematic diagram of a non-transitory computer-readable storage medium 9000 for generating an image according to an embodiment of the present disclosure.

[0110] As shown in the figure, it is a non-transitory computer-readable storage medium 9000 in which computer instructions 9100 are stored. When the computer instructions 9100 are executed by a processor, it performs steps similar to those shown in FIG. 8, including step S8100 of generating a feature amount based on the first input data, step S8200 of generating a first image based on the second input data and the feature amount, and step S8300 of discriminating the generated first image and backpropagating the output of the discriminator to a feature amount generator to train the feature amount generator.

[0111] In step S8100, generating a feature amount based on the first input data further includes extracting an initial feature amount from the first input data and performing a randomization process on the initial feature amount to generate a feature amount having a random distance from the initial feature amount.

[0112] In addition, when the computer program instructions 9100 stored in the non-transitory computer-readable storage medium 9000 are executed by the processor, the various operation steps executed by each functional unit described with reference to FIGS. 1 to 7 can also be executed, and will not be described further.

[0113] In addition, the various features, the characteristics of the various input data and output data described with reference to FIGS. 1 to 7 are similarly applicable to the embodiment of the storage medium shown in FIG. 9, unless clearly not applicable.

[0114] As can be understood by those skilled in the art, it may be fully executed by hardware, may be fully executed by software (including firmware, resident software, microcode, etc.), or may be executed by a combination of hardware and software. Any of the above hardware or software may also be referred to as a "data block", "module", "engine", "unit", "assembly" or "system". In addition, various aspects of the present application may be realized as a computer product located on one or more computer-readable media, and the product includes computer-readable program code.

[0115] This application uses specific terms to describe the embodiments of this application. For example, "one embodiment", "an embodiment", and / or "some embodiments" mean features, structures, or characteristics related to at least one embodiment of this application. Therefore, it should be emphasized and noted that "an embodiment" or "one embodiment" or "one alternative embodiment" described in two or more different places in this specification does not necessarily refer to the same embodiment. Furthermore, some features, structures, or characteristics of one or more embodiments of this application can be appropriately combined.

[0116] Unless otherwise defined, all terms (including technical and scientific terms) used herein shall have the same meaning as commonly understood by one of ordinary skill in the art. Terms such as those defined in a normal dictionary shall be interpreted to have a meaning that coincides with their meaning in the context of the relevant art, without being interpreted in an idealized or overly formalized sense, unless explicitly defined herein.

[0117] The foregoing is a description of the present disclosure and should not be regarded as limiting it. Although several exemplary embodiments of the present disclosure have been described, those skilled in the art can readily understand that many modifications can be made to the exemplary embodiments without departing from the teachings and advantages of the present disclosure. Therefore, all such modifications are included within the scope of the present disclosure. Note that the foregoing is a description of the present disclosure and should not be regarded as limited to the specific embodiments disclosed.

Claims

1. a feature generator that generates features based on first input data; an image generator that generates a first image based on second input data and the feature amount; a classifier that classifies the generated first image and back-propagates an output of the classifier to the feature generator to train the feature generator; The image generating device includes: a feature extraction unit that extracts initial features from the first input data; and a random network unit that performs a randomization process on the initial features to generate the features having a random distance from the initial features.

2. The first input data includes a category of a target object and at least one of a guide word associated with the target object; The image generation apparatus of claim 1 , wherein the second input data comprises a reference image of the target object.

3. The image generating apparatus of claim 2 , wherein the second input data further comprises a category of the target object and a bounding box in the reference image that encloses the target object.

4. the image generator is a diffusion model based image generator; The image generating apparatus according to claim 1 , wherein the feature generator and the classifier are a generator and a classifier, respectively, based on a generative adversarial network.

5. The discriminator is a classifier for generating a first probability distribution associated with the first image; a loss conversion unit that generates a second probability distribution that is a simulation of the first probability distribution based on the feature amount, The image generating device according to claim 1 , wherein the loss conversion unit determines a first loss based on the first probability distribution and the second probability distribution, and sets the first loss as the output of the discriminator.

6. 6. The image generating device according to claim 5, wherein training of the feature generator is stopped when a difference between the first probability distribution and the second probability distribution is smaller than a first threshold or when a number of training times reaches a second threshold.

7. An image generating device according to any one of claims 1 to 6; a neural network model, After the feature generator is trained, the images generated by the image generator are used as training data to train the neural network model.

8. 8. The neural network-based system of claim 7, wherein after the feature generator is trained, the image generated by the image generator, a real image including a target object, and an augmented image of the real image are used as training data to train the neural network model.

9. generating features based on first input data; generating a first image based on second input data and the feature amount; discriminating the generated first image, and back-propagating an output of the discriminator to a feature generator that generates the feature, thereby training the feature generator; Generating features based on the first input data includes: Extracting initial features from the first input data; and performing a randomization process on the initial feature amount to generate the feature amount having a random distance from the initial feature amount.

10. A non-transitory computer-readable storage medium having stored thereon computer instructions, which, when executed by a processor, comprise: generating features based on first input data; generating a first image based on second input data and the feature amount; A feature generator that generates the feature quantity is trained by discriminating the generated first image and backpropagating an output of the discriminator to the feature quantity generator, Generating a characteristic quantity based on the first input data includes: Extracting an initial characteristic amount from the first input data; and performing a randomization process on the initial features to generate features having a random distance from the initial features.