Image generation device, image generation method and program
The image generation apparatus addresses the practicality issue of controlling generative AI with large latent variable dimensions by iteratively refining the image generation process using differentiable operations and user feedback, significantly reducing the number of user interactions required.
Patent Information
- Application Number
- JP2023205505
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-17
AI Technical Summary
Conventional generative AI control methods using HITL optimization are impractical for diffusion models with a large number of latent variable dimensions, requiring an excessive number of user interactions to achieve desired image generation.
An image generation apparatus and method that involves generating multiple second latent variables from a first latent variable, converting these into higher-dimensional third latent variables using differentiable operations, and iteratively refining the search space based on user feedback to optimize image generation.
Enables efficient control of generative AI based on diffusion models with large latent variable dimensions, reducing the number of user interactions needed to generate high-quality images that meet user preferences.
Smart Images

Figure 2025090326000001_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an image generation device, an image generation method, and a program.
Background Art
[0002] There is a generative AI technology that uses a model trained with a large amount of data to generate new data having the same characteristics as the training data. Among them, there is image generation AI as a technology that learns a large amount of images and generates new images. Image generation AI can be broadly classified into three types of generation models. They are VAE (Variational Auto Encoder), GAN (Generative Adversarial Network), and diffusion models, respectively. Diffusion models can generate images with high quality such that they cannot be distinguished from real photos, paintings, or illustrations made by humans. Also, in diffusion models, a method called Text-to-Image, in which a character string (prompt) indicating an image to be generated is input and an image is output, is common.
[0003] Diffusion models learn a diffusion process of gradually adding noise to an image and an inverse diffusion process (or generation process) of gradually removing noise from the image. By using a noise (latent variable) generated by a random number as an input and calculating the inverse diffusion process using a learned model to remove the noise, a new image is generated from the noise. Non-Patent Document 1 discloses LDM (Latent Diffusion Model), which is one of the diffusion models. In LDM, in addition to the normal diffusion model, it has an encoder-decoder capable of mutually converting an image space and a latent space. In LDM, instead of calculating the diffusion process and the inverse diffusion process in the image space like the normal diffusion model, the diffusion process and the inverse diffusion process are calculated in the latent space. Further, by performing noise removal in consideration of the prompt during the calculation of the inverse diffusion process, it is possible to generate an image according to the prompt.
[0004] In the LDM, even when the same prompt is used, different images are generated from the noise (latent variables) generated with different random numbers. At this time, it is difficult to reflect detailed instructions such as the positional relationship and composition of the object in the image in the prompt. For example, when the string "A cat standing surrounded by colorful flowers" is input as the prompt for the image to be generated, an image with "colorful flowers" and "a cat" will be generated. However, regarding the composition such as the position of the flowers and the cat, the line of sight and posture of the cat, various images are generated depending on the noise (latent variables) generated by the input random numbers. Therefore, the user needs to repeatedly input the latent variables generated with different random numbers until an image with the desired composition is generated, and repeatedly generate images using the input latent variables. However, the range of values that the latent variables can take is enormous, and it is not realistic to generate images corresponding to each of the latent variables in that range. Often, no matter how many times the image generation is repeated, the user cannot generate the image that they feel is the most desirable.
[0005] On the other hand, Non-Patent Document 2 discloses a control technique for generative AI by HITL (Human-in-the-Loop) optimization that controls the generative AI by repeating intuitive and simple operations by the user. In the control of generative AI by HITL optimization, while the generative model generates products corresponding to each latent variable in the latent space, a set of latent variables (search space) on a straight line provided in the latent space is set. The user selects the product that the user feels is the most desirable from among the products corresponding to each latent variable in the search space by adjusting the slider bar corresponding to the straight line provided in the latent space. Then, in the control of generative AI by HITL optimization, the next search space is set from the latent variable corresponding to the product selected by the user. By repeating such setting of the search space and selection by the user, it is possible to search for the latent variable that can generate the product that the user feels is the most desirable without performing a full search of the latent space. That is, from the user's perspective, it is possible to obtain the most desirable product by repeating the simple operation of adjusting the slider bar.
Prior Art Documents
Non-Patent Literature
[0006]
Non-Patent Literature 1
Non-Patent Literature 2
Summary of the Invention
Problems to be Solved by the Invention
[0007] However, in the control technology of generative AI by HITL optimization, the number of dimensions of the latent variables greatly affects the number of repetitions of slider bar adjustment. For example, in Non-Patent Literature 2, in a generative model with 500-dimensional latent variables, about 50 repetitions of processing are performed. For example, in Non-Patent Literature 2, in order to handle a generative model with relatively low-dimensional latent variables, instead of a diffusion model, a GAN is used. On the other hand, in diffusion models typified by LDM, the dimension of the latent variables often exceeds 1000 dimensions. Therefore, if Non-Patent Literature 2 is directly applied to a diffusion model, it will force the user to perform an enormous number of repetitions of work, which is not realistic. Thus, in conventional generative AI control, it has been difficult to apply generative AI control by HITL optimization to diffusion models with a relatively large number of dimensions of latent variables.
[0008] The present invention has been made based on the above problems, and an object of the present invention is to provide an image generation apparatus, an image generation method, and a program capable of controlling a generative AI by HITL optimization for a generative AI based on a diffusion model in which the number of dimensions of latent variables is relatively large.
Means for Solving the Problems
[0009] The image generation apparatus of the present invention is an image generation apparatus that generates an image, and includes: a search space setting unit that generates a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable, and sets a space including the generated plurality of second latent variables as a search space; a latent variable conversion unit that converts the second latent variable into a third latent variable having a larger number of dimensions than the second latent variable; an image generation unit that generates a generated image by executing an inverse diffusion process in a diffusion model using the third latent variable; an image presentation unit that presents a plurality of the generated images generated by the image generation unit using the third latent variable converted by the latent variable conversion unit to a user corresponding to the plurality of second latent variables included in the search space; and a latent variable update unit that re-sets the second latent variable corresponding to the selected image selected by the user from the plurality of the generated images presented by the image presentation unit as the first latent variable. The latent variable conversion unit converts the second latent variable into the third latent variable using an arithmetic process composed of differentiable operations, the search space setting unit calculates a differential with respect to the generated image generated by the arithmetic process and the inverse diffusion process, and generates a plurality of the second latent variables along a direction in which the value of the differential increases in a latent space corresponding to the first latent variable.
[0010] The image generation method of the present invention is an image generation method performed by a computer which is an image generation device for generating an image. The method includes: a search space setting step of generating a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable and setting a space including the generated plurality of second latent variables as a search space; a latent variable conversion step of converting the second latent variable into a third latent variable having a larger number of dimensions than the second latent variable; an image generation step of generating a generated image by performing an inverse diffusion process in a diffusion model using the third latent variable; presenting a plurality of the generated images generated by the image generation step using the third latent variable converted by the latent variable conversion step to a user corresponding to the plurality of second latent variables included in the search space; resetting the second latent variable corresponding to the selected image selected by the user from the plurality of presented generated images as the first latent variable; in the latent variable conversion step, converting the second latent variable into the third latent variable using an arithmetic process composed of differentiable operations; in the search space setting step, calculating a differential with respect to the generated image generated by the arithmetic process and the inverse diffusion process, and generating a plurality of second latent variables along a direction in which the value of the differential increases in a latent space corresponding to the first latent variable.
[0011] The program of the present invention causes a computer, which is an image generation device that generates images, to generate a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable, and sets a space including the generated plurality of second latent variables as a search space in a search space setting step, executes a latent variable conversion step of converting the second latent variable into a third latent variable having a larger number of dimensions than the second latent variable, executes an image generation step of generating a generated image by executing an inverse diffusion process in a diffusion model using the third latent variable, causes a user to be presented with a plurality of the generated images generated in the image generation step using the third latent variable converted by the latent variable conversion step corresponding to the plurality of second latent variables included in the search space, causes the second latent variable corresponding to the selected image selected by the user from the plurality of presented generated images to be reset as the first latent variable, in the latent variable conversion step, uses an arithmetic process composed of differentiable operations to convert the second latent variable into the third latent variable, in the search space setting step, calculates a derivative with respect to the generated image generated by the arithmetic process and the inverse diffusion process, and generates a plurality of second latent variables along a direction in which the value of the derivative increases in the latent space corresponding to the first latent variable.
Advantages of the Invention
[0012] According to the present invention, it is possible to control a generative AI by HITL optimization for a generative AI based on a diffusion model in which the number of dimensions of latent variables is relatively large.
Brief Description of the Drawings
[0013]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6A
Figure 6B
Figure 7
Embodiments for Carrying Out the Invention
[0014] Hereinafter, the image generation device 1 according to the embodiment will be described with reference to the drawings.
[0015] (Overview of Each Component of the Image Generation Device 1) First, an overview of each component of the image generation device 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing an example of the configuration of the image generation device 1 according to the present embodiment. As shown in FIG. 1, the image generation device 1 includes, for example, a search space setting unit 100, a latent variable conversion unit 101, an image generation unit 102, an image presentation unit 103, a latent variable update unit 104, a user selection reception unit 105, a first latent variable storage unit 106, a second latent variable storage unit 107, a third latent variable storage unit 108, and a generated image storage unit 109.
[0016] The search space setting unit 100 sets a search space. The search space here is a space in which latent variables (low-dimensional latent variables) corresponding to each of a plurality of images presented to allow a user to select the most desirable image exist. The search space is a partial space of the latent space (a space in which all low-dimensional latent variables exist). The specific method by which the search space setting unit 100 sets the search space will be described in detail later. Here, the latent variable that serves as the basis for setting the search space by the search space setting unit 100 is an example of the first latent variable. The latent variable existing in the search space set by the search space setting unit 100 is an example of the second latent variable.
[0017] The latent variable conversion unit 101 converts the latent variable (low-dimensional latent variable) existing in the search space set by the search space setting unit 100 into a high-dimensional latent variable. There are a plurality of latent variables in the search space, and the latent variable conversion unit 101 calculates the high-dimensional latent variable corresponding to each of the plurality of latent variables (low-dimensional latent variables). That is, the latent variable conversion unit 101 calculates a plurality of latent variables (high-dimensional latent variables). The specific method by which the latent variable conversion unit 101 calculates the high-dimensional latent variable from the low-dimensional latent variable will be described in detail later. Here, the high-dimensional latent variable calculated by the latent variable conversion unit 101 is an example of the third latent variable.
[0018] The image generation unit 102 generates an image using the high-dimensional latent variable calculated by the latent variable conversion unit 101. The image generation unit 102 generates an image corresponding to each of the plurality of latent variables (high-dimensional latent variables) calculated by the latent variable conversion unit 101. That is, the image generation unit 102 generates a plurality of images. The specific method by which the image generation unit 102 generates an image from the high-dimensional latent variable will be described in detail later.
[0019] The image presentation unit 103 presents the plurality of images generated by the image generation unit 102 to the user and allows the user to select the image that the user feels is the most desirable from the presented plurality of images. The image presentation unit 103 acquires the identification information of the image selected by the user.
[0020] The latent variable update unit 104 updates the latent variable. Here, the latent variable to be updated by the latent variable update unit 104 is the latent variable (first latent variable) that serves as the basis of the search space set by the search space setting unit 100. The latent variable update unit 104 updates the first latent variable by using the low-dimensional latent variable associated with the latent variable (high-dimensional latent variable) corresponding to the image selected by the user as the latent variable (first latent variable) that serves as the basis of the search space set by the search space setting unit 100.
[0021] The first latent variable storage unit 106 stores information on the latent variable (latent variable y best ) that serves as the basis of the search space set by the search space setting unit 100. The second latent variable storage unit 107 stores information on the latent variables existing in the search space set by the search space setting unit 100. The third latent variable storage unit 108 stores information on the high-dimensional latent variables calculated by the latent variable conversion unit 101. The generated image storage unit 109 stores information on the images generated by the image generation unit 102.
[0022] (Outline of the processing flow performed by the image generation device 1) Next, an outline of the processing flow performed by the image generation device 1 will be described with reference to FIG. 2. FIG. 2 is a diagram for explaining the processing performed by the image generation device 1 of the embodiment. The image generation device 1 is configured to be able to generate an image that the user feels is most desirable by repeatedly executing the processing shown in steps S100 to S104 of FIG. 2. Also, the number of repetitions of the processing shown in steps S100 to S104 may be arbitrarily set. For example, it may be repeated until an image that satisfies the user is generated, or it may be configured to repeat until a preset upper limit number of times is reached.
[0023] Step S100: The search space setting unit 100 sets a search space. The search space setting unit 100 sets, for example, a search space based on the latent variable (first latent variable) stored in the first latent variable storage unit 106. The first latent variable is a latent variable of low dimension (for example, 100 dimensions). The search space setting unit 100 causes the second latent variable storage unit 107 to store, for example, a latent variable (second latent variable) existing in the search space set by the search space setting unit 100.
[0024] Step S101: The latent variable conversion unit 101 converts a low-dimensional latent variable into a high-dimensional latent variable. The latent variable conversion unit 101 converts, for example, each of the latent variables (second latent variables) stored in the second latent variable storage unit 107 into a latent variable of high dimension (for example, 1000 dimensions). The latent variable conversion unit 101 causes the third latent variable storage unit 108 to store, for example, the high-dimensional latent variable (third latent variable).
[0025] Step S102: The image generation unit 102 generates an image based on the high-dimensional latent variable. The image generation unit 102 generates, for example, an image (generated image) using each of the high-dimensional latent variables (third latent variables) stored in the third latent variable storage unit 108. The generated image is an image generated from the high-dimensional latent variable and is, for example, an image of 196608 dimensions. The image generation unit 102 causes the generated image storage unit 109 to store, for example, the generated image.
[0026] Step S103: The image presentation unit 103 presents a plurality of generated images generated by the image generation unit 102 to the user. The user selects one image that the user feels is the most desirable from among the plurality of presented images.
[0027] Step S104: The latent variable update unit 104 updates the first latent variable. The latent variable update unit 104 updates the first latent variable by using, as the first latent variable, the low-dimensional latent variable (second latent variable) associated with the latent variable (third latent variable) corresponding to one image selected by the user from among the plurality of images presented by the image presentation unit 103. Thereafter, the process returns to step S100, and a search space based on the updated first latent variable is set.
[0028] Here, as a comparative example, a process of generating an image by conventional HITL optimization will be described with reference to FIG. 3. FIG. 3 is a diagram showing the flow of a process of generating an image by conventional HITL optimization. The processes shown in steps S200, S203, and S204 in FIG. 3 are equivalent to the processes shown in steps S100, S103, and S104 in FIG. 2, respectively. Step S202: The image generation unit generates an image based on the low-dimensional latent variable. The image generation unit generates, for example, an image (generated image) using each of the low-dimensional latent variables (second latent variables) stored in the second latent variable storage unit 107.
[0029] Although it is possible to generate an image from a low-dimensional latent variable, it is necessary to use a model in which the dimension of the latent space is low-dimensional (for example, about 30 to 500 dimensions), such as a GAN. In a diffusion model typified by LDM, the dimension of the latent variable often exceeds 1000 dimensions. Applying it to step S202 in FIG. 3 is difficult, and even if it is applied, it is difficult to generate an image that the user feels is most desirable within a realistic number of repetitions.
[0030] As a countermeasure against this, in the present embodiment, a conversion process by the latent variable conversion unit 101 as shown in step S101 in FIG. 2 is provided. As a result, the image generation unit 102 can generate an image from a high-dimensional latent variable. By generating an image from a high-dimensional latent variable, the quality and variations of the generated image can be diversified. Therefore, the user can select an image that the user feels is most desirable from among images showing various compositions.
[0031] (Details of each component of the image generation device 1) Here, the details of each component of the image generation device 1 will be described.
[0032] (Regarding the latent variable conversion unit 101) First, the latent variable conversion unit 101 will be described with reference to FIGS. 4 to 6. FIGS. 4 to 6 are diagrams for explaining the process of calculating the third latent variable.
[0033] First, as a premise, it is desirable that the low-dimensional latent variable serving as the basis for conversion into a high-dimensional latent variable by the latent variable conversion unit 101 be a latent variable of 500 dimensions or less. Also, as the high-dimensional latent variable, a latent variable of 1000 dimensions or more is desirable. Furthermore, it is desirable that the high-dimensional latent variable follow a Gaussian noise. For example, a 100-dimensional latent variable as the low-dimensional latent variable and a 3000-dimensional latent variable as the high-dimensional latent variable can be considered as a configuration.
[0034] Also, considering the process of setting the search space performed by the search space setting unit 100, the latent variable conversion unit 101 needs to perform the conversion from the low-dimensional latent variable to the high-dimensional latent variable based on a differentiable operation. Here, examples of differentiable operations include arithmetic operations on real numbers. On the other hand, examples of non-differentiable operations include integer operations that perform arithmetic operations on real numbers (not in floating point but in integers), searching for the maximum or minimum value, and random number processing.
[0035] Also, in the process performed by the image generation unit 102, it is desirable that the high-dimensional latent variable converted by the latent variable conversion unit 101 follow a Gaussian noise. On the other hand, the process of directly calculating Gaussian noise is a random number process and thus non-differentiable, and such non-differentiable operations cannot be applied to the process of setting the search space by the search space setting unit 100.
[0036] As such a countermeasure, as shown in FIG. 4, the latent variable conversion unit 101 calculates, for example, a high-dimensional latent variable (latent variable z) from a low-dimensional latent variable (latent variable y) using a weighted sum of a plurality of pre-set high-dimensional latent variables (latent variable x). Thereby, the latent variable conversion unit 101 converts the low-dimensional latent variable into a high-dimensional latent variable.
[0037] Specifically, first, the latent variable conversion unit 101 pre-generates a set X of a plurality of high-dimensional latent variables x (fourth latent variables) that follow Gaussian noise. For example, the latent variable conversion unit 101 generates the set X of the latent variables x using Gaussian noise. The set X of the latent variables x can be represented by Equation (1). In Equation (1), X is a set of high-dimensional latent variables. The element x k is an element of the set X. The element x k is an M-dimensional latent variable for each. The element x k The number K is the size of the set X. Also, regarding the number of dimensions M of M dimensions, the relationship M>K holds.
[0038]
Equation
[0039] Next, the latent variable conversion unit 101 applies the element x k calculated by Equation (1) and the low-dimensional (K-dimensional) latent variable y to Equation (2) below to calculate the high-dimensional (M-dimensional) latent variable z. In Equation (2), z is a high-dimensional (M-dimensional) latent variable. y is a low-dimensional (K-dimensional) latent variable. x k is an element of the set X represented by Equation (1).
[0040]
Equation
[0041] Here, when the variation in each value of the latent variable y used in equation (2) is small, the variation in the value of the latent variable z calculated by applying equation (2) is likely to be small. In the process by the image generation unit 102, it is desirable that the high-dimensional latent variable z used in the diffusion model follows Gaussian noise. When the variation in the value of the latent variable z is small, the quality and variations of the image generated by the image generation unit 102 using the diffusion model tend to decrease. Therefore, it is better for the latent variable conversion unit 101 to be able to calculate a high-dimensional latent variable z with large variation.
[0042] As a countermeasure, as shown in FIG. 5, the latent variable conversion unit 101 may further perform a process of amplifying noise on the calculation result of the weighted sum, that is, the latent variable z calculated using equation (2), so as to increase the variation. Specifically, the latent variable conversion unit 101 calculates a latent variable z' with large variation by applying equation (3) to the latent variable z calculated using equation (2). In equation (3), z' is a high-dimensional (M-dimensional) latent variable. α is a constant of 1 or more indicating a magnification factor. mean(z) is the average value of each value of the latent variable z. The variation of the latent variable z' is larger than the variation of the latent variable z.
[0043]
Equation
[0044] By using equation (3), the latent variable conversion unit 101 can multiply the variation {z - mean(z)} of the value of the latent variable z by a constant factor α and add it to the average value mean(z), and can calculate a latent variable z' with an increased variation of the value of the latent variable z.
[0045] Here, the latent variable conversion unit 101 may set the magnification factor α in equation (3) according to the degree of variation in each value of the latent variable y. For example, the latent variable conversion unit 101 may be configured such that the smaller the variation in each value of the latent variable y, the larger the magnification factor α. For example, the latent variable conversion unit 101 calculates the variance of the values of the low-dimensional (K-dimensional) latent variable y used in the calculation of the latent variable z, and sets the magnification factor α to be inversely proportional to the calculated variance. This increases the variation in the latent variable z, and enables the image generation unit 102 to generate an image using a diffusion model with a large variation in the latent variable z. Therefore, it is possible to suppress a decrease in the quality and variation of the images generated in the present embodiment.
[0046] In the above, an example of calculating the high-dimensional (M-dimensional) latent variable z using equation (2) was described on the premise that the number of dimensions K of the low-dimensional latent variable y used in equation (2) and the size K of the set X of a plurality of high-dimensional (M-dimensional) latent variables x set in advance are the same value. However, it is better if the latent variable z can be calculated even when the number of dimensions K of the low-dimensional latent variable y and the size K' of the set X are different. As a countermeasure, the latent variable conversion unit 101 may calculate the weighting coefficient of the weighted sum shown in equation (2) using a learned machine learning model.
[0047] For example, as shown in FIG. 6A, the latent variable conversion unit 101 uses the value w output from the machine learning model by inputting the latent variable y into a machine learning model learned in advance as the weighting coefficient. The machine learning model is a model that has been learned in advance to output a value w of dimension K' from a latent variable y of dimension K.
[0048] Then, as shown in FIG. 6B, the latent variable conversion unit 101 applies the element x calculated in equation (1) and the weighting coefficient w calculated using the machine learning model to the following equation (4) to calculate a high-dimensional (M-dimensional) latent variable z''. In equation (4), z'' is a high-dimensional (M-dimensional) latent variable. w k and calculates a high-dimensional (M-dimensional) latent variable z''. In equation (4), z'' is a high-dimensional (M-dimensional) latent variable. w kis a weight coefficient based on a low-dimensional (K-dimensional) latent variable y. x k is an element of the set X shown in equation (1).
[0049] [Number]
[0050] Here, as the machine learning model used above, for example, a multi-layer perceptron can be used. By adopting a configuration using such a machine learning model, even when the dimensionality K of the latent variable y and the size K' of the set X are different, based on the latent variable y, a weighting coefficient w of the same dimensionality K' as the size K' of the set X can be calculated.
[0051] For example, as the machine learning model, a GAN (Generative Adversarial Network) consisting of a generator and a discriminator can be used. In this case, the generator in the GAN is a multi-layer perceptron that outputs a weight coefficient w for any vector within the range that the latent variable y can take. Also, the discriminator in the GAN is a discriminator that discriminates between a latent variable z'' calculated using the weight coefficient w generated by the generator and Gaussian noise generated from random numbers. By alternately training both the generator and the discriminator, a multi-layer perceptron (generator) that can generate a weight coefficient w capable of calculating a latent variable z'' that cannot be distinguished from Gaussian noise, that is, a latent variable z'' for which it is difficult for the discriminator to discriminate, can be generated by machine learning.
[0052] (Regarding the image generation unit 102) The image generation unit 102 generates an image from a high-dimensional latent variable x using a learned generation model. Here, as the generation model, a diffusion model is utilized, and an image is generated by applying an inverse diffusion process to the high-dimensional latent variable x. As a method for the image generation unit 102 to generate an image, the technique of Non-Patent Document 1 can be applied.
[0053] In this embodiment, for example, LDM (Latent Diffusion Model) is used as the diffusion model. At this time, considering the processing by the search space setting unit 100, the image generation unit 102 needs to generate an image using differentiable calculations.
[0054] The reverse diffusion process in a general diffusion model is configured using a neural network, and it is possible to calculate the derivative in the reverse diffusion process. Furthermore, when using LDM, in the process of generating an image, a decoder that restores the latent variable from the latent space to the image space is used. In LDM, since this decoder is also composed of a neural network, it is possible to calculate the derivative in the process of signal processing by the decoder. Note that the image generated using LDM may be a 3-channel RGB image or a 1-channel grayscale image. Here, let the number of dimensions of the image f generated by the image generation unit 102 be N dimensions.
[0055] (Regarding the search space setting unit 100) The search space setting unit 100 generates a set Y = {y} of a plurality of low-dimensional latent variables y from the low-dimensional latent variable y. Here, the latent variable y is the optimal latent variable in the current iteration step. The set Y may or may not include the latent variable y. best from, a set Y = {y n} of a plurality of low-dimensional latent variables y. Here, the latent variable y best is the optimal latent variable in the current iteration step. The set Y may include the latent variable y best or may not include the latent variable y best .
[0056] At this time, depending on the method of setting the search space, a search space may be set in which there is almost no change in the generated image in the set search space.
[0057] In particular, in image generation processing, it is known that in the latent space, changes in latent variables in a few specific directions greatly change the generated image. On the other hand, it is known that changes in latent variables in most specific directions hardly change the generated image. If a search space with few such changes is set, it is known that it takes time to generate an optimal image and the convergence tends to be slow.
[0058] In contrast, in the present embodiment, the search space setting unit 100 sets a set of variables along a straight line passing through the latent variable y in the low-dimensional latent space best as the search space. Here, the set of variables is, for example, a set of variables located at equal intervals on a straight line passing through the latent variable y best . The direction of this straight line is set in the direction in which the absolute value of the derivative with respect to the image (generated image) generated by the image generation unit 102 becomes large in the latent variable y best . The direction in which the absolute value of the derivative is large means that when the latent variable is changed in that direction, the change in the generated image becomes large. That is, by setting the search space along the direction in which the absolute value of the derivative is large, the change in the generated image in the search space becomes large and the variations of the generated image become rich. Therefore, it becomes possible to select a generated image from among rich variations, and the time required to generate an optimal image can be shortened and the convergence becomes fast.
[0059] The search space setting unit 100 sets, as the direction in which the absolute value of the derivative with respect to the generated image becomes large, the eigenvector corresponding to the element with a large singular value in the singular value decomposition of the Jacobian matrix as the direction of the search space. As a method for selecting an eigenvector, for example, a method of using the eigenvector corresponding to the element with the largest singular value can be adopted. Also, as another selection method, a method of selecting the corresponding eigenvector by a random number with the magnitude of the singular value as the generation probability can also be used.
[0060] Here, the calculation of the Jacobian matrix will be explained in detail. Let the low-dimensional latent variable be y, the transformation from the low-dimensional latent variable to the high-dimensional latent variable be z(y), and the process of generating an image from the high-dimensional latent variable be f(z). The partial derivative of the latent variable y with respect to the image f is represented by the Jacobian matrix J shown in Equation (5).
[0061]
Number
[0062] Let each element of the image f be f i and each element of the high-dimensional latent variable z be z j and each element of the low-dimensional latent variable y be y k Then, the Jacobian matrix J can be expressed by Equation (6).
[0063]
Number
[0064] Furthermore, let the Jacobian matrix with respect to the latent variable z in the image f be J fz and the Jacobian matrix with respect to the latent variable y in the latent variable z be J zy Then, the Jacobian matrix J can be expressed by Equation (7).
[0065]
Number
[0066] The search space setting unit 100 performs singular value decomposition on the Jacobian matrix J and sets the eigenvector corresponding to the element with a large singular value as the direction of the search space. That is, the search space setting unit 100 needs to calculate the Jacobian matrix J fz and the Jacobian matrix J zy in order to set the direction of the search space. In other words, the conditions for setting an appropriate search space are that the transformation process of the latent variable corresponding to the Jacobian matrix J zy is differentiable, and the Jacobian matrix J fzis equivalent to the process of generating the corresponding image being differentiable.
[0067] For the transformation z(y) from a low-dimensional latent variable to a high-dimensional latent variable to be differentiable means that the partial derivative (∂z / ∂y) is computable. That is, it is equivalent to each element (∂z j / ∂y k ) of the partial derivative (∂z / ∂y) being computable. Similarly, for the variable f(z) representing the process of generating an image from a high-dimensional latent variable to be differentiable means that the partial derivative (∂f / ∂z) is computable. That is, it is equivalent to each element (∂f i / ∂z j ) of the partial derivative (∂f / ∂z) being computable.
[0068] Here, it is explained that the process in the aforementioned latent variable transformation unit 101 enables each element (∂z j / ∂y k ) of the partial derivative (∂z / ∂y) to be computable. The transformation from a K-dimensional latent variable y to an M-dimensional latent variable z is shown in equation (8). Equation (8) is the same as equation (2).
[0069]
Equation
[0070] Here, let the elements of the Gaussian noise x k be x k、j . At this time, each element z j of the high-dimensional latent variable z is represented by the following equation (9).
[0071]
Equation
[0072] At this time, each element (∂z j / ∂y k ) of the partial derivative (∂z / ∂y) becomes the following equation (10).
[0073]
Equation
[0074] Thus, the conversion from the K-dimensional latent variable y to the M-dimensional latent variable z can be calculated for all combinations of all elements z of the latent variable z j and all elements y of the latent variable y k . Therefore, it can be said that the process by the aforementioned latent variable conversion unit 101, that is, the process of converting the K-dimensional latent variable y into the M-dimensional latent variable z, is differentiable.
[0075] Here, the case of calculating the M-dimensional latent variable z′ using the following equation (11) will be described. Equation (11) is the same as equation (3).
[0076] [Number]
[0077] (11) Equation is composed only of arithmetic operations on the latent variable z, that is, multiplication, subtraction, and addition. Also, the calculation of the average value indicated by mean(z) in equation (11) is also composed of arithmetic operations. All calculation processes consisting only of combinations of arithmetic operations are differentiable. Therefore, also in equation (11), for all combinations of all elements z of the latent variable z′ j and all elements y of the latent variable y k , the partial derivative (∂z j / ∂y k ) can be calculated, and it can be said that the process by the aforementioned latent variable conversion unit 101, that is, the process of converting the K-dimensional latent variable y into the M-dimensional latent variable z, is differentiable.
[0078] Here, a case will be described where the latent variable conversion unit 101 calculates the low-dimensional latent variable y using a pre-trained multi-layer perceptron as the weight coefficient of the weighted sum shown in Equation (2). This multi-layer perceptron updates the parameters of the learning model using the partial derivative of the output variable with respect to the input variable during learning. The calculation of this partial derivative is called error backpropagation. That is, when calculating the weight coefficient from the latent variable using a learnable multi-layer perceptron, all the calculation processes for obtaining the weight coefficient from the latent variable are differentiable. Differentiation is calculated in the process of performing error backpropagation to update the parameters of the learning model of the multi-layer perceptron.
[0079] Here, it will be explained that the partial derivative (∂f / ∂z) can be calculated in the process of generating an image by the image generation unit 102. As described above, the image generation process in the image generation unit 102 is composed of a pre-trained LDM. Here, the LDM updates the parameters of the learning model using the partial derivative of the output variable with respect to the input variable during learning. The calculation of this partial derivative is called error backpropagation. That is, all the calculation processes from the latent variable to the image performed in the process of generating an image from the latent variable using a learnable LDM are differentiable. Differentiation is calculated in the process of performing error backpropagation to update the parameters of the learning model of the LDM.
[0080] (Regarding the image presentation unit 103) The image presentation unit 103 n performs conversion to the high-dimensional latent variable z by the latent variable conversion unit 101 for each element y in the search space Y, generates an image using the latent variable z by the image generation unit 102, and generates a set F of a plurality of images. The set F of a plurality of images is shown by Equation (12).
[0081]
Equation
[0082] The image presentation unit 103 presents each image included in the set F of generated images to the user. The user visually recognizes the presented images and selects the image that the user feels is the most desirable from among the presented images. As a method of presenting images, the image presentation unit 103 may display all the images included in the set F of images on a display so that they can be selected, and cause any one of the displayed images to be selected by an operation such as clicking by the user. Also, as another image presentation method, the image presentation unit 103 may display a slider bar. In this case, each image included in the set F of images is associated with each memory of the slider bar, and the image presentation unit 103 displays, on the display, the image corresponding to each memory according to the slide operation of the slider bar by the user. The user slides the slider bar to visually recognize each image and adjusts the slider bar so that the image that the user feels is the most desirable is displayed.
[0083] (Regarding the latent variable update unit 104) The latent variable update unit 104 updates a latent variable (the first latent variable). The latent variable update unit 104 uses the low-dimensional latent variable y corresponding to the image selected by the user from among the images presented by the image presentation unit 103 next as the optimal latent variable in the next iteration step, and updates the latent variable y best . Here, the latent variable y next , and the latent variable y best are examples of the first variable. The equation for updating the latent variable y best is shown in equation (13).
[0084]
Equation
[0085] The image generation apparatus 1 of the present invention includes the setting of the search space by the search space setting unit 100, and the optimal latent variable y by the image presentation unit 103 and the latent variable update unit 104 bestBy repeatedly processing the settings, the user obtains the generated image that the user feels is the most desirable. Here, the latent variable y best The initial value of is generated as a random number. Also, the entire repetitive process may be repeated until the user is satisfied, or may be repeated until a preset upper limit number of times is reached.
[0086] (Regarding the processing flow performed by the image generation device 1) Here, the processing flow performed by the image generation device 1 will be described with reference to FIG. 7. FIG. 7 is a sequence diagram showing the processing flow performed by the image generation device 1 of the embodiment. Step S301: The search space setting unit 100 sets a search space, and outputs a plurality of latent variables (second latent variables) existing in the set search space to the latent variable conversion unit 101 via the second latent variable storage unit 107 or the like. Step S302: The latent variable conversion unit 101 calculates a high-dimensional latent variable (third latent variable) from each of the plurality of latent variables (second latent variables) acquired from the search space setting unit 100, and outputs the calculated third latent variable to the image generation unit 102 via the third latent variable storage unit 108 or the like. Step S303: The image generation unit 102 generates an image (generated image) from each of the plurality of latent variables (third latent variables) acquired from the latent variable conversion unit 101, and outputs the generated plurality of generated images to the image presentation unit 103 via the generated image storage unit 109 or the like. Step S304: The image presentation unit 103 displays the plurality of generated images acquired from the image generation unit 102 on the screen and presents them to the user. Step S305: The user selects one generated image from among the generated images presented by the image presentation unit 103. Identification information of the image (selected image) selected by the user is output to the image presentation unit 103 via an input device such as a mouse or a keyboard. Step S306: The image presentation unit 103 outputs the identification information of the selected image selected by the user to the latent variable update unit 104. Step S307: The latent variable update unit 104 updates the first latent variable based on the selected image notified from the image presentation unit 103. Step S308: When the user feels that the selected image selected from the images presented by the image presentation unit 103 is the most desirable and no further search is desired, the user performs an operation indicating the end of the process. A signal indicating the end of the process operated by the user is output to the image presentation unit 103 via an input device such as a mouse or a keyboard. Thereby, the image generation device 1 ends the process. Note that in step S308, when the operation indicating the end of the process by the user is not performed, the image generation device 1 returns to step S301 and repeatedly executes the processes shown in steps S301 to S308.
[0087] As described above, the image generation device 1 of the embodiment is a device that generates images. The image generation device 1 includes a search space setting unit 100, a latent variable conversion unit 101, an image generation unit 102, an image presentation unit 103, and a latent variable update unit 104. The search space setting unit 100 generates a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable, and sets a space including the plurality of generated second latent variables as a search space. The latent variable conversion unit 101 converts the second latent variable into a third latent variable having a larger number of dimensions than the second latent variable. The image generation unit 102 generates a generated image by executing an inverse diffusion process in the diffusion model using the third latent variable. The image presentation unit 103 presents a plurality of generated images generated by the image generation unit 102 to the user using the third latent variable. The third latent variable is a latent variable of a higher dimension than the second latent variable, which is converted by the latent variable conversion unit 101 corresponding to the plurality of second latent variables included in the search space. The latent variable update unit 104 re-sets the second latent variable corresponding to the selected image as the first latent variable. The selected image is an image selected by the user from the plurality of generated images presented by the image presentation unit 103. The latent variable conversion unit 101 converts the second latent variable into the third latent variable using an arithmetic process composed of differentiable operations. The search space setting unit 100 calculates the derivative with respect to the generated image generated by the arithmetic process (arithmetic process composed of differentiable operations) and the inverse diffusion process, and generates a plurality of second latent variables along the direction in which the value of the derivative increases in the latent space corresponding to the first latent variable. Accordingly, in the image generation device 1 of the embodiment, an image can be generated using a diffusion model with a relatively large number of dimensions of latent variables, and optimization can be executed in which selection and generation by HITL optimization are repeated for the generated image. That is, for the generative AI based on a diffusion model with a relatively large number of dimensions of latent variables, generative AI control by HITL optimization can be performed.
[0088] Also, in the image generation device 1 of the embodiment, the latent variable conversion unit 101 acquires a plurality of preset fourth latent variables for the latent variable x (a fourth latent variable having the same number of dimensions as the third latent variable). The latent variable conversion unit 101 sets each element of the second latent variable as a weight coefficient, weights the plurality of fourth latent variables, and sets the weighted sum as the third latent variable. The latent variable conversion unit 101 calculates the third latent variable using, for example, equation (2). Accordingly, in the image generation device 1 of the embodiment, the third latent variable z can be calculated by a relatively simple and differentiable operation of calculating a weighted sum.
[0089] Also, in the image generation device 1 of the embodiment, the latent variable conversion unit 101 sets, as the third latent variable, a value calculated so as to amplify the variation in the value of the weighted sum as the variation in the value of the weight coefficient becomes smaller for the weighted sum. The latent variable conversion unit 101 calculates the third latent variable using, for example, equation (3). Accordingly, in the image generation device 1 of the embodiment, even when the variation in the third latent variable z calculated using equation (2) is small, the third latent variable z' with an increased variation can be calculated.
[0090] Also, in the image generation device 1 of the embodiment, the latent variable conversion unit 101 calculates the weight coefficient w for the second latent variable through a machine learning model learned in advance. The latent variable conversion unit 101 reads a plurality of fourth latent variables x set in advance for the fourth latent variable x having the same dimension as the third latent variable z, and uses the weight coefficient w to calculate the weighted sum of the plurality of fourth latent variables x as the third latent variable z. The latent variable conversion unit 101 calculates the third latent variable z′′ using, for example, equation (4). Thus, in the image generation device 1 of the embodiment, even when the dimension number K of the latent variable y used as the weighting coefficient is different from the size K′ of the set X, the weighting coefficient w having the same dimension number as the size of the set X can be calculated based on the latent variable y.
[0091] Also, in the image generation device 1 of the embodiment, the plurality of fourth latent variables are each generated as different Gaussian noises. Thus, in the image generation device 1 of the embodiment, a high-dimensional latent variable following Gaussian noise can be calculated.
[0092] With the above configuration, even if the latent variable used in the process of generating an image by the image generation unit 102 of the image generation device 1 of the present invention is high-dimensional, generation AI control by HITL optimization is possible. Thus, even when generating an image by a diffusion model using an existing learned model, it is possible to obtain the most desirable generated image by HITL optimization.
Explanation of Signs
[0093] 1... Image generation device 100... Search space setting unit 101... Latent variable conversion unit 102... Image generation unit 103... Image presentation unit 104... Latent variable update unit 105... User selection reception unit 106... First latent variable storage unit 107... Second latent variable storage unit 108... Third latent variable storage unit 109... Generated image storage unit
Claims
1. An image generation device for generating an image, comprising: A search space setting unit that generates a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable, and sets a space including the plurality of generated second latent variables as a search space; A latent variable conversion unit that converts the second latent variable into a third latent variable having a number of dimensions larger than that of the second latent variable; An image generation unit that generates a generated image by executing an inverse diffusion process in a diffusion model using the third latent variable; An image presentation unit that presents a plurality of the generated images generated by the image generation unit using the third latent variable converted by the latent variable conversion unit to a user corresponding to the plurality of second latent variables included in the search space; A latent variable update unit that re-sets the second latent variable corresponding to the selected image selected by the user from the plurality of generated images presented by the image presentation unit as the first latent variable; and the latent variable conversion unit converts the second latent variable into the third latent variable using an arithmetic process composed of differentiable operations, the search space setting unit calculates a derivative with respect to the generated image generated by the arithmetic process and the inverse diffusion process, and generates a plurality of the second latent variables along a direction in which the value of the derivative increases in the latent space corresponding to the first latent variable; An image generation device.
2. The latent variable conversion unit acquires a plurality of preset fourth latent variables having the same number of dimensions as the third latent variable, and uses each element of the second latent variable as a weight coefficient to set a weighted sum of the plurality of fourth latent variables as the third latent variable; The image generation device according to claim 1.
3. The latent variable conversion unit calculates a value for the weighted sum such that the smaller the variation in the values of the weight coefficients with respect to the weighted sum, the greater the amplification of the variation in the value of the weighted sum, and sets the calculated value as the third latent variable. The image generation apparatus according to claim 2.
4. The latent variable conversion unit calculates weight coefficients for the second latent variable through a machine learning model that has been pre-learned, reads a plurality of the fourth latent variables that have been set in advance for the fourth latent variable having the same dimension as the third latent variable, and sets the weighted sum of the plurality of the fourth latent variables as the third latent variable using the weight coefficients. The image generation apparatus according to claim 1.
5. The plurality of fourth latent variables are each generated as different Gaussian noises. The image generation apparatus according to any one of claims 2 to 4.
6. An image generation method performed by a computer that is an image generation apparatus for generating an image, A search space setting step of generating a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable and setting a space including the generated plurality of second latent variables as a search space is executed. A latent variable conversion step of converting the second latent variable into a third latent variable having a larger number of dimensions than the second latent variable is executed. An image generation step of generating a generated image by executing a reverse diffusion process in a diffusion model using the third latent variable is executed. Corresponding to the plurality of second latent variables included in the search space, a plurality of the generated images generated in the image generation step using the third latent variable converted in the latent variable conversion step are presented to the user. The second latent variable corresponding to the selected image selected by the user from the plurality of presented generated images is reset as the first latent variable. In the latent variable conversion step, the second latent variable is converted into the third latent variable using arithmetic processing composed of differentiable operations. In the search space setting step, the arithmetic processing and the derivative with respect to the generated image generated by the reverse diffusion process are calculated, and a plurality of second latent variables are generated along the direction in which the value of the derivative increases in the latent space corresponding to the first latent variable. Image generation method. Claim 7 Causing a computer, which is an image generation device for generating an image, to execute a search space setting step of generating a plurality of second latent variables having the same number of dimensions as the first latent variable from the first latent variable and setting a space including the plurality of generated second latent variables as a search space. to execute a latent variable conversion step of converting the second latent variable into a third latent variable having a larger number of dimensions than the second latent variable. to execute an image generation step of generating a generated image by executing a reverse diffusion process in a diffusion model using the third latent variable. to cause a user to be presented with a plurality of the generated images generated by the image generation step using the third latent variable converted by the latent variable conversion step corresponding to the plurality of second latent variables included in the search space. to cause the second latent variable corresponding to the selected image selected by the user from the plurality of presented generated images to be reset as the first latent variable. In the latent variable conversion step, causing the second latent variable to be converted into the third latent variable using arithmetic processing composed of differentiable operations. In the search space setting step, calculating the arithmetic processing and the derivative with respect to the generated image generated by the reverse diffusion process, and generating a plurality of second latent variables along the direction in which the value of the derivative increases in the latent space corresponding to the first latent variable. Program.