Facial image editing and enhancement using personalized prior distribution
By identifying and utilizing personalized prior distributions in the latent vector space, the method enhances the generative model's ability to create realistic images of a specific subject, addressing the inconsistency issue in existing models and improving image editing and enhancement tasks.
Patent Information
- Application Number
- JP2025040140
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2025-03-13
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2042-01-10
Smart Images

Figure 2025106276000001_ABST
Abstract
Description
Background Art
[0001] By using a generative model, tasks can be performed that range from editing and enhancing an image of a given subject to generating a realistic image (or a portion of an image) of a given subject or a synthetically generated subject. To adequately train such a model, generally, a large number of sets of images of a large number of subjects are required. As a result, when using a generative model to edit, enhance, or fill in a portion of an image of a known subject, an image that looks realistic but resembles a different subject may be created.
Summary of the Invention
[0002] The present technology relates to a system and method for identifying a personalized prior distribution within the latent vector space of a generative model based on a set of images of a given subject. In some aspects, the present technology uses a personalized prior distribution (e.g., a convex hull defined by a set of codes generated based on a set of images of a subject) and may further include restricting the codes input into the generative model so that the identification features of the subject are reflected in the images created by the model. For example, the generative model may be configured to enhance or fill in facial features in an image of a subject where cues related to the identification of the subject are only partially present (e.g., due to motion blur, low illumination, low resolution, occlusion by other subjects). Without the present technology, the model may be able to successfully enhance or fill in such images, but may do so by creating an image that appears to be a different subject. By focusing the generative model using the present technology, the images created by the generative model may become more consistent with the appearance of the subject.
[0003] In one aspect, the present disclosure is a method executed by a computer, comprising: (1) for each given image in a set of images of a subject, using one or more processors of a processing system to test a plurality of codes to identify an optimization code for the given image, including: (a) for each code of the plurality of codes, using a generation model and the code to generate a first image, and using one or more processors to compare the first image with the given image to generate a first loss value for the code; (b) using one or more processors to compare the first loss values generated for each code of the plurality of codes, and identifying the code having the minimum first loss value as the optimization code for the given image; and (2) using one or more processors to generate a personalized prior distribution for the subject based on a convex hull including each optimization code identified for each given image in the set of images of the subject. In some aspects, the above method further comprises: (1) for each optimization code identified for each given image in the set of images of the subject, using the generation model and the optimization code to generate a second image, and using one or more processors to compare the second image with the given image to generate a second loss value; and (2) using one or more processors to modify one or more parameters of the generation model based at least in part on each generated second loss value to construct an adjusted generation model. In some aspects, the above method comprises: (1) using one or more processors to identify a plurality of coefficient sets, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull; (2) for each given coefficient set of the plurality of coefficient sets, using one or more processors to generate a third image using the adjusted generation model and a given code corresponding to the given coefficient set, and using one or more processors to compare the third image with at least a part of an input image of the subject to generate a third loss value for the third image; (3) using one or more processors to, for each generated Further including comparing third loss values and identifying a third image having the minimum third loss value as the personalized output image. In some aspects, the method described above includes: (1) using one or more processors to identify a plurality of coefficient sets, where each coefficient set of the plurality of coefficient sets corresponds to a code within a convex hull; (2) using one or more processors to identify a plurality of code sets, where each code set of the plurality of code sets includes two or more individual codes, and each individual code corresponds to a coefficient set among the plurality of coefficient sets; (3) for each given code set of the plurality of code sets, using one or more processors to generate a third image using an adjusted generation model and the given code set, where each individual code of the given code set is provided to a different layer or set of layers of the adjusted generation model; using one or more processors to compare the third image with at least a portion of an input image of a subject to generate a third loss value for the third image; and (4) using one or more processors to compare the third loss values generated for each third image and identify a third image having the minimum third loss value as the personalized output image. In some aspects, the method described above includes: (1) using one or more processors to identify a plurality of coefficient sets, where each coefficient set of the plurality of coefficient sets corresponds to a code within a convex hull; (2) for each given coefficient set of the plurality of coefficient sets, using one or more processors to generate a third image using a generation model and a given code corresponding to the given coefficient set; using one or more processors to compare the third image with at least a portion of an input image of a subject to generate a third loss value for the third image; and (3) using one or more processors to compare the third loss values generated for each third image and identify a third image having the minimum third loss value as the personalized output image.In some aspects, the above method further includes: (1) using one or more processors to identify a plurality of coefficient sets, where each coefficient set of the plurality of coefficient sets corresponds to a code within the convex hull, identifying the plurality of coefficient sets; (2) using one or more processors to identify a plurality of code sets, where each code set of the plurality of code sets includes two or more individual codes, and each individual code corresponds to a coefficient set among the plurality of coefficient sets, identifying the plurality of code sets; (3) for each given code set of the plurality of code sets, using one or more processors to generate a third image using a generative model and the given code set, where each individual code of the given code set is provided to different layers or sets of layers of the generative model, generating the third image; using one or more processors to compare the third image with at least a part of an input image of the subject to generate a third loss value for the third image; (4) using one or more processors to compare the third loss values generated for each third image and identify the third image having the minimum third loss value as the personalized output image. In some aspects, the plurality of coefficient sets includes a first coefficient set and a plurality of sets of coefficients selected using gradient descent directly or indirectly based on the first coefficient set. In some aspects, the input image of the subject includes a first portion of pixels saved from the original image of the subject and a mask in place of a second portion of pixels from the original image of the subject, and using one or more processors to compare the third image with at least a part of the input image of the subject to generate a third loss value for the third image includes comparing the third image with the first portion of pixels to generate a third loss value for the third image. In some aspects, the input image has a first resolution and the personalized output image has a second resolution higher than the first resolution. In some aspects, the plurality of codes includes a first code and a plurality of sets of codes selected using gradient descent directly or indirectly based on the first code.In some aspects, the first code represents the mean value of the latent vector space W, where the latent vector space W represents all possible codes that can be input into the generative model.
[0004] In another aspect, the present disclosure describes a processing system including a memory storing a generative model and one or more processors coupled to the memory and configured to execute any of the methods described above.
[0005] In another aspect, the present disclosure relates to a processing system comprising: (1) a memory storing a generative model; and (2) one or more processors coupled to the memory and configured to generate a personalized prior distribution for a subject for use with the generative model, the one or more processors including: (a) for each given image in a set of images of the subject, testing a plurality of codes to identify an optimization code for the given image, the testing including: (i) for each code of the plurality of codes, generating a first image using the generative model and the code, and generating a first loss value for the code by comparing the first image with the given image; and (ii) comparing the first loss values generated for each code of the plurality of codes and identifying the code having the minimum first loss value as the optimization code for the given image; and (b) generating a personalized prior distribution for the subject based on a convex hull including each optimization code identified for each given image in the set of images of the subject. In some aspects, the one or more processors are further configured to adjust the generative model, including: (1) for each optimization code identified for each given image in the set of images of the subject, generating a second image using the generative model and the optimization code, and generating a second loss value by comparing the second image with the given image using the one or more processors; and (2) using the one or more processors to modify one or more parameters of the generative model based at least in part on each generated second loss value to construct an adjusted generative model.In some embodiments, one or more processors are further configured to generate a personalized output image based on an input image of a subject, including: (1) identifying a plurality of sets of coefficients, each set of coefficients of the plurality of sets of coefficients corresponding to a code within a convex hull; (2) for each given set of coefficients of the plurality of sets of coefficients, generating a third image using an adjusted generation model and a given code corresponding to the given set of coefficients, and generating a third loss value for the third image by comparing the third image with at least a portion of the input image of the subject; and (3) comparing the third loss values generated for each third image and identifying the third image having the minimum third loss value as the personalized output image. In some embodiments, one or more processors are further configured to generate a personalized output image based on an input image of a subject, including: (1) identifying a plurality of sets of coefficients, each set of coefficients of the plurality of sets of coefficients corresponding to a code within a convex hull; (2) identifying a plurality of sets of codes, each set of codes of the plurality of sets of codes including two or more individual codes, each individual code corresponding to a set of coefficients of the plurality of sets of coefficients; (3) for each given set of codes of the plurality of sets of codes, generating a third image using an adjusted generation model and the given set of codes, where each individual code of the given set of codes is provided to a different layer or set of layers of the adjusted generation model, and generating a third loss value for the third image by comparing the third image with at least a portion of the input image of the subject; and (4) comparing the third loss values generated for each third image and identifying the third image having the minimum third loss value as the personalized output image.In some aspects, one or more processors are further configured to generate a personalized output image based on an input image of a subject, (1) identifying a plurality of sets of coefficients, wherein each set of coefficients of the plurality of sets of coefficients corresponds to a code within a convex hull, (2) for each given set of coefficients of the plurality of sets of coefficients, using a generation model and a given code corresponding to the given set of coefficients to generate a third image, comparing the third image with at least a part of the input image of the subject to generate a third loss value for the third image, and (3) comparing the third loss values generated for each third image to find the most... Identifying a third image having a small third loss value as a personalized output image. In some embodiments, one or more processors are further configured to generate a personalized output image based on an input image of a subject, including: (1) identifying a plurality of coefficient sets, where each coefficient set of the plurality of coefficient sets corresponds to a code within a convex hull; (2) identifying a plurality of code sets, where each code set of the plurality of code sets includes two or more individual codes, and each individual code corresponds to a coefficient set among the plurality of coefficient sets; (3) for each given code set of the plurality of code sets, generating a third image using a generative model and the given code set, where each individual code of the given code set is provided to a different layer or set of layers of the generative model, generating the third image, and generating a third loss value for the third image by comparing the third image with at least a portion of the input image of the subject; and (4) comparing the third loss values generated for each third image and identifying the third image having the smallest third loss value as the personalized output image. In some embodiments, the plurality of coefficient sets includes a first coefficient set and a plurality of series of coefficient sets, and one or more processors are further configured to select each coefficient set of the plurality of series of coefficient sets using gradient descent based directly or indirectly on the first coefficient set. In some embodiments, the input image of the subject includes a first portion of pixels saved from the original image of the subject and a mask instead of a second portion of pixels from the original image of the subject, and generating a third loss value for the third image by comparing the third image with at least a portion of the input image of the subject includes generating a third loss value for the third image by comparing the third image with the first portion of pixels. In some embodiments, one or more processors are configured to generate a personalized output image based on an input image of a subject, the input image has a first resolution, and the personalized output image has a second resolution higher than the first resolution.In some aspects, the plurality of codes includes a first code and a plurality of sequences of codes, and one or more processors are further configured to select each code of the plurality of sequences of codes using gradient descent based directly or indirectly on the first code. In some aspects, one or more processors are further configured to select a first code representing an average value of the latent vector space W, where the latent vector space W represents all possible codes that can be input into the generative model.
Brief Description of the Drawings
[0006]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Figure 9
Figure 10
Figure 11
DETAILED DESCRIPTION OF THE INVENTION
[0007] Next, the present technology will be described with respect to the following exemplary systems and methods. Common reference numerals between the figures illustrated and described below are intended to identify the same features.
[0008] Exemplary System FIG. 1 shows a high-level system diagram 100 of an exemplary processing system 102 for executing the methods described herein. The processing system 102 may include one or more processors 104, as well as a memory 106 that stores instructions 108 and data 110. The instructions 108 and data 110 may include a generative model (e.g., the generative models 306 of FIGS. 3A, 3B, 4A, 4B, 6, 10, and 11) described herein. Additionally, the data 110 may store training examples used for training such a generative model, data used by the generative model when generating images, a set of images of a given subject, a personalized prior distribution based on a set of images of a given subject (e.g., the personalized prior distributions 402 and 403 of FIGS. 4A, 4B, 6, 10, and 11), and / or images output by the generative model.
[0009] The processing system 102 may be present on a single computing device. For example, the processing system 102 may be a server, a personal computer, or a mobile device, and thus, the generative model and related data may be local to that single computing device. Similarly, the processing system 102 may be present on a cloud computing system or other distributed system. In such cases, the generative model and / or related data may be distributed across two or more different physical computing devices. For example, in some aspects of the present technology, the processing system may include a first computing device that stores the generative model, as well as data used by the generative model when generating an image, a set of images of a given subject, a personalized prior distribution based on the set of images of the given subject, and / or an image output by the generative model, and a second computing device that stores the image. Similarly, in some aspects of the present technology, the processing system may include a first computing device that stores layers 1 to n of a generative model having m layers, and a second computing device that stores layers n to m of the generative model.
[0010] In this regard, further, FIG. 2 shows a high-level system diagram 200. In this figure, the exemplary processing system 102 described above communicates with various websites and / or remote storage systems, including websites 210 and 218 and remote storage system 226, via one or more networks 208. In this example, each of websites 210 and 218 includes one or more servers 212a-212n and 220a-220n, respectively. Each of servers 212a-212n and 220a-220n may have one or more processors (e.g., 214 and 222), as well as associated memory (e.g., 216 and 224) that stores instructions and data, including the content of one or more web pages. Similarly, although not shown, remote storage system 226 may also include one or more processors and memory for storing instructions and data. In some aspects of the present technology, the processing system 102 is configured to obtain data, training examples, a set of images of a given subject, and / or an input image of a given subject from one or more of website 210, website 218, and / or remote storage system 226, as provided to the generative model for training or adjustment of the generative model and / or for use when generating an image.
[0011] The processing system described herein may be implemented on any type of computing device(s), such as any type of general-purpose computing device, server, or set thereof, and may further include other components typically present in a general-purpose computing device or server. Similarly, the memory of such a processing system may be of any non-transitory type capable of storing information accessible by the processor(s) of the processing system. For example, the memory may include non-transitory media such as hard drives, memory cards, optical disks, solid state, tape memory, and the like. A computing device suitable for the roles described herein may include the different combinations described above, whereby different portions of instructions and data are stored on different types of media.
[0012] In any case, the computing device described herein may further include any other components typically used in connection with a computing device, such as a user interface subsystem. The user interface subsystem may include one or more user inputs (e.g., mouse, keyboard, touch screen, and / or microphone), as well as one or more electronic displays (e.g., a monitor with a screen, or any other electrical device operable to display information). Output devices other than electronic displays, such as speakers, lights, and vibration elements, pulse elements, or tactile elements, may also be included in the computing device described herein.
[0013] One or more processors included in each computing device may be any conventional processor, such as a commercially available central processing unit (“CPU”), graphics processing unit (“GPU”), tensor processing unit (“TPU”), etc. Alternatively, the one or more processors may be a dedicated device such as an ASIC or other hardware-based processor. Each processor may have a plurality of cores that can operate in parallel. The processor(s), memory, and other elements of a single computing device may be housed within a single physical housing or may be distributed among two or more housings. Similarly, the memory of a computing device may include a hard drive or other storage medium located within a housing different from that of the processor(s), such as within an external database or a networked storage device. Thus, references to a processor or computing device are to be understood to include a collection of processors or computing devices or memories that may or may not operate in parallel, and references to one or more servers of a load-balanced server farm or cloud-based system.
[0014] The computing devices described herein may store instructions (such as machine code) directly executable by a processor(s) or instructions (such as scripts) indirectly executable by a processor(s). The computing device may also store data, which may be retrieved, stored, or modified by one or more processors according to the instructions. The instructions may be stored as computing device code on a computing device-readable medium. In that regard, the terms “instructions” and “program” may be used interchangeably herein. The instructions may also be stored in object code form for direct processing by a processor(s) or in any other computing device language including scripts or collections of independent source code modules that may be interpreted on demand or pre-compiled. Exa As such, the programming language may be C#, C++, JAVA (registered trademark), or another computer programming language. Similarly, any component of the instructions or program may be implemented in a computer scripting language such as JavaScript (registered trademark), PHP, ASP, or any other computer scripting language. Further, any one of these components may be implemented using a combination of a computer programming language and a computer scripting language.
[0015] Exemplary method FIGS. 3A and 3B show how images (308, 309) of different subjects can result from different codes within the latent vector space W(302) of the generative model 306, in accordance with aspects of the present disclosure. In these examples, the generative model 306 may be any suitable model configured to generate an output image based on an input code, and the latent vector space W(302) represents the range of all possible input codes that can be provided to the generative model 306. In that regard, for purposes of simply illustration, the latent vector space W(302) is shown as a two-dimensional space in FIGS. 3A and 3B (as well as FIGS. 4A, 4B, 6, 10, and 11). However, the present technology may be applied to any suitable generative model (e.g., a GAN or a bidirectional GAN (“BiGAN”)) configured to generate an image based on an input vector of any suitable number of dimensions. For example, in some aspects of the present technology, the generative model 306 may be an adversarial generative network (e.g., StyleGAN, StarGAN) configured to use a latent vector space W having 512 dimensions.
[0016] As shown in FIG. 3A, the code at point 304 in the latent vector space W(302) is supplied to the generative model 306, thereby creating an image 308. To illustrate the effect of the present technique, in the examples shown in FIGS. 3A, 3B, 4A, 4B, 10, and 11, images of subjects whose faces can be recognized by many people are used respectively. Therefore, in the example of FIG. 3A, it is assumed that an image 308 of Barack Obama (the 44th President of the United States) is created by the code at point 304. Similarly, in the example of FIG. 3B, it is assumed that an image 309 of Lady Gaga (an American singer-songwriter and actress, also known as Stefani Joanne Angelina Germanotta) is created by the code at point 305.
[0017] FIGS. 4A and 4B show two different personalized prior distributions 402, 403 within the latent vector space W(302) of the generative model 306 according to aspects of the present disclosure, and further show how different images of a given subject can result from points (e.g., 404 for 304, 405 for 305) within a given personalized prior distribution. The examples of FIGS. 4A and 4B both assume the use of the same latent vector space W(302) and the same generative model 306 as used in FIGS. 3A and 3B, but further illustrate the personalized prior distributions 402, 403 for individuals within the latent vector space W(302).
[0018] The personalized prior distributions 402 and 403 each represent vector spaces within the latent vector space W(302) that contain subsets of input codes capable of creating images similar to a given subject. In this case, the personalized prior distribution 402 is assumed to represent the range of codes that create images similar to Barack Obama when supplied to the generative model 306. Therefore, the personalized prior distribution 402 includes the point 304 that represents the code that created the image 308 in FIG. 3A, and another point 404 that represents the code that created a different image 406 of Barack Obama.
[0019] Similarly, the per - person prior distribution 403 is assumed to represent a range of codes that create an image similar to Lady Gaga when provided to the generative model 306. Thus, the per - person prior distribution 403 includes a point 304 representing the code to create the image 309 in FIG. 3B, and another point 405 representing the code to create a different image 407 of Lady Gaga. including the code to create a different image 407 of Lady Gaga.
[0020] Here too, for the purpose of simply simplifying the explanation, the per - person prior distributions 402 and 403 are shown as being two - dimensional spaces in FIGS. 4A and 4B (as well as FIGS. 6, 10, and 11). However, since the present technology can be applied to any suitable generative model configured to generate an image based on an input vector of any suitable number of dimensions, the per - person prior distributions 402 and 403 may similarly be vector spaces of any number of dimensions less than or equal to the number of dimensions of the latent vector space W(302). Thus, for example, if the generative model 306 is an adversarial generative network configured to use a latent vector space W having 512 dimensions, the per - person prior distributions 402 and 403 may each be vector spaces of up to 512 dimensions. Additionally, for the sake of clarity in the explanation, the exemplary vector spaces 402 and 403 are shown in a completely separated state. However, in reality, the per - person prior distributions for different subjects may intersect with each other.
[0021] FIG. 5 shows an exemplary method 500 for generating a personalized prior distribution based on a set of images of a subject, in accordance with an aspect of the present disclosure.
[0022] In step 502, the processing system (e.g., processing system 102) selects a given image from a set of images of the subject. In general, it will be understood that a larger number of images will result in a more typical prior distribution for an individual as compared to a smaller number of images. In that regard, it has generally been found that with 100 to 200 images, images can be created by the generative model that are realistic and appear to match the identity of the subject. However, other aspects can also affect how well the prior distribution for an individual represents the appearance of the subject. For example, if the appearance of the subject has changed significantly (e.g., a change in hairstyle, hair color, addition or removal of facial hair, or over a long period of time), it may be useful to tailor the set of images to a particular aspect such that the prior distribution for the individual reflects a single "look" and the images created by the generative model match that look. Similarly, for the same reason, it may be useful to limit the set of images to a single stage of life (e.g., infancy, childhood, adolescence, adulthood, etc.). On the other hand, there may be cases where the diversity of the set of images is important. For example, a set of images showing the subject's face only from the front may not be as useful as a set of images showing the subject's face from a variety of different angles, under a variety of different lighting conditions, etc. Generating a prior distribution for an individual from images that are too similar can overly limit the generative model and cause the generative model to create images that resemble the subject but do not necessarily look realistic.
[0023] As shown in step 504, after the processing system selects a given image, it repeatedly executes steps 506 - 514 to test a plurality of codes in order to identify the optimization code for the given image. In that regard, in step 506, the processing system identifies the code to be tested in the current path. This code can be identified based on any suitable selection criterion. For example, in some aspects of the present technology, the processing system blindly selects the first code (e.g., using a random selection process or a pre - selected value such as the average value of the latent vector space W), and then, using an appropriate optimization regime, is configured to select each series of codes directly or indirectly based on that first code (in each series of paths passing through step 506). Thus, in some aspects, the processing system can be configured to select each series of codes using gradient descent based on the previous code and an evaluation of how well the image generated based on the previous code matches the given image (e.g., the first loss value generated in the most recent path passing through step 510).
[0024] In step 508, the processing system uses the generation model (e.g., generation model 306) and ( the code to be tested identified in this path passing through step 506) to generate a first image. The generation model can be configured to create the first image in any suitable manner, including those described above with respect to FIGS. 3A, 3B, 4A, and 4B.
[0025] In step 510, the processing system compares the first image (generated in this path passing through step 508) with the given image (selected in step 502) to generate a first loss value for the code under test. The first loss value can be generated in any suitable manner using any suitable function. For example, in some aspects of the present technology, the first loss value may be based on a comparison between the first image and the given image using a heuristic or learned similarity metric (such as the learned perceptual image patch similarity (LPIPS), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM), L1 or L2 loss, etc.).
[0026] In step 512, the processing system determines whether to test another code. Since it is assumed that there are multiple codes to be tested, when the processing system first reaches step 512, it automatically returns to step 506 following the "yes" arrow. Thereby, a second code is identified and tested. However, whenever it returns to step 512 thereafter, the processing system can determine whether to test another code based on any suitable criterion. Thus, as described above, in some aspects of the present technology, the processing system can determine when to stop testing another code based on a suitable optimization regime such as gradient descent. In such a case, the determination in step 512 may be based on a comparison between the first loss value generated in the current path passing through step 510 (or some other evaluation of how well the first image matches the given image) and one or more of the first loss values generated in the preceding paths. For example, the processing system can be configured to stop testing a series of codes when the first loss value generated in the current path passing through step 510 is greater than or equal to the first loss value generated in a previous path passing through step 510.
[0027] Therefore, the processing system repeats steps 506 to 512 using each of the following codes until it is determined in step 512 that sufficient code has been tested. When it is determined that sufficient code has been tested, the processing system proceeds to step 514 following the "No" arrow, where it compares the first loss values generated (in step 510) for each of the plurality of codes and identifies the code having the minimum first loss value. That code for which the first loss value is minimum is selected as the optimization code for the given image.
[0028] Next, in step 516, the processing system determines whether there is a further image within the set of images of the subject. If there is such an image, the processing system proceeds to step 518 following the "Yes" arrow. There, the processing system selects the next given image to be tested. The processing system then returns to step 504 and tests the plurality of codes to identify the optimization code for this new given image. In this way, steps 504 to 518 are repeated as described above until an optimization code is identified for every image within the set of images of the subject. When the optimization code is selected (in step 514) for the last image within the set of images of the subject, the processing system determines in step 516 that there are no more images in the set and thus proceeds to step 520 following the "No" arrow.
[0029] In step 520, the processing system generates a personalized prior distribution for the subject based on the convex hull containing each of the optimization codes identified (in step 514) for each given image of the set of images of the subject. In that regard, in some aspects of the present technology, the personalized prior distribution may simply be the convex hull defined by each of the optimization codes identified in step 514 and in such a case, the n optimization codes {x1, x2, …, x nAssuming there exists a set of}, the personalized prior distribution includes any code c generated by linearly combining the optimization codes according to the following equations 1 - 3. In the equations, each of the coefficients (alpha values α1 to α n ) is non - negative, and the sum of all the coefficients is 1.
[0030]
Number
[0031] Similarly, in some aspects of the present technology, the personalized prior distribution may correspond to a subset of the convex hull defined by each optimization code identified in step 514, for example, a set of a predetermined number (e.g., 100, 500, 1000, 10000, 100000, 1000000) of codes or sets of coefficient sets corresponding to a set of sampled points within the convex hull. Further, in some aspects of the present technology, the personalized prior distribution may be a simpler hull (e.g., a hull with fewer vertices) that lies within or substantially overlaps the actual convex hull defined by each optimization code identified in step 514. Additionally, in some aspects of the present technology, the personalized prior distribution may be based on the convex hull defined by the above - mentioned equations 1 - 3 by including a wider set of codes than that defined by the above - mentioned equations 1 - 3. For example, the personalized prior distribution may include any code c generated by the above - mentioned equations 1 and 3 where the coefficients (α values α1 to α n ) are greater than or equal to a certain negative value (e.g., - 0.01, - 0.05, - 0.1).
[0032] Figure 6 shows how a given set of n images 602a - 612a can be used to generate the personalized prior distribution 402 of FIG. 4A. In that regard, each image 602a - 612a is assumed to be one of a set of n images of a subject, and the lines shown indicate how these six selected images correspond to different optimization codes 602b - 612b within the latent vector space W(302) of the generative model (e.g., generative model 306). These optimization codes 602b - 612b can each be found by the steps 502 - 518 of the exemplary method of FIG. 5, as described above. Additionally, the exemplary diagram 600 of FIG. 6 shows how each of these optimization codes 602b - 612b (along with the optimization codes for the remainder of the set of n images not shown) can be used to define the convex hull that forms the basis of the personalized prior distribution 402, as further described above with respect to step 520 of FIG. 5. Again, for simplicity of explanation, the personalized prior distribution 403 of FIG. 6 is shown as being in a two - dimensional space, and only six sample images out of the set of n images are shown. However, the personalized prior distribution 403 may be based on any suitable number of points and may be a vector space of any suitable dimensionality less than or equal to the dimensionality of the latent vector space W(302), as described above.
[0033] Figure 7 shows an exemplary method 700 for adjusting a generative model following the identification of an optimization code for each image within a set of images by the method of FIG. 5, according to an aspect of the present disclosure. In that regard, the exemplary method 700 represents a process that can be optionally executed after an optimization code has been identified ( at step 514 of FIG. 5) for each image within the set of images.
[0034] Accordingly, in step 702, it is assumed that the processing system (e.g., processing system 102) executes at least steps 502 to 518 of the exemplary method of FIG. 5 for each image in the set of images. The exemplary method of FIG. 7 shows that steps 704 to 720 are performed after step 702, but it will be understood that steps 704 to 720 can be executed in any suitable order with respect to steps 502 to 518 of the exemplary method of FIG. 5. For example, in some aspects of the present technology, the processing system identifies the optimization code for a given image in the set of images (at step 514 of FIG. 5), and then, for that given image, tests a plurality of codes (e.g., steps 504 to 514 of FIG. 5) to identify the optimization code before selecting the next given image (at step 518 of FIG. 5) or for the next given image, and may be configured to execute steps 704 to 712 in parallel. In such a case, the processing system may further be configured to continuously identify the optimization code (in a series of passes through step 518) while periodically updating the parameters of the generation model by executing step 714 in parallel with steps 504 to 518 of FIG. 5 after each batch of images has been processed by steps 704 to 712.
[0035] Regardless of timing, in step 704, the processing system selects a given image from the set of images of the subject. This set of images may be the entire set of images used in FIG. 5 or any suitable subset thereof. For example, in some aspects of the present technology, the processing system may be configured to identify the optimization code for a set of 200 images of the subject by steps 502 to 518 of FIG. 5, but may then be configured to adjust the generation model (e.g., generation model 306) based on only 100 of those images.
[0036] In step 706, the processing system generates a second image using the generation model and the optimization code identified for a given image (in step 514 of FIG. 5). Here too, the generation model can be configured to create the second image in any suitable manner, including those described above with respect to FIGS. 3A, 3B, 4A, and 4B. Additionally, although step 706 refers to a "second image", it will be understood that this second image may, in some cases, be a copy of one of the "first images" generated in step 508 of FIG. 5. Thus, in some aspects of the present technique, for any "given image" within the first batch of images, the processing system may be configured to use the "first image" (or a copy thereof) associated with the optimization code identified in step 514 of FIG. 5 as the "second image" in step 706, rather than regenerating the image in step 706. However, if the processing system modifies one or more parameters of the generation model at step 714 (described below), the generation model is changed such that for each subsequent optimization code, a "second image" is created that is different from the "second image" that the code was previously thought to generate. Thus, if method 700 is executed after method 500 (not in parallel with some or all of the steps of method 500), the processing system may be configured to generate a new "second image" in each sequence of passes through step 706 (rather than using a copy of the "first image" identified in step 508 of FIG. 5).
[0037] In step 708, the processing system compares the second image (generated in step 706) with the given image (selected in step 704) to generate a second loss value. Here too, the second loss value can be generated in any suitable manner using any suitable function. For example, in some aspects of the present technique, the second loss value uses a heuristic or learned similarity metric (e.g., learned perceptual image patch similarity ("LPIPS"), peak signal-to-noise ratio ("PSNR"), structural similarity index ("SSIM")) It may be based on a comparison between the second image and a given image. Additionally, although step 708 refers to a "second loss value", it will be understood that this second loss value may, in some cases, be a copy of the "minimum first loss value" identified in step 514 of FIG. 5 for that image. Thus, in some aspects of the present technology, for any "given image" within the first batch of images, the processing system may be configured to use the "minimum first loss value" identified in step 514 of FIG. 5 for that given image as the "second loss value" in step 708, rather than newly generating a loss value in step 708. However, as already described, when the processing system modifies one or more parameters of the generation model in step 714 (described later), the generation model is changed, and in each subsequent optimization cycle, a "second image" different from the "second image" that was considered to have been generated by that cycle before is created, and thus a different "second loss value" is created when compared to the "given image". Therefore, if method 700 is executed after method 500 (not in parallel with some or all of the steps of method 500), the processing system may be configured to generate a new "second loss value" in each sequence of passes through step 708 (rather than using a copy of the "minimum first loss value" identified in step 514 of FIG. 5).
[0038] In step 710, the processing system determines whether there are additional images within the batch. In that regard, the set of images may be retained as a whole or divided into any suitable number of batches. If the set of images is not divided and thus there is a single "batch" that includes all the images within the "set of images" of the subject, the processing system proceeds to step 712 following the "Yes" arrow, selects the next given image from the set of images of the subject, and repeats steps 706 - 710 for that newly selected image. This process is repeated until there are no more images within the set of images. When there are no more images, the processing system proceeds to step 714 following the "No" arrow. On the other hand, if the set of images is divided into two or more batches (e.g., a set of 200 images can be divided into two batches of 100 images each, four batches of 50 images each, ten batches of 20 images each, 200 single - image "batches", etc.), steps 704 - 712 are repeated for each image until the end of the batch is reached.
[0039] As shown in step 714, after a "second loss value" has been generated (in step 708) for every image within the batch, the processing system modifies one or more parameters of the generation model based at least in part on each of the generated second loss values. The processing system can be configured to modify one or more parameters in any suitable manner and at any suitable interval based on these generated second loss values. Thus, in some aspects of the present technology, each "batch" can include a single image such that the processing system performs a backpropagation step of modifying one or more parameters of the generation model each time a second loss value is generated. Similarly, if each "batch" includes two or more images, the processing system can sum (e.g., by totaling or averaging a plurality of second loss values) each of the "second loss values" generated (in step 708) for each image in that batch to obtain a total loss value and is configured to modify one or more parameters of the generation model based on that total loss value.
[0040] In step 716, the processing system determines whether there is a further batch within the set of images of the subject. If the set of images is not split, and thus there is a single "batch" that includes all the images within the "set of images" of the subject, the determination in step 716 automatically becomes "no", and then method 700 ends as shown in step 720. However, if the set of images is split into two or more batches, the processing system proceeds to step 718 following the "yes" arrow and selects the next given image from the set of images of the subject. Thereby, for each image within the next batch of images, a separate set of paths passing through steps 706 - 714 is subsequently started, and the process continues until no more batches exist. When no more batches exist, the processing system proceeds to step 720 following the "no" arrow. Method 700 is shown to end at step 720 when all images are used to adjust the generative model, but it will be understood that method 700 can be repeated any suitable number of times using the same set of images until an image is created that is close enough to each given image by the output of the generative model for each optimization code. In that regard, in some aspects of the present technology, the processing system may be configured to aggregate all of the second loss values generated during a given pass through method 700 and, based on the aggregated loss value, determine whether to repeat method 700 for the set of images. For example, in some aspects of the present technology, the processing system may be configured to repeat method 700 for the set of images if the total loss value for the most recent pass through method 700 is greater than a certain threshold value. Similarly, in some aspects, the processing system may make this determination using gradient descent, and thus may be configured to repeat method 700 for the set of images until the total loss value in a given pass through method 700 is greater than or equal to the total loss value from the previous pass.
[0041]
[0042] FIG. 8 shows an exemplary method 800 for generating a personalized output image based on an input image and a personalized prior distribution generated by the method of FIG. 5 or FIG. 7, according to an aspect of the present disclosure. In that regard, the exemplary method 800 may be optionally executed after at least a personalized prior distribution has been generated (at step 520 of FIG. 5), and may also be executed after the generation model has been further adjusted by the method of FIG. 7. Thus, at step 802, it is assumed that a processing system (e.g., processing system 102) executes at least the method 500 of FIG. 5 for each image in a set of images, and optionally executes steps 704-720 of FIG. 7.
[0043] As described above, after the processing system generates a personalized prior distribution (using the method 500 of FIG. 5) and optionally adjusts the generation model (using the method 700 of FIG. 7), it may generate different sets of candidate coefficients using the personalized prior distribution, and then generate candidate images for a given image enhancement task using the codes corresponding to the sets of candidate coefficients. Thus, as shown in step 804, for a particular input image of a subject, the processing system repeatedly executes steps 806-812 to test multiple sets of coefficients in order to identify a personalized output image. Here, each set of coefficients of the multiple sets of coefficients corresponds to a code within the convex hull (identified at step 520 of FIG. 5).
[0044] In each path passing through step 806, the processing system identifies a set of coefficients from among a plurality of sets of coefficients and uses it to generate a given code. This set of coefficients can be identified based on any suitable selection criterion. For example, in some aspects of the present technology, the processing system blindly selects a first set of coefficients (e.g., using a random selection process or a preselected value such as the mean of a vector space represented by a personalized prior distribution), and then uses an appropriate optimization regime to select each series of coefficient sets (in each series of paths passing through step 806) directly or indirectly based on that first set of coefficients. Thus, in some aspects, the processing system may be configured to select each series of coefficient sets using gradient descent based on the previous set of coefficients and an evaluation of how well the image generated based on the previous set of coefficients matches the input image (e.g., the third loss value generated in the most recent path passing through step 810).
[0045] At step 808, the processing system uses a generative model (e.g., generative model 306 or a tuned generative model obtained from one or more paths passing through steps 704 - 720 of FIG. 7) and a given code (generated in this path passing through step 806) to generate a third image. Again, the generative model can be configured to create the third image in any suitable manner, including those described above with respect to FIGS. 3A, 3B, 4A, and 4B.
[0046] In step 810, the processing system compares the third image (generated in this path passing through step 808) with the input image of the subject to generate a third loss value for the third image. The third loss value can be generated in any suitable manner using any suitable function. For example, in some aspects of the present technology, the third loss value may be based on a comparison between the third image and the input image using a heuristic or learned similarity metric (e.g., learned perceptual image patch similarity (LPIPS), peak signal-to-noise ratio (PSNR), structural similarity index (SSIM)).
[0047] In step 812, the processing system determines whether another set of coefficients should be tested. Since multiple sets of coefficients are assumed to exist to be tested, when the processing system first reaches step 812, it automatically returns to step 806 following the "yes" arrow. Thereby, the second set of coefficients is identified and tested. However, whenever it returns to step 812 thereafter, the processing system may determine whether to test another set of coefficients based on any suitable criteria. Thus, as described above, in some aspects of the present technology, the processing system may determine when to stop testing another set of coefficients based on a suitable optimization regime such as gradient descent. In such a case, the determination in step 812 may be based on a comparison between the third loss value generated in the current path passing through step 810 (or some other evaluation of how well the third image matches the input image) and one or more of the third loss values generated in the preceding paths. For example, the processing system may be configured to stop testing a series of sets of coefficients when the third loss value generated in the current path passing through step 810 is greater than or equal to the third loss value generated in the previous path passing through step 810.
[0048] Accordingly, the processing system repeats steps 806 - 812 using each successive set of coefficients until it is determined in step 812 that a sufficient set of coefficients has been tested. When it is determined that a sufficient set of coefficients has been tested, the processing system proceeds to step 814 following the "No" arrow. There, the processing system compares the third loss values generated (in step 810) for each third image and identifies the third image having the minimum third loss value. That third image with the minimum third loss value is used as the personalized output image. In this way, a personalized output image can be created using method 800. This personalized output image is generated using the code within the personalized prior distribution (and thus is likely to be more similar to the subject compared to the output image if the code were not so limited), and is optimized to closely match the input image (and thus ensures that the image created by the model remains consistent with the image that could have been obtained from the input image when performing image editing or enhancement tasks).
[0049] FIG. 9 shows another exemplary method 900 for generating a personalized output image based on an input image and a personalized prior distribution generated by the method of FIG. 5 or FIG. 7, in accordance with aspects of the present disclosure. In that regard, exemplary method 900 also represents a process that can optionally be performed after at least the personalized prior distribution has been generated (in step 520 of FIG. 5) and can also be performed after the generation model has been further adjusted by the method of FIG. 7. The only difference between method 800 of FIG. 8 and method 900 of FIG. 9 is that a given set of codes is tested in each pass through steps 906 - 912. Here, each individual code within the given set of codes is within the convex hull identified in step 520 of FIG. 5. Accordingly, the processing system repeats steps 806 - 812 using each successive set of coefficients until it is determined in step 812 that a sufficient set of coefficients has been tested. When it is determined that a sufficient set of coefficients has been tested, the processing system proceeds to step 814 following the "No" arrow. There, the processing system compares the third loss values generated (in step 810) for each third image and identifies the third image having the minimum third loss value. That third image with the minimum third loss value is used as the personalized output image. In this way, a personalized output image can be created using method 800. This personalized output image is generated using the code within the personalized prior distribution (and thus is likely to be more similar to the subject compared to the output image if the code were not so limited), and is optimized to closely match the input image (and thus ensures that the image created by the model remains consistent with the image that could have been obtained from the input image when performing image editing or enhancement tasks).
[0050] Therefore, as described above, it is assumed that in step 902, the processing system (e.g., processing system 102) executes at least method 500 of FIG. 5 for each image in the set of images and optionally executes steps 704-720 of FIG. 7. Similarly, as shown in step 904, for a particular input image of a subject, the processing system repeatedly executes steps 906-912 to test a plurality of code sets in order to identify a personalized output image. Here, each code of the code set is within the convex hull (identified in step 520 of FIG. 5).
[0051] In each pass through step 906, the processing system identifies a given code set that includes two or more individual codes. Again, this given code set can be identified based on any suitable selection criterion. For example, in some aspects of the technology, the processing system blindly selects a first given code set (e.g., using a random selection process or by assigning preselected values such as the mean of a vector space represented by a personalized prior distribution to each individual code), and then uses an appropriate optimization regime to select each series of code sets (in each series of passes through step 906) directly or indirectly based on that first given code set. Thus, in some aspects, the processing system can be configured to select each series of code sets using gradient descent based on the preceding code set and an evaluation of how well the image generated based on the preceding code set matches the input image (e.g., a third loss value generated in the most recent pass through step 910).
[0052] In step 908, the processing system generates a third image using a generation model (e.g., generation model 306, or an adjusted generation model obtained from one or more paths passing through steps 704-720 of FIG. 7) and a given code set (generated in this path passing through step 906). In that regard, in step 908, the processing system provides each individual code of the given code set to different layers or sets of layers of the generation model. Here too, the generation model can be configured to create the third image in any suitable manner, including those described above with respect to FIGS. 3A, 3B, 4A, and 4B.
[0053] In step 910, the processing system compares the third image (generated in this path passing through step 908) with the input image of the subject to generate a third loss value for the third image. Here too, the third loss value can be generated in any suitable manner using any suitable function. For example, in some aspects of the present technology, the third loss value may be based on a comparison of the third image and the input image using a heuristic or learned similarity metric (e.g., learned perceptual image patch similarity ("LPIPS"), peak signal-to-noise ratio ("PSNR"), structural similarity index ("SSIM")).
[0054] In step 912, the processing system determines whether another code set should be tested. Since multiple code sets are assumed to be tested, when the processing system first reaches step 912, it automatically returns to step 906 following the "yes" arrow. Thereby, a second code set is identified and tested. However, whenever it returns to step 912 thereafter, the processing system can determine whether to test another code set based on any suitable criteria. Thus, as described above In some aspects of the present technology, the processing system may determine when to stop testing another code set based on a suitable optimization regime such as gradient descent. In such a case, the determination in step 912 may be based on a comparison between the third loss value generated in the current pass through step 910 (or some other evaluation of how well the third image matches the input image) and one or more of the third loss values generated in the previous pass. For example, the processing system may be configured to stop testing a series of code sets when the third loss value generated in the current pass through step 910 is greater than or equal to the third loss value generated in a previous pass through step 910.
[0055] Thus, similar to method 800, the processing system repeats steps 906 - 912 using each of the next code sets until it is determined in step 912 that a sufficient code set has been tested. When it is determined that a sufficient code set has been tested, the processing system proceeds to step 914 following the "no" arrow. There, the processing system compares the third loss values generated for each third image (in step 910) to identify the third image having the minimum third loss value. Again, the third image identified as having the minimum third loss value is used as the personalized output image.
[0056] FIG. 10 is a diagram 1000 showing a comparative illustration of how a generative model (e.g., generative model 306) can complete four exemplary image enhancement tasks according to aspects of the present disclosure, with and without using the per - person prior distribution 402 of FIG. 4A.
[0057] In FIG. 10, a column of four input images 1002a - 1002d is shown in the center of diagram 1000. Since the generative model is assumed to be tasked with creating output images that resemble the input images but have improved resolution, each of the input images 1002a - 1002d is blurry.
[0058] The images 1004a - 1004d in the left column show the potential outputs of the generative model when not limited to selecting codes within any specific part of the latent vector space W(302). In that regard, the dashed lines connect each input image to the point in the latent vector space W(302) representing the code that the generative model ultimately selects (e.g., after a selection process such as the processes described above in method 800 of FIG. 8 or method 900 of FIG. 9), and the arrows connect that point to the corresponding output image created by the generative model based on that code. As can be seen, although the output images 1004a - 1004d are each visually consistent with their respective input images 1002a - 1002d, these output images do not all appear to be of the same subject.
[0059] In contrast, the images 1006a - 1006d in the right column show the potential outputs of the generative model when limited to selecting codes within the per - person prior distribution 402 of FIG. 4A. Again, the dashed lines connect each input image to the point in the per - person prior distribution 402 representing the code that the generative model ultimately selects (e.g., after a selection process such as the processes described above with respect to method 800 of FIG. 8 or method 900 of FIG. 9), and the arrows connect that point to the corresponding output image created by the generative model based on that code. As can be seen, this ultimately results in the creation of output images 1006a - 1006d. These output images are visually consistent with their respective input images 1002a - 1002d and all appear to be of the same subject. Specifically, since the personalized prior distribution 402 of FIG. 4A represents the range of codes that create images similar to Barack Obama when provided to the generative model 306 (as described above), each of the output images 1006a - 1006d appears to show an image of Barack Obama that is visually consistent with the input images 1002a - 1002d.
[0060] FIG. 11 is a diagram 1100 showing a comparative explanation of how a generative model (e.g., generative model 306) can complete four exemplary image inpainting tasks with and without using the individual prior distribution 403 of FIG. 4B according to an aspect of the present disclosure.
[0061] Here too, a column of four input images 1102a - 1102d is shown in the center of FIG. 1000. Since the generative model is assumed to be tasked with creating an output image that resembles the input image but fills in the masked parts, each of the input images 1102a - 1102d includes a black mask 1103a - 1103d representing the pixels to be replaced.
[0062] In FIG. 1100, the images 1104a - 1104d in the left column show the potential outputs of the generative model when not restricted to selecting codes within any specific portion of the latent vector space W(302). Here too, the dashed lines connect each input image to the point in the latent vector space W(302) representing the code that the generative model ultimately selects (after a selection process such as the process described above with respect to method 800 of FIG. 8 or method 900 of FIG. 9), and the arrows connect that point to the corresponding output image created by the generative model based on that code. As can be seen, although the output images 1104a - 1104d visually align with the unmasked portions of the respective input images 1102a - 1102d, these output images do not all appear to be of the same subject.
[0063] In contrast, the images 1106a - 1106d in the right column show the potential outputs of the generative model when limited to selecting codes within the personalized prior distribution 403 for individuals in Figure 4B. Here too, the dashed lines connect each input image to the points within the personalized prior distribution 403 that represent the codes ultimately selected by the generative model (after a selection process such as the processes described above with respect to Method 800 of Figure 8 or Method 900 of Figure 9), and the arrows connect that point to the corresponding output image created by the generative model based on that code. As can be seen, the output images 1106a - 1106d are ultimately created thereby. These output images are visually consistent with each input image 1102a - 1102d and appear to be of the same subject. Specifically, since the personalized prior distribution 403 in Figure 4B represents the range of codes that create images similar to Lady Gaga when provided to the generative model 306 (as described above), each of the output images 1106a - 1106d appears to show an image of Lady Gaga that is visually consistent with the unmasked portions of the input images 1102a - 1102d.
[0064] Accordingly, both Figure 1000 and Figure 1100 show examples of how a personalized prior distribution can be used to enable the generative model to create output images that are visually consistent with both the input image and the identity of a specific subject by focusing on the codes used by the generative model. Thus, when the subject of the input image is already known, the personalized prior distribution may be selected and used so that the generative model is biased towards creating more typical and thus more appropriate output images.
[0065] Unless otherwise specified, the foregoing alternative examples are not mutually exclusive and may be implemented in various combinations to achieve their respective advantages. Since these and other variations and combinations of the features described above can be utilized without departing from the subject matter defined by the claims, the foregoing description of the exemplary systems and methods should be regarded as illustrative rather than limiting the subject matter defined by the claims. In addition, the provision of the examples described herein, and phrases expressed as "such as", "including", "comprising", etc., should not be construed as limiting the subject matter of the claims to specific examples; rather, the examples are intended to illustrate only some of the many possible embodiments. Further, the same reference numerals in different drawings can identify the same or similar elements.
Claims
1. A method performed by a computer, comprising: for each given image of a set of images of a subject, using one or more processors of a processing system to test a plurality of codes to identify an optimization code for the given image, wherein the testing comprises: for each code of the plurality of codes, generating a first image using a generative model and the code; using the one or more processors to compare the first image with the given image to generate a first loss value for the code; using the one or more processors to compare the first loss values generated for each code of the plurality of codes and identify the code having the minimum first loss value as the optimization code for the given image; the method further comprises: using the one or more processors to generate a personalized prior distribution for the subject based on a convex hull comprising each optimization code identified for each given image of the set of images of the subject.
2. For each optimization code identified for each given image of the set of images of the subject, generating a second image using the generative model and the optimization code; using the one or more processors to compare the second image with the given image to generate a second loss value; using the one or more processors to modify one or more parameters of the generative model based at least in part on each generated second loss value to construct an adjusted generative model. The method according to claim 1, further comprising.
3. further comprising using the one or more processors to identify a plurality of coefficient sets, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, the method comprising: for each given coefficient set of the plurality of coefficient sets, using the one or more processors to generate a third image using the adjusted generative model and a given code corresponding to the given coefficient set; using the one or more processors to compare the third image with at least a portion of an input image of the subject to generate a third loss value for the third image; Using the one or more processors, comparing the third loss values generated for each of the third images to identify the third image having the minimum third loss value as the personalized output image, the method according to claim 2, further comprising.
4. Further comprising identifying a plurality of coefficient sets using the one or more processors, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, the method comprising: Further comprising identifying a plurality of code sets using the one or more processors, each code set of the plurality of code sets including two or more individual codes, each of the individual codes corresponding to a coefficient set among the plurality of coefficient sets, the method comprising: For each given code set of the plurality of code sets Further comprising generating a third image using the one or more processors, the adjusted generation model, and the given code set, each individual code of the given code set being provided to a different layer or set of layers of the adjusted generation model, the method comprising: For each given code set of the plurality of code sets Using the one or more processors, comparing the third image with at least A part of the input image of the subject to generate a third loss value for the third image; Using the one or more processors, comparing the third loss values generated for each of the third images to identify the third image having the minimum third loss value as the personalized output image, the method according to claim 2, further comprising.
5. Further comprising identifying a plurality of coefficient sets using the one or more processors, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, the method comprising: For each given coefficient set of the plurality of coefficient sets Using the one or more processors to generate a third image using the generation model and a given code corresponding to the given coefficient set; Using the one or more processors, comparing the third image with at least a part of the input image of the subject to generate a third loss value for the third image; Using the one or more processors, comparing the third loss values generated for each of the third images, and identifying the third image having the minimum third loss value as the personalized output image, the method according to claim 1, further comprising.
6. Further comprising identifying a plurality of coefficient sets using the one or more processors, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, the method comprising: Further comprising identifying a plurality of code sets using the one or more processors, each code set of the plurality of code sets including two or more individual codes, each said individual code corresponding to a coefficient set among the plurality of coefficient sets, the method comprising: For each given code set of the plurality of code sets, Further comprising generating a third image using the generation model and the given code set using the one or more processors, each individual code of the given code set being provided to a different layer or set of layers of the generation model, the method comprising: For each given code set of the plurality of code sets, Using the one or more processors, comparing the third image with at least a part of the input image of the subject to generate a third loss value for the third image; Using the one or more processors, comparing the third loss values generated for each of the third images, and identifying the third image having the minimum third loss value as the personalized output image, the method according to claim 1, further comprising.
7. The method according to claim 3 or claim 5, wherein the plurality of coefficient sets includes a first coefficient set and a plurality of sets of coefficients selected using gradient descent based directly or indirectly on the first coefficient set.
8. The input image of the subject includes a first portion of pixels saved from the original image of the subject and a mask in place of a second portion of pixels from the original image of the subject, Using the one or more processors, comparing the third image with at least a part of the input image of the subject to generate the third loss value for the third image includes comparing the third image with the first portion of the pixels to generate the third loss value for the third image, the method according to any one of claims 3 to 7.
9. The method according to any one of claims 3 to 7, wherein the input image has a first resolution, and the personalized output image has a second resolution higher than the first resolution.
10. The method according to any one of claims 1 to 9, wherein the plurality of codes includes a first code and a series of a plurality of codes selected using gradient descent directly or indirectly based on the first code.
11. The method according to claim 10, wherein the first code represents an average value of a latent vector space W, and the latent vector space W represents all possible codes that can be input to the generation model.
12. A memory storing a generation model; One or more processors coupled to the memory and configured to generate a personalized prior distribution for a subject for use with the generation model, the one or more processors For each given image of a set of images of the subject, testing a plurality of codes to identify an optimized code for the given image, the testing For each code of the plurality of codes, Generating a first image using the generation model and the code; Comparing the first image with the given image to generate a first loss value for the code; Comparing the first loss values generated for each code of the plurality of codes and identifying the code having the minimum first loss value as the optimized code for the given image, the testing including the above, and the one or more processors further A processing system that generates the personalized prior distribution for the subject based on a convex hull including each optimized code identified for each given image of the set of images of the subject.
13. The one or more processors are further configured to adjust the generation model, and the one or more processors For the optimized code identified for each given image of the set of images of the subject, Generating a second image using the generation model and the optimized code; Using the one or more processors to compare the second image with the given image to generate a second loss value; The processing system according to claim 12, wherein one or more parameters of the generation model are modified based at least in part on each of the generated second loss values using the one or more processors to construct an adjusted generation model.
14. The one or more processors are further configured to generate a personalized output image based on an input image of the subject, and the one or more processors identify a plurality of coefficient sets, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, and the one or more processors further for each given coefficient set of the plurality of coefficient sets generate a third image using the adjusted generation model and a given code corresponding to the given coefficient set, compare the third image with at least a portion of the input image of the subject to generate a third loss value for the third image, The processing system according to claim 13, wherein the third loss values generated for each of the third images are compared to identify the third image having the minimum third loss value as the personalized output image.
15. The one or more processors are further configured to generate a personalized output image based on an input image of the subject, and the one or more processors identify a plurality of coefficient sets, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, and the one or more processors further identify a plurality of code sets, each code set of the plurality of code sets including two or more individual codes, each of the individual codes corresponding to a coefficient set among the plurality of coefficient sets, and the one or more processors further for each given code set of the plurality of code sets generate a third image using the adjusted generation model and the given code set, each individual code of the given code set being provided to a different layer or set of layers of the adjusted generation model, and the one or more processors further for each given code set of the plurality of code sets compare the third image with at least a portion of the input image of the subject to generate a third loss value for the third image, The processing system according to claim 13, wherein the third loss value generated for each of the third images is compared, and the third image having the minimum third loss value is identified as a personalized output image.
16. The one or more processors are further configured to generate a personalized output image based on an input image of the subject, and the one or more processors identify a plurality of coefficient sets, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, and the one or more processors further for each given coefficient set of the plurality of coefficient sets generate a third image using the generation model and a given code corresponding to the given coefficient set, generate a third loss value for the third image by comparing the third image with at least a part of the input image of the subject, The processing system according to claim 12, wherein the third loss value generated for each of the third images is compared, and the third image having the minimum third loss value is identified as a personalized output image.
17. The one or more processors are further configured to generate a personalized output image based on an input image of the subject, and the one or more processors identify a plurality of coefficient sets, each coefficient set of the plurality of coefficient sets corresponding to a code within the convex hull, and the one or more processors further identify a plurality of code sets, each code set of the plurality of code sets including two or more individual codes, each of the individual codes corresponding to a coefficient set among the plurality of coefficient sets, and the one or more processors further for each given code set of the plurality of code sets generate a third image using the generation model and the given code set, each individual code of the given code set being provided to a different layer or set of layers of the generation model, and the one or more processors further for each given code set of the plurality of code sets generate a third loss value for the third image by comparing the third image with at least a part of the input image of the subject, The processing system according to claim 12, wherein the third loss value generated for each of the third images is compared, and the third image having the minimum third loss value is identified as a personalized output image.
18. The plurality of coefficient sets includes a first coefficient set and a plurality of sets of coefficients. The one or more processors are further configured to select each coefficient set of the plurality of sets of coefficients using gradient descent based directly or indirectly on the first coefficient set, the processing system according to claim 14 or claim 16.
19. The input image of the subject includes a first portion of pixels stored from the original image of the subject and a mask instead of a second portion of pixels from the original image of the subject. Generating the third loss value for the third image by comparing the third image with at least a part of the input image of the subject includes generating the third loss value for the third image by comparing the third image with the first portion of the pixels, the processing system according to any one of claims 14 to 18.
20. The one or more processors are configured to generate the personalized output image based on the input image of the subject, the input image has a first resolution, and the personalized output image has a second resolution higher than the first resolution, the processing system according to any one of claims 14 to 18.
21. The plurality of codes includes a first code and a plurality of sets of codes. The one or more processors are further configured to select each code of the plurality of sets of codes using gradient descent based directly or indirectly on the first code, the processing system according to any one of claims 12 to 20.
22. The one or more processors are further configured to select a first code representing an average value of the latent vector space W, the latent vector space W representing all possible codes that can be input to the generation model, the processing system according to claim 21.
23. A memory storing a generation model, and one or more processors coupled to the memory and configured to execute the method according to any one of claims 1 to 11. A processing system.
Citation Information
Patent Citations
Deep generation of user-customized items
US20210192594A1
Face image generation method and apparatus, device, and storage medium
US20210241521A1
Methods of generating personalized 3D head models or 3D body models
WO2017029488A2
Federated mixture models
WO2021247944A1