Synthetic image generation based on recognition images

By combining patch recognition and pose image generation techniques with control networks and diffusion models, the problem of visual feature leakage in personalized image generation is solved, generating more realistic images of multiple individual interactions and improving the accuracy and realism of the images.

CN121962329APending Publication Date: 2026-05-01FACE CUTE CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
FACE CUTE CO LTD
Filing Date
2025-10-28
Publication Date
2026-05-01

AI Technical Summary

Technical Problem

Existing attention-based personalized image generation methods are prone to visual feature leakage in synthetic images, especially when individuals are close to or interacting with each other, which leads to the incorrect mixing of visual features of specific individuals and fails to effectively preserve the real interaction between different individuals.

Method used

A computational system is employed to generate synthetic images by combining patch recognition and pose image generation techniques with a control network and a diffusion model, thus avoiding visual feature leakage and ensuring the authenticity of individual interactions. The system includes an ID extractor, a pose estimator, a patch encoder, a control network, and a diffusion model, utilizing stitched label embedding and noise reduction to generate synthetic images.

Benefits of technology

It effectively avoids the leakage of visual features, ensures that the interactions of individuals in the generated images appear realistic, and improves the visual accuracy and realism of the images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121962329A_ABST
    Figure CN121962329A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to composite image generation based on an identification image. A computing system receives an input cue and an input image, generates an identification image based on the input image, and generates identification tiles based on the identification image, respectively. The system further generates pose tile images based on the recognition tiles and the pose images, and generates word tags based on the recognition images, respectively. A tag embed is generated based on the input cue, and the word tag and the tag embed are spliced to generate a spliced tag embed. The system embeds and inputs the attitude tile images and stitched markers into a control network to generate features. The features, potential noise, and stitched indicia are then embedded into the diffusion model to generate a composite image, and an output is generated based on the composite image.
Need to check novelty before this filing date? Find Prior Art

Description

Background Technology

[0001] In the field of personalized image generation, creating visually coherent images that naturally integrate multiple concepts remains a challenging problem. One application involves generating images containing multiple distinct individuals that interact with each other in a realistic manner, each represented by multiple detected visual features derived from a reference photograph.

[0002] Current methods primarily rely on attention-based mechanisms, where the generation of visual representations of different individuals is controlled by masking attention maps at different stages of the generation process. While these techniques can ensure that different individuals are represented with a certain degree of accuracy within the same image, they are hampered by inherent limitations. Note that these mask-based methods are prone to visual feature leakage through convolutional layers, especially when two people in the synthesized image are very close or physically interacting. When this occurs, visual features associated with the first person who is very close to the second person in the image may be labeled and preserved through the convolutional layers as being associated with both the first and second people. During generation, the first person's visual features may be incorrectly represented in the masked region of the second person, causing the first person's visual features to leak into the generated image of the second person. As a concrete example, this could lead to the first person's hairstyle being incorrectly represented as the second person's hairstyle. This unintentional mixing of individual-specific visual features can result in visual outputs where different visual appearances are not well preserved, and the interactions of the individuals drawn in the image appear unrealistic. Summary of the Invention

[0003] In view of the above problems, a computational system for generating synthetic images is provided. The computational system includes processing circuitry and a memory storing instructions that, upon execution, cause the processing circuitry to receive input prompts and one or more input images, generate one or more recognition images based on the one or more input images, and generate one or more recognition patches based on the one or more recognition images. The system further generates pose patch images based on the one or more recognition patches and pose images, and generates one or more word tags based on the one or more recognition images. The tag embedding is generated based on the input prompts. One or more word tags and tag embeddings are concatenated to generate a concatenated tag embedding. The system inputs the pose patch images and the concatenated tag embeddings into a control network to generate features. Then, the features, latent noise, and the concatenated tag embeddings are input into a diffusion model to generate a synthetic image, and an output is generated based on the synthetic image.

[0004] This invention provides a simplified overview of a series of concepts, which are further described in the detailed description below. This invention is not intended to identify key or essential features of the claimed subject matter, nor is it intended to limit the scope of the claimed subject matter. Furthermore, the claimed subject matter is not limited to implementations that address any or all of the shortcomings mentioned in any part of this disclosure. Attached Figure Description

[0005] Figure 1 A schematic diagram of a first computing system according to an example of this disclosure is shown.

[0006] Figure 2 It shows Figure 1 A schematic diagram illustrating the operation of a trained machine learning diffusion model in a computing system.

[0007] Figure 3 It shows Figure 1 and Figure 2 Detailed schematic diagrams illustrating examples of the inputs and outputs of a tile encoder and a pose tile image generator based on a trained machine learning diffusion model.

[0008] Figure 4 It shows Figure 1 and Figure 2 A detailed schematic diagram of the control network of a trained machine learning diffusion model and an example of the inputs and outputs of the diffusion model.

[0009] Figure 5 A schematic diagram of a second computing system according to an example of this disclosure is shown.

[0010] Figure 6 This is a flowchart of a method for generating a synthetic image according to an example embodiment of the present disclosure.

[0011] Figure 7 An example computing environment of this disclosure is shown, in which the following can be implemented Figure 1 The first computing system or Figure 5 The second computing system. Detailed Implementation

[0012] Figure 1A schematic diagram of a first example computing system 10 is shown, which includes a computing device 100 for generating a synthetic image 130 using a trained machine learning diffusion model 128. The computing device 100 includes processing circuitry 102 (e.g., a central processing unit or “CPU”), volatile memory 104, non-volatile memory 106, input / output (I / O) module 108, a camera 110, and a display 112. The different components are operationally coupled to each other. The non-volatile memory 106 stores instructions for executing the trained machine learning diffusion model 128, which is configured to receive one or more input images 124 and input prompts 134, and to generate a synthetic image 130 of one or more individuals based at least on one or more input images 124 and input prompts 134. Although in this example, the first computing system 10 generates a synthetic image 130 comprising two individuals, it should be understood that the number of individuals depicted in the synthetic image 130 is not particularly limited. In alternative embodiments, the synthetic image 130 may depict only one individual or more than two individuals.

[0013] The trained machine learning diffusion model 128 includes a text encoder 136, an ID extractor 140, a pose estimator 144, a pose patch image generator 148, a patch encoder 150, a cue encoder 156, a stitching function, a control network 168, and a diffusion model 180. Typically, the diffusion model 180 has a latent diffusion model architecture, and the control network 168 is a neural network that takes an image as input to provide conditioning and guided image generation through the diffusion model 180. In a specific example, the diffusion model 180 may be a stable diffusion model, and the control network 168 may be a ControlNet used for stabilizing the diffusion model. The ID extractor 140 is configured to extract one or more identification images from one or more input images 124. The pose estimator 144 is configured to generate a pose image. The patch encoder 150 is configured to generate one or more identification patches based on one or more identification images, respectively. The pose patch image generator 148 is configured to generate a pose patch image based on the pose image and one or more identification patches. Text encoder 136 is configured to generate token embeddings based on input prompts. Prompt encoder 156 is configured to generate word tokens based on one or more recognized images, respectively. Concatenation function 162 is configured to concatenate the token embeddings and word tokens to generate a concatenated token embedding. Control network 168 receives the pose patch image 166 as input and the concatenated token embeddings to generate features input to diffusion model 180. The concatenated token embeddings are input to diffusion model 180 to guide the denoising process to generate a synthetic image 130 based on latent noise, and an output is generated based on the synthetic image 130. For example, synthetic image 130 can be output for rendering on display 112 and / or encoded by a video encoder to generate and output a video stream combining synthetic image 130. Synthetic image 130 can be published or shared on a social networking platform for other users of the social networking platform to view.

[0014] Figure 2 It shows Figure 1A detailed schematic diagram of the process of a trained machine learning diffusion model 128 is provided. This model is configured to receive input from one or more input images 124 and input cues 134, and to generate and output a synthetic image 130 based on the one or more input images 124 and input cues 134. The trained machine learning diffusion model 128 includes an ID extractor 140, configured to receive one or more input images 124 and generate one or more recognition images 142a, 142b, which can be collectively organized into a recognition image set 142. The recognition images 142a, 142b originate from one or more input images 124 and can take the form of cropped body features for each individual identified in the one or more input images 124. For example, each recognition image 142a, 142b can isolate and represent the face of an individual identified in the one or more input images 124.

[0015] The pose estimator 144 can be configured to generate a pose image 146 based on one or more input images 124 or reference images 125 depicting the poses of one or more individuals. The pose estimator 144 can identify one or more individuals present within one or more input images 124 or reference images 125 and determine their corresponding poses. The pose in the pose image 146 can be represented using a series of vectors connected by nodes, where each node corresponds to a key joint location, such as the shoulder, elbow, wrist, hip, knee, and ankle. The resulting pose image 146 is a pixelated image based on vector representations that depicts the simplified skeletal structure of one or more individuals, capturing the spatial arrangement and orientation of their body parts. Alternatively, the pose image 146 can be manually input by a user through manual annotation of one or more input images 124 or another image, or input by a motion capture system tracking the movements of individuals wearing specialized tracking devices (such as cameras and markers).

[0016] Turning Figure 3This document details the process by which a trained machine learning diffusion model 128 generates a pose patch image 166 from inputs of a pose image 146 and recognition images 142a and 142b. Recognition images 142a and 142b are input to a patch encoder 150 to generate corresponding recognition patches 152 and 154. In this example, the first recognition image 142a and the second recognition image 142b are cropped faces of individuals identified by an ID extractor 140 in one or more input images 124. The first recognition patch 152 corresponds to the first recognition image 142a, and the second recognition patch 154 corresponds to the second recognition image 142b. In the simplest embodiment, these recognition patches 152 and 154 may take the form of square patches. Each recognition patch 152 and 154 may encode a feature vector into pixel information, thereby utilizing the color channels of each pixel to store relevant data. In some alternative embodiments, recognition patches 152, 154 may not encode visual features; instead, recognition patches 152, 154 may represent integers or another form of non-visual data. The visual features presented in recognition patches 152, 154 may capture basic characteristics, such as facial features, from one or more input images 124.

[0017] The pose patch image generator 148 is configured to combine recognition patches 152 and 154 with the pose image 146 to generate a pose patch image 166, wherein recognition patches 152 and 154 are superimposed on the pose image 146. In this example, the first recognition patch 152 is superimposed on the head position of the individual on the left side of the pose image 146, and the second recognition patch 154 is superimposed on the head position of the individual on the right side of the pose image 146. The pose patch image generator 148 can use a combination of contextual information and predefined instructions to accurately locate the recognition patches 152 and 154 on the anatomical structures represented in the pose image 146. In one embodiment, the pose patch image generator 148 can process input prompts 134 specifying the target position of the patches, such as "place the first recognition image on the head position of the individual on the left side" and "place the second recognition image on the head position of the individual on the right side". The pose patch image generator 148 can use the instructions of the input prompt 134 to map each recognition patch 152, 154 to the corresponding position of an individual within the pose image 146.

[0018] Furthermore, the pose patch image generator 148 may include logic for determining anatomical locations such as the head, arms, torso, and legs within the pose image 146. The pose patch image generator 148 can use this logic to interpret pose vectors and nodes and to discern the spatial arrangement of different body parts. By analyzing the vectors and nodes defining each pose, the pose patch image generator 148 can identify specific anatomical regions, such as the head position based on the topmost node, or the torso position by identifying the center between the shoulder and hip nodes.

[0019] The pose patch image generator 148 can determine where to overlay the recognition patches 152, 154 in the pose image 146 based on a combination of structural analysis using logical and contextual input cues 134, to generate a pose patch image 166. For example, if the input cues 134 indicate that the recognition images 142a, 142b represent facial features of an individual, the pose patch image generator 148 can leverage its understanding of the pose structure in the pose image 146 to align the recognition patches 152, 154 with corresponding head positions in the pose image 146. For example, this alignment can be based at least on geometric center localization or scaling (e.g., adjusting the patch size to fit detected head boundaries). In the absence of the input cues 134, the pose patch image generator 148 can rely on contextual cues inferred from the pose image 146 itself, such as the relative positions of multiple individuals.

[0020] return Figure 2 Images 142a and 142b are input to a prompt encoder 156, which is configured to generate a set of corresponding word tags 158 and 160 based on the images 142a and 142b, respectively. The prompt encoder 156 can be configured as, for example, a CLIP (Contrastive Language Image Pretrained) text encoder. In this example, the first word tag 158 corresponds to the first recognized image 142a, and the second word tag 160 corresponds to the second recognized image 142b. The tag space of the prompt encoder 156 is not used to encode natural language descriptions of faces or other body parts. Instead, the tag space is used to map human-specific visual information from each recognized image 142a and 142b into the natural language tag space of the prompt encoder 156.

[0021] When the trained machine learning diffusion model 128 receives input prompt 134, the text encoder is configured to generate a token embedding 138 based on the input prompt 134, which may include a description of how to synthesize the final image 130. For example, the input prompt 134 may describe the arrangement of the recognized images 142a, 142b within the final image 130, such as "two individuals are shaking hands," "two individuals are in a glamorous ballroom," "place the first recognized image at the head position of the individual on the left," and / or "place the second recognized image at the head position of the individual on the right." The concatenation function 162 concatenates the word tokens 158, 160 generated by the prompt encoder 156 and the token embedding 138 generated by the text encoder 136 together to generate a concatenated token embedding 164. Therefore, the splicing function 162 stacks the marker features of the recognition images 142a and 142b captured by the word markers 158 and 160 and the prompt features of the input prompt 134 captured by the marker embedding 138 together, thereby integrating the marker features of the recognition images 142a and 142b and the prompt features of the input prompt 134 into a single embedding 164.

[0022] Turning Figure 4 The process of generating a final synthetic image 130 using a pose patch image 166 and a stitched labeled embedding 164 as input, performed by a trained machine learning diffusion model 128, is described in detail. The stitched labeled embedding 164 and the pose patch image 166 are used by a control network 168 to generate features 176. Latent noise 178, the generated features 176, and the stitched labeled embedding 164 are input into a diffusion model 180 to generate the final synthetic image 130. In this example, in the synthetic image 130, a first identification image 142a of a man is superimposed on the head position of the individual on the left, and a second identification image 142b of a woman is superimposed on the head position of the individual on the right. The poses of the individuals in the synthetic image 130, representing the two individuals standing and shaking hands, are arranged according to the poses depicted in the pose patch image 166. The ballroom setting of the synthetic image 130 is consistent with an input cue 134 that specifies the ballroom setting for the final synthetic image 130.

[0023] Back Figure 2 The architecture of the control network 168 and the diffusion model 180 is described in more detail. The diffusion model 180 is a pre-trained diffusion model that generates an image from latent noise 178 through an iterative denoising step, where the noise 178 is processed through a series of convolutional layers and attention mechanisms to progressively refine the image. These layers and mechanisms include an encoder 182 (which includes a first block), an intermediate block 184 (which includes a second block), and a decoder 186 (which includes a third block). The encoder 182 downsamples the latent noise 178, and the decoder 186 upsamples the latent representation back to the original resolution to generate the final image 130.

[0024] The diffusion model 180 uses a U-Net architecture, which processes noise during denoising through a series of ResNet blocks and attention layers in the encoder 182, intermediate block 184, and decoder 186, gradually refining the image to generate the final synthetic image 130. As the denoising process proceeds, the stitched label embedding 164 is input into the attention layers of the encoder 182, intermediate block 184, and / or decoder 186 of the diffusion model 180, so that the final synthetic image 130 reflects the label features of the recognized images 142a and 142b and the cue features of the input cue 134.

[0025] The control network 168 includes an encoder 170, which is a trainable copy of the encoder 182 of the diffusion model 180. The control network 168 also includes a zero-initialized convolutional layer 172 placed at the output of the encoder 170, and an intermediate block 174, which is a trainable copy of the intermediate block 184 of the diffusion model 180. A pose map patch image 166 is input to the encoder 170 of the control network 168. A concatenated label embedding 164 can be input to the attention layers of the encoder 170 and / or the intermediate block 174. The zero-initialized convolutional layer 172 is a 1×1 convolutional layer with both weights and biases introduced to zero. Before being injected into the diffusion model 180, the zero-initialized convolutional layer 172 transforms the features generated by the encoder 170 into features 176 or control signals from the control network 168. The features 176 output by the control network 168 are input to the skip connections and intermediate block 184 of the diffusion model 180. Skip connections are direct links connecting the encoder layer of encoder 182 to the corresponding decoder layer of decoder 186. Skip connections preserve spatial information that may have been lost during the downsampling process in encoder 182.

[0026] Figure 5 A schematic diagram of a second example computing system 20 is shown, which includes a computing device 200 for generating a synthetic image 230 using a trained machine learning diffusion model 228. The numbering of similar parts in this example is... Figure 1 The examples are similar and share their functionality, and for the sake of brevity, they will not be described again except as follows. The computing device 200 includes processing circuitry 202 (e.g., a central processing unit or "CPU"), volatile memory 204, non-volatile memory 206, input / output (I / O) modules 208, a camera 210, and a display 212. The different components are operationally coupled to each other. The non-volatile memory 206 stores instructions for executing the social media application 214.

[0027] Social media application 214 is configured to communicate via computer network 216 with a social networking platform 218 running on server computing system 220 of computing system 20. Social media application 214 includes a graphical user interface (GUI) 222 displayed via display 212. GUI 222 facilitates the initialization of a synthetic image generation process, which includes using social media application 214 to capture input images 224 of at least a first user's first face and a second user's second face via camera 210.

[0028] Social media application 214 can capture input images 224 of a first user and a second user in any suitable manner. In some implementations, social media application 214 displays an image capture prompt 226 in GUI 222. The image capture prompt 226 guides the first user and the second user to position their faces at a designated location within the field of view of camera 210. Social media application 214 controls camera 210 to capture images 224 of both users, at least based on detecting that the first user and the second user are positioned at the designated location within the field of view of camera 210. In other implementations, social media application 214 automatically captures images 224 of the first user and the second user during normal use of social media application 214 without explicitly displaying a prompt.

[0029] A trained machine learning diffusion model 228 is configured to receive input images 224 from a first user and a second user. The trained machine learning diffusion model 228 generates a first recognition image of at least the first face and a second recognition image of the second face based on the input images 224 by cropping the first and second faces from the input images. At least a first recognition patch and a second recognition patch are generated based on the first and second recognition images, respectively. A pose patch image is generated based on the first and second recognition patches and the pose image. The pose image may be generated based on a reference image 225 depicting the pose of one or more individuals to be used in the synthesized image 230.

[0030] First word tags and second word tags are generated based on the first and second recognition images, respectively. Furthermore, a tag embedding is generated based on input prompt 234, and then the tag embedding is concatenated with the word tags generated based on the extracted recognition images to generate a concatenated tag embedding. A trained machine learning diffusion model 228 generates a synthetic image 130 based on the pose patch image and the concatenated tag embedding.

[0031] The pose patch image and the stitched label embedding are input into the control network to generate features. Then, the features, latent noise, and the stitched label embedding are input into a diffusion model to generate a synthetic image 230 based at least on a first recognition image of a first face and a second recognition image of a second face.

[0032] The synthetic image 230 includes the faces of a first user and a second user extracted into a recognition image by a trained machine learning diffusion model 228. In the synthetic image 230, the first face of the first user and the second face of the second user are depicted on an individual posing in the same posture as in the reference image 225.

[0033] In some implementations, the trained machine learning diffusion model 228 can be executed locally on computing device 200. In other implementations, the trained machine learning diffusion model 228' can be executed on a remote computing system, such as server computing system 220. In one example, computing device 200 sends a user's image 224 to server computing system 220 via computer network 216. The trained machine learning diffusion model 228' generates a synthetic image 230, and server computing system 220 sends the synthetic image 230 to computing device 200 via computer network 216.

[0034] Social media application 214 is configured to display the user's composite image 230 in GUI 222 for the user to view. Furthermore, social media application 214 is configured to publish or share the user's composite image 230 to social networking platform 218 for other users of social networking platform 218 to view.

[0035] In one implementation where the composite image 230 is generated on the computing device 200, the computing device 200 sends the composite image 230 to the server computing system 220 via the computer network 216 for publication or sharing on the social networking platform 218. In another implementation where the composite image 230 is generated on the server computing system 220, the server computing system 200 directly publishes the composite image 230 to the social networking platform 218.

[0036] In some implementations, the social media application 214 may optionally be configured to capture video streams 232 of a first user and a second user via camera 210. The video stream 232 comprises a sequence of images of the two users. The social media application 214 is configured to display the video stream 232 of the two users in a GUI 222, where a composite image 230 combining one or more individuals is displayed. In some examples, the video stream 232 is captured before the composite image 230 is generated, and then the composite image 230 is merged into the video stream 232. For example, the composite image 230 may be merged into the background of the video stream 232. In other examples, the video stream 232 is captured after the composite image 230 is generated. For example, the video stream 232 may capture the user's reaction to viewing the composite image 230. The composite image 230 may be merged into the video stream 232 in any suitable manner. In addition, the social media application 214 may optionally publish the composite image 230 to the social network platform 218 by publishing the user's video stream 232 of the merged composite image 230 to the social network platform 218 for other users of the social network platform 218 to view.

[0037] Figure 6 A flowchart of an example method 300 for generating a synthetic image is shown. Example method 300 can be derived from... Figure 1 The computing system 10 includes a processing circuit device 102 and a memory 104. Figure 2 The computing system 20's processing circuitry 202 and memory 204 execute the process. Example method 300 includes, in step 302, receiving an input prompt and one or more input images. A first example method 300 includes, in step 304, generating one or more recognition images based on the one or more input images.

[0038] In step 306, method 300 includes generating one or more recognition patches based on one or more recognition images, respectively. In step 308, method 300 includes generating a pose patch image based on one or more recognition patches and a pose image. In step 310, method 300 includes generating one or more word tokens based on one or more recognition images, respectively. In step 312, method 300 includes generating a token embedding based on input cues. In step 314, method 300 includes concatenating one or more word tokens and token embeddings to generate a concatenated token embedding. In step 316, method 300 includes inputting the pose patch image and the concatenated token embedding into a control network to generate features. In step 318, method 300 includes inputting the features, latent noise, and the concatenated token embedding into a diffusion model to generate a synthetic image. In some examples, the diffusion model may be a latent diffusion model. In step 320, method 300 includes generating an output based on the synthetic image.

[0039] As described throughout this paper, by generating recognition patches and pose patches based on recognition images extracted from one or more input images, it is possible to synthesize images comprising multiple distinct individuals, thereby depicting their interactions in a more realistic manner. Therefore, the limitations of traditional attention-based mechanisms can be overcome by avoiding the problem of visual feature leakage, where human-specific visual features are unintentionally mixed together and the distinct identities of each individual are not well preserved.

[0040] In some embodiments, the methods and processes described herein can be attached to a computing system of one or more computing devices. In particular, such methods and processes can be implemented as computer applications or services, application programming interfaces (APIs), libraries, and / or other computer program products.

[0041] Figure 7 A non-limiting embodiment of a computing system 400 is schematically shown, which can implement one or more of the methods and processes described above. The computing system 400 is shown in a simplified form. The computing system 400 can be embodied as described above and... Figure 1 The computing system 10 shown in the figure or described above and in Figure 5 The computer system 20 shown is included. Components of the computing system 400 may be included in one or more personal computers, server computers, tablet computers, home entertainment computers, network computing devices, video game devices, mobile computing devices, mobile communication devices (e.g., smartphones) and / or other computing devices, as well as wearable computing devices such as smartwatches and head-mounted augmented reality devices.

[0042] The computing system 400 includes processing circuitry 402, volatile memory 404, and non-volatile storage device 406. The computing system 400 may optionally include a display subsystem 408, an input subsystem 410, a communication subsystem 412, and / or... Figure 7 Other components not shown.

[0043] Processing circuitry 402 typically includes one or more logic processors, which are physical devices configured to execute instructions. For example, a logic processor may be configured to execute instructions that are part of one or more application programs, programs, routines, libraries, objects, components, data structures, or other logical constructs. Such instructions may be implemented to perform tasks, implement data types, transform the state of one or more components, achieve technical effects, or otherwise obtain desired results.

[0044] The logic processor may include one or more physical processors configured to execute software instructions. Alternatively, the logic processor may include one or more hardware logic circuits or firmware devices configured to execute hardware-implemented logic or firmware instructions. The processor of the processing circuitry 402 may be single-core or multi-core, and the instructions executed thereon may be configured for sequential, parallel, and / or distributed processing. Optionally, the various components of the processing circuitry 402 may be distributed among two or more separate devices, which may be located remotely and / or configured for coordinated processing. For example, aspects of the computing system disclosed herein may be virtualized and executed by remotely accessible networked computing devices configured in a cloud computing configuration. In this case, it should be understood that these virtualized aspects run on different physical logic processors on various different machines. These different physical logic processors on different machines will be understood as being collectively included in the processing circuitry 402.

[0045] The non-volatile storage device 406 includes one or more physical devices configured to store instructions executable by the processing circuitry 402 to implement the methods and processes described herein. When implementing such methods and processes, the state of the non-volatile storage device 406 can be transformed—for example, to store different data.

[0046] Non-volatile storage device 406 may include removable and / or built-in physical devices. Non-volatile storage device 406 may include optical storage, semiconductor memory, and / or magnetic storage, or other high-capacity storage technologies. Non-volatile storage device 406 may include non-volatile, dynamic, static, read / write, read-only, sequential access, location-addressable, file-addressable, and / or content-addressable devices. It should be understood that non-volatile storage device 406 is configured to retain commands even when power to it is cut off.

[0047] Volatile memory 404 may include physical devices, including random access memory. Volatile memory 404 is typically used by processing circuitry 402 to temporarily store information during the processing of software instructions. It should be understood that when power to volatile memory 404 is cut off, it typically does not continue storing instructions.

[0048] The processing circuitry 402, volatile memory 404, and non-volatile storage device 406 can be integrated together into one or more hardware logic components. For example, such hardware logic components may include field-programmable gate arrays (FPGAs), application-specific integrated circuits (PASICs / ASICs), application-specific standard products (PSSPs / ASSPs), system-on-a-chip (SOCs), and complex programmable logic devices (CPLDs).

[0049] The terms "module," "program," and "engine" can be used to describe aspects of computing system 400, typically implemented in software by a processor, to perform specific functions using portions of volatile memory. These functions involve transformative processing specifically configured to perform the functions. Therefore, a module, program, or engine can be instantiated via processing circuitry 402 that executes instructions stored in non-volatile storage device 406 using portions of volatile memory 404. It should be understood that different modules, programs, and / or engines can be instantiated from the same application, service, code block, object, library, routine, API, function, etc. Similarly, the same module, program, and / or engine can be instantiated from different applications, services, code blocks, objects, routines, APIs, functions, etc. The terms "module," "program," and "engine" can include individual executable files, data files, libraries, drivers, scripts, database records, etc., or groups thereof.

[0050] When included, the display subsystem 408 can be used to present a visual representation of the data stored in the non-volatile storage device 406. The visual representation may take the form of a graphical user interface (GUI). As the methods and processes described herein change the data stored in the non-volatile storage device, and thus change the state of the non-volatile storage device, the state of the display subsystem 408 can also be changed to visually represent the change in the underlying data. The display subsystem 408 may include one or more display devices utilizing virtually any type of technology. Such display devices may be combined with processing circuitry 402, volatile memory 404, and / or non-volatile storage device 406 within a shared housing, or such display devices may be peripheral display devices.

[0051] When included, the input subsystem 410 may include one or more user input devices, such as a keyboard, mouse, touchscreen, camera, or microphone, or be coupled thereto.

[0052] When included, the communication subsystem 412 can be configured to communicatively couple the various computing devices described herein to each other and to other devices. The communication subsystem 412 may include wired and / or wireless communication devices compatible with one or more different communication protocols. As a non-limiting example, the communication subsystem can be configured to communicate via wired or wireless local area networks or wide area networks, broadband cellular networks, etc. In some embodiments, the communication subsystem may allow the computing system 400 to send messages to and / or receive messages from other devices via a network such as the Internet.

[0053] The following paragraphs provide additional description of the subject matter of this disclosure. One aspect provides a computational system for generating synthetic images, the computational system including processing circuitry and a memory storing instructions that, when executed, cause the processing circuitry to receive an input cue and one or more input images, generate one or more recognition images based on the one or more input images, generate one or more recognition patches based on the one or more recognition images respectively, generate a pose patch image based on the one or more recognition patches and a pose image, generate one or more word tokens based on the one or more recognition images respectively, generate a token embedding based on the input cue, concatenate the one or more word tokens and the token embedding to generate a concatenated token embedding, input the pose patch image and the concatenated token embedding into a control network to generate features, input the features, latent noise, and the concatenated token embedding into a diffusion model to generate a synthetic image, and generate an output based on the synthetic image. In this aspect, additionally or alternatively, the one or more recognition patches may each encode visual features of the one or more recognition images. In this aspect, additionally or alternatively, the one or more recognition images may be cropped faces of one or more individuals identified in one or more input images. In this regard, alternatively or additionally, the pose image may be a pixelated image of a vector representation of the skeletal structure of one or more individuals. Additionally or additionally, the pose image may be generated based on a reference image depicting the pose of one or more individuals. Additionally or additionally, the pose patch image may be generated by overlaying one or more recognition patches onto the head position of an individual in the pose image. Additionally or additionally, one or more word tags may be generated by a cue encoder that maps the visual information of each recognition image to the natural language tag space of the cue encoder. Additionally or additionally, the cue encoder may be configured as a contrastive language image pre-trained (CLIP) text encoder. Additionally or additionally, the concatenated tag embeddings may be fed into the attention layer of a diffusion model. In this respect, additionally or alternatively, the control network may include an encoder configured as a trainable copy of the encoder of the diffusion model, a zero-initialized convolutional layer placed at the output of the encoder of the control network, and an intermediate block configured as a trainable copy of the intermediate block of the diffusion model, the pose map patch image being input into the decoder of the control network, and the stitched label embedding being input into the attention layer of the encoder and the intermediate block of the control network.

[0054] On the other hand, a computational method for generating synthetic images is provided. This method includes receiving an input cue and one or more input images; generating one or more recognition images based on the one or more input images; generating one or more recognition patches based on the one or more recognition images; generating a pose patch image based on the one or more recognition patches and a pose image; generating one or more word tags based on the one or more recognition images; generating a tag embedding based on the input cue; concatenating the one or more word tags and the tag embedding to generate a concatenated tag embedding; inputting the pose patch image and the concatenated tag embedding into a control network to generate features; inputting the features, latent noise, and the concatenated tag embedding into a diffusion model to generate a synthetic image; and generating an output based on the synthetic image. In this aspect, additionally or alternatively, the one or more recognition patches may encode visual features of the one or more recognition images. In this aspect, additionally or alternatively, the one or more recognition images may be cropped faces of one or more individuals labeled in one or more input images. In this aspect, additionally or alternatively, the pose image may be a pixelated image of a vector representation of the skeletal structure of one or more individuals. In this aspect, additionally or alternatively, the pose image may be generated based on a reference image depicting the pose of one or more individuals. In this respect, additionally or alternatively, pose patch images can be generated by overlaying one or more recognition patches onto the head position of an individual in the pose image. In this respect, additionally or alternatively, one or more word tags can be generated by mapping the visual information of each recognition image to the natural language tag space of the cue encoder using a cue encoder. In this respect, additionally or alternatively, the concatenated tag embeddings can be input into the attention layer of the diffusion model. In this respect, additionally or alternatively, the control network can include an encoder configured as a trainable copy of the encoder of the diffusion model, a zero-initialized convolutional layer placed at the output of the encoder of the control network, and an intermediate block configured as a trainable copy of the intermediate block of the diffusion model, the pose patch images being input into the decoder of the control network, and the concatenated tag embeddings being input into the encoder of the control network and the attention layer of the intermediate block.

[0055] On the other hand, a computing device is provided, including a camera, a display, and a processing circuit device configured to execute instructions stored in a memory to execute a social media application. The social media application includes a graphical user interface (GUI) displayed via the display. The social media application is configured to communicate with a social networking platform running on a server computing system via a computer network. The social media application uses the camera to capture input images of at least a first face of a first user and a second face of a second user, receives input prompts, generates a first recognition image of at least the first face and a second recognition image of the second face based on the input images, and generates at least a first recognition patch and a second recognition patch based on the at least first recognition image and the second recognition image, respectively. The system generates a pose image based on a first and second recognition image and a pose image. It generates a first word tag and a second word tag based on the first and second recognition images, respectively. It generates a tag embedding based on input prompts. It concatenates the first and second word tags with the tag embedding to generate a concatenated tag embedding. It inputs the pose image and the concatenated tag embedding into a control network to generate features. It inputs the features, latent noise, and the concatenated tag embedding into a diffusion model to generate a synthetic image based at least on the first recognition image of the first face and the second recognition image of the second face. It displays the synthetic image of the first user and the second user in a GUI and publishes the synthetic image of the first user and the second user to a social networking platform for other users on the social networking platform to view.

[0056] It should be understood that the configurations and / or methods described herein are exemplary in nature, and these specific embodiments or examples should not be considered limiting, as many variations are possible. The specific routines or methods described herein may represent one or more of any number of processing strategies. Therefore, the various actions shown and / or described may be performed in the order shown and / or described, in other orders, in parallel, or omitted. Similarly, the order of the above processes may be changed.

[0057] It should be understood that the “and / or” used in this article refers to the logical disjunction operation, and therefore A and / or B have the following truth table.

[0058] The subject matter of this disclosure includes all novel and non-obvious combinations and sub-combinations of the various processes, systems and configurations disclosed herein, as well as any and all equivalent forms thereof, of other features, functions, actions and / or attributes.

Claims

1. A computational system for generating synthetic images, the computational system comprising: The processing circuitry and the memory storing instructions, which, when executed, cause the processing circuitry to: Receive input prompts and one or more input images; Based on the one or more input images, generate one or more recognition images; One or more recognition patches are generated based on the one or more recognition images, respectively; Generate a pose patch image based on the one or more recognition patches and pose image; One or more word tags are generated based on the one or more recognized images, respectively; Generate a tag embedding based on the input prompt; The one or more word tags and the tag embeddings are concatenated to generate a concatenated tag embedding; The pose patch image and the stitched markers are embedded into the control network to generate features; The features, latent noise, and the stitched markers are embedded into a diffusion model to generate the synthetic image; as well as Output is generated based on the synthesized image.

2. The computing system according to claim 1, wherein the one or more recognition blocks respectively encode the visual features of the one or more recognition images.

3. The computing system of claim 2, wherein the one or more recognition images are cropped faces of one or more individuals identified in the one or more input images.

4. The computing system according to claim 1, wherein the pose image is a pixelated image of a vector representation of the skeletal structure of one or more individuals.

5. The computing system of claim 1, wherein the pose image is generated based on a reference image depicting the pose of one or more individuals.

6. The computing system of claim 1, wherein the pose patch image is generated by overlaying the one or more recognition patches onto the head position of an individual in the pose image.

7. The computing system of claim 1, wherein the one or more word tags are generated by mapping the visual information of each recognized image to the natural language tag space of the cue encoder via a cue encoder.

8. The computing system of claim 7, wherein the cue encoder is configured as a contrastive language image pre-trained CLIP text encoder.

9. The computing system of claim 1, wherein the spliced ​​tag embedding is input into the attention layer of the diffusion model.

10. The computing system according to claim 1, wherein The control network includes: The encoder is configured as a trainable copy of the encoder of the diffusion model; A zero-initialized convolutional layer is placed at the output of the encoder of the control network; as well as Intermediate blocks are configured as trainable copies of intermediate blocks in the diffusion model, wherein The attitude patch image is input into the encoder of the control network; and The stitched marker embeddings are input into the attention layer of the encoder and the intermediate block of the control network.

11. A computational method for generating a synthetic image, the computational method comprising: Receive input prompts and one or more input images; Based on the one or more input images, generate one or more recognition images; One or more recognition patches are generated based on the one or more recognition images, respectively; Generate a pose patch image based on the one or more recognition patches and pose image; One or more word tags are generated based on the one or more recognized images, respectively; Generate a tag embedding based on the input prompt; The one or more word tags and the tag embeddings are concatenated to generate a concatenated tag embedding; The pose patch image and the stitched markers are embedded into the control network to generate features; The features, latent noise, and the stitched markers are embedded into a diffusion model to generate the synthetic image; as well as Output is generated based on the synthesized image.

12. The calculation method according to claim 11, wherein the one or more recognition blocks respectively encode the visual features of the one or more recognition images.

13. The calculation method of claim 12, wherein the one or more recognized images are cropped faces of one or more individuals identified in the one or more input images.

14. The calculation method according to claim 11, wherein the pose image is a pixelated image of a vector representation of the skeletal structure of one or more individuals.

15. The calculation method of claim 11, wherein the pose image is generated based on a reference image depicting the pose of one or more individuals.

16. The calculation method of claim 11, wherein the pose patch image is generated by overlaying the one or more recognition patches onto the head position of an individual in the pose image.

17. The computation method of claim 11, wherein the one or more word tags are generated by mapping the visual information of each recognized image to the natural language tag space of the cue encoder through a cue encoder.

18. The computational method of claim 11, wherein the spliced ​​tag embedding is input into the attention layer of the diffusion model.

19. The calculation method according to claim 11, wherein The control network includes: The encoder is configured as a trainable copy of the encoder of the diffusion model; A zero-initialized convolutional layer is placed at the output of the encoder of the control network; as well as Intermediate blocks are configured as trainable copies of intermediate blocks in the diffusion model, wherein The attitude patch image is input into the encoder of the control network; and The stitched marker embeddings are input into the attention layer of the encoder and the intermediate block of the control network.

20. A computing device, comprising: camera; monitor; as well as The processing circuit device is configured as follows: Instructions stored in the memory are executed to execute a social media application, the social media application including a graphical user interface (GUI) displayed via the display, the social media application being configured to communicate with a social networking platform running on a server computing system via a computer network; The social media application is used to capture input images of at least the first face of a first user and the second face of a second user via the camera. Receive input prompts; At least a first recognition image of the first face and a second recognition image of the second face are generated based on the input image; At least a first recognition patch and a second recognition patch are generated based on at least the first recognition image and the second recognition image, respectively. Based on the first and second recognition patches and the pose image, a pose patch image is generated; A first word tag and a second word tag are generated based on the first recognition image and the second recognition image, respectively; Generate a tag embedding based on the input prompt; The first word tag and the second word tag are concatenated with the tag embedding to generate a concatenated tag embedding; The pose patch image and the stitched markers are embedded into the control network to generate features; The features, latent noise, and the stitched markers are embedded into a diffusion model to generate a synthetic image based at least on the first recognition image of the first face and the second recognition image of the second face; The composite image of the first user and the second user is displayed in the GUI; as well as The composite image of the first user and the second user is published to the social networking platform for other users of the social networking platform to view.