Generating synthetic digital images using localized constrained text-to-image generation neural networks
By generating embedding sequences and training generative neural networks through a two-level neural network of a localized constraint system, the accuracy problem of generative neural networks in generating multiple object images under complex prompts is solved, achieving higher image generation accuracy and flexibility.
Patent Information
- Application Number
- CN202510051001.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Priority Date
- 2024-03-28
- Filing Date
- 2025-01-13
- Publication Date
- 2025-09-30
AI Technical Summary
Existing generative neural networks have difficulty accurately reflecting the constraints of textual prompts when generating synthetic digital images, especially when processing complex prompts, they are unable to accurately generate images of multiple objects with different visual attributes.
A localized constraint system is adopted, and a two-level neural network is used to generate an embedding sequence. The encoder level generates object text embedding and replaces it with the visual embedding. The decoder level is combined to generate synthetic digital images, and the generative neural network is trained using localization loss and diffusion loss.
The accuracy and flexibility of generating synthetic images are improved, and synthetic image content with correct object attribute binding can be accurately generated, solving the accuracy problem of traditional systems under complex prompts.
Smart Images

Figure CN120726151A_ABST
Abstract
Description
Background Art
[0001] Improvements in machine learning and neural network-based image processing techniques have significantly improved the ability of computing systems to generate synthetic digital image content. In particular, many entities utilize generative neural networks to generate synthetic digital images for use in a variety of different applications. For example, entities use generative neural networks to create new images, replace objects, repair images, or otherwise insert synthetic digital content into digital images. Although the quality of generative neural networks (e.g., diffusion-based models) has steadily improved in generating realistic content, ensuring that the generated content accurately reflects the constraints of the input text prompt remains a challenge in the image generation task. Therefore, traditional systems that utilize text-to-image generative neural networks lack accuracy and flexibility in generating synthetic images based on text prompts. Summary of the Invention
[0002] One or more embodiments utilize systems, methods, and non-transitory computer-readable storage media for generating digital images using localized constraints using a generative neural network to provide benefits and / or solve one or more of the above or other problems in the art. The disclosed system utilizes a two-stage neural network comprising an encoder stage and a decoder stage. In particular, the disclosed system utilizes the encoder stage to generate an embedding sequence comprising an object text embedding representing a phrase in a text prompt, the phrase indicating different objects with specific visual attributes. Additionally, the disclosed system generates visual embeddings representing images of example objects corresponding to the objects indicated in the phrases (and corresponding visual attributes), and replaces the object text embeddings with the corresponding visual embeddings to determine a modified embedding sequence. The disclosed system utilizes the decoder stage to generate a synthetic digital image comprising the objects and the corresponding visual attributes based on the modified embedding sequence comprising the visual embeddings. Thus, the disclosed system utilizes a two-stage generative neural network that accurately generates synthetic image content with correct object attribute bindings. BRIEF DESCRIPTION OF THE DRAWINGS
[0003] Various embodiments will be described and explained with additional specificity and detail through use of the accompanying drawings.
[0004] Figure 1 Illustrated is an example system environment in which a localized constraint system operates in accordance with one or more implementations.
[0005] Figure 2 An overview of a localized constraint system for generating an embedding sequence having visual embeddings for generating a synthetic digital image using a two-stage neural network is illustrated in accordance with one or more implementations.
[0006] Figure 3A diagram illustrating a localized constraint system for generating a hint embedding and multiple object text embeddings based on a text hint in accordance with one or more implementations.
[0007] Figure 4 A diagram illustrating a localized constraint system replacing textual embeddings of objects with corresponding visual embeddings in an embedding sequence in accordance with one or more implementations.
[0008] Figure 5 Illustrated is a graph showing a localized constraint system generating a visual embedding representing a synthesized object image in accordance with one or more implementations.
[0009] Figure 6 A diagram illustrating a localized constraint system generating a digital image from a modified embedding sequence in accordance with one or more implementations.
[0010] Figure 7 A diagram illustrates a localized constraint system generating a modified embedding sequence for adjusting parameters of a generative neural network in accordance with one or more implementations.
[0011] Figure 8 A diagram illustrates a localization constraint system that determines a localization loss based on a ground-truth object mask and a cross-attention map for adjusting parameters of a generative neural network in accordance with one or more implementations.
[0012] Figure 9 Illustrated is a comparison of multiple synthetic digital images generated by a localized constraint system, a localized constraint system without visual embedding, and a traditional image generation system, according to one or more implementations.
[0013] Figure 10A Illustrated are example object images provided to a localized constraint system for generating a visual embedding in accordance with one or more implementations.
[0014] Figure 10B Illustrated is a diagram of utilizing and Figure 10A Example digit images generated by the visual embedding corresponding to the object image.
[0015] Figure 11 A diagram illustrating an example of a localized constraint system in accordance with one or more implementations.
[0016] Figure 12 Illustrated is a flow diagram of a series of acts for generating a composite digital image by enforcing localization constraints via an embedding sequence including visual embeddings, in accordance with one or more implementations.
[0017] Figure 13 A block diagram of an exemplary computing device is illustrated in accordance with one or more implementations. DETAILED DESCRIPTION
[0018] One or more embodiments of the present disclosure include a localized constraint system that generates a synthetic digital image by encoding localized constraints into an embedding sequence of a text prompt. For example, the localized constraint system utilizes a two-level neural network including an encoding neural network to generate an embedding sequence that includes a prompt embedding representing a text prompt and an object text embedding representing a phrase corresponding to an object in the text prompt. In addition, the localized constraint system modifies the embedding sequence by replacing the object text embedding with a visual embedding of an object image representing the object in the text prompt. The localized constraint system utilizes a decoding neural network (e.g., a diffusion-based generative neural network) to generate a synthetic digital image based on the modified embedding sequence including the visual embedding. Therefore, the localized constraint system generates a synthetic digital image with localized constraints during encoding in a text-to-image process to accurately generate an object with correct visual properties based on the text prompt.
[0019] As mentioned, in one or more embodiments, the localized constraint system utilizes an encoding neural network to generate an embedding sequence based on textual prompts in the text-to-image process. Specifically, the localized constraint system parses the textual prompt to determine a phrase that indicates an object to be generated in the text-to-image process. For example, a phrase that indicates an object includes text that describes the object and any attributes of the object. The localized constraint system generates the embedding sequence by encoding the phrase that indicates the object as an object text embedding in a feature space. In some embodiments, the localized constraint system also encodes the textual prompt as a prompt embedding in the feature space.
[0020] According to one or more embodiments, a localized constraint system determines a modified embedding sequence by replacing the object's textual embedding with a visual embedding of an object image that includes an example object. Specifically, the localized constraint system generates a visual embedding based on the object image that includes the example object in the same feature space as the object's textual embedding. Additionally, the localized constraint system replaces the object's textual embedding with a corresponding visual embedding in the embedding sequence, thereby generating a modified embedding sequence. In some embodiments, the localized constraint system uses the modified embedding sequence to generate a composite digital image that includes the object and corresponding attributes indicated in the textual prompt.
[0021] In some embodiments, the localized constraint system also trains the generative neural network based on the ground-truth masks of the objects in the training images. Specifically, the localized constraint system determines a localized loss output of a loss function that compares the ground-truth masks with a cross-attention map generated by the generative neural network based on the modified embedding sequence for the training images. Thus, the localized constraint system uses the localized loss and the diffusion loss to adjust the parameters of the generative neural network to generate a cross-attention map that is closer to the ground-truth masks.
[0022] Some conventional systems that provide synthetic image generation utilize generative neural networks to generate digital images based on textual cues via an architecture that iteratively synthesizes images from noise patterns. For example, some conventional systems utilize generative models to generate synthetic image content via "on-the-fly" optimization of cross-attention maps to reflect prior knowledge. While such systems perform well in simple domains with simple cues, these conventional systems lack accuracy when presenting complex cues, especially cues that include multiple objects with different visual attributes. In particular, these conventional systems force the cross-attention maps to reflect certain patterns, which results in a decrease in image quality. Furthermore, such systems cannot handle cues that resolve relationships beyond attribute binding.
[0023] Some traditional systems utilize diffusion-based models that iteratively synthesize images from noise patterns. For example, traditional systems utilize text encoders with cross-attention-based conditions to generate synthetic images. These traditional systems lack accuracy because the text encoders fail to preserve the compositionality of images relative to the input text cues. Furthermore, because the outputs of such encoders are not consistent in image space with the outputs of generative neural networks, these traditional systems lack accuracy when generating image content in certain domains (e.g., humans).
[0024] Furthermore, some conventional systems utilizing diffusion-based models utilize image priors to generate synthetic digital image content. Specifically, such conventional systems generate visual features (e.g., in a single text embedding) based on text input in the prior model, and provide the text embedding to a diffusion decoder to generate synthetic image content. Because these conventional systems encode the semantic information of the textual cues into a single embedding, these conventional systems also often fail to reflect the combinatorial nature of the textual cues. Consequently, conventional systems using generative neural networks are typically unable to accurately reconstruct textual inputs that include multiple objects with different visual attributes in a synthetic digital image.
[0025] A localized constraint system provides multiple advantages in computing systems that provide digital image generation via generative neural networks. For example, the localized constraint system improves accuracy by utilizing localized constraints via multimodal embeddings in the neural network that generates the synthetic image content. In contrast to conventional systems that utilize a single embedding for textual prompts, the localized constraint system generates an embedding sequence that includes separate embeddings for phrases that refer to objects in the textual prompt. In particular, by generating an embedding sequence with separate embeddings for phrases that mention objects, the localized constraint system provides distinct embeddings to bind attributes to their corresponding objects.
[0026] Furthermore, the localized constraint system improves the accuracy of synthesized digital images via (multiple) neural networks with multimodal embeddings. In particular, the localized constraint system replaces object text embeddings representing object phrases with visual embeddings representing image objects (including instances of the corresponding objects). Thus, the localized constraint system generates synthetic digital images using an embedding sequence of text embeddings and visual embeddings in the same feature space. Consequently, compared to conventional systems that generate digital images using only text encodings, the localized constraint system provides improved priors to accurately generate synthetic image content across various domains while also providing the correct object composition relative to textual cues.
[0027] In an additional embodiment, a localized constraint system utilizes a loss function to determine a combined diffusion loss and a localization loss to improve the accuracy of the generated synthetic image content. For example, the localized constraint system utilizes a set of training images to reduce the output of the loss function based on a comparison of a cross-attention map generated by a generative neural network with a ground-truth mask of the training image. Thus, compared to existing systems that use "on-the-fly" optimization of cross-attention maps according to a specific pattern, the localized constraint system uses a localized loss to force the individual cross-attention maps of multiple different objects in a single image to align with the object mask. Thus, the localized constraint system provides an improved generative neural network that more accurately generates synthetic image content with the correct combination of objects by combining diffusion loss and localization loss.
[0028] Turning now to the accompanying drawings, Figure 1 An embodiment of a system environment 100 is included in which a localized constraint system 102 is implemented. In particular, the system environment 100 includes server device(s) 104 and client device(s) 106 communicating via a network 108. Furthermore, as shown, the server device(s) 104 include a digital imaging system 110 that includes the localized constraint system 102. Furthermore, the localized constraint system 102 includes or has access to encoder neural network(s) 112 and generative neural network 114. Although Figure 1The server device(s) 104 are shown hosting the encoder neural network(s) 112 and / or the generative neural network 114, but in alternative embodiments, the encoder neural network(s) 112 and / or the generative neural network 114 are hosted by another device or system (e.g., a third-party computing system). Additionally, the client device 106 includes a digital imaging application 116, which optionally includes the digital imaging system 110 (and the localization constraint system 102).
[0029] like Figure 1 As shown in FIG, client device 106 or server device(s) 104 include or host a digital image system 110. Digital image system 110 includes, or is part of, one or more systems that implement digital image generation or editing operations. For example, digital image system 110 provides tools for generating or editing digital images (e.g., in a composite image content task). For illustration, digital image system 110 communicates with client device 106 via network 108 to provide tools for display and interaction via digital image application 116 at client device 106. Additionally, in some embodiments, digital image system 110 receives requests to access stored digital image data (e.g., at server device(s) 104 or another device, such as a database) and / or requests to store digital image data. In some embodiments, digital image system 110 receives interaction data for viewing or performing various image processing operations, and provides the results of the interaction data (e.g., generated digital image data) for display via digital image application 116 or to a third-party system.
[0030] According to one or more embodiments, the digital image system 110 utilizes the localized constraint system 102 to generate a synthetic image via (multiple) encoder neural networks 112 with localized constraints and a generative neural network 114. In particular, the localized constraint system 102 generates a sequence of embeddings representing phrases based on a textual prompt, the phrases indicating different objects and their corresponding attributes. Additionally, the localized constraint system 102 modifies the embeddings by replacing the object textual embeddings with visual embeddings representing example objects corresponding to the objects in the textual prompt. The localized constraint system 102 utilizes the generative neural network 114 to generate synthetic digital image content based on the modified embedding sequence including the visual embeddings. Thus, the localized constraint system 102 provides accurate generation of synthetic image content via a generative neural network pipeline (e.g., utilizing a diffusion-based model) that associates attributes with the correct objects based on the textual prompt.
[0031] like Figure 1As shown in FIG, the localized constraint system 102 is implemented on the client device 106 or on the server device(s) 104. In particular, in some implementations, the localized constraint system 102 on the server device(s) 104 supports the localized constraint system 102 on the client device 106. For example, the server device(s) 104 generates or obtains the localized constraint system 102 (e.g., the encoder neural network(s) 112 and the generative neural network 114) for the client device 106 (e.g., as part of a software application or suite). The server device(s) 104 provides the localized constraint system 102 to the client device 106 to perform the digital image generation / editing process at the client device 106. In other words, the client device 106 obtains (e.g., downloads) the localized constraint system 102 from the server device(s) 104. At this point, the client device 106 is able to utilize the localized constraint system 102 to generate / edit digital images independently of the server device(s) 104.
[0032] In additional embodiments, although Figure 1 The server device(s) 104 and the client device(s) 106 are shown communicating via the network 108, but the various components of the system environment 100 communicate and / or interact via other methods (e.g., the server device(s) 104 and the client device(s) 106 communicate directly). Figure 1 The localized constraint system 102 is illustrated as being implemented by specific components and / or devices within the system environment 100, but the localized constraint system 102 is implemented in whole or in part by other computing devices and / or components in the system environment 100. For example, in some embodiments, the server device(s) 104 include or host the digital imaging system 110 and / or the localized constraint system 102.
[0033] For illustration, the localized constraint system 102 includes a web hosting application that allows a client device 106 to interact with content and services hosted on server device(s) 104 (e.g., in a software-as-a-service implementation). For illustration, in one or more implementations, the client device 106 accesses a web page hosted by the server device(s) 104. The client device 106 provides input to the server device(s) 104 to perform digital image generation, and in response, the localized constraint system 102 or digital image system 110 on the server device(s) 104 performs operations to generate a digital image via the encoder neural network(s) 112 and the generative neural network 114. The server device(s) 104 provide outputs or results of the operations to the client device 106.
[0034] In one or more embodiments, the server device(s) 104 include various computing devices, including the following references: Figure 13 Those described above. For example, the server device(s) 104 include one or more servers for storing and processing data associated with image generation and editing. In some embodiments, the server device(s) 104 also include multiple computing devices that communicate with each other, such as in a distributed storage environment. In some embodiments, the server device(s) 104 include content servers. The server device(s) 104 may also optionally include application servers, communication servers, web hosting servers, social networking servers, digital content activity servers, or digital communication management servers.
[0035] In addition, if Figure 1 As shown in FIG, the system environment 100 includes a client device 106. In one or more embodiments, the client device 106 includes but is not limited to a mobile device (e.g., a smartphone or tablet), a laptop, a desktop computer, and the like. Figure 13 Furthermore, although Figure 1 Although not shown, client device 106 may be operated by a user (e.g., a user included in or associated with system environment 100) to perform various functions. In particular, client device 106 performs functions such as, but not limited to, accessing, viewing, generating, and editing digital images. In some embodiments, client device 106 also performs functions for generating, capturing, or accessing data to provide to digital image system 110 and localized constraint system 102 in connection with editing digital images. For example, client device 106 communicates with server device(s) 104 via network 108 to provide information associated with digital images (e.g., user interactions). Although Figure 1 The system environment 100 is illustrated with a single client device, but in some embodiments, the system environment 100 includes a different number of client devices.
[0036] In addition, if Figure 1 As shown in , the system environment 100 includes a network 108. The network 108 enables communication between components of the system environment 100. In one or more embodiments, the network 108 may include the Internet or the World Wide Web. In addition, the network 108 optionally includes various types of networks using various communication technologies and protocols, such as a corporate intranet, a virtual private network (VPN), a local area network (LAN), a wireless local area network (WLAN), a cellular network, a wide area network (WAN), a metropolitan area network (MAN), or a combination of two or more such networks. In practice, the (multiple) server devices 104 and the client devices 106 communicate via the network using one or more communication platforms and technologies suitable for transmitting data and / or communication signals (including any known communication technologies, devices, media and protocols that support data communication), examples of which will be referenced. Figure 13 Provide a description.
[0037] As mentioned, the localized constraint system 102 utilizes one or more neural networks with localized constraints to generate composite image content with correct object compositionality. Figure 2 The diagram illustrates a localized constraint system 102 for generating digital images based on textual prompts using multiple neural networks. In particular, Figure 2 The localized constraint system 102 is illustrated as generating a digital image based on a sequence of embeddings representing textual cue elements.
[0038] like Figure 2 As shown in FIG, the localized constraint system 102 determines a textual prompt 200 for generating composite digital image content, the textual prompt being associated with a request to perform a text-to-image operation. For example, the textual prompt 200 includes a sentence, phrase, or combination of phrases in a textual format for generating or editing a digital image. For illustration, the textual prompt 200 includes "a red sheep and a white car," which indicates a request to generate one or more objects in a scene.
[0039] In one or more embodiments, the text prompt 200 includes multiple separate phrases for indicating multiple objects to be included in the generated image content. Additionally, in at least some embodiments, the phrases include a description of the object having one or more visual attributes of the object (e.g., color, size, shape, position, or other appearance characteristics). Thus, in various embodiments, the text prompt 200 includes one or more words for describing the compositionality of multiple objects in the scene used to generate the composite image content. For example, the first phrase indicates "a red sheep" while the second phrase indicates "a white car," with each phrase representing a separate object and corresponding attribute (or set of attributes).
[0040] In addition, in one or more embodiments, the localized constraint system 102 encodes the text prompt 200 into the feature space using one or more encoder neural networks 202. For example, the localized constraint system 102 generates an embedding sequence 204 using the encoder neural network(s) 112 that includes a text embedding representing the text prompt 200 and one or more phrases in the text prompt 200. Furthermore, the localized constraint system 102 generates a visual embedding representing the object indicated by the text prompt 200, such as by encoding, using the encoder neural network(s) 202, an example object in an object image generated by (or otherwise corresponding to) the phrases of the text prompt 200. Thus, the localized constraint system 102 generates the embedding sequence 204 to include the visual embedding 206 in the feature space of the text embedding. Figure 3-Figure 5 and the corresponding description provide more details on generating embeddings based on textual prompts.
[0041] The localized constraint system 102 also generates a digital image 210 based on the embedded sequence 204 using a generative neural network 208. In particular, Figure 2 , localized constraint system 102 utilizes generative neural network 208 to generate composite image content in digital image 210 based on text prompt 200. More particularly, localized constraint system 102 generates digital image 210 to include one or more objects indicated in text prompt 200 with correct compositionality (e.g., correctly assigning attributes to corresponding objects indicated in text prompt 200). Figure 6 and the corresponding description provide more details on generating digital images based on embedding sequences including textual embeddings and visual embeddings.
[0042] In one or more embodiments, the localized constraint system 102 generates embeddings in feature space based on textual prompts by utilizing one or more encoder neural networks. Figure 3 An example of a localized constraint system 102 generating an embedding representing a textual prompt to generate a digital image is shown. In particular, Figure 3 The localization constraint system 102 is illustrated as generating embeddings representing text in a text prompt and portions of the text for use in generating composite image content.
[0043] like Figure 3 As illustrated in FIG, the localized constraint system 102 determines a textual prompt 300 to generate a digital image including composite image content. For example, as mentioned, the textual prompt 300 includes a phrase indicating one or more objects to be included in the composite image content. In some embodiments, the textual prompt 300 includes one or more natural language phrases that include one or more sub-phrases indicating the object(s) to be included in the composite image content. Furthermore, the textual prompt 300 includes compositional properties indicating the layout of the object(s) in the scene, the visual attributes of the object(s), and / or the relationships between the multiple objects.
[0044] In one or more embodiments, the localization constraint system 102 utilizes a parser 302 to segment the text prompt 300 into phrases 304a-304n. For example, the parser 302 includes a natural language parser that identifies parts of speech and relationships between parts of speech to segment the text prompt 300 into separate text groups. To illustrate, the localization constraint system 102 utilizes the parser 302 to segment the text prompt 300 into multiple phrases by grouping words that refer to an object and its corresponding attributes into a single phrase (e.g., "a red sheep"). Thus, the multiple phrases correspond to different groups of words that correspond to separate objects and their visual attributes.
[0045] In at least some embodiments, the localized constraint system 102 utilizes an encoder neural network 306 to generate embeddings based on the text prompt 300 and phrases 304a-304n. In particular, the encoder neural network 306 comprises a text encoder neural network that converts words and / or phrases into a feature space based on features extracted from the words / phrases. Thus, the localized constraint system 102 utilizes a pre-trained text encoder neural network that encodes the phrases 304a-04n into a feature space learned based on the relationship between image content and corresponding textual content describing the image content. Thus, the feature space comprises an abstract embedding space representing features of the text and / or image content.
[0046] In one or more embodiments, the localized constraint system 102 utilizes the encoder neural network 306 to generate a prompt embedding 308 that represents the entire text prompt 300. In particular, the prompt embedding 308 includes an encoding that represents the entire text prompt 300 in a feature space. For example, the localized constraint system 102 utilizes the encoder neural network 306 to embed features of all elements of the text prompt 300 into a single embedding (e.g., a single feature vector) in the feature space. In some embodiments, the localized constraint system 102 does not generate the prompt embedding 308, but instead generates embeddings only for portions of the text prompt 300. In additional embodiments, the localized constraint system 102 generates the prompt embedding 308 and identifies specific portions of the prompt embedding 308 that correspond to the object.
[0047] In addition, if Figure 3 As illustrated in , the localized constraint system 102 generates object text embeddings 310a-310n using an encoder that represents a phrase indicating an object in the text prompt 300. In particular, the localized constraint system 102 generates the object text embeddings 310a-310n based on phrases 304a-304n extracted from the text prompt 300 using an encoder neural network 306. As an example, the localized constraint system 102 generates a first object text embedding 310a that represents a first phrase 304a indicating an object and its attributes (e.g., "a red sheep"). In an additional example, the localized constraint system 102 generates a second object text embedding 310n that represents a second phrase 304n indicating another object and its attributes (e.g., "a white car"). Thus, in one or more embodiments, the localized constraint system 102 generates a prompt embedding 308 that represents the entire text prompt 300 and one or more object text embeddings that represent individual objects (and their attributes) indicated in the text prompt 300.
[0048] In one or more embodiments, in response to generating embeddings from a text prompt, localized constraint system 102 determines a sequence of embeddings to provide for use in generating composite image content. Figure 4 An embodiment is illustrated in which the localization constraint system 102 determines an initial embedding sequence comprising text embeddings based on a text prompt. Figure 4 It is illustrated that the localization constraint system 102 determines a modified embedding sequence by replacing one or more text embeddings with visual embeddings corresponding to object images representing objects mentioned in the text prompt.
[0049] According to one or more embodiments, as mentioned, the localized constraint system 102 utilizes an encoder neural network to generate embeddings for portions of the text prompt 400. In one or more embodiments, the localized constraint system 102 determines the embedding sequence 402 by generating multiple text embeddings representing different words and / or phrases in the text prompt 400. For example, the localized constraint system 102 generates multiple text embeddings (e.g., a first text embedding 402a, a second text embedding 402b, and a third text embedding 402c) representing words (or groups of words representing various concepts) in the text prompt 400. In addition, the localized constraint system 102 generates multiple object text embeddings (e.g., a first object text embedding 402d and a second object text embedding 402e) representing words or phrases indicating objects and corresponding attributes (e.g., adjectives describing the objects) in the text prompt 400.
[0050] In some embodiments, the localization constraint system 102 determines the text embeddings and the object text embeddings in order based on the order in which the words / phrases appear in the text prompt 400. In additional embodiments, the localization constraint system 102 determines the embedding sequence 402 to include a hint embedding representing the entire text prompt 400, one or more text embeddings representing non-object words / phrases, and one or more object text embeddings representing phrases indicating objects. In still other embodiments, the localization constraint system 102 determines the embedding sequence 402 to include the hint embedding and one or more object text embeddings while excluding text embeddings representing non-objects.
[0051] like Figure 4As illustrated in , by generating the time step encoding 404 using the encoder neural network, the localized constraint system 102 further determines the embedding sequence 402 to include the time step encoding 404 corresponding to the time step parameter for use in the diffusion-based generative neural network. Additionally, as illustrated, the embedding sequence 402 also includes a noisy visual embedding 406 corresponding to the noisy input of the diffusion-based generative neural network. To illustrate, the localized constraint system 102 determines one or more portions (e.g., one or more noise patches) of the noisy image input by generating the noisy visual embedding 406 using the encoder neural network.
[0052] In one or more embodiments, the localized constraint system 102 determines a modified embedding sequence 408 by replacing one or more embeddings in the embedding sequence 402 with embeddings that represent visual image content. In particular, as illustrated, the localized constraint system 102 determines one or more object images (e.g., a first object image 410a and a second object image 410b) that include example objects corresponding to the object mentioned in the text prompt 400. For example, the localized constraint system 102 determines the object images in response to a user selection of the object images (e.g., via a graphical user interface). To illustrate, the localized constraint system 102 provides a plurality of object images that include different examples of the object in the phrase of the text prompt 400 (e.g., different versions of the object generated by a generative neural network or different versions of the object generated from a selection of images in a database). In additional embodiments, the localized constraint system 102 generates the object images using a generative neural network based on a phrase extracted from the text prompt 400, as described below with respect to Figure 5 Described in more detail.
[0053] According to one or more embodiments, the localized constraint system 102 generates visual embeddings for the object images. In particular, the localized constraint system 102 generates visual embeddings based on the object images using an encoder neural network 412 (e.g., an image encoder neural network). For example, the localized constraint system 102 generates a first visual embedding 414a representing a first object image 410a and a second visual embedding 414b representing a second object image 410b using the encoder neural network 412. In one or more embodiments, the visual embeddings are in the same feature space as the text embeddings. In other words, the text encoder neural network and the image encoder neural network generate embeddings in the same feature space such that the image and the text are represented in the same feature space.
[0054] In response to generating the visual embeddings, the localization constraint system 102 uses the visual embeddings to determine a modified embedding sequence 408. Specifically, as illustrated, the localization constraint system 102 replaces the object text embeddings with corresponding visual embeddings. For example, the localization constraint system 102 determines the position of the first object text embedding 402d in the embedding sequence 402 and removes the first object text embedding 402d.
[0055] In addition, localization constraint system 102 inserts a first visual embedding 414a corresponding to the object of first object text embedding 402d at the location of first object text embedding 402d (i.e., the location before its removal). Similarly, after removing second object text embedding 402e, localization constraint system 102 inserts a second visual embedding 414b at the location of second object text embedding 402e. Thus, localization constraint system 102 determines modified embedding sequence 408 by replacing the object text embedding with the visual embedding at the corresponding location.
[0056] In one or more embodiments, as mentioned, the localized constraint system 102 utilizes a generative neural network to generate the object image. Figure 5 An embodiment is illustrated in which the localized constraint system 102 generates a synthetic object image (including an example of the object indicated in the text prompt). In particular, the localized constraint system 102 utilizes the text prompt to generate a synthetic object image for use in generating a visual embedding to replace the object text embedding.
[0057] like Figure 5 As shown in FIG, the localized constraint system 102 parses the text prompt 500 to determine phrases 502a-502n corresponding to the objects indicated by the text prompt 500. For example, the localized constraint system 102 determines, based on the text prompt 500, a first phrase 502a corresponding to a first object and a second phrase 502n corresponding to a second object. For illustration, the first phrase 502a indicates a first object using one or more words, such as "a white hat." Furthermore, the second phrase 502n indicates a second object using one or more words, such as "a pair of blue jeans."
[0058] In addition, the localized constraint system 102 provides the phrase as a separate prompt (or in conjunction with a generation prompt) to the generative neural network 504. Thus, the localized constraint system 102 generates a first synthetic object image 506a and a second synthetic object image 506n using the generative neural network 504. Consistent with the above example, by feeding the first phrase 502a to the generative neural network, the localized constraint system 102 generates the first synthetic object image 506a including the white hat example using the generative neural network 504. By feeding the second phrase 502n to the generative neural network, the localized constraint system 102 generates the second synthetic object image 506n including the blue jeans example using the generative neural network 504.
[0059] In one or more embodiments, the localized constraint system 102 utilizes an encoder neural network 508 to generate visual embeddings based on the synthetic object images. For example, the localized constraint system 102 utilizes an image encoder neural network to generate a first visual embedding 510a representing the first synthetic object image 506a. Additionally, the localized constraint system 102 utilizes an image encoder neural network to generate a second visual embedding 510n representing the second synthetic object image 506n. As mentioned, the encoder neural network 508 generates the visual embeddings in the same feature space as the object text embeddings representing the corresponding phrases of the textual prompts.
[0060] In at least some embodiments, in response to generating the visual embedding, the localized constraint system 102 determines a modified embedding sequence for use in generating the composite digital image. Figure 6 An embodiment is illustrated in which the localized constraint system 102 utilizes a generative neural network to generate a digital image from a modified embedding sequence.
[0061] like Figure 6 As illustrated in FIG, the localized constraint system 102 determines a modified embedding sequence 600 that includes one or more non-object text embeddings (e.g., text embeddings 602a-602c), one or more visual embeddings (e.g., visual embeddings 604a-604b), a time-step embedding 606, and a noisy visual embedding 608. In one or more embodiments, the localized constraint system 102 provides the modified embedding sequence 600 to a generative neural network 610 to generate a digital image 612 including synthetic image content based on the textual prompt. For example, the localized constraint system 102 utilizes a diffusion-based neural network to generate the digital image 612.
[0062] Thus, in one or more embodiments, the localized constraint system 102 provides the modified embedding sequence 600 to one or more diffusion decoders of the generative neural network 610 to iteratively generate a digital image 612. More particularly, the localized constraint system 102 generates the digital image 612 from the text embeddings 602a-602c, the visual embeddings 604a-604b, and the noisy visual embedding 608 in a plurality of different diffusion steps based on the time-step embedding 606 using a plurality of diffusion decoders. Thus, the localized constraint system 102 generates the digital image 612 to include one or more objects mentioned in the text prompt and to have the correct combination (e.g., visual attributes) with respect to the text prompt.
[0063] According to one or more embodiments, the localized constraint system 102 determines an embedding sequence including a visual marker v i , each visual marker v i Representing or corresponding to a phrase indicating a visual object in a digital image i In particular, the localization constraint system 102 determines the sequence as the encoded text (or prompt embedding) y of the text prompt, the individual phrases p1, p2, ..., p in the text prompt, and the sequence of the encoded text (or prompt embedding) y. n Text embedding, time step t, noisy visual embedding and a learnable query sequence. In particular, the learnable query represents the visual tag v for each phrase in the text prompt i Thus, in one or more embodiments, the localized constraint system 102 determines a prior for the localized constraint system, which is expressed as:
[0064]
[0065] Figure 7-Figure 8 The figure shows a schematic diagram of a localized constraint system 102 using ground truth image data to train a generative neural network, wherein synthetic image content is generated in a text-to-image process. In particular, Figure 7 It is illustrated that the localization constraint system 102 determines an embedding sequence including a text embedding and a visual embedding based on an image-caption pair. Figure 8 The localized restraint system 102 is shown using Figure 7 The embedding sequence is used to train the generative neural network via localization loss and diffusion loss.
[0066] As mentioned, Figure 7The diagram illustrates the localization constraint system 102 determining an image-caption pair for determining an embedding sequence. Specifically, as illustrated, the localization constraint system 102 determines a digital image 700 that includes one or more objects in a scene. Furthermore, the localization constraint system 102 determines captions 702 for the digital image 700. For example, the captions 702 include text describing the digital image 700. For illustration, the captions 702 include one or more phrases describing one or more objects in the scene and various attributes of the one or more objects. In some embodiments, the captions 702 also include relative localization information for the object(s) in the digital image 700.
[0067] In at least some embodiments, the localized constraint system 102 utilizes an encoder neural network 704 to generate an embedding sequence based on the subtitles 702. In particular, as previously mentioned, the localized constraint system 102 utilizes a text encoder neural network to generate text embeddings based on the subtitles 702 in a feature space. For example, the localized constraint system 102 generates an object text embedding 708 that represents an object indicated in a single phrase of the subtitles 702.
[0068] Furthermore, in some embodiments, the localized constraint system 102 determines masked objects 710 corresponding to the digital image 700. Specifically, the localized constraint system 102 determines ground truth segmentation masks for a plurality of objects in the digital image 700. In some embodiments, the ground truth segmentation masks include binary masks to mask objects from the digital image 700 (e.g., foreground masks that isolate individual objects from the background of the digital image 700). Thus, the localized constraint system 102 determines the masked objects 710, including pixel values representing foreground objects without background information, to serve as the ground truth object image.
[0069] like Figure 7 , the localized constraint system 102 utilizes an encoder neural network 712 to generate a visual embedding 714 representing the masked object 710. In one or more embodiments, the encoder neural network 712 is an image encoder neural network that encodes the masked object 710 into the visual embedding 714 in the same feature space as the object text embedding 708. Thus, the localized constraint system 102 utilizes the encoder neural network 712 to encode the visual information of the masked object 710 into the same feature space as the text of the caption 702 describing the masked object 710.
[0070] In one or more embodiments, localization constraint system 102 determines a modified embedding sequence 716 based on embedding sequence 706 and visual embedding 714. Specifically, localization constraint system 102 replaces the object text embedding 708 in embedding sequence 706 with visual embedding 714 to create modified embedding sequence 716. Thus, localization constraint system 102 replaces the text encoding representing the object in subtitle 702 with an image embedding representing the masked object 710 from digital image 700.
[0071] like Figure 8 As shown in FIG, the localized constraint system 102 trains a generative neural network 804 using a modified embedding sequence 800 including a visual embedding 802. In particular, the localized constraint system 102 determines the modified embedding sequence 800, such as Figure 7 In addition, the localized constraint system 102 generates digital image content via one or more image generation steps using a generative neural network 804. For example, the generative neural network 804 includes a diffusion-based model that iteratively generates digital image content via multiple diffusion steps that create cross-attention maps corresponding to portions of the digital image content being generated.
[0072] In one or more embodiments, the localized constraint system 102 utilizes a generative neural network 804 to generate a plurality of cross-attention maps 806a-806n based on the visual embedding 802. For example, the localized constraint system 102 utilizes cross-attention in conjunction with denoising a noisy input (e.g., a noisy visual label) based on the corresponding visual embedding 802 and / or text embedding. In some examples, the generative neural network 804 generates the cross-attention maps 806a-806n at each diffusion step to iteratively generate a digital image output.
[0073] According to some embodiments, the localization constraint system 102 enforces alignment of the cross-attention maps 806a-806n of the visual embedding 802 with the corresponding object masks. In particular, the localization constraint system 102 utilizes object masks 808a-808n corresponding to the masked objects represented by the visual embedding 802 as ground truth masks for comparison with the cross-attention maps. For example, the localization constraint system 102 determines the binary masks (or alpha masks) used to generate the masked objects as the object masks 808a-808n. For illustration, the first object mask 808a corresponds to the mask of the first object (e.g., the mask of the pants), and the second object mask 808n corresponds to the mask of the second object (e.g., the mask of the hat).
[0074] In addition, the localization constraint system 102 compares the cross-attention maps 806a-806n with the object masks 808a-808n to determine a localization loss 810. For example, the localization constraint system 102 determines the localization loss 810 by utilizing a loss function that determines the difference between the cross-attention maps 806a-806n and the object masks 808a-808n. In addition, in some embodiments, the localization constraint system 102 determines a diffusion loss 812 (e.g., mean squared error) based on the diffusion decoder of the generative neural network 804. Thus, the localization constraint system 102 utilizes a loss function that combines the localization loss 810 and the diffusion loss 812 into a total loss 814.
[0075] In an additional embodiment, the localization constraint system 102 trains the generative neural network 804 using the total loss 814. Specifically, the localization constraint system 102 updates the parameters of the generative neural network 804 to reduce the output of the loss function (e.g., reduce the localization loss 810, the diffusion loss 812, and / or the total loss 814). Thus, the localization constraint system 102 adjusts the parameters of the generative neural network 804 to enforce alignment of the cross-attention maps 806a-806n of the visual embedding 802 with the object masks 808a-808n according to the localization loss 810.
[0076] According to one or more embodiments, the localization constraint system 102 determines the total loss L as the localized loss and diffusion / noise loss In particular, the localized constraint system 102 determines the total loss based on the loss function as Furthermore, in one or more embodiments, the localization constraint system 102 determines the localization loss as Among them, M i represents the cross attention map, and Denotes the corresponding ground-truth object mask. Therefore, the localization constraint system 102 determines the localization loss based on the sum of the differences between the cross-attention map and its corresponding ground-truth object mask. In addition, as indicated, the localization constraint system 102 determines the total loss by summing the localization loss with the diffusion loss.
[0077] Figure 9 The diagram shows a comparison of digital images generated using the localized constraint system 102 and a traditional image generation system based on text prompts. In particular, the text prompt includes text for generating a digital image, which includes "child in white shirt and black hat posing for camera." Figure 9 The localized constraint system 102 is illustrated as generating a first digital image 900 by generating an embedding sequence with visual markers representing example objects based on a textual prompt. Figure 9Illustrated is a second digital image 902 generated by the localized constraint system 102 including a generative neural network trained based on visual embeddings, but without replacing the object text embeddings with visual embeddings based on textual cues (ie, in an ablation study). Figure 9 Also illustrated is a third digital image 904 generated by a conventional image generation system that includes a diffusion-based model without using visual embeddings in inference or training.
[0078] As shown, the localized constraint system 102 generates a first digital image 900 having an accurate composition of objects with respect to the textual prompt, such that the generated objects have correct visual attributes. Figure 9 It is also illustrated that not using visual embeddings to train a generative neural network can produce a second digital image 902 that does not accurately combine objects with the correct visual attributes. Additionally, a third digital image 904 does not include the correct combination of objects that would be generated by a traditional image generation system. Figure 9 As illustrated in , the localized constraint system 102 generates synthetic image content with improved accuracy by utilizing a modified embedding sequence including visual embeddings.
[0079] In one or more embodiments, the localized constraint system 102 also provides improved flexibility in generating composite image content by providing options for customizing object features. In particular, Figures 10A-10B An example of a localized constraint system 102 is illustrated that maintains visual features of an example object from an object image in generated composite image content. For example, the localized constraint system 102 utilizes features of the example object to adjust the generated object in the composite image content.
[0080] In particular, if Figure 10A As shown in FIG, the first object image 1000 includes an example masked object with a noisy background for generating a dog in the composite image content. Figure 10A A first set of visual features 1002a and a second set of visual features 1002b of a first object are illustrated. For example, the first set of visual features 1002a and the second set of visual features 1002b indicate white fur or other distinct features of a black dog. Figure 10A A second object image 1004 is illustrated, indicating an example masked object used to generate a green field in the composite image content.
[0081] In one or more embodiments, the localized constraint system 102 determines an object mask based on a digital image provided or otherwise selected by a user's client device in association with a text prompt. For example, a user interacts with a client device to generate, upload, or select (e.g., from an image database) one or more digital images that include example objects. To illustrate, the localized constraint system 102 determines an object mentioned in a text prompt and provides multiple different versions of the example object in multiple different object images. For example, in response to the prompt "A black dog standing on a rock with a green field behind it," the localized constraint system 102 generates (or otherwise obtains) images of various types of black dogs selected by the user. The localized constraint system 102 generates an object mask for each object based on the digital image indicated in association with the text prompt.
[0082] Additionally, the localized constraint system 102 generates a visual embedding representing each masked object for use in generating a sequence of embeddings to provide to the generative neural network. Figure 10B The figure shows a comparison of digital images generated by the localized constraint system 102 and a conventional image generation system. In particular, as shown in the figure, the localized constraint system 102 is based on text prompts and uses representations Figure 10A The first digital image 1006 is generated by visually embedding the object image of the dog. As shown in the figure, the first digital image 1006 includes visual features 1008 of the dog included in the object image.
[0083] Figure 10B Also illustrated is the localized constraint system 102 generating a second digital image 1010 based on the textual prompt, without using a visual embedding corresponding to the object image. Figure 10B Further illustrated is a conventional image generation system (including a diffusion-based model) generating a third digital image 1012 based on the textual prompt in the ablation study. Since the second digital image 1010 and the third digital image 1012 are generated without using visual embeddings, the resulting objects do not retain Figure 10A Thus, as shown, the localized constraint system 102 increases the flexibility of the synthesized image content by providing the option of preserving the visual features of the example object during the digital image generation process.
[0084] Figure 11 FIGURE 1 illustrates a detailed schematic diagram of an embodiment of the localized constraint system 102. As shown, the localized constraint system 102 is implemented on (a plurality of) computing devices 1100 (e.g., Figure 1 The client device and / or server device described in, and as described below with respect to Figure 1310. In one or more embodiments, the localized constraint system 102 is implemented on a digital image system 110 on a server (described further herein). Additionally, the localized constraint system 102 includes, but is not limited to, a prompt manager 1102, an encoding manager 1104, an embedding manager 1106, an image generator 1108, a neural network manager 1110, and a data storage manager 1112. In one or more embodiments, the localized constraint system 102 is implemented on any number of computing devices. For example, in one or more embodiments, the localized constraint system 102 is implemented in a distributed system of server devices for digital image generation. Alternatively, the localized constraint system 102 is also implemented within one or more additional systems. For example, in one or more embodiments, the localized constraint system 102 is implemented on a single computing device (such as a single client device).
[0085] In one or more embodiments, each component of the localized restraint system 102 communicates with the other components using any suitable communication technology. In addition, the components of the localized restraint system 102 can communicate with one or more other devices, including other computing devices of the user, server devices (e.g., cloud storage devices), various servers, or other devices / systems. It will be appreciated that although the components of the localized restraint system 102 are Figure 13 Although shown as separate, any subcomponents may be combined into fewer components, such as into a single component, or divided into more components for a particular implementation. Figure 13 The components of are described in conjunction with the localized restraint system 102, but at least some of the components used to perform operations with the localized restraint system 102 described herein may be implemented on other devices within the environment.
[0086] In some embodiments, the components of the localized restraint system 102 include software, hardware, or both. For example, the components of the localized restraint system 102 include one or more instructions stored on a computer-readable storage medium and executable by a processor of one or more computing devices (e.g., computing device(s) 1100). When executed by one or more processors, the computer-executable instructions of the localized restraint system 102 cause the computing device(s) 1100 to perform the operations described herein. Alternatively, the components of the localized restraint system 102 include hardware, such as a dedicated processing device for performing a particular function or group of functions. Additionally or alternatively, the components of the localized restraint system 102 include a combination of computer-executable instructions and hardware.
[0087] Furthermore, the components of the localized constraint system 102 that perform the functions described herein with respect to the localized constraint system 102 can be implemented, for example, as part of a standalone application, as a module of an application, as a plug-in to an application, as a library function or function that can be called by other applications, and / or as a cloud computing model. Thus, the components of the localized constraint system 102 can be implemented as part of a standalone application on a personal computing device or mobile device. Alternatively or additionally, the components of the localized constraint system 102 can be implemented in any application that provides differential captioning of digital images, including but not limited to and CREATIVE software.
[0088] As illustrated, the localized constraint system 102 includes a prompt manager 1102 for managing text prompts for generating digital images. Specifically, the prompt manager 1102 determines text prompts for generating composite image content based on user input to a user interface. Additionally, the prompt manager 1102 parses the text prompts to determine phrases indicating objects to be included in the digital image.
[0089] In one or more embodiments, the localized constraint system 102 includes an encoding manager 1104 for encoding information from textual prompts and / or images into a feature space. For example, the encoding manager 1104 utilizes an encoder neural network to generate text embeddings (e.g., object text embeddings) based on phrases in the textual prompts. Additionally, the encoding manager 1104 utilizes an encoder neural network to generate visual embeddings based on object images.
[0090] The localization constraint system 102 also utilizes the embedding manager 1106 to generate and modify embedding sequences. Specifically, the embedding manager 1106 determines an initial embedding sequence, including object text embeddings from textual prompts. Additionally, the embedding manager 1106 determines a modified embedding sequence by replacing the object text embeddings with corresponding visual embeddings.
[0091] The localized constraint system 102 includes an image generator 1108 to generate synthetic image content. Specifically, the image generator 1108 utilizes a generative neural network (e.g., a diffusion-based model) to generate a digital image based on an embedding sequence. More specifically, the image generator 1108 utilizes a generative neural network to generate synthetic image content based on a modified embedding sequence including a visual embedding.
[0092] In one or more embodiments, the localization constraint system 102 utilizes the neural network manager 1110 to train one or more neural networks. For example, the neural network manager 1110 determines a localization loss based on a cross-attention map generated by the generative neural network and a ground-truth object mask from a training image. Furthermore, the neural network manager 1110 utilizes the localization loss to adjust parameters of the generative neural network.
[0093] The localized constraint system 102 also includes a data storage manager 1112 (which includes non-transitory computer memory) that stores and maintains data associated with generating a synthesized digital image. For example, the data storage manager 1112 stores data associated with synthesizing digital image content with localized constraints using visual embeddings in an embedding sequence. In some embodiments, the data storage manager 1112 stores text embeddings, visual embeddings, and synthesized digital image data. The data storage manager 1112 also stores data associated with training and utilizing various neural networks, including encoder neural networks and various decoder neural networks in a generative neural network.
[0094] Now go to Figure 12 , which shows a flow chart of a series of actions 1200 for generating a synthetic digital image by enforcing localization constraints via an embedding sequence comprising visual embeddings. Figure 12 The actions according to one embodiment are illustrated, but alternative embodiments may omit, add to, reorder, and / or modify Figure 12 Any action shown in . Figure 12 Alternatively, the non-transitory computer readable medium includes instructions that, when executed by one or more processors, cause the one or more processors to perform Figure 12 In another embodiment, the system includes a device configured to perform Figure 12 The processor or server of the action.
[0095] As shown, a series of actions 1200 includes an action 1202 of generating an embedding sequence for a prompt. Action 1200 includes a sub-action 1202a of generating a prompt embedding representing a textual prompt and a sub-action 1202b of generating an object text embedding representing a phrase indicating an object. The series of actions 1200 includes an action 1204 of generating a visual embedding representing an object image corresponding to the object. The series of actions 1200 also includes an action 1206 of determining a modified embedding sequence including the visual embedding. Furthermore, the series of actions 1200 includes an action 1208 of generating a composite digital image based on the modified embedding sequence.
[0096] In one or more embodiments, act 1202 involves generating, using one or more encoder neural networks, an embedding sequence comprising a cue embedding representing a text cue and an object text embedding representing a phrase indicating an object in the text cue. Additionally, act 1204 involves generating, using one or more encoder neural networks, a visual embedding representing an object image corresponding to the object. Act 1206 involves determining a modified embedding sequence by replacing the object text embedding with the visual embedding in the embedding sequence. Act 1208 further involves generating, using a generative neural network, a synthetic digital image based on the modified embedding sequence comprising the visual embedding.
[0097] In at least some embodiments, a series of actions 1200 includes generating an embedding sequence by determining, based on a text prompt, a plurality of phrases indicating a plurality of objects in the text prompt. For example, the series of actions 1200 includes generating an object text embedding indicating a phrase indicating an object based on the plurality of phrases. The series of actions 1200 also includes generating an additional object text embedding indicating an additional phrase indicating an additional object based on the plurality of phrases.
[0098] Furthermore, the series of actions 1200 includes generating an object image with the example object based on the phrase indicating the object. For example, the series of actions 1200 includes generating additional object images including additional example objects based on additional phrases indicating the additional objects in the text prompt. Furthermore, the series of actions 1200 includes generating additional visual embeddings representing the additional object images corresponding to the additional objects.
[0099] The series of actions 1200 also includes determining a modified embedding sequence by determining a position of the object text embedding in the embedding sequence. The series of actions 1200 also includes removing the object text embedding from the embedding sequence. The series of actions 1200 also includes inserting the visual embedding into the embedding sequence at the position.
[0100] Furthermore, the sequence of actions 1200 includes generating an embedding sequence including generating an object text embedding in a feature space. The sequence of actions 1200 also includes generating a visual embedding including generating a visual embedding in a feature space of the object text embedding.
[0101] The series of actions 1200 also includes generating an additional embedding sequence for the caption corresponding to the additional digital image, the additional embedding sequence including a plurality of object text embeddings representing a plurality of phrases in the caption, the plurality of phrases indicating a plurality of objects in the additional digital image. The series of actions 1200 also includes generating a plurality of visual embeddings representing the plurality of objects in the additional digital image. Furthermore, the series of actions 1200 includes determining an additional modified embedding sequence by replacing the plurality of object text embeddings with the plurality of visual embeddings. The series of actions 1200 also includes adjusting parameters of the generative neural network by reducing an output of a loss function based on the additional modified embedding sequence.
[0102] In one or more embodiments, the series of actions 1200 includes adjusting a generative neural network by determining a plurality of ground truth masks corresponding to a plurality of objects of an attached digital image. The series of actions 1200 also includes determining a localization loss based on a comparison between the plurality of ground truth masks and a plurality of cross-attention maps corresponding to a plurality of visual embeddings using a loss function. The series of actions 1200 also includes adjusting parameters of the generative neural network based on the localization loss.
[0103] In one or more embodiments, a series of actions 1200 includes generating, using one or more encoder neural networks, an embedding sequence comprising a first object text embedding representing a first phrase indicating a first object in a text prompt and a second object text embedding representing a second phrase indicating a second object in the text prompt. The series of actions 1200 also includes generating, using one or more encoder neural networks, a first visual embedding representing a first object image corresponding to the first object and a second visual embedding representing a second object image corresponding to the second object. The series of actions 1200 also includes determining a modified embedding sequence by replacing the first object text embedding with the first visual embedding and replacing the second object text embedding with the second visual embedding in the embedding sequence. Additionally, the series of actions 1200 also includes generating, using a generative neural network, a synthetic digital image based on the modified embedding sequence comprising the first visual embedding and the second visual embedding.
[0104] In one or more embodiments, the series of actions 1200 further includes generating an embedding sequence by generating a prompt embedding representing a text prompt in a feature space corresponding to the first object text embedding and the second object text embedding using one or more encoder neural networks.
[0105] In some embodiments, the sequence of actions 1200 includes generating an embedding sequence by parsing a text prompt to: determine a first object and one or more visual attributes of the first object; and determine a second object and one or more visual attributes of the second object. For example, the sequence of actions 1200 includes determining, based on a first phrase, a first object image that includes a first example object, the first example object including one or more visual attributes of the first object. The sequence of actions 1200 also includes determining, based on a second phrase, a second object image that includes a second example object, the second example object including one or more visual attributes of the second object.
[0106] The series of actions 1200 also includes determining a modified embedding sequence by determining a first position in the embedding sequence corresponding to the first object text embedding and a second position in the embedding sequence corresponding to the second object text embedding. The series of actions 1200 also includes removing the first object text embedding and the second object text embedding from the embedding sequence. In addition, the series of actions 1200 includes inserting a first visual embedding at the first position and inserting a second visual embedding at the second position.
[0107] The series of actions 1200 also includes generating a synthetic digital image by providing the modified embedding sequence with the noisy image embedding to the generative neural network.
[0108] In addition, in some embodiments, a series of actions 1200 includes generating a plurality of visual embeddings from a plurality of object text embeddings of phrases in captions representing objects in training digital images. A series of actions 1200 includes determining an additional modified embedding sequence by replacing the plurality of object text embeddings with the plurality of visual embeddings in the corresponding embedding sequence. A series of actions 1200 includes adjusting parameters of a generative neural network by reducing an output of a loss function based on the additional modified embedding sequence.
[0109] In addition, the series of actions 1200 includes determining a plurality of ground-truth masks corresponding to objects of the training digital images. The series of actions 1200 also includes determining a cross-attention map generated by the generative neural network for the plurality of visual embeddings appended with the modified embedding sequence. The series of actions 1200 also includes adjusting parameters of the generative neural network based on a comparison between the plurality of ground-truth masks and the cross-attention map to reduce an output of the loss function.
[0110] In some embodiments, a series of actions 1200 includes determining a plurality of phrases corresponding to a plurality of objects by parsing a textual prompt used to generate or modify a digital image. The series of actions 1200 also includes generating, using one or more encoder neural networks, an embedding sequence comprising a plurality of object text embeddings representing a plurality of phrases corresponding to the plurality of objects. In addition, the series of actions 1200 includes generating, using one or more encoder neural networks, a plurality of visual embeddings representing a plurality of object images corresponding to the plurality of objects. The series of actions 1200 also includes determining a modified embedding sequence by replacing the plurality of object text embeddings with the plurality of visual embeddings in the embedding sequence. The series of actions 1200 also includes generating, using a generative neural network, a composite digital image based on the modified embedding sequence comprising the plurality of visual embeddings.
[0111] In one or more embodiments, a series of actions 1200 includes parsing a textual prompt by: determining a first phrase corresponding to a first object of a digital image and one or more visual attributes of the first object; and determining a second phrase corresponding to a second object of the digital image and one or more visual attributes of the second object. The series of actions 1200 also includes generating an embedding sequence, including generating a first object text embedding representing the first phrase and a second object text embedding representing the second phrase.
[0112] In some embodiments, the sequence of actions 1200 includes generating a plurality of visual embeddings by determining a first object image comprising a first example object, the first example object comprising one or more visual attributes of the first object. Additionally, the sequence of actions 1200 includes determining a second object image comprising a second example object, the second example object comprising one or more visual attributes of the second object. The sequence of actions 1200 also includes generating, using one or more encoder neural networks, a first visual embedding representing the first object image and a second visual embedding representing the second object image.
[0113] In one or more embodiments, a series of actions 1200 includes determining a modified embedding sequence by determining positions in the embedding sequence corresponding to the plurality of object text embeddings. The series of actions 1200 includes replacing the plurality of object text embeddings with the plurality of visual embeddings at the positions in the embedding sequence.
[0114] Embodiments of the present disclosure may include or utilize a special-purpose or general-purpose computer, including computer hardware, such as, for example, one or more processors and system memory, as discussed in more detail below. Embodiments within the scope of the present disclosure also include physical and other computer-readable media for carrying or storing computer-executable instructions and / or data structures. In particular, one or more processes described herein may be implemented at least in part as instructions embodied in a non-transitory computer-readable medium and executable by one or more computing devices (e.g., any media content access device described herein). In general, a processor (e.g., a microprocessor) receives instructions from a non-transitory computer-readable medium (e.g., a memory, etc.) and executes these instructions to perform one or more processes, including one or more processes described herein.
[0115] Computer-readable media can be any available medium that can be accessed by a general-purpose or special-purpose computer system. A computer-readable medium that stores computer-executable instructions is a non-transitory computer-readable storage medium (device). A computer-readable medium that carries computer-executable instructions is a transmission medium. Therefore, by way of example and not limitation, embodiments of the present disclosure may include at least two distinct types of computer-readable media: a non-transitory computer-readable storage medium (device) and a transmission medium.
[0116] Non-transitory computer-readable storage media (devices) include RAM, ROM, EEPROM, CD-ROM, solid-state drives (“SSD”) (e.g., RAM-based), flash memory, phase-change memory (“PCM”), other types of memory, other optical disk storage, magnetic disk storage or other magnetic storage devices, or any other medium that can be used to store desired program code means in the form of computer-executable instructions or data structures and that can be accessed by a general-purpose or special-purpose computer.
[0117] A "network" is defined as one or more data links that enable electronic data to be transmitted between computer systems and / or modules and / or other electronic devices. When information is transmitted or provided to a computer via a network or other communication connection (wired, wireless, or a combination of wired and wireless), the computer views the connection as a transmission medium. Transmission media may include networks and / or data links that can be used to carry desired program code components in the form of computer-executable instructions or data structures and that can be accessed by general-purpose or special-purpose computers. Combinations of the above should also be included within the scope of computer-readable media.
[0118] In addition, upon reaching various computer system components, program code components in the form of computer-executable instructions or data structures can be automatically transferred from the transmission medium to the non-transitory computer-readable storage medium (device) (or vice versa). For example, computer-executable instructions or data structures received over a network or data link can be buffered in RAM within a network interface module (e.g., a "NIC") and ultimately transferred to RAM and / or a less volatile computer storage medium (device) at the computer system. Therefore, it should be understood that non-transitory computer-readable storage media (devices) can be included in computer system components that also (or even primarily) utilize transmission media.
[0119] Computer executable instructions include, for example, instructions and data, which, when executed on a processor, cause a general-purpose computer, a special-purpose computer, or a dedicated processing device to perform a specific function or function group. In some embodiments, computer executable instructions are executed on a general-purpose computer to turn a general-purpose computer into a special-purpose computer that implements an element of the present disclosure. Computer executable instructions can be, for example, binary files, intermediate format instructions (such as assembly language), or even source code. Although subject matter has been described in a language specific to structural features and / or method actions, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the features or actions described above. On the contrary, the features and actions are disclosed as example forms for implementing claims.
[0120] Those skilled in the art will appreciate that the present disclosure can be practiced in a network computing environment with various types of computer system configurations, including personal computers, desktop computers, notebook computers, message processors, handheld devices, multiprocessor systems, microprocessor-based or programmable consumer electronics, network PCs, minicomputers, mainframe computers, mobile phones, PDAs, tablet computers, pagers, routers, switches, and the like. The present disclosure can also be practiced in a distributed system environment in which local and remote computer systems connected (by hardwired data links, wireless data links, or by a combination of hardwired and wireless data links) over a network all perform tasks. In a distributed system environment, program modules can be located in local and remote memory storage devices.
[0121] Embodiments of the present disclosure may also be implemented in a cloud computing environment. In this specification, "cloud computing" is defined as a model for enabling on-demand network access to a shared pool of configurable computing resources. For example, cloud computing can be adopted in the marketplace to provide ubiquitous and convenient on-demand access to a shared pool of configurable computing resources. This shared pool of configurable computing resources can be rapidly provisioned via virtualization, published with minimal management effort or service provider interaction, and scaled accordingly.
[0122] The cloud computing model can be composed of various characteristics, such as, for example, on-demand self-service, broad network access, resource pooling, rapid elasticity, and measured services. The cloud computing model can also expose various service models, such as, for example, software as a service ("SaaS"), platform as a service ("PaaS"), and infrastructure as a service ("IaaS"). The cloud computing model can also be deployed using different deployment models, such as private cloud, community cloud, public cloud, hybrid cloud, etc. In this specification and claims, a "cloud computing environment" is an environment in which cloud computing is employed.
[0123] Figure 13 A block diagram of an exemplary computing device 1300 is shown, which may be configured to perform one or more of the processes described above. It will be appreciated that one or more computing devices (such as computing device 1300) may implement Figure 1 (multiple) systems. Figure 13 As shown in FIG, computing device 1300 may include a processor 1302, a memory 1304, a storage device 1306, an I / O interface 1308, and a communication interface 1310, which may be communicatively coupled via a communication infrastructure 1312. In some embodiments, computing device 1300 may include a processor 1302, a memory 1304, a storage device 1306, an I / O interface 1308, and a communication interface 1310. Figure 13 The following will now be described in more detail. Figure 13 Components of computing device 1300 are shown in FIG.
[0124] In one or more embodiments, the processor 1302 includes hardware for executing instructions, such as instructions that constitute a computer program. By way of example and not limitation, to execute instructions for dynamically modifying a workflow, the processor 1302 may retrieve (or fetch) instructions from an internal register, an internal cache, a memory 1304, or a storage device 1306 and decode and execute them. The memory 1304 may be a volatile or non-volatile memory used to store data, metadata, and programs for execution by the processor. The storage device 1306 includes a memory, such as a hard disk, a flash drive, or other digital storage device, for storing data or instructions for executing the methods described herein.
[0125] The I / O interface 1308 allows a user to provide input to the computing device 1300, receive output from it, and otherwise transmit data to and receive data from the computing device 1300. The I / O interface 1308 may include a mouse, a keypad or keyboard, a touch screen, a camera, an optical scanner, a network interface, a modem, other known I / O devices, or a combination of these I / O interfaces. The I / O interface 1308 may include one or more devices for presenting output to the user, including but not limited to a graphics engine, a display (e.g., a display screen), one or more output drivers (e.g., a display driver), one or more audio speakers, and one or more audio drivers. In certain embodiments, the I / O interface 1308 is configured to provide graphical data to the display for presentation to the user. The graphical data may represent one or more graphical user interfaces and / or any other graphical content that may be used for a particular implementation.
[0126] The communication interface 1310 may include hardware, software, or both. In any case, the communication interface 1310 may provide one or more interfaces for communication between the computing device 1300 and one or more other computing devices or networks (such as, for example, packet-based communication). By way of example and not limitation, the communication interface 1310 may include a network interface controller (NIC) or network adapter for communicating with an Ethernet or other wired network, or a wireless NIC (WNIC) or wireless adapter for communicating with a wireless network (such as WI-FI).
[0127] In addition, the communication interface 1310 can facilitate communication with various types of wired or wireless networks. The communication interface 1310 can also facilitate communication using various communication protocols. The communication infrastructure 1312 can also include hardware, software, or both that couple the components of the computing device 1300 to each other. For example, the communication interface 1310 can use one or more networks and / or protocols to enable multiple computing devices connected by a particular infrastructure to communicate with each other to perform one or more aspects of the processes described herein. To illustrate, the digital content campaign management process can allow multiple devices (e.g., client devices and server devices) to exchange information using various communication networks and protocols to share information such as electronic messages, user interaction information, engagement metrics, or campaign management resources.
[0128] In the foregoing description, the present disclosure has been described with reference to specific exemplary embodiments thereof. Various embodiments and aspects of the present disclosure have been described with reference to the details discussed herein, and the accompanying drawings illustrate various embodiments. The foregoing description and drawings are illustrative of the present disclosure and should not be construed as limiting the present disclosure. Numerous specific details have been described to provide a thorough understanding of the various embodiments of the present disclosure.
[0129] The present disclosure may be embodied in other specific forms without departing from its spirit or essential characteristics. The described embodiments should be considered in all respects to be merely illustrative and not restrictive. For example, the methods described herein may be performed with fewer or more steps / actions, or the steps / actions may be performed in a different order. In addition, the steps / actions described herein may be repeated or performed in parallel with each other, or performed in parallel with different instances of the same or similar steps / actions. Therefore, the scope of this application is indicated by the appended claims rather than the foregoing description. All variations of the equivalent meaning and scope of the claims should be included within their scope.
Claims
1. A computer-implemented method comprising: generating, using one or more encoder neural networks, a sequence of embeddings comprising a prompt embedding representing a text prompt and an object text embedding representing a phrase indicating an object in the text prompt; generating, using the one or more encoder neural networks, a visual embedding representing an object image corresponding to the object; determining a modified embedding sequence by replacing the object text embedding with the visual embedding in the embedding sequence; as well as A synthetic digital image is generated from the modified embedding sequence including the visual embedding using a generative neural network.
2. The computer-implemented method of claim 1 , wherein generating the embedded sequence comprises: determining, from the text prompt, a plurality of phrases indicating a plurality of objects in the text prompt; generating the object text embedding representing the phrase indicating the object based on the plurality of phrases; as well as Additional object text embeddings representing additional phrases indicating additional objects are generated based on the plurality of phrases.
3. The computer-implemented method of claim 1 , wherein generating the visual embedding comprises generating the object image including an example object based on the phrase indicating the object.
4. The computer-implemented method of claim 3 , further comprising: generating an additional object image including an additional example object based on an additional phrase in the text prompt indicating an additional object; as well as An additional visual embedding representing the additional object image corresponding to the additional object is generated.
5. The computer-implemented method of claim 1 , wherein determining the modified embedding sequence comprises: Determining a position where the object text is embedded in the embedding sequence; removing the object text embedding from the embedding sequence; as well as The visual embedding is inserted into the embedding sequence at the position.
6. The computer-implemented method of claim 1 , wherein: Generating the embedding sequence includes: generating the object text embedding in a feature space; and Generating the visual embedding includes generating the visual embedding in the feature space of the object text embedding.
7. The computer-implemented method of claim 1 , further comprising: generating an additional embedding sequence for a caption corresponding to an additional digital image, the additional embedding sequence comprising a plurality of object text embeddings representing a plurality of phrases in the caption, the plurality of phrases referring to a plurality of objects in the additional digital image; generating a plurality of visual embeddings representing the plurality of objects in the additional digital image; determining an additional modified embedding sequence by replacing the plurality of object text embeddings with the plurality of visual embeddings; as well as Parameters of the generative neural network are adjusted by reducing an output of a loss function based on the additional modified embedding sequence.
8. The computer-implemented method of claim 7, wherein adjusting the parameters of the generative neural network comprises: determining a plurality of ground truth masks corresponding to the plurality of objects of the additional digital image; determining a localization loss based on a comparison between the plurality of ground-truth masks and a plurality of cross-attention maps corresponding to the plurality of visual embeddings using the loss function; as well as The parameters of the generative neural network are adjusted according to the localization loss.
9. A system comprising: one or more memory devices; as well as one or more processors configured to cause the system to: generating, using one or more encoder neural networks, an embedding sequence comprising a first object text embedding and a second object text embedding, the first object text embedding representing a first phrase indicating a first object in a text prompt, the second object text embedding representing a second phrase indicating a second object in the text prompt; generating, using the one or more encoder neural networks, a first visual embedding representing a first object image corresponding to the first object and a second visual embedding representing a second object image corresponding to the second object; determining a modified embedding sequence by replacing the first object text embedding with the first visual embedding and replacing the second object text embedding with the second visual embedding in the embedding sequence; as well as A synthetic digital image is generated from the modified embedding sequence including the first visual embedding and the second visual embedding using a generative neural network.
10. The system of claim 9, wherein the one or more processors are further configured to generate the embedding sequence by generating a prompt embedding representing the text prompt in a feature space corresponding to the first object text embedding and the second object text embedding using the one or more encoder neural networks.
11. The system of claim 9, wherein the one or more processors are further configured to generate the embedded sequence by: Parse the text prompt to: determining the first object and one or more visual attributes of the first object; and The second object and one or more visual attributes of the second object are determined.
12. The system of claim 11, wherein the one or more processors are further configured to: determining, based on the first phrase, the first object image including a first example object, the first example object including the one or more visual attributes of the first object; and The second object image is determined to include a second example object based on the second phrase, the second example object including the one or more visual attributes of the second object.
13. The system of claim 9, wherein the one or more processors are further configured to determine the modified embedding sequence by: determining a first position in the embedding sequence corresponding to the first object text embedding and a second position in the embedding sequence corresponding to the second object text embedding; removing the first object text embedding and the second object text embedding from the embedding sequence; as well as The first visual embedding is inserted at the first position and the second visual embedding is inserted at the second position.
14. The system of claim 9, wherein the one or more processors are further configured to generate the synthetic digital image by providing the modified embedding sequence with noisy image embeddings to the generative neural network.
15. The system of claim 9, wherein the one or more processors are further configured to: generating a plurality of visual embeddings from a plurality of object text embeddings representing phrases in captions of objects in training digital images; determining an additional modified embedding sequence by replacing the plurality of object text embeddings with the plurality of visual embeddings in corresponding embedding sequences; as well as Parameters of the generative neural network are adjusted by reducing an output of a loss function based on the additional modified embedding sequence.
16. The system of claim 15, wherein the one or more processors are further configured to adjust the parameters of the generative neural network, comprising: determining a plurality of ground truth masks corresponding to the objects of the training digital images; determining a cross-attention map generated by the generative neural network for the plurality of visual embeddings of the additional modified embedding sequence; as well as The parameters of the generative neural network are adjusted based on a comparison between the plurality of ground truth masks and the criss-cross attention map to reduce the output of the loss function.
17. A non-transitory computer-readable medium storing executable instructions that, when executed by at least one processing device, cause the at least one processing device to perform operations comprising: determining a plurality of phrases corresponding to a plurality of objects by parsing a textual prompt used to generate or modify a digital image; generating, using one or more encoder neural networks, an embedding sequence comprising a plurality of object text embeddings representing the plurality of phrases corresponding to the plurality of objects; generating, using the one or more encoder neural networks, a plurality of visual embeddings representing a plurality of object images corresponding to the plurality of objects; determining a modified embedding sequence by replacing the plurality of object text embeddings with the plurality of visual embeddings in the embedding sequence; as well as A synthetic digital image is generated from the modified embedding sequence including the plurality of visual embeddings using a generative neural network.
18. The non-transitory computer readable medium of claim 17, wherein: Parsing the text prompt includes: determining a first phrase corresponding to a first object of the digital image and one or more visual attributes of the first object; and determining a second phrase corresponding to a second object of the digital image and one or more visual attributes of the second object; and Generating the embedding sequence includes generating a first object text embedding representing the first phrase and a second object text embedding representing the second phrase.
19. The non-transitory computer-readable medium of claim 18, wherein generating the plurality of visual embeddings comprises: determining a first object image comprising a first example object including the one or more visual attributes of the first object; determining a second object image comprising a second example object including the one or more visual attributes of the second object; as well as The one or more encoder neural networks are utilized to generate a first visual embedding representing the first object image and a second visual embedding representing the second object image.
20. The non-transitory computer-readable medium of claim 18, wherein determining the modified embedding sequence comprises: determining positions in the embedding sequence corresponding to the plurality of object text embeddings; as well as The plurality of object text embeddings are replaced with the plurality of visual embeddings at the positions in the embedding sequence.