Real-time text deentanglement-based real image editing

By combining the inversion model and the image generation model, the problems of low efficiency and poor quality of image editing in the existing technology are solved, and accurate image editing can be generated in a small number of steps while keeping other elements of the input image unchanged.

CN120655778APending Publication Date: 2025-09-16ADOBE INC
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510158341.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Priority Date
2024-10-24
Filing Date
2025-02-13
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

Existing image editing systems are inefficient and suffer from poor image quality when generating modified images. They also struggle to accurately disentangle elements in the input image. Conventional techniques often require multiple diffusion steps and are prone to introducing artifacts.

Method used

A combination of an inversion model and an image generation model is adopted. The inversion model generates an intermediate output to keep other elements of the input image unchanged, and the image generation model generates a synthetic image based on the intermediate output and modification prompts, achieving accurate image editing in a small number of steps.

Benefits of technology

The efficiency and quality of image editing are improved, and accurate synthetic images can be generated in a small number of steps, keeping other elements of the input image unchanged and avoiding the appearance of artifacts.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120655778A_ABST
    Figure CN120655778A_ABST
Patent Text Reader

Abstract

The embodiment of the invention relates to real-time text deentanglement-based real image editing. A method, apparatus, non-transitory computer readable medium and system for image processing includes obtaining an input image depicting a first element, a textual description of the input image, and a modification prompt describing a second element, the second element being different from the first element; generating an intermediate output based on the input image and the textual description, where the intermediate output represents the first element; and generating a composite image based on the intermediate output and the modification cue, where the composite image replaces the first element from the input image with a second element from the modification cue.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] CROSS-REFERENCE TO RELATED APPLICATIONS

[0002] This application claims priority to U.S. Provisional Application No. 63 / 564,930, filed in the U.S. Patent and Trademark Office on March 13, 2024, and corresponding U.S. Non-Provisional Application No. 18 / 925,290, filed on October 24, 2024, the disclosures of which are incorporated herein by reference in their entireties. Technical Field

[0003] The following content relates generally to image processing, and more specifically to image editing using machine learning models. Image processing refers to the use of computers to edit images using algorithms or processing networks. In some cases, image processing software can be used for various image processing tasks, such as image restoration, image detection, image generation, image compositing, and image editing. For example, image editing involves using machine learning models to conditionally edit an input image to generate an output image. Background Art

[0004] In the field of image editing, an input image and a text prompt are provided to a machine learning model to generate a modified image. In some cases, the text prompt includes user instructions that describe changes to image elements from the input image. In some cases, the modified image depicts the changes to the image elements described by the text prompt. However, in some cases, the modified image may depict undesirable changes to other image elements depicted in the input image. Summary of the Invention

[0005] Aspects of the present disclosure provide a method and system for image editing. In one aspect, the system receives an input image and a modification hint, the input image depicting an image element, the modification hint depicting a modification to the image element; and generates a composite image, the composite image depicting the modification. According to some aspects, the system includes an inversion model that is trained to perform image inversion and generate intermediate features (or latent features), the intermediate features representing an original image including the image element to be modified. In some aspects, the system includes an image generation model that is configured to generate a composite image based on the intermediate features and the modification hint, the modification hint describing the change from the image element to a different image element. The intermediate features generated by the inversion model are provided to the image generation model to ensure that the target image element is modified while maintaining the remaining image elements of the input image.

[0006] A method, apparatus, non-transitory computer-readable medium, and system for image processing are described. One or more aspects of the method, apparatus, non-transitory computer-readable medium, and system include obtaining an input image and a modification cue, wherein the input image depicts a first element and the modification cue describes a second element, the second element being different from the first element; generating an intermediate output based on the input image using an inversion model, wherein the intermediate output includes image features representing the image; and generating a composite image based on the intermediate output and the modification cue using an image generation model, wherein the composite image replaces a first element from the input image with the second element from the modification cue.

[0007] A method, apparatus, non-transitory computer-readable medium, and system for image processing are described. One or more aspects of the method, apparatus, non-transitory computer-readable medium, and system include: obtaining a training set comprising training images and training descriptions of the images; using an inversion model to generate an intermediate output based on the training images and the text descriptions; using an image generation model to generate a reconstructed image based on the intermediate output and the text descriptions; and training the inversion model based on the training images and the reconstructed image.

[0008] An apparatus and system for image processing are described. One or more aspects of the apparatus and system include at least one processor; at least one memory storing instructions executable by the processor; an inversion model comprising parameters stored in the at least one memory and trained to generate an intermediate output based on an input image and a textual description of the input image; and an image generation model comprising parameters stored in the at least one memory and trained to generate a modified image based on the intermediate output and a modification prompt, wherein the modified image retains elements of the input image and includes modifications based on the modification prompt. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] Figure 1 An example of an image processing system according to aspects of the present disclosure is shown.

[0010] Figure 2 An example of a method for editing an image according to aspects of the present disclosure is shown.

[0011] Figure 3 Examples of text-based image editing according to aspects of the present disclosure are shown.

[0012] Figure 4 An example of image editing based on image interpolation according to aspects of the present disclosure is shown.

[0013] Figure 5 An example of an image processing apparatus according to aspects of the present disclosure is shown.

[0014] Figure 6 Examples of machine learning models according to aspects of the present disclosure are shown.

[0015] Figure 7 Examples of inversion models according to aspects of the present disclosure are shown.

[0016] Figure 8 Examples of image generation models according to aspects of the present disclosure are shown.

[0017] Figure 9 An example of a U-Net architecture according to aspects of the present disclosure is shown.

[0018] Figure 10 An example of a diffusion process according to aspects of the present disclosure is shown.

[0019] Figure 11 An example of a method for generating a modified image according to aspects of the present disclosure is shown.

[0020] Figure 12 Examples of methods for training machine learning models according to aspects of the present disclosure are shown.

[0021] Figure 13 An example of a method for training an inversion model according to aspects of the present disclosure is shown.

[0022] Figure 14 An example of a flowchart depicting an algorithm as a step-by-step procedure in an example operational implementation that may be performed for training a machine learning model in accordance with aspects of the present disclosure is shown.

[0023] Figure 15 An example of a method for training a diffusion model according to aspects of the present disclosure is shown.

[0024] Figure 16 An example of a computing device according to aspects of the present disclosure is shown. DETAILED DESCRIPTION

[0025] Various aspects of the present disclosure relate to image editing using generative machine learning. Some embodiments of the present disclosure relate to an image generation system that accurately and efficiently generates an image depicting modifications to an image element from an original image. In some aspects, the system includes an inversion model that is trained to perform image inversion and generate intermediate features (or latent features), the intermediate features representing the original image including the image element to be modified. In some aspects, the system includes an image generation model that is configured to generate a modified image based on the intermediate features and a modification hint that describes the change from the image element to a different image element. The intermediate features generated by the inversion model are provided to the image generation model to ensure that the target image element is modified while maintaining the rest of the image elements of the input image.

[0026] In the field of image editing, machine learning systems are used to edit synthetic images generated by image generation models (e.g., diffusion models). In some cases, a text prompt can correspond to multiple images, for example, the prompt "cat" can refer to different images depicting various cats. Therefore, image inversion involves mapping the real image onto the reverse diffusion trajectory. In some cases, image inversion involves finding the noise pattern used during the forward diffusion process that represents or is identical to the input image.

[0027] However, conventional diffusion techniques require many diffusion steps (e.g., 50 or more steps) to restore the true image and an additional 20-50 steps to generate a new edit of the input image. Consequently, the efficiency of generating modified images in image editing is reduced. Furthermore, conventional image editing systems encounter challenges such as low image quality and slow processing speeds.

[0028] In some cases, conventional image editing systems are unable to disentangle elements in the input image. For example, disentanglement refers to edits that change one attribute of the input image instead of one or more attributes. Some systems use techniques such as frozen attention maps to attempt to address this issue. However, attention control methods impose overly restrictive influences on the generation process, resulting in insufficient changes in the image space or introducing artifacts. Other techniques involve expensive optimization steps or fine-tune the synthetic paired edit data from attribute blends at edit time.

[0029] Therefore, the present disclosure provides a system and method for improving conventional image generation systems by accurately and efficiently generating a synthetic image that depicts the modification described by the modification hint. This is achieved using a system that includes an inversion model and an image generation model, the inversion model being trained to generate an intermediate output (i.e., image features), and the image generation model being configured to generate a synthetic image based on the intermediate output.

[0030] According to embodiments of the present disclosure, a machine learning system is trained to accurately reconstruct real images and efficiently generate image edits. For example, given an input image and a textual prompt (or modification prompt) describing a modification to an element of the input image, the system is able to accurately and efficiently generate a modified image (or a composite image) that depicts the modification described by the textual prompt.

[0031] According to some aspects, the inversion model generates an intermediate output based on an input image. In some cases, a previously reconstructed image is also used to generate the intermediate output. In some embodiments, the inversion model is trained to iteratively generate a reconstruction of a previously reconstructed image so that the subsequently reconstructed image becomes increasingly visually similar or identical to the input image within 2-4 steps. For example, an image generator iteratively receives the intermediate output and the previously reconstructed image to generate the next reconstructed image (or the final reconstructed image at the final step). In some embodiments, the intermediate output is used as input to an image generation model to generate a modified image at each step. In some cases, the inversion model includes a diffusion-based inversion network that is trained to generate an intermediate output based on an input image and a text description.

[0032] According to some aspects, the image generation model receives the intermediate output (generated from the inversion model) and the modified text description (or modification prompt) to generate a modified image. In some embodiments, the caption generation model is configured to generate a text description based on the input image. For example, the modified text description includes the modification prompt. By using the modified text description, the image generation model is able to disentangle the attributes in the input image. By changing one attribute in the text description, the corresponding element in the input image can be changed without changing other elements in the input image. Therefore, the other elements of the input image are retained and the image quality can be maintained. Therefore, the text prompt can be edited to obtain the desired disentangled, modified image.

[0033] refer to Figure 1 and Figure 16 An example system of the present disclosure in image processing is provided. Figure 2-Figure 4 An example application of the present disclosure in image processing is provided. Figure 5-Figure 9 Provides detailed information about the architecture of the image processing device. Figure 10-11 Provides examples of procedures used for image processing. Figure 12-15 A description of an example training procedure is provided.

[0034] Therefore, the present disclosure provides a system and method for improving conventional image editing systems by generating synthetic images that depict accurate edits in fewer time steps. For example, by training the system using modified text descriptions that include modification hints, the system is able to disentangle attributes in an input image. In some cases, elements described by modification hints depicted in the input image can be edited without requiring other elements in the input image to be edited. By using intermediate outputs generated by an inversion model, the system is able to generate modified images in fewer steps.

[0035] Image Editing

[0036] exist Figures 1-4 and Figure 10-11 In the present invention, a method, apparatus, non-transitory computer-readable medium, and system for image processing include: obtaining an input image depicting a first element, a text description of the input image, and a modification prompt describing a second element, the second element being different from the first element; using an inversion model to generate an intermediate output based on the input image and the text description, wherein the intermediate output represents the first element; and using an image generation model to generate a composite image based on the intermediate output and the modification prompt, wherein the composite image replaces the first element from the input image with the second element from the modification prompt.

[0037] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include generating a text description based on the input image. In some aspects, the modification prompt includes editing the text description.

[0038] Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include iteratively alternating between generating successive intermediate outputs using the inversion model and the image generation model. Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include generating a reconstructed image based on the intermediate output, wherein the reconstructed image depicts the first element; and generating a subsequent intermediate output based on the reconstructed image, wherein the composite image is based on the subsequent intermediate output.

[0039] Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include obtaining a noisy input; and denoising the noisy input based on the intermediate output. In some aspects, the inversion model is trained using a training set comprising training images and training descriptions of the training images.

[0040] According to some embodiments, methods, apparatus, non-transitory computer-readable media, and systems for image processing include: obtaining an input image, a text description of the input image, and a modification prompt; using an inversion model to generate an intermediate output based on the input image and the text description; and using an image generation model to generate a modified image based on the intermediate output and the modification prompt, wherein the modified image retains elements of the input image and includes modifications based on the modification prompt.

[0041] In some aspects, the modification prompt includes an edit to the text description. In some aspects, the modification prompt includes a description of the change to the input image. Some examples of the method, apparatus, non-transitory computer-readable medium, and system further include iteratively alternating between generating successive intermediate outputs and corresponding modified images.

[0042] Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include generating a reconstructed image based on the intermediate output and the text description using the image generation model. Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include generating a text description based on the input image.

[0043] Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include adding the intermediate output to an input image to obtain a noisy image, wherein the image generation model receives the noisy image as input. Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include performing a single-pass diffusion process.

[0044] Figure 1 An example of an image processing system according to aspects of the present disclosure is shown. The example shown includes a user 100, a user device 105, an image processing apparatus 110, a cloud 115, and a database 120. The image processing apparatus 110 is a reference Figure 8 Examples of elements described or including references Figure 8 Describes various aspects of an element.

[0045] refer to Figure 1, the user 100 provides an input image and a modification prompt to the image processing apparatus 110 via the user device 105 and the cloud 115. For example, the input image depicts a pharaoh. In some cases, the modification prompt states "Man2Fox". For example, in some cases, the modification prompt indicates a change of an element from a man (or a man's face) to a fox (or a fox's face). In some cases, the modification prompt includes a change of an element within the input image. According to some embodiments, the image processing apparatus 110 includes an inversion model that is trained to generate an intermediate output based on the input image and the original text prompt. In some embodiments, the image processing apparatus includes an image generation model that is configured to generate a modified image based on the intermediate output and the modification prompt. In some cases, the image processing apparatus 110 generates a modified image depicting the change and displays the modified image to the user 100 via the user device 105 and the cloud 115.

[0046] User device 105 can be a personal computer, laptop computer, mainframe computer, PDA, personal assistant, mobile device, or any other suitable processing device. In some examples, user device 105 includes software that incorporates an image processing application. In some examples, the image processing application on user device 105 can include the functionality of image processing device 110.

[0047] The user interface can enable the user 100 to interact with the user device 105. In some embodiments, the user interface can include an audio device such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that interfaces with the user interface directly or via an I / O controller module). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the user interface can be represented in code, where the code is sent to the user device 105 and rendered locally by the browser. Figure 2 Further description.

[0048] The image processing device 110 is a reference Figure 5 Examples of corresponding elements described or including references Figure 5 According to some aspects, the image processing device 110 includes a computer-implemented network including a machine learning model, an inversion model, an image generation model, and a caption generation model. The image processing device 110 also includes a processor unit, a memory unit, an I / O module, and a training component. In some embodiments, the image processing device 110 also includes a computer-implemented network including a machine learning model, an inversion model, an image generation model, and a caption generation model. Figure 16 Additionally, the image processing apparatus 110 communicates with the user device 105 and the database 120 via the cloud 115. Figure 2 Additional details regarding the operation of the image processing apparatus 110 are provided.

[0049] In some cases, the image processing device 110 is implemented on a server. The server provides one or more functions for the user of one or more network links in various networks. In some cases, the server includes a single microprocessor board, which includes a microprocessor responsible for controlling the various aspects of the server. In some cases, the server uses a microprocessor and protocol to exchange data with other devices / users on one or more networks via Hypertext Transfer Protocol (HTTP) and Simple Mail Transfer Protocol (SMTP), but other protocols such as File Transfer Protocol (FTP) and Simple Network Management Protocol (SNMP) can also be used. In some cases, the server is configured to send and receive files (for example, for displaying web pages) in Hypertext Markup Language (HTML) format. In various embodiments, the server includes a general-purpose computing device, a personal computer, a notebook computer, a mainframe computer, a supercomputer, or any other suitable processing device.

[0050] Cloud 115 is a computer network configured to provide on-demand availability of computer system resources (such as data storage and computing power). In some examples, cloud 115 provides resources without active management by a user (e.g., user 100). The term cloud is sometimes used to describe a data center that is available to many users via the Internet. Some large cloud networks have functions distributed across multiple locations from a central server. A server is designated as an edge server if it has a direct or close connection to a user. In some cases, cloud 115 is limited to a single organization. In other examples, cloud 115 is available to many organizations. In one example, cloud 115 includes a multi-layer communication network with multiple edge routers and core routers. In another example, cloud 115 is based on a collection of local switches in a single physical location.

[0051] According to some aspects, database 120 stores training data (or training sets), which include training images and training descriptions of the training images. Database 120 is an organized collection of data. For example, database 120 stores data in a specified format called a schema. Database 120 can be structured as a single database, a distributed database, multiple distributed databases, or an emergency backup database. In some cases, a database controller can manage the storage and processing of data in database 120. In some cases, a user (e.g., user 100) interacts with the database controller. In other cases, the database controller can operate automatically without user interaction.

[0052] Figure 2An example of a method 200 for editing an image according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system comprising a processor that executes a code set to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or performed in conjunction with other operations.

[0053] At operation 205, the system provides the input image and modification prompts. In some cases, the operation of this step refers to the operation of Figure 1 The user described may be referred to as Figure 1 For example, the user inputs an image depicting a pharaoh. For example, the user provides the image processing apparatus with a modification prompt "Man2Fox" via a user interface, which is provided by the image processing apparatus on a user device (e.g., a reference image). Figure 1 In some cases, the system can generate a text description based on the input image and provide the text description to the user to provide the target modification. In some cases, the modified text description is used as a modification prompt.

[0054] At operation 210, the system generates a conditional guidance feature. In some cases, the operation of this step is referred to as Figure 1 and Figure 5 The image processing device may be configured as shown in FIG. Figure 1 and Figure 5 The image processing apparatus described herein is used to perform the conditional guided feature. For example, the conditional guided feature includes an intermediate output. In some cases, the system includes an inversion model that is trained to generate the intermediate output based on the input image and a previous reconstruction of the input image (or noise at the first step). In some cases, the intermediate output is generated based on a textual description of the input image. More details regarding the intermediate output are described with reference to the figures.

[0055] At operation 215, the system initializes the noise input. In some cases, the operation of this step is referred to as Figure 1 and Figure 5 The image processing device may be configured as shown in FIG. Figure 1 and Figure 5 The image processing apparatus is described. In some cases, a noise input comprising random noise is initialized. The noise input can be in a latent space. By initializing the image generation model using random noise, different variations of the synthesized image can be generated, each of which includes content described by the text condition (e.g., a text prompt).

[0056] At operation 220, the system generates media content. In some cases, the operation of this step is referred to as referring to Figure 1 and Figure 5 The image processing device may be configured as shown in FIG. Figure 1 and Figure 5 The image processing apparatus may be used to perform the above-described processing. For example, in some cases, the media content includes a composite image or a modified image. For example, in some cases, the media content includes a composite image or a modified image. For example, the composite image includes image pixels generated by an image generation model. For example, the modified image includes image pixels from an input image and image pixels generated by the image generation model. In some cases, the image generation model receives intermediate outputs and modification cues to generate the composite image.

[0057] Figure 3 An example of text-based image editing according to aspects of the present disclosure is shown. The illustrated example includes an image editing system 300, a modification prompt 305, an input image 310, a machine learning model 315, a modified image 320, and a conventional output image 325.

[0058] refer to Figure 3 , the machine learning model 315 receives the input image 310 and the modification prompt 305 to generate the modified image 320. For example, the modification prompt 305 states "Cat2Tiger", instructing the machine learning model 315 to change the object "cat" depicted in the input image 310 to a different object "tiger". According to some embodiments, the machine learning model 315 can generate the modified image 320 using the novel real-time text disentanglement-based real-image editing described in the present disclosure. In some cases, the machine learning model 315 can process photorealistic and artistic images, control the intensity of the edits, and perform attribute editing with large structures.

[0059] In some cases, a conventional image generation system may generate a conventional output image 325 based on the modification prompt 305 and the input image 310. For example, the conventional output image 325 depicts a tiger with artifacts (such as overlapping facial elements). Therefore, the conventional system is unable to accurately generate an output image based on a given input.

[0060] According to some embodiments, an input image is provided to a caption generation model to generate a text description describing the input image. For example, in some cases, the text description states "a young man with black hair, wearing a brown windbreaker (hat) and a gray T-shirt, standing in front of subtropical flowers (in heavy snow)." He looks directly at the camera, giving a sense of focus and determination. The coat is open, revealing the man's attire underneath. The entire scene is well-lit, and the man is the main subject of the image. In some cases, a modification hint is used to replace italic text in the text description to generate one or more modified images.

[0061] Conventional techniques in diffusion models include attention-based image editing to maintain structural similarity between the source and target images. For example, conventional techniques freeze the self-attention layers and cross-attention layers in the U-Net of the diffusion model. However, the effectiveness of these methods is related to the number of time steps over which attention control is applied. In some cases, applying attention control for many time steps results in a target image that is identical to the source image but lacks the target attributes described by the textual cues. On the other hand, applying attention control over a small number of time steps results in a target image that contains the desired attributes but deviates significantly from the source image. Therefore, conventional techniques need to calculate the optimal number of steps for applying attention control to tune based on the specific situation to achieve a balance between structure preservation and editability.

[0062] In the field of few-step diffusion models, the range of tuning this parameter is constrained due to the limited number of time steps (e.g., 1-4 time steps). In some cases, when the edit requires a large amount of structural modification (e.g., cat to tiger transformation), conventional models (e.g., one-step and four-step diffusion models) are unable to generate satisfactory edited images.

[0063] In some cases, the sampling trajectory in the diffusion model is subject to the initial noise x T , the injected noise ∈ t and the influence of text condition c. In some cases, the initial noise x T and the injected noise ∈ t Random sampling from a Gaussian distribution with spatial dimensions, which affects the image layout. Conventional techniques suggest that only the initial noise x is frozen T and the injected noise ∈ t Not sufficient to preserve the structure of the modified image.

[0064] Therefore, when the text prompt is very detailed and covers semantic information across a variety of attributes, modifying a single attribute in the text prompt results in a small change in the text embedding. Therefore, the two sampled trajectories can remain close enough, which indicates that the modified image and input can be almost identical except for the modified attributes. In some cases, a lengthy text prompt may not be provided. Therefore, a pre-trained language model can be used to expand short text prompts into long descriptions. For example, the language model is configured to "Please describe the image of {short caption} in detail", where the language model generates a long and detailed text description consisting of 50 to 100 words. The attributes can then be modified based on the text description, for example, changing the corresponding instance of "cat" to "tiger". By using the same random seed (the same random seed means the same initial noise x T and the injected noise ∈ t ), the source text prompt and the target text prompt can generate substantially the same image, which differs in the modified properties.

[0065] In some cases, in a single diffusion model, attention control techniques can exert an overly restrictive influence on the generation process, resulting in insufficient changes in image space (e.g., horse to unicorn) or the introduction of artifacts (e.g., fox to dog). In some cases, particularly where edits require significant structural modifications, attention control can result in insufficient preservation of structure or the appearance of artifacts exhibited by conventional output image 325.

[0066] Figure 4 An example of image editing based on image interpolation according to aspects of the present disclosure is shown. The shown example includes an image generation system 400, a first input image 405, a second input image 410, a machine learning model 415, and a modified image 420.

[0067] refer to Figure 4 , first input image 405 and second input image 410 are provided to machine learning model 415 to generate modified image 420, which includes multiple images depicting the transformation from first input image 405 to second input image 410. In some cases, for example, modified image 420 depicts edits to one or more image elements from first input image 405 or second input image 410.

[0068] System Architecture

[0069] exist Figure 5-Figure 9 and Figure 16

[0014] An apparatus and system for image processing are described. One or more aspects of the apparatus and system include: at least one processor; at least one memory storing instructions executable by the processor; an inversion model including parameters stored in the at least one memory and trained to generate an intermediate output based on an input image and a textual description of the input image; and an image generation model including parameters stored in the at least one memory and trained to generate a modified image based on the intermediate output and a modification prompt, wherein the modified image retains elements of the input image and includes modifications based on the modification prompt.

[0070] In some aspects, the image generation model includes a one-way or short-range diffusion model. In some aspects, the inversion model has the same architecture as the image generation model. Some examples of the apparatus and system also include a caption generation model configured to generate a text description based on an input image.

[0071] Some examples of the apparatus and system further include a user interface comprising an input image display element, a modified image display element, a text description field, and a modification prompt field. In some aspects, the user interface further includes a selection element indicating a balance between the text description and the modification prompt.

[0072] Figure 5 An example of an image processing apparatus 500 according to aspects of the present disclosure is shown. The example shown includes the image processing apparatus 500, a processor unit 505, an I / O module 510, a memory unit 515, and a training component 535. In one aspect, the memory unit 515 includes an inversion model 520, an image generation model 525, and a caption generation model 530. In some cases, the processor unit 505 may be referred to as a processing device. In some cases, the memory unit 515 may be referred to as a memory component.

[0073] According to some embodiments of the present disclosure, the image processing device 500 includes a computer-implemented artificial neural network (ANN). An ANN is a hardware or software component that includes several connected nodes (e.g., artificial neurons), which loosely correspond to neurons in the human brain. Each connection or edge transmits a signal from one node to another (just like a physical synapse in the brain). When a node receives a signal, the node processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between nodes include real numbers and the output of each node is calculated by the sum of the inputs. In some examples, the nodes can use other mathematical algorithms (e.g., selecting the maximum value from the input as the output) or any other suitable algorithm to determine the output to activate the node. Each node and edge is associated with one or more node weights, which determine how the signal is processed and transmitted. The image processing device 500 is a reference Figure 1 Examples of corresponding elements described or including references Figure 1 Describes aspects of the corresponding element.

[0074] The processor unit 505 is an intelligent hardware device (e.g., a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), a programmable logic device, a separate gate or transistor logic component, a separate hardware component, or any combination thereof). In some cases, the processor unit 505 is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into the processor. In some cases, the processor unit 505 is configured to execute computer-readable instructions stored in the memory to perform various functions. In some embodiments, the processor unit 505 includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing.

[0075] The I / O module 510 (e.g., input / output interface) may include an I / O controller. The I / O controller may manage input and output signals for the device. The I / O controller may also manage peripheral devices that are not integrated into the device. In some cases, the I / O controller may represent a physical connection or port to an external peripheral device. In some cases, the I / O controller may utilize an operating system, such as or other known operating systems. In other cases, an I / O controller may represent or interact with a modem, keyboard, mouse, touch screen, or similar device. In some cases, an I / O controller may be implemented as part of a processor. In some cases, a user may interact with a device via the I / O controller or via hardware components controlled by the I / O controller.

[0076] In some examples, the I / O module 510 includes a user interface. The user interface can enable a user to interact with the device. In some embodiments, the user interface can include an audio device such as an external speaker system, an external display device such as a display screen, or an input device (e.g., a remote control device that interfaces with the user interface directly or via an I / O controller module). In some cases, the user interface can be a graphical user interface (GUI). In some examples, the communication interface operates at the boundary between the communication entity and the channel and can also record and process communications. A communication interface is provided herein to enable a processing system coupled to a transceiver (e.g., a transmitter and / or receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna. The I / O module 510 is a reference Figure 16 Examples of I / O interfaces described or included with reference Figure 16 Describes various aspects of the I / O interface.

[0077] Examples of memory unit 515 include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory unit 515 include solid-state memory and a hard disk drive. In some examples, memory unit 515 is used to store computer-readable, computer-executable software, which includes instructions that, when executed, cause the processor to perform the various functions described herein.

[0078] In some cases, memory unit 515 includes, among other things, a basic input / output system (BIOS) that controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller operates the memory units. For example, a memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within memory unit 515 store information in the form of logical states.

[0079] In one aspect, the memory unit 515 includes a machine learning model 520, an inversion model 520, an image generation model 525, and a caption generation model 530. The memory unit 515 is a reference Figure 16 Examples of memory subsystems described or including reference Figure 16Aspects of the memory subsystem are described.

[0080] In some cases, a machine learning model is a computational algorithm, model, or system designed to recognize patterns, make predictions, or perform a specific task (e.g., image processing) without being explicitly programmed. According to some aspects, the machine learning model is implemented as software stored in memory unit 515 and can be executed by processor unit 505 as firmware, as one or more hardware circuits, or as a combination thereof.

[0081] According to some embodiments of the present disclosure, the machine learning model includes an ANN, which is a hardware or software component that includes several connected nodes (e.g., artificial neurons), which loosely correspond to neurons in the human brain. Each connection or edge transmits a signal from one node to another (just like a physical synapse in the brain). When a node receives a signal, the node processes the signal and then transmits the processed signal to other connected nodes. In some cases, the signals between nodes include real numbers and the output of each node is calculated by the sum of the inputs. In some examples, the nodes can use other mathematical algorithms (e.g., selecting the maximum value from the input as the output) or any other suitable algorithm to determine the output to activate the node. Each node and edge is associated with one or more node weights, which determine how the signal is processed and transmitted.

[0082] During the training process, one or more node weights are adjusted to improve the accuracy of the results (e.g., by minimizing a loss function that corresponds in some way to the difference between the current result and the target result). The weights of the edges increase or decrease the strength of the signal transmitted between the nodes. In some cases, the nodes have a threshold below which the signal is not transmitted at all. In some examples, the nodes are grouped into layers. Different layers perform different transformations on their respective inputs. The initial layer is called the input layer and the last layer is called the output layer. In some cases, the signal passes through certain layers multiple times.

[0083] According to some embodiments, the machine learning model includes a computer-implemented convolutional neural network (CNN). CNN is a class of neural networks commonly used in computer vision or image classification systems. In some cases, CNN can enable digital images to be processed with minimal preprocessing. CNN is characterized by using convolutional (or cross-correlation) hidden layers. These layers apply convolution operations to the input before sending the results through signals to the next layer. Each convolution node can process data for a limited input field (e.g., a receptive field). During the forward pass of the CNN, the filter at each layer can be convolved across the input volume, thereby calculating the dot product between the filter and the input. During the training process, the filter can be modified so that the filter is activated when the filter detects specific features within the input.

[0084] In one aspect, a machine learning model includes machine learning parameters. Machine learning parameters, also known as model parameters or weights, are variables that provide the behavior and characteristics of a machine learning model. Machine learning parameters can be learned or estimated from training data and used to make predictions or perform tasks based on the patterns and relationships learned in the data.

[0085] Machine learning parameters are adjusted during the training process to minimize a loss function or maximize a performance metric. The goal of the training process is to find optimal values ​​for the parameters that enable the machine learning model to make accurate predictions or perform well on a given task.

[0086] For example, during the training process, the algorithm adjusts the machine learning parameters according to an optimization technique such as gradient descent, stochastic gradient descent, or other optimization algorithms to minimize the error or loss between the predicted output and the actual target. Once the machine learning parameters are learned from the training data, the machine learning parameters can be used to make predictions on new, unseen data.

[0087] According to some embodiments, the machine learning model includes a computer-implemented recurrent neural network (RNN). An RNN is a type of ANN in which the connections between nodes form a directed graph along an ordered (e.g., time) sequence. This enables the RNN to model temporal dynamic behavior, such as predicting which element in a sequence will appear next. Therefore, RNNs are suitable for tasks involving ordered sequences, such as text recognition (the ordering of words in a sentence). In some cases, the RNN includes one or more finite impulse recurrent networks (characterized by nodes forming a directed acyclic graph), one or more infinite impulse recurrent networks (characterized by nodes forming a directed cyclic graph), or a combination thereof.

[0088] According to some embodiments, the machine learning model includes a transformer (or transformer model, or transformer network), wherein the transformer is a type of neural network model for natural language processing tasks. The transformer network uses an encoder and a decoder to transform one sequence into another sequence. The encoder and decoder include modules that can be stacked on top of each other multiple times. The module includes multi-head attention and feedforward layers. The input and output (target sentence) are first embedded in an n-dimensional space. The position encoding of different words (for example, giving a relative position for each word / part in the sequence, because the sequence is related to the order of elements) is added to the embedding representation (n-dimensional vector) of each word. In some examples, the transformer network includes an attention mechanism, wherein the attention looks at the input sequence and decides at each step which other parts of the sequence are important. The attention mechanism involves queries, keys, and values, represented by Q, K, and V, respectively. Q is a matrix containing the query (a vector representation of a word in the sequence), K is the key (a vector representation of the word in the sequence), and V is the value, which is in turn a vector representation of the word in the sequence. For the encoder and decoder, multi-head attention module, V consists of the same word sequence as Q. However, for attention modules that consider encoder and decoder sequences, the sequences represented by V and Q are different. In some cases, the values ​​in V are multiplied and summed with some attention weights a.

[0089] In the field of machine learning, an attention mechanism (e.g., implemented in one or more ANNs) is a method of placing different levels of importance on different input elements. Calculating attention can involve three basic steps. First, the similarity between the query and the key vector obtained from the input is calculated to generate attention weights. Similarity functions used for this process can include dot products, concatenation, detectors, etc. Next, a softmax function is used to normalize the attention weights. Finally, the attention weights are weighted together using the corresponding values. In the context of attention networks, keys and values ​​are vectors or matrices used to represent the input data. Keys are used to determine which parts of the input the attention mechanism should focus on, while values ​​are used to represent the actual data being processed.

[0090] Attention mechanisms are a key component in some ANN architectures, particularly those used in natural language processing (NLP) and sequence-to-sequence tasks, that enable ANNs to focus on different parts of an input sequence when making predictions or generating outputs. Some sequence models, such as RNNs, process the input sequence sequentially, maintaining an internal hidden state to capture information from previous steps. However, in some cases, this sequential processing makes it difficult to capture long-range dependencies or focus on specific parts of the input sequence.

[0091] The attention mechanism addresses these difficulties by enabling the ANN to selectively focus on different parts of the input sequence, assigning different levels of importance or attention to each part. The attention mechanism achieves selective attention by considering the relevance of each input element relative to the current state of the ANN.

[0092] The term "self-attention" refers to a machine learning model in which representations of the input interact with each other to determine attention weights for the input. Because the attention weights are determined at least in part by the input, self-attention can be distinguished from other attention models.

[0093] According to some aspects, the machine learning model receives an input image, a text description of the input image, and a modification prompt. The machine learning model is a reference Figure 3 and Figure 4 Examples of corresponding elements described or including references Figure 3 and Figure 4 Describes aspects of the corresponding element.

[0094] According to some aspects, the inversion model 520 is implemented as software stored in the memory unit 515 and executable by the processor unit 505, as firmware, as one or more hardware circuits, or a combination thereof. According to some aspects, the inversion model 520 generates an intermediate output based on the input image and the text description. In some examples, the inversion model 520 iteratively alternates between generating successive intermediate outputs and corresponding modified images. According to some aspects, the inversion model 520 generates the intermediate output based on the training image and the text description.

[0095] According to some aspects, the inversion model 520 includes parameters stored in at least one memory and is trained to generate an intermediate output based on an input image and a text description of the input image. In some aspects, the inversion model 520 has the same architecture as the image generation model 525. The inversion model 520 is a reference Figure 6 and Figure 7 Examples of corresponding elements described or including references Figure 6 and Figure 7 Describes aspects of the corresponding element.

[0096] In some aspects, the image generation model 525 is implemented as software stored in the memory unit 515 and executable by the processor unit 505, as firmware, as one or more hardware circuits, or as a combination thereof. In some aspects, the image generation model 525 generates a modified image based on the intermediate output and the modification prompt, wherein the modified image retains elements of the input image and includes modifications based on the modification prompt. In some examples, the image generation model 525 generates a reconstructed image based on the intermediate output and the text description. In some examples, the image generation model 525 adds the intermediate output to the input image to obtain a noisy image, wherein the image generation model 525 receives the noisy image as input. In some examples, the image generation model 525 performs a single-pass diffusion process.

[0097] According to some aspects, the image generation model 525 generates a reconstructed image based on the intermediate output and the text description. In some examples, the image generation model 525 generates a modified image based on the intermediate output and the modification prompt. According to some aspects, the image generation model 525 includes parameters stored in at least one memory and is trained to generate a modified image based on the intermediate output and the modification prompt, wherein the modified image retains elements of the input image and includes modifications based on the modification prompt. In some aspects, the image generation model 525 includes a single-pass or multi-pass diffusion model. The image generation model 525 is a reference Figure 6 and Figure 7 Examples of corresponding elements described or including references Figure 6 and Figure 7 Describes aspects of the corresponding element.

[0098] According to some aspects, the caption generation model 530 is implemented as software stored in the memory unit 515 and executable by the processor unit 505, is implemented as firmware, one or more hardware circuits, or a combination thereof. In some aspects, the modification prompt includes an edit to the text description. In some aspects, the modification prompt includes a description of the change to the input image. In some examples, the caption generation model 530 generates a text description based on the input image. According to some aspects, the caption generation model 530 is configured to generate a text description based on the input image. The caption generation model 530 is a reference Figure 6 Examples of corresponding elements described or including references Figure 6 Describes aspects of the corresponding element.

[0099] According to some aspects, the training component 535 is implemented as software stored in the memory unit 515 and executable by the processor unit 505, as firmware, as one or more hardware circuits, or a combination thereof. According to some embodiments, the training component 535 is implemented as software stored in the memory unit and executable by a processor in the processor unit of a separate computing device, as firmware in a separate computing device, as one or more hardware circuits of a separate computing device, or a combination thereof. In some examples, the training component 535 is part of another device other than the image processing device 500 and communicates with the image processing device 500. In some examples, the training component 535 is part of the image processing device 500.

[0100] According to some aspects, the training component 535 obtains a training set comprising training images and training descriptions of the images. In some examples, the training component 535 trains the inversion model 520 based on the training images and the reconstructed images. In some examples, the training component 535 generates training descriptions based on the training images. In some examples, the training component 535 calculates a reconstruction loss based on the training images and the reconstructed images. In some examples, the training component 535 updates the parameters of the inversion model 520 based on the reconstruction loss. In some examples, the training component 535 calculates a modification loss based on the modified images and the true modified images. In some examples, the training component 535 updates the parameters of the inversion model 520 based on the modification loss. In some examples, the training component 535 initializes the inversion model 520 using parameters from the image generation model 525. In some aspects, the image generation model 525 is fixed during the training of the inversion model 520.

[0101] Figure 6 An example of a machine learning model 600 according to aspects of the present disclosure is shown. The example shown includes the machine learning model 600, an input image 605, a caption generation model 610, a text description 615, a modification hint 620, a noisy input 625, an original input 630, an inversion model 635, a first reconstructed image 640, a first intermediate output 645, a second reconstructed image 650, a second intermediate output 655, a modified input 660, an image generation model 665, a first modified image 670, and a final modified image 675.

[0102] refer to Figure 6, the machine learning model 600 receives the input image 605 and the modification hint 620 and generates the final modified image 375. According to some embodiments, the caption generation model 610 receives the input image 605 and generates a text description 615. In some cases, the input image 605 depicts a plate of food. In some cases, the text description 615 describes the input image 605 as "an image primarily of a white plate filled with delicious dishes." The plate is topped with various items of food, including a piece of fish, asparagus, and tomatoes. The fish is placed in the center of the plate, while the asparagus and tomatoes are scattered around the plate. The arrangement of the food items creates a visually appealing and appetizing display. In some cases, the modification hint 620 modifies an element based on the text description 615. For example, the modification hint 620 replaces the fish with a steak. For example, the modification hint 620 provides input to the machine learning model 600 to generate a composite image (e.g., the final modified image 375) that depicts a change from fish on a white plate to steak on the same plate, while maintaining other elements of the input image 605.

[0103] In some embodiments, the inversion model 635 receives the noise input 625 and the original input 630 to generate a first reconstructed image 640 at a first time step. For example, the noise input 625 may include random noise or may be a noise map. For example, the original input 630 includes the input image 605, the text description 615, and the current time step. In some cases, the first reconstructed image 640 may include visual features / elements that are substantially the same as the visual features / elements of the input image 605. In some embodiments, the inversion model 635 generates a first intermediate output 645 based on the noise input 625 and the original input 630. In some cases, the first intermediate output 645 may be a latent representation of the first reconstructed image 640 (e.g., latent features, visual features, latent code, or a combination thereof).

[0104] In some embodiments, during the second time step (or a subsequent time step), the inversion model 635 receives the first reconstructed image 640 and the original input 630 to generate a second reconstructed image 650 at the second time step or a subsequent time step. In some cases, for example, the second reconstructed image 650 may include visual features / elements that are substantially the same as the visual features / elements of the input image 605. In some embodiments, the inversion model 635 generates a second intermediate output 655 based on the first reconstructed image 640 and the original input 630. In some cases, the second intermediate output 655 may be a potential representation of the second reconstructed image 650.

[0105] In some embodiments, the image generation model 665 is configured to generate a composite image (e.g., a first modified image 670 or a final modified image 675) based on the modified input 660 and the corresponding intermediate output (e.g., the first intermediate output 645 or the second intermediate output 655). For example, the image generation model 665 receives the modified input 660 and the first intermediate output 645 to generate the first modified image 670. In some cases, the modified input 660 includes a modification hint 620 and a corresponding time step. For example, the first intermediate output 645 is provided to the image generation model 665 to initiate the image generation process, and the modification hint 620 of the modified input 660 can be used to guide the image generation process (e.g., a back diffusion process). Reference Figure 8 to describe more details about the backdiffusion process.

[0106] In some embodiments, the image generation model 665 receives the first modified image 670, the modified input 660, and the second intermediate output 655 to generate a final modified image 675. In some cases, the second intermediate output 655 is provided to the image generation model 665 to initiate the image generation process. In some cases, for example, the first modified image 670 and the modified input 660 are provided to a multimodal encoder to generate an embedding (or concatenated embedding), where the embedding is used to guide the back diffusion process of the image generation model 665 to generate the final modified image 675.

[0107] In the field of diffusion model, by adding Gaussian noise n t Iteratively added to the clear image, the forward diffusion process gradually transforms the clear image x0 into white Gaussian noise x T .

[0108]

[0109] where β t represents a noisy schedule. In some cases, the process can be rewritten as:

[0110]

[0111] where α t =1-β t , And ∈ t represents Gaussian noise. is trained to use the loss function, given x t , text prompt t and time step t to generate ∈ t :

[0112]

[0113] During the sampling process, the sample x t-1 We can use the following equation to calculate the value of x t And generate:

[0114] in In some cases, the network can generate the predicted x0 at time step t using the following equation:

[0115]

[0116] In some cases, the model is derived from sampled Gaussian noise x t It takes 20-50 steps to get to the clean image x0. In some cases, the image generation model 665 of the present disclosure can obtain a high-quality image in 1-4 steps.

[0117] Given an input real image x0 (e.g., input image 605), the caption generation model 610 generates a detailed caption c (e.g., text description 615). In some cases, attributes in c may be modified to create a new text prompt c' (e.g., modified prompt 620). Given the detailed nature of c and the subtle nature of the attribute modifications, c and c' may be substantially similar. The inversion process is performed by combining x0, c, the current time step t, and the previously reconstructed image x 0,t+1 It starts by feeding ∑(x) into the inversion model 635 (initially as a zero matrix). The inversion model 635 then generates the noise ∈ t (e.g., first intermediate output 645), noise ∈ t is fed into the image generation model 665 to generate a new constructed image x 0,0 (eg, the first reconstructed image 640), the new constructed image x 0,0 The inversion process is similar to the input image x0. It iterates from t = T to smaller t by first encoding the semantic information and then capturing finer details. Noise ∈ t includes spatial information that is not explicitly encoded in c. Given the final inverted noise ∈ t Together with c, the image generation model 665 is used to generate the inversion trajectory and reconstruct the image x 0,0 , image x 0,0 Similar to the input image x0. Using the same noise ∈ t With a slightly different textual cue c' (eg, modification cue 620), starting from t=T to smaller t, the editing trajectory may be similar to the inversion trajectory and the modified image may be very similar to the input image 605, but differ in the specified attributes in c'.

[0118] Machine Learning Model 600 is a reference Figure 3 and Figure 5 Examples of corresponding elements described or including references Figure 3 and Figure 5 The input image 605 is a reference to the corresponding elements. Figure 3 and Figure 7 Examples of corresponding elements described or including references Figure 3 and Figure 7 The description text generation model 610 is a reference to the corresponding elements. Figure 5 Examples of corresponding elements described or including references Figure 5 The modification hint 620 is a reference to the corresponding elements of the description. Figure 3 Examples of corresponding elements described or including references Figure 3 Describes aspects of the corresponding element.

[0119] Inversion model 635 is reference Figure 5 and Figure 7 Examples of corresponding elements described or including references Figure 5 and Figure 7 The first reconstructed image 640 is a reference to the corresponding elements of Figure 7 Examples of corresponding elements described or including references Figure 7 The first intermediate output 645 is a reference to the corresponding elements of Figure 7 Examples of corresponding elements described or including references Figure 7 The image generation model 665 is a reference to the corresponding elements of the description. Figure 5 and Figure 7 Examples of corresponding elements described or including references Figure 5 and Figure 7 Describes aspects of the corresponding element.

[0120] Figure 7 An example of an inversion model 700 according to aspects of the present disclosure is shown. The example shown includes the inversion model 700, an input image 705, a first reconstructed image 710, an inversion network 715, an intermediate output 720, an image generation model 725, and a second reconstructed image 730. In one aspect, the inversion model 700 includes the inversion network 715. In an embodiment, the inversion model 700 includes the inversion network 715 and the image generation model 725.

[0121] refer to Figure 7 , the generator (e.g., image generation model 725) receives the time step t, the text prompt c, and the noisy image x t =x0+∈ t And output the intermediate reconstructed image x 0,t (e.g., second reconstructed image 730). In some cases, the model can generate a clean image x from a noisy version using the following equation0,t :

[0122] x 0,t =G(t,c,x t ) (6)

[0123] In some embodiments, the inversion network 715 is used in a single-step approach, where t = T. Given a ground-truth image x0 and a corresponding textual cue c, the inversion network 715F single is trained to generate ∈ T (e.g., intermediate output 720), such that when ∈ T When fed to G (e.g., the image generation model 725), x 0,t can be the same as x0 using the following loss function:

[0124] in

[0125]

[0126] The inversion network 715 is initialized from G (e.g., the image generation model 725), where G is frozen during training. The information of the input image x0 is stored in the contextual prompt c (e.g., global information) and ∈ T =F single (T, c, x0) (e.g., spatial information). Then, to perform image editing, the modification hint c' is used, where the modified image can be represented as:

[0127]

[0128] According to some embodiments, the single-step encoder approach described above performs semantic editing while preserving background details.

[0129] According to some embodiments, the inversion process is performed iteratively to refine the reconstructed image. For example, the inversion network 715 is used to take the input image x0 together with the image from the previous step size x 0,t+1 and generates the predicted noise ∈ for the current time step t (e.g., intermediate output 720). The injected noise ∈ t With the previous reconstruction x 0,t+1 Combine (via concatenation) to generate a new noisy image x using the image generation model 725 t , thus obtaining the new constructed image x 0,t Additionally, the multi-step training loss can be formulated as follows:

[0130] in

[0131]

[0132] In some embodiments, the image generation model 725 takes the first reconstructed image 710 as input, and thus the loss function drives the inversion network 715 to output noise ∈ t , noise∈ t The first reconstructed image 710 is enhanced with respect to the input image 705. During training, at time step t=T, a zero matrix is ​​used as the first reconstructed image 710.

[0133] In some embodiments, a reparameterization method is used to constrain the injected noise to a distribution close to a standard Gaussian to prevent high predicted noise values ​​and excessive structural information from the input image, which may cause additional artifacts in the reconstructed image. In some cases, the inversion network 715 generates a mean and variance for each pixel from which the injected noise can be sampled. The KL loss for this modification can be expressed as:

[0134]

[0135] Therefore, the total loss can be expressed as:

[0136] L(E)=L MSE (F)+λ*L KL (F)) (11)

[0137] In some cases, the hyperparameter λ is set to λ=10 -6 .

[0138] According to some aspects, after training a machine learning model, the model can perform the following Figure 6 In some cases, the inversion model 700 and the image generation model 725 are iterated to output ∈ t In some cases, ∈ t includes different levels of spatial information of the input image x0. For example, ∈ t Including spatial semantic information, and with small t∈ t Include fine details. Use the generated intermediate output ∈ t and a new textual prompt c′, the machine learning model can generate a new image (e.g., a modified image) that is similar to the input image x0 and includes the target attributes described in c′.

[0139] Figure 8An example of an image generation model according to aspects of the present disclosure is shown. The example shown includes a diffusion model 800, an original image 805, a pixel space 810, an image encoder 815, original image features 820, a latent space 825, a forward diffusion process 830, noise features 835, a backward diffusion process 840, denoised image features 845, an image decoder 850, an output image 855, a text prompt 860, a text encoder 865, a guide feature 870, and a guide space 875.

[0140] Diffusion models are a type of generative neural network that can be trained to generate new data with features similar to those found in the training data. Specifically, diffusion models can be used to generate new images. Diffusion models can be used for a variety of image generation tasks, including image super-resolution, image generation using perceptual metrics, conditional generation (e.g., based on text guidance, color guidance, style guidance, and image guidance), image interior filling, and image manipulation.

[0141] Types of diffusion models include denoising diffusion probabilistic models (DDPMs) and denoising diffusion implicit models (DDIMs). In DDPMs, the generation process involves inverting a stochastic Markov diffusion process. DDIMs, on the other hand, use a deterministic process so that the same input produces the same output. Diffusion models can also be characterized by whether the noise is added to the image or to image features generated by the encoder (e.g., latent diffusion).

[0142] The diffusion model works by iteratively adding noise to the data during the forward process and then learning to recover the data by denoising the data during the reverse process. For example, during training, the diffusion model 800 can take as input an original image 805 in pixel space 810 and apply an image encoder 815 to convert the original image 805 into original image features 820 in a latent space 825. Then, a forward diffusion process 830 gradually adds noise to the original image features 820 to obtain noise features 835 (also in the latent space 825) at different noise levels.

[0143] Next, a back diffusion process 840 (e.g., a U-Net ANN) gradually removes noise from the noise features 835 at various noise levels to obtain denoised image features 845 in the latent space 825. In some examples, the denoised image features 845 are compared to the original image features 820 at each of the various noise levels and the parameters of the back diffusion process 840 of the diffusion model are updated based on the comparison. Finally, an image decoder 850 decodes the denoised image features 845 to obtain an output image 855 in the pixel space 810. In some cases, the output image 855 is created at each of the various noise levels. The output image 855 can be compared to the original image 805 to train the back diffusion process 840. In some cases, the output image 855 refers to a synthesized image (e.g., a reference image). Figure 3 、 Figure 4 and Figure 6 described)

[0144] In some cases, the image encoder 815 and the image decoder 850 are pre-trained before training the backdiffusion process 840. In some examples, the image encoder 815 and the image decoder 850 are jointly trained or the image encoder 815 and the image decoder 850 are jointly fine-tuned with the backdiffusion process 840.

[0145] The back diffusion process 840 can also be guided based on textual cues 860 or other guiding cues such as images, layouts, styles, colors, segmentation maps, etc. The textual cues 860 can be encoded using a text encoder 865 (e.g., a multimodal encoder) to obtain guided features 870 in a guided space 875. The guided features 870 can be combined with the noise features 835 at one or more layers of the back diffusion process 840 to ensure that the output image 855 includes the content described by the textual cues 860. For example, the guided features 870 can be combined with the noise features 835 using a cross-attention block within the back diffusion process 840.

[0146] Cross-attention, also known as multi-head attention, is an extension of the attention mechanism used in some ANNs, for example for NLP tasks. In some cases, cross-attention focuses on multiple parts of the input sequence simultaneously, thereby capturing interactions and correlations between different elements. In cross-attention, there are two input sequences: a query sequence and a key-value sequence. The query sequence represents the elements that require attention, while the key-value sequence contains the elements to be attended to. In some cases, to compute cross-attention, the cross-attention block transforms (e.g., using linear projection) each element in the query sequence into a "query" representation, while the elements in the key-value sequence are transformed into "key" and "value" representations.

[0147] The crisscross attention block computes an attention score by measuring the similarity between each query representation and the key representation, where a higher similarity indicates more attention is paid to the key element. The attention score indicates the importance or relevance of each key element to the corresponding query element.

[0148] The crisscross attention block then normalizes the attention scores to obtain attention weights (e.g., using a softmax function), where the attention weights determine how much information from each value element is incorporated into the final attention representation. By simultaneously focusing on different parts of the key-value sequence, the crisscross attention block captures relationships and correlations across the input sequence, enabling the machine learning model to understand the context and generate more accurate and contextually relevant outputs.

[0149] In some examples, the diffusion model is based on a neural network architecture called U-Net. U-Net takes input features having an initial resolution and an initial number of channels and processes the input features using an initial neural network layer (e.g., a convolutional network layer) to generate intermediate features. The intermediate features are then downsampled using a downsampling layer so that the downsampled features have a resolution smaller than the initial resolution and a number of channels larger than the initial number of channels.

[0150] This process is repeated multiple times and then the process is reversed. For example, the downsampled features are upsampled using the upsampling process to obtain upsampled features. The upsampled features can be combined with intermediate features of the same resolution and number of channels via skip connections. These inputs are processed using the final neural network layer to produce output features. In some cases, the output features have the same resolution as the initial resolution and the same number of channels as the initial number of channels.

[0151] In some cases, U-Net requires additional input features to produce conditionally generated outputs. For example, the additional input features may include vector representations of input prompts. The additional input features may be combined with intermediate features at one or more layers within the neural network. For example, a cross-attention module may be used to combine the additional input features with the intermediate features. For more details on U-Net, see Figure 9 To describe.

[0152] The diffusion process can also be modified based on conditional guidance. In some cases, the user provides a text prompt (e.g., text prompt 860) that describes the content included in the generated image. In some examples, the guidance can be provided in a form other than text, such as via an image, sketch, color, style, or layout. The system converts the text prompt 860 (or other guidance) into a conditional guidance vector or other multidimensional representation. For example, the text can be converted into a vector or a series of vectors using a transformer model or a multimodal encoder. In some cases, the conditional guidance encoder is trained independently of the diffusion model.

[0153] A noise map comprising random noise is initialized. The noise map can be in pixel space or latent space. By initializing the image with random noise, different variations of the image including the content described by the conditional guidance can be generated. The diffusion model 800 then generates the image based on the noise map and the conditional guidance vector.

[0154] The diffusion process may include a forward diffusion process 830 for adding noise to an image (e.g., original image 805) or a feature (e.g., original image feature 820) in a latent space 825 and a backward diffusion process 840 for denoising the image (or feature) to obtain a denoised image (e.g., output image 855). The forward diffusion process 830 may be represented as q(x t ∣x t-1 ), the reverse diffusion process 840 can be represented as p θ (x t-1 ∣x t For further details on the diffusion process, see Figure 10 To describe.

[0155] The diffusion model 800 can be trained using both the forward diffusion process 830 and the backward diffusion process 840. In one example, an untrained model is initialized. Initialization can include defining the architecture of the model and establishing initial values ​​for the model parameters. In some cases, initialization can include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, the location of skip connections, etc.

[0156] The system then uses an N-stage forward diffusion process 830 to add noise to the training image. In some cases, the forward diffusion process 830 is a fixed process in which Gaussian noise is continuously added to the image. In a latent diffusion model, Gaussian noise can be continuously added to features in the latent space 825 (e.g., original image features 820).

[0157] At each stage n, starting from stage N, the back diffusion process 840 is used to predict the image or image features at stage n-1. For example, the back diffusion process 840 can predict the noise added by the forward diffusion process 830, and the predicted noise can be removed from the image to obtain a predicted image. In some cases, the original image 805 is predicted at each stage of the training process.

[0158] Training components (e.g., reference Figure 7 The training component described herein compares the predicted image (or image features) at stage n-1 with the actual image (or image features) (such as the image at stage n-1 or the original input image). For example, given observation data x, the diffusion model 800 can be trained to compare the negative log-likelihood of the training data -log p θ The variational upper bound of (x) is minimized. The training component then updates the parameters of the diffusion model 800 based on the comparison. For example, the parameters of the U-Net can be updated using gradient descent. The time-dependent parameters of the Gaussian transformation can also be learned. For further details on training the diffusion model, see Figure 15 To describe.

[0159] Figure 9 An example of a U-Net 900 architecture according to aspects of the present disclosure is shown. The example shown includes U-Net 900, input features 905, initial neural network layer 910, intermediate features 915, downsampling layer 920, downsampled features 925, upsampling process 930, upsampled features 935, skip connections 940, final neural network layer 945, and output features 950.

[0160] In some examples, U-Net 900 is the execution reference Figure 8 Examples of components of the reverse diffusion process 840 of the diffusion model 800 are described and include reference to Figure 5 Describes the architectural elements of the image generation model. Figure 9 The U-Net 900 depicted in the reference Figure 8 Examples of architectures used in the backdiffusion process described or included in reference Figure 10 Describes various aspects of the architecture used in the backdiffusion process.

[0161] In some examples, the diffusion model is based on a neural network architecture called U-Net. U-Net 900 takes input features 905 having an initial resolution and an initial number of channels and processes the input features 905 using an initial neural network layer 910 (e.g., a convolutional network layer) to produce intermediate features 915. The intermediate features 915 are then downsampled using a downsampling layer 920 so that the downsampled features 925 have a resolution smaller than the initial resolution and a number of channels larger than the initial number of channels.

[0162] This process is repeated multiple times and then the process is reversed. For example, the downsampled features 925 are upsampled using the upsampling process 930 to obtain upsampled features 935. The upsampled features 935 can be combined with the intermediate features 915 of the same resolution and number of channels via skip connections 940. These inputs are processed using the final neural network layer 945 to produce output features 950. In some cases, the output features 950 have the same resolution as the initial resolution and have the same number of channels as the initial number of channels.

[0163] In some cases, U-Net 900 uses additional input features to produce conditionally generated outputs. For example, the additional input features can include vector representations of input prompts. The additional input features can be combined with intermediate features 915 at one or more layers within the neural network. For example, a cross-attention module can be used to combine the additional input features with the intermediate features 915.

[0164] Diffusion process

[0165] Figure 10 An example of a diffusion process 1000 according to aspects of the present disclosure is shown. The example shown includes the diffusion process 1000, a forward diffusion process 1005, a backward diffusion process 1010, a noise image 1015, a first intermediate image 1020, a second intermediate image 1025, and an original image 1030.

[0166] The diffusion process 1000 may include a forward diffusion process 1005 for transferring the original image 1030 (e.g., reference image 1030) in the latent space. Figure 8 The original image 805 described) or features (e.g., reference Figure 8 In some aspects, the diffusion process 1000 includes a back diffusion process 1010 for denoising the noisy image 1015 (or image features) to obtain a denoised image (or original image 1030). The forward diffusion process 1005 can be represented as q(x t ∣x t-1 ) and the reverse diffusion process 1010 can be represented as p θ (x t-1 ∣x t In some cases, the forward diffusion process 1005 is used during training to generate images with continuously increasing noise, and the neural network is trained to perform the backward diffusion process 1010 (eg, to continuously remove the noise).

[0167] In the context of potential diffusion models (e.g., ref. Figure 8In the forward diffusion process 1005 of the diffusion model 800 described above, the diffusion model uses a Markov chain to map the observed variable x0 (in pixel space or latent space) to obtain intermediate variables x1, ..., x T When the latent variable is passed through a neural network such as U-Net, the Markov chain gradually adds Gaussian noise to the data to obtain an approximate posterior q(x 1:T |x0), where x1,…,x T has the same dimensions as x0.

[0168] A neural network can be trained to perform the back diffusion process 1010. During the back diffusion process 1010, the diffusion model utilizes the noise data x T (such as the noisy image 1215), and denoise the data to obtain p θ (x t-1 ∣x t ). At each step t-1, the back diffusion process 1010 uses x t (such as the first intermediate image 1020) and t as input. Here, t represents a step in the transformation sequence associated with different noise levels. The back diffusion process 1010 iteratively outputs x t-1 , such as the second intermediate image 1025, until x T Restore to x0, that is, the original image 1030. The back diffusion process 1010 can be expressed as:

[0169] p θ (x t-1 ∣x t ):=N(x t-1 ;μ θ (x t ,t),Σ θ (x t ,t)), (12)

[0170] The joint probability of a sequence of samples in a Markov chain can be written as the product of the conditional probability and the marginal probability:

[0171]

[0172] where p(x T )=N(x T 0, I) is the pure noise distribution when the reverse diffusion process 1010 uses the result of the forward diffusion process 1005 (pure noise sample) as input and Represents the Gaussian transformed sequence corresponding to a sequence with Gaussian noise added to the samples.

[0173] At inference time, the observation data x0 in pixel space can be mapped into the latent space as input and the generated data It can be mapped back from the latent space to the pixel space as output. In some examples, x0 represents the original input image with low image quality, and the latent variables x1,…,x T represents a noisy image and Indicates the generated image has high image quality.

[0174] Image Editing

[0175] Figure 11 An example of a method 1100 for generating a modified image according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a code set to control functional elements of a device. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or performed in conjunction with other operations.

[0176] At operation 1105, the system obtains an input image depicting a first element, a text description of the input image, and a modification prompt, wherein the modification prompt describes a second element that is different from the first element. In some cases, the operation of this step is referred to as reference. Figure 3 、 Figure 8 and Figure 9 The machine learning model can be obtained by referring to Figure 3 、 Figure 8 and Figure 9 In some cases, the text description and the modification hint may be substantially the same. In some cases, the modification hint may be a simple phrase describing the change to the element depicted in the input image.

[0177] In some cases, the first element and the second element may include one or more image elements that are the same or different. For example, an image element is an image component or image feature that constitutes the overall composite of an image, such as an object, entity, subject, shape, color, texture, pattern, background scene, visual attribute, and / or style. For example, an image element can be an animal (such as a cat or dog), a person, an object (such as a hat or table), a scene (such as a beach or a mountaintop), or a combination thereof. For example, the first element can refer to a fish and the second element can refer to a steak.

[0178] At operation 1110, the system generates an intermediate output based on the input image and the text description using the inversion model, wherein the intermediate output represents the first element. In some cases, the operation of this step is referred to as reference Figure 5-Figure 7 The inversion model can be obtained by referring to Figure 5-Figure 7 In some cases, the intermediate output may include a latent representation, such as latent features, visual features, and / or a latent code of an image (e.g., a reconstructed image). In some cases, the intermediate output may be a predicted noise generated based on an input image and a textual description of the input image using the inversion model.

[0179] At operation 1115, the system generates a composite image using the image generation model based on the intermediate output and the modification prompt, wherein the composite image replaces the first element from the input image with the second element from the modification prompt. In some cases, the operation of this step is referred to as reference Figure 5-Figure 7 The image generation model can be obtained by referring to Figure 5-Figure 7 The image generation model is used to perform the modification. For example, the synthesized image includes image pixels generated by the image generation model. For example, the modified image includes image pixels from the input image and image pixels generated by the image generation model.

[0180] Training and evaluation

[0181] exist Figure 12-15 In the present invention, a method, apparatus, non-transitory computer-readable medium, and system for image processing include: obtaining a training set, the training set including training images and training descriptions of the training images; using an inversion model to generate an intermediate output based on the training images and the training descriptions; using an image generation model to generate a reconstructed image based on the intermediate output and the training descriptions; and training the inversion model to perform image inversion.

[0182] According to some embodiments, methods, apparatus, non-transitory computer-readable media, and systems for image processing include: obtaining a training set comprising training images and training descriptions of the images; using an inversion model to generate an intermediate output based on the training images and the text descriptions; using an image generation model to generate a reconstructed image based on the intermediate output and the text descriptions; and training the inversion model based on the training images and the reconstructed images.

[0183] Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include generating a training description based on the training images. Some examples of the methods, apparatuses, non-transitory computer-readable media, and systems further include calculating a reconstruction loss based on the training images and the reconstructed images. Some examples further include updating parameters of the inversion model based on the reconstruction loss.

[0184] Some examples of methods, apparatuses, non-transitory computer-readable media, and systems further include generating a modified image based on the intermediate output and the modification hint. Some examples further include calculating a modification loss based on the modified image and the true modified image. Some examples further include updating parameters of an inversion model based on the modification loss. Some examples of methods, apparatuses, non-transitory computer-readable media, and systems further include initializing the inversion model using parameters from the image generation model. In some aspects, the image generation model is fixed during training of the inversion model.

[0185] Figure 12 An example of a method 1200 for training a machine learning model according to various aspects of the present disclosure is shown. In some examples, these operations are performed by a system that includes a processor that executes a code set to control functional elements of a device. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or performed in conjunction with other operations.

[0186] At operation 1205, the system obtains a training set, which includes training images and training descriptions of the training images. In some cases, the operation of this step is referred to as reference Figure 8 The training component may be provided by reference to Figure 8 In some cases, the training set is stored in a Figure 1 in the database.

[0187] At operation 1210, the system generates an intermediate output based on the training images and the training description using the inversion model. In some cases, the operation of this step is referred to as Figure 5-Figure 7 The inversion model can be obtained by referring to Figure 5-Figure 7 In some cases, the intermediate output includes a reconstructed predicted noise of the input image. Figure 6-Figure 7 Describes more details about the intermediate output.

[0188] At operation 1215, the system generates a reconstructed image based on the intermediate output and the training description using the image generation model. In some cases, the operation of this step is referred to as reference Figure 6-Figure 7 The image generation model can be obtained by referring to Figure 6-Figure 7 In some cases, the reconstructed image can be substantially the same as the input image. Figure 6-Figure 7 to describe more details about the reconstructed image.

[0189] At operation 1220, the system trains the inversion model to perform image inversion. In some cases, the operation of this step is referred to as Figure 5 The training component may be provided by reference to Figure 5 The training component is performed. In some cases, the inversion network of the inversion model is initialized from the SDXL-Turbo model, and the image generator of the inversion network is fixed throughout the training. In some aspects, a learning rate of 10 and a batch size of 10 are used to train the machine learning model.

[0190] Figure 13 An example of a method 1300 for training an inversion model according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system comprising a processor that executes a code set to control functional elements of a device. Additionally or alternatively, certain processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or performed in conjunction with other operations.

[0191] At operation 1305, the system generates a modified image based on the intermediate output and the modification hint. In some cases, the operation of this step is referred to as Figure 5-Figure 7 The image generation model can be obtained by referring to Figure 5-Figure 7 At operation 1310, the system calculates the modification loss based on the modified image and the real modified image. In some cases, the operation of this step is referred to as Figure 5 The training component may be provided by reference to Figure 5 At operation 1315, the system updates the parameters of the inversion model based on the modified loss. In some cases, the operation of this step is referred to as Figure 5 The training component may be provided by reference to Figure 5 The training component is used to execute the Figure 7 More details describing the loss function used to update the parameters of the inversion model.

[0192] Figure 14 An example of a flowchart depicting an algorithm as a step-by-step procedure in an example operational implementation that may be performed for training a machine learning model according to aspects of the present disclosure is shown. In some embodiments, the procedure 1400 describes the operation of the training component 535, which is described for configuring as described in reference Figure 5 The first image generation model 525 and / or the second image generation model 520. The procedure 1400 provides one or more examples of generating training data, using the training data to train a machine learning model, and using the trained machine learning model to perform a task.

[0193] To get started in this example, the machine learning system collects training data (block 1402), which is used as the basis for training the machine learning model, i.e., it defines what is being modeled. Training data can be collected by the machine learning system from a variety of sources. Examples of training data sources include public datasets, service provider system platforms that expose application programming interfaces (e.g., social media platforms), user data collection systems (e.g., digital surveys and online crowdsourcing systems), etc. Training data collection can also include data augmentation and synthetic data generation techniques for expanding and diversifying the available training data, balancing techniques for balancing the number of positive and negative examples, etc.

[0194] The machine learning system may also be configured to identify and / or determine features relevant to the type of task for which the machine learning model is to be trained (block 1404). Examples of tasks include classification, natural language processing, generative artificial intelligence, recommendation engines, reinforcement learning, clustering, and the like. To this end, the machine learning system collects training data based on the identified features and / or, after collection, filters the training data based on the identified features. The training data is then used to train the machine learning model.

[0195] To train the machine learning model in the illustrated example, the machine learning model is first initialized (block 1406). Initialization of the machine learning model includes selecting a model architecture to be trained (block 1408). Examples of model architectures include neural networks, convolutional neural networks (CNNs), long short-term memory (LSTM) neural networks, generative adversarial networks (GANs), decision trees, support vector machines, linear regression, logistic regression, Bayesian networks, random forest learning, dimensionality reduction algorithms, boosting algorithms, deep learning neural networks, U-Net architectures, and the like.

[0196] A loss function is also selected (block 1410). The loss function is used to measure the difference between the output of the machine learning model (e.g., the model prediction) and the target value (e.g., represented by the training data) used to train the machine learning model. Additionally, an optimization algorithm is selected (block 1412) that is used in conjunction with the loss function to optimize the parameters of the machine learning model during training. Examples of optimization algorithms include gradient descent, stochastic gradient descent (SGD), and the like.

[0197] Initialization of the machine learning model also includes setting initial values ​​for the machine learning model (block 1416), examples of which include initializing the weights and biases of the nodes to increase training efficiency and the computational resources consumed as part of the training. Hyperparameters for controlling the training of the machine learning model are also set (block 1414), examples of which include regularization parameters, model parameters (e.g., the number of layers in a neural network), learning rates, batch sizes selected from the training data, and the like. Hyperparameters are set using various techniques, including randomization techniques, heuristics learned from other training scenarios, and the like.

[0198] The machine learning model is then trained by the machine learning system using the training data (block 1418). A machine learning model refers to a computer representation that can be tuned (e.g., trained and retrained) to approximate an unknown function based on the input of the training data. Specifically, the term machine learning model can include models that learn and predict from known data using an algorithm (e.g., using the model architecture described above) to learn and relearn by analyzing the training data to generate outputs that reflect the patterns and properties expressed by the training data.

[0199] Examples of training types include supervised learning using labeled data, unsupervised learning involving finding underlying structures or patterns in training data, reinforcement learning based on optimization functions (e.g., rewards and / or penalties), using nodes as part of "deep learning," and the like. A machine learning model can be configured, for example, to include multiple nodes that collectively form multiple layers. A layer can be configured, for example, to include an input layer, an output layer, and one or more hidden layers. Computations are performed by the nodes within the layer using hidden states through a system of weighted connections that are "learned" during training, for example, by using a selected loss function and backpropagation to optimize the performance of the machine learning model to perform the associated task.

[0200] As part of training the machine learning model, a determination is made as to whether a stopping criterion is met (decision block 1420), i.e., the stopping criterion is used to validate the machine learning model. The stopping criterion can be used to reduce overfitting of the machine learning model, reduce consumption of computational resources, and improve the ability of the machine learning model to handle unseen data (data not included as examples in the training data). Examples of stopping criteria include, but are not limited to, a predefined number of epochs, verifying loss stability, achieving a performance improvement threshold, whether a threshold level of accuracy is met, or based on performance metrics such as precision and recall. If the stopping criterion is not met ("No" from decision block 1420), the process 1400 continues to train the machine learning model using the training data in the example (block 1418).

[0201] If the stopping criteria is met ("yes" from decision block 1420), the trained machine learning model is then used to generate an output based on subsequent data (block 1422). The trained machine learning model is, for example, trained to perform the task described above and thus, once trained, is configured to perform the task based on subsequent data received as input and processed by the machine learning model.

[0202] Figure 15 An example of a method 1500 for training a diffusion model according to aspects of the present disclosure is shown. In some examples, these operations are performed by a system including a processor that executes a code set to control functional elements of a device. Additionally or alternatively, some processes are performed using dedicated hardware. Generally, these operations are performed according to the methods and processes described in various aspects of the present disclosure. In some cases, the operations described herein are composed of various sub-steps or performed in conjunction with other operations.

[0203] In some embodiments, method 1500 describes the operation of training component 535, which is described as being used to train a Figure 5 The inversion model 520. Method 1500 represents training as described above. Figure 10 In some examples, these operations are performed by a system that includes a processor that executes a code set to control functional elements of a device, such as Figure 5 The inversion model and / or image generation model described in .

[0204] At operation 1505, the system initializes the untrained model. In some cases, the operation of this step is referred to as Figure 5 The training component may be provided by reference to Figure 5 Initialization can include defining the architecture of the model and establishing initial values ​​for the model parameters. In some cases, initialization can include defining hyperparameters such as the number of layers, the resolution and channels of each layer block, the location of skip connections, etc.

[0205] At operation 1510, the system adds noise to the media item using a forward diffusion process with N stages. In some cases, the operation of this step is as described in reference to Figure 5 The training component may be provided by reference to Figure 5 In some cases, the forward diffusion process is a fixed process in which Gaussian noise is continuously added to the media item (such as the original image). In a latent diffusion model, Gaussian noise can be continuously added to features in the latent space.

[0206] At operation 1515, the system predicts the media items for stage n-1 at each stage n, starting from stage N. In some cases, the operation of this step refers to or can be performed by referring to Figure 5 The training component is performed as described. In some cases, the media item is a synthetic image generated using an image generation model. For example, the backward diffusion process can predict the noise added by the forward diffusion process, and the predicted noise can be removed from the noise input to obtain a predicted output. In some cases, the original media item is predicted at each stage of the training process.

[0207] At operation 1520, the system compares the media item (or feature) predicted at stage n-1 with the media item at stage n-1. For example, in some cases, the system compares the synthesized image (or predicted image feature) at stage n-1 with the real image (or real feature) at stage n-1. In some cases, the operation of this step refers to or can be performed by referring to Figure 5 For example, given observation data x, the diffusion model can be trained to transform the negative log-likelihood of the training data into -log p θ Minimize the variational upper bound of (x).

[0208] At operation 1525, the system updates the parameters of the model based on the comparison. In some cases, the operation of this step refers to or can be performed by referring to Figure 5 The training components described above can be used to perform this. For example, the parameters of the U-Net can be updated using gradient descent. The time-dependent parameters of the Gaussian transform can also be learned.

[0209] computing devices

[0210] Figure 16 An example of a computing device 1600 according to aspects of the present disclosure is shown. The illustrated example includes computing device 1600, processor 1605, memory subsystem 1610, communication interface 1615, I / O interface 1620, user interface component 1625, and channel 1630.

[0211] In some embodiments, computing device 1600 is a reference Figure 1 and Figure 5 Examples of the image processing apparatus described herein include reference Figure 1 and Figure 5 In some embodiments, computing device 1600 includes a processor 1605 that can execute instructions stored in memory subsystem 1610 to obtain an input image, a text description, and a modification prompt, generate an intermediate output based on the input image and the text description, and generate a composite image based on the intermediate output and the modification prompt.

[0212] According to some embodiments, processor 1605 includes one or more processors. In some cases, processor 1605 is an intelligent hardware device, such as a general-purpose processing component, a digital signal processor (DSP), a central processing unit (CPU), a graphics processing unit (GPU), a microcontroller, an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), a programmable logic device, a discrete gate or transistor logic component, a discrete hardware component, or a combination thereof. In some cases, processor 1605 is configured to operate a memory array using a memory controller. In other cases, the memory controller is integrated into processor 1605. In some cases, processor 1605 is configured to execute computer-readable instructions stored in memory to perform various functions. In some embodiments, processor 1605 includes dedicated components for modem processing, baseband processing, digital signal processing, or transmission processing. Processor 1605 is a reference Figure 5 Examples of processor units described or including reference Figure 5 Aspects of the processor unit are described.

[0213] According to some embodiments, the memory subsystem 1610 includes one or more memory devices. Examples of memory devices include random access memory (RAM), read-only memory (ROM), or a hard disk. Examples of memory devices include solid-state memory and hard disk drives. In some examples, the memory is used to store computer-readable, computer-executable software, which includes instructions that, when executed, cause the processor to perform the various functions described herein. In some cases, the memory includes, for example, a basic input / output system (BIOS), which controls basic hardware or software operations, such as interaction with peripheral components or devices. In some cases, a memory controller operates the memory cells. For example, the memory controller may include a row decoder, a column decoder, or both. In some cases, the memory cells within the memory store information in the form of logical states. The memory subsystem 1610 is a reference to Figure 5 Examples of memory cells described or including reference Figure 5 Aspects of a memory cell are described.

[0214] According to some embodiments, the communication interface 1615 operates at the boundary between the communication entity (such as the computing device 1600, one or more user devices, the cloud, and one or more databases) and the channel 1630 and can record and process communications. In some cases, the communication interface 1615 is provided to enable a processing system coupled to a transceiver (e.g., a transmitter and / or receiver). In some examples, the transceiver is configured to transmit (or send) and receive signals for the communication device via an antenna. In some cases, a bus is used in the communication interface 1615.

[0215] According to some embodiments, I / O interface 1620 is controlled by an I / O controller to manage input and output signals for computing device 1600. In some cases, I / O interface 1620 manages peripheral devices that are not integrated into computing device 1600. In some cases, I / O interface 1620 represents a physical connection or port to an external peripheral device. In some cases, an I / O controller uses a controller such as 1620 or other known operating systems. In some cases, an I / O controller represents or interacts with a modem, keyboard, mouse, touch screen, or similar device. In some cases, an I / O controller is implemented as a component of a processor. In some cases, a user interacts with a device via I / O interface 1620 or a hardware component controlled by an I / O controller. I / O interface 1620 is a reference to Figure 5 Examples of I / O modules described or included with reference Figure 5 Describes various aspects of the I / O module.

[0216] According to some embodiments, user interface components 1625 enable a user to interact with computing device 1600. In some cases, user interface components 1625 include an audio device such as an external speaker system, an external display device such as a display screen, an input device (e.g., a remote control device that interfaces with the user interface directly or via an I / O controller), or a combination thereof.

[0217] The performance of the apparatus, system, and method of the present disclosure has been evaluated and the results indicate that the embodiments of the present disclosure have achieved higher performance than conventional techniques (e.g., conventional image generation models). Example experiments show that the image processing apparatus based on the present disclosure outperforms conventional image generation models. For details on example use cases based on embodiments of the present disclosure, refer to Figure 3 and Figure 4 To describe.

[0218] The descriptions and drawings described herein represent example configurations and do not represent all implementations within the scope of the claims. For example, operations and steps may be rearranged, combined, or otherwise modified. In addition, structures and devices may be represented in block diagram form to illustrate the relationships between components and avoid obscuring the concepts being described. Similar components or features may have the same name but may have different reference numerals corresponding to different figures.

[0219] Some modifications of the present disclosure will be apparent to those skilled in the art and the principles defined herein may be applied to other variations without departing from the scope of the present disclosure. Therefore, the present disclosure is not limited to the examples and designs described herein, but should be accorded the widest scope consistent with the principles and novel features disclosed herein.

[0220] The methods described may be implemented or performed by a device including a general purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic device, discrete gate or transistor logic, discrete hardware components, or any combination thereof. A general purpose processor may be a microprocessor, a conventional processor, a controller, a microcontroller, or a state machine. A processor may also be implemented as a combination of computing devices (e.g., a combination of a DSP and a microprocessor, a plurality of microprocessors, one or more microprocessors in combination with a DSP core, or any other such configuration). Thus, the functions described herein may be implemented in hardware or software and may be performed by a processor, firmware, or any combination thereof. If implemented in software executed by a processor, the functions may be stored on a computer-readable medium in the form of instructions or code.

[0221] Computer-readable media include non-transitory computer storage media and communication media, including any media that facilitate the transmission of code or data. Non-transitory storage media can be any available media that can be accessed by a computer. For example, non-transitory computer-readable media can include random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), compact discs (CDs) or other optical disc storage devices, magnetic disk storage devices, or any other non-transitory media for carrying or storing data or code.

[0222] Additionally, a connecting component may be properly referred to as a computer-readable medium. For example, if code or data is transmitted from a website, server, or other remote source using a coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), or wireless technology (such as infrared, radio, or microwave signals), the coaxial cable, fiber optic cable, twisted pair, DSL, or wireless technology is included in the definition of medium. Combinations of media are also included within the scope of computer-readable media.

[0223] In this disclosure and the appended claims, the word "or" indicates an inclusive list, so that, for example, a list of X, Y, or Z means X or Y or Z or XY or XZ or YZ or XYZ. Furthermore, the phrase "based on" is not used to indicate a closed set of conditions. For example, a step described as "based on condition A" can be based on both condition A and condition B. In other words, the phrase "based on" should be interpreted as "based at least in part on." Furthermore, the words "a" or "an" indicate "at least one."

Claims

1. A method comprising: Obtaining an input image and a modification hint, the input image depicting a first element, the modification hint describing a second element different from the first element; generating an intermediate output based on the input image using an inversion model, wherein the intermediate output includes image features representing the image; as well as A composite image is generated based on the intermediate output and the modification hint using an image generation model, wherein the composite image replaces the first element from the input image with the second element from the modification hint.

2. The method according to claim 1, further comprising: The text description of the input image is obtained, wherein the intermediate output is generated based on the text description.

3. The method according to claim 1, wherein: The modification hint includes an edit to a text description of the input image.

4. The method of claim 1 , wherein generating the composite image comprises: Using the inversion model and the image generation model, iteratively alternates between generating successive intermediate outputs.

5. The method according to claim 1, further comprising: generating a reconstructed image based on the intermediate output, wherein the reconstructed image depicts the first element; as well as Based on the reconstructed image, a subsequent intermediate output is generated, wherein the composite image is based on the subsequent intermediate output.

6. The method of claim 1 , wherein generating the composite image comprises: Get noise input; as well as The noise input is denoised based on the intermediate output.

7. The method according to claim 1, wherein: The inversion model is trained using a training set comprising training images and training descriptions of the training images.

8. A method for training a machine learning model, the method comprising: Obtaining a training set, the training set comprising an input image and a text description of the training image; generating an intermediate output based on the input image and the text description; generating, using an image generation model, a reconstructed image based on the intermediate output and the text description; as well as The inversion model is trained to perform image inversion using the training set and the reconstructed image.

9. The method according to claim 8, wherein obtaining the training set comprises: Based on the input image, the text description is generated.

10. The method of claim 8, wherein training the inversion model comprises: Calculating a reconstruction loss based on the input image and the reconstructed image; as well as Based on the reconstruction loss, parameters of the inversion model are updated.

11. The method of claim 8, wherein training the inversion model comprises: generating a modified image based on the intermediate output and the modification hint; calculating a modification loss based on the modified image and the true modified image; as well as Parameters of the inversion model are updated based on the modified loss.

12. The method according to claim 8, further comprising: The inversion model is initialized using parameters from the image generation model.

13. The method according to claim 8, wherein: The image generation model is frozen during the training of the inversion model.

14. An apparatus comprising: at least one memory component; at least one processing device coupled to the at least one memory component; an inversion model comprising parameters stored in the at least one memory component and trained to generate an intermediate output based on an input image and a text description, wherein the intermediate output represents a first element of the input image; as well as An image generation model comprising parameters stored in the at least one memory component and trained to generate a composite image based on the intermediate output and a modified prompt, wherein the composite image replaces the first element from the input image with a second element from the modified prompt.

15. The apparatus according to claim 14, further comprising: A caption generation model is configured to generate the text description based on the input image.

16. The apparatus according to claim 14, wherein: The modification prompt includes editing the text description.

17. The apparatus of claim 14, wherein generating the composite image comprises: Using the inversion model and the image generation model, iteratively alternates between generating successive intermediate outputs.

18. The apparatus of claim 14, wherein generating the composite image comprises: generating a reconstructed image based on the intermediate output, wherein the reconstructed image depicts the first element; as well as Based on the reconstructed image, a subsequent intermediate output is generated, wherein the composite image is based on the subsequent intermediate output.

19. The apparatus of claim 14, wherein generating the composite image comprises: Get noise input; as well as The noise input is denoised based on the intermediate output.

20. The apparatus of claim 18, wherein: The inversion model includes a diffusion model.