Image editing with selected machine learning model
By rewriting prompts using a large language model (LLM) and selecting an appropriate machine learning model for image editing, the problems of ambiguous user prompts and high computational costs in existing technologies are solved, achieving efficient and accurate image modification results.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- GOOGLE LLC
- Filing Date
- 2025-08-11
- Publication Date
- 2026-07-31
AI Technical Summary
Existing generative AI models face challenges in image editing, including ambiguous user prompts, high computational costs, and difficulty in selecting the appropriate model for high-quality modification based on user intent.
By rewriting user prompts using a large language model (LLM) and selecting different types of machine learning models for image editing based on the rewritten prompts, including structure-preserving, shape-preserving, and non-structure/non-shape-preserving models, more accurate and efficient image modification can be achieved.
It improves the accuracy and efficiency of image editing, can generate high-quality output images according to user intent, and reduces computing costs.
Smart Images

Figure FT_1 
Figure FT_2 
Figure FT_3
Abstract
Description
[0001] Cross-references to related applications
[0002] This application is an international application claiming priority to U.S. Provisional Patent Application No. 63 / 682,231, filed August 12, 2024, entitled “Selection of Machine-Learning Model for Image Editing,” the entire contents of which are hereby incorporated by reference. Background Technology
[0003] Generative artificial intelligence (AI) can be used to generate images based on text prompts. Generative AI models have different strengths and weaknesses. For example, when a user provides a prompt to change an initial image pointing to a pyramid to a beach in Bali, some generative AI models generate output images with pyramid artifacts because the shape of the pyramid does not correspond to the pixel positions where the beach was added in the output image.
[0004] The background description provided herein is for the purpose of presenting the overall context of this disclosure. The work of the currently nominated inventors (to the extent described in this background section) and aspects of the specification that may not be considered prior art at the time of filing are neither expressly nor impliedly acknowledged as prior art of this disclosure. Summary of the Invention
[0005] A computer-implemented method includes receiving an initial image and an original prompt from a user, wherein the original prompt includes a request to modify the initial image. The method further includes selecting a machine learning model from a set of machine learning models based on the original prompt. The method further includes providing the original prompt and the initial image as input to a large language model (LLM). The method further includes receiving a rewritten prompt from the LLM based on the original prompt and the initial image. The method further includes providing the rewritten prompt and the initial image as input to the selected machine learning model. The method further includes generating an output image from the selected machine learning model that satisfies the rewritten prompt.
[0006] In some embodiments, the method further includes receiving user input identifying one or more objects or regions in an initial image, wherein the rewriting prompt is further based on the identification of one or more objects or regions in the initial image to be modified. In some embodiments, the set of machine learning models includes a structure-preserving machine learning model, a shape-preserving machine learning model, and non-structure and non-shape-preserving machine learning models. In some embodiments, selecting a machine learning model includes selecting a structure-preserving machine learning model based on the rewriting prompt including a command to modify the one or more objects or regions in the initial image while preserving their structure. In some embodiments, providing the rewriting prompt and the initial image as input to the selected machine learning model further includes providing the rewriting prompt, the initial image, and a depth map of the initial image to the structure-preserving machine learning model. In some embodiments, selecting a machine learning model includes selecting a shape-preserving machine learning model based on the rewriting prompt including a command to modify the one or more objects or regions in the initial image while preserving their shape. In some embodiments, selecting a machine learning model includes selecting a non-structure and non-shape-preserving machine learning model based on the rewriting prompt including a command to replace one or more objects or regions in the initial image with one or more new objects or new regions. In some embodiments, in response to selecting an unstructured and non-shape-preserving machine learning model, the method further includes: generating a minimum bounding box around one or more selected objects in an initial image; generating a bounding box mask based on the minimum bounding box in response to selecting the unstructured and non-shape-preserving machine learning model; and providing the bounding box mask, along with a rewritten prompt and the initial image, as input to the unstructured and non-shape-preserving machine learning model. In some embodiments, selecting the machine learning model includes: selecting the unstructured and non-shape-preserving machine learning model based on a rewritten prompt including a command to generate additional objects to be added to the initial image. In some embodiments, the method further includes: generating a user interface including an initial image and options for applying a preset to modify the initial image; and, in response to receiving a selection of the preset, outputting an output image by the machine learning model that satisfies a command associated with the preset. In some embodiments, the preset includes at least one option selected from the group consisting of: removing fences from the initial image, erasing objects from the initial image, adding new objects to the initial image, changing the material or color of objects in the initial image, enhancing the initial image, replacing the background of the initial image, changing the subject in the initial image (e.g., changing the subject's expression, changing the subject's features, changing the subject's clothing, etc.), and combinations thereof.
[0007] A non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform an operation or control the execution of an operation. The operation includes: receiving an initial image and an original prompt from a user, wherein the original prompt includes a request to modify the initial image; selecting a machine learning model from a set of machine learning models based on the original prompt; providing the original prompt and the initial image as input to the LLM; receiving a rewritten prompt from the LLM based on the original prompt and the initial image; providing the rewritten prompt and the initial image as input to the selected machine learning model; and generating an output image from the selected machine learning model that satisfies the rewritten prompt.
[0008] In some embodiments, the operation further includes receiving user input identifying one or more objects or regions in an initial image, wherein the rewriting prompt is further based on the identification of one or more objects or regions in the initial image to be modified. In some embodiments, the set of machine learning models includes structure-preserving machine learning models, shape-preserving machine learning models, and non-structure and non-shape-preserving machine learning models. In some embodiments, the operation further includes: providing an option to regenerate an output image, receiving a subsequent prompt from the user, and generating a subsequent output image based on the subsequent prompt.
[0009] A system includes one or more processors and one or more computer-readable media coupled to the one or more processors, the one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations or control operations. The operations include: receiving an initial image and an original prompt from a user, wherein the original prompt includes a request to modify the initial image; selecting a machine learning model from a set of machine learning models based on the original prompt; providing the original prompt and the initial image as input to the LLM; receiving a rewritten prompt from the LLM based on the original prompt and the initial image; providing the rewritten prompt and the initial image as input to the selected machine learning model; and generating an output image from the selected machine learning model that satisfies the rewritten prompt.
[0010] In some embodiments, the operation further includes receiving user input identifying one or more objects or regions in an initial image, wherein the rewriting prompt is further based on the identification of one or more objects or regions in the initial image to be modified. In some embodiments, the set of machine learning models includes a structure-preserving machine learning model, a shape-preserving machine learning model, and non-structure and non-shape-preserving machine learning models. In some embodiments, selecting a machine learning model includes selecting a structure-preserving machine learning model based on a rewriting prompt that includes a command to modify one or more objects or regions in the initial image while preserving their structure. In some embodiments, selecting a machine learning model includes selecting a shape-preserving machine learning model based on a rewriting prompt that includes a command to modify one or more objects or regions in the initial image while preserving their shape. Attached Figure Description
[0011] Figure 1 This is a block diagram illustrating an example network environment according to some embodiments described herein.
[0012] Figure 2 This is a block diagram of an exemplary computing device according to some embodiments described herein.
[0013] Figure 3A An example user interface according to some embodiments described herein is illustrated, which has an initial image of a fence and a preset for removing the fence.
[0014] Figure 3B Example user interfaces with output images without fences are illustrated according to some embodiments described herein.
[0015] Figures 4A to 4B Example user interfaces with automatic suggestions for modifying regions of an initial image according to some embodiments described herein, and example user interfaces with output images generated by selecting automatic suggestions, are illustrated.
[0016] Figures 5A to 5C Example user interfaces including an initial image, an example user interface receiving an original prompt from a user, and an example user interface displaying an output image in response to a rewritten prompt are illustrated according to some embodiments described herein.
[0017] Figures 6A to 6C Example initial images of cats, including bounding boxes, and example output images are illustrated according to some embodiments described herein.
[0018] Figures 7A to 7COther example user interfaces including an initial image, other example user interfaces receiving an original prompt from a user, and other example user interfaces displaying an output image in response to a rewritten prompt, according to some embodiments described herein, are illustrated respectively.
[0019] Figures 8A to 8C Other example user interfaces including an initial image, other example user interfaces receiving an original prompt from a user, and other example user interfaces displaying an output image in response to a rewritten prompt, according to some embodiments described herein, are illustrated respectively.
[0020] Figures 9A to 9C The architecture of example machine learning models according to some embodiments described herein is illustrated.
[0021] Figure 10 A flowchart illustrating a method for generating an output image based on a rewritten prompt and an initial image according to some embodiments described herein is provided. Detailed Implementation
[0022] Overview
[0023] With the widespread use of digital cameras and smartphones, users can easily capture, store, and share a large number of digital images. As image editing software becomes more readily available and more sophisticated, users increasingly desire to modify their images in creative and complex ways. Traditional image editing tools often require significant technical skills and manual work to achieve desired results, such as changing the color of an object, altering its texture, or completely replacing it.
[0024] Recent advances in generative artificial intelligence (AI), particularly in the field of image generation, have introduced new possibilities for image manipulation. Generative AI models can generate or modify images based on text prompts. They can receive a text request from a user describing the desired change in natural language, and then attempt to produce the corresponding visual output. For example, a user could provide an image of a cat and a prompt like "make the cat orange," and the system would generate a new image with an orange cat.
[0025] However, existing generative AI models for image editing face several challenges. A significant problem is the ambiguity of user prompts. Users may provide short, context-poor prompts such as "make it shiny" or "wavy." Without understanding the image's context and the user's intent, generative AI models may misinterpret the request, leading to unexpected or meaningless results. For example, when editing an image of a car, the prompt "brand new" might be misapplied by the generative AI model, causing it to generate something other than a brand-new car. Furthermore, using these traditional generative AI models is computationally expensive, as users may have to request multiple iterations of image generation until they are satisfied with the result.
[0026] Another challenge lies in controlling the extent and nature of modifications. Sometimes users may want to change the appearance of an object while preserving its underlying structure and shape (e.g., changing the material of a car in the initial image from metal to wood). In other cases, users may want to retain the general shape of a region but change its internal structure (e.g., making a calm lake in the initial image appear to be undulating). In yet another scenario, users may want to completely replace an object with a new one, regardless of both the original shape and structure (e.g., replacing a cat with a dog). A single generative AI model trained to output a specific type of image is often insufficient to effectively handle such broad user intentions, as a model optimized for structure preservation may perform poorly in object replacement, and vice versa. A method is needed to select from multiple generative AI models based on user intentions to produce high-quality, relevant results.
[0027] The techniques described herein advantageously address these and other problems by using large language models (LLMs) to rewrite prompts and by employing different machine learning models based on the rewritten prompts. For example, the technique includes receiving an initial image and an original prompt, wherein the original prompt includes a request to modify the initial image. The original prompt defines one or more image modification tasks to be performed with respect to the initial image. The original prompt may include limited information, such as "reimagine to gold." User input may also be provided, such as selecting objects (e.g., the user taps different objects in the initial image until the object the user wants to modify is highlighted, thus selecting the object, etc.).
[0028] An initial image and original prompts (and optionally user input) are provided to an LLM or other text generation model. In some embodiments, the LLM or other text generation model may be a multimodal model that can process text, images, videos, gesture input, or other types of input as input. The LLM rewrites the prompts. The rewritten prompts correspond to the corresponding original prompts, i.e., specifying one or more of the same (image modification) tasks for modifying the initial image, but the rewritten prompts further satisfy at least one of the following criteria: the rewritten prompts are clearer instructions for the machine learning model, the rewritten prompts are more concise instructions for the machine learning model, and / or the wording / instructions of the rewritten prompts improve the performance of the machine learning model. The LLM is trained to rewrite the prompts such that the rewritten prompts satisfy at least one of the above criteria, i.e., the LLM rewrites the prompts such that the rewritten prompts satisfy at least one of the above criteria. For example, continuing the example above, if the initial image is an eagle, the original prompt "reimagine gold" can be rewritten as "reimagine to a golden statue of an eagle". If no user input is provided and / or the user does not select an object in the initial image, the LLM can generate a rewritten prompt that associates the original prompt with a unique object in the image or the most prominent object in the image (e.g., identifying an object in the foreground when other objects are in the background).
[0029] A machine learning model is selected from a set of machine learning models to generate the output image. In some embodiments, the machine learning model is selected by the media application based on the original prompt. For example, if the user selects a cat and the original prompt is "make pink," the media application selects a shape-preserving machine learning model. In some embodiments, the machine learning model is selected by an LLM (Locally Minimum Model). For example, continuing the same example, the rewritten prompt could be "make the selected region pink, preserving the shape and texture." As a result, the shape-preserving machine learning model is selected.
[0030] Different generative models may have different capabilities and / or limitations (e.g., recoloring images, adding / removing objects, artistic effects, known failure modes, preserving image structure and / or object shape, etc.). For example, if the rewrite prompt requests changing the color of an object, a structure-preserving machine learning model is selected. If the rewrite prompt requests changing the water under the bridge to cold water under the bridge, a shape-preserving machine learning model is selected. If the rewrite prompt requests replacing an object with a different object, neither structure-preserving nor shape-preserving machine learning models are selected.
[0031] Network environment
[0032] Figure 1 A block diagram illustrating an example environment 100 is shown. In some embodiments, environment 100 includes a media server 101, user devices 115a and 115n, and a large language model (LLM) 120, each coupled to a network 105. Users 125a and 125n may be associated with their respective user devices 115a and 115n. In some embodiments, environment 100 may include... Figure 1 Other servers or devices not shown in the diagram. Figure 1 In the remaining figures, reference numerals followed by a letter (e.g., "115a") indicate a reference to an element having that particular reference numeral. Reference numerals without a letter following them (e.g., "115") indicate a general reference to an embodiment of the element having that reference numeral.
[0033] Media server 101 may include a processor, memory, and network communication hardware. In some embodiments, media server 101 is a hardware server. Media server 101 is communicatively coupled to network 105 via signal line 102. Signal line 102 may be a wired connection, such as Ethernet, coaxial cable, fiber optic cable, etc., or a wireless connection, such as Wi-Fi®, Bluetooth®, or other wireless technologies. In some embodiments, media server 101 sends data to and receives data from one or more user devices 115a, 115n via network 105. Media server 101 may include media application 103a and database 199.
[0034] Database 199 can store machine learning models, training datasets, images, etc. Database 199 can also store social network data associated with user 125, user preferences, etc.
[0035] User device 115 may be a computing device including memory coupled to a hardware processor. For example, user device 115 may include a mobile device, tablet computer, mobile phone, wearable device, head-mounted display, mobile email device, portable game player, portable music player, e-reader device, or another electronic device capable of accessing network 105.
[0036] In the illustrated embodiment, user device 115a is coupled to network 105 via signal line 108, and user device 115n is coupled to network 105 via signal line 110. Media application 103 may be stored on user device 115a as media application 103b and / or on user device 115n as media application 103c. Signal lines 108 and 110 may be wired connections, such as Ethernet, coaxial cable, fiber optic cable, etc., or wireless connections, such as Wi-Fi®, Bluetooth®, or other wireless technologies. User devices 115a and 115n are accessed by users 125a and 125n, respectively. Figure 1 User devices 115a and 115n are used only as examples. Although Figure 1 Two user units 115a and 115n are illustrated, but this disclosure applies to system architectures having one or more user units 115.
[0037] Media application 103 may be stored on media server 101 and / or user device 115. In some embodiments, the operations described herein are performed on media server 101 or user device 115. For example, media application 103b on user device 115a may receive an initial image captured by user device 115a and generate an output image. In some embodiments, some operations may be performed on media server 101 and some operations may be performed on user device 115. For example, an initial image may be captured by user device 115a and transmitted to media application 103a on media server 101 along with user input and prompts, which generates an output image, which is transmitted to media application 103b on user device 115a for display.
[0038] The execution of operations is based on user settings. For example, user 125a may specify the following settings: operations should be performed on their respective user device 115a instead of on media server 101. Under such settings, the operations described herein are performed entirely on user device 115a, and no operations are performed on media server 101. Furthermore, user 125a may specify that the user's images and / or other data should be stored locally only on user device 115a and not on media server 101. Under such settings, no user data is transferred to or stored on media server 101. The transfer of user data to media server 101, any temporary or permanent storage of such data by media server 101, and the execution of operations on such data by media server 101 are only performed if the user has consented to the transfer, storage, and operation execution via media server 101. Users are provided with the option to change the settings at any time, such as enabling or disabling the use of media server 101.
[0039] If a machine learning model (e.g., a diffusion model or other type of model) is used for one or more operations, the machine learning model is stored and utilized locally on user device 115 with specific user permission. Server-side models are used only with user permission. Furthermore, a trained model can be provided for use on user device 115. During such use, on-device training of the model can be performed with user permission. Updated model parameters can be transferred to media server 101 with user permission, for example, to enable federated learning. The model parameters do not include any user data.
[0040] Media application 103 receives an initial image and an original prompt from the user. The original prompt includes a request to modify the initial image. In some embodiments, media application 103 also receives user input identifying one or more objects or regions in the initial image. For example, the user can select an object in the initial image and provide a text request to change the object's color to a different color, change the features of a region to a different feature, or replace the original object with a new object.
[0041] Media application 103 provides the original prompt as input to LLM 120. The LLM is trained / deployed to rewrite (e.g., optimize) the prompt so that the machine learning model selected to execute it can execute it correctly and more accurately; that is, the prompt is not only human-understandable but also optimized input for the machine learning model executing it. Prompt rewriting can be achieved through various methods, such as supervised learning and reinforcement learning, reinforcement learning-based prompt rewriting, instruction and example-based prompt rewriting, meta-prompts and few-shot demonstrations, automatic multi-round iterative rewriting, or any other suitable method, where any combination of these methods can also be used to achieve prompt rewriting. Rewriting the prompt improves the effectiveness and quality of the machine learning model's response. Although... Figure 1 It is exemplified as including LLM 120, but other text generation models can be used. LLM 120 in Figure 1 The LLM 120 is illustrated as separate from the media application 103; however, in some embodiments, the LLM 120 is part of the media application 103. The media application 103 receives a rewritten prompt from the LLM 120 based on the original prompt and the initial image. In some embodiments, the rewritten prompt is also based on user input, such as the identification of objects or regions in the initial image.
[0042] Media application 103 selects a machine learning model from a set of machine learning models. For example, media application 103 may select a structure-preserving machine learning model for a rewrite prompt requesting a change in the color of an object, a shape-preserving machine learning model for a rewrite prompt requesting a change in the appearance of an ocean from calm to undulating, or a non-structure- and non-shape-preserving machine learning model for a rewrite prompt requesting a replacement of an original object with a new object or the addition of an additional object to the initial image. The selected machine learning model generates an output image in response to the rewrite prompt.
[0043] In some embodiments, media application 103 may be implemented using hardware including a central processing unit (CPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a machine learning processor / coprocessor, any other type of processor, or a combination thereof. In some embodiments, media application 103a may be implemented using a combination of hardware and software.
[0044] Computing device
[0045] Figure 2This is a block diagram of an example computing device 200 that can be used to implement one or more features described herein. The computing device 200 can be any suitable computer system, server, or other electronic or hardware device. In one example, the computing device 200 is a media server 101 for implementing media application 103a. In another example, the computing device 200 is a user device 115.
[0046] In some embodiments, the computing device 200 includes a processor 235, a memory 237, an input / output (I / O) interface 239, a display 241, a camera 243, and a storage device 245, all coupled via a bus 218. The processor 235 may be coupled to the bus 218 via signal line 222, the memory 237 may be coupled to the bus 218 via signal line 224, the I / O interface 239 may be coupled to the bus 218 via signal line 226, the display 241 may be coupled to the bus 218 via signal line 228, the camera 243 may be coupled to the bus 218 via signal line 230, and the storage device 245 may be coupled to the bus 218 via signal line 232.
[0047] Processor 235 may be one or more processors and / or processing circuits for executing program code and controlling the basic operations of computing device 200. "Processor" includes any suitable hardware system, mechanism, or component that processes data, signals, or other information. A processor may include systems having: a general-purpose central processing unit (CPU) with one or more cores (e.g., in a single-core, dual-core, or multi-core configuration), multiple processing units (e.g., in a multiprocessor configuration), a graphics processing unit (GPU), a field-programmable gate array (FPGA), an application-specific integrated circuit (ASIC), a complex programmable logic device (CPLD), a dedicated circuit system for implementing functionality, a dedicated processor for implementing processing based on neural network models, neural circuits, a processor optimized for matrix computations (e.g., matrix multiplication), or other systems. In some embodiments, processor 235 may include one or more coprocessors implementing neural network processing. In some embodiments, processor 235 may be a processor that processes data to produce probabilistic outputs (e.g., the output produced by processor 235 may be inaccurate or may be accurate within a range from the expected output). Processing is not required to be geographically restricted or time-limited. For example, a processor may perform its functions in real-time, offline, in batch mode, etc. The different parts of the process can be executed at different times and in different locations by different (or the same) processing systems. The computer can be any processor that communicates with memory.
[0048] Memory 237 is typically provided in computing device 200 for access by processor 235 and can be any suitable processor-readable storage medium suitable for storing instructions for execution by processor or multiple processors and located separately from and / or integrated with processor 235, such as random access memory (RAM), read-only memory (ROM), electrically erasable read-only memory (EEPROM), flash memory, etc. Memory 237 may store software operated on computing device 200 by processor 235, including media application 103.
[0049] The memory 237 may include an operating system 262, other applications 264, and application data 266. Other applications 264 may include, for example, image library applications, image management applications, image gallery applications, communication applications, web hosting engines or applications, media sharing applications, etc. One or more methods disclosed herein can operate in various environments and platforms, for example, as a standalone computer program that can run on any type of computing device, as a web application with web pages, as a mobile application (“app”) running on a mobile computing device, etc.
[0050] Application data 266 may be data generated by other applications 264 or the hardware of computing device 200. For example, application data 266 may include images used by an image gallery application and user actions identified by other applications 264 (e.g., social networking applications, etc.).
[0051] I / O interface 239 provides functionality that enables the computing device 200 to interface with other systems and devices. The interfaced device may be included as part of the computing device 200, or it may be separate and communicate with the computing device 200. For example, network communication devices, storage devices (e.g., memory 237 and / or storage device 245), and input / output devices may communicate via I / O interface 239. In some embodiments, I / O interface 239 may be connected to interface devices such as input devices (keyboard, pointing device, touchscreen, microphone, scanner, sensor, etc.) and / or output devices (display device, speaker device, printer, monitor, etc.).
[0052] Some examples of docked devices that can be connected to I / O interface 239 may include display 241, which can be used to display content (e.g., images, videos, and / or user interfaces for output applications as described herein) and receive touch (or gesture) input from a user. For example, display 241 may be used to display a user interface including graphical guidance on a viewfinder. Display 241 may include any suitable display device, such as a liquid crystal display (LCD), a light-emitting diode (LED) or plasma display, a cathode ray tube (CRT), a television, a monitor, a touch screen, a 3D display, or other visual display device. For example, display 241 may be a flat panel display mounted on a mobile device, multiple displays embedded in an eyeglass shape factor or head-mounted device, or a monitor screen of a computer device.
[0053] Camera 243 can be any type of image capture device capable of capturing images and / or videos. In some embodiments, camera 243 captures images or videos, and I / O interface 239 transmits the images or videos to media application 103.
[0054] Storage device 245 stores data related to media application 103. For example, storage device 245 may store training datasets including labeled images, machine learning models, outputs from machine learning models, etc.
[0055] Figure 2 An example media application 103 stored in memory 237 is shown. The media application includes a user interface module 202, a splitter 204, a prompting engine 206, and a machine learning module 208. The user interface module 202, the splitter 204, the prompting engine 206, and the machine learning module 208 may be implemented as code or other computer-readable instructions executable by one or more processors, such as processor 235.
[0056] The user interface module 202 generates graphical data for displaying a user interface including images. The user interface module 202 receives an initial image. The initial image can be received from the camera 243 of the computing device 200 or from the media server 101 via the I / O interface 239.
[0057] Before the initial image is processed, the user interface presents the user with a request for their consent to modify the image. In some embodiments, such consent may be obtained by the media application 103 for all future images at once. The user is provided with an option to withdraw such one-time consent as well as an option to request consent for each image. Unless the user provides their consent, the user interface module 202 does not collect or utilize user information.
[0058] The initial image includes one or more objects. In some embodiments, the initial image also includes one or more human subjects (e.g., one or more objects in the initial image may correspond to human subjects, such as a human face, body, etc.). In some embodiments, the user interface module 202 receives user input that selects one or more objects in the initial image. User input may include moving a finger around one or more objects in the initial image (e.g., by drawing a circle or other shape that at least approximately surrounds the object), moving a finger over one or more objects, tapping one or more objects in the initial image, providing text recognition for one or more images, etc.
[0059] The user interface may highlight one or more objects in response to receiving user input. In some embodiments, where a tap can be associated with multiple objects, different numbers of taps may cause the user interface to highlight different objects. For example, in an initial image that is a beach scene with a bucket in front of a sandcastle, a first tap on the bucket / sandcastle area highlights the bucket first, a second tap highlights the sandcastle, and a third tap highlights both the bucket and the sandcastle.
[0060] The user interface includes options for providing text requests associated with one or more selected objects in the initial image. For example, the user interface may include text fields for users to directly enter text requests (also known as raw prompts), text fields with presets, microphone buttons for providing audio input that is converted into text requests, etc.
[0061] In some embodiments, the user interface module 202 generates presets that are displayed along with the initial image. The user interface module 202 generates presets as selectable icons that, when selected, cause the generation of an output image that satisfies the description in the preset. In some embodiments, the user interface module 202 provides the same set of presets in response to a user selecting an edit button and / or a suggestion button. In some embodiments, the set of presets is customized based on parameters such as the type of objects and regions in the initial image. The user interface module 202 can receive segmentation information from the segmenter 204 that divides the initial image into different parts. The user interface module 202 can generate different presets based on the segmentation. In some embodiments, the user interface module 202 performs object recognition to identify the types of objects in different segments of the initial image. For example, the initial image may be segmented into a background and have presets related to the background (e.g., changing the sky to a different type of sky, changing buildings to a different type of building, changing water bodies to a different type of water body, etc.), one or more objects, etc.
[0062] In some embodiments, the preset includes selectable buttons or links for erasing objects in the initial image, adding new objects to the initial image, changing the material or color of objects in the initial image, enhancing the initial image (e.g., by correcting the hue of the initial image, deblurring objects in the initial image, removing reflections in the initial image, etc.), replacing the background of the initial image, changing the subject in the initial image (e.g., changing the subject's expression, changing the subject's features, changing the subject's clothing, etc.), and removing fences from the initial image.
[0063] The user interface module 202 generates graphical data for displaying the output image. In some embodiments, the user interface module 202 includes options for enabling multiple edits of the initial image. For example, a user can provide a first initial prompt and receive a first output image, a user can provide a second initial prompt and receive a second output image, and so on, until the user is satisfied with the result. The user interface may also include options for sharing the output image, adding the output image to an album, adding a title to the output image, etc.
[0064] In some embodiments, the user interface module 202 generates a text response that is displayed along with the output image based on the original prompt and the rewritten prompt. For example, if the user provides the original prompt "make it silver" and the rewritten prompt is "make the tree silver by using a structure-preserving machine-learning model", then the text response displayed along with the output image is "we have changed the color of the tree to silver".
[0065] Figure 3AAn example user interface 300 according to some embodiments described herein is illustrated, having an initial image 302 including a fence 304 in front of a body 306 and a preset 308 for removing the fence. User interface module 202 performs object recognition on the initial image 302 and identifies that the initial image 302 includes the fence 304 and the body 306. User interface module 202 generates a “fence removal” preset 308, which, when selected, instructs a machine learning model to generate an output image without the fence 304. In some embodiments, in response to the user selecting the “fence removal” preset 308, prompt engine 206 (via LLM in some embodiments) generates a rewritten prompt with instructions to remove the fence 304 using an unstructured and non-shape-preserving machine learning model.
[0066] Figure 3B An example user interface 350 according to some embodiments described herein is illustrated, the example user interface having an output image 355, the output image including a subject 357 (corresponding to...). Figure 3A The main body 306) but excluding Figure 3A The fence 304. The user interface module 202 receives output from the machine learning module 208 (e.g., from an unstructured and non-shape-preserving machine learning model) and displays the output image 355. The user can continue to edit the output image 355, or select the "save" button 359 to save the output image.
[0067] In some embodiments, the user interface module 202 generates automatic suggestions. Automatic suggestions differ from preset suggestions in that they include suggestions that can be modified. In some embodiments, the user interface module 202 generates automatic suggestions based on objects and / or regions in an initial image, based on the most frequently suggested requests (based on a specific user or based on all users), etc.
[0068] Figures 4A to 4B Example user interface 400 with automatic suggestions for modifying regions of an initial image 402 according to some embodiments described herein, and example user interface 450 with an output image 452 generated by selecting the automatic suggestions are illustrated.
[0069] The user interface module 202 receives segmentation information from the segmenter 204 and divides the initial image 402 into a background region 404 and a foreground region 406. The background region 404 includes clouds 408 and is delineated with lines 410 to indicate the areas affected by the changes. The foreground region 406 includes the subject 412.
[0070] User interface module 202 generates suggestion 414 "Reimagine as clear blue skies", where "clear blue skies" is determined by user interface module 202 based on the identification of a sky with clouds 408 in the background area 404. Suggestion 414 is editable, allowing the user to change it, for example, from "clear blue skies" to "sunset", "dark and stormy", etc.
[0071] In response to user Figure 4A Select suggestion 414, and user interface module 202 generates a display. Figure 4B The output image 452 contains graphic data. Output image 452 has a foreground region 456 and a background region 454. The background region 454 has a clear blue sky, and... Figure 4A The cloud 408 in the image is not part of the output image 452. The foreground region 456 includes the same subject 462.
[0072] Figure 5A An example user interface 500, including an initial image 502, is illustrated according to some embodiments described herein. The initial image 502 includes a human subject 504 and a white dog 506. The user interface 500 also includes a reimagine button 508. Selecting the reimagine button 508 causes the user interface module 202 to generate... Figure 5B The user interface 525 is shown in the example.
[0073] Figure 5B An example user interface 525 according to some embodiments described herein is illustrated, which receives user input 531 on an initial image 527 and a text field for receiving an initial prompt 533 from the user. The user provides user input 531 by circling the dog and adding "Pink" to the text field to create the initial prompt 533. By circling the dog and adding "Pink" to the text field 533, the user indicates that they want to change the dog to a pink dog. The user selects an arrow button 545 to generate an output image.
[0074] As described in more detail below, the prompting engine 206 receives an original prompt and an initial image provided by the user. In embodiments where the user provides user input, the prompting engine 206 also receives user input. In some embodiments, the prompting engine 206 specifies a selected machine learning model. The LLM generates a rewritten prompt based on the initial image, the original prompt, and user input (if available). Continuing Figure 5A and Figure 5BIn the example above, prompt engine 206 rewrites the original prompt to combine the user input 531 (selecting a dog) with the original prompt 533, resulting in a rewritten prompt "A pink dog". In some embodiments, the rewritten prompt also specifies the selected machine learning model. For example, the rewritten prompt could include "a pinkdog generated by a structure-preserving machine-learning model". In some embodiments, the rewritten prompt is not visible to the user. In some embodiments, the rewritten prompt is visible to the user as guidance on how to draft future requests.
[0075] In embodiments where the user's face is used as part of the original prompt and / or rewritten prompt, the user is provided with guidance on the use of user information, how the user information can be used to generate images (e.g., including generated images including the face), and how the user information can be stored. If the user chooses to accept the applicable terms and conditions and grants permission, the process of generating the output image is initiated. The user may choose not to use user features, in which case no image is captured. User information is only part of the creation process in specific states / countries where the creation, storage, and use of user information are permitted and comply with applicable regulations. In some embodiments, the user's image is uploaded for use in creating the output image. Once the output image is generated, the machine learning module 208 removes the captured image of the user. In some embodiments, user-associated identification information is removed from the output image. The output image is stored locally on the user's device and used exclusively with the user's permission and in accordance with applicable regulations.
[0076] Figure 5C An example user interface 550 is illustrated according to some embodiments described herein, showing an output image 552 that displays prompts for rewriting. The output image 552 includes a person 554 and a pink dog 556. In this example, the user interface 550 also includes the statement 555 "we have changed the dog to pink" and a "Reimagine" button 557, allowing the user to further modify the output image if they are not satisfied with the result. The user can save a copy of the output image by selecting the "Save a copy" link 558, undo the changes by selecting the "Undo" button 560, or select the "Done" button 562.
[0077] Segmenter 204 segments the initial image. In some embodiments where the user selects one or more objects or regions, segmenter 204 generates a user-selected mask. In some embodiments, segmenter 204 generates a segmentation mask that identifies object pixels or region pixels associated with one or more objects or regions based on segmenting one or more objects or regions.
[0078] Segmenter 204 can automatically or in response to user input to segment one or more objects in an initial image. For example, segmenter 204 can automatically segment different objects and / or regions in an initial image to create a segmentation mask. In another example, the user interface receives user input identifying objects to be modified, removed, and / or replaced, and segmenter 204 segments the objects in response to object selection to create a user-selected mask. Segmentation refers to determining the pixels of an image that belong to a particular object. In some embodiments, segmenter 204 generates a segmentation map that associates identity with each pixel in the initial image as belonging to a particular object or a portion thereof (e.g., face, body, object, etc.).
[0079] Segmenter 204 can perform segmentation by detecting objects in the initial image. Objects can be people, animals, cars, buildings, etc. People can be the main subject of the initial image, or they may not be (e.g., a bystander captured in the initial image). Bystanders can include people walking, running, cycling, standing behind a subject, or otherwise located within the initial image. In different examples, bystanders may be in the foreground (e.g., a person walking in front of the camera), at the same depth as the subject (e.g., a person standing next to the subject), or in the background. In some examples, there may be more than one bystander in the initial image. Bystanders can be humans in any pose (e.g., standing, sitting, crouching, lying down, jumping, etc.). Bystanders can be facing the camera, at an angle to the camera, or with their backs to the camera.
[0080] The segmenter 204 can detect the type of an object by performing object recognition, comparing it with prior objects such as people, vehicles, and buildings, in order to identify the expected shape of the object and thus determine whether a pixel is associated with the selected object or with the background.
[0081] In some embodiments, segmenter 204 generates a segmentation mask or a user-selected mask based on the segmentation, which indicates the pixels to be modified. The segmentation mask or user-selected mask is used by a machine learning model to determine the pixels in the initial image to be modified based on the rewritten cue. In some embodiments, the segmentation mask or user-selected mask corresponds to the segmentation such that the mask identifies the selected object or the selected region. In some embodiments where the original cue provided by the user includes a request to replace an object, segmenter 204 generates a segmentation mask corresponding to a bounding box having x, y coordinates and a scale. The bounding box may be a minimum bounding box, defined as the smallest rectangle that captures all pixels associated with the object.
[0082] Figure 6A An example initial image 600 of a cat 605 according to some embodiments described herein is illustrated. The user provides the following prompt in text field 610: “Change the cat into a turtle.” Segmenter 204 generates a minimum bounding box corresponding to the cat and generates a segmentation mask from this minimum bounding box. Segmenter 204 generates a bounding box mask from the minimum bounding box, which indicates the region in the initial image to be replaced by a second object in the output image. The second object is not limited to the structure and / or shape of the first object.
[0083] Figure 6B Here is an example initial image 625 of cat 630 and a minimum bounding box 635 including cat 630. The minimum bounding box 635 includes all pixels associated with cat 630, which result in the formation of the box. Using bounding boxes to label pixels associated with a replacement object in the output image advantageously identifies regions of the replacement object without limiting the replacement object to characteristics associated with the original object. For example, if a machine learning model receives an image of cat 630... Figure 6A The segmentation mask corresponds to the pixels of the cat in the image, and the turtle can have the attributes of the cat (e.g., shape, texture, etc.). Conversely, Figure 6C Example output image 650 of a turtle 655 having attributes of a turtle rather than a cat, according to some embodiments described herein, is shown. The user can save a copy of the output image (not shown), undo changes (not shown), or select the finish button (not shown).
[0084] In some embodiments, segmenter 204 generates a depth map of the initial image. The depth map is a representation of distance or depth information for each pixel in the initial image. The depth map may be a two-dimensional array where each pixel contains a value representing the distance from a camera (e.g., camera 243 if computing device 200 captured the initial image) to a corresponding point in the scene. The depth map provides a continuous representation of the depth information of the scene captured in the initial image. The depth map can be generated using a depth sensor (either as metadata generated during image capture if available in the initial image, or by deriving depth from pixel values using depth estimation techniques).
[0085] Segmenter 204 can generate a user-selected mask or segmentation mask by generating superpixels for the image and matching the superpixel centroids with depth map values to cluster the detections based on depth. More specifically, depth ranges can be determined using depth values in the masked regions, and superpixels falling within these depth ranges can be identified. Another technique for generating the user-selected mask or segmentation mask includes weighting the depth values based on their proximity to the user-selected mask or segmentation mask, where the weights are represented by a distance transform map.
[0086] In some embodiments, the segmenter 204 generates a retention mask that identifies pixels to be retained in the initial image. In some embodiments, the retention mask is generated for pixels corresponding to a portion of a subject (such as a face, hand, whole body, etc.).
[0087] In some embodiments, the segmenter 204 may specify a circuit configuration that enables the processor 235 to apply a machine learning model (e.g., for a programmable processor, for a field-programmable gate array (FPGA), etc.). In some embodiments, the segmenter 204 may include software instructions, hardware instructions, or a combination thereof. In some embodiments, the segmenter 204 may provide an application programming interface (API) that can be used by the operating system 262 and / or other applications 264 to invoke the segmenter 204 (e.g., to apply a machine learning model to application data 266 to output a mask).
[0088] The segmenter 204 uses training data to generate a trained machine learning model. For example, training data for generating segmentation masks may include pairs of initial images with one or more objects or regions and output images with one or more segmentation masks. Training data for generating user-selected masks may include pairs of initial images with user-selected objects or regions and output images with one or more user-selected masks. Training data for generating preserved masks may include pairs of initial images with one or more subjects and output images with one or more preserved masks.
[0089] Training data can be obtained from any source (e.g., a data repository specifically tagged for training, data licensed for use as machine learning training data, etc.). In some embodiments, the training may occur on a media server 101 that provides training data directly to user device 115, the training may occur locally on user device 115, or a combination of both.
[0090] In some embodiments, segmenter 204 uses weights obtained from another application and not edited / transmitted. For example, in these embodiments, a trained model may be generated (e.g., on different devices) and provided as part of segmenter 204. In various embodiments, the trained model may be provided as a data file including a model structure or form (e.g., defining the number and type of neural network nodes, the connectivity between nodes, and the organization of nodes into multiple layers) and associated weights. Segmenter 204 may read the data file of the trained model and implement a neural network with node connectivity, layers, and weights based on the model structure or form specified in the trained model.
[0091] A trained machine learning model can include one or more model forms or structures. For example, a model form or structure can include any type of neural network, such as a linear network, a deep learning neural network that implements multiple layers (e.g., "hidden layers" between the input and output layers, where each layer is a linear network), a convolutional neural network (e.g., a network that splits or partitions input data into multiple parts or tiles, uses one or more neural network layers to process each tile separately and aggregates the results of the processing from each tile), a sequence-to-sequence neural network (e.g., a network that takes sequential data such as words in a sentence, frames in a video, etc., as input and produces a sequence of results as output), and so on.
[0092] The model form or structure can specify the connectivity between various nodes and the organization of nodes into layers. For example, nodes in a first layer (e.g., an input layer) can receive data as input data or application data. Such data may include, for example, one or more pixels per node (e.g., when a trained model is used for analysis of, for example, an initial image). Subsequent intermediate layers can receive the outputs of nodes in the previous layer as input, according to the connectivity specified in the model form or structure. These layers may also be referred to as hidden layers. For example, a first layer may output a segmentation between foreground and background. The final layer (e.g., an output layer) produces the output of the machine learning model. For example, an output layer may receive the initial image to segment the foreground and background, and output whether a pixel is part of a mask. In some embodiments, the model form or structure also specifies the number and / or type of nodes in each layer.
[0093] In various embodiments, the trained model may include one or more models. One or more models may include multiple nodes arranged in layers according to a model structure or form. In some embodiments, a node may be a computation node without memory (e.g., configured to process an input unit to produce an output unit). Computations performed by a node may include, for example, multiplying each node input in a plurality of node inputs by a weight to obtain a weighted sum, and adjusting the weighted sum with a bias or intercept value to produce a node output. In some embodiments, computations performed by a node may also include applying a step / activation function to the adjusted weighted sum. In some embodiments, the step / activation function may be a nonlinear function. In various embodiments, such computations may include operations such as matrix multiplication. In some embodiments, computations performed by multiple nodes may be performed in parallel, for example, using multiple processor cores of a multi-core processor, an individual processing unit using a graphics processing unit (GPU), or a dedicated neural circuit system. In some embodiments, a node may include memory (e.g., capable of storing and using one or more earlier inputs while processing subsequent inputs). For example, a node with memory may include a Long Short-Term Memory (LSTM) node. An LSTM node may use memory to maintain a “state” that allows the node to act like a finite state machine (FSM).
[0094] In some embodiments, the trained model may include embeddings or weights for individual nodes. For example, the model may be started as a plurality of nodes organized into layers as specified by the model form or structure. Upon initialization, corresponding weights may be applied to the connections between each pair of nodes connected in the model form—e.g., nodes in consecutive layers of a neural network. For example, the corresponding weights may be randomly assigned or initialized to default values. The model can then be trained (e.g., using training data) to produce results.
[0095] Training can include the application of supervised learning techniques. In supervised learning, training data can include multiple inputs (e.g., an initial image, user input, etc.) and the corresponding ground truth output for each input (e.g., a user-selected mask that correctly identifies the ground truth of pixels corresponding to a selected object, a segmentation mask that correctly identifies the ground truth of pixels corresponding to an object or region, or a ground truth preservation mask that correctly identifies a portion of a subject in each image (such as the subject's face)). Based on a comparison between the model's output and the ground truth output, the values of the weights are automatically adjusted (e.g., in a way that increases the probability that the model will produce a ground truth output for the image).
[0096] In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In some embodiments, the trained model may include, for example, a fixed set of weights downloaded from a server providing the weights. In various embodiments, the trained model includes a set of weights or embeddings corresponding to the model structure. In embodiments where data is omitted, the segmenter 204 may generate a trained model based on previous training (e.g., by the developer of segmenter 204, by a third party, etc.).
[0097] In some embodiments, a trained machine learning model receives an initial image having one or more selected objects. In some embodiments, the trained machine learning model outputs one or more user-selected masks that identify object pixels in the initial image associated with one or more objects. In some embodiments, the trained machine learning model receives the initial image and outputs one or more segmentation masks. In some embodiments, if the initial image includes one or more human subjects, the trained machine learning model generates one or more preservation masks corresponding to the one or more human subjects. For example, the one or more preservation masks may be for the faces of one or more subjects.
[0098] The prompting engine 206 receives an initial image and a basic prompt from the user interface module 202. In some embodiments, the prompting engine 206 also receives user input from the user interface module 202, such as selection of one or more objects and / or regions.
[0099] The prompting engine 206 (e.g., implemented using an LLM or another text generation model as a backend) generates a rewritten prompt based on an initial image, the original prompt, and user input (if applicable). The rewritten prompt is designed to make user requests for the output image compatible with the machine learning image generation model (e.g., including generation context, ensuring the prompt is within model constraints, including constraints on generation, etc.). In some embodiments, the prompting engine 206 adds the name of a selected object and / or region to the rewritten prompt. For example, the prompting engine 206 receives an initial image of an eagle and the original prompt stating "Reimagine to a cartoon look," and outputs a rewritten prompt stating "Reimage to a cartoon eagle."
[0100] In some embodiments, the description of the selected object can be specific. For example, prompting engine 206 receives an original prompt stating "ice" and an initial image of a seal in the water, and outputs a rewritten prompt stating "replace the background to water surface covered in broken ice". In some embodiments, the rewritten prompt may include commands for multiple images. For example, prompting engine 206 receives an original prompt about a man riding a bicycle on a steep slope, stating "cliff and ominous clouds". Prompting engine 206 rewrites the prompt to "replace the background to the cliff of a mountain with a very sharp drop under a sky with ominous clouds".
[0101] In some embodiments, the prompting engine 206 implements a machine learning model, such as an LLM (e.g., a text generation LLM, a multimodal LLM, etc.), which uses natural language processing (NLP) to provide a conversational response to a text query. In some embodiments, the LLM is stored on the computing device 220 or on a separate server (such as...). Figure 1 On LLM 120).
[0102] In some embodiments, the machine learning model includes an encoder that generates representations of the original cue, an initial image, and user input. For example, the encoder receives an initial image of the Golden Gate Bridge and the original cue “Reimagine to icy,” as well as user input selecting water in the initial image. The machine learning model also includes a transformer for generating embeddings of the original cue, the initial image, and the user input, and a self-attention mechanism for aggregating information from the embeddings to generate the rewritten cue. Continuing the example above, the transformer outputs the rewritten cue “Reimagine to icy water beneath a bridge on a cold winter day.”
[0103] In some embodiments, the prompting engine 206 includes a multilingual LLM capable of receiving input in languages other than English and outputting prompts in the language of the original prompt or in a language compatible with the image generation machine learning model.
[0104] The prompting engine 206 selects a machine learning model from a set of machine learning models to generate an output image based on the original prompt and / or rewritten prompt. In some embodiments, the prompting engine 206 includes a base LLM for selecting the machine learning model. In some embodiments, the prompting engine 206 uses an LLM that also generates rewritten prompts.
[0105] In some embodiments, the rewritten prompt includes a command specifying which machine learning model from the set of machine learning models to use. In some embodiments, the set of machine learning models includes three types of machine learning models: structure-preserving machine learning models, shape-preserving machine learning models, and unstructured and non-shape-preserving machine learning models. In various embodiments, two, three, four, or any other number of machine learning models may be utilized. Different image generation machine learning models may be implemented using different techniques (e.g., diffusion models, models trained using generative adversarial networks, or other types of models). In different embodiments, different models may have different reliability, different image generation capabilities, different computational costs, etc., and the selection of a model may be based on one or more of these model properties.
[0106] In some embodiments, when prompted to rewrite one or more objects or regions in an initial image while preserving the structure and shape of one or more objects or regions, prompt engine 206 selects structure preservation machine learning model. Figure 5C This includes an example of a rewritten prompt that requests a modification to the dog, changing its color from white to pink.
[0107] Structure-preserving machine learning models are used to change the color of objects because they are trained to preserve the structure of the objects modified for the output image. These models use depth control as a parameter during image generation. In some embodiments, the model is trained to learn a joint embedding space where the feature vectors of the input text are closely correlated with the feature vectors of the initial image, and images with similar meanings are close to each other in the learned latent space.
[0108] If a rewritten prompt requests modification of one or more objects or regions in the initial image, thereby altering the structure of one or more objects or regions, then the structure-preserving machine learning model does not satisfy the rewritten prompt. For example, if a prompt requests changing an image of a lizard found in nature to a cartoon lizard, then while the shape of the lizard remains unchanged, details such as the lizard's texture are altered.
[0109] For a rewritten prompt requesting modification of one or more objects or regions in an initial image while preserving the shape of one or more objects or regions, prompt engine 206 selects a shape-preserving machine learning model. In some embodiments, the shape-preserving machine learning model modifies the structure of one or more objects or regions while preserving the shape and without using depth control.
[0110] Go to Figure 7A An example user interface 700 according to some embodiments described herein is illustrated, which includes an initial image 702. The initial image 702 includes a sailboat 704 and calm water 706. The user selects a reimagine button 708 to initiate a process of modifying the initial image 702 using a machine learning model.
[0111] Figure 7B User interface 725 is illustrated, which includes an initial image 727 and a text field 735 where the user has entered "Wavy". Hint engine 206 generates a rewritten hint from the original hint, associating wave undulations with water rather than a sailboat, because wave undulations are a property typically associated with water rather than sailboats. The user selects arrow button 745 to generate an output image.
[0112] In various embodiments, an LLM can perform inference tasks to generate rewritten suggestions. For example, an LLM could provide the query "The user has provided a prompt that states..." wavyThe prompt is in the context of an image modification request. The initial image is a sailboat in calm water in an ocean. There are no other objects in the image. Please rewrite the user prompt based on this information. In response, the LLM can perform reasoning (e.g., determine that the state "wavy" is often associated with water, including the ocean or lake where the sailboat might be sailing, rather than with the sailboat itself), and thus determine that the rewritten prompt indicates that the ocean should be wavy in the output image. In contrast, if the user enters the text statement "sails full," the LLM can infer that this text corresponds to a sailboat's sails being fully inflated (e.g., due to strong winds) and rewrite the prompt as "a sailboat in the ocean having its sails full." In another example, if the user enters the text statement "topsy-turvy," If the prompt is "ride (a bumpy journey)," then LLM can rewrite it as "a sailboat in strong ocean waves, the boat not level with the ocean surface." LLM can perform such reasoning tasks by mapping user input text (with additional context) into a latent space to generate output text that responds to the reasoning tasks included in the input to LLM.
[0113] Figure 7C A user interface 750 according to some embodiments described herein is illustrated, the user interface including an output image 752 that satisfies a rewrite prompt. In this example, the rewrite prompt is illustrated in a text field 757 as “A wavy ocean beneath a boat”, which is visible to the user; however, in some embodiments, the rewrite prompt is used as part of the image generation process and is not shown to the user.
[0114] The output image 752 responds to the rewrite prompt, as it includes the undulating ocean 754 beneath the boat 756. If the user is satisfied with the output image 752, the user can select the "save a copy" link 758. If the user wants to undo or redo the generation, the user can select the arrow 760. The user can also select the "done" button 762 to complete the editing of the output image 752.
[0115] In this example, prompt engine 206 selects the shape-preserving machine learning model to generate the output image because the shape of the water is preserved, while the structure of the water changes from calm to undulating waves. The shape-preserving machine learning model does not use depth control as a parameter because changing the structure of a region also causes a change in the depth of that region.
[0116] If a rewrite prompt requests the replacement of one or more objects or regions in the initial image, structure-preserving and shape-preserving machine learning models do not satisfy the rewrite prompt because the shape and structure of one or more objects or regions in the initial image can be modified. For example, if a user requests to replace a glass with a mug, the glass and the mug have different shapes and structures. If a structure-preserving or shape-preserving machine learning model is used to generate the output image, the output image could include two mugs stacked to resemble the shape of a glass. Conversely, if a non-structure-preserving and non-shape-preserving machine learning model is used to generate the output image, the output image includes a mug with the shape and structure of a mug, unconstrained by the properties of the glass in the image.
[0117] In some embodiments, when a rewritten prompt requests that one or more objects or regions in the initial image be replaced with one or more new objects or regions, the prompt engine 206 selects a non-structure and non-shape-preserving machine learning model. In some embodiments, when a rewritten prompt requests that additional objects be added to the initial image, the prompt engine 206 selects a non-structure and non-shape-preserving machine learning model.
[0118] Figure 8A An example user interface 800 according to some embodiments described herein is illustrated, the example user interface including an initial image 802 of a car 806. A user can select a reimagine button 808 to initiate a process of generating an output image using a machine learning model.
[0119] Figure 8B An example user interface 825 is illustrated, which includes an initial image 827 and a text field 835 in which the user has provided the following original prompt: "A blue flowered bush". The user selects the arrow button 845 to generate an output image.
[0120] The prompt engine 206 generates a rewritten prompt with the following content: "A car replaced with a blueflowered bush". Figure 8C An example user interface 850 is illustrated, which includes an output image 852 with a blue flowering shrub 854 and a rewrite prompt in a text field 855. In some embodiments, the rewrite prompt is not provided for the user to view. The user can save a copy of the output image by selecting the "Save a copy" link 858, undo changes by selecting the undo button 860, or select the finish button 862.
[0121] In some embodiments, the user interface includes a request to the user to confirm that the output image meets the original prompt. For example, the original prompt might be "Change the sky to cloudy". The user interface module 202 may provide the output image with an option to regenerate the output image using a regenerate button, text fields (such as...) Figure 8C The text field 855 and / or the statement "I changed the sky, is it OK?" can be used. The user can provide follow-up suggestions, such as "No, I meant feather clouds." The suggestion engine 206 can generate a subsequent rewritten suggestion based on the follow-up suggestion. The selected machine learning model generates a subsequent output image based on the follow-up suggestion or the rewritten suggestion, and the user interface module 202 provides the subsequent output image to the user. The user can continue to modify the subsequent output image until the user is satisfied.
[0122] In some embodiments, the prompting engine 206 generates a rewritten prompt for a preset. For example, if the user selects a preset stating "fence removal," the prompting engine 206 can generate a rewritten prompt specific to the initial image. Figure 7A If the fence removal prompt 308 is generated, the prompt engine 206 can generate a rewritten prompt stating "remove a fence from the image so that the baseball player is visible using the non-structure and non-shape preserving machine-learning model".
[0123] Machine learning module 208 trains a machine learning model to generate an output image based on a rewritten prompt and an initial image. In some embodiments, machine learning module 208 receives commands from prompt engine 206 to generate an output image based on a machine learning model selected by prompt engine 206, an initial image, a rewritten prompt, and user input (if available). In some embodiments, the machine learning model is selected from a structure-preserving machine learning model, a shape-preserving machine learning model, or a non-structure and non-shape-preserving machine learning model.
[0124] The machine learning module 208 trains and implements machine learning models to receive an initial image and a text request to generate an output image; it also uses a segmentation mask or a user-selected mask as input and / or retains a mask.
[0125] The diffusion model generates an output image that satisfies the text request and does not include object pixels associated with human subjects. In some embodiments, the diffusion model receives an empty mask as input, which identifies all pixels in the initial image as not associated with humans (regardless of whether the initial image includes humans). Due to the use of the empty mask, the machine learning module 208 generates an output image that does not include human pixels.
[0126] In some embodiments where the initial image includes a human subject (either as a selected object or present in the image), the machine learning model also receives a preserving mask from the segmenter 204. The preserving mask is used to prevent the machine learning model from modifying the human subject during the generation of the output image.
[0127] In some embodiments, the machine learning model is a diffusion model, and the machine learning module 208 trains the diffusion model using a two-step process to generate an output image. First, the diffusion model is trained to perform a forward diffusion process on an initial image, where Gaussian noise with variance is added to obtain a noisy image. Gaussian noise with variance is added to obtain an image with progressively increasing noise until the final noisy image is achieved. Second, the diffusion model is trained to perform a reverse diffusion process, which uses a convolutional neural network (CNN) to transform the final noisy image into a meaningful output (e.g., an output image).
[0128] The machine learning module 208 trains a diffusion model to perform forward diffusion using training data including the initial image. The machine learning module 208 converts the initial image into a tensor. A tensor is an array of bytes with any number of dimensions. Tensors can be described as having arbitrary shapes because they can have any number of dimensions. The machine learning module 208 parses the bytes in the tensor to convert them into pixel data for the red, green, and blue (RGB) color channels.
[0129] Machine learning module 208 can sample noise to match the shape (dimension) of the initial image. Machine learning module 208 can sample random diffusion times and use these random diffusion times to generate noise and signal rates according to a diffusion timetable. Machine learning module 208 applies weights to the initial image to generate a noisy image. In some embodiments where the diffusion model is used to generate an output image from text, each forward diffusion step predicts noise from the noisy image and text embeddings generated from the text.
[0130] Machine learning module 208 calculates the loss (e.g., mean absolute error) between the predicted noise and the noise from the ground truth image, and takes a gradient step with respect to this loss function. After the gradient step, the neural network weights of the (trained) diffusion model are updated to a weighted average of the existing weights and the trained neural network weights.
[0131] The machine learning module 208 can train a diffusion model to perform inverse diffusion and denoise noisy images, such that it satisfies text requests by instructing a neural network to predict noise and then undoing the noise addition operation using the noise rate and signal rate. The diffusion model includes a CNN comprising convolutional layers, the output of which is used as input to subsequent layers. The convolutional layers include downsampling blocks and upsampling blocks, in which the initial image is spatially compressed but channel-wise expanded, and in the upsampling blocks, the representation is spatially expanded while reducing the number of channels.
[0132] Machine learning module 208 provides noise variance and a noisy image, as described by a tensor, as input to the first convolutional layer in the CNN to increase the number of channels. The noise variance and the noisy image are cascaded across channels. In some embodiments, machine learning module 208 includes skip connections between the outputs of convolutional layers performing downsampling and upsampling to achieve an equivalent spatial shaping layer in the network. The final convolutional layer can reduce the number of channels to three RGB channels.
[0133] During training for the reverse diffusion process, machine learning module 208 predicts noise in order to remove noise from the noisy image to achieve an initial image. Machine learning module 208 performs predictions through multiple steps, and these multiple steps may differ from those used during training for the forward diffusion process.
[0134] Structure-preserving machine learning models
[0135] Figure 9A The architecture of an example structure-preserving machine learning model according to some embodiments described herein is illustrated. In some embodiments, the structure-preserving machine learning model is a diffusion model 900. The diffusion model 900 may be... Figure 1 Media applications 103 and / or Figure 2 Part of machine learning model 208.
[0136] The diffusion model 900 is trained using training data including an initial image 902 and conditions 905. In some embodiments, the training data includes ground truth output images, such as output images that satisfy a text request and have modifications to one or more objects or regions, including the same structure and shape. For example, the initial image may include objects with a first color (e.g., a green trampoline), and the ground truth image includes objects with a second color (e.g., a purple trampoline). In some embodiments, the training data further includes pairs of ground truth images and corresponding images of randomly masked portions of the ground truth images.
[0137] Condition 905 includes a text encoder 907, a temporal encoder 909, an optional user-selected mask 911, a depth map 913, an optional retention mask 914, an optional segmentation mask 915, and a classifier-free guide 916. The text encoder 907 encodes the text request (i.e., the text condition) by converting the text into tokens used to represent the text request in a vector space (embedding space). The temporal encoder 909 uses positional encoding to encode the diffusion timestamp.
[0138] The user-selected mask 911 identifies object pixels associated with one or more objects or regions selected by the user in the initial image. During inference (i.e., during the generation of the output image), the user-selected mask 911 identifies regions in the output image to be modified. The user-selected mask 911 can identify object pixels associated with one or more selected objects.
[0139] Depth map 913 identifies the depth of one or more image pixels in the initial image. Depth map 913 is fed as input to CNN 912 to preserve the relative depth of various objects in the initial image in the output image. For example, if the selected image includes a door with a handle, depth map 913 is used to preserve the structure of the door and retain the handle in the output image. Depth map 913 is used for user requests where the user desires a photorealistic output image.
[0140] The retention mask 914 identifies pixels corresponding to the human subject in the initial image and to be retained during the generation of the output image 957. For example, the retention mask may include the human subject's hair (if the user instructs the hair to remain the same (or more generally, if no changes to the hair are specified in condition 905)), the human subject's fingers, the subject's entire body (to prevent excessive modification of the pet in the case of the subject being a pet), etc. In some embodiments where the output image modifies the human subject's clothing, the retention mask excludes pixels of the human subject's clothing and instead includes the remaining pixels associated with the human subject to prevent modifications to the human subject by the diffusion model 900. In some embodiments, multiple different generative machine learning diffusion models may be trained and used for image generation (e.g., shape-preserving models, structure-preserving models, etc.). In some embodiments, instead of using the retention mask 914, condition 905 may include an empty mask that identifies all pixels in the initial image 902 as not associated with a human.
[0141] Segmentation mask 915 identifies one or more objects or regions in the initial image 902. In some embodiments, segmentation mask 915 is used if user-selected mask 911 is not used. In some embodiments, segmentation mask 915 is used in addition to user-selected mask 911 to improve the recognition of user-selected mask 911.
[0142] In some embodiments, a classifier-free guide 916 is used to control the depth in the output image. Classifier guidance controls the categories generated by a classification model. The classifier-free guide 916 trains a diffusion model 900 with conditional dropout, i.e., conditions are removed at a certain percentage of time. In some embodiments, the removed conditions are replaced with specific input values representing the absence of conditional information. Higher conditional dropout values preserve more structure of one or more objects in the initial image than lower conditional dropout values. A drawback of higher conditional dropout values is that the increased structure may come at the cost of reduced diversity in the output image.
[0143] An initial image 902 is provided as input to the first layer of CNN 912, and conditions 905 are provided as input to each block within CNN 912. CNN 912 includes encoder blocks 917, 920, 925, and 930; an intermediate block 935; and decoder blocks 940, 945, 950, and 955 with skip connections. In some embodiments, the model is a diffusion model 900 and contains 25 blocks, of which 8 blocks are downsampled convolutional layers or upsampled convolutional layers. Although Figure 9A Four encoder blocks and four decoder blocks are shown, but in various embodiments, fewer or more encoder blocks and / or decoder blocks may be used (and the number of encoder blocks and the number of decoder blocks may be different).
[0144] The denoising process can occur in pixel space or in the latent space of the diffusion model 900. In some embodiments, during training, the machine learning module 208 performs preprocessing on the initial image 902 to transform the initial image 902 from a pixel-space image into a latent space (e.g., a vector representation of the image in a high-dimensional vector space). The machine learning module 208 performs training by transforming one or more conditions in condition 905 from the input size to a feature space vector that matches the size of the CNN 912.
[0145] Machine learning module 208 trains diffusion model 900 to receive initial image 902 and progressively adds noise to initial image 902 with each iteration of diffusion model 900 to produce a noisy image. Given a set of conditions 905 including time generated by time encoder 909, text request encoded by text encoder 907, and other task-specific conditions (e.g., user-selected mask 911, depth map 913, preservation mask 914, segmentation mask 915, and classifier-free guidance 916), the image diffusion model is trained to predict the noise to be added to the noisy image. Machine learning module 208 trains diffusion model 900 to generate multiple output images that satisfy the text request and do not include human pixels by progressively removing noise (via a denoising process). In some embodiments, denoising during training includes approximately 10,000 optimization steps to minimize the loss between the generated output image and the ground truth output image.
[0146] In some embodiments, the machine learning module 208 trains the diffusion model using three different versions of varying text requests and depth values. For example, the machine learning module 208 may run a first version of the diffusion model without text requests and depth values, a second version of the diffusion model with text requests and no depth values, and a third version of the diffusion model with text requests and depth values. Training each version of the diffusion model may include multiple iterations.
[0147] Once the diffusion model is trained, it receives a text request to generate an output image, a corresponding depth map, and a user-selected mask and / or segmentation mask, wherein the diffusion model is trained to generate output pixels unrelated to the human subject. The diffusion model performs a diffusion process on an initial image to generate a noisy image based on the initial image. In some embodiments, the diffusion model performs an inverse diffusion process, such as DDIM inversion, to generate an output image from the noisy image, wherein the output image is generated according to condition 905. The diffusion model performs inverse diffusion by predicting the noise added to the noisy image and generating an output image that satisfies the text request.
[0148] Shape Preservation Machine Learning Model
[0149] Figure 9B The architecture of an example shape-preserving machine learning model according to some embodiments described herein is illustrated. In some embodiments, the shape-preserving machine learning model is a diffusion model 958. The diffusion model 958 may be... Figure 1 Media applications 103 and / or Figure 2 Part of machine learning model 208.
[0150] The diffusion model 958 is trained using training data including an initial image 959 and conditions 960. In some embodiments, the training data includes ground truth output images, such as output images that satisfy a text request and have modifications to one or more objects or regions, including the same shape. For example, the initial image may include an object with a first texture (e.g., a realistic cat), and the ground truth includes an object with a second texture (e.g., a cartoon version of a cat). In some embodiments, the training data further includes pairs of ground truth images and corresponding images of randomly masked portions of the ground truth images.
[0151] In some embodiments, the architecture of diffusion model 958 is similar to that of a structure-preserving machine learning model, except that the shape-preserving machine learning model does not include a depth map as input. Conditions 960 include a text encoder 961, a temporal encoder 962, an optional user-selected mask 963, an optional preservation mask 964, an optional segmentation mask 965, and a classifier-free guidance 966. Because these conditions 960 are similar to reference... Figure 9A The condition described is 905, so further details will not be repeated here.
[0152] The initial image 959 is fed as input to the first layer of CNN 967, and condition 960 is fed as input to each block within CNN 967. CNN 967 includes encoder blocks 968, 969, 970, and 971; intermediate block 972; and decoder blocks 973, 974, 975, and 976 with skip connections. Because CNN 967 is similar to the reference... Figure 9A The CNN 912 is described, so further details will not be repeated here. The diffusion model 958 is trained to generate output images 977 that satisfy the rewrite prompts.
[0153] Unstructured and non-shape-preserving machine learning models
[0154] Figure 9C The architectures of example unstructured and non-shape-preserving machine learning models according to some embodiments described herein are illustrated. In some embodiments, the unstructured and non-shape-preserving machine learning model is a diffusion model 978. The diffusion model 978 may be... Figure 1 Media applications 103 and / or Figure 2 Part of machine learning model 208.
[0155] The diffusion model 978 is trained using training data including an initial image 986 and conditions 979. In some embodiments, the training data includes ground truth output images, such as output images that satisfy a text request and have modifications to one or more objects or regions that do not include identical structures or shapes. For example, the initial image may include a first object (e.g., a dog), and the ground truth image includes an object with a second object (e.g., a cat). In some embodiments, the training data further includes the initial image, and the ground truth image includes objects not present in the initial image. In some embodiments, the training data further includes pairs of ground truth images and corresponding images of randomly masked portions of the ground truth images.
[0156] In some embodiments, the architecture of the diffusion model 978 is similar to that of a structure-preserving machine learning model, except that the non-structure and non-shape-preserving machine learning model does not include a depth map, a user-selected mask, or a segmentation mask as condition 979. Additionally, for example, in the case where the first object is replaced by a second object, the condition includes a bounding box mask 984 indicating the location where the second object will be located. Condition 979 additionally includes a text encoder 980, a temporal encoder 981, an optional preservation mask 983, and a classifier-free guide 985. Because these conditions 979 are similar to reference... Figure 9A The condition described is 905, so further details will not be repeated here.
[0157] The initial image 986 is fed as input to the first layer of CNN 987, and condition 979 is fed as input to each block within CNN 987. CNN 987 includes encoder blocks 988, 989, 990, and 991; intermediate block 992; and decoder blocks 993, 994, 995, and 996 with skip connections. Because CNN 987 is similar to the reference... Figure 9A The CNN 912 is described, so further details will not be repeated here. The diffusion model 958 is trained to generate output images 997 that satisfy the rewrite prompts.
[0158] method
[0159] Figure 10 An example method 1000 for generating an output image based on rewritten prompts is illustrated. Method 1000 can be derived from... Figure 2 The method is executed by the computing device 200. In some embodiments, the method 1000 is performed by... Figure 1 The operation is performed on user device 115 or media server 101, or partly on user device 115 and partly on media server 101.
[0160] Figure 10 Method 1000 may begin at box 1002. At box 1002, an initial image and an original prompt are received from the user. In some embodiments, only the original prompt may be received (e.g., to generate a completely new image in response to the original prompt). In some embodiments, the initial image and the original prompt may be received, for example, to generate a modified image that retains some aspects of the initial image (e.g., shape, structure, color palette, objects, etc.) in the modified image, while also generating the modified image in response to the original prompt. Box 1002 may be followed by box 1004.
[0161] At box 1004, it is determined whether permission to modify the original image has been granted. For example, a request for permission is presented to the user. If permission is not granted, box 1004 may be followed by box 1006, where method 1000 ends. If permission is granted, box 1004 may be followed by box 1008.
[0162] At box 1008, a machine learning model is selected from the set of machine learning models based on the original prompt. In some embodiments, the machine learning model is selected from a base LLM that is part of computing device 200 or an LLM that is not part of computing device 200 (such as...). Figure 1 The selection is made from the LLM 120 (exemplified in the example). In some embodiments where the machine learning model is selected from the LLM 120, the selection is further based on a rewritten prompt. For example, the LLM 120 may generate a rewritten prompt that includes commands for using the selected machine learning model.
[0163] In some embodiments, model selection may be performed by an LLM or a separate model selection module (e.g., a prompting engine, different machine learning models, a classifier, or other selection algorithms). In embodiments utilizing an LLM or other machine learning model for model selection, the rewritten prompts and information about the properties of available generative models may be provided to the LLM as input along with a command that instructs the LLM output to indicate that a specific model among the available generative models will be used for image generation based on the rewritten prompts. In some embodiments, the set of machine learning models includes structure-preserving machine learning models, shape-preserving machine learning models, and non-structure and non-shape-preserving machine learning models, such as... Figures 9A to 9C As described above.
[0164] A structure-preserving machine learning model can be selected based on rewriting prompts, including commands to modify one or more objects or regions while preserving their structure in the initial image. Providing the rewritten prompts and the initial image as input to the structure-preserving machine learning model can further include providing the rewritten prompts, the initial image, and a depth map of the initial image to the structure-preserving machine learning model.
[0165] A shape-preserving machine learning model can be selected based on rewrite prompts that include commands to modify one or more objects or regions while preserving their shapes in the initial image.
[0166] The non-structure and non-shape-preserving machine learning model can be selected based on a rewritten prompt, including a command to replace one or more objects or regions in the initial image with one or more new objects or regions. In some embodiments, method 1000 further includes: generating a minimum bounding box around one or more selected objects in the initial image; generating a bounding box mask based on the minimum bounding box in response to selecting the non-structure and non-shape-preserving machine learning model; and providing the bounding box mask, along with the rewritten prompt and the initial image, as input to the non-structure and non-shape-preserving machine learning model. The non-structure and non-shape-preserving machine learning model can be selected based on a rewritten prompt, including a command to generate additional objects to be added to the initial image. Box 1008 may be followed by box 1010.
[0167] At box 1010, the original prompt and initial image are provided as input to the LLM (e.g., Figure 1 LLM 120 or as Figure 2 (This is part of an LLM or another text generation model, such as the prompt engine 206). In some embodiments, the LLM also receives user input that identifies one or more objects or regions in the initial image. Box 1010 may be followed by box 1012.
[0168] At box 1012, a rewritten prompt is received from the LLM based on the original prompt and the initial image. In some embodiments, the rewritten prompt is also based on the identification of one or more objects or regions in the initial image to be modified. In various embodiments, prompt rewriting by the LLM (or other text generation model) ensures that the input to the generative model is carefully crafted so that the model output specifically includes images that meet criteria (e.g., higher or lower levels of realism, artistic effects specified in the prompt), ensuring that the output image complies with applicable regulations and is safe to view, etc. In some embodiments, the rewritten prompt may include a user-provided initial image or a representation of that initial image (e.g., an embedding of the initial image). Box 1012 may be followed by box 1014.
[0169] At box 1014, the rewritten prompt and the initial image (or an embedding representing the initial image) are provided as input to the selected machine learning model. Box 1014 may be followed by box 1016.
[0170] At box 1016, the machine learning model outputs an image that satisfies the rewrite prompt.
[0171] In some embodiments, the method 1000 further includes: generating a user interface including an initial image and options for applying a preset to modify the initial image; and, in response to receiving a selection of the preset, outputting an output image by a machine learning model that satisfies a command associated with the preset. In some embodiments, the preset includes at least one option selected from the group consisting of: removing a fence from the initial image, erasing an object from the initial image, adding a new object to the initial image, changing the material or color of an object in the initial image, enhancing the initial image, replacing the background of the initial image, changing a subject in the initial image (e.g., changing the subject's expression, changing the subject's features, changing the subject's clothing, etc.), and combinations thereof. In some embodiments, the method 1000 further includes: providing an option to regenerate the output image, receiving a subsequent prompt from the user, and generating a subsequent output image based on the subsequent prompt.
[0172] In various embodiments, the original prompts from the user and / or the rewritten prompts from the LLM can be processed by one or more filters to ensure that the generated output image conforms to applicable rules and standards. For example, the filters may detect text requests that prevent certain modifications to the image (e.g., adding objects of prohibited categories, changing objects in the image that meet certain criteria, etc.). In response to such detection, the user is provided with guidance on the types of text requests that are not permitted. Additionally, the user may be provided with guidance on structured text requests for specifying their requirements for the output image.
[0173] In addition to the description above, users may be provided with controls that allow them to choose whether and when a system, program, or feature described herein can enable the collection of user information (e.g., information about the user's social networks, social behaviors or activities, occupation, user preferences, or the user's current location) and whether to send content or communications to the user from a server. Furthermore, certain data may be processed in one or more ways before being stored or used to remove personally identifiable information. For example, a user's identity may be processed to make it impossible to determine the user's personally identifiable information, or the user's geographic location may be generalized (e.g., to a city, zip code, or state level) if location information is available, making it impossible to determine the user's specific location. Therefore, users can control what information is collected from them, how that information is used, and what information is provided to them.
[0174] Some parts of the detailed description above are presented in terms of algorithms and symbolic representations of operations on data bits within computer memory. These algorithmic descriptions and representations are used by those of ordinary skill in the field of data processing to most effectively communicate the essence of their work to others. An algorithm here generally refers to a self-consistent sequence of steps that leads to a desired result. A step refers to a step that requires physical manipulation of physical quantities. Although not essential, these quantities are usually in the form of electrical or magnetic data that can be stored, transmitted, combined, compared, and otherwise manipulated. It has been found that, primarily for reasons of general use, these data may sometimes be appropriately referred to as bits, values, elements, symbols, characters, items, numbers, etc.
[0175] However, it should be remembered that all these and similar terms will be associated with appropriate physical quantities and are merely convenient labels applied to those quantities. Unless otherwise specifically indicated, as will be apparent from the following discussion, it should be understood that throughout the description, the use of terms including “processing” or “calculating” or “operating” or “determining” or “displaying” refers to the actions and processes of a computer system or similar electronic computing device that manipulates data represented as physical (electronic) quantities in the registers and memory of the computer system and transforms that data into other data similarly represented as physical quantities in the computer system's memory or registers or other such information storage, transmission, or display devices.
[0176] Embodiments of this specification may also relate to a processor for performing one or more steps of the methods described above. This processor may be a dedicated processor that is selectively activated or reconfigured by a computer program stored in a computer. Such a computer program may be stored in a non-transitory computer-readable storage medium, including but not limited to: any type of disk (including optical disks), ROM, CD-ROM, magnetic disk, RAM, EPROM, EEPROM, magnetic cards or optical cards, flash memory (including USB flash drives with non-volatile memory), or any type of medium suitable for storing electronic instructions, each coupled to a computer system bus.
[0177] This specification may take the form of some entirely hardware embodiments, some entirely software embodiments, or some embodiments that include both hardware and software elements. In some embodiments, this specification is implemented in software, including but not limited to firmware, resident software, microcode, etc.
[0178] Furthermore, this description may take the form of a computer program product accessible from a computer-usable or computer-readable medium that provides program code for use by or in conjunction with a computer or any instruction execution system. For the purposes of this description, a computer-usable or computer-readable medium may be any device capable of containing, storing, communicating, propagating, or transmitting a program for use by or in conjunction with an instruction execution system, device, or apparatus.
[0179] A data processing system suitable for storing or executing program code will include at least one processor directly or indirectly coupled to memory elements via a system bus. Memory elements may include local memory, mass storage, and cache memory used during the actual execution of the program code, the cache memory providing temporary storage for at least some of the program code to reduce the number of times code must be retrieved from mass storage during execution.
Claims
1. A computer-implemented method, comprising: Receive an initial image and a prompt from the user, wherein the prompt includes a request to modify the initial image; Select a machine learning model from the set of machine learning models based on the original prompts; The original prompt and the initial image are provided as input to a large language model LLM. Based on the original prompt and the initial image, a rewritten prompt is received from the LLM; The rewritten prompt and the initial image are provided as input to the selected machine learning model; as well as The selected machine learning model generates an output image that satisfies the rewrite prompts.
2. The method of claim 1, further comprising receiving user input identifying one or more objects or regions in the initial image, wherein the rewrite prompt is further based on the identification of the one or more objects or regions in the initial image to be modified.
3. The method according to claim 2, wherein, The set of machine learning models includes structure-preserving machine learning models, shape-preserving machine learning models, and non-structure and non-shape-preserving machine learning models.
4. The method according to claim 3, wherein, Selecting the machine learning model includes: based on the rewrite prompt, including commands to modify the one or more objects or regions while preserving the structure of the initial image, selecting the structure to retain the machine learning model.
5. The method according to claim 4, wherein, Providing the rewritten prompt and the initial image as input to the selected machine learning model further includes providing the rewritten prompt, the initial image, and a depth map of the initial image to the structure-preserving machine learning model.
6. The method according to claim 3, wherein, Selecting the machine learning model includes: based on the rewrite prompt, including a command to modify the shape of the one or more objects or regions in the initial image while retaining the shape of the machine learning model.
7. The method according to claim 3, wherein, Selecting the machine learning model includes: based on the rewrite prompt including a command to replace one or more objects or regions in the initial image with one or more new objects or regions, selecting the non-structure and non-shape preserving machine learning model.
8. The method of claim 7, further comprising: Generate a minimum bounding box around one or more selected objects in the initial image; In response to selecting the non-structure and non-shape-preserving machine learning model, a bounding box mask is generated based on the minimum bounding box; as well as The bounding box mask, along with the rewritten cue and the initial image, are provided as input to the unstructured and non-shape-preserving machine learning model.
9. The method according to claim 3, wherein, Selecting the machine learning model includes: based on the rewrite prompt, including a command to generate additional objects to be added to the initial image, selecting the non-structure and non-shape-preserving machine learning model.
10. The method of claim 1, further comprising: Generate a user interface, the user interface including the initial image and options for applying presets to modify the initial image; as well as In response to receiving a selection of the preset, the machine learning model outputs an output image that satisfies the command associated with the preset.
11. The method according to claim 10, wherein, The preset includes at least one option selected from the group consisting of: removing a fence from the initial image, erasing objects from the initial image, adding new objects to the initial image, changing the material or color of objects in the initial image, enhancing the initial image, replacing the background of the initial image, changing the subject in the initial image (e.g., changing the subject's expression, changing the subject's features, changing the subject's clothing, etc.), and combinations thereof.
12. A non-transitory computer-readable medium storing instructions that, when executed by one or more computers, cause the one or more computers to perform an operation or control operation, the operation comprising: Receive an initial image and a prompt from the user, wherein the prompt includes a request to modify the initial image; Select a machine learning model from the set of machine learning models based on the original prompts; The original prompt and the initial image are provided as input to a large language model LLM. Based on the original prompt and the initial image, a rewritten prompt is received from the LLM; The rewritten prompt and the initial image are provided as input to the selected machine learning model; as well as The selected machine learning model generates an output image that satisfies the rewrite prompts.
13. The non-transitory computer-readable medium according to claim 12, wherein, The operation further includes receiving user input that identifies one or more objects or regions in the initial image, wherein the rewrite prompt is further based on the identification of the one or more objects or regions in the initial image to be modified.
14. The non-transitory computer-readable medium according to claim 13, wherein, The set of machine learning models includes structure-preserving machine learning models, shape-preserving machine learning models, and non-structure and non-shape-preserving machine learning models.
15. The non-transitory computer-readable medium according to claim 12, wherein, The operation further includes: Provides an option to regenerate the output image; Receive subsequent prompts from the user; and The subsequent output image is generated based on the aforementioned prompts.
16. A system comprising: One or more processors; as well as One or more computer-readable media coupled to the one or more processors, the one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the one or more processors to perform operations or control operations, the operations including: Receive an initial image and a prompt from the user, wherein the prompt includes a request to modify the initial image; Select a machine learning model from the set of machine learning models based on the original prompts; The original prompt and the initial image are provided as input to a large language model LLM. Based on the original prompt and the initial image, a rewritten prompt is received from the LLM; The rewritten prompt and the initial image are provided as input to the selected machine learning model; and The selected machine learning model generates an output image that satisfies the rewrite prompts.
17. The system according to claim 16, wherein, The operation further includes receiving user input that identifies one or more objects or regions in the initial image, wherein the rewrite prompt is further based on the identification of the one or more objects or regions in the initial image to be modified.
18. The system according to claim 17, wherein, The set of machine learning models includes structure-preserving machine learning models, shape-preserving machine learning models, and non-structure and non-shape-preserving machine learning models.
19. The system according to claim 18, wherein, Selecting the machine learning model includes: based on the rewrite prompt, including commands to modify the one or more objects or regions while preserving the structure of the initial image, selecting the structure to retain the machine learning model.
20. The system according to claim 18, wherein, Selecting the machine learning model includes: based on the rewrite prompt, including a command to modify the shape of the one or more objects or regions in the initial image while retaining the shape of the machine learning model.