Image editing methods, apparatus, electronic devices, and storage media
The multi-round interactive image editing method addresses the complexity of existing image editing technologies by using historical dialogue and a large-scale language model to efficiently generate target images, improving user experience and efficiency.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2024-11-20
- Publication Date
- 2026-03-25
AI Technical Summary
Existing image editing technologies, including conventional tools and generative models, require specialized knowledge and complex operations, leading to low efficiency and difficulty in meeting user needs.
A multi-round interactive image editing method that utilizes historical dialogue information to accurately understand user editing commands, employing a large-scale language model to determine and edit source images, generating target images through a text-to-image diffusion model.
Significantly reduces user operation complexity and enhances editing efficiency and experience by naturally understanding user needs through multi-round dialogue.
Smart Images

Figure 0007835828000001 
Figure 0007835828000002 
Figure 0007835828000003
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of artificial intelligence, particularly to technical fields such as natural language processing, computer vision, deep learning, etc., and specifically relates to an image editing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.
Background Art
[0002] Artificial Intelligence (AI) is a subject that studies how to simulate some human thinking processes and intelligent behaviors (such as learning, reasoning, thinking, planning, etc.) on a computer. There are both hardware technologies and software technologies. The hardware technologies of artificial intelligence generally include technologies such as sensors, artificial intelligence dedicated chips, cloud computing, distributed storage, and big data processing. The artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.
[0003] A Large Language Model (LLM, also called a large-scale model) is a deep learning model trained using a large amount of text data, which can generate natural language text or understand the meaning of natural language text. Large language models can process multiple types of natural language tasks, such as dialogue, text classification, text generation, etc., and are one of the important paths to artificial intelligence. Some large language models further have multimodal data processing capabilities, for example, they can process multimodal data such as text, images, videos, etc.
[0004] The methods described in this section are not necessarily previously conceived or adopted. Unless otherwise specified, none of the methods described in this section should be considered prior art simply because they are included in this section. Similarly, unless otherwise specified, the problems mentioned in this section should not be considered to be recognized in the prior art. [Overview of the project]
[0005] This disclosure provides an image editing method and apparatus, electronic equipment, a computer-readable storage medium, and a computer program product.
[0006] According to one aspect of the present disclosure, an image editing method is provided, which includes obtaining an editing command entered by a user in a current round of dialogue and history dialogue information in a history round of dialogue, wherein the history dialogue information includes history dialogue text and at least one history image; determining a source image to be edited from the at least one history image based on the editing command and the history dialogue information; and editing the source image based on the editing command to generate a target image.
[0007] According to one aspect of the present disclosure, an image editing device is provided, comprising an acquisition module configured to acquire an editing command entered by a user in a current round of dialogue and history dialogue information in a history round of dialogue, wherein the history dialogue information includes history dialogue text and at least one history image; a determination module configured to determine a source image to be edited from the at least one history image based on the editing command and the history dialogue information; and an editing module configured to edit the source image to generate a target image based on the editing command.
[0008] According to one aspect of the present disclosure, an electronic device is provided which includes at least one processor and a memory communicated to the at least one processor, wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to perform the above method.
[0009] According to one aspect of the present disclosure, a non-temporary computer-readable storage medium is provided which stores computer instructions, said computer instructions are used to cause a computer to perform the above method.
[0010] According to one aspect of this disclosure, a computer program product including computer program instructions is provided, and the above method is realized when the computer program instructions are executed by a processor.
[0011] According to one or more embodiments of this disclosure, multi-round interactive image editing can be realized, significantly reducing the complexity of user operations and improving image editing efficiency and user experience.
[0012] It should be understood that the content described in this section is not intended to identify the essential or important features of the embodiments of this disclosure, nor is it intended to limit the scope of this disclosure. Other features of this disclosure will be readily apparent from the following specification. [Brief explanation of the drawing]
[0013] The drawings illustrate the embodiments and constitute part of the specification, and are used to illustrate exemplary embodiments of the embodiments in conjunction with the textual description of the specification. The embodiments shown are for illustrative purposes only and do not limit the scope of the claims. In all drawings, the same reference numerals refer to elements that are similar but not necessarily identical. [Figure 1]This is a schematic diagram of an exemplary system in which each of the methods described herein can be carried out according to the embodiments of this disclosure. [Figure 2] This is a flowchart of the image editing method according to the embodiments of this disclosure. [Figure 3] This is a schematic diagram of the t-th iteration of the text-to-image diffusion model according to an embodiment of the present disclosure. [Figure 4] This is a schematic diagram of the image editing process according to an embodiment of the present disclosure. [Figure 5] This is a schematic diagram of an example of image editing according to the embodiments of the present disclosure. [Figure 6] This is a schematic diagram of the multi-round interactive image editing effect according to the embodiments of the present disclosure. [Figure 7] This is a structural block diagram of an image editing device according to an embodiment of the present disclosure. [Figure 8] This is an exemplary structural block diagram of an electronic device that can be used to implement embodiments of the present disclosure. [Modes for carrying out the invention]
[0014] The following description illustrates exemplary embodiments of the disclosure, accompanied by drawings, and includes various details of the embodiments for the sake of ease of understanding; however, these should be considered merely illustrative. Therefore, as those skilled in the art should recognize, various changes and modifications can be made to the embodiments described herein without departing from the scope of the disclosure. Similarly, for clarity and brevity, descriptions of known functions and structures are omitted in the following description.
[0015] In this disclosure, unless otherwise specified, terms such as “first,” “second,” etc., used to describe various elements are not intended to limit the spatial, timing, or importance relationships of these elements. Such terms are used solely to distinguish one element from another. In some examples, the first and second elements may refer to the same example of that element, or, depending on the context, to different examples.
[0016] The terms used in describing the various examples in this disclosure are for illustrative purposes only and are not intended to limit them. Unless otherwise explicitly indicated in the context, such elements may be one or more, unless the number of elements is specifically limited. The terms "and / or" as used in this disclosure cover any one of the listed items and all possible combinations thereof. "Multiple" means two or more.
[0017] In the proposed technology disclosed herein, the acquisition, storage, and application of relevant user personal information all comply with the provisions of relevant laws and regulations and do not violate public order and morality.
[0018] Image editing is the process of obtaining a new image by making specified changes to an existing image, such as adjusting brightness and contrast, or adding, modifying, or deleting elements within the image.
[0019] In related technologies, images are often edited using conventional image editing tools, such as graphic software like Photoshop and CorelDRAW. These image editing tools have high usage requirements, necessitate special training for users, and have complex and cumbersome operation flows, resulting in low image editing efficiency, high costs, and difficulty in meeting user needs.
[0020] With the development of artificial intelligence technology, generative image editing technologies, such as image inpainting models and image outpainting models, etc., have shown great potential in image editing tasks. Compared with traditional image editing tools, generative image editing technologies have already effectively improved the image editing efficiency. However, in order to obtain a desirable editing effect, users still need to perform complicated and specialized operation steps. For example, for a generative image inpainting model, the user needs to specify the image area to be edited in a painting way and then input an empirically and carefully crafted text (prompt), and the usage method varies according to the editing purpose (such as adding, modifying, deleting elements, etc.). In order to obtain a desirable editing effect, the user needs to understand specialized knowledge such as relevant algorithm principles and parameter settings. Therefore, generative image editing technologies are still difficult for users to use, have low image editing efficiency, and are difficult to meet the needs of users.
[0021] As can be seen from this, the image editing solutions in related technologies lack generality, have high usage requirements for users, have a complicated and cumbersome operation flow, have low image editing efficiency, and are difficult to meet the needs of users.
[0022] In response to the above problems, the embodiments of the present disclosure provide a multi-round interactive image editing method. By associating historical dialogue information, the editing object (i.e., the source image) targeted by the user's current editing instruction is accurately understood, and further, based on the user's editing instruction, the source image is edited to generate a target image. The embodiments of the present disclosure understand and satisfy the user's image editing needs in a unified and natural multi-round dialogue manner, significantly reduce the complexity of user operations, and improve the image editing efficiency and user experience.
[0023] Hereinafter, the embodiments of the present disclosure will be described in detail with reference to the drawings.
[0024] Figure 1 shows a schematic diagram of an exemplary system 100 in which various methods and apparatus described herein can be implemented according to embodiments of the present disclosure. Referring to Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105 and 106, a server 120, and one or more communication networks 110 that connect one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105 and 106 can be configured to run one or more applications.
[0025] In embodiments of this disclosure, client devices 101, 102, 103, 104, 105 and 106 and server 120 can be operated to run an image editing method or one or more image editing service or software application.
[0026] In some embodiments, the server 120 may further provide other services or software applications, which may include non-virtual and virtual environments. In some embodiments, these services may be provided as web-based services or cloud services, for example, to users of client devices 101, 102, 103, 104, 105 and / or 106 in a Software as a Service (SaaS) model.
[0027] In the configuration shown in Figure 1, the server 120 may include one or more assemblies that implement the functions performed by the server 120. These assemblies may include software assemblies, hardware assemblies, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105 and / or 106 can utilize the services provided by these assemblies by sequentially using one or more client applications to interact with the server 120. It should be understood that various different system configurations are possible and may differ from system 100. Therefore, Figure 1 is an example of a system for implementing the various methods described herein and is not intended to limit it.
[0028] Client devices 101, 102, 103, 104, 105, and / or 106 may provide a user of the client device with an interface that allows interaction with the client device. The client device may further output information to the user through the interface. Although only six client devices are shown in Figure 1, as will be understood by those skilled in the art, this disclosure can support any number of client devices.
[0029] Client devices 101, 102, 103, 104, 105 and / or 106 may include various types of computer equipment, such as portable handheld devices, general-purpose computers (e.g., personal computers and laptop computers), workstation computers, wearable devices, smartscreen devices, self-service terminal equipment, service robots, in-vehicle equipment, game systems, thin clients, various message sending and receiving devices, sensors, or other sensing devices. These computer devices may run various types and versions of software applications and operating systems, such as MICROSOFT Windows, APPLE iOS, UNIX-like operating systems, Linux, or Linux-based operating systems, or may include various mobile operating systems, such as MICROSOFT Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablet computers, and personal digital assistants (PDAs). Wearable devices may include head-mounted displays (e.g., smart glasses) and other devices. Game systems may include various handheld game devices, internet-enabled game devices, and so on. Client devices can run various applications related to the Internet, communication applications (such as email applications), and short message service (SMS) applications, and can use various communication protocols.
[0030] Network 110 may be any type of network known to those skilled in the art, and it may use any one of several available protocols (including but not limited to TCP / IP, SNA, IPX, etc.) to support data communication. For example, one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token loop, a wide area network (WAN), the internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, Wi-Fi), and / or any combination of these and / or other networks.
[0031] Server 120 may include one or more general-purpose computers, dedicated server computers (e.g., PC (personal computer) servers, UNIX servers, midrange servers), blade servers, large computers, server clusters, or any other suitable configuration and / or combination. Server 120 may include one or more virtual machines running a virtual operating system, or other computing architectures related to virtualization (e.g., one or more flexible pools of logical memory devices that can be virtualized to maintain the server's virtual memory devices). In various embodiments, Server 120 may run one or more services or software applications that provide the functions described below.
[0032] The computing units in server 120 may run one or more operating systems, including any of the above-mentioned operating systems and any commercially available server operating systems. Server 120 may also run any one of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, Java servers, database servers, etc.
[0033] In some embodiments, the server 120 may include one or more applications for analyzing and integrating data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105 and / or 106. The server 120 may further include one or more applications for displaying data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105 and / or 106.
[0034] In some embodiments, server 120 may be a server in a distributed system or a server incorporating blockchain. Server 120 may be a cloud server or a smart cloud computing server or smart cloud host equipped with artificial intelligence technology. A cloud server is a host product in a cloud computing service system, thereby solving the shortcomings of conventional physical hosts and virtual private server (VPS) services, which are difficult to manage and have low scalability.
[0035] The system 100 may further include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information on audio files and video files. The databases 130 may be located in various locations. For example, a database used by server 120 may be located locally at server 120, or it may be located away from server 120 and communicate with server 120 via a network or a dedicated connection. The databases 130 may be of various types. In some embodiments, a database used by server 120 may be, for example, a relational database. One or more of these databases may store, update, and retrieve data from the database in response to commands.
[0036] In some embodiments, one or more of the databases 130 may be used by an application to store application data. The databases used by the application may be of different types, such as a key-value repository, an object repository, or a general-purpose repository supported by a file system.
[0037] The system 100 in Figure 1 may be configured and operated in various ways so as to allow the application of the various methods and apparatus described herein.
[0038] According to several embodiments, client devices 101-106 can perform the image editing method of the embodiments of this disclosure and provide an immersive image editing service for a user. Specifically, a user can communicate their image creation or image editing needs by operating client devices 101-106 (e.g., by operating input devices such as a mouse, keyboard, or touchscreen) to input image processing commands (i.e., queries) for each round of dialogue. If a user communicates an image creation need, client devices 101-106 generate a new target image for the user. If a user communicates an image editing need, client devices 101-106 generate a target image by performing the image editing method of the embodiments of this disclosure, determining the source image to be edited from a previously generated history image or a history image uploaded by the user, and editing the source image. Client devices 101-106 further output the generated target image to the user as response data (i.e., response) for the current round of dialogue (e.g., output via a display). According to some embodiments, the response data for the current round of dialogue may further include interpretive text for the target image, which may be, for example, descriptive text for describing the screen content of the target image, or text for describing the process and logic of the system for generating the target image.
[0039] In some embodiments, the image editing method according to the embodiments of this disclosure may be executed on the server 120. Specifically, the user can communicate their need for image creation or image editing by operating client devices 101-106 (for example, by operating input devices such as a mouse, keyboard, or touchscreen) to input image processing commands for each round of dialogue. Client devices 101-106 send the user's image processing commands for each round of dialogue to the server 120. If the user's image processing command communicates a need for image editing, the server 120 executes the image editing method according to the embodiments of this disclosure to determine the source image to be edited from a history image already generated in the current dialogue or a history image uploaded by the user, generates a target image by editing the source image, and outputs the target image to the client devices 101-106 as response data for the current round of dialogue. Client devices 101-106 further output the response data to the user (for example, via a display). According to some embodiments, the response data for the current round of dialogue may further include interpretive text for the target image, which may be, for example, descriptive text for describing the screen content of the target image, or text for describing the process and logic of the system for generating the target image.
[0040] Figure 2 shows a flowchart of the image editing method 200 according to an embodiment of the present disclosure. As described above, the entity executing the method 200 may be a client device such as the client devices 101 to 106 shown in Figure 1, or a server such as the server 120 shown in Figure 1.
[0041] As shown in Figure 2, method 200 includes steps S210 to S230.
[0042] In step S210, the editing command entered by the user in the current round of dialogue and the history dialogue information from the history round of dialogue are obtained. The history dialogue information includes the history dialogue text and at least one history image.
[0043] In step S220, based on the editing command and history dialogue information, the source image to be edited is determined from at least one history image.
[0044] In step S230, the source image is edited based on the editing command to generate the target image.
[0045] The embodiments of this disclosure provide a multi-round interactive image editing method. By linking historical dialogue information, the system accurately understands the editing object (i.e., the source image) targeted by the user's current editing command, and further edits the source image and generates a target image based on the user's editing command. The embodiments of this disclosure understand and satisfy the user's image editing needs through a unified and natural multi-round dialogue method, significantly reducing the complexity of user operations and improving image editing efficiency and user experience.
[0046] Each step of Method 200 is described in detail below.
[0047] In step S210, the editing command entered by the user in the current round's dialogue and the history dialogue information from the history round's dialogue are obtained.
[0048] In the embodiments of this disclosure, a dialogue refers to an interaction process in which a user inputs a query and an AI image generation system outputs a response. Depending on the number of interactions between the user and the AI image generation system, a dialogue can be divided into a single-round dialogue and a multi-round dialogue. In a single-round dialogue, the user interacts with the AI image generation system only once. The user inputs one question, receives a response from the system, and then the dialogue ends. In a multi-round dialogue, the user interacts with the AI image generation system multiple times. Each interaction is called a "round" of the dialogue and includes a question input by the user and a response output by the system for that question.
[0049] In the embodiments of this disclosure, the current round of dialogue may be any round other than the first round of a multi-round dialogue, for example, the second round of dialogue, the third round of dialogue, and so on.
[0050] The history dialogue information includes history dialogue text and at least one history image. The history dialogue text includes user input text and system response text in the dialogue of the history round. At least one history image includes user input image and system response image in the dialogue of the history round.
[0051] The editing commands entered by the user during the current round of dialogue are used to communicate the user's current image editing needs. Editing commands may directly or indirectly refer to past images. For example, editing commands may include "Change the first image to something more high-tech," "Put a hat on the dog in the picture," or "Redraw another set of similar images."
[0052] According to some embodiments, the text entered by the user in the current round of dialogue may be used as an edit command.
[0053] In some other embodiments, speech recognition may be performed on the speech entered by the user in the current round of dialogue to obtain the resulting text, and this resulting text may be used as an editing command.
[0054] It should be explained that user input data in the current round of dialogue does not always convey a need for image editing; it may also convey other needs, such as a need for image creation, casual conversation, or inquiries about how to use the various functions of the AI image generation system. To understand this, user input data is considered an editing command only if it conveys a need for image editing. If user input data conveys a need other than image editing, that input data is not considered an editing command.
[0055] According to some embodiments, intent identification may be performed on the input data of the user's current round of interaction to determine whether the user currently has a need for image editing, that is, to determine whether the current input data is an editing command.
[0056] Specifically, user input data from the current round of dialogue and historical dialogue information from previous rounds of dialogue may be obtained. Based on the input data and historical dialogue information, the user's intent in the current round of dialogue is identified. In response to the intent being an image editing intent, the input data from the current round of dialogue is converted into an editing command.
[0057] According to the above embodiment, it is possible to improve image editing efficiency and the user's image editing experience by accurately identifying the user's image editing needs and avoiding unnecessary image editing processes.
[0058] According to some embodiments, user intent may be identified using a pre-configured large-scale language model. For example, input data from the user's current round of dialogue and historical dialogue information may be populated into a pre-configured prompt template to obtain input information for the large-scale language model. This input information is then input into the large-scale language model to obtain intent identification results output from the large-scale language model.
[0059] In step S220, based on the editing command and history dialogue information, the source image to be edited is determined from at least one history image.
[0060] In multi-round interactions, the editing commands currently entered by the user are generally related to the content of the history dialogue, and the editing needs may gradually become clearer through the multi-round input. Therefore, it is necessary to link the history dialogue information to accurately understand the editing object (i.e., the source image to be edited) that the user's current editing command is targeting.
[0061] According to some embodiments, in step S220, the source image may be determined using a pre-configured large-scale language model. Step S220 may include steps S221 to S223.
[0062] In step S221, a pre-configured prompt template is obtained. The prompt template includes guide text and slots to be filled in to guide the language model to determine the source image to be edited from at least one history image.
[0063] In step S222, the edit command and history dialogue information are filled into the slots to obtain the input information.
[0064] In step S223, input information is fed into the language model, and the source image output from the language model is obtained. Specifically, the language model can output the identifier (ID) of the source image.
[0065] According to the above embodiment, by utilizing the language comprehension capabilities of a large-scale language model to accurately understand the user's image editing needs, the accuracy of image editing can be ensured.
[0066] According to some embodiments, the prompt template may be, for example, "User's current edit command: {edit command}\nHistory dialogue text: {history dialogue text}\nHistory image: {history image}\nImage to be edited by the user: {source image}". In the above prompt template, {edit command}, {history dialogue text}, and {history image} are slots for filling in the edit command, history dialogue text, and history image, respectively, and {source image} is the output result of a large-scale language model.
[0067] According to some embodiments, before executing steps S221 to S223, high-quality annotation data consisting of "edit command - history dialogue information - source image" can be manually created, and the accuracy of the source image recognition of the large-scale language model can be improved by using the annotation data to fine-tune the pre-trained large-scale language model.
[0068] According to some embodiments, a trained image language model may be used to determine the source image. The image language model includes a text encoder and an image encoder, which can encode text and images into the same semantic space.
[0069] Specifically, in step S220, the editing command may be encoded using a text encoder in the image language model to obtain a vector representation of the editing command. For each history image, the history image is encoded using an image encoder to obtain an initial vector representation of the history image. If no interpreted text exists for the history image, the initial vector representation becomes the final vector representation of the history image. If interpreted text exists for the history image, the interpreted text is further encoded using a text encoder to obtain a vector representation of the interpreted text, and the initial vector representation of the history image and the vector representation of the interpreted text are merged (for example, by calculating the sum, calculating the average, or merging using an attention mechanism) to obtain the final vector representation of the history image. The similarity (e.g., cosine similarity) between the vector representation of the editing command and the vector representation of each history image is calculated, and the history image with the greatest similarity is determined to be the source image to be edited.
[0070] After determining the source image to be edited in step S220, step S230 is executed. In step S230, the source image is edited based on the editing command to generate the target image.
[0071] According to some embodiments, step S230 may include steps S231 to S233.
[0072] In step S231, the source description text of the source image is obtained.
[0073] In step S232, the target description text for the target image is determined based on the source description text and editing instructions.
[0074] In step S233, a target image is generated based on the target description text. Here, the source description text is used to control the above process of generating the target image.
[0075] According to the above embodiment, by controlling the target image generation process using source description text, it is possible to achieve the image editing effect expected by the user, while maintaining the target image and source image to be as similar as possible.
[0076] According to some embodiments, the descriptive text for each history image may be stored in advance. Accordingly, in step S231, the descriptive text corresponding to the source image may be obtained from the stored descriptive text and used as the source descriptive text.
[0077] In some embodiments, the source image may be an image generated by an AI image generation system calling a text-to-image model in a history round of dialogue; that is, the source image is obtained by inputting a specified descriptive text (prompt text) into the text-to-image model and outputting the source image from the text-to-image model. In this case, the source descriptive text is the descriptive text input into the text-to-image model for generating the source image.
[0078] In some other embodiments, the source image may be an image voluntarily uploaded by the user during a history round of interaction. In this case, when the user uploads the image, a descriptive text for the image, i.e., the source description text, may be generated by calling a large-scale language model.
[0079] According to some embodiments, in step S232, a large-scale language model may be used to generate the target description text. Specifically, based on editing instructions, the language model may be used to rewrite the source description text and obtain the target description text. According to these embodiments, intelligent rewriting can be achieved by utilizing the language understanding and text generation capabilities of the large-scale language model, and the image editing effect is improved by deeply understanding the user's image editing needs.
[0080] According to some embodiments, a prompt template may be obtained to guide a large language model to generate pre-configured target description text. The prompt template includes slots to be filled in. The prompt template may be, for example, "Source description text:{source description text}\nEdit command:{edit command}\nTarget description text:{target description text}", "Source description text:{source description text}\nEdit command:{edit command}\nHistory dialogue text:{history dialogue text}\nHistory image:{history image}\nTarget description text:{target description text}", etc.
[0081] Input information for the large-scale language model is obtained by filling the corresponding slots with source description text and editing commands (and, depending on the prompt template's requirements, history dialogue information may also be filled in). This input information is then input into the large-scale language model to obtain the target description text output from the large-scale language model.
[0082] In some embodiments, in step S233, a target image may be generated using a pre-configured text-to-image diffusion model. The text-to-image diffusion model includes a text encoder and a noise generation network. The model receives a text input and generates an image that matches the text conditions.
[0083] The text-to-image diffusion model generates an image through a reverse diffusion process, which involves multiple iterations. Through these iterations, the text-to-image diffusion model gradually removes noise from the initial image (which may be a purely noised image), ultimately obtaining a sharp result image. In the reverse diffusion process, a given text is encoded into a high-dimensional text feature vector via a text encoder, which is used to guide a noise generation network to generate noise that matches the text conditions. The noise generated in each iteration is then sequentially subtracted from the initial image to obtain an image that matches the text conditions. Each iteration in the reverse diffusion process requires sampling of several random variables (typically random noise), thus each iteration possesses a certain degree of randomness.
[0084] According to some embodiments, step S233 may include steps S2331 and S2332.
[0085] In step S2331, based on the source description text, noise is removed from the first initial image by performing multiple first iterations using a text-to-image diffusion model, and the random variable values sampled in each of the multiple first iterations are recorded.
[0086] In step S2332, based on the target description text, the target image is generated by removing noise from the second initial image by performing multiple second iterations using a text-to-image diffusion model. Here, each second iteration in the multiple second iterations is multiplexed with the random variable values sampled in the first iteration of the corresponding round.
[0087] According to the above embodiment, a text-to-image diffusion model is used to simulate the nose removal process (dediffusion process) for each iteration of the source image, and the random factors hidden in the source image, i.e., what random variables are sampled in each iteration when the source image is generated by the text-to-image diffusion model, are estimated. These random variables contain a large amount of information about the source image.
[0088] In the process of generating the target image, the target image can retain details of the source image's content by multiplexing random variable values sampled when the source image was generated. Furthermore, by generating the target image based on the target description text, the editing effect expected by the user can be achieved.
[0089] The text-to-image diffusion model in the above embodiment may be any existing text-to-image diffusion model. The above embodiment can achieve the editing effect expected by the user, assuming that the target image and source image are kept as similar as possible, simply by controlling the inference process of the text-to-image diffusion model. There is no need to separately create large-scale training data to train the text-to-image diffusion model, so general-purpose image editing capabilities can be realized at low cost, and it also has transitionability between various types of text-to-image diffusion models.
[0090] It should be explained that the terms "first" and "second" in the above embodiment are used to distinguish between the image generation process conditioned on source description text and the image generation process conditioned on target description text. "First iteration" refers to one iteration in the image generation process conditioned on source description text, and "second iteration" refers to one iteration in the image generation process conditioned on target description text.
[0091] According to some embodiments, the text-to-image diffusion model includes a text encoder and a noise generation network.
[0092] According to some embodiments, in step S2331, the first initial image may be a random noise image. This image may be generated by randomly sampling in a predetermined noise distribution (e.g., a Gaussian distribution).
[0093] According to some embodiments, each first iteration in step S2331 may include the following steps S23311 to S23314.
[0094] In step S23311, the source description text is input to a text encoder to generate a source text vector corresponding to the source description text.
[0095] In step S23312, sampling is performed using a pre-set noise distribution (e.g., a Gaussian distribution) to obtain the random variable values sampled in this first iteration, i.e., the initial sampling noise. The sampled random variable values may be random variable images of the same size as the first initial image.
[0096] In step S23313, the source text vector and the sampled random variable values are input to the noise generation network to obtain predicted noise for the current image. The predicted noise may be a noise image with the same size as the current image.
[0097] In step S23314, prediction noise is removed from the current image to obtain the result image for the first iteration.
[0098] To make it clear, the current image in the first iteration is the initial image, and the current image in the second and subsequent first iterations is the resulting image generated in the previous first iteration.
[0099] It should be explained that step S2331 (including steps S23311 to S23314) is used to simulate the denoising process for each iteration of the source image, but the image finally generated in this step (i.e., the resulting image generated in the first and final iteration) is not necessarily the same as the source image.
[0100] According to some embodiments, in step S2332, the second initial image may be generated based on the source image. This allows the second initial image to include information from the source image, thereby ensuring that the generated target image matches the source image as closely as possible.
[0101] According to some embodiments, the second initial image may be the source image itself, thereby retaining all the information of the source image.
[0102] According to some other embodiments, the second initial image may be obtained by adding noise (e.g., random noise fitted to a Gaussian distribution) to the source image. This preserves the overall visual effect (including color, composition, style, etc.) of the target image and the source image, while also improving the diversity and richness of the target image.
[0103] According to some embodiments, each second iteration in step S2332 may include steps S23321 to S23323.
[0104] In step S23321, the target description text is input to a text encoder to generate a target text vector corresponding to the target description text.
[0105] In step S23322, the target text vector and the random variable values sampled in the first iteration of the round are input to the noise generation network to obtain the predicted noise for the current image.
[0106] In step S23323, prediction noise is removed from the current image to obtain the result image for the current second iteration. Here, the current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.
[0107] According to the above embodiment, in the process of generating a target image based on the target description text, the editing effect expected by the user can be achieved by multiplexing the random variable values sampled when the source image was generated, while maintaining the target image and source image to be as similar as possible.
[0108] Figure 3 shows a schematic diagram of the t-th iteration of the text-to-image diffusion model 300 according to an embodiment of the present disclosure. The text-to-image diffusion model 300 includes a text encoder 310 and a noise generation network 320. The upper half of Figure 3 shows the t-th iteration, i.e., the t-th first iteration, in the dediffusion process conditioned on source description text t1, and the lower half of Figure 3 shows the t-th iteration, i.e., the t-th second iteration, in the dediffusion process conditioned on target description text.
[0109] As shown in the upper half of Figure 3, in the tth iteration of the despreading process conditioned on the source description text t1, the source description text t1 is input to the text encoder 310 to obtain the source text vector e1. The sampling noise y(t) for the current iteration is obtained by sampling using a pre-set noise distribution (e.g., a Gaussian distribution). The source text vector e1 and the sampling noise y(t) are input to the noise generation network 320 to obtain the predicted noise z1(t) for the current iteration. The current image c1(t) and the predicted noise z1(t) are subtracted to obtain the resulting image r1(t). As can be understood, the resulting image r1(t) for the current iteration becomes the current image c1(t+1) for the next iteration.
[0110] As shown in the lower half of Figure 3, in the t-th iteration of the despreading process conditioned on the target description text t2, the target description text t2 is input to the text encoder 310 to obtain the target text vector e2. The sampling noise y(t) from the t-th iteration of the despreading process conditioned on the source description text t1 is multiplexed, and the target text vector e2 and the sampling noise y(t) are input to the noise generation network 320 to obtain the predicted noise z2(t) for the current iteration. The current image c2(t) and the predicted noise z2(t) are subtracted to obtain the resulting image r2(t). As can be understood, the resulting image r2(t) for the current iteration becomes the current image c2(t+1) for the next iteration. After all T iterations are completed, the generated resulting image is the target image.
[0111] According to some embodiments, in step S230, a target image may be generated using an image editing model driven by trained language instructions. Specifically, editing instructions and a source image are input to the image editing model to obtain a target image output from the image editing model.
[0112] It should be explained that the image editing model in the above embodiment is obtained by training it with a large amount of sample data consisting of "editing command - source image - target image". Annotating the sample data and training the model requires a large amount of manual labor and time cost, resulting in low versatility and transitionability.
[0113] Figure 4 shows a schematic diagram of an image editing process according to an embodiment of the present disclosure. The image editing process shown in Figure 4 is implemented by an AI image generation system. The AI image generation system includes a context-related intent understanding module 410, a text difference-driven image editing module 420, and a history context recording module 430. The history context recording module 430 is used to store relevant information of a round of dialogue (including dialogue text, generated images, image descriptions, etc.) as history dialogue information when each round of dialogue is completed.
[0114] As shown in Figure 4, the context-related intent understanding module 410 understands the user's image editing intent by acquiring the user's dialogue input (i.e., edit command) in the current round, acquiring historical dialogue information from the history context recording module 430, and calling a large-scale language model, thereby determining the identifier (ID) of the source image to be edited and the description (i.e., target description text) of the target image. Based on the source image identifier, the source image and the description (i.e., source description text) of the source image are acquired from the history context recording module 430.
[0115] The text difference-driven image editing module 420 takes the source image description, target image description, and source image as input, generates a target image, and returns it to the user. By controlling the inference process of the text difference-driven image editing module 420, the user can achieve the editing effect expected by the user, on the premise that the generated target image and the source image are kept as similar as possible.
[0116] After the dialogue in this round is completed, the description of the target image of the user dialogue in this round and the generated target image are stored in the history context recording module 430 and used in the dialogue of subsequent rounds.
[0117] Figure 5 shows a schematic diagram of an example of image editing according to an embodiment of the present disclosure. As shown in Figure 5, the editing command text entered by the user in the current round of dialogue is "Replace the oranges with tennis balls." The intent understanding module 510 first calls a large language model to determine the image 501 to be edited (i.e., the source image) from the history images generated in the current dialogue, and obtains the original text (i.e., the source description text) "a basket full of oranges" corresponding to the image 501 to be edited.
[0118] The intent understanding module 510 uses a large-scale language model (i.e., a large-scale language model) to rewrite the original text "a basket full of oranges" based on the current edit command text "replace the oranges with tennis balls" to obtain the target text (i.e., the target description text) "a basket full of tennis balls".
[0119] The image to be edited 501, the original text "a basket full of oranges," and the target text "a basket full of tennis balls" are input to the image editing module 520, which is driven by text differences, to generate the resulting image (i.e., the target image) 502.
[0120] The text difference-driven image editing module 520 includes a text-to-image diffusion model 521 and a diffusion process control module 522.
[0121] The text-to-image diffusion model 521 includes a text encoder and a noise generation network. The model receives text input and generates an image that meets the text conditions.
[0122] The diffusion process control module 522 controls the randomness of the dediffusion process of the text-to-image diffusion model 521, thereby maximizing the similarity between the resulting image 502 and the image 501 to be edited, while also meeting the user's editing needs. The diffusion process control module 522 may include a randomness estimation submodule and an image-to-image submodule.
[0123] The Randomness Estimation submodule aims to inversely estimate the random factors hidden in the image to be edited. In other words, if the image to be edited is generated by the text-to-image diffusion model 521, each iteration in the dediffusion process samples what random variables (random noise). These random variables contain a large amount of information hidden in the image to be edited. The Randomness Estimation submodule starts with a random noise image, uses the original text "a basket full of oranges" as a condition, simulates the denoising process for each iteration of the image to be edited 501, and records the random variable values sampled in each iteration.
[0124] The image-to-image submodule aims to generate an image that is similar to the image 501 to be edited and satisfies the user's editing needs. To preserve the image to be edited, the dediffusion process is initialized by adding noise to the image 501 to be edited, thereby preserving the overall visual effects of the image 501, such as its color and composition. At the same time, more content details of the image 501 are preserved by multiplexing the random variable values sampled in the randomness estimation submodule during the dediffusion process of the resulting image 502. To satisfy the user's editing needs and achieve the editing effect the user expects, the text condition is replaced with the target text "a basket full of tennis balls".
[0125] In some other examples, the editing command text entered by the user in the current round's dialogue may be "Change it to an oil painting style." Through the corresponding processing by the intent understanding module 510 and the image editing module 520 driven by the text difference, the image 501 to be edited is determined, it is modified to an oil painting style, and the resulting image 503 is obtained.
[0126] Figure 6 shows a schematic diagram of the multi-round interactive image editing effect according to an embodiment of the present disclosure. In Figure 6, U and AI represent the interaction, where U represents the user and AI represents the AI image generation system (which may also be called the "AI drawing assistant").
[0127] As shown in Figure 6, in the first round of dialogue, the user inputs a natural language command 610, "Please draw a cat among flowers." Command 610 conveys the user's need for image creation. In response to command 610, the AI image generation system generates a new image 622 for the user, along with interpretive text, "This is a work I generated for you. Click on the picture to see a larger version." The combination of image 622 and its interpretive text becomes the response 620 to user command 610.
[0128] In the second round of dialogue, the user inputs a natural language command 630, "Replace the cat with a dog." Command 630 conveys the user's need for image editing. In response to command 630, the AI image generation system uses the method 200 of the embodiments of this disclosure to determine that the image to be edited is image 622, edits image 622, and generates image 642. Furthermore, by calling a large language model, it may generate interpretive text for image 642, "Replaced with a dog. Click the picture to see a larger version," and the combination of image 642 and its interpretive text may be used as the response 640 to user command 630. As shown in Figure 6, the newly generated image 642 closely matches the original image 622 in terms of color, style, and the position of elements (flowers, animals, etc.), and satisfies the user's editing need by replacing the cat in the original image 622 with a dog.
[0129] In the third round of dialogue, the user inputs a natural language command 650, "Please draw a happy expression." Command 650 conveys the user's need for image editing. In response to command 650, the AI image generation system uses the method 200 of the embodiments of this disclosure to determine that the image to be edited is image 642, edits image 642, and generates image 662. Furthermore, by calling a large language model, it may generate interpretive text for image 662, "This is the edited work. Click the picture to see a larger version," and the combination of image 662 and its interpretive text may be used as the response 660 to user command 650. As shown in Figure 6, the newly generated image 662 closely matches the original image 642 in terms of color, style, composition, etc., and modifies the dog in the original image 642 from a sad expression to a happy expression, thus satisfying the user's editing need.
[0130] In the fourth round of dialogue, the user inputs a natural language command 670, "Please remove the flowers." Command 670 conveys the user's need for image editing. In response to command 670, the AI image generation system uses the method 200 of the embodiments of this disclosure to determine that the image to be edited is image 662, edits image 662, and generates image 682. Furthermore, by calling a large language model, it may generate interpretive text for image 682, "This is a work generated for you. Click the picture to see a larger version," and the combination of image 682 and its interpretive text may be used as the response 680 to user command 670. As shown in Figure 6, the newly generated image 682 closely matches the original image 662 in terms of color, style, composition, etc., and the flowers in the original image 662 have been removed, thus satisfying the user's editing need.
[0131] The embodiments of this disclosure further provide an image editing apparatus. Figure 7 shows a structural block diagram of an image editing apparatus 700 according to an embodiment of this disclosure. As shown in Figure 7, the apparatus 700 includes an acquisition module 710, a finalization module 720, and an editing module 730.
[0132] The acquisition module 710 is configured to acquire editing commands entered by the user in the current round of interaction and history interaction information in the history round of interaction, wherein the history interaction information includes history interaction text and at least one history image.
[0133] The confirmation module 720 is configured to confirm the source image to be edited from at least one history image based on the editing command and the history dialogue information.
[0134] The editing module 730 is configured to edit the source image based on the editing command and generate a target image.
[0135] The embodiments of this disclosure provide a multi-round interactive image editing device. By linking historical dialogue information, the device accurately understands the editing object (i.e., the source image) targeted by the user's current editing command, and further edits the source image and generates a target image based on the user's editing command. The embodiments of this disclosure understand and satisfy the user's image editing needs through a unified and natural multi-round dialogue method, significantly reducing the complexity of user operation and improving image editing efficiency and user experience.
[0136] According to some embodiments, the confirmation module includes a first acquisition unit configured to acquire a pre-configured prompt template, wherein the prompt template includes guide text and slots to be filled in for guiding a language model to determine a source image to be edited from the at least one history image; a filling unit configured to fill the slots with the editing command and the history dialogue information to obtain input information; and an input unit configured to input the input information to the language model to obtain the source image output from the language model.
[0137] According to some embodiments, the editing module includes a second acquisition unit configured to acquire source description text for the source image, a confirmation unit configured to confirm target description text for the target image based on the source description text and the editing instructions, and a generation unit configured to generate the target image based on the target description text, wherein the source description text is used to control the process of generating the target image.
[0138] According to some embodiments, the determination unit is further configured to rewrite the source description text using a language model based on the editing instruction to obtain the target description text.
[0139] According to some embodiments, the generation unit includes a simulation subunit configured to denoise a first initial image by performing multiple first iterations using a text-to-image diffusion model based on the source description text, and to record random variable values sampled in each first iteration of the multiple first iterations; and a generation subunit configured to generate the target image by denoising a second initial image by performing multiple second iterations using the text-to-image diffusion model based on the target description text, wherein each second iteration of the multiple second iterations multiplexes the random variable values sampled in the first iteration of the round.
[0140] According to some embodiments, the text-to-image diffusion model includes a text encoder and a noise generation network, wherein each second iteration of the plurality of second iterations includes inputting the target description text into the text encoder to generate a target text vector corresponding to the target description text, inputting the target text vector and random variable values sampled in the first iteration of the round into the noise generation network to obtain predicted noise for the current image, and removing the predicted noise from the current image to obtain the result image for the current second iteration, wherein the current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.
[0141] According to some embodiments, the second initial image is generated based on the source image.
[0142] According to some embodiments, the second initial image is obtained by adding noise to the source image.
[0143] It should be understood that each module and unit of the apparatus 700 shown in Figure 7 can correspond to each step in Method 200 described with reference to Figure 2. Therefore, the operations, features, and advantages described above for Method 200 are also applicable to the apparatus 700 and the modules and units contained therein. For the sake of brevity, some operations, features, and advantages are omitted here.
[0144] While specific functions were discussed above by referring to specific modules, it should be noted that the functions of each module discussed herein may be divided into multiple modules, and / or at least some functions of multiple modules may be combined into a single module.
[0145] Furthermore, it should be understood that this specification can describe various technologies in the general context of software, hardware elements, or program modules. Each unit described with respect to Figure 7 may be implemented in hardware, or in hardware combined with software and / or firmware. For example, these units can be implemented as computer program code / instructions configured to run on one or more processors and be stored in a computer-readable storage medium. Selectively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of modules 710-730 may be implemented together in a System on Chip (SoC). The SoC may include an integrated circuit chip (e.g., a processor (e.g., a Central Processing Unit, CPU), a microcontroller, a microprocessor, a Digital Signal Processor, DSP, etc.), memory, one or more communication interfaces, and / or one or more components in other circuits), and optionally include the execution of received program code and / or embedded firmware, thereby enabling it to perform functions.
[0146] According to embodiments of the present disclosure, an electronic device is provided which includes at least one processor and a memory that is communicated to the at least one processor, the memory storing instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to execute the image editing method of the embodiments of the present disclosure.
[0147] Embodiments of the present disclosure further provide a non-temporary computer-readable storage medium in which computer instructions are stored, which are used to cause a computer to execute an image editing method of the embodiments of the present disclosure.
[0148] According to embodiments of the present disclosure, a computer program product including computer program instructions is provided, and when the computer program instructions are executed by a processor, an image editing method of the embodiments of the present disclosure is realized.
[0149] Referring to Figure 8, a structural block diagram of an electronic device 800 that can be used as a server or client in this disclosure is described, which is an example of hardware equipment that can be applied to each aspect of this disclosure. Electronic devices represent various forms of digital electronic computer equipment, such as laptop computers, desktop computers, stages, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. Electronic devices may further represent various forms of mobile devices, such as personal digital assistants, cellular phones, smartphones, wearable devices, and other similar computing devices. The components, their connections and their functions shown herein are illustrative only and are not intended to limit the realization of the disclosure described and / or claimed herein.
[0150] As shown in Figure 8, the electronic device 800 includes a computing unit 801, which can perform various appropriate operations and processes based on computer programs stored in read-only memory (ROM) 802 or computer programs loaded from storage unit 808 into random access memory (RAM) 803. RAM 803 may further store various programs and data necessary for the operation of the electronic device 800. The computing unit 801, ROM 802, and RAM 803 are connected to each other via bus 804. An input / output (I / O) interface 805 is also connected to bus 804.
[0151] Multiple components in the electronic device 800 are connected to the I / O interface 805 and include an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. The input unit 806 may be any type of device capable of inputting information into the electronic device 800, and may receive input numeric or character information and generate key signal inputs related to user settings and / or function control of the electronic device, and may include, but is not limited to, a mouse, keyboard, touchscreen, trackboard, trackball, lever, microphone, and / or remote control. The output unit 807 may be any type of device capable of presenting information, and may include, but is not limited to, a display, speaker, video / audio output terminal, vibrator, and / or printer. The storage unit 808 may include, but is not limited to, a magnetic disk or an optical disk. The communication unit 809 enables the electronic device 800 to exchange information / data with other devices via computer networks, such as the Internet, and / or various telecommunication networks, and may include, but is not limited to, modems, network cards, infrared communication devices, wireless communication transceivers, and / or chipsets, such as Bluetooth devices, 802.11 devices, Wi-Fi devices, WiMAX devices, cellular communication devices, and / or similar devices.
[0152] The computing unit 801 may be a variety of general-purpose and / or dedicated processing assemblies having processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 801 performs each of the methods and processes described in the preceding paragraph, for example, method 200. For example, in some embodiments, method 200 may be implemented as a computer software program tangibly contained in a machine-readable medium, for example, a storage unit 808. In some embodiments, part or all of the computer program may be loaded and / or installed into the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of method 200 described in the preceding paragraph can be performed. Optionally, in other embodiments, the computing unit 801 may be configured to perform method 200 in any other suitable manner (for example, by firmware).
[0153] Various embodiments of the systems and technologies described herein may be implemented in digital electronic circuit systems, integrated circuit systems, field-programmable gate arrays (FPGAs), application-specific integrated circuits (ASICs), application-specific standard products (ASSPs), systems on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments may include: being implemented in one or more computer programs, which may be executed and / or interpreted on a programmable system including at least one programmable processor, which may be a dedicated or general-purpose programmable processor, which may receive data and instructions from a storage system, at least one input device, and at least one output device, and which may transmit data and instructions to the storage system, at least one input device, and at least one output device.
[0154] Program code for carrying out the methods of this disclosure may be written in any combination of one or more programming languages. This program code may be provided to a processor or controller of a general-purpose computer, a dedicated computer, or other programmable data processing device, so that when the program code is executed by the processor or controller, the functions / operations defined in the flowcharts and / or block diagrams are performed. The program code may be executed entirely by machine, partially by machine, partially by machine and partially by remote machine as a standalone software package, or entirely by remote machine or server.
[0155] In the context of this disclosure, a machine-readable medium may be a tangible medium that contains or stores a program for use by or in combination with an instruction execution system, device, or apparatus. A machine-readable medium may be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium may include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device, or apparatus, or any appropriate combination of the above. More specific examples of machine-readable storage media include one or more wire-based electrical connections, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any appropriate combination of the above.
[0156] To provide user interaction, a computer may implement the systems and techniques described herein, the computer having a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitoring monitor), and a keyboard and pointing device (e.g., a mouse or trackball), the user may provide input to the computer using the keyboard and pointing device. Other types of devices may further provide user interaction, for example, the feedback provided to the user may be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback), and input from the user may be received in any form (including voice input, haptic input).
[0157] The systems and technologies described herein may be implemented in computing systems including background components (e.g., data servers), computing systems including middleware components (e.g., application servers), computing systems including front-end components (e.g., user computers having a graphical user interface or web browser, through which users can interact with embodiments of the systems and technologies described herein), or in computing systems including any combination of such background components, middleware components, or front-end components. The components of the system may be interconnected by digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (LANs), wide area networks (WANs), the internet, and blockchain networks.
[0158] A computer system may include a client and a server. The client and server are generally geographically distant from each other and typically interact via a communication network. The client-server relationship is created by running computer programs on the relevant computers that have a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server incorporating blockchain technology.
[0159] It should be understood that the steps may be reordered, added, or deleted using the various forms of flows described above. For example, each step described herein may be performed in parallel, sequentially, or in a different order, as long as it achieves the desired results of the proposed technology disclosed herein, and this specification does not limit this.
[0160] While embodiments or examples of the present disclosure have been described with reference to the drawings, it should be understood that the above methods, systems, and apparatus are merely illustrative embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but is limited only by the authorized claims and their equivalents. Various elements in the embodiments or examples may be omitted or replaced by equivalent elements. Furthermore, each step may be performed in an order different from the order described in the present disclosure. In addition, various elements in the embodiments or examples may be combined in various ways. In essence, as technology advances, many of the elements described herein may be replaced by equivalent elements appearing later in the present disclosure.
Claims
1. This is an image editing method, The process involves obtaining the editing commands entered by the user in the current round of dialogue and the history dialogue information from the history round of dialogue, wherein the history dialogue information includes the history dialogue text and at least one history image. Based on the editing command and the history dialogue information, the source image to be edited is determined from the at least one history image, This includes editing the source image based on the editing command to generate a target image, Editing the source image based on the aforementioned editing command is: Obtaining the source description text of the aforementioned source image, Based on the source description text and the editing command, the target description text of the target image is determined, An image editing method comprising generating a target image based on the target description text, wherein the source description text is used to control the process of generating the target image.
2. Based on the aforementioned editing command and the history dialogue information, determining the source image to be edited from at least one history image is: Obtaining a pre-configured prompt template, wherein the prompt template includes guide text and slots to be filled in for guiding the language model to determine the source image to be edited from the at least one history image. The editing command and the history dialogue information are filled into the slot to obtain input information, The method according to claim 1, comprising inputting the input information into the language model and obtaining the source image output from the language model.
3. Based on the aforementioned source description text and editing instructions, determining the target description text for the target image is: The method according to claim 1, comprising rewriting the source description text using a language model based on the editing command to obtain the target description text.
4. Generating the target image based on the aforementioned target description text is: Based on the source description text, noise is removed from the first initial image by performing multiple first iterations using a text-to-image diffusion model, and the random variable values sampled in each of the multiple first iterations are recorded. The process includes generating the target image by removing noise from a second initial image by performing multiple second iterations using the text-to-image diffusion model based on the target description text, The method according to claim 1, wherein each second iteration of the multiple second iterations multiplexes the random variable values sampled in the first iteration of the round.
5. The text-to-image diffusion model includes a text encoder and a noise generation network, wherein each second iteration of the plurality of second iterations is The target description text is input to the text encoder to generate a target text vector corresponding to the target description text, The target text vector and the random variable values sampled in the first iteration of the round are input to the noise generation network to obtain the predicted noise of the current image. This includes removing the prediction noise from the current image to obtain the result image of the second iteration, The method according to claim 4, wherein the current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.
6. The method according to claim 4, wherein the second initial image is generated based on the source image.
7. The method according to claim 6, wherein the second initial image is obtained by adding noise to the source image.
8. It is an image editing device, An acquisition module configured to acquire edit commands entered by a user in the current round of dialogue and history dialogue information in the history round of dialogue, wherein the history dialogue information includes history dialogue text and at least one history image, A confirmation module configured to determine the source image to be edited from at least one history image based on the editing command and the history dialogue information, Includes an editing module configured to edit the source image based on the editing command to generate a target image, The aforementioned editing module is A second acquisition unit configured to acquire the source description text of the aforementioned source image, A confirmation unit configured to confirm the target description text of the target image based on the source description text and the editing command, An image editing apparatus comprising a generation unit configured to generate a target image based on the target description text, wherein the source description text is used to control the process of generating the target image.
9. The aforementioned confirmation module is A first acquisition unit configured to acquire a pre-configured prompt template, wherein the prompt template includes guide text and slots to be filled in for guiding a language model to determine a source image to be edited from at least one history image, A filling unit configured to obtain input information by filling the slot with the editing command and the history dialogue information, The apparatus according to claim 8, further comprising an input unit configured to input the input information into the language model and obtain the source image output from the language model.
10. The aforementioned confirmation unit further, The apparatus according to claim 8, configured to rewrite the source description text using a language model based on the editing command to obtain the target description text.
11. The aforementioned generation unit is Based on the aforementioned source description text, a simulated subunit is configured to remove noise from a first initial image by performing multiple first iterations using a text-to-image diffusion model, and to record random variable values sampled in each of the multiple first iterations. Includes a generation subunit configured to generate the target image by removing noise from a second initial image by performing multiple second iterations using the text-to-image diffusion model based on the target description text, The apparatus according to claim 8, wherein each second iteration of the multiple second iterations multiplexes the random variable values sampled in the first iteration of the round.
12. The text-to-image diffusion model includes a text encoder and a noise generation network, wherein each second iteration of the plurality of second iterations is The target description text is input to the text encoder to generate a target text vector corresponding to the target description text, The target text vector and the random variable values sampled in the first iteration of the round are input to the noise generation network to obtain the predicted noise of the current image. This includes removing the prediction noise from the current image to obtain the result image of the second iteration, The apparatus according to claim 11, wherein the current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.
13. The apparatus according to claim 11, wherein the second initial image is generated based on the source image.
14. The apparatus according to claim 13, wherein the second initial image is obtained by adding noise to the source image.
15. At least one processor, An electronic device including a memory that is communicated to at least one processor, Electronic device wherein the memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to cause the at least one processor to perform the method according to any one of claims 1 to 7.
16. A non-temporary computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause a computer to perform the method described in any one of claims 1 to 7.
17. A computer program product comprising computer program instructions, wherein when the computer program instructions are executed by a processor, the method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Image generation method and device, electronic equipment and storage medium
CN116843795A
Digital Media Environment for Conversational Image Editing and Enhancement
US20200066261A1