Image editing method, device, electronic device and storage medium

Through the multi-round dialogue image editing method, combined with historical dialogue information and large language models, the problems of high user threshold and complex operation in the existing technology are solved, and efficient image editing is achieved.

CN117765117BActive Publication Date: 2025-08-08BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202410024309.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-01-05
Publication Date
2025-08-08
Estimated Expiration
2044-01-05

AI Technical Summary

Technical Problem

The existing image editing technology has a high user threshold, complex operation process, difficult to meet user needs, and low image editing efficiency.

Method used

Through multi-round dialogue image editing methods, users' editing instructions are accurately understood in combination with historical dialogue information, and target images are generated using large language models and literary image diffusion models.

Benefits of technology

It significantly reduces the complexity of user operations and improves image editing efficiency and user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117765117B_ABST
    Figure CN117765117B_ABST
Patent Text Reader

Abstract

The present disclosure provides an image editing method, apparatus, electronic device, and storage medium, relating to the fields of artificial intelligence technology, particularly natural language processing, computer vision, deep learning, and other technical fields. The image editing method includes: obtaining an editing instruction input by a user in a current round of conversation and historical conversation information from previous rounds of conversation, the historical conversation information comprising historical conversation text and at least one historical image; determining a source image to be edited from the at least one historical image based on the editing instruction and the historical conversation information; and editing the source image based on the editing instruction to generate a target image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as natural language processing, computer vision, and deep learning, and specifically to an image editing method and device, an electronic device, a computer-readable storage medium, and a computer program product. Background Art

[0002] Artificial Intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). This discipline encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0003] Large language models (LLMs) are deep learning models trained using large amounts of text data. They can generate natural language text or understand its meaning. Large language models can handle a variety of natural language tasks, such as conversation, text classification, and text generation, and are an important path to artificial intelligence. Some large language models also have multimodal data processing capabilities, such as the ability to process text, images, and video.

[0004] The approaches described in this section are not necessarily approaches that have been previously conceived or employed. Unless otherwise indicated, it should not be assumed that any approach described in this section is prior art simply by virtue of its inclusion in this section. Similarly, unless otherwise indicated, the issues raised in this section should not be considered as having been recognized in any prior art. Summary of the Invention

[0005] The present disclosure provides an image editing method and apparatus, an electronic device, a computer-readable storage medium, and a computer program product.

[0006] According to one aspect of the present disclosure, there is provided an image editing method, comprising: obtaining an editing instruction input by a user in a current round of dialogue and historical dialogue information in a historical round of dialogue, wherein the historical dialogue information includes a historical dialogue text and at least one historical image; based on the editing instruction and the historical dialogue information, determining a source image to be edited from the at least one historical image; and editing the source image based on the editing instruction to generate a target image.

[0007] According to one aspect of the present disclosure, an image editing device is provided, comprising: an acquisition module configured to acquire editing instructions input by a user in a current round of dialogue and historical dialogue information in historical rounds of dialogue, wherein the historical dialogue information includes historical dialogue text and at least one historical image; a determination module configured to determine a source image to be edited from the at least one historical image based on the editing instructions and the historical dialogue information; and an editing module configured to edit the source image based on the editing instructions to generate a target image.

[0008] According to one aspect of the present disclosure, an electronic device is provided, comprising: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the above method.

[0009] According to one aspect of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to cause a computer to execute the above method.

[0010] According to one aspect of the present disclosure, a computer program product is provided, comprising computer program instructions, which implement the above method when executed by a processor.

[0011] According to one or more embodiments of the present disclosure, multi-round conversational image editing can be implemented, which significantly reduces the user's operation complexity and improves image editing efficiency and user experience.

[0012] It should be understood that the contents described in this section are not intended to identify the key or important features of the embodiments of the present disclosure, nor are they intended to limit the scope of the present disclosure. Other features of the present disclosure will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0013] The accompanying drawings illustrate exemplary embodiments and constitute a part of the specification. Together with the description of the specification, they serve to explain exemplary implementation of the embodiments. The illustrated embodiments are for illustrative purposes only and do not limit the scope of the claims. Throughout the drawings, the same reference numerals designate similar, but not necessarily identical, elements.

[0014] Figure 1 A schematic diagram illustrating an exemplary system in which the various methods described herein may be implemented according to an embodiment of the present disclosure;

[0015] Figure 2 A flowchart of an image editing method according to an embodiment of the present disclosure is shown;

[0016] Figure 3 Schematic diagram showing the t-th iteration of the Vincent graph diffusion model according to an embodiment of the present disclosure;

[0017] Figure 4 A schematic diagram illustrating an image editing process according to an embodiment of the present disclosure is shown;

[0018] Figure 5 A schematic diagram showing an example of image editing according to an embodiment of the present disclosure;

[0019] Figure 6 A schematic diagram showing a multi-round conversational image editing effect according to an embodiment of the present disclosure is shown;

[0020] Figure 7 shows a structural block diagram of an image editing device according to an embodiment of the present disclosure; and

[0021] Figure 8 A structural block diagram of an exemplary electronic device that can be used to implement the embodiments of the present disclosure is shown. DETAILED DESCRIPTION

[0022] The following description of exemplary embodiments of the present disclosure is made in conjunction with the accompanying drawings, including various details of the embodiments of the present disclosure to facilitate understanding, which should be considered as merely exemplary. Therefore, it should be appreciated by those skilled in the art that various changes and modifications may be made to the embodiments described herein without departing from the scope of the present disclosure. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0023] In this disclosure, unless otherwise specified, the use of terms such as "first" and "second" to describe various elements is not intended to limit the positional relationship, temporal relationship, or importance relationship of these elements. Such terms are only used to distinguish one element from another. In some examples, the first element and the second element may refer to the same instance of the element, while in some cases, based on the context of the description, they may also refer to different instances.

[0024] The terms used in the descriptions of the various examples in this disclosure are for the purpose of describing specific examples only and are not intended to be limiting. Unless the context clearly indicates otherwise, if the number of elements is not specifically limited, the element can be one or more. In addition, the term "and / or" used in this disclosure covers any one and all possible combinations of the listed items. "Multiple" refers to two or more.

[0025] In the technical solutions disclosed herein, the acquisition, storage, and application of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0026] Image editing refers to making specified changes to an existing image, such as adjusting brightness and contrast, adding, modifying, or deleting elements in the image to obtain a new image.

[0027] In the related art, traditional image editing tools, such as Photoshop and CorelDRAW, are often used to edit images. These image editing tools have high barriers to entry, require specialized user training, and have complex and tedious operation processes. This results in low image editing efficiency and high costs, making it difficult to meet user needs.

[0028] With the development of artificial intelligence technology, generative image editing technologies, such as image completion (inpainting) models and image extension (outpainting) models, have shown great potential in image editing tasks. Although generative image editing technologies have effectively improved image editing efficiency compared to traditional image editing tools, they still require users to perform tedious and specialized operation steps to achieve the desired editing effect. For example, for generative image completion models, users first need to specify the image area to be edited by smearing, and then enter a piece of empirical, carefully constructed text (prompt). Different editing purposes (such as adding, modifying, deleting elements, etc.) use different methods. Users need to understand the relevant algorithm principles, parameter settings, and other professional knowledge to achieve the desired editing effect. Therefore, generative image editing technology is still difficult for users to use, the image editing efficiency is low, and it is difficult to meet user needs.

[0029] As can be seen from the above, the image editing solutions in the related art are not universal, have high user barriers, complex and cumbersome operation processes, low image editing efficiency, and are difficult to meet user needs.

[0030] To address the above issues, the disclosed embodiments provide a multi-round conversational image editing method. By combining historical conversation information, the editing object (i.e., the source image) targeted by the user's current editing instruction is accurately understood. Based on the user's editing instruction, the source image is then edited to generate the target image. The disclosed embodiments understand and meet the user's image editing needs through a unified and natural multi-round conversational approach, significantly reducing the user's operational complexity and improving image editing efficiency and user experience.

[0031] The embodiments of the present disclosure will be described in detail below with reference to the accompanying drawings.

[0032] Figure 1 FIG2 is a schematic diagram of an exemplary system 100 in which the various methods and apparatuses described herein may be implemented according to an embodiment of the present disclosure. Figure 1, the system 100 includes one or more client devices 101, 102, 103, 104, 105, and 106, a server 120, and one or more communication networks 110 coupling the one or more client devices to the server 120. The client devices 101, 102, 103, 104, 105, and 106 can be configured to execute one or more applications.

[0033] In an embodiment of the present disclosure, the client devices 101 , 102 , 103 , 104 , 105 , and 106 and the server 120 may run one or more services or software applications that enable execution of the image editing method or the image editing method.

[0034] In some embodiments, server 120 may also provide other services or software applications, which may include non-virtualized environments and virtualized environments. In some embodiments, these services may be provided as web-based services or cloud services, such as provided to users of client devices 101, 102, 103, 104, 105, and / or 106 under a software as a service (SaaS) model.

[0035] exist Figure 1 In the configuration shown, the server 120 may include one or more components that implement the functions performed by the server 120. These components may include software components, hardware components, or a combination thereof that can be executed by one or more processors. Users operating client devices 101, 102, 103, 104, 105, and / or 106 may, in turn, utilize one or more client applications to interact with the server 120 to utilize the services provided by these components. It should be understood that a variety of different system configurations are possible, which may differ from the system 100. Therefore, Figure 1 is one example of a system for implementing the various methods described herein and is not intended to be limiting.

[0036] Client devices 101, 102, 103, 104, 105, and / or 106 may provide an interface that enables a user of the client device to interact with the client device. The client device may also output information to the user via the interface. Figure 1 Only six client devices are depicted, but one skilled in the art will appreciate that the present disclosure can support any number of client devices.

[0037] Client devices 101, 102, 103, 104, 105, and / or 106 may include various types of computer devices, such as portable handheld devices, general-purpose computers (such as personal computers and laptops), workstation computers, wearable devices, smart screen devices, self-service terminal devices, service robots, in-vehicle devices, gaming systems, thin clients, various messaging devices, sensors or other sensing devices, and the like. These computer devices may run various types and versions of software applications and operating systems, such as Microsoft Windows, Apple iOS, UNIX-like operating systems, Linux, or Linux-like operating systems; or various mobile operating systems, such as Microsoft Windows Mobile OS, iOS, Windows Phone, and Android. Portable handheld devices may include cellular phones, smartphones, tablet computers, personal digital assistants (PDAs), and the like. Wearable devices may include head-mounted displays (such as smart glasses) and other devices. Gaming systems may include various handheld gaming devices and internet-enabled gaming devices. Client devices are capable of executing a variety of different applications, such as various internet-related applications, communication applications (such as email applications), and short message service (SMS) applications, and may utilize various communication protocols.

[0038] The network 110 may be any type of network known to those skilled in the art that can support data communications using any of a variety of available protocols, including but not limited to TCP / IP, SNA, IPX, etc. By way of example only, the one or more networks 110 may be a local area network (LAN), an Ethernet-based network, a token ring, a wide area network (WAN), the Internet, a virtual network, a virtual private network (VPN), an intranet, an extranet, a blockchain network, a public switched telephone network (PSTN), an infrared network, a wireless network (e.g., Bluetooth, Wi-Fi), and / or any combination of these and / or other networks.

[0039] Server 120 may include one or more general-purpose computers, specialized server computers (e.g., PC (personal computer) servers, UNIX servers, mid-range servers), blade servers, mainframe computers, server clusters, or any other suitable arrangement and / or combination. Server 120 may include one or more virtual machines running virtual operating systems, or other computing architectures involving virtualization (e.g., one or more flexible pools of logical storage devices that may be virtualized to maintain a server's virtual storage device). In various embodiments, server 120 may run one or more services or software applications that provide the functionality described below.

[0040] The computing units in the server 120 may run one or more operating systems including any of the operating systems described above as well as any commercially available server operating systems. The server 120 may also run any of a variety of additional server applications and / or middle-tier applications, including HTTP servers, FTP servers, CGI servers, JAVA servers, database servers, and the like.

[0041] In some implementations, server 120 may include one or more applications to analyze and consolidate data feeds and / or event updates received from users of client devices 101, 102, 103, 104, 105, and / or 106. Server 120 may also include one or more applications to display the data feeds and / or real-time events via one or more display devices of client devices 101, 102, 103, 104, 105, and / or 106.

[0042] In some embodiments, server 120 may be a distributed system server or a server integrated with blockchain. Server 120 may also be a cloud server, or an intelligent cloud computing server or intelligent cloud host equipped with artificial intelligence technology. A cloud server is a host product within the cloud computing service system that addresses the management difficulties and poor scalability of traditional physical hosts and virtual private servers (VPS) services.

[0043] The system 100 may also include one or more databases 130. In some embodiments, these databases may be used to store data and other information. For example, one or more of the databases 130 may be used to store information such as audio files and video files. The databases 130 may reside in a variety of locations. For example, the database used by the server 120 may be local to the server 120, or may be remote from the server 120 and communicate with the server 120 via a network-based or dedicated connection. The databases 130 may be of different types. In some embodiments, the databases used by the server 120 may be, for example, relational databases. One or more of these databases may store, update, and retrieve data to and from the databases in response to commands.

[0044] In some embodiments, one or more of the databases 130 may also be used by applications to store application data. The databases used by the applications may be different types of databases, such as a key-value store, an object store, or a conventional store backed by a file system.

[0045] Figure 1 The system 100 may be configured and operated in various ways to enable application of the various methods and apparatuses described in accordance with the present disclosure.

[0046] According to some embodiments, the client devices 101-106 can execute the image editing method of the embodiment of the present disclosure to provide users with immersive image editing services. Specifically, the user can input the image processing instructions (i.e., query) for each round of dialogue by operating the client devices 101-106 (for example, operating input devices such as a mouse, keyboard, and touch screen) to express their own image creation or image editing needs. In the case where the user expresses an image creation need, the client devices 101-106 generate a new target image for the user. In the case where the user expresses an image editing need, the client devices 101-106 determine the source image to be edited from the generated historical images or the historical images uploaded by the user by executing the image editing method of the embodiment of the present disclosure, and generate a target image by editing the source image. The client devices 101-106 further output the generated target image to the user as the response data (i.e., response) of the current round of dialogue (for example, output through a display). According to some embodiments, the response data of the current round of dialogue may also include explanatory text of the target image. The explanatory text may be, for example, descriptive text for describing the screen content of the target image, text for describing the process and logic of the system generating the target image, etc.

[0047] According to some embodiments, the server 120 may also execute the image editing method according to the embodiments of the present disclosure. Specifically, the user can enter image processing instructions for each round of conversation by operating the client devices 101-106 (for example, operating an input device such as a mouse, keyboard, or touch screen) to express their image creation or image editing needs. The client devices 101-106 send the user's image processing instructions for each round of conversation to the server 120. When the user's image processing instructions express image editing needs, the server 120 executes the image editing method according to the embodiments of the present disclosure to determine the source image to be edited from historical images generated in the current conversation or historical images uploaded by the user, edit the source image to generate a target image, and output the target image as the response data for the current round of conversation to the client devices 101-106. The client devices 101-106 further output the response data to the user (for example, via a display). According to some embodiments, the response data for the current round of conversation may also include explanatory text for the target image. The explanatory text may include, for example, a description of the target image's content, text describing the system's process and logic for generating the target image, etc.

[0048] Figure 2 1 shows a flow chart of an image editing method 200 according to an embodiment of the present disclosure. As mentioned above, the execution subject of the method 200 may be a client device, such as Figure 1The client devices 101-106 shown in FIG; can also be servers, such as Figure 1 The server 120 shown in FIG.

[0049] like Figure 2 As shown, the method 200 includes steps S210 - S230 .

[0050] In step S210, the editing instruction input by the user in the current round of dialogue and the historical dialogue information in the previous round of dialogue are obtained. The historical dialogue information includes the historical dialogue text and at least one historical image.

[0051] In step S220 , a source image to be edited is determined from at least one historical image based on the editing instruction and the historical conversation information.

[0052] In step S230 , the source image is edited based on the editing instruction to generate a target image.

[0053] According to an embodiment of the present disclosure, a multi-round conversational image editing method is provided. By combining historical conversation information, the editing object (i.e., the source image) targeted by the user's current editing instruction is accurately understood. Based on the user's editing instruction, the source image is edited to generate the target image. The embodiments of the present disclosure understand and meet the user's image editing needs through a unified and natural multi-round conversational approach, significantly reducing the user's operational complexity and improving image editing efficiency and user experience.

[0054] The following describes each step of method 200 in detail.

[0055] In step S210, the editing instructions input by the user in the current round of dialogue and the historical dialogue information in the previous rounds of dialogue are obtained.

[0056] In the embodiments of the present disclosure, a dialogue refers to an interactive process in which a user inputs a question (query) and an AI image generation system outputs a response (response). Depending on the number of interactions between the user and the AI image generation system, the dialogue can be divided into a single-round dialogue and a multi-round dialogue. In a single-round dialogue, the user interacts with the AI image generation system only once. After the user inputs a question and gets a response output by the system, the dialogue ends. In a multi-round dialogue, the user interacts with the AI image generation multiple times. Each interaction is called a "round" in the dialogue, including the question input by the user and the response output by the system to the question.

[0057] In the embodiments of the present disclosure, the current round of dialogue may be any round of dialogue except the first round among multiple rounds of dialogue, for example, the second round of dialogue, the third round of dialogue, etc.

[0058] The historical conversation information includes historical conversation text and at least one historical image. The historical conversation text includes the user input text and system response text in the historical conversation rounds. The at least one historical image includes the user input image and system response image in the historical conversation rounds.

[0059] The editing instructions entered by the user in the current conversation express their current image editing needs. These instructions can directly or indirectly reference historical images. For example, editing instructions could include "make the first image more technological," "put a hat on the puppy in the image," "draw another similar set," and so on.

[0060] According to some embodiments, the text input by the user in the current round of dialogue may be used as the editing instruction.

[0061] According to other embodiments, speech recognition may be performed on the speech input by the user in the current round of dialogue to obtain result text, and the result text may be used as the editing instruction.

[0062] It should be noted that the user's input data in the current round of conversation does not always express image editing needs. It may also express other needs, such as image creation needs, casual chatting, and inquiries about how to use various functions of the AI image generation system. It is understood that only when the user's input data expresses an image editing need does it constitute an editing instruction. If the user's input data expresses a non-image editing need, it does not constitute an editing instruction.

[0063] According to some embodiments, the user's input data in the current round of dialogue may be used to perform intent recognition to determine whether the user currently has an image editing requirement, that is, to determine whether the current input data is an editing instruction.

[0064] Specifically, the user's input data in the current round of conversation and historical conversation information from previous rounds of conversation can be obtained. Based on the input data and historical conversation information, the user's intention in the current round of conversation is identified. If the intention is an image editing intention, the input data in the current round of conversation is used as an editing instruction.

[0065] According to the above embodiments, the user's image editing needs can be accurately identified, unnecessary image editing processing can be avoided, and thus image editing efficiency and the user's image editing experience can be improved.

[0066] According to some embodiments, a preset large language model can be used to identify user intent. For example, the user's current conversation input data and historical conversation information can be entered into a preset prompt template to obtain input information for the large language model. This input information is then fed into the large language model to obtain the intent recognition result output by the large language model.

[0067] In step S220 , a source image to be edited is determined from at least one historical image based on the editing instruction and the historical conversation information.

[0068] In a multi-turn conversation, the user's current edit command is often related to the previous conversation content, and their editing needs may gradually become clear over multiple rounds of input. Therefore, it is necessary to combine historical conversation information to accurately understand the editing object (i.e., the source image to be edited) targeted by the user's current edit command.

[0069] According to some embodiments, in step S220, a preset large language model may be used to determine the source image. Step S220 may include steps S221-S223.

[0070] In step S221, a preset prompt template is obtained. The prompt template includes a guide text and slots to be filled for guiding the language model to determine a source image to be edited from at least one historical image.

[0071] In step S222, the editing instructions and historical conversation information are filled into the slots to obtain input information.

[0072] In step S223, the input information is input into the language model to obtain the source image output by the language model. Specifically, the language model may output an ID of the source image.

[0073] According to the above embodiment, the language understanding capability of the large language model can be used to accurately understand the user's image editing needs, thereby ensuring the accuracy of image editing.

[0074] According to some embodiments, a prompt template may be, for example, "User's current editing instruction: {Edit instruction}\nHistorical conversation text: {Historical conversation text}\nHistorical image: {Historical image}\nThe image to be edited by the user is: {Source image}." In this prompt template, {Edit instruction}, {Historical conversation text}, and {Historical image} are slots for filling in editing instructions, historical conversation text, and historical images, respectively, and {Source image} is the output result of the large language model.

[0075] According to some embodiments, before executing steps S221-S223, high-quality annotated data of "editing instructions-historical dialogue information-source image" can be manually constructed, and the annotated data can be used to fine-tune a pre-trained large language model to improve the accuracy of the large language model in recognizing the source image.

[0076] According to some embodiments, a trained image-text matching model may also be used to determine the source image. The image-text matching model includes a text encoder and an image encoder, which can encode text and images into the same semantic space.

[0077] Specifically, in step S220, the text encoder in the image-text matching model can be used to encode the editing instruction to obtain a vector representation of the editing instruction. For each historical image, the image encoder is used to encode the historical image to obtain an initial vector representation of the historical image. If there is no explanatory text for the historical image, the initial vector representation is the final vector representation of the historical image. If there is an explanatory text for the historical image, the text encoder is further used to encode the explanatory text to obtain a vector representation of the explanatory text, and then the initial vector representation of the historical image is fused with the vector representation of the explanatory text (for example, summing, averaging, fusion through an attention mechanism, etc.) to obtain the final vector representation of the historical image. The similarity (for example, cosine similarity) between the vector representation of the editing instruction and the vector representation of each historical image is calculated, and the historical image with the greatest similarity is determined as the source image to be edited.

[0078] After the source image to be edited is determined in step S220, step S230 is executed. In step S230, the source image is edited based on the editing instruction to generate a target image.

[0079] According to some embodiments, step S230 may include steps S231 - S233 .

[0080] In step S231 , the source description text of the source image is obtained.

[0081] In step S232 , a target description text of the target image is determined based on the source description text and the editing instruction.

[0082] In step S233, a target image is generated based on the target description text, wherein the source description text is used to control the process of generating the target image.

[0083] According to the above embodiment, the generation process of the target image is controlled by using the source description text, so that the image editing effect desired by the user can be achieved while keeping the target image as similar as possible to the source image.

[0084] According to some embodiments, the description text of each historical image may be pre-stored. Accordingly, in step S231, the description text corresponding to the source image may be obtained from the stored description texts as the source description text.

[0085] According to some embodiments, the source image can be an image generated by the AI image generation system by invoking a text graph model during a historical round dialogue. That is, the source image output by the text graph model is obtained by inputting specified description text (prompt text) into the text graph model. In this case, the source description text is the description text input into the text graph model used to generate the source image.

[0086] According to some other embodiments, the source image may also be an image that the user actively uploaded in a historical round of dialogue. In this case, when the user uploads the image, a description text of the image, i.e., the source description text, can be generated by calling the large language model.

[0087] According to some embodiments, in step S232, a large language model can be used to generate a target description text. Specifically, based on the editing instructions, the language model can be used to rewrite the source description text to obtain the target description text. According to this embodiment, the language understanding and text generation capabilities of the large language model can be utilized to achieve intelligent rewriting, deeply understand the user's image editing needs, and thus improve the image editing effect.

[0088] According to some embodiments, a preset prompt template can be obtained for guiding a large language model to generate target description text. The prompt template includes slots to be filled. For example, the prompt template can be "Source description text: {source description text}\nEdit instruction: {edit instruction}\nTarget description text: {target description text}," "Source description text: {source description text}\nEdit instruction: {edit instruction}\nHistorical conversation text: {historical conversation text}\nHistorical image: {historical image}\nTarget description text: {target description text}," etc.

[0089] By filling the corresponding slots with the source description text and editing instructions (and sometimes also with historical conversation information depending on the needs of the prompt template), the input information for the large language model is obtained. This input information is then fed into the large language model to obtain the target description text as output.

[0090] According to some embodiments, in step S233, a preset text-to-image diffusion model may be used to generate a target image. The text-to-image diffusion model includes a text encoder and a noise generation network. The model receives text input and generates an image that meets the text conditions.

[0091] The Wensheng graph diffusion model generates images through a reverse diffusion process, which involves multiple iterations. Through these multiple iterations, the Wensheng graph diffusion model gradually denoises the initial image (which can be a pure noise image), ultimately producing a clear result image. During the reverse diffusion process, a given text is encoded into a high-dimensional text feature vector by a text encoder. This vector guides the noise generation network to generate noise that matches the text. The image that matches the text is then generated by sequentially subtracting the noise generated at each iteration from the initial image. Each iteration of the reverse diffusion process requires sampling the value of some random variable (typically random noise), resulting in a certain degree of randomness.

[0092] According to some embodiments, step S233 may include steps S2331 and S2332.

[0093] In step S2331 , based on the source description text, the first initial image is denoised by performing multiple first iterations using the Vincent graph diffusion model, and random variable values sampled in each of the multiple first iterations are recorded.

[0094] In step S2332, based on the target description text, the second initial image is denoised by performing multiple second iterations using the Vincent graph diffusion model to generate a target image, wherein each of the multiple second iterations reuses the random variable value sampled in the first iteration of the corresponding round.

[0095] According to the above embodiment, the Vincent graph diffusion model is used to simulate the denoising process (inverse diffusion process) of each iteration of the source image, and the random factors implicit in the source image are estimated. In other words, if the source image is generated by the Vincent graph diffusion model, what random variables should be sampled in each iteration. These random variables implicitly contain a large amount of information about the source image.

[0096] When generating the target image, the random variable values sampled when generating the source image are reused, allowing the target image to retain the content details of the source image. At the same time, generating the target image based on the target description text can achieve the editing effect desired by the user.

[0097] Furthermore, the Vue graph diffusion model in the above embodiment can be any existing Vue graph diffusion model. By controlling the inference process of the Vue graph diffusion model, the above embodiment can achieve the user's desired editing effect while maintaining the target image and the source image as similar as possible. Without the need to construct additional large-scale training data to train the Vue graph diffusion model, it can achieve universal image editing capabilities at a low cost and be transferable between multiple Vue graph diffusion models.

[0098] It should be noted that the terms "first" and "second" in the above embodiments are used to distinguish between the image generation process conditioned on the source description text and the image generation process conditioned on the target description text. The "first iteration" refers to one iteration of the image generation process conditioned on the source description text, and the "second iteration" refers to one iteration of the image generation process conditioned on the target description text.

[0099] According to some embodiments, a text graph diffusion model includes a text encoder and a noise generation network.

[0100] According to some embodiments, in step S2331, the first initial image may be a random noise image, which may be generated by random sampling in a preset noise distribution (eg, Gaussian distribution).

[0101] According to some embodiments, each first iteration of step S2331 may include the following steps S23311 - S23314 .

[0102] In step S23311, the source description text is input into the text encoder to generate a source text vector corresponding to the source description text.

[0103] In step S23312, a random variable value of the first iteration sampling, i.e., the initial sampling noise, is obtained by sampling in a preset noise distribution (e.g., a Gaussian distribution). The sampled random variable value can be a random variable image of the same size as the first initial image.

[0104] In step S23313, the source text vector and the sampled random variable value are input into the noise generation network to obtain the predicted noise of the current image. The predicted noise can be a noise image with the same size as the current image.

[0105] In step S23314, the predicted noise is removed from the current image to obtain the result image of the first iteration.

[0106] It can be understood that the current image in the first first iteration is the first initial image, and the current image in the second and subsequent first iterations is the result image generated in the previous first iteration.

[0107] It should be noted that step S2331 (including steps S23311-S23314) is used to simulate the denoising process of each iteration of the source image, but the image finally generated by this step (i.e., the result image generated by the last first iteration) is not necessarily the same as the source image.

[0108] According to some embodiments, in step S2332, the second initial image may be generated based on the source image, so that the second initial image can include information of the source image, thereby making the generated target image as consistent as possible with the source image.

[0109] According to some embodiments, the second initial image may be the source image itself, so that the second initial image retains all information of the source image.

[0110] According to other embodiments, the second initial image can be obtained by adding noise (e.g., random noise conforming to a Gaussian distribution) to the source image. This allows the target image to maintain similarity in overall visual effects (including color, composition, style, etc.) to the source image, while increasing the diversity and richness of the target image.

[0111] According to some embodiments, each second iteration of step S2332 may include steps S23321 - S23323 .

[0112] In step S23321, the target description text is input into the text encoder to generate a target text vector corresponding to the target description text.

[0113] In step S23322, the target text vector and the random variable value sampled in the first iteration of the corresponding round are input into the noise generation network to obtain the predicted noise of the current image.

[0114] In step S23323, the prediction noise is removed from the current image to obtain the result image of the second iteration. The current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.

[0115] According to the above embodiment, in the process of generating a target image based on the target description text, the random variable values sampled when generating the source image are reused, so that the user's desired editing effect can be achieved while keeping the target image as similar as possible to the source image.

[0116] Figure 3 FIG2 is a schematic diagram showing the t-th iteration of the text graph diffusion model 300 according to an embodiment of the present disclosure. The text graph diffusion model 300 includes a text encoder 310 and a noise generation network 320 . Figure 3 The upper part of shows the t-th iteration in the reverse diffusion process conditioned on the source description text t1, i.e., the t-th first iteration; Figure 3 The second part of FIG. 1 shows the t-th iteration in the inverse diffusion process conditioned on the target description text, that is, the t-th second iteration.

[0117] like Figure 3As shown in the upper part, in the tth iteration of the inverse diffusion process conditioned on the source description text t1, the source description text t1 is input into the text encoder 310 to obtain the source text vector e1. The sampling noise y(t) of this iteration is generated by sampling in a preset noise distribution (e.g., a Gaussian distribution). The source text vector e1 and the sampling noise y(t) are input into the noise generation network 320 to obtain the predicted noise z1(t) of this iteration. The current image c1(t) is subtracted from the predicted noise z1(t) to obtain the result image r1(t). It can be understood that the result image r1(t) of this iteration is the current image c1(t+1) of the next iteration.

[0118] like Figure 3 As shown in the lower half, in the tth iteration of the inverse diffusion process conditioned on the target description text t2, the target description text t2 is input into the text encoder 310 to obtain the target text vector e2. The sampling noise y(t) in the tth iteration of the inverse diffusion process conditioned on the source description text t1 is reused, and the target text vector e2 and the sampling noise y(t) are input into the noise generation network 320 to obtain the predicted noise z2(t) of this iteration. The current image c2(t) is subtracted from the predicted noise z2(t) to obtain the result image r2(t). It can be understood that the result image r2(t) of this iteration is the current image c2(t+1) of the next iteration. After completing all T iterations, the generated result image is the target image.

[0119] According to some embodiments, in step S230, a trained language instruction driven image editing model may also be used to generate a target image. Specifically, the editing instruction and the source image are input into the image editing model, and the target image is output by the image editing model.

[0120] It should be noted that the image editing model in the above embodiment is trained using a large amount of sample data of "editing instructions-source images-target images". Labeling the sample data and training the model require a lot of manpower and time, resulting in poor versatility and transferability.

[0121] Figure 4 A schematic diagram of an image editing process according to an embodiment of the present disclosure is shown. Figure 4 The image editing process shown is implemented by an AI image generation system. This AI image generation system includes a context-sensitive intent understanding module 410, a text difference-driven image editing module 420, and a historical context recording module 430. At the end of each conversation round, the historical context recording module 430 is used to store the relevant information of the conversation round (including the conversation text, generated images, image descriptions, etc.) as historical conversation information.

[0122] like Figure 4As shown, the context-dependent intent understanding module 410 obtains the user's current dialogue input (i.e., editing instructions), obtains historical dialogue information from the historical context recording module 430, and uses the large language model to understand the user's image editing intent. It then determines the ID of the source image to be edited and the description of the target image (i.e., the target description text). Based on the source image ID, the source image and its description (i.e., the source description text) are obtained from the historical context recording module 430.

[0123] The text-difference-driven image editing module 420 takes the source image description, the target image description, and the source image as input, generates a target image, and returns it to the user. By controlling the inference process of the text-image diffusion model, the text-difference-driven image editing module 420 achieves the user's desired editing effect while maintaining the generated target image's resemblance to the source image.

[0124] After the current round of dialogue ends, the user's current round of dialogue input, the description of the target image, and the generated target image are stored in the historical context recording module 430 for use in subsequent rounds of dialogue.

[0125] Figure 5 Schematic diagram showing an example of image editing according to an embodiment of the present disclosure. Figure 5 As shown, the user enters the editing instruction text "Replace the oranges with tennis balls" in the current conversation. Intent understanding module 510 first calls the large language model to determine image 501 to be edited (i.e., the source image) from the historical images generated in this conversation, and obtains the original text (i.e., the source description text) corresponding to image 501 to be edited, "a basket of oranges."

[0126] The intention understanding module 510 uses a large-scale language model (i.e., a large language model) to rewrite the original text "a basket of oranges" based on the current editing instruction text "replace oranges with tennis balls" to obtain the target text (i.e., target description text) "a basket of tennis balls".

[0127] The image to be edited 501 , the original text “a basket of oranges” and the target text “a basket of tennis balls” are input into a text difference driven image editing module 520 to generate a result image (ie, a target image) 502 .

[0128] The text difference driven image editing module 520 includes a text image diffusion model 521 and a diffusion process control module 522 .

[0129] The text-image diffusion model 521 includes a text encoder and a noise generation network. The model receives text input and generates an image that meets the text conditions.

[0130] Diffusion process control module 522 controls the randomness of the inverse diffusion process of the text-generated graph diffusion model 521 to maximize the similarity between the result image 502 and the image to be edited 501 while meeting the user's editing requirements. Diffusion process control module 522 may include a randomness estimation submodule and a graph-generated graph submodule.

[0131] The randomness estimation submodule is designed to reversely estimate the randomness factors implicit in the image to be edited. In other words, if the image to be edited is generated by the Wensheng graph diffusion model 521, what random variables (random noise) should be sampled at each iteration of the reverse diffusion process? These random variables contain a large amount of information about the image to be edited. The randomness estimation submodule uses the random noise image as a starting point and the original text "a basket of oranges" as a condition to simulate the denoising process of each iteration of the image to be edited 501 and record the value of the random variables sampled in each iteration.

[0132] The image generation submodule aims to generate an image similar to the image to be edited 501 while also meeting the user's editing requirements. To maintain the image to be edited, the inverse diffusion process is initialized by adding noise to the image to be edited 501, thereby preserving the overall visual quality of the image to be edited, such as its color and composition. Furthermore, the random variable values sampled by the randomness estimation submodule are reused during the inverse diffusion process of the resulting image 502, thereby preserving more content details of the image to be edited 501. To meet the user's editing needs and achieve the desired editing effect, the text condition is replaced with the target text "a basket of tennis balls."

[0133] In other examples, the editing instruction text entered by the user in the current round of dialogue may also be "change to oil painting style." Through the corresponding processing of the intention understanding module 510 and the text difference-driven image editing module 520, the image to be edited 501 can be determined and modified to the oil painting style, resulting in a result image 503.

[0134] Figure 6 A schematic diagram showing the effect of multi-round dialogue-style image editing according to an embodiment of the present disclosure is shown. Figure 6 In the figure, U and AI are the two parties in the dialogue, where U represents the user and AI represents the AI image generation system (also known as "AI painting assistant").

[0135] like Figure 6 As shown, in the first round of conversation, the user enters a natural language instruction 610, "Draw a cat in a field of flowers." Instruction 610 expresses the user's desire to create an image. In response to instruction 610, the AI image generation system generates a new image 622 for the user, along with its explanatory text, "This is a drawing generated for you. Click on the image to see a larger version." The combination of image 622 and its explanatory text serves as response 620 to user instruction 610.

[0136] In the second round of dialogue, the user inputs a natural language instruction 630 "Replace the kitten with a puppy". Instruction 630 expresses the user's image editing needs. In response to instruction 630, the AI image generation system uses the method 200 of the embodiment of the present disclosure to determine that the image to be edited is image 622, and edits image 622 to generate image 642. Furthermore, the explanatory text of image 642 can be generated by calling the large language model, "Replaced with a puppy, click on the picture to view the larger image~", and the combination of image 642 and its explanatory text is used as a response 640 to the user instruction 630. Figure 6 As shown, the newly generated image 642 is highly consistent with the original image 622 in terms of color, style, and position of elements (flowers, animals, etc.). At the same time, the kitten in the original image 622 is replaced with a puppy, meeting the user's editing needs.

[0137] In the third round of dialogue, the user inputs a natural language instruction 650 "draw a happy expression". Instruction 650 expresses the user's image editing needs. In response to instruction 650, the AI image generation system uses the method 200 of the embodiment of the present disclosure to determine that the image to be edited is image 642, and edits image 642 to generate image 662. Furthermore, the explanatory text of image 662 can be generated by calling the large language model "This is the edited painting, click on the picture to see the large picture~", and the combination of image 662 and its explanatory text is used as the response 660 to the user instruction 650. Figure 6 As shown, the newly generated image 662 is highly consistent with the original image 642 in terms of color, style, composition, etc., and the expression of the puppy in the original image 642 is changed from sad to happy, meeting the user's editing needs.

[0138] In the fourth round of dialogue, the user inputs a natural language instruction 670 "Remove the flowers". Instruction 670 expresses the user's image editing needs. In response to instruction 670, the AI image generation system uses the method 200 of the embodiment of the present disclosure to determine that the image to be edited is image 662, and edits image 662 to generate image 682. Furthermore, the explanatory text of image 682 can be generated by calling the large language model "This is a painting generated for you, click on the picture to see the large picture~", and the combination of image 682 and its explanatory text is used as the response 680 to the user instruction 670. Figure 6 As shown, the newly generated image 682 is highly consistent with the original image 662 in terms of color, style, composition, etc., while the flowers in the original image 662 are removed, meeting the user's editing needs.

[0139] According to an embodiment of the present disclosure, an image editing device is also provided. Figure 7FIG. 7 shows a structural block diagram of an image editing device 700 according to an embodiment of the present disclosure. Figure 7 As shown, the apparatus 700 includes an acquisition module 710 , a determination module 720 and an editing module 730 .

[0140] The acquisition module 710 is configured to acquire the editing instructions input by the user in the current round of dialogue and the historical dialogue information in the previous rounds of dialogue, wherein the historical dialogue information includes historical dialogue text and at least one historical image.

[0141] The determination module 720 is configured to determine a source image to be edited from the at least one historical image based on the editing instruction and the historical conversation information.

[0142] The editing module 730 is configured to edit the source image based on the editing instruction to generate a target image.

[0143] According to an embodiment of the present disclosure, a multi-round conversational image editing device is provided. By combining historical conversation information, the device accurately understands the editing object (i.e., the source image) targeted by the user's current editing instruction. Based on the user's editing instruction, the source image is edited to generate the target image. The embodiments of the present disclosure understand and meet the user's image editing needs through a unified and natural multi-round conversational approach, significantly reducing the user's operational complexity and improving image editing efficiency and user experience.

[0144] According to some embodiments, the determination module includes: a first acquisition unit, configured to acquire a preset prompt template, wherein the prompt template includes a guide text and a slot to be filled for guiding the language model to determine the source image to be edited from the at least one historical image; a filling unit, configured to fill the editing instruction and the historical conversation information into the slot to obtain input information; and an input unit, configured to input the input information into the language model to obtain the source image output by the language model.

[0145] According to some embodiments, the editing module includes: a second acquisition unit, configured to acquire the source description text of the source image; a determination unit, configured to determine the target description text of the target image based on the source description text and the editing instruction; and a generation unit, configured to generate the target image based on the target description text, wherein the source description text is used to control the process of generating the target image.

[0146] According to some embodiments, the determining unit is further configured to: rewrite the source description text using a language model based on the editing instruction to obtain the target description text.

[0147] According to some embodiments, the generation unit includes: a simulation subunit, configured to denoise a first initial image based on the source description text by performing multiple first iterations using a Vincent graph diffusion model, and record the random variable values sampled in each of the multiple first iterations; and a generation subunit, configured to denoise a second initial image based on the target description text by performing multiple second iterations using the Vincent graph diffusion model to generate the target image, wherein each of the multiple second iterations reuses the random variable values sampled in the first iteration of the corresponding round.

[0148] According to some embodiments, the text-graph diffusion model includes a text encoder and a noise generation network, and wherein each second iteration of the multiple second iterations includes: inputting the target description text into the text encoder to generate a target text vector corresponding to the target description text; inputting the target text vector and the random variable value sampled in the first iteration of the corresponding round into the noise generation network to obtain the predicted noise of the current image; and removing the predicted noise from the current image to obtain the result image of this second iteration, wherein the current image in the first second iteration is the second initial image, and the current image in the second and each subsequent second iteration is the result image generated in the previous second iteration.

[0149] According to some embodiments, the second initial image is generated based on the source image.

[0150] According to some embodiments, the second initial image is obtained by adding noise to the source image.

[0151] It should be understood that Figure 7 The various modules and units of the apparatus 700 shown in FIG. 7 can be compared with those in FIG. Figure 2 The steps in the method 200 described above correspond to each other. Therefore, the operations, features and advantages described above for the method 200 are also applicable to the apparatus 700 and the modules and units included therein. For the sake of brevity, some operations, features and advantages are not repeated here.

[0152] Although specific functionality is discussed above with reference to specific modules, it should be noted that the functionality of the various modules discussed herein may be separated into multiple modules, and / or at least some functionality of multiple modules may be combined into a single module.

[0153] It should also be understood that various techniques may be described herein in the general context of software hardware elements or program modules. Figure 7The various units described can be implemented in hardware or in hardware in combination with software and / or firmware. For example, these units can be implemented as computer program code / instructions, which are configured to be executed in one or more processors and stored in a computer-readable storage medium. Alternatively, these units can be implemented as hardware logic / circuits. For example, in some embodiments, one or more of modules 710-730 can be implemented together in a system on chip (SoC). SoC can include an integrated circuit chip (which includes a processor (e.g., a central processing unit (CPU), a microcontroller, a microprocessor, a digital signal processor (DSP), etc.), a memory, one or more communication interfaces, and / or one or more components in other circuits), and can optionally execute the received program code and / or include embedded firmware to perform functions.

[0154] According to an embodiment of the present disclosure, an electronic device is also provided, including: at least one processor; and a memory communicatively connected to the at least one processor, the memory storing instructions executable by the at least one processor, the instructions being executed by the at least one processor so that the at least one processor can execute the image editing method of the embodiment of the present disclosure.

[0155] According to an embodiment of the present disclosure, a non-transitory computer-readable storage medium storing computer instructions is further provided. The computer instructions are used to enable a computer to execute the image editing method of the embodiment of the present disclosure.

[0156] According to an embodiment of the present disclosure, a computer program product is further provided, including computer program instructions, which implement the image editing method of the embodiment of the present disclosure when executed by a processor.

[0157] refer to Figure 8 , a block diagram of an electronic device 800 that can serve as a server or client of the present disclosure will now be described, which is an example of a hardware device that can be applied to various aspects of the present disclosure. The electronic device is intended to represent various forms of digital electronic computer devices, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present disclosure described and / or required herein.

[0158] like Figure 8 As shown, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0159] Multiple components within electronic device 800 are connected to I / O interface 805, including an input unit 806, an output unit 807, a storage unit 808, and a communication unit 809. Input unit 806 can be any type of device capable of inputting information into electronic device 800. Input unit 806 can receive input numeric or character information and generate key signal input related to user settings and / or function control of the electronic device. It may include, but is not limited to, a mouse, keyboard, touch screen, trackpad, trackball, joystick, microphone, and / or remote control. Output unit 807 can be any type of device capable of presenting information, and may include, but is not limited to, a display, a speaker, a video / audio output terminal, a vibrator, and / or a printer. Storage unit 808 may include, but is not limited to, a magnetic disk or an optical disk. Communication unit 809 allows electronic device 800 to exchange information / data with other devices via computer networks such as the Internet and / or various telecommunication networks. It may include, but is not limited to, a modem, a network card, an infrared communication device, a wireless communication transceiver and / or chipset, such as a Bluetooth device, an 802.11 device, a Wi-Fi device, a WiMAX device, a cellular communication device, and / or the like.

[0160] The computing unit 801 can be a variety of general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the various methods and processes described above, such as method 200. For example, in some embodiments, method 200 can be implemented as a computer software program that is tangibly contained in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via ROM 802 and / or communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the method 200 described above can be performed. Alternatively, in other embodiments, the computing unit 801 can be configured to perform method 200 in any other appropriate manner (e.g., by means of firmware).

[0161] Various embodiments of the systems and techniques described above can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0162] The program code for implementing the method of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / operations specified in the flow chart and / or block diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0163] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in conjunction with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, an electronic, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or equipment, or any suitable combination of the foregoing. A more specific example of a machine-readable storage medium can include an electrical connection based on one or more lines, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0164] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0165] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or a web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), the Internet, and a blockchain network.

[0166] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact through a communication network. The client-server relationship arises through computer programs running on the respective computers and having a client-server relationship with each other. The server may be a cloud server, a server in a distributed system, or a server integrated with a blockchain.

[0167] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in this disclosure can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved. This is not limited herein.

[0168] Although the embodiments or examples of the present disclosure have been described with reference to the accompanying drawings, it should be understood that the above-mentioned methods, systems and devices are merely exemplary embodiments or examples, and the scope of the present disclosure is not limited by these embodiments or examples, but is only limited by the claims after authorization and their equivalents. Various elements in the embodiments or examples can be omitted or replaced by their equivalents. In addition, the steps can be performed in an order different from that described in the present disclosure. Further, the various elements in the embodiments or examples can be combined in various ways. It is important that as technology evolves, many of the elements described here can be replaced by equivalent elements that appear after the present disclosure.

Claims

1. An image editing method, comprising: Obtaining the editing instructions input by the user in the current round of dialogue and historical dialogue information in previous rounds of dialogue, wherein the historical dialogue information includes historical dialogue text and at least one historical image; determining a source image to be edited from the at least one historical image based on the editing instruction and the historical conversation information; Obtaining source description text of the source image; Determining a target description text of a target image based on the source description text and the editing instruction; Based on the source description text, denoising the first initial image by performing a plurality of first iterations using a Vincent graph diffusion model, and recording a random variable value sampled at each first iteration of the plurality of first iterations, the Vincent graph diffusion model including a text encoder and a noise generation network, each of the plurality of first iterations including: Input the source description text into the text encoder to generate a source text vector corresponding to the source description text; By sampling in a preset noise distribution, the random variable value of the first iterative sampling is obtained; Input the source text vector and the sampled random variable value into the noise generation network to obtain the predicted noise of the current image; and Removing the predicted noise from the current image to obtain a result image of this first iteration, where the current image in the first first iteration is the first initial image, which is a random noise image, and the current image in the second and subsequent first iterations is the result image generated in the previous first iteration; Based on the target description text, a second initial image is denoised by performing multiple second iterations using the Vincent graph diffusion model to generate the target image, wherein the second initial image is generated based on the source image, and each second iteration of the multiple second iterations reuses the random variable value sampled by the first iteration of the corresponding round.

2. The method according to claim 1, wherein Determining a source image to be edited from the at least one historical image based on the editing instruction and the historical conversation information includes: Obtaining a preset prompt template, wherein the prompt template includes a guide text for guiding the language model to determine a source image to be edited from the at least one historical image and a slot to be filled; Filling the editing instruction and the historical conversation information into the slot to obtain input information; and The input information is input into the language model to obtain the source image output by the language model.

3. The method according to claim 1, wherein The determining of a target description text of a target image based on the source description text and the editing instruction comprises: Based on the editing instruction, the source description text is rewritten using a language model to obtain the target description text.

4. The method according to claim 1, wherein Each second iteration of the plurality of second iterations comprises: Inputting the target description text into the text encoder to generate a target text vector corresponding to the target description text; Inputting the target text vector and the random variable value sampled in the first iteration of the corresponding round into the noise generation network to obtain the predicted noise of the current image; and removing the predicted noise from the current image to obtain a result image of the second iteration, The current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.

5. The method according to claim 1, wherein The second initial image is obtained by adding noise to the source image.

6. An image editing device comprising: an acquisition module configured to acquire the editing instructions input by the user in the current round of dialogue and historical dialogue information in previous rounds of dialogue, wherein the historical dialogue information includes historical dialogue text and at least one historical image; a determining module configured to determine a source image to be edited from the at least one historical image based on the editing instruction and the historical conversation information; and an editing module configured to: Obtaining source description text of the source image; Determining a target description text of a target image based on the source description text and the editing instruction; Based on the source description text, denoising the first initial image by performing a plurality of first iterations using a Vincent graph diffusion model, and recording a random variable value sampled at each first iteration of the plurality of first iterations, the Vincent graph diffusion model including a text encoder and a noise generation network, each of the plurality of first iterations including: Input the source description text into the text encoder to generate a source text vector corresponding to the source description text; By sampling in a preset noise distribution, the random variable value of the first iterative sampling is obtained; Input the source text vector and the sampled random variable value into the noise generation network to obtain the predicted noise of the current image; and Removing the predicted noise from the current image to obtain a result image of this first iteration, where the current image in the first first iteration is the first initial image, which is a random noise image, and the current image in the second and subsequent first iterations is the result image generated in the previous first iteration; Based on the target description text, a second initial image is denoised by performing multiple second iterations using the Vincent graph diffusion model to generate the target image, wherein the second initial image is generated based on the source image, and each second iteration of the multiple second iterations reuses the random variable value sampled by the first iteration of the corresponding round.

7. The device according to claim 6, wherein The determination module includes: a first acquiring unit configured to acquire a preset prompt template, wherein the prompt template includes a guide text for guiding the language model to determine a source image to be edited from the at least one historical image and a slot to be filled; a filling unit configured to fill the editing instruction and the historical conversation information into the slot to obtain input information; and An input unit is configured to input the input information into the language model to obtain the source image output by the language model.

8. The device according to claim 6, wherein The determining module is further configured to: Based on the editing instruction, the source description text is rewritten using a language model to obtain the target description text.

9. The device according to claim 6, wherein Each second iteration of the plurality of second iterations comprises: Inputting the target description text into the text encoder to generate a target text vector corresponding to the target description text; Inputting the target text vector and the random variable value sampled in the first iteration of the corresponding round into the noise generation network to obtain the predicted noise of the current image; and removing the predicted noise from the current image to obtain a result image of the second iteration, The current image in the first second iteration is the second initial image, and the current image in the second and subsequent second iterations is the result image generated in the previous second iteration.

10. The device according to claim 6, wherein The second initial image is obtained by adding noise to the source image.

11. An electronic device comprising: at least one processor; as well as a memory communicatively coupled to the at least one processor; in The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform the method according to any one of claims 1 to 5.

12. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method according to any one of claims 1 to 5.

13. A computer program product comprising computer program instructions, wherein: When the computer program instructions are executed by a processor, the method of any one of claims 1 to 5 is implemented.

Citation Information

Patent Citations

  • Image generation method and device, electronic equipment and storage medium

    CN116843795A

  • Image generation model training method and device, equipment and storage medium

    CN116958324A

  • Image generation method and device, equipment and storage medium

    CN117173284A