Image editing method and device, storage medium and electronic device
By combining a multimodal large model and a diffusion model, and using visual and textual lexical units as guiding conditions, the problem of uncontrollable image editing effects is solved, and image editing effects that are closer to user expectations are achieved.
Patent Information
- Application Number
- CN202411603258.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-11
- Publication Date
- 2025-11-18
- Estimated Expiration
- 2044-11-11
AI Technical Summary
In existing image editing methods, the editing effect is uncontrollable, and the effect of the edited image is inconsistent with the user's expectations.
The system generates target image lexical units using a pre-trained multimodal large model, and uses visual lexical units and text lexical units corresponding to editing instructions as guiding conditions. It then uses a pre-trained diffusion model for image editing, adding a multi-condition guiding architecture to improve the consistency of image editing results.
The generated target image is closer to the user's editing instructions, improving the consistency between the image editing effect and the editing instructions.
Smart Images

Figure CN119399327B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the fields of computer vision and artificial intelligence, and in particular to an image editing method, apparatus, storage medium, and electronic device. Background Technology
[0002] Image editing refers to the process of modifying, adjusting, and processing digital images using various techniques and tools to achieve specific visual effects, improve image quality, or alter image content. Image editing can involve local or overall changes to an image, including but not limited to color adjustment, brightness and contrast adjustment, image compositing, removing or adding elements, shape deformation, texture compositing, and image restoration.
[0003] In related technologies, a diffusion model can be used to edit the input image based on user-inputted editing instructions, such as a text command. However, this image editing method suffers from uncontrollable editing results, and the edited image often does not match the user's expectations. Summary of the Invention
[0004] Embodiments of this disclosure provide an image editing method, apparatus, storage medium, and electronic device.
[0005] According to one aspect of an embodiment of the present disclosure, an image editing method is provided, the method comprising: acquiring an image to be edited and editing instructions; generating lexical units of a target image based on the image to be edited and the editing instructions using a pre-trained multimodal large model; extracting visual lexical units of the target image from the lexical units of the target image; generating a latent space representation of the target image based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instructions, and the visual lexical units using a pre-trained diffusion model, wherein the pre-trained diffusion model has a multi-conditional guided architecture; and decoding the latent space representation of the target image to obtain the target image.
[0006] According to another aspect of the present disclosure, an image editing apparatus is provided, comprising: a first acquisition module for acquiring an image to be edited and an editing instruction; a first generation module for generating lexical units of a target image based on the image to be edited and the editing instruction using a pre-trained multimodal large model; a reading module for extracting visual lexical units of the target image from the lexical units of the target image; a second generation module for generating a latent space representation of the target image based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instruction, and the visual lexical units using a pre-trained diffusion model, wherein the pre-trained diffusion model has a multi-conditional guided architecture; and a decoding module for decoding the latent space representation of the target image to obtain the target image.
[0007] According to another aspect of the present disclosure, a computer-readable storage medium is provided, which stores a computer program for performing the image editing method described above.
[0008] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; a memory for storing processor-executable instructions; and a processor for reading executable instructions from the memory and executing the instructions to implement the image editing method described above.
[0009] Based on the image editing method, apparatus, storage medium, and electronic device provided in the above embodiments of this disclosure, when image editing is required, a pre-trained multimodal large model can be used to generate target image lexical units based on the image to be edited and the editing instructions. Visual lexical units of the target image are extracted from these lexical units. When using a pre-trained diffusion model for image editing, the visual lexical units of the target image and the text lexical units corresponding to the editing instructions can be used as guiding conditions to guide the diffusion model in image editing. The pre-trained diffusion model has a multi-condition guiding architecture. This technical solution, by adding visual lexical units generated by a multimodal large model as constraints, makes the generated target image more closely resemble the user's editing instructions, improving the editing effect of the target image and its consistency with the editing instructions.
[0010] The technical solutions of this disclosure will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0011] The above and other objects, features, and advantages of this disclosure will become more apparent from the more detailed description of the embodiments thereof in conjunction with the accompanying drawings. The drawings are provided to further illustrate the embodiments of this disclosure and form part of the specification. They are used together with the embodiments of this disclosure to explain the disclosure and do not constitute a limitation thereof. In the drawings, the same reference numerals generally represent the same components or steps.
[0012] Figure 1 This is a schematic flowchart of an image editing method provided in an exemplary embodiment of this disclosure.
[0013] Figure 2 This is a schematic diagram of an image editing generation framework provided by an exemplary embodiment of the present disclosure.
[0014] Figure 3 This is a schematic diagram of a model training process provided in an exemplary embodiment of this disclosure.
[0015] Figure 4 This is a schematic diagram of a model training framework provided in an exemplary embodiment of this disclosure.
[0016] Figure 5 This is a flowchart illustrating an image editing method provided in another exemplary embodiment of this disclosure.
[0017] Figure 6 This is a flowchart illustrating an image editing method provided in another exemplary embodiment of this disclosure.
[0018] Figure 7 This is a schematic diagram of an image editing generation framework provided by another exemplary embodiment of this disclosure.
[0019] Figure 8 This is a reference diagram illustrating the generation of lexical terms provided in an exemplary embodiment of this disclosure.
[0020] Figure 9 This is a schematic diagram of the image editing process of a diffusion model provided in an exemplary embodiment of this disclosure.
[0021] Figure 10 This is a schematic diagram of the structure of an image editing apparatus provided in an exemplary embodiment of the present disclosure.
[0022] Figure 11 This is a schematic diagram of the structure of an image editing apparatus provided in another exemplary embodiment of the present disclosure.
[0023] Figure 12 This is a structural diagram of an electronic device provided in an exemplary embodiment of this disclosure. Detailed Implementation
[0024] Hereinafter, exemplary embodiments according to the present disclosure will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are merely some embodiments of the present disclosure, and not all embodiments of the present disclosure, and it should be understood that the present disclosure is not limited to the exemplary embodiments described herein.
[0025] It should be noted that, unless otherwise specifically stated, the relative arrangement, numerical expressions, and values of the components and steps set forth in these embodiments do not limit the scope of this disclosure.
[0026] Those skilled in the art will understand that the terms "first," "second," etc., in the embodiments of this disclosure are only used to distinguish different steps, devices, or modules, and do not represent any specific technical meaning, nor do they indicate a necessary logical order between them.
[0027] It should also be understood that in the embodiments disclosed herein, "a plurality of" may refer to two or more, and "at least one" may refer to one, two or more.
[0028] It should also be understood that any component, data or structure mentioned in the embodiments of this disclosure can generally be understood as one or more unless expressly defined or given to the contrary in the context.
[0029] Furthermore, the term "and / or" in this disclosure is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this disclosure generally indicates that the preceding and following related objects have an "or" relationship.
[0030] It should also be understood that the description of the various embodiments in this disclosure emphasizes the differences between the various embodiments, and the similarities or similarities can be referred to each other. For the sake of brevity, they will not be described in detail.
[0031] At the same time, it should be understood that, for ease of description, the dimensions of the various parts shown in the accompanying drawings are not drawn according to actual scale.
[0032] The following description of at least one exemplary embodiment is merely illustrative and is in no way intended to limit this disclosure or its application or use.
[0033] Techniques, methods, and equipment known to those skilled in the art may not be discussed in detail, but where appropriate, such techniques, methods, and equipment should be considered part of the specification.
[0034] It should be noted that similar labels and letters in the following figures indicate similar items; therefore, once an item is defined in one figure, it does not need to be discussed further in subsequent figures.
[0035] The embodiments disclosed herein can be applied to electronic devices such as terminal devices, computer systems, and servers, and can operate together with a wide range of other general-purpose or special-purpose computing system environments or configurations. Examples of well-known terminal devices, computing systems, environments, and / or configurations suitable for use with electronic devices such as terminal devices, computer systems, and servers include, but are not limited to: personal computer systems, server computer systems, thin clients, thick clients, handheld or laptop devices, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputer systems, mainframe computer systems, and distributed cloud computing environments including any of the above systems, etc.
[0036] Electronic devices such as terminal devices, computer systems, and servers can be described in the general context of computer system executable instructions (such as program modules) executed by a computer system. Typically, program modules can include routines, programs, object programs, components, logic, data structures, etc., which perform specific tasks or implement specific abstract data types. Computer systems / servers can be implemented in distributed cloud computing environments, where tasks are executed by remote processing devices linked through communication networks. In distributed cloud computing environments, program modules can reside on local or remote computing system storage media, including storage devices.
[0037] This disclosure outlines
[0038] Current image editing solutions typically utilize diffusion models, directly executing the edits specified in the user-inputted commands on the image to be edited. However, editing an image solely based on these commands results in uncontrollable image quality. This disclosed technical solution uses a multimodal large model to generate visual lexical units for the target image based on the edit commands and the image to be edited. These visual lexical units are then used as constraints on the diffusion model, making the generated target image more closely resemble the user's edit commands and improving the image editing effect.
[0039] In addition, reference graph terms can be obtained from the reference image and used as constraints for the diffusion model, which can further improve the image editing effect of the diffusion model.
[0040] Exemplary methods
[0041] Figure 1 This is a schematic flowchart illustrating an exemplary embodiment of the image editing method provided in this disclosure. This embodiment can be applied to electronic devices, such as servers or mobile terminals, such as... Figure 1 As shown, it includes the following steps:
[0042] Step 101: Obtain the image to be edited and the editing instructions.
[0043] In this disclosure, the image to be edited is used to indicate the image that needs to be modified or adjusted, and the editing instructions are used to indicate the instructions to edit the image, such as adjusting the image color, removing or adding image elements, or replacing a certain image element in the image.
[0044] The editing instructions can be text instructions or voice instructions. Voice instructions can be recognized as text instructions through a voice recognition algorithm.
[0045] Step 102: Using a pre-trained multimodal large model, generate the lexical units of the target image based on the image to be edited and the editing instructions.
[0046] In this disclosure, a pre-trained multimodal large model is used to indicate a model trained by combining multimodal information such as text, images, and audio. The training process of the multimodal large model can be found in [reference needed]. Figure 3 The embodiments shown are not described in detail here.
[0047] In this embodiment, the trained multimodal large model can automatically generate the target image after editing based on the user's input editing instructions and the image to be edited.
[0048] The target image is the image obtained after editing the image to be edited according to the editing instructions.
[0049] Step 103: Extract visual words from the words of the target image.
[0050] The visual terms in this disclosure are used to indicate that the target image is broken down into smaller feature vectors. The visual terms can be used to calculate and match the multidimensional vectors of the latent space representation of the image to be edited by the diffusion model, thereby guiding the diffusion model to perform image editing and denoising to generate the target image.
[0051] In practice, textual and visual words can be extracted from the words in the target image obtained in step 102. Textual words are used to represent text features and are text used to describe the target image.
[0052] Step 104: Using a pre-trained diffusion model, a latent space representation of the target image is generated based on the latent space representation of the image to be edited, the text terms corresponding to the editing instructions, and the visual terms. The pre-trained diffusion model has a multi-conditional guided architecture.
[0053] In this disclosure, the pre-trained diffusion model is a pre-trained model for text-to-image editing. The diffusion model can generate the latent space representation of the target image based on the latent space representation of the image to be edited. During the process of generating the latent space representation of the target image using the diffusion model, visual terms and text terms corresponding to editing instructions are used as guiding conditions to guide the diffusion model to generate an edited image that better meets the user's needs.
[0054] Among them, the multi-condition guidance architecture refers to the pre-trained diffusion model using multiple guidance conditions to guide the image editing process.
[0055] The diffusion model mainly includes two processes: the diffusion process and the reversal process. The diffusion process involves gradually adding noise to the image to be edited until the image to be edited becomes completely noisy. This process is usually described using a stochastic process, such as Brownian motion. The reversal process involves starting with completely noisy data and gradually removing the noise in order to generate new data (the edited target image) that is similar to the image to be edited.
[0056] In this disclosure, during image editing using a pre-trained diffusion model, visual lexical units generated based on a multimodal large model and text lexical units corresponding to editing instructions can be used as conditional guidance to calculate cross-attention feature representations of the latent space representation of the image by the cross-attention module of the pre-trained diffusion model, and the target image's latent space representation is obtained based on these cross-attention feature representations. Specifically, in generating the cross-attention feature representations, weights (i.e., weighted fusion weights) are assigned to the visual lexical units and the text lexical units corresponding to editing instructions respectively to obtain the cross-attention feature representations. Optionally, these weighted fusion weights are automatically learned during the pre-training of the diffusion model.
[0057] In this embodiment, the latent space of the image to be edited is represented by the latent space expression obtained after encoding the image using an image encoder, such as a variational autoencoder (VAE).
[0058] Step 105: Decode the latent space representation of the target image to obtain the target image.
[0059] In this embodiment, VAE can be used to decode the latent space representation of the target image to obtain the pixel space representation of the target image.
[0060] For example, see Figure 2 This diagram illustrates the image editing generation framework. After the image to be edited and the editing instructions are processed by a multimodal large model, the target image's lexical units are obtained. Based on these lexical units, the target image's visual lexical units can be extracted. The image to be edited is then encoded by an image encoder to obtain its latent space representation, which is then input into a diffusion model. The diffusion model uses the textual lexical units corresponding to the editing instructions and the visual lexical units of the target image as guiding conditions to output the latent space representation of the target image expected by the user. Finally, the image is decoded by an image decoder to output the target image.
[0061] The method provided in the above embodiments of this disclosure, when image editing is required, can first generate target image lexical units based on the image to be edited and editing instructions using a pre-trained multimodal large model. Visual lexical units of the target image are extracted from these lexical units. When using a pre-trained diffusion model for image editing, the visual lexical units of the target image and the textual lexical units corresponding to the editing instructions can be used as guiding conditions to guide the diffusion model in image editing, improving the accuracy of the latent space representation of the generated target image, thereby improving the editing effect of the target image. The technical solution of this disclosure, by adding visual lexical units generated by a multimodal large model as constraints, makes the generated target image closer to the user's editing instructions, improving the editing effect of the target image and its consistency with the editing instructions.
[0062] In some optional implementations, a pre-trained diffusion model is used to generate the latent space representation of the target image based on the latent space representation of the image to be edited, the text terms corresponding to the editing instructions, and the visual terms. This includes: obtaining cross-attention feature representations based on the text terms corresponding to the editing instructions, the visual terms, and the latent space representation of the image to be edited; and removing noise to be removed from each time step of the image to be edited based on the cross-attention feature representations using the pre-trained diffusion model.
[0063] The cross-attention mechanism assigns attention weights to each feature based on the latent space representation of the input image to be edited, the text terms corresponding to the editing instructions, and the visual terms. This allows the pre-trained diffusion model to focus more on key features with higher weights (i.e., higher feature correlation) during the generation process. In this way, the pre-trained diffusion model can better conform to the user's editing instructions and preserve the data features of the original data (the image to be edited) when generating the edited target image, thereby improving the image quality of the edited target image.
[0064] The attention weights assigned to each feature are automatically learned by the pre-trained diffusion model.
[0065] In some alternative implementations, multimodal large models and diffusion models can be obtained through training, such as... Figure 3 As shown, the training process includes the following steps.
[0066] Step 301: Obtain a training sample set. The training sample set includes at least one set of training samples. Each set of training samples includes pre-edit image samples, editing instruction samples, real post-edit image samples, and editing result description samples.
[0067] The editing instruction samples can be text or audio samples. If they are audio samples, speech recognition can be performed to obtain the corresponding text. The actual edited image sample is the image sample corresponding to the image sample before editing. The editing result description sample is information describing the effect of the actual edited image sample in the sample.
[0068] The edit result description sample can be output by a pre-trained text generation model (Generative Pre-Trained, or GPT) based on the image sample before editing and the real image sample after editing (the real image sample after editing is obtained by the user after editing the image sample before editing according to the editing instruction sample).
[0069] Step 302: For each group of training samples, the image samples before editing and the editing instruction samples are used to obtain the corresponding image words using the first model to be trained. The image words include the first text words and the first visual words. And, based on the words of the editing result description sample and the first text words, the first loss value of the first model to be trained is obtained.
[0070] In this disclosure, the first model to be trained is capable of receiving input from at least one modality of data and predicting the corresponding image terms based on the image samples before editing and the editing instruction samples.
[0071] In this embodiment, the image sample before editing can be preprocessed, and the preprocessed image can be converted into a tensor. Specifically, the image sample before editing can be cropped, normalized, etc. Cropping can crop images of different sizes to the same size, while normalization can normalize the pixels of an image with pixel values of 0-255 to between -1 and 1, which helps to convert the image into a tensor of size (C,H,W), where C is the number of channels used to represent the number of color channels of the image, H is the height used to represent the vertical size of the image, i.e., the height of the image, and W is the width used to represent the horizontal size of the image, i.e., the width of the image.
[0072] For editing instruction samples, lexicalization can be performed to obtain the corresponding text words. Lexicalization can be understood as the process of segmenting the editing instruction sample into a sequence of words that the first training model can recognize, thus providing the first training model with recognizable and easily identifiable text data. By performing lexicalization on the editing instructions, the first training model can better understand the text.
[0073] Converting the pre-edit image samples into tensors and the editing instruction samples into text terms helps the first model to be trained to process them.
[0074] In order to enable the first model to be trained to edit the edit instruction sample and the image sample before editing, an instruction format acceptable to the multimodal large model can be constructed based on the edit instruction sample and the image sample before editing, and the instruction can be input into the first model to be trained so that the first model to be trained can perform label prediction.
[0075] The first text lexical unit is used to represent text features, which are features of descriptive text used to express the edited image; the first visual lexical unit is used to represent visual features, which are image features used to express the edited image.
[0076] In practical implementation, the first set number of characters of an image word can be used as the first text word, and the last set number of characters as the first visual word; alternatively, the last set number of characters of an image word can be used as the first text word, and the first set number of characters as the first visual word. An image word consists of text words and visual words, and their positions within the image word can be learned.
[0077] In this embodiment, the edit result description samples are lexicalized to obtain the lexical units of the edit result description samples. Then, the cross-entropy loss function can be used to obtain the first loss value for training the first model based on the lexical units of each first text and each edit result description sample.
[0078] In other implementations, each first text word can be decoded using a text decoder, and meaningless placeholders and markers can be removed from the decoded text through filtering to obtain each first text. Then, a text encoder, such as a multimodal pre-trained neural network (Contrastive Language-Image Pre-Training, or CLIP model for short), is used to encode the first text to obtain the corresponding second text words. Finally, the cross-entropy loss function is used to obtain each first loss value based on each second text word and the words in each edit result description sample.
[0079] For example, see Figure 4 After the image sample before editing and the editing instruction sample are labeled by the first training model, image tokens can be obtained. After the image tokens are truncated, the first text tokens (token) and the first visual tokens (token) can be obtained. After the first text tokens are decoded and filtered, the first text can be obtained. After the first text is encoded, the second text tokens can be obtained. Based on the tokens of the editing result description sample and the second text tokens, the cross-entropy loss (first loss value) can be obtained.
[0080] Step 303: For each set of training samples, based on the corresponding first visual lexical units, the text lexical units of the editing instruction samples, and the latent space representation of the image samples before editing, the second loss value of the second model to be trained is obtained; wherein, the second model to be trained has a diffusion model architecture guided by multiple conditions.
[0081] The second model to be trained is an initial diffusion model. The image editing process for the diffusion model can be found in [link to relevant documentation]. Figure 9 First, the original image is transformed from pixel space to latent space using a VAE encoder. Then, noise is added to the latent space representation of the image to obtain z. T Then, in the denoising process, a cross-attention mechanism is introduced, using the text vector representation corresponding to the editing instruction, reference graph lexical units (which may or may not contain reference graph lexical units), and the aforementioned first visual lexical units as constraints to constrain the process of the diffusion model in generating the denoised image, generating the latent space representation of the predicted edited image, and then decoding the latent space representation of the predicted edited image through the VAE encoder to obtain the output image.
[0082] The cross-attention mechanism assigns a weight to each of the input full-image features (the latent space representation of the image sample before editing), the text terms corresponding to each editing instruction sample, and each first visual term, enabling the model to focus more on key regions with higher weights (i.e., higher feature correlation) during image generation. This allows the model to better preserve and reproduce the key features of the second image sample when generating the edited image, thereby improving the quality and similarity of the generated image. The calculation method of the cross-attention mechanism is shown in Equation (1).
[0083]
[0084] In Equation (1), Attn(Q,K,V) represents the output of the attention mechanism, Q represents the weight matrix of the full image features, K and V represent the weight matrices of the constraints (reference image lexical, visual lexical, and text vector representation of editing instructions), and the softmax function is used to represent the normalization process.
[0085] In this disclosure, the guiding conditions of the cross-attention mechanism may include text lexical units corresponding to each editing instruction sample, first visual lexical units obtained based on the multimodal large model, and reference graph lexical units (which may or may not include reference graph lexical units). The guiding conditions can jointly guide the image editing process of the diffusion model, making image editing more accurate and effective, and meeting the user's editing needs.
[0086] In this disclosure, the second loss value may include the image mean square error value determined based on the first loss function and the noise mean square error value determined based on the second loss function.
[0087] The first loss function is used to determine the image mean square error (MSE) of the latent space representation of each predicted edited image and the latent space representation of each real edited image sample. Each real edited image sample is a labeled sample contained in the training sample set, obtained by the user editing the unedited image sample based on editing instructions. The latent space representation of the predicted edited image is the predicted sample output by the second training model based on the aforementioned guiding conditions and the unedited image sample. Determining the latent space representation of both the real and predicted edited image samples using the first loss function determines the image quality.
[0088] The second loss function is a commonly used function in diffusion models to determine the loss value based on noise. Using the second loss function, the latent space representation of the predicted edited image and the noise mean square error of the actual edited image samples can be determined.
[0089] The second loss value can be obtained based on the image mean square error value determined by the first loss function and the noise mean square error value determined by the second loss function.
[0090] It is understandable that if all training samples in the training sample set contain a reference image, the trained diffusion model needs to use the reference image terminology of the reference image as a guiding condition to guide the image generation process; if some training samples in the second training sample set contain a reference image and some training samples do not contain a reference image, the trained diffusion model can use the reference image terminology of the reference image as a guiding condition to guide the image generation process, or it can choose not to use the reference image terminology of the reference image as a guiding condition.
[0091] Step 304: Based on the first loss value and the second loss value, guide the training of the first model to be trained and the second model to be trained to obtain the pre-trained multimodal large model and the pre-trained diffusion model.
[0092] The method provided in the above embodiments of this disclosure improves the accuracy of the multimodal large model by using the edit result description text output by the trained text generation model as the sample label for the training samples of the multimodal large model during training. This helps to generate more accurate constraints (visual lexical units) for the diffusion model, thereby improving the image editing effect of the diffusion model. In addition, by introducing a cross-attention mechanism, the trained diffusion model can pay more attention to key regions with higher weights when editing images, improving the quality of the generated images. Furthermore, by adding the loss calculation of the image mean square error value to the original diffusion model's loss based on noise determination during training, the images generated by the trained diffusion model are closer to the user's editing instructions, thereby improving the image editing effect of the diffusion model.
[0093] Figure 5 This is a flowchart illustrating an image editing method provided in another exemplary embodiment of this disclosure, such as... Figure 5 As shown, the image editing method includes the following steps.
[0094] Step 501: Obtain the image to be edited and the editing instructions.
[0095] Step 502: Perform lexicalization on the editing instructions to obtain instruction lexicals, and convert the preprocessed image into a tensor after preprocessing the image to be edited.
[0096] In this embodiment, lexicalization of editing instructions can be understood as the process of segmenting the editing instructions into a sequence of lexical terms that can be recognized by the multimodal large model, thereby providing the multimodal large model with recognizable and easily identifiable text data. Lexicalization of editing instructions helps the multimodal large model understand and generate text.
[0097] In this embodiment, the preprocessing of the image to be edited may include operations such as cropping and normalization. Cropping can cut images of different sizes into the same size, while normalization can normalize the pixels of an image with a pixel value of 0-255 to between -1 and 1, which helps to convert the image into a tensor of size (C,H,W).
[0098] Step 503: Generate model query instructions based on instruction terms and tensors.
[0099] In this disclosure, the model query instruction is an instruction format acceptable to multimodal large models. The model query instruction is generated by using instruction terms and image tensors. Both the editing instruction and the image to be edited can be constructed in the instruction so that the multimodal large model can generate and output the results.
[0100] Exemplarily, the model query instruction is as follows: Conversation(system="A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions.", roles=('USER', 'ASSISTANT'), messages=[['USER', "what will this image be like if′'remove bananas and add grapes'(in a short paragraph)\n<im_start><im_patch>......<im_patch><im_end>"], ['ASSISTANT', None]], offset=0, sep_style=<SeparatorStyle.TWO:2>, sep=",", sep2="", version='Unknown', skip_next=False), and input this instruction into the multimodal large model so that the multimodal large model can generate and output results.
[0101] The Chinese interpretation of the above model query instruction is "Conversation(system='A chat between a curious user and an artificial intelligence assistant. The assistant gives helpful, detailed, and polite answers to the user's questions.', roles=('user', 'assistant'), messages=[['user', "what will this image be like if'remove bananas and add grapes'(in a short paragraph)\n<im_start><im_patch>…<im_patch><im_end>"], ['assistant', None]], offset=0, sep_style=<SeparatorStyle.TWO:2>, sep='', sep2='', version='Unknown', skip_next=False)", that is, it is hoped that the multimodal large model will output an instruction to remove the banana element and add the grape element in the image.
[0102] Step 504, using the pre-trained multimodal large model, based on the model query instruction, generate the tokens of the target image, and intercept the visual tokens of the target image from the tokens of the target image.
[0103] The pre-trained multimodal large model generates lexical units for the target image. Based on these lexical units, textual and visual lexical units of the target image can be obtained through image decomposition. The visual lexical units can be used in step 605 to generate the latent space representation of the target image using the constrained diffusion model. The textual lexical units are used to describe the result after image editing. For example, after inputting the instruction shown in step 603, the multimodal large model can output "Remove the bananas from the image and add grapes. Edit the image to add the grapes and arrange them on the fruitstand. Finished."
[0104] Step 505: Using a pre-trained diffusion model, generate the latent space representation of the target image based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instructions, and the visual lexical units of the target image.
[0105] Step 506: Decode the latent space representation of the target image to obtain the target image.
[0106] In some embodiments, the implementation of steps 505 to 506 can be found in [reference needed]. Figure 1 The descriptions of steps 104 to 105 in the illustrated embodiment will not be detailed here.
[0107] The method provided in the above embodiments of this disclosure discloses a specific implementation of generating visual lexical units using a multimodal large model, which helps to constrain the process of generating target images using a diffusion model through visual lexical units generated by a multimodal large model.
[0108] In some other implementations, an additional reference image can be added to guide the diffusion model in generating the target image. Figure 6 This is a schematic flowchart of an image editing method provided in another exemplary embodiment of this disclosure. This embodiment exemplifies how to edit an image by combining a reference image, visual terms of the target image, and text terms corresponding to editing instructions. Figure 6 As shown, the image editing method includes the following steps.
[0109] Step 601: Obtain the image to be edited, editing instructions, and at least one reference image.
[0110] In this disclosure, the reference image includes at least one of the following: local image information to be added or removed in the image to be edited, image style information of the image to be edited, and image filter information of the image to be edited.
[0111] For example, when it's necessary to adjust the background color, partial color of an object, or the shape of a person's face or the image style of an image to be edited, a reference image can be acquired simultaneously with the image to be edited and the editing instructions. For instance, to adjust the background color of the image to be edited to the blue background of an ID photo, a reference image with a blue background can be acquired; similarly, to adjust the image style of the image to be edited to a Disney style or a pastoral style, a Disney-style or pastoral-style reference image can be selected. The reference image can also be an image of an element to be added; for example, to add or remove a dog image element from the image to be edited, an image of a dog can be selected as a reference image. In this embodiment, by adding the constraint of the reference image, the accuracy and effectiveness of image editing are enhanced.
[0112] Furthermore, reference image lexical units corresponding to the reference image can be obtained through step 602, and visual lexical units of the target image can be obtained through step 603. This disclosure does not limit the execution order of steps 602 and 603; step 602 can be executed first and then step 603, or step 603 can be executed first and then step 602, or steps 602 and 603 can be executed simultaneously.
[0113] Step 602: Extract features from at least one reference image to obtain reference image features, and generate reference image words based on the reference image features using a multilayer perceptron.
[0114] In this disclosure, the reference image features are feature vectors obtained by encoding a reference image using a feature extractor based on a convolutional neural network. The reference image terms are image-text features that can be used in a diffusion model, obtained by capturing, refining, combining, and mapping the reference image features using a multilayer perceptron.
[0115] In practice, the reference image can be preprocessed first, such as resizing, cropping, and standardizing. Since different reference images may have different sizes, resizing and cropping can be used to make the images the same size, thereby improving the model's feature extraction efficiency. After image preprocessing, the image can be input into a feature extractor, such as a Visual Geometry Group Network (VGG), for encoding to obtain the reference image features.
[0116] In this embodiment, to enable the reference image features to be used as guiding conditions for the diffusion model and to guide the image generation process of the diffusion model, a mapping process can be further performed on the extracted reference image features. For specific implementation details, see [link to implementation details]. Figure 8This can be implemented using a Multilayer Perceptron (MLP), which is a fully connected neural network structure containing multiple hidden layers (such as...). Figure 8 Hidden layers 1, 2, 3, and 4 in the multilayer perceptron help capture reference graph features at different levels of abstraction. Then, through layer-by-layer propagation and weighting, features are progressively refined and combined to obtain higher-level representations of reference graph attributes and characteristics. By training and fine-tuning the parameters of the multilayer perceptron, reference graph features can be mapped from the original feature space to the graph-text feature space of the diffusion model, enabling more effective supervision of the diffusion model training.
[0117] Step 603: Using a pre-trained multimodal large model, generate the target image's lexical units based on the image to be edited and the editing instructions, and extract the target image's visual lexical units from the target image's lexical units.
[0118] The implementation method of step 603 can be found in [reference needed]. Figure 1 The descriptions of steps 102-103 in the illustrated embodiment will not be detailed here.
[0119] Step 604: Using a pre-trained diffusion model, generate the latent space representation of the target image based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instructions, the visual lexical units of the target image, and the lexical units of the reference image.
[0120] In this disclosure, the latent space representation of the image to be edited is used to indicate the representation of the image to be edited from the pixel space to the latent space, which helps to improve the computational efficiency of the diffusion model. In specific implementation, an image encoder can be used to encode the image to be edited to obtain the latent space representation of the image to be edited.
[0121] In this disclosure, the text lexical units corresponding to editing instructions are used to indicate the transformation of the editing instructions from instruction text into a text vector representation that can be combined with the latent space representation of an image. The editing instructions are encoded using a text encoder to obtain the text lexical units corresponding to the editing instructions. Specifically, the CLIP model can be used to encode the text corresponding to the editing instructions to obtain the text lexical units corresponding to the editing instructions.
[0122] The CLIP model is a deep learning model that combines image and text information. It can be trained by simultaneously considering the image and its associated text description, thereby representing them in a unified latent space. Using the CLIP model, a text vector representation can be output based on the input text. This text vector representation is used in the diffusion model, and it can be combined with the image's VAE latent space representation to obtain an embedding that simultaneously considers image and text information. This embedding is then used as a guiding condition for the diffusion model, constraining the image generation process of the diffusion model.
[0123] Step 605: Decode the latent space representation of the target image to obtain the target image.
[0124] For example, see Figure 7 This diagram illustrates the generative framework for image editing. A reference image, after feature extraction and mapping, yields reference image lexical units. The image to be edited and the editing instructions are processed by a multimodal large model to obtain target image lexical units. Visual lexical units of the target image can be extracted from these lexical units. The image to be edited is encoded by an image encoder to obtain its latent space representation, which is then input into a diffusion model. During the diffusion process, noise is progressively added to the latent space representation of the image to be edited, resulting in a noisy image. In the reverse process, the noisy image is progressively denoised, constrained by the text vector representation of the editing instructions, visual lexical units, and reference image lexical units, to obtain the latent space representation of the target image. Finally, the image is decoded by an image decoder to output the target image.
[0125] The method provided in the above embodiments of this disclosure, when image editing is required, can first utilize a pre-trained multimodal large model to generate target image lexical units based on the image to be edited and editing instructions. Visual lexical units of the target image are extracted from the target image lexical units, and reference image lexical units are generated based on a reference image. Thus, when using a pre-trained diffusion model for image editing, the visual lexical units of the target image and the reference image lexical units of the reference image can serve as guiding conditions to guide the image editing process, improving the accuracy of the latent space representation of the generated target image, thereby improving the editing effect of the target image. The technical solution of this disclosure can add visual lexical units and reference image lexical units generated by a multimodal large model as guiding conditions, more effectively supervising and guiding the image editing process of the diffusion model, making the generated target image closer to the user's editing instructions, improving the editing effect of the target image and its consistency with the editing instructions.
[0126] Exemplary device
[0127] Figure 10 This is a schematic diagram of the structure of an image editing apparatus provided in an exemplary embodiment of the present disclosure, such as... Figure 10As shown, the device may include:
[0128] The first acquisition module 111 is used to acquire the image to be edited and the editing instructions;
[0129] The first generation module 112 is used to generate lexical units of the target image based on the image to be edited and the editing instructions using a pre-trained multimodal large model;
[0130] The reading module 113 is used to extract visual words from the words of the target image;
[0131] The second generation module 114 is used to generate the latent space representation of the target image based on the latent space representation of the image to be edited, the text words corresponding to the editing instructions, and the visual words using a pre-trained diffusion model. The pre-trained diffusion model has a multi-condition guided architecture.
[0132] The decoding module 115 is used to decode the latent space representation of the target image to obtain the target image.
[0133] Figure 11 This is a schematic diagram of the structure of an image editing apparatus provided in another exemplary embodiment of this disclosure, such as... Figure 11 As shown, in Figure 10 Based on the illustrated embodiment, in some implementations, the second generation module 114 includes:
[0134] The cross-attention submodule 1141 is used to obtain cross-attention feature representation based on the cross-attention mechanism, according to the text lexical units, visual lexical units, and latent space representation of the image to be edited corresponding to the editing instruction;
[0135] The denoising submodule 1142 is used to remove the noise to be removed from the image to be edited at each time step based on the cross-attention feature representation using a pre-trained diffusion model, thereby obtaining the latent space representation of the target image.
[0136] In some implementations, it also includes: a model training module 116 for training to obtain a pre-trained multimodal large model and a pre-trained diffusion model;
[0137] Model training module 116 includes:
[0138] The first acquisition submodule 1161 is used to acquire a training sample set, which includes at least one set of training samples. Each set of training samples includes a pre-edit image sample, an editing instruction sample, a real post-edit image sample, and an editing result description sample.
[0139] The first prediction submodule 1162 is used to obtain corresponding image words for each pre-edit image sample and editing instruction sample in each training sample using the first model to be trained, wherein the image words include first text words and first visual words; and to obtain a first loss value for training the first model to be trained based on the words of the editing result description sample and the first text words.
[0140] The second prediction submodule 1163 is used to obtain a second loss value of the second model to be trained for each set of training samples based on the corresponding first visual lexical unit, the text lexical unit of the editing instruction sample, and the latent space representation of the image sample before editing; wherein the second model to be trained has a multi-condition guided diffusion model architecture.
[0141] The training submodule 1164 is used to guide the training of the first model to be trained and the second model to be trained based on the first loss value and the second loss value, so as to obtain the pre-trained multimodal large model and the pre-trained diffusion model.
[0142] In some implementations, reference image samples are also included in response to each set of training samples;
[0143] The second prediction submodule 1163 is used to obtain the second loss value of the second model to be trained for each group of training samples based on the corresponding first visual lexical unit, the text lexical unit of the editing instruction sample, the latent space representation of the image sample before editing, and the reference graph lexical unit.
[0144] In some embodiments, the apparatus further includes:
[0145] The third generation module 117 is used to generate model query instructions based on instruction lexical units and tensors;
[0146] The first generation module 112 is used to generate lexical units of the target image based on the model query command using a pre-trained multimodal large model.
[0147] In some embodiments, the apparatus further includes:
[0148] The second acquisition module 118 is used to acquire at least one reference image; wherein the reference image includes at least one of the following information: local image information to be added or removed in the image to be edited, image style information of the image to be edited, and image filter information of the image to be edited;
[0149] Feature extraction module 119 is used to extract features from at least one reference image to obtain reference image features;
[0150] The fourth generation module 120 is used to generate reference graph lexical units based on the features of the reference graph using a multilayer perceptron.
[0151] In some implementations, the second generation module 114 is used to generate a latent space representation of the target image based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instructions, the visual lexical units of the target image, and the reference graph lexical units, using a pre-trained diffusion model.
[0152] It should be noted that the modules in this device can be disassembled and / or recombined, and these disassemblies and / or recombinations should be considered as equivalent solutions of this device.
[0153] The exemplary embodiments of this device correspond to the exemplary method section described above, and the relevant content can be referenced and cited interchangeably. The beneficial technical effects corresponding to the exemplary embodiments of this device can be found in the corresponding beneficial technical effects of the exemplary method section described above, and will not be repeated here.
[0154] Exemplary electronic devices
[0155] Figure 12 A structural diagram of an electronic device provided in an embodiment of this disclosure includes at least one processor 121 and a memory 122.
[0156] The processor 121 may be a central processing unit (CPU) or other form of processing unit with data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 12 to perform desired functions.
[0157] The memory 122 may include one or more computer program products, which may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. Volatile memory may include, for example, random access memory (RAM) and / or cache memory. Non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 121 may execute one or more computer program instructions to implement the vehicle pose detection method and / or other desired functions of the various embodiments of this disclosure described above.
[0158] In one example, the electronic device may also include an input device 123 and an output device 124, which are interconnected via a bus system and / or other forms of connection mechanism (not shown).
[0159] The input device 123 may also include, for example, a keyboard, a mouse, a touch screen, a pickup device (such as a microphone array), etc.
[0160] The output device 124 can output various information to the outside, including, for example, a display, a speaker, a printer, and a communication network and its connected remote output devices, etc.
[0161] Of course, for the sake of simplicity, Figure 12 Only some of the components of the electronic device relevant to this disclosure are shown, omitting components such as buses, input / output interfaces, etc. In addition, the electronic device may include any other suitable components depending on the specific application.
[0162] Exemplary computer program products and computer-readable storage media
[0163] In addition to the methods and apparatus described above, embodiments of this disclosure may also be computer program products, including computer program instructions that, when executed by a processor, cause the processor to perform the steps in the image editing methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section of this specification.
[0164] Computer program products can be written in any combination of one or more programming languages to perform the operations of embodiments of this disclosure. The programming languages include object-oriented programming languages such as Java and C++, as well as conventional procedural programming languages such as C or similar languages. The program code can be executed entirely on a user's computing device, partially on a user's computing device, as a standalone software package, partially on a user's computing device and partially on a remote computing device, or entirely on a remote computing device or server.
[0165] Furthermore, embodiments of this disclosure may also be computer-readable storage media storing computer program instructions thereon, which, when executed by a processor, cause the processor to perform the steps in the image editing methods according to various embodiments of this disclosure as described in the "Exemplary Methods" section above.
[0166] Computer-readable storage media may take the form of any combination of one or more readable media. A readable medium may be a readable signal medium or a readable storage medium. A readable storage medium may, for example, include, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples of readable storage media (a non-exhaustive list) include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0167] The basic principles of this disclosure have been described above with reference to specific embodiments. However, it should be noted that the advantages, benefits, and effects mentioned in this disclosure are merely examples and not limitations, and should not be considered as essential features of each embodiment of this disclosure. Furthermore, the specific details disclosed above are for illustrative and facilitative purposes only, and are not limitations. These details do not limit the scope of this disclosure to the necessity of employing the aforementioned specific details for implementation.
[0168] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For system embodiments, since they largely correspond to method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0169] The block diagrams of devices, apparatuses, and devices involved in this disclosure are merely illustrative examples and are not intended to require or imply that they must be connected, arranged, or configured in the manner shown in the block diagrams. As those skilled in the art will recognize, these devices, apparatuses, and devices can be connected, arranged, and configured in any manner. Words such as “comprising,” “including,” “having,” etc., are open-ended terms meaning “including but not limited to,” and are used interchangeably with them. The terms “or” and “and” as used herein refer to the terms “and / or,” and are used interchangeably with them unless the context clearly indicates otherwise. The term “such as” as used herein refers to the phrase “such as but not limited to,” and is used interchangeably with it.
[0170] The methods and apparatus of this disclosure may be implemented in many ways. For example, they may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above-described order of steps for the method is for illustrative purposes only, and the steps of the method of this disclosure are not limited to the order specifically described above, unless otherwise specifically stated. Furthermore, in some embodiments, this disclosure may also be implemented as a program recorded on a recording medium, the program including machine-readable instructions for implementing the method according to this disclosure. Thus, this disclosure also covers recording media storing programs for performing the method according to this disclosure.
[0171] It should also be noted that in the apparatus, devices, and methods of this disclosure, the components or steps can be disassembled and / or recombined. These disassemblies and / or recombinations should be considered as equivalent solutions to this disclosure.
[0172] The above description of the disclosed aspects is provided to enable any person skilled in the art to make or use this disclosure. Various modifications to these aspects will be readily apparent to those skilled in the art, and the general principles defined herein may be applied to other aspects without departing from the scope of this disclosure. Therefore, this disclosure is not intended to be limited to the aspects shown herein, but rather to be carried out within the widest scope consistent with the principles and novel features disclosed herein.
[0173] The above description has been given for purposes of illustration and description. Furthermore, this description is not intended to limit the embodiments of this disclosure to the forms disclosed herein. Although numerous exemplary aspects and embodiments have been discussed above, those skilled in the art will recognize certain variations, modifications, alterations, additions, and sub-combinations thereof.
Claims
1. An image editing method, characterized in that, include: Obtain the image to be edited and the editing instructions; Using a pre-trained multimodal large model, based on the image to be edited and the editing instructions, lexical units of the target image are generated; wherein, the lexical units of the target image include text lexical units and visual lexical units, the text lexical units are used to express the features of the descriptive text of the edited image; the visual lexical units are used to represent the image features of the edited image; Visual words of the target image are extracted from the words of the target image, wherein the visual words of the target image are used to calculate and match with the multidimensional vector of the latent space representation of the image to be edited, so as to guide the pre-trained diffusion model to perform image editing and denoising to generate the target image. Using the pre-trained diffusion model, the latent space representation of the target image is generated based on the latent space representation of the image to be edited, the text terms corresponding to the editing instructions, and the visual terms; wherein, the pre-trained diffusion model has a multi-condition guided architecture; The latent space representation of the target image is decoded to obtain the target image.
2. The method according to claim 1, characterized in that, The step of generating the latent space representation of the target image using the pre-trained diffusion model, based on the latent space representation of the image to be edited, the text terms corresponding to the editing instructions, and the visual terms, includes: Based on the cross-attention mechanism, the cross-attention feature representation is obtained according to the text lexical units corresponding to the editing instruction, the visual lexical units, and the latent space representation of the image to be edited; Using the pre-trained diffusion model and based on the cross-attention feature representation, the noise to be removed at each time step of the image to be edited is removed to obtain the latent space representation of the target image.
3. The method according to any one of claims 1-2, characterized in that, Before generating the target image's lexical units based on the image to be edited and the editing instructions using a pre-trained multimodal large model, the process further includes: Based on the image to be edited and the editing instructions, generate a model query instruction; The step of generating lexical units for the target image using a pre-trained multimodal large model, based on the image to be edited and the editing instructions, includes: Using the pre-trained multimodal large model, the target image's lexical units are generated based on the model's query instructions.
4. The method according to any one of claims 1-3, characterized in that, The process of acquiring the image to be edited and the editing instructions also includes: Obtain at least one reference image; wherein the reference image includes at least one of the following: local image information to be added or removed in the image to be edited, image style information of the image to be edited, and image filter information of the image to be edited; Feature extraction is performed on the at least one reference image to obtain reference image features; Using a multilayer perceptron, reference graph lexical units are generated based on the features of the reference graph.
5. The method according to claim 4, characterized in that, The step of generating the latent space representation of the target image using a pre-trained diffusion model, based on the latent space representation of the image to be edited, the text terms corresponding to the editing instructions, and the visual terms, includes: Using a pre-trained diffusion model, the latent space representation of the target image is generated based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instructions, the visual lexical units of the target image, and the lexical units of the reference image.
6. The method according to any one of claims 1-3, characterized in that, The pre-trained multimodal large model and the pre-trained diffusion model are obtained through the following steps: Obtain a training sample set, which includes at least one set of training samples. Each set of training samples includes pre-edit image samples, editing instruction samples, real post-edit image samples, and editing result description samples. For each set of training samples, the image units before editing and the editing instruction samples are obtained using the first model to be trained. The image units include the first text units and the first visual units. And, based on the word units of the edited sample and the first text word units, a first loss value for training the first model to be trained is obtained; For each set of training samples, a second loss value for the second model to be trained is obtained based on the corresponding first visual lexical unit, the text lexical unit of the editing instruction sample, and the latent space representation of the image sample before editing; wherein, the second model to be trained has a multi-condition guided diffusion model architecture; Based on the first loss value and the second loss value, the first model to be trained and the second model to be trained are guided to obtain the pre-trained multimodal large model and the pre-trained diffusion model.
7. The method according to claim 6, characterized in that, In response to each set of training samples also including reference image samples, reference image terms of the reference image samples are obtained; The process of obtaining the second loss value of the second model to be trained includes: For each set of training samples, a second loss value for the second model to be trained is obtained based on the corresponding first visual lexical unit, the text lexical unit of the editing instruction sample, the latent space representation of the image sample before editing, and the reference graph lexical unit.
8. An image editing device, characterized in that, include: The first acquisition module is used to acquire the image to be edited and the editing instructions; The first generation module is used to generate lexical units of the target image based on the image to be edited and the editing instructions using a pre-trained multimodal large model; wherein, the lexical units of the target image include text lexical units and visual lexical units, the text lexical units are used to express the features of the descriptive text of the edited image; the visual lexical units are used to represent the image features of the edited image; The reading module is used to extract visual words from the words of the target image. The visual words of the target image are used to calculate and match with the multidimensional vector of the latent space representation of the image to be edited, so as to guide the pre-trained diffusion model to perform image editing and denoising to generate the target image. The second generation module is used to generate the latent space representation of the target image based on the latent space representation of the image to be edited, the text lexical units corresponding to the editing instructions, and the visual lexical units, using the pre-trained diffusion model; wherein the pre-trained diffusion model has a multi-condition guided architecture. The decoding module is used to decode the latent space representation of the target image to obtain the target image.
9. A computer-readable storage medium storing a computer program for performing the method according to any one of claims 1-7.
10. An electronic device, the electronic device comprising: processor; Memory used to store the processor's executable instructions; The processor is configured to read the executable instructions from the memory and execute the instructions to implement the method described in any one of claims 1-7.
Citation Information
Patent Citations
Method for editing face image through text, terminal and storage medium
CN117576257A