Image editing processing method and device, electronic equipment and storage medium

By segmenting and cropping images, combining style and content description text, using the image diffusion model to generate high-quality image editing results, solving the problem of instability in the prior art and achieving high-quality image editing effects.

CN120298545AActive Publication Date: 2025-07-11HUNAN HAPPLY SUNSHINE INTERACTIVE ENTERTAINMENT MEDIA CO LTD

Patent Information

Application Number
CN202510686418.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-27
Publication Date
2025-07-11
Estimated Expiration
2045-05-27

AI Technical Summary

Technical Problem

现有技术难以稳定实现高质量的图像编辑结果,尤其是在图像内容复杂或改动较大的情况下,无法满足用户的实际需求。

Method used

By obtaining the editing request input by the user, the target mask image is segmented and the image is cropped, and the style description text, content description text and task trigger words are combined, and the image diffusion model is used to process it to generate high-quality cropped result images.

Benefits of technology

It realizes precise positioning of the objects to be edited under editing requests, maintains the overall consistency of the image style before and after editing, has broad applicability and extremely high editing quality, and meets the actual needs of users.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120298545A_ABST
    Figure CN120298545A_ABST
Patent Text Reader

Abstract

The invention provides an image editing processing method and device, electronic equipment and a storage medium, and relates to the technical field of image processing, and the method comprises the steps: obtaining an editing request input by a user, and carrying out the segmentation processing of a to-be-edited object click coordinate in an initial image and editing data, and obtaining a target mask image; determining a cutting image according to the target mask image and the initial image; performing style description on the initial image to obtain a corresponding style description text; processing the text instruction to obtain a content description text and a task trigger word; processing based on the cut image, the target mask image, the style description text, the content description text and the task trigger word to obtain a cut result image; and according to the target mask image and the initial coordinate point and width and height data of the cutting area, combining the cutting result image with the original image to obtain a result image with the same resolution as the original image. Therefore, a high-quality editing result is realized, and actual requirements of users are met.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and in particular, to an image editing and processing method, apparatus, electronic device, and storage medium. Background Art

[0002] Image editing refers to modifying and adjusting images using computer algorithms and technologies. More specifically, image editing involves modifying the appearance, structure, and content of images: from subtle changes in image color and texture to operations such as adding, replacing, and removing image objects to create complex image content.

[0003] Current image editing methods edit images through a simple text and image input corresponding image editing model to generate a target image corresponding to the descriptive text. Since only a single model can be used to edit and process text and images, it is difficult to stably achieve high-quality editing results for tasks with complex image content or large content changes, and cannot meet the actual needs of users. Summary of the Invention

[0004] In view of this, embodiments of the present invention provide an image editing and processing method, apparatus, electronic device, and storage medium to solve the problem in the prior art that it is difficult to stably achieve high-quality editing results and cannot meet the actual needs of users.

[0005] To achieve the above object, embodiments of the present invention provide the following technical solutions:

[0006] A first aspect of an embodiment of the present invention shows an image editing and processing method, the method including:

[0007] Obtain an editing request input by a user, the editing request including an initial image, a text instruction, and editing data;

[0008] Perform segmentation processing on the initial image and the selected coordinates of the object to be edited in the editing data to obtain a target mask image;

[0009] Determine a cropped image according to the target mask image and the initial image;

[0010] Describe the style of the initial image to obtain a corresponding style description text;

[0011] Process the text instruction to obtain a content description text and a task trigger word;

[0012] Perform processing based on the cropped image, target mask image, style description text, content description text, and task trigger word to obtain a cropped result image;

[0013] Perform image processing on the cropped result image, the target mask image, and the initial image to obtain the target image.

[0014] Optionally, perform segmentation processing on the initial image and the selected coordinates of the object to be edited in the edit data to obtain the target mask image, including:

[0015] Call a pre-trained image segmentation model to process the initial image according to the selected coordinates of the object to be edited in the edit request to obtain an initial mask image, where the pre-trained image segmentation model is trained based on a sample set;

[0016] Respond to an adjustment request corresponding to the initial mask image to adjust the initial mask image based on the adjustment request to obtain the target mask image.

[0017] Optionally, determine the cropped image according to the target mask image and the initial image, including:

[0018] Calculate the position coordinates of the target mask image in the initial image and use them as the cropping region;

[0019] Crop the image corresponding to the cropping region from the initial image to obtain the cropped image.

[0020] Optionally, process the text instruction to obtain the content description text and the task trigger word, including:

[0021] Extract instructions from the text instruction to obtain the content description text corresponding to the scene of the initial image;

[0022] Identify the intent of the text instruction to obtain the edit task category;

[0023] Determine the task trigger word corresponding to the edit task category.

[0024] Optionally, extract instructions from the text instruction to obtain the content description text corresponding to the scene of the initial image, including:

[0025] Use a preset large language model to remove the text irrelevant to the content description in the text instruction to obtain the target text instruction, where the preset large language model is trained based on multiple historical text instructions, the initial images corresponding to the text instructions, and the content description texts corresponding to the scenes;

[0026] Expand and enhance the target text instruction to determine the content description text corresponding to the scene of the initial image.

[0027] Optionally, perform processing based on the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain the cropped result image, including:

[0028] Trigger the corresponding editing task based on the task trigger word to load the corresponding image diffusion model and the corresponding pre-trained editing task;

[0029] Adjust the model parameters of the image diffusion model based on the pre-trained editing task to obtain an adjusted image diffusion model;

[0030] Encode the cropped image using the image encoder in the adjusted image diffusion model, and use the encoded cropped image as an image latent variable;

[0031] Add noise to the image latent variable in the corresponding mask area according to the target mask map using the adjusted image diffusion model to obtain a noise-added image latent variable;

[0032] Encode the merged style description text and content description text using the image encoder in the adjusted image diffusion model to obtain a text latent variable;

[0033] Process the merged text latent variable and the noise-added image latent variable based on the adjusted image diffusion model to obtain a cropped result image.

[0034] Optionally, perform image processing on the cropped result image, the target mask map, and the initial image to obtain a target image, including:

[0035] Expand the mask area in the target mask map and blur the expanded part to obtain a blurred mask map;

[0036] Merge the cropped result image corresponding to the mask area in the blurred mask map into the initial image to obtain a target image.

[0037] A second aspect of the embodiments of the present invention shows an image editing processing device, the device includes:

[0038] An input module, configured to obtain an editing request input by a user, where the editing request includes an initial image, a text instruction, and editing data;

[0039] A mask generation and processing module, configured to perform segmentation processing on the initial image and the selected coordinates of the object to be edited in the editing data to obtain a target mask map; determine a cropped image according to the target mask map and the initial image;

[0040] A style reverse deduction module, configured to describe the style of the initial image to obtain a corresponding style description text;

[0041] A text processing module for processing the text instruction to obtain a content description text and a task trigger word;

[0042] An image redrawing module for processing based on the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain a cropped result image;

[0043] A post-processing module for performing image processing on the cropped result image, the target mask image, and the initial image to obtain a target image.

[0044] A third aspect of the embodiments of the present invention shows an electronic device, which includes a processor and a memory. The memory is used to store program codes and data for data generation, and the processor is used to call the program instructions in the memory to execute the image editing processing method as shown in the first aspect of the embodiments of the present invention.

[0045] A fourth aspect of the embodiments of the present invention shows a storage medium, which includes a stored program. When the program runs, it controls the device where the storage medium is located to execute the image editing processing method as shown in the first aspect of the embodiments of the present invention.

[0046] Based on the image editing and processing method, apparatus, electronic device, and storage medium provided in the embodiments of the present invention described above, the method includes: obtaining an editing request input by a user, where the editing request includes an initial image, a text instruction, and editing data; performing segmentation processing on the click coordinates of the object to be edited in the initial image and the editing data to obtain a target mask image; determining a cropped image according to the target mask image and the initial image; performing style description on the initial image to obtain a corresponding style description text; processing the text instruction to obtain a content description text and a task trigger word; performing processing based on the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain a cropped result image; performing image processing according to the cropped result image, the target mask image, and the initial image to obtain a target image. In the embodiments of the present invention, first, a target mask image corresponding to the editing request needs to be segmented, and then a cropped image corresponding to the target mask image is cropped from the initial image; then, a style description is performed on the initial image, and a content description text and a task trigger word corresponding to the file instruction are determined; then, the cropped image, the target mask image, the style description text, the content description text, and the task trigger word are processed to obtain a cropped result image; then, according to the target mask image and the starting coordinates and width and height data of the cropping area, the cropped result image is merged with the original image to obtain a result image with the same resolution as the original image. Thus, precise positioning of the object to be edited under the editing request is achieved, maintaining the overall consistency of the image style before and after editing, having wide applicability and extremely high editing quality, thus achieving high-quality editing results and meeting the actual needs of users. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] In order to more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for use in the description of the embodiments or the prior art. Obviously, the drawings in the following description are only the embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained according to the provided drawings.

[0048] Figure 1 It is a schematic flowchart of an image editing and processing method shown in the embodiments of the present invention;

[0049] Figure 2 It is an example diagram of mask cropping and image cropping shown in the embodiments of the present invention;

[0050] Figure 3 It is an example diagram of the diffusion process (adding noise) shown in the embodiments of the present invention;

[0051] Figure 4 It is an example diagram of the reverse diffusion process (denoising) shown in the embodiments of the present invention;

[0052] Figure 5 It is a diagram showing the image editing effect demonstrated in the embodiments of the present invention;

[0053] Figure 6 It is a visualization example diagram of the image editing process shown in the embodiments of the present invention;

[0054] Figure 7 It is a schematic structural diagram of an image editing processing device shown in the embodiments of the present invention;

[0055] Figure 8 It is a processing schematic diagram of each module shown in the embodiments of the present invention. Detailed implementation manners

[0056] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0057] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of this application and the above-mentioned drawings are used to distinguish similar objects, and do not have to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments described here can be implemented in an order other than that shown or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0058] It should be noted that the descriptions involving "first", "second", etc. in the present invention are only for descriptive purposes, and cannot be understood as indicating or implying their relative importance or implicitly indicating the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include at least one of such features. In addition, the technical solutions between various embodiments can be combined with each other, but it must be based on the fact that those of ordinary skill in the art can implement it. When the combination of technical solutions results in contradictions or cannot be implemented, it should be considered that such a combination of technical solutions does not exist and is not within the protection scope required by the present invention.

[0059] In this application, the term "comprise", "include" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.

[0060] See Figure 1 , which shows a schematic flow diagram of an image editing and processing method according to an embodiment of the present invention. The method includes:

[0061] Step S101: Obtain an editing request input by a user, where the editing request includes an initial image, a text instruction, and editing data;

[0062] In the process of specifically implementing step S101, in response to the user's input operation, obtain the image input by the user and use it as the initial image; then, display the initial image in a visualization interface so that the user can select the coordinates of the object to be segmented, that is, the click coordinates of the object to be edited, and obtain the text instruction input by the user.

[0063] It should be noted that the number of click coordinates of the object to be edited can be one or more.

[0064] The click coordinates include two sets of coordinates: a positive feedback coordinate set and a negative feedback coordinate set; the positive feedback coordinate set is a set of one or more coordinate points, indicating that the object to be edited includes this coordinate point; the negative feedback coordinate set is a set of zero or more coordinate points, indicating that the object to be edited does not include this coordinate point.

[0065] The text instruction includes the editing intention input by the user and the text description of the object to be edited.

[0066] The editing data includes the click coordinates of the object to be edited.

[0067] Step S102: Perform segmentation processing on the initial image and the click coordinates of the object to be edited in the editing data to obtain a target mask image.

[0068] It should be noted that the process of specifically implementing step S102 includes the following steps:

[0069] Step S11: Call a pre-trained image segmentation model to process the initial image according to the click coordinates of the object to be edited in the editing request to obtain an initial mask image.

[0070] Among them, the pre-trained image segmentation model is trained based on a sample set.

[0071] It should be noted that the pre-training process of the specific image segmentation model includes:

[0072] Collecting different historical images and the click coordinates of different objects selected by the user;

[0073] Training the model based on the click coordinates of different objects and historical images, and using the trained model to test the test set to obtain a training mask; if the training mask meets the expectations, it is determined that the trained image segmentation model is obtained; if not, re-training is continued using the click coordinates of different objects and historical images until the output training mask meets the expectations.

[0074] It should be noted that the test set refers to some historical images and the click coordinates of different objects during training.

[0075] The expectations are set in advance by technicians according to historical images and the click coordinates of different objects.

[0076] Step S12: Respond to an adjustment request corresponding to the initial mask, and adjust the initial mask based on the adjustment request to obtain a target mask.

[0077] In the process of specifically implementing step S12, the initial mask is displayed to the user, and the user is allowed to manually modify and adjust the mask to obtain an object mask consistent with the user's expectations; respond to the click coordinates of the mask adjusted by the user to crop the image to obtain a target mask, and use it as the first input.

[0078] Step S103: Determine a cropped image according to the target mask and the initial image.

[0079] It should be noted that in the process of specifically implementing step S103, the following steps are included:

[0080] Step S21: Calculate the position coordinates of the target mask in the initial image and use them as the cropping area.

[0081] In the process of specifically implementing step S103, calculate the adjusted initial mask, that is, the position coordinates x, y and width and height of the target mask in the initial image, and take N times the maximum value of the width and height as the width and height of the cropping area; the center point of the target mask is used as the center point of the cropping area, and the center point of the cropping area and the width and height of the cropping area are used to draw the cropping area on the initial image.

[0082] It should be noted that N is set in advance by technicians according to multiple experiments and can be specifically set as a positive number.

[0083] Step S22: Crop the image corresponding to the cropping area from the initial image to obtain a cropped image.

[0084] In the process of specifically implementing step S22, crop the sanctioned cropping area drawn on the initial image to obtain a cropped image corresponding to the object to be edited, and use it as the second input to reduce the size of the redrawn image area and improve the editing speed. Retain the starting coordinate points and width and height data of the cropping area for subsequent processing after editing.

[0085] It should be noted that in the process of specifically implementing steps S102 and S103, it can be illustrated by Figure 2 for explanation.

[0086] First, call the pre-trained image segmentation model to segment and mask the initial image according to the selected coordinates of the object to be edited in the editing request to obtain an initial mask image; then, respond to the adjustment request corresponding to the initial mask image to expand the cropping area corresponding to the initial mask image to obtain a target mask image, and perform mask cropping according to the target mask image;

[0087] Next, crop the image corresponding to the cropping area from the initial image according to the target mask image to obtain a cropped image.

[0088] Step S104: Describe the style of the initial image to obtain a corresponding style description text.

[0089] Since the text instruction does not contain descriptions related to the image style, and most non-professional users cannot accurately describe the image style in detail; therefore, in the process of specifically implementing step S104, input the initial image into a preset multimodal large model so that the multimodal large model can reverse-infer the style-related descriptions of the initial image, such as art category, color, brightness, depth of field, etc., to obtain a style description text, and use it as the third input.

[0090] It should be noted that the style description text is only used to constrain the style of the edited image and does not contain text descriptions unrelated to the style, such as scenes and characters; the length of the style description text is not limited, but to ensure the degree of detail of the description, it is usually set to be more than 256 characters.

[0091] The pre-set multi-modal large model uses multiple historical images and style-related descriptive texts corresponding to each image, such as art category, color, brightness, depth of field, etc. as the sample set; and divides the sample set into a training set and a test set, uses the training set to train the model to obtain a multi-modal large model; then, uses the test set to test the multi-modal large model to obtain a test result. If the test result is consistent with the style-related descriptive text in the test set, it is determined that the training of the multi-modal large model is completed; if not, re-train the model based on the training set until the test result is consistent with the style-related descriptive text in the test set.

[0092] Step S105: Process the text instruction to obtain a content description text and a task trigger word.

[0093] It should be noted that in the specific implementation process of step S105, the following steps are included:

[0094] Step S31: Extract instructions from the text instruction to obtain a content description text corresponding to the scene of the initial image.

[0095] It should be noted that in the specific implementation process of step S31, the following steps are included.

[0096] Step S41: Use a pre-set large language model to remove the text irrelevant to the content description in the text instruction to obtain a target text instruction;

[0097] Since the text instruction usually contains text irrelevant to the editing target, directly inputting it as editing information into the image diffusion model will reduce the generation quality; in the specific implementation process of step S41, use a large language model to filter the input text instruction, remove the text irrelevant to the content description, and only retain the content description text, that is, obtain the target text instruction.

[0098] Step S42: Use a pre-set large language model to expand and enhance the target text instruction to determine a content description text corresponding to the scene of the initial image;

[0099] In the specific implementation process of step S42, use a pre-set large language model to expand and enhance the target text instruction to obtain a content description text that is more detailed and fits the scene of the initial image, and use the content description text as the fourth input.

[0100] To better understand the content shown in step S41 and step S42, the following is an example for illustration.

[0101] For example, using a pre-set large language model for the target text instruction "I want to change the girl's top into overalls", after instruction extraction, we get "A girl wearing overalls", and then expand it to "A girl wearing blue overalls with rose patterns on them".

[0102] It should be noted that the preset large language model here is trained based on multiple historical text instructions, the initial images corresponding to the text instructions, and the content description texts corresponding to the scenarios.

[0103] The specific training process is as follows: multiple historical text instructions, the corresponding content description texts, and the initial images corresponding to the text instructions are used as a sample set; the sample set is divided to obtain a training set and a test set; the large language model is trained using the training set, and the large language model is tested using the test set. If the output target text instruction is consistent with the corresponding content description text in the test set, and the output description text is consistent with the description text corresponding to the scenario in the test set, a trained large language model is determined; if any one is inconsistent, the model is continuously trained using the training set until a trained large language model is obtained.

[0104] Step S32: Perform intent recognition on the text instruction to obtain an editing task category.

[0105] It should be noted that text instructions are unpredictable, and a single rigid task assignment strategy is difficult to adapt to complex and variable user inputs. Therefore, in addition to the above training, the preset large language model also needs to be trained on text instructions and the corresponding editing task categories. Specifically, extract the key features in the text instruction and construct the corresponding relationship between different key features and editing task categories.

[0106] Among them, the editing task categories include but are not limited to categories such as adding objects, replacing objects, removing objects, and detail enhancement.

[0107] Step S33: Determine the task trigger word corresponding to the editing task category.

[0108] It should be noted that the editing task category and the task trigger word do not need to be exactly the same, only a one-to-one mapping is required, that is, a one-to-one mapping relationship between the editing task category and the task trigger word is set in advance.

[0109] In the specific process of implementing step S32, using the large language model, perform intent recognition on the text instruction input by the user, and intelligently analyze categories such as enhancement that the user wants to perform; specifically, the large language model extracts the key features of the text instruction, searches for the editing task category corresponding to the key features, and then uses the task trigger word corresponding to the editing task category as the next input for assigning the editing task, which can effectively improve the applicability of the editing method, and use the task category as the fifth input.

[0110] For example, the text instruction is "I want to change the girl's top into overalls". The large language model identifies the key feature of this text instruction as "change into" through intention recognition, looks up the editing task category corresponding to this key feature "change into" as the replacement object, and then determines the task trigger word corresponding to the editing task category as "replacement object".

[0111] Among them, the task category name and the trigger word name do not need to be exactly the same, and only need to be mapped one by one.

[0112] Step S106: Process based on the cropped image, target mask image, style description text, content description text, and task trigger word to obtain a cropped result image.

[0113] It should be noted that in the process of specifically implementing step S106, the following steps are included:

[0114] Step S51: Trigger the corresponding editing task based on the task trigger word to load the corresponding image diffusion model and the corresponding pre-trained editing task LoRA;

[0115] It should be noted that the corresponding relationship between different task trigger words and editing tasks is set in advance, and each editing task includes an image diffusion model and the corresponding pre-trained editing task LoRA.

[0116] (Low-Rank Adaptation of Large Language Models, LoRA) is a low-rank adaptation technology for fine-tuning large models. It realizes the fine-tuning of the model by training low-rank matrices and then injecting these parameters into the original model, while maintaining the original capabilities of the pre-trained model.

[0117] The construction of the image diffusion model mainly includes the forward diffusion process and the reverse denoising process. The forward diffusion process is to gradually add noise to the original image until the image becomes pure noise; the reverse denoising process is to start from pure noise and gradually remove the noise to restore the image. The specific model needs to learn parameters through a neural network so that the predicted noise is as close as possible to the actual noise, that is, to obtain a trained image diffusion model.

[0118] The number of image diffusion models is multiple, and each image diffusion model corresponds to a task trigger word.

[0119] Step S52: Adjust the model parameters of the image diffusion model based on the pre-trained editing task LoRA to obtain an adjusted image diffusion model.

[0120] In the process of specifically implementing step S52, the pre-trained editing task LoRA fine-tunes the model by training its own low-rank matrix and then injecting these parameters into the image diffusion model, resulting in an adjusted image diffusion model.

[0121] Step S53: Encode the cropped image using the image encoder in the adjusted image diffusion model, and use the encoded cropped image as an image latent variable.

[0122] In the process of specifically implementing step S53, the built-in variational autoencoder of the diffusion model is used to encode the second input to obtain a cropped image encoding, which is used as an image latent variable.

[0123] Step S54: Use the adjusted image diffusion model to add noise to the image latent variable within the mask region corresponding to the target mask map, resulting in a noise-added image latent variable.

[0124] In the process of specifically implementing step S54, the image diffusion model determines the mask region corresponding to the cropped mask map in the image latent variable and adds noise within this mask region. As Figure 3 shown, the image within the mask region becomes a noise image after adding noise, that is, a noise-added image latent variable is obtained.

[0125] Among them, the noise distribution conforms to a Gaussian distribution.

[0126] Step S55: Encode the merged style description text and content description text using the image encoder in the adjusted image diffusion model to obtain a text latent variable.

[0127] In the process of specifically implementing step S55, the style description text and content description text are merged, and the built-in text encoder of the image diffusion model is used to encode the merged style description text and content description text to obtain a text latent variable.

[0128] Step S56: Process the merged text latent variable and the noise-added image latent variable based on the adjusted image diffusion model to obtain a cropped result image.

[0129] In the process of specifically implementing step S56, the noise-added image latent variable and the text latent variable are merged in a preset latent space to obtain a latent variable; then the merged latent variable undergoes a reverse diffusion process to gradually remove noise to generate a denoised image latent variable, and the built-in image decoder of the image diffusion model is used to decode the denoised image latent variable into a cropped result image.

[0130] It should be noted that the process of gradually denoising the merged latent variables through the reverse diffusion process to generate the latent variables of the denoised image can be as follows Figure 4 shown.

[0131] Step S107: Perform image processing on the cropped result image, the target mask image, and the initial image to obtain the target image.

[0132] It should be noted that the specific implementation process of step S107 includes the following steps:

[0133] Step S61: Expand the mask area within the target mask image and blur the expanded part to obtain a blurred mask image;

[0134] In the specific implementation process of step S61, expand the mask area within the target mask image by M pixels, and then blur the part of the mask area expanded by M pixels to obtain a blurred mask image;

[0135] It should be noted that M is set in advance by technicians according to multiple experiments or experience.

[0136] The range of mask expansion is not limited and can change with the size of the mask.

[0137] Step S62: Merge the cropped result image corresponding to the mask area in the blurred mask image into the initial image to obtain the target image.

[0138] In the specific implementation process of step S62, determine the mask area corresponding to the mask area in the blurred mask image in the cropped result image, and merge the mask area into the corresponding mask position of the initial image to obtain the target image.

[0139] For the initial image, that is, the input image Q, different file instructions will output different target images, as Figure 5 shown.

[0140] Suppose the text instruction 1 is to add a desk and a schoolbag. After the processing of the above steps S101 to S107, the corresponding target image output is Q1; suppose the text instruction 2 is to help me change the boy's coat into overalls. After the processing of the above steps S101 to S107, the corresponding target image output is Q2; suppose the text instruction 3 is to remove the girl. After the processing of the above steps S101 to S107, the corresponding target image output is Q3; suppose the text instruction 4 is detail enhancement. After the processing of the above steps S101 to S107, the corresponding target image output is Q4.

[0141] To better understand the image editing process shown in the above embodiments of the present invention, as Figure 6 shown, the following is described with an example.

[0142] Obtain an editing request input by the user, where the editing request includes an initial image a, a clicked coordinate, and a text instruction. The text instruction is "Help me change the boy's top into overalls."

[0143] Call a pre-trained image segmentation model to process the initial image a according to the clicked coordinate of the object to be edited in the editing request. That is, after the processing of step S102, obtain a target mask image a1 and use it as the first input. Crop an image corresponding to the cropping area of the target mask image a1 from the initial image. That is, after the processing of step S103, obtain a cropped image a2 and use it as the second input.

[0144] Use a preset multi-modal large model to perform risk backpropagation on the initial image. That is, after the processing of step S104, obtain a style description text "This painting is in the style of Makoto Shinkai's anime with a simple background" and use it as the third input.

[0145] Use a preset large language model to extract instructions from the text instruction, obtaining a content description text "A boy wearing overalls in a classroom". That is, after the processing of step S105, and use it as the fourth input.

[0146] Use the preset large language model to perform intention recognition on the text instruction, and then determine the corresponding task trigger word as "replace object". That is, after the processing of step S105, and use it as the fifth input.

[0147] Allocate tasks for the first input, the second input, the third input, the fourth input, and the fifth input, and perform image redrawing through the corresponding image diffusion model. That is, after the processing of step S106, obtain a cropped result image a3.

[0148] Finally, perform image mask merging on the target mask image a1, the cropped result image a3, and the initial image a. That is, after the processing of step S107, output a result image, that is, the target image.

[0149] In the embodiments of the present invention, first, an image segmentation model is used to segment a target mask image corresponding to an editing request, and then a cropped image corresponding to the target mask image is cropped from the initial image; then, a multi-modal large model describes the style of the initial image, and a large language model is used to determine a content description text and a task trigger word corresponding to the file instruction; then, an image diffusion model is used to process the cropped image, the target mask image, the style description text, the content description text, and the task trigger word obtained from the above processing to obtain a cropped result image; then, according to the target mask image and the starting coordinate points and width and height data of the cropped area, the cropped result image is merged with the original image to obtain a result image with the same resolution as the original image. The present invention combines an image segmentation model, a multi-modal large model, a large language model, and an image diffusion model, thereby realizing the precise positioning of the object to be edited under the editing request, maintaining the overall consistency of the image style before and after editing, having wide applicability and extremely high editing quality, thus realizing high-quality editing results and meeting the actual needs of users. Furthermore, it is widely applicable to different image editing tasks such as modifying the appearance, structure, and content of images, and stably generates high-quality editing results.

[0150] Based on the image editing method shown in the above embodiments of the present invention, correspondingly, the embodiments of the present invention also correspondingly show a schematic flowchart of an image editing device, as Figure 7 shown, the device includes:

[0151] An input module 701, configured to obtain an editing request input by a user, where the editing request includes an initial image, a text instruction, and editing data;

[0152] A mask generation processing module 702, configured to perform segmentation processing on the initial image and the selected coordinate points of the object to be edited in the editing data to obtain a target mask image; determine a cropped image according to the target mask image and the initial image;

[0153] A style reverse inference module 703, configured to describe the style of the initial image to obtain a corresponding style description text;

[0154] A text processing module 704, configured to process the text instruction to obtain a content description text and a task trigger word;

[0155] An image redrawing module 705, configured to perform processing based on the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain a cropped result image;

[0156] A post-processing module 706, configured to perform image processing according to the cropped result image, the target mask image, and the initial image to obtain a target image.

[0157] The specific principles and execution processes of each unit in the image editing and processing device disclosed in the above embodiments of the present invention are the same as the corresponding content in the image editing and processing method provided in the above embodiments of the present invention. For details, please refer to the corresponding parts in the image editing and processing method disclosed in the above embodiments of the present invention, and will not be elaborated here.

[0158] The present invention combines an image segmentation model, a multi-modal large model, a large language model, and an image diffusion model, thereby achieving precise positioning of the object to be edited under an editing request, maintaining the overall consistency of the image style before and after editing, having wide applicability and extremely high editing quality, thus achieving high-quality editing results and meeting the actual needs of users. Furthermore, it is widely applicable to different image editing tasks such as modifying the appearance, structure, and content of images, and stably generates high-quality editing results.

[0159] Optionally, based on the image editing and processing device shown in the above embodiments of the present invention, the mask generation processing module 702 for performing segmentation processing on the point selection coordinates of the object to be edited in the initial image and the editing data is specifically configured to:

[0160] Call a pre-trained image segmentation model to process the initial image according to the point selection coordinates of the object to be edited in the editing request, and obtain an initial mask image. The pre-trained image segmentation model is trained based on a sample set;

[0161] Respond to an adjustment request corresponding to the initial mask image, and adjust the initial mask image based on the adjustment request to obtain a target mask image.

[0162] Optionally, based on the image editing and processing device shown in the above embodiments of the present invention, the mask generation processing module 702 for determining a cropped image according to the target mask image and the initial image is specifically configured to:

[0163] Calculate the position coordinates of the target mask image in the initial image, and use them as the cropping area;

[0164] Crop an image corresponding to the cropping area from the initial image to obtain a cropped image.

[0165] Optionally, based on the image editing and processing device shown in the above embodiments of the present invention, the text processing module 704 includes an instruction extraction module and an intention recognition module;

[0166] The instruction extraction module is configured to extract instructions from the text instruction to obtain a content description text corresponding to the scene of the initial image;

[0167] The intention recognition module is configured to recognize the intention of the text instruction to obtain an editing task category; and determine a task trigger word corresponding to the editing task category.

[0168] Optionally, based on the image editing processing device shown in the above embodiments of the present invention, the image redrawing module 705 is specifically configured to:

[0169] Use a preset large language model to remove the text unrelated to the content description in the text instruction to obtain a target text instruction, where the preset large language model is trained based on multiple historical text instructions;

[0170] Expand and enhance the target text instruction to determine the content description text corresponding to the scene of the initial image.

[0171] Optionally, based on the image editing processing device shown in the above embodiments of the present invention, the post-processing module 706 is specifically configured to:

[0172] Process the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain a cropped result image, including:

[0173] Trigger a corresponding editing task based on the task trigger word to load a corresponding image diffusion model and a corresponding pre-trained editing task;

[0174] Adjust the model parameters of the image diffusion model based on the pre-trained editing task to obtain an adjusted image diffusion model;

[0175] Use the image encoder in the adjusted image diffusion model to encode the cropped image, and use the encoded cropped image as an image latent variable;

[0176] Use the adjusted image diffusion model to add noise to the masked area corresponding to the image latent variable according to the target mask image to obtain an image latent variable with added noise;

[0177] Use the image encoder in the adjusted image diffusion model to encode the merged style description text and content description text to obtain a text latent variable;

[0178] Process the merged text latent variable and the image latent variable with added noise based on the adjusted image diffusion model to obtain a cropped result image.

[0179] Among them, performing image processing on the cropped result image, the target mask image, and the initial image to obtain a target image includes:

[0180] Expand the masked area in the target mask image and perform blurring processing on the expanded part to obtain a blurred mask image;

[0181] Merge the cropped result image corresponding to the masked area in the blurred mask image into the original image to obtain the target image.

[0182] Optionally, based on the processes of implementing image editing and processing by each module shown in the embodiments of the present invention above, it can be as Figure 8 shown.

[0183] An embodiment of the present application provides an electronic device, which includes a processor and a memory. The memory is used to store the program code and data for image editing and processing, and the processor is used to call the program instructions in the memory to execute the steps shown in the image editing and processing method in the above embodiments.

[0184] An embodiment of the present invention provides a storage medium, which includes the electronic device provided in the embodiment of the present application above, and this electronic device is used to execute the image editing and processing method disclosed in the embodiment of the present application.

[0185] Each embodiment in this specification is described in a progressive manner. The same or similar parts among the embodiments can be referred to each other, and the key points of each embodiment are the differences from other embodiments. In particular, for a system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. A person of ordinary skill in the art can understand and implement it without creative work.

[0186] Those skilled in the art can further realize that the units and algorithm steps of each example described in combination with the embodiments disclosed in this article can be implemented by electronic hardware, computer software, or a combination of the two. To clearly illustrate the interchangeability of hardware and software, the composition and steps of each example have been generally described according to functions in the above description. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of the present invention.

[0187] The foregoing description of the disclosed embodiments enables those skilled in the art to practice or use the present invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Thus, the present invention is not intended to be limited to the embodiments shown herein but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. An image editing and processing method, characterized in that, The method includes: Obtaining an editing request input by a user, where the editing request includes an initial image, a text instruction, and editing data; Performing segmentation processing on the selected coordinates of the object to be edited in the initial image and the editing data to obtain a target mask image; Determining a cropped image based on the target mask image and the initial image; Describing the style of the initial image to obtain a corresponding style description text; Processing the text instruction to obtain a content description text and a task trigger word; Performing processing based on the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain a cropped result image; Performing image processing based on the cropped result image, the target mask image, and the initial image to obtain a target image.

2. The method according to claim 1, wherein Performing segmentation processing on the selected coordinates of the object to be edited in the initial image and the editing data to obtain a target mask image, including: Invoking a pre-trained image segmentation model to process the initial image according to the selected coordinates of the object to be edited in the editing request to obtain an initial mask image, where the pre-trained image segmentation model is trained based on a sample set; Responding to an adjustment request corresponding to the initial mask image to adjust the initial mask image based on the adjustment request to obtain a target mask image.

3. The method according to claim 1, characterized in that, Determining a cropped image based on the target mask image and the initial image, including: Calculating the position coordinates of the target mask image in the initial image and using them as a cropping region; Cropping an image corresponding to the cropping region from the initial image to obtain a cropped image.

4. The method according to claim 1, wherein Processing the text instruction to obtain a content description text and a task trigger word, including: Performing instruction extraction on the text instruction to obtain a content description text corresponding to the scene of the initial image; Performing intention recognition on the text instruction to obtain an editing task category; Determining a task trigger word corresponding to the editing task category.

5. The method according to claim 4, wherein Performing instruction extraction on the text instruction to obtain a content description text corresponding to the scene of the initial image, including: Using a preset large language model to remove the text irrelevant to the content description in the text instruction to obtain a target text instruction, where the preset large language model is trained based on multiple historical text instructions, the initial images corresponding to the text instructions, and the content description texts corresponding to the scenes; Performing expansion and enhancement on the target text instruction to determine a content description text corresponding to the scene of the initial image.

6. The method according to claim 1, characterized in that, Performing processing based on the cropped image, the target mask image, the style description text, the content description text, and the task trigger word to obtain a cropped result image, including: Triggering a corresponding editing task based on the task trigger word to load a corresponding image diffusion model and a corresponding pre-trained editing task; Adjusting the model parameters of the image diffusion model based on the pre-trained editing task to obtain an adjusted image diffusion model; Encoding the cropped image using the image encoder in the adjusted image diffusion model and using the encoded cropped image as an image latent variable; Adding noise to the masked area corresponding to the image latent variable according to the target mask map by using the adjusted image diffusion model to obtain the image latent variable after adding noise; Encoding the merged style description text and content description text by using the image encoder in the adjusted image diffusion model to obtain a text latent variable; Processing the merged text latent variable and the image latent variable after adding noise based on the adjusted image diffusion model to obtain a cropped result image.

7. The method according to claim 1, characterized in that, Performing image processing on the cropped result image, the target mask map, and the initial image to obtain a target image, including: Expanding the masked area in the target mask map and blurring the expanded part to obtain a blurred mask map; Merging the cropped result image corresponding to the masked area in the blurred mask map into the initial image to obtain a target image.

8. An image editing and processing device, characterized in that, The device includes: An input module, configured to obtain an editing request input by a user, where the editing request includes an initial image, a text instruction, and editing data; A mask generation and processing module, configured to perform segmentation processing on the initial image and the selected coordinates of the object to be edited in the editing data to obtain a target mask map; determining a cropped image according to the target mask map and the initial image; A style inversion module, configured to perform style description on the initial image to obtain a corresponding style description text; A text processing module, configured to process the text instruction to obtain a content description text and a task trigger word; An image redrawing module, configured to perform processing based on the cropped image, the target mask map, the style description text, the content description text, and the task trigger word to obtain a cropped result image; A post-processing module, configured to perform image processing on the cropped result image, the target mask map, and the initial image to obtain a target image.

9. An electronic device, characterized in that, The electronic device includes a processor and a memory, where the memory is used to store program codes and data generated by the data, and the processor is used to call the program instructions in the memory to execute the image editing processing method according to any one of claims 1-7.

10. A storage medium, characterized in that, The storage medium includes a stored program, where, when the program runs, it controls the device where the storage medium is located to execute the image editing processing method according to any one of claims 1-7.

Citation Information

Patent Citations

  • Image processing method and device, computer readable storage medium and electronic equipment

    CN116485944A

  • Image editing method and device, equipment, storage medium and program product

    CN117611709A

  • Image editing and generating method and device, electronic equipment and storage medium

    CN119205982A

  • Image editing method and device, storage medium and electronic equipment

    CN119399327A

  • Shielding object moving and editing method and system based on diffusion model

    CN119810263A

Cited By

  • Image editing method and related device

    CN120931769A