Model training method, image editing method, device, equipment, medium and product
By acquiring and evaluating information after image editing and updating the image editing model, the problems of poor quality and unnaturalness in image editing are solved, and better editing results are achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-09-27
- Publication Date
- 2026-03-31
AI Technical Summary
Existing technologies are insufficient to effectively address defects such as poor image quality and unnatural appearance in image editing, thus affecting the editing results.
By acquiring evaluation information from the original image, edited descriptive text, and the edited image, the image editing model is used for processing, and the model is updated based on the evaluation information to improve image editing performance.
The performance of the image editing model has been improved, enabling better image editing results and overcoming the interference of defects in the edited images in the training data.
Smart Images

Figure CN121767769A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of image processing technology, and in particular to a model training method, an image editing method, an apparatus, a device, a medium, and a product. Background Technology
[0002] In some scenarios, such as image retouching or instruction-based image editing, the following requirements exist: to edit an existing image according to the user's needs, such as adjusting the color of the hat in the image to blue, to obtain the edited image.
[0003] However, how to achieve the above editing process has become a technical problem that urgently needs to be solved. Summary of the Invention
[0004] This application provides a model training method, an image editing method, an apparatus, a device, a medium, and a product, which are beneficial for improving image editing results.
[0005] To achieve the above objectives, the technical solution provided in this application is as follows:
[0006] This application provides a model training method, the method comprising: acquiring an original image, an edit description text corresponding to the original image, an edited image corresponding to the original image under the edit description text, and evaluation information of the edited image, wherein the evaluation information is used to describe the state of the edited image under at least one evaluation item; processing the original image, the edit description text, and the evaluation information using an image editing model to obtain an image editing result corresponding to the original image; and updating the image editing model based on the difference between the image editing result and the edited image.
[0007] In one possible implementation, the at least one evaluation item includes one or more of a text following evaluation item, an image preservation evaluation item, and an image quality evaluation item; the text following evaluation item describes the matching state between the edited image and the edited description text on a first content, the first content being determined based on at least one editing instruction described by the edited description text; the image preservation evaluation item describes the matching state between the edited image and the original image on a second content, the second content being determined based on other content in the original image besides the first content; the image quality evaluation item describes the quality change of the edited image relative to the original image.
[0008] In one possible implementation, the evaluation information includes scores for each of the evaluation items, and / or, the evaluation information includes defect description text for some or all of the evaluation items; for any one of the evaluation items, the score of the evaluation item is used to characterize the level achieved by the edited image under that evaluation item, and the defect description text of the evaluation item is used to describe the defects present in the edited image under that evaluation item.
[0009] In one possible implementation, the at least one evaluation item includes one or more of a text following evaluation item, an image preservation evaluation item, and an image quality evaluation item; the score of the text following evaluation item describes the similarity between the content changes of the edited image relative to the original image and the content changes described by the edit description text; the score of the image preservation evaluation item describes the similarity between the content preserved by the edited image relative to the original image and other content in the original image besides the edited content specified by the edit description text; the score of the image quality evaluation item describes the quality changes of the edited image relative to the original image; the defect description text of the text following evaluation item describes at least one following error of the edited image relative to the edit description text; the defect description text of the image preservation evaluation item describes at least one preservation error of the edited image relative to the original image; and the defect description text of the image quality evaluation item describes at least one quality defect of the edited image relative to the original image.
[0010] In one possible implementation, the evaluation information is introduced as conditional information into the denoising network of the image editing model.
[0011] In one possible implementation, for any of the cross-attention layers in the denoising network, the input data of the data fusion layer corresponding to the cross-attention layer is determined based on the evaluation information and the output data of the cross-attention layer, and the input data of the cross-attention layer includes the encoding result of the edited description text.
[0012] In one possible implementation, the input data of the data fusion layer is determined based on the feature vector of the evaluation information; the evaluation information includes at least one type of data, and the feature vector of the evaluation information is determined based on the vectorization results of each type of data and the vectorization results of each type.
[0013] In one possible implementation, the data fusion layer is used to: perform residual calculation or sum calculation based on the mapping result of the feature vector of the cross-attention layer and the evaluation information in a first feature space, wherein the first feature space is the feature space to which the output data of the cross-attention layer belongs.
[0014] In one possible implementation, the input data of the denoising network in the image editing model includes first data and second data. The first data is input to the denoising network as conditional information. The input data of the first network layer in the denoising network includes the second data. The first data is different from the second data. The second data is determined based on the original image and the evaluation information.
[0015] In one possible implementation, the process of determining the second data includes: concatenating the image features of the original image with the noise-added result of the image features of the edited image to obtain a concatenated result; performing convolution processing on the concatenated result to obtain a convolution result; and performing at least one cross-attention processing based on the mapping result of the convolution result and the feature vector of the evaluation information in a second feature space to obtain the second data, wherein the second feature space is the feature space to which the convolution result belongs.
[0016] In one possible implementation, the image editing model includes a first module, a second module, a third module, a linear layer corresponding to the third module, a denoising network, a linear layer corresponding to the denoising network, and a decoding module; the first module is used to obtain the splicing result between the image features of the original image and the denoised result of the image features of the edited image; the second module is used to obtain the feature vector of the evaluation information; the linear layer corresponding to the third module and the linear layer corresponding to the denoising network are both used to process the feature vector of the evaluation information; the third module is used to perform at least one cross-attention process based on the splicing result and the output data of the linear layer corresponding to the third module; the denoising network is used to perform denoising processing based on the output data of the third module, the encoding result of the edit description text, and the output data of the linear layer corresponding to the denoising network; the decoding module is used to decode the output data of the denoising network to obtain the image editing result.
[0017] In one possible implementation, updating the image editing model includes: updating a portion of the modules in the image editing model, wherein the portion of the modules includes some or all of the network layers in the second module, the third module, the linear layer corresponding to the third module, the denoising network, and the linear layer corresponding to the denoising network.
[0018] This application provides an image editing method, the method comprising: acquiring a target image, an editing description text corresponding to the target image, and preset constraint information; processing the target image, the editing description text, and the preset constraint information using an image editing model to obtain an image editing result corresponding to the target image, wherein the preset constraint information is used to describe the state of the image editing result under at least one evaluation item, and the image editing model is determined using the model training method provided in this application.
[0019] This application provides a model training apparatus, comprising: a first acquisition unit, configured to acquire an original image, an edit description text corresponding to the original image, an edited image corresponding to the original image under the edit description text, and evaluation information of the edited image, wherein the evaluation information is used to describe the state of the edited image under at least one evaluation item; a first processing unit, configured to process the original image, the edit description text, and the evaluation information using an image editing model to obtain an image editing result corresponding to the original image; and a model updating unit, configured to update the image editing model based on the difference between the image editing result and the edited image.
[0020] This application provides an image editing apparatus, comprising: a second acquisition unit for acquiring a target image, an editing description text corresponding to the target image, and preset constraint information; and a second processing unit for processing the target image, the editing description text, and the preset constraint information using an image editing model to obtain an image editing result corresponding to the target image, wherein the preset constraint information describes the state of the image editing result under at least one evaluation item, and the image editing model is determined using the model training method provided in this application.
[0021] This application provides an electronic device, the device comprising: a processor and a memory; the memory for storing instructions or computer programs; the processor for executing the instructions or computer programs in the memory, so that the electronic device performs the model training method or image editing method provided in this application.
[0022] This application provides a computer-readable medium storing instructions or computer programs that, when executed on a device, cause the device to perform the model training method or image editing method provided in this application.
[0023] This application provides a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the model training method or image editing method provided in this application.
[0024] Compared with related technologies, this application has at least the following advantages:
[0025] The technical solution provided in this application first obtains an original image, an edit description text corresponding to the original image, an edited image corresponding to the original image under the edit description text, and evaluation information of the edited image. This evaluation information is used to describe the state of the edited image in at least one evaluation item, thereby enabling the evaluation information to indicate the deficiencies of the edited image and to indicate the required level of image editing processing for the original image. Then, an image editing model is used to process the original image, the edit description text, and the evaluation information to obtain the image editing result corresponding to the original image. Finally, based on the difference between the image editing result and the edited image, the image editing model is updated to give the updated model better image editing performance, thereby enabling the image editing processing based on the model to achieve better results.
[0026] The image editing model performs image editing based on the aforementioned evaluation information. This allows the model to accurately determine the required level of image editing for the original image and the shortcomings of the resulting image. This enables the model to learn how to perform image editing while meeting the constraints described by the evaluation information. Furthermore, it allows the model to learn how to flexibly control the state of its output image in at least one evaluation term. This effectively overcomes interference caused by defects in the edited images within the training data, thereby improving the model's performance and ultimately achieving better results in image editing. Attached Figure Description
[0027] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0028] Figure 1 A flowchart illustrating a model training method provided in this application embodiment;
[0029] Figure 2 A schematic diagram of a diffusion model provided in an embodiment of this application;
[0030] Figure 3This is a schematic diagram of another diffusion model provided in an embodiment of this application;
[0031] Figure 4 This is a schematic diagram of another diffusion model provided in an embodiment of this application;
[0032] Figure 5 This is a schematic diagram of another diffusion model provided in an embodiment of this application;
[0033] Figure 6 A flowchart illustrating an image editing method provided in this application embodiment;
[0034] Figure 7 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;
[0035] Figure 8 This is a schematic diagram of the structure of an image editing device provided in an embodiment of this application;
[0036] Figure 9 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0037] Research has revealed that some image editing solutions involve: first, collecting and constructing training data, such that the training data includes a triple of {original image, editing instructions text, and edited image}; then, using this training data, training a machine learning model to obtain a model with image editing capabilities, so that the model can subsequently perform image editing based on the original image and the editing instructions text.
[0038] The study also found that the implementation scheme shown above has the following defects: because the edited image is generated in a certain way, it may have certain defects, such as poor quality or unnaturalness. As a result, the model trained using the edited image as ground truth information has relatively poor performance, which in turn affects the editing effect.
[0039] Based on the above two studies, in order to better improve the editing effect, this application provides a model training method. The method includes: firstly, acquiring an original image, an editing description text corresponding to the original image, an edited image corresponding to the original image under the editing description text, and evaluation information of the edited image, so that the evaluation information is used to describe the state of the edited image in at least one evaluation item, thereby enabling the evaluation information to represent the deficiencies of the edited image and to represent the level that the image editing processing for the original image needs to achieve; then, using an image editing model to process the original image, the editing description text, and the evaluation information to obtain the image editing result corresponding to the original image; then, based on the difference between the image editing result and the edited image, updating the image editing model so that the updated model has better image editing performance, thereby enabling the image editing processing based on the model to achieve better results.
[0040] The image editing model performs image editing based on the aforementioned evaluation information. This allows the model to accurately determine the required level of image editing for the original image and the shortcomings of the resulting image. This enables the model to learn how to perform image editing while meeting the constraints described by the evaluation information. Furthermore, it allows the model to learn how to flexibly control the state of its output image in at least one evaluation term. This effectively overcomes interference caused by defects in the edited images within the training data, thereby improving the model's performance and ultimately achieving better results in image editing.
[0041] Furthermore, this application does not limit the executing entity of the model training method provided in the embodiments of this application. For example, the model training method provided in the embodiments of this application can be applied to a terminal device or a server. Alternatively, the model training method provided in the embodiments of this application can also be implemented through the data interaction process between the terminal device and the server. The terminal device can be a smartphone, computer, personal digital assistant (PDA), tablet computer, etc. The server can be a standalone server, a cluster server, or a cloud server.
[0042] To enable those skilled in the art to better understand the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present application, and not all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present application.
[0043] To better understand the technical solution provided in this application, the model training method provided in this application will be explained below with reference to some accompanying figures. For example... Figure 1 As shown, the model training method provided in this application includes S101-S103 as described below. Wherein, the... Figure 1 This is a flowchart of a model training method provided in an embodiment of this application.
[0044] S101: Obtain the original image, the edit description text corresponding to the original image, the edited image corresponding to the original image under the edit description text, and the evaluation information of the edited image, wherein the evaluation information is used to describe the state of the edited image under at least one evaluation item.
[0045] Here, the original image refers to an image in the training data that requires image editing processing, such as... Figures 2-5 The original image shown in any one of these examples is used to provide the image editing process with content other than what needs to be edited, so that the result obtained by the image editing process is as consistent as possible with the original image in terms of that other content.
[0046] It should be noted that this application does not limit the implementation method of maintaining consistency. For example, in some scenarios, such as those with high accuracy requirements, maintaining consistency means being completely identical, so that the actual meaning of the statement "Information 1 and Information 2 are consistent" is: Information 1 and Information 2 are completely identical. Conversely, in some scenarios, such as those with low accuracy requirements, maintaining consistency means having a high degree of similarity, so that the actual meaning of the statement "Information 1 and Information 2 are consistent" is: the similarity between Information 1 and Information 2 is higher than a preset similarity threshold.
[0047] It should also be noted that, for the training data involved in the model training process provided in this application, the training data is implemented using a quadruple of {original image, edit description text corresponding to the original image, edited image corresponding to the original image under the edit description text, and evaluation information of the edited image}, so that the training data can not only represent the ground truth image value required for image editing processing of the original image, but also represent the level that the image editing processing of the original image needs to achieve. This can effectively overcome the adverse effects caused by the defects in the ground truth image, thereby improving the model performance.
[0048] It should also be noted that, for all the training data involved in the model training process provided in this application, this application does not limit the characteristics of these training data. For example, they can at least satisfy the following constraints: all the training data involved in the model training process should traverse various situations as much as possible (e.g., various states under various evaluation items) to ensure that the model obtained based on these training data can learn as much as possible how to flexibly control the state of its output image under at least one evaluation item, which is beneficial to improving image editing performance.
[0049] Furthermore, this application does not limit the process of acquiring the original image; for example, it can be implemented using any existing or future image acquisition method.
[0050] The edit description text corresponding to the original image refers to the text existing in the training data that needs to be referenced when performing image editing processing on that original image, such as... Figures 2-5 The text of any of the edit instructions shown is such that it describes what kind of edits to be performed on the original image.
[0051] Therefore, in one possible implementation, the edit description text corresponding to the original image can be used to describe at least one editing instruction for the original image. As an example, when the edit description text includes the content "adjust the color of the hat to blue", the edit description text can at least describe the target instruction, and the target instruction is used to indicate the editing requirement of "adjusting the color of the hat to blue".
[0052] Furthermore, this application does not limit the process of obtaining the edited description text corresponding to the original image. For example, it can be implemented using any existing or future method of obtaining text, such as a method where the user manually provides the text or a method where the text is automatically generated using certain machine learning models.
[0053] The edited image corresponding to the original image under the edit description text refers to the image present in the training data used as ground truth information. This edited image represents the actual result of image editing the original image according to the edit description text, thus enabling it to guide the image editing results of the original image under the edit description text during model training. Figures 2-5 The image editing result shown in any one of the options enables the edited image to guide the model in learning how to perform image editing processes.
[0054] Furthermore, this application does not limit the process of obtaining the edited image as shown in the preceding paragraph. For example, it can be implemented using any existing or future method for obtaining edited images. Alternatively, it can be implemented using manual annotation. Furthermore, the process of obtaining the edited image can specifically involve a machine learning model with image generation capabilities performing image generation processing based on the original image and the corresponding edit description text to obtain the edited image. It should be noted that this application does not limit the implementation method of the machine learning model.
[0055] Furthermore, for the edited image corresponding to the original image under the edit description text, the evaluation information of the edited image is obtained by performing evaluation processing on the edited image, so that the evaluation information is used to describe the state of the edited image under at least one evaluation item, thereby enabling the evaluation information to represent the quality level of the edited image, and further enabling the evaluation information to represent the deficiencies of the edited image, such as poor quality, etc. In this way, the evaluation information can represent the level that the image editing processing of the original image needs to achieve, so that subsequent image editing processing of the original image can be constrained based on the evaluation information.
[0056] Furthermore, this application does not limit the implementation of at least one of the above evaluation items. For example, it can be implemented using any existing or future index that can characterize the state of an image, such as quality, naturalness, light distribution, etc.
[0057] Research has found that the desired goals of image editing of an original image include at least the following: the result of the image editing process should, as far as possible, meet the editing requirements described in the corresponding editing description text of the original image, and the result of the image editing process should be as consistent as possible with the original image in terms of content other than the content that needs to be edited. Therefore, in order to improve the image editing effect, these goals can be used as evaluation items to measure the quality of an image.
[0058] Based on the above research, to improve the results, at least one of the evaluation items mentioned above can include one or more of the following evaluation item, the preserving evaluation item, and the quality evaluation item, so that these evaluation items can be used to better measure the quality of an image. For ease of understanding, these three evaluation items are introduced below.
[0059] For the text-following evaluation item, it is used to evaluate whether the edited image can meet the content editing requirements described by the edit description text, so that the text-following evaluation item can indicate whether the edited image can accurately and comprehensively cover the editing instructions described by the edit description text, and thus the text-following evaluation item can indicate the matching state between the edited image and the edit description text. In this way, the text-following evaluation item is used to describe the matching state of the edited image and the edit description text in the first content.
[0060] The first content refers to the content that needs to be considered when evaluating text-following evaluation items; and the first content may be determined based on at least one editing instruction described by the editing description text, so that the first content can represent the content involved by these editing instructions, such as the content of "hat color", thereby enabling the first content to include the editing content specified by the editing description text, and thus enabling the first content to represent the content that needs to be edited.
[0061] For the image preservation evaluation item, it is used to evaluate whether the edited image can meet the content preservation requirements of the original image. This allows the image preservation evaluation item to indicate whether the edited image can accurately and comprehensively retain all content in the original image except for the edited content specified by the edit description text corresponding to the original image. In this way, the image preservation evaluation item can indicate the matching status between the edited image and the original image, thus enabling the image preservation evaluation item to describe the matching status between the edited image and the original image in terms of the second content.
[0062] The second content refers to the content that needs to be considered when evaluating the image preservation evaluation item; and the second content is determined based on the content in the original image other than the first content, so that the first content can represent the content that exists in the original image but is not covered by the corresponding edit description text of the original image, so that the second content can include the difference between the set of content described by the original image and the set of content described by the edit description text, thereby enabling the second content to represent the content that does not need to be edited, that is, the content that needs to be preserved.
[0063] The image quality assessment item is used to evaluate the image quality of the edited image. Moreover, this application does not limit the implementation of the image quality assessment item. For example, in some scenarios, the image quality assessment item of the edited image can be implemented using any method that can evaluate the quality of an image, so that the evaluation result of the image quality assessment item of the edited image can represent the quality level achieved by the edited image itself.
[0064] Research has found that in some scenarios, the low quality of the original image itself leads to a low quality of the edited image generated from it. For example, if the original image contains hand deformities such as extra fingers, then based on the principle of content preservation, the edited image obtained by editing the original image will theoretically still have this problem, thus affecting the quality of the edited image.
[0065] Based on the above research, it is known that the quality of the original image itself will affect the quality of the edited image. Therefore, in order to better overcome the interference caused by this effect, this application provides a possible implementation of the image quality evaluation item. In this way, the image quality evaluation item can be used to evaluate the relative difference in quality between the edited image and the original image, so that the image quality evaluation item can be used to describe the quality change of the edited image relative to the original image, thereby enabling the image quality evaluation item to more accurately represent the level of quality achieved by the image editing process.
[0066] Furthermore, this application does not limit the implementation of the above evaluation information. For example, when the evaluation information is used to describe the state of the edited image under at least one evaluation item, the evaluation information may include the scores of each evaluation item so that the evaluation information can better represent the level achieved by the edited image under each evaluation item. Specifically, for any evaluation item, the score of that evaluation item is used to characterize the level achieved by the edited image under that evaluation item.
[0067] For example, when the aforementioned evaluation information is used to describe the state of the edited image under at least one evaluation item, the evaluation information may include defect description text for some or all of the evaluation items, so that the evaluation information can better represent the defects of the edited image under these evaluation items, such as content 1 not following accurately, content 2 not being maintained accurately, hand deformities, etc. Specifically, for any evaluation item, the defect description text for that evaluation item is used to describe the defects of the edited image under that evaluation item.
[0068] For example, to improve the effectiveness, when the above evaluation information is used to describe the state of the image after the above editing under at least one evaluation item, the evaluation information may include the scores of each evaluation item, as well as the defect description text of some or all of the evaluation items.
[0069] Based on the aforementioned evaluation information, in one possible implementation, the evaluation information may include: the score of the text following evaluation item, the score of the image preservation evaluation item, the score of the image quality evaluation item, the defect description text of the text following evaluation item, the defect description text of the image preservation evaluation item, and part or all of the defect description text of the image quality evaluation item.
[0070] Regarding the "score for the text-following evaluation item" mentioned above, this score describes the similarity between the content changes of the edited image relative to the original image and the content changes described in the edit description text. This allows the "score for the text-following evaluation item" to represent the degree to which the edited image follows the content of the edit description text, thus indicating to some extent whether the edited image accurately and comprehensively meets the content editing requirements described in the edit description text, and consequently, the level of image editing in terms of follow-up. Specifically, the "content changes of the edited image relative to the original image" describes the differences between the edited image and the original image. The "content changes described in the edit description text" refers to the changes described in the edit description text.
[0071] Regarding the "image preservation evaluation score" mentioned above, this score describes the similarity between the content preserved in the edited image relative to the original image and other content in the original image besides the editable content specified by the edit description text. This allows the "image preservation evaluation score" to represent the degree of content preservation of the edited image relative to the original image, thus indicating to some extent whether the edited image accurately and comprehensively retains the content of the original image that does not require editing. Consequently, the "image preservation evaluation score" can represent the level of preservation achieved by the image editing process. Specifically, the "content preserved in the edited image relative to the original image" describes the content that is identical between the edited image and the original image. The "editable content specified by the edit description text" refers to the content described in the edit description text that requires editing.
[0072] Regarding the "image quality assessment score" mentioned above, this "image quality assessment score" is used to describe the quality changes (such as quality fluctuations) of the edited image relative to the original image, so that the "image quality assessment score" can represent the difference between the quality of the edited image and the quality of the original image, thereby enabling the "image quality assessment score" to represent the level of quality achieved by the image editing process.
[0073] Regarding the "defect description text of the text following the evaluation item" mentioned above, this "defect description text of the text following the evaluation item" is used to describe at least one following error that exists in the edited image relative to the editing description text, such as the error of "incorrectly changing the hat color to red". This is so that the "defect description text of the text following the evaluation item" can represent the content in the edited image that does not achieve the editing goal described by the editing description text, such as content that has been edited incorrectly, content that was forgotten to be edited, etc., thereby enabling the "defect description text of the text following the evaluation item" to represent the defects that occur in the image editing process in terms of following.
[0074] Regarding the "defect description text of the image preservation evaluation item" mentioned above, this "defect description text of the image preservation evaluation item" is used to describe at least one preservation error that exists in the edited image relative to the original image, such as the error of "not preserving skin color". This is so that the "defect description text of the image preservation evaluation item" can represent additional changes in the content changes of the edited image relative to the original image that are not part of the content changes described by the editing description text. Thus, the "defect description text of the image preservation evaluation item" can represent the content in the edited image that does not achieve the preservation target described by the original image, and further, the "defect description text of the image preservation evaluation item" can represent the defects that occur in the preservation aspect of the image editing process.
[0075] Regarding the "defect description text of image quality assessment item" mentioned above, this "defect description text of image quality assessment item" is used to describe at least one quality defect in the edited image relative to the original image, such as unnatural lighting transitions, so that the "defect description text of image quality assessment item" can represent the influencing factors in the edited image that cause the quality of the edited image to be inferior to that of the original image, thereby enabling the "defect description text of image quality assessment item" to represent the defects in image editing processing in terms of quality.
[0076] Furthermore, this application does not limit the method of obtaining the above evaluation information. For example, it can use any machine learning model with image evaluation performance, so that the machine learning model can evaluate the original image and the corresponding edit description text of the original image, and obtain and output the evaluation information of the edited image, so that the evaluation information can indicate the state of the edited image under at least one evaluation item, such as score and / or defect description text.
[0077] Based on the relevant content of S101 above, in some scenarios, after obtaining the original image, the corresponding edit description text, and the edited image corresponding to the original image under the edit description text, the edited image can be evaluated based on the original image and the edit instruction text to obtain the evaluation information of the edited image. This evaluation information can represent the quality of the edited image, thereby supplementing the edited image with some features. This allows the model to be trained using the training data {original image, edit description text, edited image, evaluation information}. This enables the model to comprehensively influence the training process using the edited image and its evaluation information, helping the model to identify and understand the deficiencies of the edited image in the training data, thus improving the model's performance.
[0078] S102: The image editing model is used to process the evaluation information of the original image, the editing description text corresponding to the original image, and the edited image corresponding to the original image under the editing description text, so as to obtain the image editing result corresponding to the original image.
[0079] The image editing model is used to perform image editing processing based on the input data of the image editing model, such as... Figure 3 or Figure 5 The image editing process shown.
[0080] Furthermore, this application does not limit the implementation method of the image editing model described above. For example, the image editing model can be based on any diffusion model, such as Instructpix2pix. Figure 2 The diffusion model 1 shown or Figure 4 The diffusion model 3 shown is constructed so that the image editing model includes all network layers in the diffusion model, thereby enabling the image editing model to have image generation capabilities, so that the image editing model can subsequently perform image editing processing by means of image generation.
[0081] As can be seen, in one possible implementation, the image editing model described above can at least include a denoising network, such as... Figure 3 or Figure 5 The denoising network shown is used to enable subsequent image generation processing.
[0082] In addition, in order to better improve the image editing effect, the above image editing model can at least satisfy the following constraints: the above evaluation information is introduced as conditional information into the denoising network of the image editing model so that the denoising network can perform denoising processing under the constraint of the evaluation information.
[0083] As can be seen, in one possible implementation, the working principle of the above image editing model may include at least the following: introducing the editing description text corresponding to the original image and the above evaluation information as conditional information into the denoising network of the image editing model, so that the image editing model can perform image editing processing on the original image under the constraints of these two conditions, thereby making the image output by the image editing model satisfy these two conditions as much as possible.
[0084] Furthermore, this application does not limit the way the conditional information of the above-mentioned edited descriptive text is introduced. For example, it can be implemented using any existing or future method of introducing textual conditions into the denoising network, such as cross-attention.
[0085] It can be seen that, in one possible implementation, the above image editing model can at least satisfy the following constraints: the conditional information of the above editing description text is introduced into the denoising network of the image editing model through cross-attention, so that the denoising network can perform denoising processing under the constraint of the editing description text.
[0086] Furthermore, this application does not limit the way the conditional information of the above evaluation information is introduced. For example, it can be implemented using any existing or future method that can introduce conditional information into the denoising network, such as cross-attention.
[0087] Research has shown that, in order to maximize the impact of the aforementioned evaluation information on image editing, a direct data fusion approach can be adopted to introduce this evaluation information—a conditional piece of information—into the denoising network of the image editing model.
[0088] It can be seen that, in one possible implementation, the above image editing model can at least satisfy the following constraints: the image editing model includes a denoising network, and for any cross-attention layer in the denoising network, the input data of the data fusion layer corresponding to the cross-attention layer is determined based on the evaluation information and the output data of the cross-attention layer, and the input data of the cross-attention layer includes the encoding result of the above editing description text.
[0089] For any cross-attention layer in the denoising network, this layer is used to introduce conditional information related to the preceding editable text. Specifically, this layer performs cross-attention processing on the encoded result of the editable text and the output data of the preceding network layer. The preceding network layer refers to a network layer in the denoising network that is adjacent to the cross-attention layer but ranks higher than it. The encoded result of the editable text is obtained through text encoding processing to better represent the information carried by the editable text.
[0090] It should be noted that this application does not limit the implementation method of text encoding processing. For example, it can be implemented using any existing or future editor with text encoding function, such as a text encoder based on a Contrastive Language-Image Pre-training (CLIP) model.
[0091] Furthermore, for any cross-attention layer in the denoising network, the corresponding data fusion layer refers to a network layer in the denoising network that is adjacent to the cross-attention layer but positioned later than it, such as a residual layer or... Figure 5 The denoising network shown in the diagram has a network layer marked with a "+", and the data fusion layer can be specifically used to: fuse the output data of the cross-attention layer with the evaluation information above, so that the data fusion layer can realize the process of introducing conditional information for the evaluation information.
[0092] Furthermore, this application does not limit the implementation of the above data fusion layer. For example, for any cross-attention layer in the denoising network, the data fusion layer corresponding to the cross-attention layer can at least satisfy the following constraint: the input data of the data fusion layer is determined based on the feature vector of the above evaluation information.
[0093] The feature vector of the evaluation information above is used to better represent the content described by the evaluation information; and this application does not limit the method of determining the feature vector.
[0094] Furthermore, to improve the effectiveness, when the evaluation information includes at least one type of data, the feature vector of the evaluation information is determined based on the vectorization results of each type of data and the vectorization results of each type, so that the feature vector can better represent the content described by the evaluation information. For ease of understanding, an example is provided below.
[0095] As an example, when the above evaluation information includes the scores of the text following evaluation item, the scores of the image preservation evaluation item, the scores of the image quality evaluation item, the defect description text of the text following evaluation item, the defect description text of the image preservation evaluation item, and the defect description text of the image quality evaluation item, the process of determining the feature vector of the evaluation information may include steps 11-15 below.
[0096] Step 11: Vectorize the scores of the text following evaluation item, the image preservation evaluation item, and the image quality evaluation item to obtain the vectorized results of the following score, the preservation score, and the quality score, so that the vectorized results of the following score can better represent the score of the text following evaluation item, the vectorized results of the preservation score can better represent the score of the image preservation evaluation item, and the vectorized results of the quality score can better represent the score of the image quality evaluation item.
[0097] It should be noted that this application does not limit the implementation method of numerical vectorization processing. For example, it can be implemented using any existing or future method that can map numerical values to numerical embeddings, such as position encoding.
[0098] Step 12: Perform text vectorization on the defect description text of the text following the evaluation item, the defect description text of the image preservation evaluation item, and the defect description text of the image quality evaluation item, respectively, to obtain the text vectorization results of the following defects, the preservation defects, and the quality defects, so that the text vectorization results of the following defects can better represent the defects described by the defect description text of the text following the evaluation item, the text vectorization results of the preservation defects can better represent the defects described by the defect description text of the image preservation evaluation item, and the text vectorization results of the quality defects can better represent the defects described by the defect description text of the image quality evaluation item.
[0099] It should be noted that this application does not limit the implementation method of text vectorization processing. For example, it can be implemented using any existing or future method that can map text to text embedding, such as any text encoding process.
[0100] It should also be noted that this application does not limit the execution time of step 12 above, but only needs to ensure that the execution time of step 12 is earlier than the execution time of step 14 below.
[0101] Step 13: Map the vectorized results of the following score, the maintaining score, and the quality score to the target feature space to obtain the mapping results corresponding to the following score, the maintaining score, and the quality score. This ensures that the mapping results corresponding to the following score represent the state of the text following evaluation item's score in the target feature space, the maintaining score represent the state of the image maintaining evaluation item's score in the target feature space, and the quality score represent the state of the image quality evaluation item's score in the target feature space. This helps avoid defects caused by inconsistencies in the feature space. Here, the target feature space refers to the feature space where the above text vectorized results are located.
[0102] It should be noted that this application does not limit the implementation of step 13 above. For example, it can be implemented using any method that can map data in one feature space to another feature space, such as a linear layer or a multilayer perceptron (MLP).
[0103] Step 14: Concatenate the text vectorization results of the following defects, the text vectorization results of the maintaining defects, the text vectorization results of the quality defects, the mapping results corresponding to the following scores, the mapping results corresponding to the maintaining scores, and the mapping results corresponding to the quality scores to obtain the first concatenation result.
[0104] It should be noted that this application does not limit the implementation of step 14 above. For example, it can be implemented using any existing or future method with splicing function, such as concat network layer.
[0105] Step 15: Perform residual calculation or sum calculation on the data type vectorization result of the above evaluation information and the first concatenation result to obtain the feature vector of the above evaluation information. The data type vectorization result includes the vectorization results of each type involved in the evaluation information, so that the type vectorization result can describe the type to which different parts of the first concatenation result belong, such as numerical or text types, so that the type vectorization result can be used to distinguish different parts of the first concatenation result.
[0106] Based on the relevant content of steps 11 to 15 above, for the evaluation information mentioned above, when the evaluation information includes numerical data (e.g., scores for text following evaluation items, scores for image preservation evaluation items, and scores for image quality evaluation items) and text data (e.g., defect description text for text following evaluation items, defect description text for image preservation evaluation items, and defect description text for image quality evaluation items), the numerical data can first be vectorized to obtain the vectorized result of the numerical data, and the text data can be vectorized (e.g., text encoding processing) to obtain the vectorized result of the text data; then the numerical data can be vectorized... The data vectorization result of the text type is mapped to the feature space of the data vectorization result of the text type to obtain the data mapping result corresponding to the numerical type, so that the data mapping result and the data vectorization result of the text type belong to the same feature space; then, the data vectorization result of the text type and the data mapping result corresponding to the numerical type are concatenated to obtain the concatenated result; finally, the concatenated result and the data vectorization result of the data type of the evaluation information are subjected to residual calculation or sum calculation to obtain the feature vector of the evaluation information, so that the feature vector can better represent the content described by the evaluation information, thereby making the image editing processing based on the evaluation information better.
[0107] Furthermore, to better enhance the impact of the evaluation information, this application also provides an implementation of the aforementioned data fusion layer. In this implementation, for any cross-attention layer in the denoising network, the corresponding data fusion layer can at least satisfy the following constraints: the data fusion layer is used to perform residual calculation or sum calculation based on the mapping result of the output data of the cross-attention layer and the feature vector of the evaluation information in a first feature space. The first feature space is the feature space to which the output data of the cross-attention layer belongs. This allows for more effective utilization of the content described by the evaluation information, thereby improving model performance. The mapping result represents the state of the evaluation information in the first feature space; moreover, this application does not limit the method of determining the mapping result. For example, the mapping result can be obtained by a linear layer, such as... Figure 5 The linear layer 2 shown is obtained by processing the feature vector of this evaluation information.
[0108] Research has revealed that the input data of a denoising network consists of two parts. One part is used as conditional information so that it can only affect the processing of certain network layers in the denoising network. However, the other part includes noise information so that it can be used as the object of processing, thereby enabling it to affect the processing of all network layers in the denoising network.
[0109] Based on the above research, this application provides a possible implementation of the image editing model described above. In this implementation, the image editing model can at least satisfy the following constraints: the input data of the denoising network in the image editing model includes first data and second data. The first data is input to the denoising network as conditional information. The input data of the first network layer in the denoising network includes the second data. The first data is different from the second data. The second data is determined based on the original image and the evaluation information described above.
[0110] Here, the first data refers to the portion of the input data of the denoising network that serves as conditional information. Furthermore, this application does not limit the first data; for example, the first data can be determined based on the edit description text corresponding to the original image, enabling the denoising network to perform denoising processing under the constraint of the edit description text. Alternatively, to further improve the effect, the first data can also be determined based on the aforementioned evaluation information, enabling the denoising network to perform denoising processing under the constraints of both the edit description text and the evaluation information. Therefore, in one possible implementation, the first data may include the encoding result of the aforementioned edit description text and the mapping result of the feature vector of the aforementioned evaluation information in the first feature space.
[0111] The second data refers to the portion of the input data of the denoising network that affects all network layers in the denoising network; and this second data is determined based on the original image and the evaluation information above.
[0112] In addition, this application does not limit the method of determining the second data mentioned above. For example, it may specifically include steps 21-23 below.
[0113] Step 21: Combine the image features of the original image with the noise-added image features of the edited image to obtain the combined result.
[0114] The image features of the original image refer to those obtained by performing image feature extraction on the original image, so that the image features can represent the information carried by the original image.
[0115] It should be noted that this application does not limit the implementation method of image feature extraction processing. For example, it can be implemented using any existing or future image feature extraction method, such as using a variational auto (VAE) encoder.
[0116] Image features of an edited image refer to those obtained by performing image feature extraction processing on the edited image, so that the image features can represent the information carried by the edited image.
[0117] The noise addition result of the image features of the edited image is obtained by performing noise addition processing on the image features of the edited image. It should be noted that this application does not limit the implementation method of the noise addition processing. For example, it can be implemented using any noise addition method involved in any existing or future diffusion model.
[0118] Based on the relevant content of step 21 above, for some scenarios, after obtaining the original image and the edited image, the VAE encoder can be used to process the two images separately to obtain the image features of the original image and the image features of the edited image; then, noise processing is performed on the image features of the edited image to obtain the noise-added result of the image features of the edited image; then, the image features of the original image and the noise-added result are spliced together to obtain the spliced result.
[0119] Step 22: Perform convolution processing on the above splicing result to reduce the number of channels in the splicing result from 8 to 4, and obtain the convolution result.
[0120] Step 23: Perform at least one cross-attention process based on the mapping result of the feature vectors of the convolution result and the evaluation information in the second feature space to obtain the second data. The second feature space is the feature space to which the convolution result belongs. The mapping result represents the state of the evaluation information in the second feature space; moreover, this application does not limit the method of determining the mapping result. For example, the mapping result can be obtained by a linear layer, such as... Figure 5 The linear layer 1 shown is obtained by processing the feature vector of this evaluation information.
[0121] Based on the relevant content of steps 21 to 23 above, it is known that for some scenarios, after obtaining the training data {original image, the corresponding edit description text of the original image, the edited image corresponding to the original image under the edit description text, and the evaluation information of the edited image}, the noise addition result (such as the image features of the original image and the image features of the edited image) can be used. Figure 5 Using the noise data shown and the feature vector of the evaluation information, the second data required as input to the denoising network in the image editing model is determined, so that the denoising network can subsequently perform denoising processing on the second data under the constraint of the first data required as input to the denoising network.
[0122] In addition, to better improve the image editing effect, the image editing model mentioned above may include a first module, a second module, a third module, a linear layer corresponding to the third module, a denoising network, a linear layer corresponding to the denoising network, and a decoding module.
[0123] Regarding the first module mentioned above, this first module is used to process the first part of the input data of the image editing model, such as the original image, so that the first module can obtain the splicing result between the image features of the original image and the noise-added result of the image features of the edited image. Furthermore, this application does not limit the implementation of the first module. For example, the first module can be used to: first obtain the image features of the original image and the noise-added result of the image features of the edited image; then splice the image features of the original image and the noise-added result of the image features of the edited image to obtain the splicing result.
[0124] Regarding the second module above, such as Figure 5 As shown in the Reward module, the second module is used to process the second part of the input data of the image editing model, such as evaluation information; and this application does not limit the implementation of the second module. For example, the second module can be used to perform the process of determining the feature vector of the evaluation information, as shown in steps 11 to 15 above.
[0125] For the linear layer corresponding to the third module above, such as Figure 5 As shown in the linear layer 1, this linear layer is used to map the output data of the second module above to another feature space (such as the second feature space above or the feature space to which the output data of the first module above belongs), so as to achieve feature space unification. It should be noted that this application does not limit the implementation of this linear layer. For example, it can be implemented using a fully connected layer.
[0126] Regarding the third module above, such as Figure 5 In the processing module shown, the third module is used to perform at least one cross-attention processing based on the output data of the first module (or the convolution processing result of the output data of the first module) and the output data of the linear layer corresponding to the third module. This enables the fusion of image information and evaluation information, so that the evaluation information can guide the denoising process according to the influence of the noise addition result, thereby allowing the influence of the evaluation information to run through the entire denoising process, which is beneficial to improving the denoising performance.
[0127] For the linear layer corresponding to the denoising network mentioned above, such as Figure 5 As shown in the linear layer 2, this linear layer is used to map the output data of the second module above (such as the feature vector of the evaluation information) to the feature space (such as the first feature space above) to which the output data of the cross-attention layer in the denoising network belongs, so as to achieve feature space unification. It should be noted that this application does not limit the implementation of this linear layer. For example, it can be implemented using a fully connected layer.
[0128] Regarding the denoising network mentioned above, such as Figure 5The denoising network shown is used to perform denoising processing based on the output data of the third module, the encoding result of the above-mentioned edit description text, and the output data of the corresponding linear layer of the denoising network, so as to achieve denoising processing of the output data of the third module under the constraints of the edit description text and the above-mentioned evaluation information.
[0129] Regarding the decoding module mentioned above, such as Figure 5 In the decoder shown, this decoding module is used to decode the output data of the denoising network to obtain the image editing result. It should be noted that this application does not limit the implementation of this decoding module; for example, it can be implemented using a VAE decoder.
[0130] Based on the content related to S102 above, for some scenarios, after obtaining the training data {original image, the corresponding edit description text of the original image, the edited image corresponding to the original image under the edit description text, and the evaluation information of the edited image}, an image editing model can be used, such as... Figure 3 The diffusion model 2 shown or Figure 5 The diffusion model 4 shown processes the original image, the edit description text, and the evaluation information to obtain the image editing result corresponding to the original image, so that the image editing result can represent the predicted state of the original image under the edit description text.
[0131] S103: Update the image editing model based on the difference between the image editing result corresponding to the original image and the edited image corresponding to the original image under the editing description text.
[0132] It should be noted that this application does not limit the implementation method of S103 above. For example, it can specifically be: first, based on the difference between the image editing result corresponding to the original image and the edited image corresponding to the original image under the edit description text, determine the model loss of the image editing model so that the model loss can represent the performance of the image editing model; then, update the image editing model based on the model loss. It should be noted that this application does not limit the calculation method of the model loss.
[0133] Research has shown that in some scenarios, some modules in the image editing model can be implemented directly using existing models. Therefore, in order to improve efficiency, this application also provides a possible implementation of S103 above. In this method, when the image editing model includes a first module, a second module, a third module, a linear layer corresponding to the third module, a denoising network, a linear layer corresponding to the denoising network, and a decoding module, S103 may specifically include: updating some modules in the image editing model based on the difference between the image editing result corresponding to the original image and the edited image corresponding to the original image under the edit description text, and the partial modules include some or all network layers in the second module, the third module, the linear layer corresponding to the third module, the denoising network, and the linear layer corresponding to the denoising network.
[0134] It should be noted that this application does not limit the implementation of "some or all of the network layers in the second module" mentioned above. For example, when the second module adopts... Figure 5 When the Reward module shown is implemented, the "part or all of the network layers in the second module" can be MLPs.
[0135] In addition, in order to improve the model performance, this application also provides a possible implementation of S103 above. In this way, S103 can specifically be: based on the difference between the image editing result corresponding to the original image and the edited image corresponding to the original image under the edit description text, update some or all modules in the image editing model, and return to continue executing S101 and its subsequent steps until the preset stopping condition is reached, and end the iterative training process for the image editing model.
[0136] The preset stopping condition refers to the condition required to end the iterative training process of the image editing model. This application does not limit the implementation of the preset stopping condition. For example, the preset stopping condition may include: the model loss of the image editing model is lower than a preset loss threshold. Alternatively, the preset stopping condition may include: the rate of change of the model loss of the image editing model is lower than a preset rate of change threshold. Yet another example is: the number of updates to the image editing model reaches a preset number of updates threshold.
[0137] Based on the relevant content in S101 to S103 above, the model training method provided in this application first obtains the original image, the corresponding editing description text, the edited image corresponding to the original image under the editing description text, and the evaluation information of the edited image. This evaluation information is used to describe the state of the edited image in at least one evaluation item, thereby enabling the evaluation information to represent the deficiencies of the edited image and to represent the level that the image editing processing for the original image needs to achieve. Then, the image editing model is used to process the original image, the editing description text, and the evaluation information to obtain the image editing result corresponding to the original image. Finally, based on the difference between the image editing result and the edited image, the image editing model is updated so that the updated model has better image editing performance, thereby enabling the image editing processing based on the model to achieve better results.
[0138] The image editing model performs image editing based on the aforementioned evaluation information. This allows the model to accurately determine the required level of image editing for the original image and the shortcomings of the resulting image. This enables the model to learn how to perform image editing while meeting the constraints described by the evaluation information. Furthermore, it allows the model to learn how to flexibly control the state of its output image in at least one evaluation term. This effectively overcomes interference caused by defects in the edited images within the training data, thereby improving the model's performance and ultimately achieving better results in image editing.
[0139] Based on the above-mentioned image editing model, this application also provides an image editing method, such as... Figure 6 As shown, the image editing method includes S601-S602 as described below. Among them, the... Figure 6 A flowchart illustrating an image editing method provided in an embodiment of this application.
[0140] S601: Obtain the target image, the corresponding edit description text, and preset constraint information.
[0141] The target image refers to an image provided by the user that requires image editing processing; and this application does not limit the method of obtaining the target image.
[0142] The edit description text corresponding to the target image refers to the text provided by the user that describes what kind of editing processing is performed on the target image; moreover, this application does not limit the method of obtaining the edit description text. It should be noted that the implementation method of the edit description text corresponding to the target image is similar to the implementation method of the edit description text corresponding to the original image mentioned above.
[0143] The preset constraint information is used to describe the level achieved by the image editing processing of the target image under at least one evaluation item, so that the preset constraint information represents the state of the image editing result obtained by the target image under the corresponding editing description text under at least one evaluation item.
[0144] Furthermore, this application does not limit the implementation of the preset constraint information mentioned above. For example, the implementation of the preset constraint information may be similar to the implementation of the evaluation information mentioned above.
[0145] Furthermore, this application does not limit the method of obtaining the aforementioned pre-defined constraint information. For ease of understanding, the following explanation will be provided in conjunction with two scenarios.
[0146] In scenario one, where high-quality image editing is required, the aforementioned pre-defined constraints can be implemented using the pre-set parameters {5, 5, 5, None, None, None} for the trained image editing model. This ensures that the image editing process is performed at the highest possible level, thereby optimizing the image editing result and improving the overall image editing performance. Here, the first 5 represents the maximum score for the text-following evaluation item; the second 5 represents the maximum score for the image-preserving evaluation item; the third 5 represents the maximum score for the image-quality evaluation item; the first None indicates that the image editing process has no defects in the text-following evaluation item; the second None indicates that the image editing process has no defects in the image-preserving evaluation item; and the third None indicates that the image editing process has no defects in the image-quality evaluation item.
[0147] In scenario two, where high flexibility in image editing is required, the aforementioned pre-defined constraint information can be implemented using editing level description information provided by the user for the target image. This editing level description information refers to user-specified information describing the level of image editing processing achieved under at least one evaluation criterion. This allows the editing level description information to describe the user's specified editing level requirements for the target image, ensuring that the image editing processing based on this information meets those requirements, thereby improving image editing flexibility.
[0148] S602: The target image, the edit description text, and the preset constraint information are processed using an image editing model to obtain the image editing result corresponding to the target image. The preset constraint information is used to describe the state of the image editing result under at least one evaluation item. The image editing model is determined using any implementation of the model training method provided in this application.
[0149] It should be noted that the implementation of S602 is similar to the implementation of S102 above, and will not be repeated here for the sake of brevity.
[0150] Based on the relevant content of S601 to S602 above, it can be seen that for the image editing method provided in this application, after obtaining the target image, the editing description text corresponding to the target image, and the preset constraint information, the image editing model is used to process these data to obtain the image editing result corresponding to the target image, so that the image editing result can satisfy the content editing constraints described by the editing description text, the editing level constraints described by the preset constraint information, and the content preservation constraints involved in the target image, which is beneficial to improving the image editing effect.
[0151] Furthermore, this application does not limit the executing entity of the image editing method provided in the embodiments of this application. For example, the image editing method provided in the embodiments of this application can be applied to a terminal device or a server. Alternatively, the image editing method provided in the embodiments of this application can also be implemented through the data interaction process between the terminal device and the server.
[0152] Based on the model training method provided in the embodiments of this application, the embodiments of this application also provide a model training device, which is described below in conjunction with... Figure 7 Explanation and clarification will be provided. Among them, Figure 7 This is a schematic diagram of a model training device provided in an embodiment of this application. It should be noted that for technical details of the model training device provided in this embodiment, please refer to the relevant content of the model training method above.
[0153] like Figure 7 As shown, the model training device 700 provided in this application embodiment includes:
[0154] The first acquisition unit 701 is used to acquire an original image, an edit description text corresponding to the original image, an edited image corresponding to the original image under the edit description text, and evaluation information of the edited image, wherein the evaluation information is used to describe the state of the edited image under at least one evaluation item;
[0155] The first processing unit 702 is used to process the original image, the edit description text, and the evaluation information using an image editing model to obtain the image editing result corresponding to the original image;
[0156] The model update unit 703 is used to update the image editing model based on the difference between the image editing result and the edited image.
[0157] In one possible implementation, the at least one evaluation item includes one or more of a text following evaluation item, an image preservation evaluation item, and an image quality evaluation item; the text following evaluation item describes the matching state between the edited image and the edited description text on a first content, the first content being determined based on at least one editing instruction described by the edited description text; the image preservation evaluation item describes the matching state between the edited image and the original image on a second content, the second content being determined based on other content in the original image besides the first content; the image quality evaluation item describes the quality change of the edited image relative to the original image.
[0158] In one possible implementation, the evaluation information includes scores for each of the evaluation items, and / or, the evaluation information includes defect description text for some or all of the evaluation items; for any one of the evaluation items, the score of the evaluation item is used to characterize the level achieved by the edited image under that evaluation item, and the defect description text of the evaluation item is used to describe the defects present in the edited image under that evaluation item.
[0159] In one possible implementation, the at least one evaluation item includes one or more of a text following evaluation item, an image preservation evaluation item, and an image quality evaluation item; the score of the text following evaluation item describes the similarity between the content changes of the edited image relative to the original image and the content changes described by the edit description text; the score of the image preservation evaluation item describes the similarity between the content preserved by the edited image relative to the original image and other content in the original image besides the edited content specified by the edit description text; the score of the image quality evaluation item describes the quality changes of the edited image relative to the original image; the defect description text of the text following evaluation item describes at least one following error of the edited image relative to the edit description text; the defect description text of the image preservation evaluation item describes at least one preservation error of the edited image relative to the original image; and the defect description text of the image quality evaluation item describes at least one quality defect of the edited image relative to the original image.
[0160] In one possible implementation, the evaluation information is introduced as conditional information into the denoising network of the image editing model.
[0161] In one possible implementation, for any of the cross-attention layers in the denoising network, the input data of the data fusion layer corresponding to the cross-attention layer is determined based on the evaluation information and the output data of the cross-attention layer, and the input data of the cross-attention layer includes the encoding result of the edited description text.
[0162] In one possible implementation, the input data of the data fusion layer is determined based on the feature vector of the evaluation information; the evaluation information includes at least one type of data, and the feature vector of the evaluation information is determined based on the vectorization results of each type of data and the vectorization results of each type.
[0163] In one possible implementation, the data fusion layer is used to: perform residual calculation or sum calculation based on the mapping result of the feature vector of the cross-attention layer and the evaluation information in a first feature space, wherein the first feature space is the feature space to which the output data of the cross-attention layer belongs.
[0164] In one possible implementation, the input data of the denoising network in the image editing model includes first data and second data. The first data is input to the denoising network as conditional information. The input data of the first network layer in the denoising network includes the second data. The first data is different from the second data. The second data is determined based on the original image and the evaluation information.
[0165] In one possible implementation, the process of determining the second data includes: concatenating the image features of the original image with the noise-added result of the image features of the edited image to obtain a concatenated result; performing convolution processing on the concatenated result to obtain a convolution result; and performing at least one cross-attention processing based on the mapping result of the convolution result and the feature vector of the evaluation information in a second feature space to obtain the second data, wherein the second feature space is the feature space to which the convolution result belongs.
[0166] In one possible implementation, the image editing model includes a first module, a second module, a third module, a linear layer corresponding to the third module, a denoising network, a linear layer corresponding to the denoising network, and a decoding module; the first module is used to obtain the splicing result between the image features of the original image and the denoised result of the image features of the edited image; the second module is used to obtain the feature vector of the evaluation information; the linear layer corresponding to the third module and the linear layer corresponding to the denoising network are both used to process the feature vector of the evaluation information; the third module is used to perform at least one cross-attention process based on the splicing result and the output data of the linear layer corresponding to the third module; the denoising network is used to perform denoising processing based on the output data of the third module, the encoding result of the edit description text, and the output data of the linear layer corresponding to the denoising network; the decoding module is used to decode the output data of the denoising network to obtain the image editing result.
[0167] In one possible implementation, the model update unit 703 is specifically used to: update a portion of the modules in the image editing model based on the difference between the image editing result and the edited image. The portion of the modules includes some or all of the network layers in the second module, the third module, the linear layer corresponding to the third module, the denoising network, and the linear layer corresponding to the denoising network.
[0168] Based on the aforementioned content of the model training device 700, the working principle of the model training device 700 provided in this application includes: firstly, acquiring an original image, an edit description text corresponding to the original image, an edited image corresponding to the original image under the edit description text, and evaluation information of the edited image, so that the evaluation information is used to describe the state of the edited image in at least one evaluation item, thereby enabling the evaluation information to represent the deficiencies of the edited image and to represent the level that the image editing processing for the original image needs to achieve; then, using an image editing model to process the original image, the edit description text, and the evaluation information to obtain the image editing result corresponding to the original image; then, based on the difference between the image editing result and the edited image, updating the image editing model so that the updated model has better image editing performance, thereby enabling the image editing processing based on the model to achieve better results. The image editing model performs image editing based on the aforementioned evaluation information. This allows the model to accurately determine the required level of image editing for the original image and the shortcomings of the resulting image. This enables the model to learn how to perform image editing while meeting the constraints described by the evaluation information. Furthermore, it allows the model to learn how to flexibly control the state of its output image in at least one evaluation term. This effectively overcomes interference caused by defects in the edited images within the training data, thereby improving the model's performance and ultimately achieving better results in image editing.
[0169] Based on the image editing method provided in the embodiments of this application, the embodiments of this application also provide an image editing device, which is described below in conjunction with... Figure 8 Explanation and clarification will be provided. Among them, Figure 8 This is a schematic diagram of an image editing device provided in an embodiment of this application. It should be noted that for technical details of the image editing device provided in this application embodiment, please refer to the relevant content of the image editing method above.
[0170] like Figure 8 As shown, the image editing device 800 provided in this application embodiment includes:
[0171] The second acquisition unit 801 is used to acquire the target image, the edit description text corresponding to the target image, and preset constraint information;
[0172] The second processing unit 802 is used to process the target image, the edit description text, and the preset constraint information using an image editing model to obtain an image editing result corresponding to the target image. The preset constraint information is used to describe the state of the image editing result under at least one evaluation item. The image editing model is determined using any embodiment of the model training method provided in this application.
[0173] Based on the above-mentioned content of the image editing device 800, it can be understood that the working principle of the image editing device 800 provided in this application is as follows: after acquiring the target image, the editing description text corresponding to the target image, and the preset constraint information, the image editing model is used to process these data to obtain the image editing result corresponding to the target image, so that the image editing result can meet the content editing constraints described by the editing description text, the editing level constraints described by the preset constraint information, and the content preservation constraints involved in the target image, which is beneficial to improving the image editing effect.
[0174] In addition, this application embodiment also provides an electronic device, the device including a processor and a memory: the memory is used to store instructions or computer programs; the processor is used to execute the instructions or computer programs in the memory, so that the electronic device performs any implementation of the model training method provided in this application embodiment, or performs any implementation of the image editing method provided in this application embodiment.
[0175] See Figure 9 This document illustrates a structural schematic diagram of an electronic device 900 suitable for implementing embodiments of the present disclosure. The terminal devices in these embodiments may include, but are not limited to, mobile terminals such as mobile phones, laptops, digital broadcast receivers, PDAs (personal digital assistants), PADs (tablet computers), PMPs (portable multimedia players), in-vehicle terminals (e.g., in-vehicle navigation terminals), and fixed terminals such as digital TVs and desktop computers. Figure 9 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments disclosed herein.
[0176] like Figure 9As shown, the electronic device 900 may include a processing unit (e.g., a central processing unit, a graphics processor, etc.) 901, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 902 or a program loaded from a storage device 908 into a random access memory (RAM) 903. The RAM 903 also stores various programs and data required for the operation of the electronic device 900. The processing unit 901, ROM 902, and RAM 903 are interconnected via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.
[0177] Typically, the following devices can be connected to I / O interface 905: input devices 906 including, for example, touchscreens, touchpads, keyboards, mice, cameras, microphones, accelerometers, gyroscopes, etc.; output devices 907 including, for example, liquid crystal displays (LCDs), speakers, vibrators, etc.; storage devices 908 including, for example, magnetic tapes, hard disks, etc.; and communication devices 909. Communication device 909 allows electronic device 900 to communicate wirelessly or wiredly with other devices to exchange data. Although Figure 9 An electronic device 900 with various devices is shown; however, it should be understood that it is not required to implement or possess all of the devices shown. More or fewer devices may be implemented or possessed alternatively.
[0178] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a non-transitory computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device 909, or installed from a storage device 908, or installed from a ROM 902. When the computer program is executed by a processing device 901, it performs the functions defined in the methods of embodiments of this disclosure.
[0179] The electronic device provided in this embodiment belongs to the same inventive concept as the method provided in the above embodiments. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0180] This application also provides a computer-readable medium storing instructions or computer programs that, when executed on a device, cause the device to perform any implementation of the model training method provided in this application, or any implementation of the image editing method provided in this application.
[0181] It should be noted that the computer-readable medium described in this disclosure can be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium can be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in connection with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium can include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable program code. Such propagated data signals can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. A computer-readable signal medium can be any computer-readable medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the computer-readable medium can be transmitted using any suitable medium, including but not limited to: wires, optical fibers, RF (radio frequency), etc., or any suitable combination thereof.
[0182] In some implementations, clients and servers can communicate using any currently known or future-developed network protocol such as HTTP (Hypertext Transfer Protocol) and can interconnect with digital data communication (e.g., communication networks) of any form or medium. Examples of communication networks include local area networks (“LANs”), wide area networks (“WANs”), the Internet (e.g., the Internet of Things), and end-to-end networks (e.g., ad hoc end-to-end networks), as well as any currently known or future-developed networks.
[0183] The aforementioned computer-readable medium may be included in the aforementioned electronic device; or it may exist independently and not assembled into the electronic device.
[0184] The aforementioned computer-readable medium carries one or more programs, which, when executed by the electronic device, enable the electronic device to perform the aforementioned methods.
[0185] Computer program code for performing the operations of this disclosure can be written in one or more programming languages or a combination thereof, including but not limited to object-oriented programming languages such as Java, Smalltalk, and C++, as well as conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a local area network (LAN) or a wide area network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).
[0186] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this disclosure. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.
[0187] The units described in the embodiments of this disclosure can be implemented in software or hardware. The names of the units / modules do not necessarily limit the specific unit itself.
[0188] The functions described above in this document can be performed, at least in part, by one or more hardware logic components. For example, exemplary types of hardware logic components that can be used, without limitation, include: Field Programmable Gate Arrays (FPGAs), Application-Specific Integrated Circuits (ASICs), Application Standard Products (ASSPs), System-on-Chip (SoCs), Complex Programmable Logic Devices (CPLDs), and so on.
[0189] In the context of this disclosure, a machine-readable medium can be a tangible medium that may contain or store a program for use by or in conjunction with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can be, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of machine-readable storage media include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0190] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
[0191] It should be understood that in this application, "at least one (item)" means one or more, and "more than" means two or more. "And / or" is used to describe the relationship between related objects, indicating that three relationships can exist. For example, "A and / or B" can represent three cases: only A exists, only B exists, and both A and B exist simultaneously, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one (item) of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one (item) of a, b, or c can represent: a, b, c, "a and b", "a and c", "b and c", or "a and b and c", where a, b, and c can be single or multiple.
[0192] It should also be noted that, in this document, relational terms such as "first" and "second" are used only to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.
[0193] The steps of the methods or algorithms described in conjunction with the embodiments disclosed herein can be implemented directly by hardware, a software module executed by a processor, or a combination of both. The software module can be located in random access memory (RAM), main memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disk, removable disk, CD-ROM, or any other form of storage medium known in the art.
[0194] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A model training method, characterized in that, The method comprises: obtaining an original image, an editing description text corresponding to the original image, an edited image corresponding to the original image under the editing description text, and evaluation information of the edited image, the evaluation information being used to describe a state of the edited image under at least one evaluation item; processing the original image, the editing description text and the evaluation information by using an image editing model to obtain an image editing result corresponding to the original image; updating the image editing model according to a difference between the image editing result and the edited image.
2. The method of claim 1, wherein, The at least one evaluation item comprises one or more of a text following evaluation item, an image keeping evaluation item and an image quality evaluation item; The text following evaluation item is used to describe a matching state of the edited image and the editing description text on a first content, the first content being determined according to at least one editing instruction described by the editing description text; The image keeping evaluation item is used to describe a matching state of the edited image and the original image on a second content, the second content being determined according to other content in the original image except the first content; The image quality evaluation item is used to describe a quality change of the edited image relative to the original image.
3. The method of claim 1, wherein, The evaluation information comprises scores of the evaluation items, and / or the evaluation information comprises defect description texts of part or all of the evaluation items; For any evaluation item, the score of the evaluation item is used to represent a level reached by the edited image under the evaluation item, and the defect description text of the evaluation item is used to describe a defect existing in the edited image under the evaluation item.
4. The method of claim 3, wherein, The at least one evaluation item comprises one or more of a text following evaluation item, an image keeping evaluation item and an image quality evaluation item; The score of the text following evaluation item is used to describe a similarity degree between a content change of the edited image relative to the original image and a content change described by the editing description text; The score of the image keeping evaluation item is used to describe a similarity degree between a kept content of the edited image relative to the original image and other content in the original image except an edited content specified by the editing description text; The score of the image quality evaluation item is used to describe a quality change existing in the edited image relative to the original image; The defect description text of the text following evaluation item is used to describe at least one following error existing in the edited image relative to the editing description text; The defect description text of the image keeping evaluation item is used to describe at least one keeping error existing in the edited image relative to the original image; The defect description text of the image quality evaluation item is used to describe at least one quality defect existing in the edited image relative to the original image.
5. The method of claim 1, wherein, The evaluation information is introduced into a denoising network of the image editing model as condition information.
6. The method of claim 5, wherein, For any cross attention layer in the denoising network, input data of a data fusion layer corresponding to the cross attention layer is determined according to the evaluation information and output data of the cross attention layer, and the input data of the cross attention layer includes the encoding result of the editing description text.
7. The method of claim 6, wherein, The input data of the data fusion layer is determined according to the feature vector of the evaluation information; The evaluation information includes at least one type of data, and the feature vector of the evaluation information is determined according to the vectorization result of each type of data and each type of vectorization result.
8. The method of claim 6, wherein, The data fusion layer is configured to perform residual calculation or sum calculation according to the output data of the cross attention layer and the mapping result of the feature vector of the evaluation information in a first feature space, and the first feature space is a feature space to which the output data of the cross attention layer belongs.
9. The method of claim 1, wherein, The input data of the denoising network in the image editing model includes first data and second data, the first data is input into the denoising network as conditional information, the input data of a first network layer in the denoising network includes the second data, and the first data is different from the second data. The second data is determined according to the original image and the evaluation information.
10. The method of claim 9, wherein, The determination process of the second data includes: Splicing the image features of the original image and the noise-added image features of the edited image to obtain a splicing result; Convolution processing is performed on the splicing result to obtain a convolution result; At least one cross attention processing is performed according to the mapping result of the convolution result and the feature vector of the evaluation information in a second feature space to obtain the second data, and the second feature space is a feature space to which the convolution result belongs.
11. The method according to any one of claims 1 to 10, characterized in that, The image editing model includes a first module, a second module, a third module, a linear layer corresponding to the third module, a denoising network, a linear layer corresponding to the denoising network, and a decoding module; The first module is configured to obtain a splicing result between the image features of the original image and the noise-added image features of the edited image; The second module is configured to obtain a feature vector of the evaluation information; The linear layer corresponding to the third module and the linear layer corresponding to the denoising network are both configured to process the feature vector of the evaluation information; The third module is configured to perform at least one cross attention processing on the splicing result and the output data of the linear layer corresponding to the third module; The denoising network is configured to perform denoising processing on the output data of the third module, the encoding result of the editing description text, and the output data of the linear layer corresponding to the denoising network; The decoding module is configured to perform decoding processing on the output data of the denoising network to obtain the image editing result.
12. The method of any one of claims 11, wherein, The updating of the image editing model includes: Updating part of the modules in the image editing model, and the part of the modules includes part or all of the network layers in the second module, the third module, the linear layer corresponding to the third module, the denoising network, and the linear layer corresponding to the denoising network.
13. An image editing method characterized by, The method includes: obtain a target image, an editing description text corresponding to the target image, and preset constraint information; perform processing on the target image, the editing description text, and the preset constraint information by using an image editing model to obtain an image editing result corresponding to the target image, the preset constraint information being used to describe a state of the image editing result under at least one evaluation item, the image editing model being determined by using the model training method in any one of claims 1-12.
14. A model training apparatus, comprising: comprising: a first obtaining unit configured to obtain an original image, an editing description text corresponding to the original image, an edited image corresponding to the original image under the editing description text, and evaluation information of the edited image, the evaluation information being used to describe a state of the edited image under at least one evaluation item; a first processing unit configured to perform processing on the original image, the editing description text, and the evaluation information by using an image editing model to obtain an image editing result corresponding to the original image; a model updating unit configured to update the image editing model according to a difference between the image editing result and the edited image.
15. An image editing apparatus characterized by comprising: comprising: a second obtaining unit configured to obtain a target image, an editing description text corresponding to the target image, and preset constraint information; a second processing unit configured to perform processing on the target image, the editing description text, and the preset constraint information by using an image editing model to obtain an image editing result corresponding to the target image, the preset constraint information being used to describe a state of the image editing result under at least one evaluation item, the image editing model being determined by using the model training method in any one of claims 1-12.
16. An electronic device, comprising: The device comprises a processor and a memory; The memory is configured to store instructions or a computer program; The processor is configured to execute the instructions or the computer program in the memory, so that the electronic device executes the method in any one of claims 1-13.
17. A computer readable medium characterized by The computer readable medium stores instructions or a computer program, when the instructions or the computer program run on a device, make the device execute the method in any one of claims 1-13.
18. A computer program product, characterised in that, It comprises a computer program carried on a non-transitory computer readable medium, the computer program contains program code for executing the method in any one of claims 1-13.