Text-driven image editing quality evaluation method and device and storage medium

Through the text-driven image editing quality evaluation method, the image connection score, editing quality score and text consistency score are comprehensively obtained, which solves the problem of difficulty in comprehensively evaluating image editing effects in the prior art, and achieves more accurate and convincing image editing quality evaluation.

CN119991573AActive Publication Date: 2025-05-13PEKING UNIV SHENZHEN GRADUATE SCHOOL
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411986262.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-12-31
Publication Date
2025-05-13
Estimated Expiration
2044-12-31

AI Technical Summary

Technical Problem

The prior art is difficult to comprehensively and accurately evaluate the overall effect of image editing, and cannot fully reflect the true status of image editing quality.

Method used

A text-driven image editing quality evaluation method is proposed, and the image editing quality is comprehensively evaluated by obtaining image contact scores, editing quality scores and text consistency scores. The method includes inputting the source image and the features of the edited image into the interactive module to confirm the image connection, inputting the native quality evaluation module to perform quality evaluation, and inputting the edited text and image features into the multimodal attention interaction module to confirm text consistency.

Benefits of technology

Through multi-dimensional evaluation, the one-sidedness of single-angle evaluation is avoided, and the actual effect of image editing can be reflected more comprehensively and accurately, making the evaluation results more convincing and better aligned with human perception.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119991573A_ABST
    Figure CN119991573A_ABST
Patent Text Reader

Abstract

The invention discloses a text-driven image editing quality evaluation method and device and a storage medium, and relates to the technical field of artificial intelligence, and the method comprises the steps: obtaining a first image feature of a source image and a second image feature of an editing image, inputting to a source image and edited image interaction module for contact confirmation of the source image and the edited image to obtain an image contact score; inputting the second image feature into a native quality evaluation module to perform quality evaluation of the edited image to obtain an edited quality score; inputting the text feature of the edited text and the third image feature of the edited image into a multi-modal attention interaction module for consistency confirmation of the text and the image to obtain a text consistency score; and obtaining an image editing quality score based on the image contact score, the editing quality score and the text consistency score. According to the method, the image editing quality is scored from multiple dimensions, and the actual effect of image editing can be reflected more comprehensively and accurately.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of artificial intelligence technology, and in particular to a text-driven image editing quality evaluation method, device and storage medium. Background Art

[0002] After the source image is edited by editing text instructions, text-driven image editing quality evaluation is required to measure whether the edited image meets the expectations of the editing text instructions for image editing, and whether the edited image has achieved the expected changes compared with the source image.

[0003] Among the existing evaluation methods, CLIP-V (Visual Localization Model Based on CLIP) measures the effect of image editing by calculating the cosine similarity between the edited image and the source image. FID (Fréchet Inception Distance) measures the similarity of two images in the feature space by calculating the Fréchet distance between the source image and the edited image. The two indicators MSE (Mean Squared Error) and SSIM (Structural Similarity Index) represent the variance (MSE) and overall structural similarity (SSIM) between the edited image and the source image at the pixel level, respectively. However, the overall effect of image editing is the result of the combined effect of multiple factors. The existing evaluation methods only start from a single dimension and do not cover dimensions such as the alignment relationship between text and image. It is difficult to conduct a comprehensive and integrated evaluation of the overall effect of image editing, and it is impossible to fully reflect the true status of image editing quality.

[0004] The above contents are only used to assist in understanding the technical solution of the present application and do not constitute an admission that the above contents are prior art. Summary of the invention

[0005] The main purpose of this application is to provide a text-driven image editing quality evaluation method, device and storage medium, aiming to solve the technical problem of how to comprehensively and accurately evaluate the actual effect of image editing.

[0006] To achieve the above objectives, the present application proposes a text-driven image editing quality evaluation method, the method comprising:

[0007] Inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image, and obtaining an image connection score;

[0008] Inputting the second image feature into a native quality evaluation module to perform quality evaluation on the edited image to obtain an edit quality score;

[0009] Inputting the text feature of the edited text and the third image feature of the edited image into a multimodal attention interaction module to confirm the consistency of the text and the image, and obtaining a text consistency score;

[0010] An image editing quality score is obtained based on the image connection score, the editing quality score and the text consistency score.

[0011] In one embodiment, before the step of inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image to obtain the image connection score, the step further includes:

[0012] Inputting the source image into a pre-trained first visual backbone network model to obtain the first image feature;

[0013] The edited image is input into a pre-trained second visual backbone network model to obtain the second image features.

[0014] In one embodiment, the step of inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image to obtain the image connection score includes:

[0015] Splicing the first image feature and the second image feature in a channel dimension to obtain a first splicing feature;

[0016] Performing feature fusion on the first concatenated features through a self-attention mechanism and a forward network to obtain a first fused feature;

[0017] Based on a preset first weight parameter, the first fusion features are linearly combined and weighted to obtain the image connection score.

[0018] In one embodiment, the step of inputting the second image feature into a native quality evaluation module to evaluate the quality of the edited image to obtain an edit quality score comprises:

[0019] Construct a multi-layer feed-forward network consisting of linear layers;

[0020] The second image feature is input into the multi-layer forward network for regression calculation to obtain the editing quality score.

[0021] In one embodiment, before the step of inputting the text feature of the edited text and the third image feature of the edited image into the multimodal attention interaction module to confirm the consistency of the text and the image to obtain a text consistency score, the step further includes:

[0022] Inputting the edited text into a pre-trained text feature extractor for feature extraction to obtain the text features;

[0023] The edited image is input into a pre-trained convolutional neural network model for feature extraction to obtain the third image feature.

[0024] In one embodiment, the step of inputting the text feature of the edited text and the third image feature of the edited image into the multimodal attention interaction module to confirm the consistency of the text and the image to obtain a text consistency score includes:

[0025] Flattening the third image feature in the spatial dimension, and flattening the text feature in the length dimension;

[0026] The third image feature and the text feature that have been flattened are concatenated in a channel dimension to obtain a second concatenated feature;

[0027] Based on a preset second weight parameter, linear combination and weighted calculation are performed on the second splicing features to obtain the text consistency score.

[0028] In one embodiment, the step of obtaining the image editing quality score based on the image connection score, the editing quality score and the text consistency score comprises:

[0029] Performing a splicing operation on the image connection score, the editing quality score and the text consistency score to obtain a third splicing feature;

[0030] Based on a preset third weight parameter, the third splicing feature is linearly combined and weightedly calculated to obtain the image editing quality score.

[0031] In one embodiment, the method further comprises:

[0032] Build a source image dataset;

[0033] Editing a source image in the source image data set according to a preset editing text instruction set and an image editing method to obtain a first edited image;

[0034] An edited image dataset is constructed based on each of the first edited images and observers' evaluation scores of the first edited images.

[0035] In addition, to achieve the above-mentioned objectives, the present application also proposes a text-driven image editing quality assessment device, which includes: a memory, a processor, and a computer program stored on the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text-driven image editing quality assessment method as described above.

[0036] In addition, to achieve the above-mentioned purpose, the present application also proposes a storage medium, which is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the text-driven image editing quality evaluation method described above are implemented.

[0037] The present application provides a text-driven image editing quality evaluation method, which can measure the image editing quality from multiple different and key dimensions by obtaining the image connection score, the editing quality score and the text consistency score respectively. The image connection score focuses on the relationship between the source image and the edited image, the editing quality score focuses on the quality of the edited image itself, and the text consistency score focuses on the degree of fit between the edited text and the edited image at the semantic level. Combining the scores of these dimensions to obtain the image editing quality score avoids the one-sidedness of evaluating from only a single perspective, can more comprehensively and accurately reflect the actual effect of image editing, and make the final evaluation results more convincing. In addition, the way of obtaining scores of different dimensions in modules and combining them is closer to the subjective evaluation ideas of humans, so that the image editing quality score output by the model can be better aligned with human perception, improving the rationality and credibility of the evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0039] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative labor.

[0040] Figure 1 A flowchart diagram of the first embodiment of the text-driven image editing quality evaluation method of the present application;

[0041] Figure 2 A schematic diagram of the overall details of the process provided for the text-driven image editing quality evaluation method of this application;

[0042] Figure 3A flowchart diagram for the second implementation of the text-driven image editing quality assessment method of the present application;

[0043] Figure 4 A flowchart diagram of a third embodiment of the text-driven image editing quality evaluation method of the present application;

[0044] Figure 5 Schematic diagram of the device structure of the hardware operating environment involved in the text-driven image editing quality evaluation method in the embodiment of the present application. DETAILED DESCRIPTION

[0045] It should be understood that the specific embodiments described herein are only used to explain the technical solutions of the present application and are not used to limit the present application.

[0046] In order to better understand the technical solution of the present application, a detailed description will be given below in conjunction with the accompanying drawings and specific implementation methods.

[0047] The main solution of the embodiment of the present application is: input the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image, and obtain an image connection score; input the second image feature into the native quality evaluation module to perform quality evaluation of the edited image, and obtain an editing quality score; input the text feature of the edited text and the third image feature of the edited image into the multimodal attention interaction module to confirm the consistency of the text and the image, and obtain a text consistency score; based on the image connection score, the editing quality score and the text consistency score, obtain an image editing quality score.

[0048] After the source image is edited by editing text instructions, text-driven image editing quality evaluation is required to measure whether the edited image meets the expectations of the editing text instructions for image editing, and whether the edited image has achieved the expected changes compared with the source image.

[0049] Among the existing evaluation methods, CLIP-V measures the effect of image editing by calculating the cosine similarity between the edited image and the source image. FID measures the similarity of two images in the feature space by calculating the Fréchet distance between the source image and the edited image. The two indicators MSE and SSIM represent the variance (MSE) at the pixel level and the similarity (SSIM) in the overall structure between the edited image and the source image, respectively. However, the overall effect of image editing is the result of the combined effect of multiple factors. The existing evaluation methods only start from a single dimension and do not include dimensions such as the alignment relationship between text and image. It is difficult to conduct a comprehensive and integrated evaluation of the overall effect of image editing, and it is impossible to fully reflect the true status of image editing quality.

[0050] The present application provides a text-driven image editing quality evaluation method, which can measure the image editing quality from multiple different and key dimensions by obtaining the image connection score, the editing quality score and the text consistency score respectively. The image connection score focuses on the relationship between the source image and the edited image, the editing quality score focuses on the quality of the edited image itself, and the text consistency score focuses on the degree of fit between the edited text and the edited image at the semantic level. Combining the scores of these dimensions to obtain the image editing quality score avoids the one-sidedness of evaluating from only a single perspective, can more comprehensively and accurately reflect the actual effect of image editing, and make the final evaluation results more convincing. In addition, the way of obtaining scores of different dimensions in modules and combining them is closer to the subjective evaluation ideas of humans, so that the image editing quality score output by the model can be better aligned with human perception, improving the rationality and credibility of the evaluation results.

[0051] It should be noted that the execution subject of this embodiment can be a computing service device with network communication and program running functions, such as a tablet computer, a personal computer, a mobile phone, etc., or an electronic device or device capable of realizing the above functions. The following takes a text-driven image editing quality evaluation device as an example to illustrate this embodiment and the following embodiments.

[0052] Based on this, the present application embodiment provides a text-driven image editing quality evaluation method, referring to Figure 1 , Figure 1 This is a flowchart diagram of the first embodiment of the text-driven image editing quality assessment method of the present application.

[0053] In this embodiment, the text-driven image editing quality evaluation method includes steps S100 to S400:

[0054] Step S100: input the first image feature of the source image and the second image feature of the edited image into a source image and edited image interaction module to confirm the connection between the source image and the edited image, and obtain an image connection score.

[0055] It should be noted that the source image refers to the original image that has not been edited, and the edited image refers to the image that has been edited or processed in some way. This editing can be cropping, adjusting brightness or contrast, adding filters, repairing defects, etc., or more complex operations such as image synthesis, style transfer, etc. The edited image is the result of the source image after processing, and contains features or information different from the source image.

[0056] In this embodiment, the first image feature of the source image and the second image feature of the edited image are input into the source image and edited image interaction module to confirm the connection between the source image and the edited image and obtain the image connection score. The source image and edited image interaction module is a linear layer and a self-attention layer connected in series, which can dynamically model the source feature and the target feature. The image connection score is usually used to determine whether the edited image still retains the important features of the source image.

[0057] In a feasible implementation manner, before step S100, the method further includes:

[0058] Inputting the source image into a pre-trained first visual backbone network model to obtain the first image feature;

[0059] The edited image is input into a pre-trained second visual backbone network model to obtain the second image features.

[0060] It should be noted that the Vision Transformer (ViT) backbone network model is a network model based on the Transformer architecture and applied to the visual field. It treats images as a series of image patches, just like processing word sequences in text. Each image patch is converted into a vector representation through linear embedding and other operations, and then input into the Transformer encoder. The Transformer encoder contains components such as multi-head attention mechanism, feedforward neural network, and layer normalization. The multi-head attention mechanism enables the model to pay attention to the relationship between different parts of the image from multiple different "perspectives", and explore the importance of each element in the image and their mutual influence. For example, in an image containing people and background buildings, it can determine the degree of correlation between the main body of the person and surrounding buildings, so as to comprehensively extract image features.

[0061] In this embodiment, the source image and the edited image are input into the quality assessment model. The source image and the edited image are respectively input into two visual quality encoders for image feature extraction. The visual quality encoder is a visual backbone network ViT pre-trained on a large-scale dataset. The visual backbone network is pre-trained on a large number of image datasets such as ImageNet, which can ensure the correct extraction of the representation of the source image and the edited image in the latent space. The two visual quality encoders have the same structure, but do not share parameters.

[0062] Step S200: input the second image feature into a native quality evaluation module to perform quality evaluation on the edited image to obtain an edit quality score.

[0063] In this embodiment, the image visual encoder CLIP (Contrastive Language-Image Pre-training) is used to extract the latent space features (i.e., the second image features) of the input edited image by utilizing its prior knowledge and excellent multimodal alignment properties on the pre-trained image quality evaluation dataset. The obtained latent space features are input into the native quality evaluation module of the linear layer to obtain an evaluation of the native quality dimension. Optionally, the image encoder of CLIP can be a pre-trained model such as ViT or CNN.

[0064] In this implementation, edited images in an edited image dataset, which contains edited images and their corresponding native quality scores, are used for model training and validation. The native quality scores are obtained by conducting subjective experiments and collecting human feedback. When using an edited image dataset containing quality scores, CLIP can associate the general visual features of images (such as object shape, color distribution, texture structure, etc.) learned in pre-training with specific quality evaluation information. For example, CLIP may already know that certain specific texture structures or color combinations are generally associated with high-quality images, and in this image quality evaluation dataset, if these features correspond to high native quality scores, CLIP can further strengthen this association, thereby better utilizing its prior knowledge to understand the essential characteristics of image quality.

[0065] In this embodiment, the edited image is input into the image visual encoder of the CLIP model. The image encoder of the CLIP model processes the image. For example, when ViT is used as the image encoder, the image is segmented into image blocks, and after linear projection and position encoding, feature extraction is performed through the Transformer Encoder to finally obtain the latent space feature representation of the image.

[0066] Step S300, input the text features of the edited text and the third image features of the edited image into a multimodal attention interaction module to confirm the consistency of the text and the image, and obtain a text consistency score.

[0067] In this implementation, the edited text is passed through the text encoder of CLIP to obtain text features aligned with the visual information, and the edited edited image is passed through the corresponding ResNet50 to obtain visual features aligned with the text information (i.e., the third image features). Then, the text features and the third image features are input into the multimodal attention interaction module, which is the self-attention and pooling module based on IP-IQA introduced below. Based on the self-attention mechanism and the pooling layer, the text features and the visual features are respectively passed through the Transformer block and flattened in the spatial dimension (the text features correspond to the length dimension) and input into the pooling layer, and the compressed text features and visual features are spliced ​​together in the channel dimension through the pooling layer, and finally the text-edited image semantic consistency measurement score is obtained through linear layer regression.

[0068] In a feasible embodiment, before step S300, the following steps may also be included:

[0069] Inputting the edited text into a pre-trained text feature extractor for feature extraction to obtain the text features;

[0070] The edited image is input into a pre-trained convolutional neural network model for feature extraction to obtain the third image feature.

[0071] It should be noted that CLIP (Contrastive Language-Image Pretraining) is a multimodal machine learning model. It aims to learn to understand the content of images and match them with corresponding natural language descriptions through training on a large number of text-image pairs.

[0072] In addition, it should be noted that ResNet50 is a deep convolutional neural network architecture with powerful image feature extraction capabilities.

[0073] In this embodiment, the edited text is input into the text encoder of CLIP, and the input edited text is mapped into a specific vector representation form, namely, text features. In this process, the text encoder converts the vocabulary, grammatical structure and other information in the text into a feature vector that can be compared and associated with the visual information based on its learned language semantic understanding and the pattern aligned with the visual information.

[0074] For example, for the edited text "Replace this man with Iron Man", the semantic information in the text is first parsed, and the semantics contained in key nouns such as "man" and "Iron Man", as well as the word "replace with", which indicates a change in action, are converted into corresponding feature vectors. The text encoder will extract feature vectors corresponding to the key semantic change from the original man's image to the Iron Man's image based on the language understanding patterns learned in the past and the alignment with the visual information. Moreover, these vectors are intrinsically related to the differences in the visual presentation of the man and Iron Man in reality. For example, the features corresponding to the visual elements such as Iron Man's unique armor shape and color will be reflected in the vectors, preparing for the subsequent matching with the visual features of the edited image.

[0075] In this embodiment, the edited result image corresponding to the edited text is input into ResNet50 to obtain the visual features aligned with the text information. For the edited result image, ResNet50 gradually extracts features of different levels in the image according to a series of operations such as convolution layer and pooling layer. From low-level edge and texture features to high-level object shape and semantic category related features, all will be mined.

[0076] For example, for edited images (images showing the image of Iron Man), ResNet50 begins to exert its powerful image feature extraction capabilities. Through the convolution layer, local features at different positions in the image are captured, such as the texture on Iron Man's armor, the edges of various components and other low-level features. As the network goes deeper, through operations such as the pooling layer, high-level features are gradually integrated. For example, the outline of the entire Iron Man, his iconic energy reactor and other features related to semantic categories can be recognized, and information such as Iron Man's position in the picture and his relative relationship with the background can be extracted. These extracted visual features can correspond to the "Iron Man" related semantic elements mentioned in the edited text. For example, the appearance features of the armor echo the image of Iron Man described in the text, which is convenient for subsequent matching considerations with text features.

[0077] Step S400: obtaining an image editing quality score based on the image connection score, the editing quality score and the text consistency score.

[0078] In this embodiment, the image connection score, the editing quality score and the text consistency score are input into a multi-dimensional comprehensive regressor. The multi-dimensional comprehensive regressor is a linear layer that inputs the concatenation of features of different dimensions and obtains the final score through regression.

[0079] In an optional implementation, step S400 includes:

[0080] Performing a splicing operation on the image connection score, the editing quality score and the text consistency score to obtain a third splicing feature;

[0081] Based on a preset third weight parameter, the third splicing feature is linearly combined and weightedly calculated to obtain the image editing quality score.

[0082] In this embodiment, the scores of three different dimensions, namely, the image connection score, the editing quality score and the text consistency score, are concatenated to form a comprehensive feature vector. The concatenated feature vector is used as input and sent to the multidimensional comprehensive regressor. After receiving the concatenated feature vector, the multidimensional comprehensive regressor performs linear combination and weighted calculation on the features of each dimension according to the pre-trained weight parameters, and finally outputs a scalar value, which is the image editing quality score obtained by regression of the multidimensional comprehensive regressor.

[0083] Please refer to Figure 2 In this embodiment, two visual quality encoders are used to extract the image features of the source image and the edited image respectively, and the two image features are input into the source image and edited image interaction module to obtain an image connection score. The image features of the edited image are input into the native quality evaluation module to obtain the editing quality score. The image features of the edited image extracted by the visual feature extractor and the text features of the edited text extracted by the text feature extractor are input into the multimodal attention interaction module to obtain the text consistency score. By obtaining the image connection score, the editing quality score and the text consistency score respectively, the image editing quality is measured from multiple different and key dimensions, which can more comprehensively and accurately reflect the overall quality level of the image after editing.

[0084] Based on the first embodiment of the present application, in the second embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 3 , the load parameters are the load value of the central processing unit and the load value of the graphics processing unit, and step S100 may include steps S110 to S130:

[0085] Step S110: splicing the first image feature and the second image feature in the channel dimension to obtain a first splicing feature.

[0086] In this embodiment, when the visual backbone network model ViT extracts corresponding correlation features from the source image and the edited image respectively, the feature vector shape of the first image feature extracted from the source image is [batch_size, num_channels1, height1, width1], where batch_size represents the number of images processed at a time, num_channels1 represents the number of channels of the feature, and height1 and width1 represent the size of the feature in the spatial dimension. The feature vector shape of the second image feature extracted from the edited image is [batch_size, num_channels2, height2, width2]. After extracting the first image feature and the second image feature, a splicing operation is performed in the channel dimension to merge the two features together along the channel direction to form a new feature vector, whose shape becomes [batch_size, num_channels1+num_channels2, height, width]. Height and width can take the appropriate spatial dimension size of the two or the size after certain adjustment.

[0087] For example, if the first image feature has 32 channels and the second image feature has 64 channels, a new feature vector with 96 channels will be obtained after splicing. This new feature combines the feature information of the source image and the edited image, and integrates the two in the same dimension, so that they can be input as a whole into the subsequent network layer for unified processing, and explore the deeper connections between them and the fused feature representation.

[0088] Step S120, performing feature fusion on the first concatenated features through a self-attention mechanism and a forward network to obtain a first fused feature.

[0089] In this embodiment, the first spliced ​​feature obtained after splicing first enters the linear layer, which is essentially a fully connected neural network layer. It will perform a linear transformation on the input feature vector, and change the dimension and numerical distribution of the feature vector according to certain rules through the pre-learned weight matrix and bias vector. For example, it can map the high-dimensional splicing feature to a space of intermediate dimensions, perform preliminary integration and adjustment on the features, extract feature representations that are more suitable for subsequent self-attention mechanism processing, and also help the transition and adaptation of features between different network layers, so that the features meet the expected input form of the entire network structure. The features processed by the linear layer then enter the self-attention module. The self-attention mechanism can focus on the correlation between different elements in the feature vector again, and redistribute weights according to the correlation between the elements, so that the model pays more attention to the important parts for the fusion features. For example, in the image features after splicing, the self-attention mechanism will give higher weights to important feature elements such as changes in the main body of the image and differences in key areas, further highlight these key information, and strengthen the fusion effect between features, so that the model can better capture the association and changes between the source image and the edited image.

[0090] In this embodiment, the features output by the self-attention module are further processed by a feed-forward network, which is usually composed of multiple fully connected layers, activation functions, etc. It can perform nonlinear transformation on the features adjusted by self-attention, further explore the complex relationship between features, refine and enrich the representation of fused features, so that the final fused features can fully contain various valuable information generated by the source image and the edited image in the fusion process, such as the similarities and differences between the two, etc. This information is critical for the subsequent accurate measurement of dynamic modeling.

[0091] Step S130: Based on a preset first weight parameter, linear combination and weighted calculation are performed on the first fusion features to obtain the image connection score.

[0092] In this embodiment, the fused features obtained through the previous steps will be input into the last linear layer for linear transformation operations, and the fused features will be mapped to a specific value, which is the final dynamic modeling score. The linear layer calculates the input fused features based on the weight parameters learned in advance during the training process and compresses them into a scalar value. For example, this score can be set within a certain interval, such as between 0 and 1, where 0 may indicate that there is almost no expected dynamic association or consistency between the source image and the edited image, and 1 indicates that there is a very high dynamic association between the two, that is, the edited image is highly consistent with the source image in terms of dynamic changes, etc. The intermediate values ​​correspond to different degrees of dynamic modeling, reflecting the degree of consistency between the source image and the edited image in terms of dynamic features.

[0093] In this embodiment, ViT is used to extract image features, the image features are spliced ​​and fused, and a dynamic modeling score is obtained through linear regression, thereby achieving quantitative analysis and evaluation of the correlation characteristics between the source image and the edited image.

[0094] Based on the first embodiment of the present application, in the third embodiment of the present application, the same or similar contents as those in the above-mentioned embodiment 1 can be referred to the above introduction, and will not be repeated in the following. Figure 3 , step S200 may include steps S210 to S220:

[0095] Step S210, constructing a multi-layer forward network consisting of linear layers.

[0096] Step S220: input the second image feature into the multi-layer forward network for regression calculation to obtain the editing quality score.

[0097] In this implementation, a multi-layer feed-forward network consisting of linear layers is constructed. The input of the network is the second image feature extracted by the CLIP model, and the output is the predicted native quality evaluation. The number of layers of the network and the number of neurons in each layer can be designed according to actual conditions. For example, 2-3 layers of linear layers can be used, and the middle layer can use activation functions such as ReLU to increase the nonlinear expression ability of the network.

[0098] In this embodiment, a suitable loss function is selected to measure the difference between the predicted native quality and the true quality label. Exemplarily, a mean square error (MSE) loss function is selected to calculate the average square error between the predicted value and the true value in the regression task to guide the model training process so that the model can learn accurate quality evaluation prediction capabilities. The edited image data set is divided into a training set and a validation set. During the training process, for each batch of images in the training set, the CLIP model is input to extract features, and then the features are input into the multi-layer forward network to obtain the predicted quality, and the value of the loss function is calculated based on the predicted quality and the true quality label. The parameters of the multi-layer forward network are updated according to the gradient, such as stochastic gradient descent, Adam, etc., and hyperparameters such as the learning rate are adjusted to optimize the model training process so that the model gradually converges to a better performance state. After each training cycle, the performance of the model is evaluated using the validation set, such as calculating the loss value or other relevant indicators (such as mean absolute error, etc.) on the validation set, monitoring whether the model is overfitting or underfitting, and adjusting the training strategy according to the evaluation results, such as adjusting the learning rate, increasing or decreasing the number of training rounds, etc. After training is completed, the model is fully evaluated using an independent test set, and various evaluation metrics (such as mean square error, coefficient of determination, etc.) are calculated to measure the accuracy and generalization ability of the model in predicting native quality.

[0099] Based on the first embodiment of the present application, in the fourth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be described in detail later. Figure 3 , step S300 may include steps S310 to S330:

[0100] Step S310: flatten the third image feature in the spatial dimension, and flatten the text feature in the length dimension.

[0101] In this embodiment, the text features and visual features are further converted and enhanced feature representation by the Transformer block. The Transformer block has structures such as multi-head attention mechanism and feedforward neural network. The multi-head attention mechanism can capture the relationship between features from multiple angles, and further refine and optimize the feature representation. After the Transformer block, the text features and visual features are more effectively organized and refined in terms of semantic expression. Secondly, the pooling layer operation is performed, and the features after the Transformer block are flattened in the spatial dimension (the length dimension corresponds to the text feature), and the multi-dimensional feature matrix is ​​converted into a one-dimensional vector form, and then input into the pooling layer. The pooling layer usually adopts methods such as maximum pooling or average pooling. For example, the maximum pooling will select the maximum value in each local area in the feature vector as the output, which can further compress the features, extract the most representative feature information, reduce the amount of data while retaining key semantic features, and facilitate subsequent splicing and calculation.

[0102] Step S320: splicing the third image feature and the text feature after the flattening operation in the channel dimension to obtain a second splicing feature.

[0103] In this embodiment, text features and visual features are spliced ​​in the channel dimension. The text features and visual features compressed by the pooling layer are spliced ​​in the channel dimension. For example, if the text feature is a vector with a shape of [1,n] (n is the feature length) after the previous processing, and the visual feature is a vector with a shape of [1,m] (m is the feature length), then they will be spliced ​​along the channel dimension into a new vector with a shape of [1,n+m]. In this way, the key semantic features of both text and image are fused together to form a comprehensive feature representation, which contains the core information of both text description and image presentation, which is convenient for subsequent unified regression analysis to measure the semantic consistency of the two.

[0104] Step S330: Based on a preset second weight parameter, linear combination and weighted calculation are performed on the second concatenation features to obtain the text consistency score.

[0105] In this embodiment, the spliced ​​comprehensive feature vector will be input into the linear layer. The linear layer is actually a simple fully connected neural network layer. It will perform a linear transformation on the input feature vector according to the pre-learned weight parameters and map it to a specific numerical range. This numerical value is the final measurement score representing the semantic consistency between the text and the edited image. For example, if after training, the score output by the linear layer is between 0 and 1, 0 may indicate that the text and image semantics are completely inconsistent, 1 indicates complete consistency, and the values ​​in between correspond to different degrees of semantic consistency. The model learns the appropriate weights by training on a large amount of annotated text-image pair data to accurately give a reasonable consistency measurement score based on the input features. Using the mapping ability of the linear layer to output a specific measurement score can intuitively reflect the degree of semantic consistency between the text and the edited image.

[0106] Based on the first embodiment of the present application, in the fifth embodiment of the present application, the same or similar contents as those in the first embodiment can be referred to the above description, and will not be described in detail later. Figure 4 , before step S100, steps S01 to S03 may be included:

[0107] Step S01, constructing a source image data set;

[0108] In this embodiment, during the source image construction process, the content of the source image data set includes but is not limited to real-world scenes, computer graphics rendering images, text-driven generated images, and works of art. Considering the current wide range of application scenarios, real-world scenes occupy a large proportion. When collecting source images, for considerations such as copyright, watermarking, and resolution, this embodiment selects several images from four data sets: ADE20k data set, WIKIArt (WIKIArt Dataset), COCO (Common Objects in Context), ReasonEdit, and other network resources.

[0109] In this embodiment, when collecting source images from the above four datasets, the data source (real-world dataset, computer graphics, AIGC (Artificial Intelligence Generated Content), and works of art) is first confirmed, and then each sample is checked one by one, and attributes such as "[landscape / object / animal / human]" and "[action type]" are marked. Samples of the same type (for example, landscape / object / animal or human action) or with insufficient resolution are skipped until the dataset is fully checked or the number of samples of the relevant category is sufficient. For smaller datasets, priority is given to checking, and more samples are extracted from larger datasets. Finally, suitable images are selected from the Internet for supplementation. For example, although many datasets contain landscapes such as grasslands and snow-capped mountains, scenes such as auroras, lava flows, and lightning are less common. Similarly, although the current dataset provides a rich category of action, there are still certain deficiencies in significantly changing action patterns.

[0110] Step S02, editing the source image in the source image data set according to a preset editing text instruction set and an image editing method to obtain a first edited image;

[0111] In this embodiment, in order to ensure the specificity and diversity of the instructions, corresponding text instructions are generated for each source image. Editing text instructions can be divided into three categories: (1) style editing, including editing of color, texture or overall atmosphere; (2) semantic editing, including background editing and local editing, such as adding, replacing or removing specific objects; (3) structural editing, including changes in object size, posture, action, etc.

[0112] In this embodiment, in order to ensure the distribution of edited image quality, different multiple image editing methods are selected to edit the source images in the source image dataset. The image editing method can be constructed by the latest deep learning technology and a large amount of training data, aiming to provide a performance model of high-quality image editing effects.

[0113] Exemplarily, in this embodiment, editing methods of different basic models are selected, covering models from SD (Stable Diffusion) 1-4 to SD2-1, to improve the diversity of editing results. In addition, in order to ensure the diversity of editing content, the selection includes 0-shot methods and methods that require fine-tuning. Models are also selected according to different editing paradigms, including effective editing strategies such as Instruct-P2P (Instruct Point-to-Point), Prompt-to-prompt, MasaCtrl (Tuning-Free Mutual Self-Attention Control), etc. Instruct-P2P is a technology that precisely controls specific areas or elements in an image for point-to-point editing through instructions. The Prompt-to-prompt strategy allows users to gradually adjust the content of the generated image by modifying the prompt words, providing a flexible editing method. MasaCtrl is a method for image synthesis and editing, which achieves consistent image generation and complex non-rigid image editing without fine-tuning by converting self-attention in a diffusion model to mutual self-attention. It is able to query relevant local content and texture from the source image to ensure the consistency of the generated image.

[0114] Step S03: constructing an edited image dataset according to each of the first edited images and the observer's evaluation score of the first edited image.

[0115] In this embodiment, subjective experiments are conducted to collect human feedback to construct an edited image dataset with human feedback. According to the ITU standard (International Telecommunication Union Standards), the number of participants in a subjective experiment should be at least 15 to ensure that the variance of the results is within a controllable range. Exemplarily, 25 experimental participants with diverse backgrounds are recruited. During the experiment, participants were asked to comprehensively evaluate the consistency of text and image, the fidelity of source images and edited images, and the quality of edited images based on subjective impressions. All participants were over 18 years old and had at least a bachelor's degree. Their backgrounds covered multiple fields such as business, engineering, science, and law, and they had independent judgment. Before the experiment began, all participants underwent face-to-face training, in which examples of good and poor editing that were not included in the dataset were shown. In the experiment, each participant evaluated all video samples and was forced to rest for 5 minutes every 15 minutes of work to avoid fatigue.

[0116] In this embodiment, by conducting subjective experiments to collect human feedback and constructing an edited image dataset with human feedback, human subjective evaluation can be aligned with machine evaluation, thereby improving the reliability and practicality of the evaluation results.

[0117] The present application provides a text-driven image editing quality assessment device, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the text-driven image editing quality assessment method in the above-mentioned embodiment one.

[0118] Reference below Figure 5 , which shows a schematic diagram of the structure of a text-driven image editing quality evaluation device suitable for implementing the embodiment of the present application. The text-driven image editing quality evaluation device in the embodiment of the present application may include but is not limited to mobile terminals such as mobile phones, laptop computers, digital broadcast receivers, PDAs (Personal Digital Assistants), PADs (portable android devices), PMPs (Portable Media Players), vehicle-mounted terminals (such as vehicle-mounted navigation terminals), etc., and fixed terminals such as digital TVs, desktop computers, etc. Figure 5 The text-driven image editing quality assessment device shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present application.

[0119] like Figure 5As shown, the text-driven image editing quality assessment device may include a processing device 1001 (e.g., a central processing unit, a graphics processor, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM: Read Only Memory) 1002 or a program loaded from a storage device 1003 to a random access memory (RAM: Random Access Memory) 1004. Various programs and data required for the operation of the text-driven image editing quality assessment device are also stored in the RAM 1004. The processing device 1001, the ROM 1002, and the RAM 1004 are connected to each other via a bus 1005. An input / output (I / O) interface 1006 is also connected to the bus. Typically, the following systems can be connected to the I / O interface 1006: input devices 1007 including, for example, a touch screen, a touchpad, a keyboard, a mouse, an image sensor, a microphone, an accelerometer, a gyroscope, etc.; output devices 1008 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; storage devices 1003 including, for example, a magnetic tape, a hard disk, etc.; and communication devices 1009. The communication device 1009 can allow the text-driven image editing quality assessment device to communicate wirelessly or wired with other devices to exchange data. Although the figure shows a text-driven image editing quality assessment device with various systems, it should be understood that it is not required to implement or have all the systems shown. More or fewer systems may be implemented or have instead.

[0120] In particular, according to the embodiments disclosed in the present application, the process described above with reference to the flowchart can be implemented as a computer software program. For example, the embodiments disclosed in the present application include a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network through a communication device, or installed from a storage device 1003, or installed from a ROM 1002. When the computer program is executed by the processing device 1001, the above-mentioned functions defined in the method of the embodiment disclosed in the present application are executed.

[0121] The text-driven image editing quality evaluation device provided by the present application adopts the text-driven image editing quality evaluation method in the above-mentioned embodiment, which can solve the technical problem of how to improve the comprehensive and accurate evaluation of the actual effect of image editing on long texts. Compared with the prior art, the beneficial effects of the text-driven image editing quality evaluation device provided by the present application are the same as the beneficial effects of the text-driven image editing quality evaluation method provided by the above-mentioned embodiment, and the other technical features in the text-driven image editing quality evaluation device are the same as the features disclosed in the method of the previous embodiment, which will not be repeated here.

[0122] It should be understood that the various parts disclosed in this application can be implemented by hardware, software, firmware or a combination thereof. In the description of the above embodiments, specific features, structures, materials or characteristics can be combined in any one or more embodiments or examples in a suitable manner.

[0123] The above is only a specific implementation of the present application, but the protection scope of the present application is not limited thereto. Any technician familiar with the technical field can easily think of changes or substitutions within the technical scope disclosed in the present application, which should be included in the protection scope of the present application. Therefore, the protection scope of the present application should be based on the protection scope of the claims.

[0124] The present application provides a computer-readable storage medium having computer-readable program instructions (ie, computer programs) stored thereon, and the computer-readable program instructions are used to execute the text-driven image editing quality assessment method in the above-mentioned embodiment.

[0125] The computer-readable storage medium provided in the present application may be, for example, a USB flash drive, but is not limited to electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, systems or devices, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, system or device. The program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination of the above.

[0126] The computer-readable storage medium may be included in the text-driven image editing quality assessment device; or may exist independently without being assembled into the text-driven image editing quality assessment device.

[0127] The above-mentioned computer-readable storage medium carries one or more programs. When the above-mentioned one or more programs are executed by a text-driven image editing quality assessment device, the text-driven image editing quality assessment device: inputs the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image, and obtains an image connection score; inputs the second image feature into the native quality assessment module to perform quality assessment of the edited image, and obtains an editing quality score; inputs the text feature of the edited text and the third image feature of the edited image into the multimodal attention interaction module to confirm the consistency of the text and the image, and obtains a text consistency score; and obtains an image editing quality score based on the image connection score, the editing quality score and the text consistency score.

[0128] Computer program code for performing the operations of the present application may be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a separate software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer may be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0129] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, function and operation of the system, method and computer program product according to various embodiments of the present application. In this regard, each square box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the module, the program segment or a part of the code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the square box can also occur in a sequence different from that marked in the accompanying drawings. For example, two square boxes represented in succession can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each square box in the block diagram and / or flow chart, and the combination of the square boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operation, or can be implemented with a combination of dedicated hardware and computer instructions.

[0130] The modules involved in the embodiments described in this application may be implemented by software or hardware, wherein the name of the module does not constitute a limitation on the unit itself in some cases.

[0131] The readable storage medium provided by the present application is a computer-readable storage medium, which stores computer-readable program instructions (i.e., computer programs) for executing the above-mentioned text-driven image editing quality evaluation method, and can solve the technical problem of how to improve the comprehensive and accurate evaluation of the actual effect of image editing on long texts. Compared with the prior art, the beneficial effects of the computer-readable storage medium provided by the present application are the same as the beneficial effects of the text-driven image editing quality evaluation method provided by the above-mentioned embodiment, and will not be repeated here.

[0132] The above descriptions are only some embodiments of the present application, and are not intended to limit the patent scope of the present application. All equivalent structural changes made using the contents of the present application specification and drawings under the technical concept of the present application, or direct / indirect applications in other related technical fields are included in the patent protection scope of the present application.

Claims

1. A text-driven image editing quality evaluation method, characterized in that: The method includes: Inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image, and obtaining an image connection score; Inputting the second image feature into a native quality evaluation module to perform quality evaluation on the edited image to obtain an edit quality score; Inputting the text feature of the edited text and the third image feature of the edited image into a multimodal attention interaction module to confirm the consistency of the text and the image, and obtaining a text consistency score; An image editing quality score is obtained based on the image connection score, the editing quality score and the text consistency score.

2. The text-driven image editing quality evaluation method according to claim 1, characterized in that: Before the step of inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image to obtain the image connection score, the step further includes: Inputting the source image into a pre-trained first visual backbone network model to obtain the first image feature; The edited image is input into a pre-trained second visual backbone network model to obtain the second image features.

3. The text-driven image editing quality evaluation method according to claim 1, characterized in that: The step of inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image to obtain the image connection score comprises: Splicing the first image feature and the second image feature in a channel dimension to obtain a first splicing feature; Performing feature fusion on the first concatenated features through a self-attention mechanism and a forward network to obtain a first fused feature; Based on a preset first weight parameter, the first fusion features are linearly combined and weighted to obtain the image connection score.

4. The text-driven image editing quality evaluation method according to claim 1, characterized in that: The step of inputting the second image feature into the native quality evaluation module to evaluate the quality of the edited image to obtain an edit quality score comprises: Construct a multi-layer feed-forward network consisting of linear layers; The second image feature is input into the multi-layer forward network for regression calculation to obtain the editing quality score.

5. The text-driven image editing quality assessment method according to claim 1, characterized in that: Before the step of inputting the text feature of the edited text and the third image feature of the edited image into the multimodal attention interaction module to confirm the consistency of the text and the image to obtain a text consistency score, the step further includes: Inputting the edited text into a pre-trained text feature extractor for feature extraction to obtain the text features; The edited image is input into a pre-trained convolutional neural network model for feature extraction to obtain the third image feature.

6. The text-driven image editing quality assessment method according to claim 1, characterized in that: The step of inputting the text feature of the edited text and the third image feature of the edited image into the multimodal attention interaction module to confirm the consistency of the text and the image to obtain a text consistency score comprises: Flattening the third image feature in the spatial dimension, and flattening the text feature in the length dimension; The third image feature and the text feature that have been flattened are concatenated in a channel dimension to obtain a second concatenated feature; Based on a preset second weight parameter, linear combination and weighted calculation are performed on the second splicing features to obtain the text consistency score.

7. The text-driven image editing quality assessment method according to claim 1, characterized in that: The step of obtaining the image editing quality score based on the image connection score, the editing quality score and the text consistency score comprises: Performing a splicing operation on the image connection score, the editing quality score and the text consistency score to obtain a third splicing feature; Based on a preset third weight parameter, the third splicing feature is linearly combined and weightedly calculated to obtain the image editing quality score.

8. The text-driven image editing quality assessment method according to claim 1, characterized in that: Before the step of inputting the first image feature of the source image and the second image feature of the edited image into the source image and edited image interaction module to confirm the connection between the source image and the edited image to obtain the image connection score, the step further includes: Build a source image dataset; Editing a source image in the source image data set according to a preset editing text instruction set and an image editing method to obtain a first edited image; An edited image dataset is constructed based on each of the first edited images and observers' evaluation scores of the first edited images.

9. A text-driven image editing quality assessment device, characterized in that: The device comprises: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is configured to implement the steps of the text-driven image editing quality assessment method according to any one of claims 1 to 8.

10. A storage medium, characterized in that: The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, the steps of the text-driven image editing quality assessment method according to any one of claims 1 to 8 are implemented.

Citation Information

Patent Citations

  • Image quality evaluation method, system and equipment of AI image and medium

    CN118154571A

  • Performing global image editing using editing operations determined from natural language requests

    US20220399017A1