Multi-modal Image Aesthetic Quality Evaluation Method Combining Local and Global Image Features
Through a multimodal image aesthetic quality evaluation method that integrates local and global image features, combined with autoencoder and cross encoder, the problem of limited manual features in the prior art is solved, and a more accurate image aesthetic quality evaluation is achieved.
Patent Information
- Application Number
- CN202211554654.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-05
- Publication Date
- 2025-07-25
- Estimated Expiration
- 2042-12-05
AI Technical Summary
In the evaluation of image aesthetic quality, the characteristics of hand-design are limited and cannot fully represent aesthetic characteristics, resulting in insufficient validity of the evaluation results and strong subjectivity, making it difficult to automatically evaluate image aesthetic quality.
A multimodal image aesthetic quality evaluation method that fuses local and global image features is designed, and an image feature extraction subnet is combined with a text feature extraction subnet to build a multimodal image aesthetic quality evaluation network, using an autoencoder and a cross encoder to perform feature fusion, and output aesthetic score distribution through a full connection layer.
Effectively fuse the local and global features of the image, and use the text features in user comments to improve the accuracy and objectivity of image aesthetic quality evaluation, achieving better aesthetic feature extraction and evaluation.
Smart Images

Figure CN116012300B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical fields of image processing and computer vision, and in particular to a multi-modal image aesthetic quality evaluation method that fuses local and global image features. Background Art
[0002] With the popularization of Internet technology, information such as images and videos has increased exponentially. Among them, image information is the most intuitive and contains a large amount of information. However, due to the increasing demand for beauty, the quality of image aesthetics has become the focus of people's attention. The generation of aesthetic value is the pursuit of aesthetic feelings by people in terms of vision and spirit. Evaluating images from an aesthetic perspective is an important manifestation of their development in the spiritual direction. The level of image aesthetic quality measures the visual attractiveness of an image in the eyes of humans. Therefore, people usually hope that the images they obtain have a high visual aesthetic quality. Image aesthetic quality evaluation refers to using a computer to imitate people's aesthetic process for images, enabling the computer to discover and understand the beauty of images, so as to screen out images with higher aesthetic quality. Image aesthetic quality evaluation has been applied in applications such as aesthetic-assisted image search, automatic photo enhancement, photo screening, and album management. However, the subjectivity of visual aesthetic feelings is relatively strong, often involving subjective factors such as emotions and personal tastes, which makes it a very challenging task to automatically evaluate image aesthetic quality using a computer.
[0003] Image aesthetic quality evaluation methods are generally divided into a feature extraction stage and a decision-making stage. In the feature extraction stage, features can be extracted manually or through deep learning. However, in the decision-making stage, a classifier or regression model for decision-making is trained using the aesthetic features obtained in the feature extraction stage. Therefore, image aesthetic quality evaluation methods can be divided into methods based on manually extracted features and methods based on deep learning. Methods based on manually extracted features require manually designing various image features related to aesthetic quality, and then combining effective machine learning algorithms for aesthetic classification or regression. However, manually designed features have their limitations. First, the scope of manually designed features is limited and cannot comprehensively represent aesthetic features. Second, these manually designed features are only approximations of these rules and cannot guarantee the effectiveness of these features. Summary of the Invention
[0004] In view of this, the purpose of the present invention is to provide a multi-modal image aesthetic quality evaluation method that fuses local and global image features, which can effectively fuse the local features, global features, and text aesthetic features of aesthetic images and improve the performance of image aesthetic quality evaluation algorithms.
[0005] To achieve the above purpose, the present invention adopts the following technical solutions: A multi-modal image aesthetic quality evaluation method that fuses local and global image features, comprising the following steps:
[0006] Step S1: preprocess the data in the aesthetic image dataset to obtain a fixed-size aesthetic image and a text encoding matrix of the corresponding comments, and divide the dataset into a training set and a test set;
[0007] Step S2: Design an image feature extraction subnetwork that integrates local features and global features;
[0008] Step S3: design a text feature extraction subnetwork;
[0009] Step S4: designing a multimodal image aesthetic quality evaluation network that integrates local and global image features, and using the designed network to train a multimodal image aesthetic quality score distribution prediction model that integrates local and global image features;
[0010] Step S5: Input the test image into the trained multimodal image aesthetic quality score distribution prediction model that integrates local and global image features, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.
[0011] In a preferred embodiment, the step S1 specifically includes the following steps:
[0012] Step S11: convert all words in the comment texts in the aesthetic image dataset into lowercase, and remove stop words and numbers; then use the Glove pre-trained word vector to encode all words and punctuation marks to obtain the encoding matrix of all comment texts;
[0013] Step S12: fix the size of the coding matrix of the comment text obtained in step S11 to D×K; delete the part of the coding matrix of the comment text whose number of rows exceeds D, otherwise, fill it with 0; delete the part of the coding matrix of the comment text whose number of columns exceeds K, otherwise, fill it with 0, to obtain the final coding matrix of the comment text;
[0014] Step S13: randomly crop all images in the aesthetic dataset and scale them to a fixed size H×W;
[0015] Step S14: The preprocessed images of the aesthetic dataset and the text encoding matrix of the corresponding comments are uniformly divided into a training set and a test set according to a certain ratio.
[0016] In a preferred embodiment, in step S2, an image feature extraction subnetwork integrating local features and global features is designed; the following steps are included:
[0017] Step S21: Assume that the input image of the image feature extraction subnetwork that integrates local features and global features is I in, with a dimension of 3×H×W, where H and W are the height and width of the image respectively; remove the last layer of the pre-trained ResNet50 network, and the modified network is used to extract the local features of the input image I in The output features of the last four stages of this network, the output features of the i-th stage are denoted as i = 1, 2, 3, 4, with a dimension of where and i = 1, 2, 3, 4, are the number of channels, height, and width of the feature respectively; then the feature i = 1, 2, 3, 4, is reduced in dimension through a 1×1 convolution, and the dimension after reduction is i = 1, 2, 3, 4, where c is the number of channels after reduction, and the feature after reduction is added to the randomly initialized position feature to obtain the feature and both have a dimension of i = 1, 2, 3, 4; then is dimensionally adjusted through a Reshape operation to obtain the feature with a dimension of where i = 1, 2, 3, 4; the specific calculation formula is as follows:
[0018]
[0019]
[0020] where i = 1, 2, 3, 4; Conv 1×1 (·) represents a 1×1 convolution, + represents matrix addition operation, and Reshape(·) represents dimension adjustment operation;
[0021] Step S22: Use the same input image I in as in step S21, with a dimension of 3×H×W; perform downsampling through a 32×32 convolution, and after downsampling, add it to the randomly initialized position feature P G to obtain the feature P G and both have a dimension of c×h×w, where Then is dimensionally adjusted through a Reshape operation to obtain the feature with a dimension of c×s G , where s G = h×w; the specific calculation formula is as follows:
[0022]
[0023]
[0024] Among them, Conv 32×32 (·) represents a 32×32 convolution, + represents matrix addition operation, and Reshape(·) represents a dimensionality adjustment operation;
[0025] Step S23: Construct an autoencoder SEncoder, which consists of multi-head self-attention, layer normalization, and a fully connected layer; let the input feature of the autoencoder be x, and its dimension be c×s. First, it is input into the multi-head self-attention. The output of the multi-head self-attention is added to x, and layer normalization is performed on it, denoted as to obtain the intermediate output feature x′ of the autoencoder, and then it is input into two fully connected layers, denoted as MLP s (·). The output of the two fully connected layers is added to x′ again, and layer normalization is performed on it, denoted as Finally, the output feature x″ is obtained, and its dimension is still c×s;
[0026] The formula of the autoencoder SEncoder is x″ = SEncoder(x), where SEncoder(·) represents the calculation of the autoencoder, and the specific calculation formula is as follows:
[0027]
[0028]
[0029] Among them, MHSA(·) represents multi-head self-attention, and + represents matrix addition operation;
[0030] Step S24: Construct a cross encoder CEncoder, which consists of multi-head cross-attention, layer normalization, and a fully connected layer; let the features input into the cross encoder be q and k, and the dimensions of q and k are both c×s. First, they are input into the multi-head cross-attention. The output of the multi-head cross-attention is added to q, and layer normalization is performed on it, denoted as to obtain the intermediate output feature z of the cross encoder, and then it is input into two fully connected layers, denoted as MLP c (·). The output of the two fully connected layers is added to z again, and layer normalization is performed on it, denoted as Finally, the output feature z′ is obtained, and its dimension is c×s;
[0031] The formula of the cross encoder CEncoder is z′ = CEncoder(q, k), where CEncoder(·,·) represents the calculation of the cross encoder, and the specific calculation formula is as follows:
[0032]
[0033]
[0034] Among them, MHCA(·, ·) represents multi-head cross-attention, and + represents matrix addition operation;
[0035] Step S25: The output features after step S22 Are input into the autoencoder constructed in step S23 to obtain the global image features The specific calculation formula is as follows:
[0036]
[0037] Among them, SEncoder(·) represents the autoencoder;
[0038] Step S26: Using the features after step S21 With the dimension of f = 1, 2, 3, 4, and the global image features after step S25 With the dimension of c×s C , and the intermediate process output features are used as the input of the cross-encoder constructed in step S24; specifically, first, And Are input into the cross-encoder CEncoder1 to obtain the feature Whose dimension is Second, And Are input into the cross-encoder CEncoder2 to obtain the feature Whose dimension is c×s G ; Third, And Are input into the cross-encoder CEncoder3 to obtain the feature Whose dimension is Fourth, And Are input into the cross-encoder CEncoder4 to obtain the feature Whose dimension is c×s G ; Fifth, And Are input into the cross-encoder CEncoder5 to obtain the feature Whose dimension is Sixth, And Are input into the cross-encoder CEncoder6 to obtain the feature Whose dimension is c×s G ; Seventh, And Input into the cross-encoder CEncoder7 to obtain the final image fusion features Its dimension is The specific calculation formula is as follows:
[0039]
[0040]
[0041]
[0042]
[0043]
[0044]
[0045]
[0046] Among them, CEncoder1, CEncoder2,..., CEncoder7 are all cross-encoders constructed in step S24.
[0047] In a preferred embodiment, a text feature extraction sub-network is designed. The specific steps of step S3 are as follows:
[0048] Step S31: Input the text encoding matrix obtained in step S12 into the bidirectional gated recurrent unit network to obtain the text feature F T , F T The dimension of is c×d, where c is the number of channels of F T and d is the sequence length of F T ;
[0049] Step S32: First, add the text feature F T obtained in step S31 to the randomly initialized position feature P T to obtain the feature F′ T . Finally, the feature F′ T successively passes through the two self-encoders SEncoder1 and SEncoder2 constructed in step S23 to obtain the text enhanced feature F″ T ; among them, the dimensions of P T , F′ T and F″ T are all c×d; the specific calculation formula is as follows:
[0050] F′ T =F T +P T
[0051] F″ T= SEncoder2(SEncoder1(F′ T ))
[0052] where + represents matrix addition operation.
[0053] In a preferred embodiment, in step S4, a multi-modal image aesthetic quality evaluation network that fuses local and global image features is designed, and the designed network is used to train a multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features; the method includes the following steps:
[0054] Step S41: Input each batch of images in the training set that has passed through step S1 into the image feature extraction sub-network that fuses local features and global features in step S2 to obtain the final image fusion features;
[0055] Step S42: Input the text encoding matrix obtained through step S12 corresponding to the input images of the same batch in step S41
[0056] into the text feature extraction sub-network in step S3 to obtain the final text enhancement features;
[0057] Step S43: Input the image fusion features obtained through step S41 and the text enhancement features obtained through step S42 into the cross-encoder constructed in step S24 to obtain multi-modal fusion features; finally, input the multi-modal fusion features into the fully connected layer to obtain the image aesthetic score distribution; the number of classifications output by the fully connected layer is N, and N is the number of scores in the aesthetic score set;
[0058] Step S44: Design the loss function of the multi-modal image aesthetic quality evaluation network that fuses local and global image features. The specific formula is as follows:
[0059]
[0060] where and y i respectively represent the probabilities corresponding to the i-th value of the aesthetic score in the score distribution predicted by the multi-modal image aesthetic quality evaluation network that fuses local and global image features and the true distribution of the labels. i corresponds to the aesthetic score values 1, 2,..., N, where N is the number of score values in the dataset;
[0061] Step S45: Repeat the above steps S41 to S44 in batches until the loss value calculated in step S44 converges and stabilizes, save the network parameters, and complete the training process of the multi-modal image aesthetic quality evaluation network that fuses local and global image features.
[0062] In a preferred embodiment, in step S5, the test image is input into a trained multi-modal image aesthetic quality scoring distribution prediction model that fuses local and global image features, and the corresponding image aesthetic scoring distribution is output. Finally, the average value of the aesthetic scoring distribution is calculated as the image aesthetic quality score, including the following steps:
[0063] Step S51: Input the test images in the test set into a trained multi-modal image aesthetic quality evaluation network model that fuses local and global image features, and output the corresponding image aesthetic scoring distribution p;
[0064] Step S52: Calculate the average value of the image aesthetic scoring distribution p to obtain the image aesthetic quality score score. The calculation formula is as follows:
[0065]
[0066] where represents the probability that the score in the predicted image aesthetic scoring distribution p is s i s i represents the i-th aesthetic scoring value, and N represents the number of scores.
[0067] Compared with the prior art, the present invention has the following beneficial effects: The present invention proposes a multi-modal image aesthetic quality evaluation method that fuses local and global image features, which can not only well fuse the local features and global features of aesthetic images, but also effectively make full use of and mine the text features in the user comments corresponding to the images, realizing the mutual guidance and fusion of image aesthetic features and text aesthetic features. BRIEF DESCRIPTION OF THE DRAWINGS
[0068] Figure 1 is the implementation flowchart of the method of the preferred embodiment of the present invention.
[0069] Figure 2 is the network model structure diagram in the preferred embodiment of the present invention.
[0070] Figure 3 is the structure diagram of the image feature extraction sub-network that fuses local features and global features in the preferred embodiment of the present invention.
[0071] Figure 4 is the structure diagram of the text feature extraction sub-network in the preferred embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0072] The present invention will be further described below with reference to the drawings and embodiments.
[0073] It should be noted that the following detailed description is illustrative and aims to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present application belongs.
[0074] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application; as used herein, unless the context clearly indicates otherwise, the singular forms are also intended to include the plural forms. In addition, it should also be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0075] The present invention provides a multi-modal image aesthetic quality evaluation method that fuses local and global image features. Referring to Figures 1 to 4 , it includes the following steps:
[0076] Step S1: Perform data preprocessing on the data in the aesthetic image dataset. After processing, obtain aesthetic images of a fixed size and the text encoding matrix corresponding to their comments, and divide the dataset into a training set and a test set;
[0077] Step S2: Design an image feature extraction sub-network that fuses local features and global features;
[0078] Step S3: Design a text feature extraction sub-network;
[0079] Step S4: Design a multi-modal image aesthetic quality evaluation network that fuses local and global image features, and use the designed network to train a multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features;
[0080] Step S5: Input the test image into the trained multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.
[0081] The following is the specific implementation process of the present invention.
[0082] As Figure 1 shown, a multi-modal image aesthetic quality evaluation method of the present invention that fuses local and global image features includes the following steps:
[0083] Step S1: Perform data preprocessing on the data in the aesthetic image dataset. After processing, obtain aesthetic images of a fixed size and the text encoding matrix corresponding to their comments, and divide the dataset into a training set and a test set;
[0084] Step S2: Design an image feature extraction sub-network that fuses local and global features;
[0085] Step S3: Design a text feature extraction sub-network;
[0086] Step S4: Design a multi-modal image aesthetic quality evaluation network that fuses local and global image features, and use the designed network to train a multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features;
[0087] Step S5: Input the image into the trained multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.
[0088] In this embodiment, the step S1 specifically includes the following steps:
[0089] Step S11: Convert all words in the review text in the aesthetic image dataset to lowercase, and remove stop words and numbers. Then use Glove pre-trained word vectors to encode all words and punctuation marks to obtain an encoding matrix of all review texts.
[0090] Step S12: Fix the size of the encoding matrix of the review text obtained in step S11 to D×K. Delete the part of the encoding matrix of the review text whose number of rows exceeds D, and vice versa, fill it with 0; delete the part of the encoding matrix of the review text whose number of columns exceeds K, and vice versa, fill it with 0 to obtain the final encoding matrix of the review text.
[0091] Step S13: Randomly crop and scale all images in the aesthetic dataset to a fixed size of H×W.
[0092] Step S14: Uniformly divide the preprocessed images in the aesthetic dataset and the text encoding matrix of the corresponding reviews into a training set and a test set according to a certain ratio.
[0093] In this embodiment, the step S2 specifically includes the following steps:
[0094] Step S21: Let the input image of the image feature extraction sub-network that fuses local and global features be I in , whose dimension is 3×H×W, and H and W are the height and width of the image respectively. Remove the last layer of the pre-trained ResNet50 network, and the modified network is used to extract the local features of the input image I in . For the output features of the last four stages of this network, the output feature of the i-th stage is denoted as The dimension is where and are the number of channels, height, and width of the feature respectively. Then the feature is dimensionally reduced through a 1×1 convolution, and the dimension after dimensional reduction is where c is the number of channels after dimensional reduction, and the feature after dimensional reduction is added to the randomly initialized position feature to obtain the feature and both have dimensions of Then the dimension is adjusted through a Reshape operation to obtain the feature whose dimension is where the specific calculation formula is as follows:
[0095]
[0096]
[0097] where f = 1, 2, 3, 4. Conv 1×1 (·) represents a 1×1 convolution, + represents matrix addition operation, and Reshape(·) represents dimension adjustment operation.
[0098] Step S22: The same input image I as in step S21 in (with dimensions 3×H×W) is downsampled through a 32×32 convolution, and after downsampling, it is added to the randomly initialized position feature P G to obtain the feature P G and both have dimensions of c×h×w, (where ), and then the dimension is adjusted through a Reshape operation to obtain the feature whose dimension is c×s G . The specific calculation formula is as follows:
[0099]
[0100]
[0101] where Conv 32×32 (·) represents a 32×32 convolution, + represents matrix addition operation, and Reshape(·) represents dimension adjustment operation.
[0102] Step S23: Construct an autoencoder SEncoder, which consists of multi-head self-attention, layer normalization, and fully connected layers. Let the input feature of the autoencoder be x, with its dimension being c×s. First, it is input into the multi-head self-attention. The output of the multi-head self-attention is added to x, and layer normalization is performed on it (denoted as ), to obtain the intermediate output feature x' of the autoencoder. Then, it is input into two fully connected layers (denoted as MLP s (·)). The output of the two fully connected layers is added to x' again, and layer normalization is performed on it (denoted as ), and finally, the output feature x″ is obtained, with its dimension being c×s.
[0103] The formula of the autoencoder SEncoder is x″ = SEncoder(x), where SEncoder(·) represents the calculation of the autoencoder, and the specific calculation formula is as follows:
[0104]
[0105]
[0106] Among them, MHSA(·) represents multi-head self-attention, and + represents matrix addition operation.
[0107] Step S24: Construct a cross encoder CEncoder, which consists of multi-head cross-attention, layer normalization, and fully connected layers. Let the inputs be q and k, with the dimensions of both q and k being c×s. First, they are input into the multi-head cross-attention. The output of the multi-head cross-attention is added to q, and layer normalization is performed on it (denoted as ), to obtain the intermediate output feature z of the cross encoder. Then, it is input into two fully connected layers (denoted as ), the output of the two fully connected layers is added to z again, and layer normalization is performed on it (denoted as ), and finally, the output feature z' is obtained, with its dimension being c×s.
[0108] The formula of the cross encoder CEncoder is z' = CEncoder(q, k), where CEncoder(·, ·) represents the calculation of the cross encoder, and the specific calculation formula is as follows:
[0109]
[0110]
[0111] Among them, MHCA(·, ·) represents multi-head cross-attention, and + represents matrix addition operation.
[0112] Step S25: The output feature after Step S22 Input into the autoencoder constructed in step S23 to obtain the global image features The specific calculation formula is as follows:
[0113]
[0114] Among them, SEncoder(·) represents the autoencoder.
[0115] Step S26: Use the features after step S21 (with dimension f = 1, 2, 3, 4) and the global image features after step S25 (with dimension c×s G ) and the intermediate process output features as the input of the cross-encoder constructed in step S24. Specifically, first, input and into the cross-encoder CEncoder1 to obtain the feature whose dimension is Second, input and into the cross-encoder CEncoder2 to obtain the feature whose dimension is c×s G ; Third, input and into the cross-encoder CEncoder3 to obtain the feature whose dimension is Fourth, input and into the cross-encoder CEncoder4 to obtain the feature whose dimension is c×s G ; Fifth, input and into the cross-encoder CEncoder5 to obtain the feature whose dimension is Sixth, input and into the cross-encoder CEncoder6 to obtain the feature whose dimension is c×s G ; Seventh, input and into the cross-encoder CEncoder7 to obtain the final image fusion feature whose dimension is The specific calculation formula is as follows:
[0116]
[0117]
[0118]
[0119]
[0120]
[0121]
[0122]
[0123] Among them, CEncoder1, CEncoder2, ..., CEncoder7 are all cross-encoders constructed in step S24.
[0124] In this embodiment, step S3 specifically includes the following steps:
[0125] Step S31: Input the text encoding matrix obtained in step S12 into a Bidirectional Gate Recurrent Unit (BiGRU) network to obtain text feature F T , F T The dimension of F is c×d, where c is the number of channels of F T and d is the sequence length of F T .
[0126] Step S32: First, add the text feature F T obtained in step S31 to the randomly initialized position feature P T to obtain feature F′ T . Finally, feature F′ T passes through the autoencoders SEncoder1 and SEncoder2 of two step S23 in sequence to obtain text enhancement feature F″ T . Among them, the dimensions of P T , F′ T and F″ T are all c×D. The specific calculation formula is as follows:
[0127] F′ T = F T + P T
[0128] F″ T = SEncoder2(SEncoder1(F′ T ))
[0129] Among them, + represents matrix addition operation.
[0130] In this embodiment, step S4 specifically includes the following steps:
[0131] Step S41: Input each batch of images in the training set that has gone through Step S1 into the image feature extraction sub-network that fuses local and global features in Step S2 to obtain the final image fusion features.
[0132] Step S42: Input the text encoding matrix obtained through Step S12 corresponding to the input images of the same batch in Step S41 into the text feature extraction sub-network in Step S3 to obtain the final text enhancement features.
[0133] Step S43: Input the image fusion features obtained through Step S41 and the text enhancement features obtained through Step S42 into the cross-encoder constructed in Step S24 to obtain multi-modal fusion features. Finally, input the multi-modal fusion features into the fully connected layer to obtain the image aesthetics score distribution. The number of classifications output by the fully connected layer is N, where N is the number of scores in the aesthetics score set. For example, when the score set is {1, 2,..., 10}, N is 10.
[0134] Step S44: Design the loss function of the multi-modal image aesthetics quality evaluation network that fuses local and global image features. The specific formula is as follows:
[0135]
[0136] where and y i respectively represent the probabilities corresponding to the i-th value of the aesthetics score in the predicted score distribution and the true distribution of the labels of the multi-modal image aesthetics quality evaluation network that fuses local and global image features. i corresponds to the aesthetics score values 1, 2,..., N, and N is the number of score values in the dataset.
[0137] Step S45: Repeat the above Steps S41 to S44 in batches until the loss value calculated in Step S44 converges and stabilizes. Save the network parameters to complete the training process of the multi-modal image aesthetics quality evaluation network that fuses local and global image features.
[0138] In this embodiment, Step S5 specifically includes the following steps:
[0139] Step S51: Input the test images in the test set into the trained multi-modal image aesthetics quality evaluation network model that fuses local and global image features, and output the corresponding image aesthetics score distribution p.
[0140] Step S52: Calculate the average value of the image aesthetics score distribution p to obtain the image aesthetics quality score score. The calculation formula is as follows:
[0141]
[0142] Among them, represents the probability that the score in the predicted image aesthetic score distribution p is s i , s i represents the i-th score, and N represents the number of scores.
Claims
1. A multi-modal image aesthetic quality evaluation method that fuses local and global image features, characterized in that, The following steps are involved: Step S1: preprocess the data in the aesthetic image dataset to obtain a fixed-size aesthetic image and a text encoding matrix of the corresponding comments, and divide the dataset into a training set and a test set; Step S2: Design an image feature extraction subnetwork that integrates local features and global features; Step S3: design a text feature extraction subnetwork; Step S4: designing a multimodal image aesthetic quality evaluation network that integrates local and global image features, and using the designed network to train a multimodal image aesthetic quality score distribution prediction model that integrates local and global image features; Step S5: input the test image into the trained multimodal image aesthetic quality score distribution prediction model that integrates local and global image features, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score; Step 4 includes: Step S44: designing a loss function of a multimodal image aesthetic quality evaluation network that integrates local and global image features. The specific formula is as follows: wherein, and y i respectively represent the probabilities corresponding to the i-th value of the aesthetic score in the scoring distribution predicted by the multi-modal image aesthetic quality evaluation network that fuses local and global image features and the true distribution of labels, where i corresponds to the aesthetic score values 1, 2,..., N, and N is the number of score values in the dataset; In the step S2, an image feature extraction subnetwork integrating local features and global features is designed; the steps include: Step S21: Set the input image of the image feature extraction sub-network that fuses local features and global features as I in , whose dimension is 3×H×W, where H and W are the height and width of the image respectively; remove the last layer of the pre-trained ResNet50 network, and the modified network is used to extract the local features of the input image I in . For the output features of the last four stages of this network, the output feature of the i-th stage is denoted as with the dimension of where and are the number of channels, height and width of the feature respectively; then the feature is dimensionally reduced through a 1×1 convolution, and the dimension after dimensional reduction is where c is the number of channels after dimensional reduction. The feature after dimensional reduction is added to the randomly initialized position feature to obtain the feature and both have the dimension of Then undergoes a Reshape operation for dimension adjustment to obtain the feature with the dimension of where The specific calculation formula is as follows: where \(i = 1, 2, 3, 4\); Conv 1×1 (·) represents a 1×1 convolution, + represents matrix addition operation, and Reshape(·) represents a dimension adjustment operation; Step S22: Use the same input image I as in Step S21 in , with a dimension of 3×H×W; perform downsampling through a 32×32 convolution, and after downsampling, add it to the randomly initialized position feature P G to obtain the feature P G and both have a dimension of c×h×w, where Then perform a Reshape operation to adjust the dimension to obtain the feature with a dimension of c×s G , where s G =h×w; the specific calculation formula is as follows: Among them, Conv 32×32 (·) represents a 32×32 convolution, + represents matrix addition operation, and Reshape(·) represents a dimension adjustment operation; Step S23: Construct an autoencoder SEncoder, which consists of multi-head self-attention, layer normalization, and fully connected layers; let the input feature of the autoencoder be x, whose dimension is c×s. First, it is input into the multi-head self-attention. The output of the multi-head self-attention is added to x, and layer normalization is performed on it, denoted as to obtain the intermediate output feature x′ of the autoencoder, and then it is input into two fully connected layers, denoted as MLP s (·). The output of the two fully connected layers is added to x′, and layer normalization is performed on it, denoted as Finally, the output feature x” is obtained, whose dimension is still c×s; The formula of the autoencoder SEncoder is x'=SEncoder(x), where SEncoder(·) represents the calculation of the autoencoder. The specific calculation formula is as follows: Among them, MHSA(·) represents multi-head self-attention, + represents matrix addition operation; Step S24: Construct a cross encoder CEncoder, which consists of multi-head cross attention, layer normalization, and a fully connected layer; assume the features input to the cross encoder are q and k, and the dimensions of both q and k are c×s. First, they are input to the multi-head cross attention. The output of the multi-head cross attention is added to q, and layer normalization is performed on it, denoted as to obtain the intermediate output feature z of the cross encoder, which is then input into two fully connected layers, denoted as MLP c (·). The output of the two fully connected layers is added to z, and layer normalization is performed on it, denoted as Finally, the output feature z′ is obtained, and its dimension is c×s; The formula of the cross encoder CEncoder is z′=CEncoder(q,k), where CEncoder(·,·) represents the calculation of the cross encoder. The specific calculation formula is as follows: Among them, MHCA(·,·) represents multi-head cross attention, + represents matrix addition operation; Step S25: The output features after Step S22 are input into the autoencoder constructed in Step S23 to obtain the global image features The specific calculation formula is as follows: Where, SEncoder(·) represents the autoencoder; Step S26: Using the feature passed through step S21 with dimension and the global image feature passed through step S25 with dimension c×s G , and the intermediate process output feature as the input of the cross-encoder constructed in step S24; specifically, first, input and into the cross-encoder CEncoder1 to obtain a feature whose dimension is Second, input and into the cross-encoder CEncoder2 to obtain a feature whose dimension is c×s G ; Third, input and into the cross-encoder CEncoder3 to obtain a feature whose dimension is Fourth, input and into the cross-encoder CEncoder4 to obtain a feature whose dimension is c×s G ; Fifth, input and into the cross-encoder CEncoder5 to obtain a feature whose dimension is Sixth, input and into the cross-encoder CEncoder6 to obtain a feature whose dimension is c×s G ; Seventh, input and into the cross-encoder CEncoder7 to obtain the final image fusion feature whose dimension is The specific calculation formula is as follows: Among them, CEncoder1, CEncoder2, ..., CEncoder7 are all cross encoders constructed in step S24.
2. The multi-modal image aesthetic quality evaluation method for fusing local and global image features according to claim 1, characterized in that The step S1 specifically includes the following steps: Step S11: convert all words in the comment texts in the aesthetic image dataset into lowercase, and remove stop words and numbers; then use the Glove pre-trained word vector to encode all words and punctuation marks to obtain the encoding matrix of all comment texts; Step S12: fix the size of the coding matrix of the comment text obtained in step S11 to D×K; delete the part of the coding matrix of the comment text whose number of rows exceeds D, otherwise, fill it with 0; delete the part of the coding matrix of the comment text whose number of columns exceeds K, otherwise, fill it with 0, to obtain the final coding matrix of the comment text; Step S13: randomly crop all images in the aesthetic dataset and scale them to a fixed size H×W; Step S14: The preprocessed images of the aesthetic dataset and the text encoding matrix of the corresponding comments are uniformly divided into a training set and a test set according to a certain ratio.
3. The multi-modal image aesthetic quality evaluation method for fusing local and global image features according to claim 1, characterized in that Design a text feature extraction subnetwork, and the step S3 specifically includes the following steps: Step S31: Input the text encoding matrix obtained in Step S12 into a bidirectional gated recurrent unit network to obtain text feature F T , F T has a dimension of c×d, where c is the number of channels of F T and d is the sequence length of F T ; Step S32: First, add the text feature F obtained in step S31 T to the randomly initialized position feature P T to obtain the feature F'. T Finally, the feature F' T successively passes through the autoencoders SEncoder1 and SEncoder2 constructed in two steps S23 to obtain the text enhancement feature F'' T ; where the dimensions of P T , F' T and F'' T are all c×d; the specific calculation formula is as follows: F′ T = F T + P T F″ T = SEncoder2(SEncoder1(F′ T )) Among them, + represents the matrix addition operation.
4. The multi-modal image aesthetic quality evaluation method for fusing local and global image features according to claim 1, wherein In step S4, a multi-modal image aesthetic quality evaluation network that fuses local and global image features is designed, and the designed network is used to train a multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features; the method includes the following steps: Step S41: Input each batch of images in the training set that has gone through step S1 into the image feature extraction sub-network that fuses local features and global features in step S2 to obtain the final image fusion features; Step S42: Input the text encoding matrix obtained through step S12 corresponding to the input images of the same batch in step S41 into the text feature extraction sub-network in step S3 to obtain the final text enhanced features; Step S43: Input the image fusion features obtained through step S41 and the text enhanced features obtained through step S42 into the cross encoder constructed in step S24 to obtain multi-modal fusion features; finally, input the multi-modal fusion features into the fully connected layer to obtain the image aesthetic score distribution; the number of classifications output by the fully connected layer is N; Step S45: Repeat the above steps S41 to S44 in batches until the loss value calculated in step S44 converges and stabilizes, save the network parameters, and complete the training process of the multi-modal image aesthetic quality evaluation network that fuses local and global image features.
5. The multi-modal image aesthetic quality evaluation method for fusing local and global image features according to claim 1, characterized in that In step S5, input the test image into the trained multi-modal image aesthetic quality score distribution prediction model that fuses local and global image features, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score; the method includes the following steps: Step S51: Input the test images in the test set into the trained multi-modal image aesthetic quality evaluation network model that fuses local and global image features, and output the corresponding image aesthetic score distribution p; Step S52: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score score; the calculation formula is as follows: Among them, represents the probability that the score in the predicted image aesthetics score distribution p is s i , s i represents the i-th aesthetic score value, and N represents the number of score values in the dataset.
Citation Information
Patent Citations
Image aesthetic quality evaluation method fused with multi-modal attention mechanism
CN113657380A
Quantifying and visualizing changes over time to health and wellness
WO2022169886A1