Image aesthetic quality evaluation method and system integrating scene features and multimodal attention mechanism

Through the image aesthetic quality evaluation method that integrates scene features and multimodal attention mechanism, the problem of difficult to effectively integrate image scene features and aesthetic features in the prior art is solved, and more efficient image aesthetic quality evaluation performance is achieved.

CN115908979BActive Publication Date: 2025-05-09FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202211503996.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-11-28
Publication Date
2025-05-09
Estimated Expiration
2042-11-28

AI Technical Summary

Technical Problem

The prior art is difficult to effectively integrate scene features and aesthetic features in images, resulting in insufficient performance in image aesthetic quality evaluation.

Method used

A method of image aesthetic quality evaluation is proposed that integrates scene features and multimodal attention mechanisms. Through the hierarchical image feature fusion module and multimodal attention mechanism module, the image scene features and aesthetic features are fused to generate image aesthetic score distribution.

Benefits of technology

Effectively utilize text features in user comments and fuse image scene features with aesthetic features, improving the performance of image aesthetic quality evaluation algorithm.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115908979B_ABST
    Figure CN115908979B_ABST
Patent Text Reader

Abstract

The present invention proposes an image aesthetic quality evaluation method and system integrating scene features and a multimodal attention mechanism, comprising the following steps: step S1: performing data preprocessing on data in an aesthetic image data set to extract text features of comments corresponding to the aesthetic image; step S2: training an image aesthetic score distribution prediction model integrating scene features and a multimodal attention mechanism; the image aesthetic score distribution prediction model is obtained by training an image aesthetic quality evaluation network integrating scene features and a multimodal attention mechanism, and comprises a hierarchical image feature fusion module and a multimodal attention mechanism module integrating text features and image features; step S3: inputting an image into the trained image aesthetic quality score distribution prediction model integrating scene features and a multimodal attention mechanism, outputting a corresponding image aesthetic score distribution, and finally calculating an average value of the aesthetic score distribution as an image aesthetic quality score.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of image processing and computer vision technology, and specifically relates to an image aesthetic quality evaluation method and system integrating scene features and a multimodal attention mechanism. Background Art

[0002] With the popularization of Internet technology, information such as images and videos is increasing rapidly. Among them, image information is the most intuitive and contains a large amount of information. However, due to the increasing demand for beauty, the quality of image aesthetics has become the focus of people's attention. The emergence of aesthetic value is people's pursuit of aesthetic feelings in visual and spiritual aspects. Evaluating images from an aesthetic perspective is an important manifestation of developing them in a spiritual direction. The level of image aesthetic quality measures the strength of an image's visual appeal in the eyes of humans. Therefore, people usually hope that the images they obtain have high visual aesthetic quality. Image aesthetic quality evaluation refers to the use of computers to imitate people's aesthetic process of images, so that computers can discover the beauty of images and understand the beauty of images, thereby screening out images with high aesthetic quality. Image aesthetic quality evaluation has been applied in applications such as aesthetic-assisted image search, automatic photo enhancement, photo screening, and album management. However, visual aesthetic feelings are highly subjective, and they often involve subjective factors such as emotions and personal tastes, which makes it a very challenging task to use computers to automatically evaluate image aesthetic quality.

[0003] Image aesthetic quality evaluation methods are generally divided into feature extraction stage and decision-making stage. In the feature extraction stage, two methods can be used: manual feature extraction and deep learning. However, the decision-making stage uses the aesthetic features obtained in the feature extraction stage to train a classifier or regression model for decision-making. Therefore, image aesthetic quality evaluation methods can be divided into methods based on manual feature extraction and methods based on deep learning. The method based on manual feature extraction requires manual design of multiple image features related to aesthetic quality, and then combines effective machine learning algorithms for aesthetic classification or regression. However, manually designed features have their limitations. First, the range of manually designed features is limited and cannot fully represent aesthetic features; second, these manually designed features are only approximations of these rules, and the validity of these features cannot be guaranteed.

[0004] At present, the most advanced image aesthetic quality evaluation methods are based on deep learning. Deep learning has a powerful automatic feature learning ability, does not require people to have rich knowledge of image aesthetics and psychology, and is much more efficient and accurate than methods based on manual feature extraction. Therefore, more and more researchers have successfully solved the problem of image aesthetic quality evaluation using deep convolutional neural networks and achieved good performance. However, most of the image aesthetic quality evaluation methods based on deep learning are currently limited to learning image aesthetic features. Summary of the invention

[0005] Considering that the user comments corresponding to the images in the aesthetic data set often explain their reasons for evaluating the image quality, this contains important information related to the image and can be used to assist in the evaluation of the image aesthetic quality. At the same time, it is found that when people judge the aesthetic quality of an image, they will be affected by the scene factors in the image. For example, when the image scene is the sky, people will evaluate the image as having a higher aesthetic quality; and when the image scene is a desert, people will evaluate the image as having a lower aesthetic quality. Therefore, the purpose of this invention is to fuse the scene features in the image with the aesthetic features to better evaluate the aesthetic quality of the image.

[0006] In summary, in order to better improve the performance of the image aesthetic quality evaluation method, the present invention proposes an image aesthetic quality evaluation method and system that integrates scene features and multimodal attention mechanism, which can not only effectively and fully utilize and mine the text features in the user comments corresponding to the image, but also can well integrate the image scene features with the image aesthetic features, and realize the mutual guidance and integration of visual features and aesthetic text features. It can effectively integrate the scene features in the image and the aesthetic features reflected in the text, and improve the performance of the image aesthetic quality evaluation algorithm.

[0007] The technical solution adopted by the present invention to solve the technical problem is:

[0008] A method for evaluating image aesthetic quality by integrating scene features and multimodal attention mechanism, characterized in that it comprises the following steps:

[0009] Step S1: preprocess the data in the aesthetic image dataset, extract the text features of the comments corresponding to the aesthetic images, and divide the dataset into a training set and a test set;

[0010] Step S2: training an image aesthetic score distribution prediction model integrating scene features and a multimodal attention mechanism; the image aesthetic score distribution prediction model is obtained by training an image aesthetic quality evaluation network integrating scene features and a multimodal attention mechanism, including a hierarchical image feature fusion module and a multimodal attention mechanism module integrating text features and image features;

[0011] Step S3: Input the image into the trained image aesthetic quality score distribution prediction model that integrates scene features and multimodal attention mechanism, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.

[0012] Furthermore, step S1 specifically includes the following steps:

[0013] Step S11: performing data cleaning on the comment text dataset corresponding to the aesthetic image, and deleting the comment text that is empty, has only one word, or is all numbers;

[0014] Step S12: convert all words in the comment text obtained in step S11 into lowercase, and remove stop words and numbers; then use the GloVe pre-trained word vector to encode all words and punctuation marks to obtain the encoding matrix of all comment texts;

[0015] Step S13: fix the size of the coding matrix of the comment text obtained in step S12 to D×L; delete the part of the coding matrix of the comment text whose height exceeds D, otherwise, fill it with 0; delete the part of the coding matrix of the comment text whose width exceeds L, otherwise, fill it with 0, to obtain the final coding matrix of the comment text;

[0016] Step S14: input the text encoding matrix obtained in step S13 into the bidirectional gate controlled recurrent unit network to obtain text features with a size of C×D;

[0017] Step S15: randomly crop all images in the aesthetic dataset and scale them to a fixed size H×W;

[0018] Step S16: Divide the preprocessed images of the aesthetic dataset and the corresponding comment text features into a training set and a test set.

[0019] Furthermore, the hierarchical image feature fusion module is composed of a structure R for changing feature dimensions and matrix calculation;

[0020] The structure R consists of two 1×1 convolutions and one 3×3 convolution, with input image features X and dimension C x ×H x ×W x ,The feature X reduces the feature width and height and increases the number of channels through the structure R. The calculation formula is as follows:

[0021] f=w1(X)+b1

[0022] f′=w2(f)+b2

[0023] f″=w3(f′)+b3

[0024] Among them, w1 and b1 are the weight and bias of the first 1×1 convolutional layer, and the dimension of feature f is w2 and b2 are the weights and biases of the 3×3 convolutional layer, and the dimension of feature f′ is w3 and b3 are the weights and biases of the second 1×1 convolutional layer, and the dimension of feature f” is

[0025] Assume that the input feature of the i-th level of the hierarchical image feature fusion module is F i , i=1,2,3,4, dimension is C i ×H i ×W i ,in C i+1 =2C i ;

[0026] First, feature F1 is added to feature F2 after passing through structure R to obtain feature The dimensions are C2×H2×W2, After passing through structure R, it is added to feature F3 to obtain feature The dimensions are C3×H3×W3, After the structure R, the feature F′1 is obtained, with the dimension of C4×H4×W4;

[0027] Secondly, feature F2 is added to feature F3 after passing through structure R to obtain feature The dimensions are C3×H3×W3, After the structure R, the feature F′2 is obtained, with the dimension of C4×H4×W4;

[0028] Again, feature F3 is transformed into feature F′3 after structure R, with the dimension of C4×H4×W4;

[0029] Finally, features F′1, F′2, F′3, and F4 are added to obtain the image fusion feature F, whose dimension is C4×H4×W4. The specific calculation formula of the above process is as follows:

[0030]

[0031] F3=R(F3)

[0032] F=F′1+F′2+F′3+F4.

[0033] Furthermore, in the multimodal attention mechanism module that integrates text features and image features:

[0034] The image feature F with the dimension C×H×W I Input into two 1×1 convolutional layers respectively to obtain the key point feature K and image projection feature V, and the dimension remains unchanged. The specific formula is as follows:

[0035] K=w1(F I )+b1

[0036] V=w2(F I )+b2

[0037] Among them, F Iis the input image feature; w1 and b1 are the weight and bias of the 1×1 convolution layer used to extract the key point feature K; w2 and b2 are the weight and bias of the 1×1 convolution layer used to extract the image projection feature V;

[0038] Then adjust the dimensions of the key point feature K and the image projection feature V, both of which have dimensions C×H×W, to heads×c×H×W, where C=heads×c;

[0039] Step S33: First, the text feature F with a dimension of C×D obtained in step S1 is T After the dimension is adjusted, the adjusted dimension is C×H×W, and then after a 1×1 convolution layer, the dimension is adjusted again. The adjusted dimension is heads×c×H×W to obtain the text feature Q, where D=H×W. The specific formula is as follows:

[0040] Q = reshape(w3(reshape(F T ))+b3)

[0041] Among them, F T is the input text feature, reshape(·) represents the dimension adjustment operation, w3 and b3 are the weight and bias of the 1×1 convolution layer used to extract the text feature Q;

[0042] Then randomly initialize the position feature P, whose dimension is heads×c×H×W;

[0043] The obtained key point feature K, image projection feature V, text feature Q, and S position feature P are calculated through activation functions and multiple matrix operations to obtain the multimodal fusion feature Z. The specific formula is as follows:

[0044] z=Softmax((Q+P)×K T )×V

[0045] Where T represents the transpose of the matrix, Softmax(·) represents the Softmax activation function, + represents the matrix addition operation, and × represents the matrix multiplication operation;

[0046] Finally, the multimodal fusion feature Z with the dimension of heads×c×H×W is dimensionally adjusted to C×H×W, where C=heads×c.

[0047] Furthermore, the process of training the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism to obtain the image aesthetic score distribution prediction model integrating scene features and multimodal attention mechanism comprises the following steps:

[0048] Step S21: using two pre-trained ResNet50s as feature extraction subnetworks, removing the last layers of the two ResNet50 networks respectively as image scene feature extraction subnetwork and image aesthetic feature extraction subnetwork; the image scene feature extraction subnetwork uses weights pre-trained on the Places365 dataset as initial parameters, and the image aesthetic feature extraction subnetwork uses weights pre-trained on the ImageNet dataset as initial parameters;

[0049] Step S22: Input each batch of images in the training set of step S1 into the two sub-networks in step S21; suppose that the image scene features and image aesthetic features of the i-th corresponding layer output by the last four corresponding layers of the two ResNet50 networks are and i=1,2,3,4; first and Feature concatenation is performed according to the channel dimension, and then dimensionality reduction is performed through 1×1 convolution. The specific formula is as follows:

[0050]

[0051] F′ i =w i (F i )+b i

[0052] Where i = 1, 2, 3, 4, and They are the image scene features and image aesthetic features output by the i-th corresponding layer, and their dimensions are both C i ×H i ×W i ; Concat(·) means that features are concatenated according to the channel dimension, F i yes and The output feature after concatenation has a dimension of 2C i ×H i ×W i ;w i and b i is the weight and bias of the 1×1 convolutional layer used by the i-th corresponding layer; F′ i Yes F i The output feature after the 1×1 convolution layer has a dimension of C i ×H i ×W i ;

[0053] Step S23: Output feature F′ obtained in step S22 i, i=1,2,3,4 are input into the hierarchical image feature fusion module to obtain image fusion features, and then the image fusion features and the text features corresponding to the same batch of images after step S1 are input into the multimodal attention mechanism module to obtain multimodal fusion features; finally, the multimodal fusion features are input into the fully connected layer to obtain the image aesthetic score distribution; the number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set;

[0054] The specific formula of the network's loss function during training is as follows:

[0055]

[0056] in, and i They represent the probability corresponding to the i-th value of the aesthetic score in the score distribution predicted by the image aesthetic quality evaluation network that integrates scene features and multimodal attention mechanism and the true distribution of the label, respectively. i corresponds to the aesthetic score value of 1, 2, ...N, and N is the number of score values ​​in the dataset;

[0057] Step S24: Repeat the above steps S21 to S23 in batches until the loss value calculated in step S23 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism.

[0058] Furthermore, step S3 specifically includes the following steps:

[0059] Step S31: input the images in the test set into the trained image aesthetic quality evaluation network model integrating scene features and multimodal attention mechanism, and output the corresponding image aesthetic score distribution p;

[0060] Step S32: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score score; the calculation formula is as follows:

[0061]

[0062] in, Indicates a rating of s i The probability of i Represents the value of the i-th aesthetic score.

[0063] And, an image aesthetic quality evaluation system integrating scene features and multimodal attention mechanism, characterized in that it is based on a computer system and includes:

[0064] Preprocessing module: used to preprocess the data in the aesthetic image dataset, extract the text features of the comments corresponding to the aesthetic images, and divide the dataset into a training set and a test set;

[0065] Evaluation network: an image aesthetic score distribution prediction model obtained by training a fusion scene feature and a multimodal attention mechanism; the image aesthetic score distribution prediction model is obtained by training an image aesthetic quality evaluation network that integrates scene features and a multimodal attention mechanism, including a hierarchical image feature fusion module and a multimodal attention mechanism module that integrates text features and image features;

[0066] Scoring model: It is used to input the image into the trained image aesthetic quality score distribution prediction model that integrates scene features and multimodal attention mechanism, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.

[0067] And, an image aesthetic quality evaluation system that integrates scene features and multimodal attention mechanism, characterized in that it includes: a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that when the processor executes the computer program, it implements the image aesthetic quality evaluation method that integrates scene features and multimodal attention mechanism as described above.

[0068] And, a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: when the computer program is executed by a processor, the image aesthetic quality evaluation method integrating scene features and a multimodal attention mechanism as described above is implemented.

[0069] Compared with the prior art, the present invention and its preferred solution can predict the distribution of image aesthetic scores and improve the performance of image aesthetic quality assessment algorithms. BRIEF DESCRIPTION OF THE DRAWINGS

[0070] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0071] Figure 1 It is a design and implementation flow chart of the embodiment of the present invention.

[0072] Figure 2 It is a structural diagram of a network model in an embodiment of the present invention.

[0073] Figure 3 It is a structural diagram of the hierarchical image feature fusion module in an embodiment of the present invention.

[0074] Figure 4 It is a structural diagram of the multimodal attention mechanism module in an embodiment of the present invention. DETAILED DESCRIPTION

[0075] In order to make the features and advantages of this patent more obvious and easy to understand, the following embodiments are specifically described in detail as follows:

[0076] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in this specification have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.

[0077] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.

[0078] like Figure 1-Figure 4 As shown, the present invention provides an overall design process and implementation process of an image aesthetic quality evaluation method integrating scene features and a multimodal attention mechanism, comprising the following steps:

[0079] Step 1: Preprocess the data in the aesthetic image dataset, extract the text features of the comments corresponding to the aesthetic images, and divide the dataset into a training set and a test set;

[0080] Step 2: Design a hierarchical image feature fusion module;

[0081] Step 3: Design a multimodal attention mechanism module that integrates text features and image features;

[0082] Step 4: Design an image aesthetic quality evaluation network that integrates scene features and multimodal attention mechanism, and use the designed network to train an image aesthetic score distribution prediction model that integrates scene features and multimodal attention mechanism;

[0083] Step 5: Input the image into the trained image aesthetic quality score distribution prediction model that integrates scene features and multimodal attention mechanism, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.

[0084] The following is the specific implementation process of the present invention.

[0085] like Figure 1 As shown, the present invention provides an image aesthetic quality evaluation method integrating scene features and multimodal attention mechanism, comprising the following steps:

[0086] Step 1: Preprocess the data in the aesthetic image dataset, extract the text features of the comments corresponding to the aesthetic images, and divide the dataset into a training set and a test set;

[0087] Step 2: Design a hierarchical image feature fusion module;

[0088] Step 3: Design a multimodal attention mechanism module that integrates text features and image features;

[0089] Step 4: Design an image aesthetic quality evaluation network that integrates scene features and multimodal attention mechanism, and use the designed network to train an image aesthetic score distribution prediction model that integrates scene features and multimodal attention mechanism;

[0090] Step 5: Input the image into the trained image aesthetic quality score distribution prediction model that integrates scene features and multimodal attention mechanism, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.

[0091] In this embodiment, step 1 specifically includes the following steps:

[0092] Step 11: Clean the comment text dataset corresponding to the aesthetic images and delete the comment texts that are empty, have only one word, or are all numbers;

[0093] Step 12: Convert all words in the comment text obtained in step 11 to lowercase, and remove stop words and numbers. Then use the Glove pre-trained word vector to encode all words and punctuation marks to obtain the encoding matrix of all comment texts.

[0094] Step 13: Fix the size of the encoding matrix of the comment text obtained in step 12 to D × L. Delete the part of the encoding matrix of the comment text whose height exceeds D, otherwise fill it with 0; delete the part of the encoding matrix of the comment text whose width exceeds L, otherwise fill it with 0, and get the final comment text encoding matrix.

[0095] Step 14: Input the text encoding matrix obtained in step 13 into the Bidirectional Gate Recurrent Unit (BiGRU) network to obtain text features with a size of C×D.

[0096] Step 15: All images in the aesthetic dataset are randomly cropped and scaled to a fixed size H×W.

[0097] Step 16: The preprocessed images of the aesthetic dataset and the corresponding comment text features are uniformly divided into training sets and test sets according to a certain ratio.

[0098] In this embodiment, step 2 specifically includes the following steps:

[0099] Step 21: Design a structure R that changes the feature dimension. The structure R consists of two 1×1 convolutions and one 3×3 convolution. The input image feature X has a dimension of C. x ×H x ×W x ,The feature X reduces the feature width and height and increases the number of channels through the structure R. The calculation formula is as follows:

[0100] f=w1(X)+b1

[0101] f′=w2(f)+b2

[0102] f″=w3(f′)+b3

[0103] Among them, w1 and b1 are the weight and bias of the first 1×1 convolutional layer, and the dimension of feature f is w2 and b2 are the weights and biases of the 3×3 convolutional layer, and the dimension of feature f′ is w3 and b3 are the weights and biases of the second 1×1 convolutional layer, and the dimension of feature f” is

[0104] Step 22: The hierarchical image feature fusion module consists of multiple structures R and matrix calculations designed in step 21. Suppose the input feature of the hierarchical image feature fusion module is F i (i=1,2,3,4) dimension is C i ×H i ×W i ,in C i+1 =2C i .

[0105] First, feature F1 is added to feature F2 after passing through structure R to obtain feature (dimensions are C2×H2×W2), After passing through structure R, it is added to feature F3 to obtain feature (dimensions are C3×H3×W3), After the structure R, the feature F′1 (dimension is C4×H4×W4) is obtained.

[0106] Secondly, feature F2 is added to feature F3 after passing through structure R to obtain feature (dimensions are C3×H3×W3), After the structure R, the feature F′2 (dimension is C4×H4×W4) is obtained.

[0107] Again, feature F3 passes through structure R to obtain feature F′3 (with dimensions of C4×H4×W4).

[0108] Finally, features F′1, F′2, F′3, and F4 are added together to obtain the image fusion feature F, whose dimension is C4×H4×W4.

[0109] The specific calculation formula for the above process is as follows:

[0110]

[0111] F′3=R(F3)

[0112] F=F′1+F′2+F′3+F4

[0113] In this embodiment, step 3 specifically includes the following steps:

[0114] Step 31: Transform the image feature F of dimension C×H×W I Input into two 1×1 convolutional layers respectively to obtain the key point feature K and image projection feature V, and the dimension remains unchanged. The specific formula is as follows:

[0115] K=w1(F I )+b1

[0116] V=w2(F I )+b2

[0117] Among them, F I is the input image feature. w1 and b1 are the weight and bias of the 1×1 convolution layer used to extract the key point feature K. w2 and b2 are the weight and bias of the 1×1 convolution layer used to extract the image projection feature V.

[0118] Step 32: The key point feature K and the image projection feature V obtained in step 31, both of which have dimensions of C×H×W, are dimensionally adjusted to heads×c×H×W, where C=heads×c.

[0119] Step 33: First, the text feature F with dimension C×D obtained in step 14 is T After the dimension is adjusted, the adjusted dimension is C×H×W, and then after a 1×1 convolution layer, the dimension is adjusted again, and the adjusted dimension is heads×c×H×W to obtain the text feature Q, where D=H×W and C=heads×c. The specific formula is as follows:

[0120] Q = reshape(w3(reshape(F T ))+b3)

[0121] Among them, F T is the input text feature, reshape(·) represents the dimension adjustment operation, w3 and b3 are the weights and bias of the 1×1 convolutional layer used to extract the text feature Q.

[0122] Step 34: Randomly initialize the position feature P with the dimension of heads×c×H×W.

[0123] Step 35: The key point feature K and image projection feature V obtained in step 31, the text feature Q obtained in step 32, and the position feature P obtained in step 33 are calculated through activation functions and multiple matrix operations to obtain a multimodal fusion feature Z. The specific formula is as follows:

[0124] Z = Softmax((Q+P)×K T )×V

[0125] Among them, T represents the transpose of the matrix, Softmax(·) represents the Softmax activation function, + represents the matrix addition operation, and × represents the matrix multiplication operation.

[0126] Finally, the multimodal fusion feature Z with the dimension of heads×c×H×W is adjusted to C×H×W, where C=heads×c.

[0127] In this embodiment, step 4 specifically includes the following steps:

[0128] Step 41: Use two pre-trained ReNet50 as feature extraction subnetworks, remove the last layer of the two ReNet50 networks and use them as image scene feature extraction subnetwork and image aesthetic feature extraction subnetwork. The image scene feature extraction subnetwork uses the weights pre-trained on the Place365 dataset as initial parameters, and the image aesthetic feature extraction subnetwork uses the weights pre-trained on the ImageNet dataset as initial parameters.

[0129] Step 42: Input each batch of images in the training set in step 1 into the two sub-networks in step 41. Suppose the image scene features and image aesthetic features of the i-th corresponding layer output by the last four corresponding layers of the two ReNet50 networks are and First, and The features are concatenated according to the channel dimension, and then the dimension is reduced by 1×1 convolution. The specific formula is as follows:

[0130]

[0131] F′ i =w i (F i )+b i

[0132] Among them, i=1,2,3,4. and They are the image scene features and image aesthetic features output by the i-th corresponding layer, and their dimensions are both C i ×H i ×W i Concat(·) means that the features are concatenated according to the channel dimension. i yes and The output feature after concatenation has a dimension of 2C i ×H i ×W i .w i and b i are the weights and biases of the 1×1 convolutional layer used by the i-th corresponding layer. i Yes F i The output feature after the 1×1 convolution layer has a dimension of C i ×H i ×W i .

[0133] Step 43: Output feature F′ obtained in step 42 i (i=1,2,3,4) is input into the hierarchical image feature fusion module designed in step 2 to obtain the image fusion feature, and then the image fusion feature and the text features corresponding to the same batch of images after step 1 are input into the multimodal attention mechanism module designed in step 3 to obtain the multimodal fusion feature. Finally, the multimodal fusion feature is input into the fully connected layer to obtain the image aesthetic score distribution. The number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set. For example, when the score set is {1, 2, …, 10}, N is 10.

[0134] Step 44: Design a loss function for the image aesthetic quality evaluation network that integrates scene features and multimodal attention mechanism. The specific formula is as follows:

[0135]

[0136] in, and i They respectively represent the probability corresponding to the i-th aesthetic score value in the score distribution predicted by the image aesthetic quality assessment network that integrates scene features and multimodal attention mechanism and the true distribution of the label. i corresponds to an aesthetic score value of 1, 2, …N, and N is the number of score values ​​in the dataset.

[0137] Step 45: Repeat the above steps 41 to 44 in batches until the loss value calculated in step 44 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism.

[0138] In this embodiment, step 5 specifically includes the following steps:

[0139] Step 51: Input the images in the test set into the trained image aesthetic quality assessment network model that integrates scene features and multimodal attention mechanism, and output the corresponding image aesthetic score distribution p.

[0140] Step 52: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score score. The calculation formula is as follows:

[0141]

[0142] in, Indicates a rating of s i The probability of i represents the i-th score, and N represents the number of scores.

[0143] Those skilled in the art will appreciate that the embodiments of the present application may be provided as methods, systems, or computer program products. Therefore, the present application may adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application may adopt the form of a computer program product implemented in one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) that include computer-usable program code.

[0144] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems), and computer program products according to the embodiments of the present application. It should be understood that each process and / or box in the flowchart and / or block diagram, as well as the combination of the processes and / or boxes in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.

[0145] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.

[0146] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.

[0147] The above is only a preferred embodiment of the present invention, and does not limit the present invention in other forms. Any technician familiar with the profession may use the above disclosed technical content to change or modify it into an equivalent embodiment with equivalent changes. However, any simple modification, equivalent change and modification made to the above embodiment according to the technical essence of the present invention without departing from the technical solution of the present invention still belongs to the protection scope of the technical solution of the present invention.

[0148] This patent is not limited to the above-mentioned optimal implementation mode. Anyone can derive various other forms of image aesthetic quality evaluation methods and systems that integrate scene features and multimodal attention mechanisms under the inspiration of this patent. All equivalent changes and modifications made within the scope of the patent application of the present invention shall fall within the scope of this patent.

Claims

1. A method for evaluating image aesthetic quality by integrating scene features and multimodal attention mechanism, characterized in that: The following steps are involved: Step S1: preprocess the data in the aesthetic image dataset, extract the text features of the comments corresponding to the aesthetic images, and divide the dataset into a training set and a test set; Step S2: training an image aesthetic score distribution prediction model that integrates scene features and a multimodal attention mechanism; the image aesthetic score distribution prediction model is obtained by training an image aesthetic quality evaluation network that integrates scene features and a multimodal attention mechanism, including a hierarchical image feature fusion module and a multimodal attention mechanism module that integrates text features and image features; Step S3: input the image into the trained image aesthetic quality score distribution prediction model that integrates scene features and multimodal attention mechanism, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score; The process of training the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism to obtain the image aesthetic score distribution prediction model integrating scene features and multimodal attention mechanism comprises the following steps: Step S21: using two pre-trained ResNet50s as feature extraction subnetworks, removing the last layers of the two ResNet50 networks respectively as image scene feature extraction subnetwork and image aesthetic feature extraction subnetwork; the image scene feature extraction subnetwork uses weights pre-trained on the Places365 dataset as initial parameters, and the image aesthetic feature extraction subnetwork uses weights pre-trained on the ImageNet dataset as initial parameters; Step S22: Input each batch of images in the training set of step S1 into the two sub-networks in step S21; suppose that the image scene features and image aesthetic features of the i-th corresponding layer output by the last four corresponding layers of the two ResNet50 networks are and i=1,2,3,4; first and Feature concatenation is performed according to the channel dimension, and then dimensionality reduction is performed through 1×1 convolution. The specific formula is as follows: F′ i =w i (F i )+b i Where i = 1, 2, 3, 4, and They are the image scene features and image aesthetic features output by the i-th corresponding layer, and their dimensions are both C i ×H i ×W i ; Concat(·) means that features are concatenated according to the channel dimension, F i yes and The output feature after concatenation has a dimension of 2C i ×H i ×W i ;w i and b i is the weight and bias of the 1×1 convolutional layer used by the i-th corresponding layer; F′ i Yes F i The output feature after the 1×1 convolution layer has a dimension of C i ×H i ×W i ; Step S23: Output feature F′ obtained in step S22 i , i=1,2,3,4 are input into the hierarchical image feature fusion module to obtain image fusion features, and then the image fusion features and the text features corresponding to the same batch of images after step S1 are input into the multimodal attention mechanism module to obtain multimodal fusion features; finally, the multimodal fusion features are input into the fully connected layer to obtain the image aesthetic score distribution; the number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set; The specific formula of the network's loss function during training is as follows: in, and i They represent the probability corresponding to the i-th value of the aesthetic score in the score distribution predicted by the image aesthetic quality evaluation network that integrates scene features and multimodal attention mechanism and the true distribution of the label, respectively. i corresponds to the aesthetic score value of 1, 2, ...N, and N is the number of score values ​​in the dataset; Step S24: Repeat the above steps S21 to S23 in batches until the loss value calculated in step S23 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism.

2. The image aesthetic quality evaluation method integrating scene features and multimodal attention mechanism according to claim 1, characterized in that: Step S1 specifically includes the following steps: Step S11: performing data cleaning on the comment text dataset corresponding to the aesthetic image, and deleting the comment text that is empty, has only one word, or is all numbers; Step S12: convert all words in the comment text obtained in step S11 into lowercase, and remove stop words and numbers; then use the GloVe pre-trained word vector to encode all words and punctuation marks to obtain the encoding matrix of all comment texts; Step S13: fix the size of the coding matrix of the comment text obtained in step S12 to D×L; delete the part of the coding matrix of the comment text whose height exceeds D, otherwise, fill it with 0; delete the part of the coding matrix of the comment text whose width exceeds L, otherwise, fill it with 0, to obtain the final coding matrix of the comment text; Step S14: input the text encoding matrix obtained in step S13 into the bidirectional gate controlled recurrent unit network to obtain text features with a size of C×D; Step S15: randomly crop all images in the aesthetic dataset and scale them to a fixed size H×W; Step S16: Divide the preprocessed images of the aesthetic dataset and the corresponding comment text features into a training set and a test set.

3. The image aesthetic quality evaluation method integrating scene features and multimodal attention mechanism according to claim 2 is characterized by: The hierarchical image feature fusion module consists of a structure R that changes the feature dimension and matrix calculation; The structure R consists of two 1×1 convolutions and one 3×3 convolution, with input image features X and dimension C x ×H x ×W x ,The feature X reduces the feature width and height and increases the number of channels through the structure R. The calculation formula is as follows: f=w1(X)+b1 f′=w2(f)+b2 f″=w3(f′)+b3 Among them, w1 and b1 are the weight and bias of the first 1×1 convolutional layer, and the dimension of feature f is w2 and b2 are the weights and biases of the 3×3 convolutional layer, and the dimension of feature f′ is w3 and b3 are the weights and biases of the second 1×1 convolutional layer, and the dimension of feature f” is Assume that the input feature of the i-th level of the hierarchical image feature fusion module is F i , i=1,2,3,4, dimension is C i ×H i ×W i ,in C i+1 =2C i ; First, feature F1 is added to feature F2 after passing through structure R to obtain feature The dimensions are C2×H2×W2, After passing through structure R, it is added to feature F3 to obtain feature The dimensions are C3×H3×W3, After the structure R, the feature F′1 is obtained, with the dimension of C4×H4×W4; Secondly, feature F2 is added to feature F3 after passing through structure R to obtain feature The dimensions are C3×H3×W3, After the structure R, the feature F′2 is obtained, with the dimension of C4×H4×W4; Again, feature F3 is transformed into feature F′3 after structure R, with the dimension of C4×H4×W4; Finally, features F′1, F′2, F′3, and F4 are added together to obtain the image fusion feature F, whose dimension is C4×H4×W4; The specific calculation formula for the above process is as follows: F′3=R(F3) F=F′1+F′2+F′3+F4.

4. The image aesthetic quality evaluation method integrating scene features and multimodal attention mechanism according to claim 3 is characterized by: In the multimodal attention mechanism module that integrates text features and image features: The image feature F with the dimension C×H×W I Input into two 1×1 convolutional layers respectively to obtain the key point feature K and image projection feature V, and the dimension remains unchanged. The specific formula is as follows: K=w1(F I )+b1 <h2 style=";text-align:left;direction:ltr">V = w2(F<h2 style=";text-align:left;direction:ltr"> I <h2 style=";text-align:left;direction:ltr"> )+b2 Among them, F I is the input image feature; w1 and b1 are the weight and bias of the 1×1 convolution layer used to extract the key point feature K; w2 and b2 are the weight and bias of the 1×1 convolution layer used to extract the image projection feature V; Then adjust the dimensions of the key point feature K and the image projection feature V, both of which have dimensions C×H×W, to heads×c×H×W, where C=heads×c; Step S33: First, the text feature F with a dimension of C×D obtained in step S1 is T After the dimension is adjusted, the adjusted dimension is C×H×W, and then after a 1×1 convolution layer, the dimension is adjusted again. The adjusted dimension is heads×c×H×W to obtain the text feature Q, where D=H×W. The specific formula is as follows: Q=rewhape(w3(reshape(F T ))+b3) Among them, F T is the input text feature, reshape(·) represents the dimension adjustment operation, w3 and b3 are the weight and bias of the 1×1 convolution layer used to extract the text feature Q; Then randomly initialize the position feature P, whose dimension is heads×c×H×W; The obtained key point feature K, image projection feature V, text feature Q, and S position feature P are calculated through activation functions and multiple matrix operations to obtain the multimodal fusion feature Z. The specific formula is as follows: Z =Softmax((Q+P)×K T )×V Where T represents the transpose of the matrix, Softmax(·) represents the Softmax activation function, + represents the matrix addition operation, and × represents the matrix multiplication operation; Finally, the multimodal fusion feature Z with the dimension of heads×c×H×W is dimensionally adjusted to C×H×W, where C=heads×c.

5. The image aesthetic quality evaluation method integrating scene features and multimodal attention mechanism according to claim 4 is characterized by: Step S3 specifically includes the following steps: Step S31: input the images in the test set into the trained image aesthetic quality evaluation network model integrating scene features and multimodal attention mechanism, and output the corresponding image aesthetic score distribution p; Step S32: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score core; the calculation formula is as follows: in, Indicates a rating of s i The probability of i Represents the value of the i-th aesthetic score.

6. An image aesthetic quality evaluation system integrating scene features and multimodal attention mechanism, characterized in that: Computer-based systems, including: Preprocessing module: used to preprocess the data in the aesthetic image dataset, extract the text features of the comments corresponding to the aesthetic images, and divide the dataset into a training set and a test set; Evaluation network: an image aesthetic score distribution prediction model obtained by training a fusion scene feature and a multimodal attention mechanism; the image aesthetic score distribution prediction model is obtained by training an image aesthetic quality evaluation network that integrates scene features and a multimodal attention mechanism, including a hierarchical image feature fusion module and a multimodal attention mechanism module that integrates text features and image features; Scoring model: It is used to input the image into the trained image aesthetic quality score distribution prediction model that integrates scene features and multimodal attention mechanism, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score; The process of training the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism to obtain the image aesthetic score distribution prediction model integrating scene features and multimodal attention mechanism comprises the following steps: Step S21: using two pre-trained ResNet50s as feature extraction subnetworks, removing the last layers of the two ResNet50 networks respectively as image scene feature extraction subnetwork and image aesthetic feature extraction subnetwork; the image scene feature extraction subnetwork uses weights pre-trained on the Places365 dataset as initial parameters, and the image aesthetic feature extraction subnetwork uses weights pre-trained on the ImageNet dataset as initial parameters; Step S22: Input each batch of images in the training set of step S1 into the two sub-networks in step S21; suppose that the image scene features and image aesthetic features of the i-th corresponding layer output by the last four corresponding layers of the two ResNet50 networks are and , i = 1, 2, 3, 4; first and Feature concatenation is performed according to the channel dimension, and then dimensionality reduction is performed through 1×1 convolution. The specific formula is as follows: F′ i =w i (F i )+b i Where i = 1, 2, 3, 4, and They are the image scene features and image aesthetic features output by the i-th corresponding layer, and their dimensions are both C i ×H i ×W i ; CDoncat(·) indicates that features are concatenated according to the channel dimension, F i yes and The output feature after concatenation has a dimension of 2C i ×H i ×W i ;w i and b i is the weight and bias of the 1×1 convolutional layer used by the i-th corresponding layer; F′ i Yes F i The output feature after the 1×1 convolution layer has a dimension of C i ×H i ×W i ; Step S23: Output feature F′ obtained in step S22 i , i=1,2,3,4 are input into the hierarchical image feature fusion module to obtain image fusion features, and then the image fusion features and the text features corresponding to the same batch of images after step S1 are input into the multimodal attention mechanism module to obtain multimodal fusion features; finally, the multimodal fusion features are input into the fully connected layer to obtain the image aesthetic score distribution; the number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set; The specific formula of the network's loss function during training is as follows: in, and i They represent the probability corresponding to the i-th value of the aesthetic score in the score distribution predicted by the image aesthetic quality evaluation network that integrates scene features and multimodal attention mechanism and the true distribution of the label, respectively. i corresponds to the aesthetic score value of 1, 2, ...N, and N is the number of score values ​​in the dataset; Step S24: Repeat the above steps S21 to S23 in batches until the loss value calculated in step S23 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network integrating scene features and multimodal attention mechanism.

7. An image aesthetic quality assessment system integrating scene features and multimodal attention mechanism, characterized by: The method comprises a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the image aesthetic quality evaluation method integrating scene features and a multimodal attention mechanism as described in any one of claims 1 to 5 when executing the computer program.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the image aesthetic quality evaluation method integrating scene features and a multimodal attention mechanism as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Image aesthetics quality evaluation method based on multi-domain knowledge driving

    CN111950655A

  • Image aesthetic quality evaluation method fused with multi-modal attention mechanism

    CN113657380A