A method for image aesthetic quality evaluation based on composition perception

By combining the fusion method of compositional features and aesthetic features, the problem of inaccurate image aesthetic quality evaluation in the prior art is solved, and a more efficient and accurate image aesthetic quality evaluation is achieved.

CN116342569BActive Publication Date: 2025-05-16FUZHOU UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310347918.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-04
Publication Date
2025-05-16
Estimated Expiration
2043-04-04

AI Technical Summary

Technical Problem

The prior art is difficult to fully represent image aesthetic features, resulting in inaccurate evaluation of image aesthetic quality, especially the composition information cannot be fully utilized in existing methods.

Method used

Using the image aesthetic quality evaluation method based on composition perception, the image aesthetic quality evaluation data set and image composition quality evaluation data set formed by subjective evaluation is designed, and a pyramid-type multi-scale feature fusion module and image composition quality evaluation network are integrated, combining composition features and aesthetic features to improve the performance of image aesthetic quality evaluation.

Benefits of technology

Effectively utilize image composition information to improve the accuracy and performance of image aesthetic quality evaluation, and can more comprehensively evaluate the aesthetic quality of images.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116342569B_ABST
    Figure CN116342569B_ABST
Patent Text Reader

Abstract

The present invention proposes an image aesthetic quality evaluation method based on composition perception, comprising the following steps: step S1: forming an aesthetic quality evaluation training set and an aesthetic quality evaluation test set, a composition quality evaluation training set and a composition quality evaluation test set from data after processing an image aesthetic quality evaluation data set and an image composition quality evaluation data set; step S2: designing a pyramid-type multi-scale feature fusion module; step S3: designing an image composition quality evaluation network, and training to obtain an image composition quality evaluation model; step S4: designing an image aesthetic quality evaluation network, and training to obtain an image aesthetic quality evaluation model; step S5: inputting an image in an aesthetic quality evaluation test set into the image aesthetic quality evaluation model, outputting a corresponding score distribution, and calculating an average value as an image aesthetic quality score; the present invention can effectively assist in realizing image aesthetic evaluation with the help of composition information in an image, and further improve the performance of an image aesthetic quality evaluation algorithm.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention relates to the fields of image processing and computer vision, and in particular to an image aesthetic quality evaluation method based on composition perception. Background Art

[0002] With the popularity of mobile devices and the Internet, the number of images is increasing rapidly. As the number of images increases, people's demand for the beauty of images also increases, and the quality of image aesthetics has become the focus of people's attention. The quality of image aesthetics measures the visual appeal of an image in the eyes of humans. People usually hope that the images they obtain have high aesthetic quality. Image aesthetic quality evaluation technology refers to the computer automatically evaluating the beauty of an image by calculating the quality of the image and imitating the process of people's perception and cognition of the image. Image aesthetic quality evaluation technology is widely used in daily life. For example, for multiple similar photos taken by mobile phones, this technology can help people select the most "beautiful" photo to overcome the fear of choice; for multiple different video covers, this technology can help the video select the most "beautiful" cover to increase its click-through rate. Image aesthetic quality evaluation technology can not only select pictures with high aesthetic quality, but also the computer can automatically beautify the image according to its own understanding. This technology has not only promoted the progress of the design industry, the beauty industry, and the film and television industry, but also promoted the development of science and technology. However, since visual aesthetics is a subjective attribute that often involves subjective factors such as emotions and personal taste, it makes it a very challenging task to automatically evaluate the aesthetic quality of images using computers.

[0003] Image aesthetic quality evaluation methods are generally divided into methods based on manual feature extraction and methods based on deep learning. The method based on manual feature extraction first manually designs a variety of image features related to aesthetic quality according to aesthetic expertise, then extracts these manually designed features on the image aesthetic quality evaluation dataset, and then uses these features in combination with effective machine learning algorithms to train classifiers or regressors to classify or regress the aesthetic quality of the image. However, manually designed features are often inspired by photographic factors or psychology, and the range of manually designed features is limited, and they cannot fully represent aesthetic features, and thus cannot fully evaluate the aesthetics of the image.

[0004] At present, the advanced image aesthetic quality evaluation methods are all based on deep learning methods. Deep learning has a powerful automatic feature learning ability, does not require people to have rich knowledge of image aesthetics and psychology, and is much more efficient and accurate than methods based on manual feature extraction. Therefore, more and more researchers have successfully solved the problem of image aesthetic quality evaluation using deep convolutional neural networks and achieved good performance. However, most of the image aesthetic quality evaluation methods based on deep learning are currently limited to learning image aesthetic features. We found that the composition of an image will directly affect people's evaluation of the aesthetic quality of the image. Usually, when the composition quality of an image is relatively high, the aesthetic quality of the image will also be relatively high, and the two are proportional. Therefore, we believe that the composition information of the image is an important influencing factor in image aesthetic evaluation, and making full use of the composition information of the image can well assist in aesthetic evaluation. We propose an image aesthetic quality evaluation method guided by composition attributes, which can well integrate the image composition features with the image aesthetic features, and further improve the performance of the image aesthetic quality evaluation method. Summary of the invention

[0005] The present invention proposes an image aesthetic quality evaluation method based on composition perception, which can effectively use the composition information in the image to assist in image aesthetic evaluation and further improve the performance of the image aesthetic quality evaluation algorithm.

[0006] The present invention adopts the following technical solutions.

[0007] A method for evaluating image aesthetic quality based on composition perception, the method uses an image aesthetic quality evaluation dataset and an image composition quality evaluation dataset formed by human subjective evaluation, and guides image aesthetic quality based on the composition attributes of the image, comprising the following steps:

[0008] Step S1: preprocessing the data in the image aesthetic quality evaluation data set and the image composition quality evaluation data set to form an aesthetic quality evaluation training set and an aesthetic quality evaluation test set, and a composition quality evaluation training set and a composition quality evaluation test set;

[0009] Step S2: Design a pyramid-type multi-scale feature fusion module;

[0010] Step S3: designing an image composition quality evaluation network, and using the designed network training to obtain an image composition quality evaluation model;

[0011] Step S4: designing an image aesthetic quality evaluation network based on composition perception, and using the designed network to train to obtain an image aesthetic quality evaluation model based on composition perception;

[0012] Step S5: Input the images in the aesthetic quality evaluation test set into the trained image aesthetic quality evaluation model based on composition perception, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.

[0013] The step S1 specifically includes the following steps:

[0014] Step S11: Pairing images in the image aesthetic quality assessment dataset and their corresponding labels, and images in the image composition quality assessment dataset and their corresponding labels;

[0015] Step S12: dividing the images in the image aesthetic quality assessment dataset and the image composition quality assessment dataset into an aesthetic quality assessment training set and an aesthetic quality assessment test set, and a composition quality assessment training set and a composition quality assessment test set according to a required ratio;

[0016] Step S13: scaling the images in the aesthetic quality assessment training set and the aesthetic quality assessment test set, and the composition quality assessment training set and the composition quality assessment test set to a fixed size H×W respectively;

[0017] Step S14: performing random horizontal flipping operations on the images in the aesthetic quality evaluation training set and the composition quality evaluation training set respectively, for training set data enhancement;

[0018] Step S15: normalizing the images in the aesthetic quality evaluation training set and the aesthetic quality evaluation test set, and the composition quality evaluation training set and the composition quality evaluation test set respectively.

[0019] The step S2 specifically includes the following steps:

[0020] Step S21: Design a pyramid-type multi-scale feature fusion module to change the structure Z of the image feature dimension. The structure Z consists of 1 1×1 convolution, 1 3×3 convolution and 1 1×1 convolution. The input image feature X has a dimension of C. x ×H x ×W x , C x , H x and W x are the number of channels, height and width of the image feature X respectively; feature X reduces the feature width and height and increases the number of channels through structure Z. The calculation formula is as follows:

[0021] f=w1(X)+b1 Formula 1;

[0022] f′=w2(f)+b2 Formula 2;

[0023] f″=w3(f′)+b3 Formula 3;

[0024] Among them, w1 and b1 are the weight and bias of the first 1×1 convolutional layer, and feature f is the feature after the first 1×1 convolutional layer, and its dimension is w2 and b2 are the weights and biases of the 3×3 convolutional layer, and feature f′ is the feature after the 3×3 convolutional layer, and its dimension is w3 and b3 are the weights and biases of the second 1×1 convolutional layer, and feature f″ is the feature after the second 1×1 convolutional layer, and its dimension is

[0025] Step S22: The pyramid multi-scale feature fusion module consists of the six structures Z designed in step S21, six feature concatenation operations based on channel dimensions, and one 1×1 convolution. If the i-th input feature of the pyramid multi-scale feature fusion module is F i (i=1, 2, 3, 4), dimension is C i ×H i ×W i , then C i+1 =2C i ;

[0026] First, feature F1 passes through structure Z and is concatenated with feature F2 according to the channel dimension to obtain feature Its dimension is 2C2×H2×W2. After passing through structure Z, feature F2 is concatenated with feature F3 according to the channel dimension to obtain feature Its dimensions are 2C3×H3×W3,

[0027] After passing through structure Z, feature F3 is concatenated with feature F4 according to the channel dimension to obtain feature Its dimensions are 2C4×H4×W4;

[0028] Secondly, the characteristics After structure Z and characteristics Concatenate by channel dimension to get features Its dimensions are 4C3×H3×W3, and its features are After structure Z and characteristics Concatenate by channel dimension to get features Its dimensions are 4C4×H4×W4;

[0029] Again, features After structure Z and characteristics Concatenate by channel dimension to get features Its dimensions are 8C4×H4×W4;

[0030] Finally, the feature with dimension 8C4×H4×W4 After a 1×1 convolution, the dimension is reduced according to the channel dimension. After the dimension reduction, the final image fusion feature F is obtained, and its dimension is C4×H4×W4;

[0031] The specific calculation formula for the above process is as follows:

[0032] F1 1 =Concat(Z(F1), F2) Formula 4;

[0033] F1 2 =Concat(Z(F2), F3) Formula 5;

[0034] F1 3 =Concat(Z(F3), F4) Formula 6;

[0035] F2 1 =Concat(Z(F1 1 ), F1 2 ) Formula 7;

[0036] F2 2 =Concat(Z(F1 2 ), F1 3 ) Formula 8;

[0037] F3 1 =Concat(Z(F2 1 ), F2 2 ) Formula 9;

[0038] F=Conv 1×1 (F3 1 ) Formula 10;

[0039] Among them, Concat(·,·) represents the feature concatenation operation according to the channel dimension, Conv 1×1 (·) denotes a 1×1 convolution.

[0040] The step S3 comprises the following steps:

[0041] Step S31: The ResNet_v2_50 network that has been pre-trained and has the last layer removed is used as a composition feature extraction network, and the weights pre-trained on the ImageNet dataset are used as initial parameters;

[0042] Step S32: Input each batch of images of the composition quality evaluation training set obtained in step S1 into the composition feature extraction network in step S31 to obtain the output features of the last four stages of the ResNet_v2_50 network, and set the output feature of the i-th stage to be The dimension is

[0043] Step S33: Output features obtained in step S32 Input to the pyramid multi-scale feature fusion module designed in step S2 to obtain the final image composition feature F C , whose dimensions are Last feature F C Input to the fully connected layer to get the image composition quality score;

[0044] Step S34: Design the loss function of the image composition quality evaluation network. The specific calculation formula is as follows:

[0045]

[0046] Among them, m is the number of samples, x i is the true composition quality score of the i-th image, is the predicted composition quality score of the i-th image;

[0047] Step S35: Repeat the above steps S31 to S34 in batches until the loss value calculated in step S34 converges and stabilizes, save the network parameters, and complete the training process of the image composition quality evaluation network.

[0048] The step S4 specifically includes the following steps:

[0049] Step S41: first, the image composition quality evaluation network designed in step S3 is used as the image composition feature extraction sub-network after removing the fully connected layer, and the network weights trained in step S3 are used as the parameters of the image composition feature extraction sub-network. This part of the parameters does not participate in the training of the image aesthetic quality evaluation network based on composition perception; then the pre-trained SwinTransformer network is used and the last layer is removed as the image aesthetic feature extraction sub-network, which uses the weights pre-trained on the ImageNet dataset as the initial parameters;

[0050] Step S42: Input each batch of images in the aesthetic training set in step S1 into the two sub-networks in step S41 in sequence; suppose the aesthetic image is passed through the image composition feature extraction sub-network to obtain the image composition feature F C , whose dimensions are And after the image aesthetic feature extraction sub-network, the output features of the four stages are obtained as follows:

[0051] Step S43: Output features obtained in step S42 Input to the pyramid multi-scale feature fusion module designed in step S2 to obtain the final image aesthetic feature F A , whose dimensions are

[0052] Step S44: Construct a cross encoder, which consists of multi-head cross attention, layer normalization and fully connected layers. Assume that the input features of the cross encoder are q, k and v, and the dimensions of q, k and v are all c×s. First, input them to the multi-head cross attention. The output of the multi-head cross attention is added to v, and the layer normalization is performed, which is recorded as The intermediate output feature r of the cross encoder is obtained and then input into two fully connected layers, denoted as MLP c (·), the outputs of the two fully connected layers are added to r, and then layer normalized, denoted as Finally, the output feature r′ is obtained, whose dimension is c×s;

[0053] The formula of the cross encoder Encoder is r′=Encoder(q, k, v), where Encoder(·, ·, ·) represents the calculation of the cross encoder. The specific calculation formula is as follows:

[0054]

[0055]

[0056] Among them, MHCA(·,·,·) represents multi-head cross attention, + represents matrix addition operation;

[0057] Step S45: First, the image composition feature F obtained in step S42 is C , whose dimensions are After a 1×1 convolution, the dimension is reduced according to the channel dimension to obtain the reduced-dimensional composition feature F′ C , whose dimensions are Then the feature F′ C and the image aesthetic feature F obtained in step S43 A , whose dimensions are After the Reshape operation, the dimensions are adjusted to obtain the feature F″ with a dimension of C×S. C and the feature F′ of dimension C×S A ;in, The specific calculation formula is as follows:

[0058] F′ C =Conv 1×1 (F C ) Formula 14;

[0059] F″ C =Reshape(F′ C ) Formula 15;

[0060] F′ A =Reshape(FA ) Formula 16;

[0061] Among them, Conv 1×1 (·) represents 1×1 convolution, and Reshape(·) represents dimension adjustment operation;

[0062] Step S46: The feature F″ obtained in step S45 C and feature F′ A Input to the cross encoder constructed in step S44, F″ C As the input feature q of the cross encoder, F′ A As the input features k and v of the cross encoder, the image aesthetic features fused with the composition features are obtained, and then the image aesthetic features fused with the composition features are input into the fully connected layer to obtain the image aesthetic score distribution;

[0063] The number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set; when the score set is set to {1, 2, …, 10}, then N is 10;

[0064] Step S47: Design a loss function for the image aesthetic quality evaluation network based on composition perception. The specific formula is as follows:

[0065]

[0066] in, and i They represent the probability corresponding to the i-th value of the aesthetic score in the score distribution predicted by the image aesthetic quality evaluation network based on composition perception and the true distribution of the label, respectively. i corresponds to the aesthetic score value of 1, 2, ...N, and N is the number of score values ​​in the dataset;

[0067] Step S48: Repeat the above steps S41 to S47 in batches until the loss value calculated in step S47 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network based on composition perception.

[0068] Step S5 specifically includes the following steps:

[0069] Step S51: input the images in the aesthetic quality evaluation test set into the trained image aesthetic quality evaluation network model based on composition perception, and output the corresponding image aesthetic score distribution p;

[0070] Step S52: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score score; the calculation formula is as follows:

[0071]

[0072] in, Indicates a rating of s i The probability of i represents the i-th aesthetic score value, and N represents the number of scores.

[0073] In the image aesthetic quality evaluation data set and the image composition quality evaluation data set, the image composition quality evaluation value is proportional to the image aesthetic evaluation value.

[0074] The present invention can effectively use the composition information in the image to assist in image aesthetic evaluation, and further improve the performance of the image aesthetic quality evaluation algorithm. BRIEF DESCRIPTION OF THE DRAWINGS

[0075] The present invention is further described in detail below with reference to the accompanying drawings and specific embodiments:

[0076] Attached Figure 1 It is a schematic diagram of the implementation flow of the method of the present invention;

[0077] Attached Figure 2 is a schematic diagram of the network model structure in an embodiment of the present invention;

[0078] Attached Figure 3 It is a schematic diagram of the structure of a pyramid-type multi-scale feature fusion module in an embodiment of the present invention. DETAILED DESCRIPTION

[0079] As shown in the figure, a method for evaluating image aesthetic quality based on composition perception is provided. The method uses an image aesthetic quality evaluation dataset and an image composition quality evaluation dataset formed by subjective evaluation by humans, and guides image aesthetic quality based on the composition attributes of the image, and includes the following steps:

[0080] Step S1: preprocessing the data in the image aesthetic quality evaluation data set and the image composition quality evaluation data set to form an aesthetic quality evaluation training set and an aesthetic quality evaluation test set, and a composition quality evaluation training set and a composition quality evaluation test set;

[0081] Step S2: Design a pyramid-type multi-scale feature fusion module;

[0082] Step S3: designing an image composition quality evaluation network, and using the designed network training to obtain an image composition quality evaluation model;

[0083] Step S4: designing an image aesthetic quality evaluation network based on composition perception, and using the designed network to train to obtain an image aesthetic quality evaluation model based on composition perception;

[0084] Step S5: Input the images in the aesthetic quality evaluation test set into the trained image aesthetic quality evaluation model based on composition perception, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score.

[0085] The step S1 specifically includes the following steps:

[0086] Step S11: Pairing images in the image aesthetic quality assessment dataset and their corresponding labels, and images in the image composition quality assessment dataset and their corresponding labels;

[0087] Step S12: dividing the images in the image aesthetic quality assessment dataset and the image composition quality assessment dataset into an aesthetic quality assessment training set and an aesthetic quality assessment test set, and a composition quality assessment training set and a composition quality assessment test set according to a required ratio;

[0088] Step S13: scaling the images in the aesthetic quality assessment training set and the aesthetic quality assessment test set, and the composition quality assessment training set and the composition quality assessment test set to a fixed size H×W respectively;

[0089] Step S14: performing random horizontal flipping operations on the images in the aesthetic quality evaluation training set and the composition quality evaluation training set respectively, for training set data enhancement;

[0090] Step S15: normalizing the images in the aesthetic quality evaluation training set and the aesthetic quality evaluation test set, and the composition quality evaluation training set and the composition quality evaluation test set respectively.

[0091] The step S2 specifically includes the following steps:

[0092] Step S21: Design a pyramid-type multi-scale feature fusion module to change the structure Z of the image feature dimension. The structure Z consists of 1 1×1 convolution, 1 3×3 convolution and 1 1×1 convolution. The input image feature X has a dimension of C. x ×H x ×W x ,C x , H x and W x are the number of channels, height and width of the image feature X respectively; feature X reduces the feature width and height and increases the number of channels through structure Z. The calculation formula is as follows:

[0093] f=w1(X)+b1 Formula 1;

[0094] f′=w2(f)+b2 Formula 2;

[0095] f″=w3(f′)+b3 Formula 3;

[0096] Among them, w1 and b1 are the weight and bias of the first 1×1 convolutional layer, and feature f is the feature after the first 1×1 convolutional layer, and its dimension is w2 and b2 are the weights and biases of the 3×3 convolutional layer, and feature f′ is the feature after the 3×3 convolutional layer, and its dimension is

[0097] w3 and b3 are the weights and biases of the second 1×1 convolutional layer. Feature F is the feature after the second 1×1 convolutional layer, and its dimension is

[0098] Step S22: The pyramid multi-scale feature fusion module consists of the six structures Z designed in step S21, six feature concatenation operations based on channel dimensions, and one 1×1 convolution. If the i-th input feature of the pyramid multi-scale feature fusion module is F i (i=1, 2, 3, 4), dimension is C i ×H i ×W i , then C i+1 =2C i ;

[0099] First, feature F1 passes through structure Z and is concatenated with feature F2 according to the channel dimension to obtain feature Its dimension is 2C2×H2×W2. After passing through structure Z, feature F2 is concatenated with feature F3 according to the channel dimension to obtain feature Its dimensions are 2C3×H3×W3,

[0100] After passing through structure Z, feature F3 is concatenated with feature F4 according to the channel dimension to obtain feature Its dimensions are 2C4×H4×W4;

[0101] Secondly, the characteristics After structure Z and characteristics Concatenate by channel dimension to get features Its dimensions are 4C3×H3×W3, and its features are After structure Z and characteristics Concatenate by channel dimension to get features Its dimensions are 4C4×H4×W4;

[0102] Again, features After structure Z and characteristics Concatenate by channel dimension to get features Its dimensions are 8C4×H4×W4;

[0103] Finally, the feature with dimension 8C4×H4×W4 After a 1×1 convolution, the dimension is reduced according to the channel dimension. After the dimension reduction, the final image fusion feature F is obtained, and its dimension is C4×H4×W4;

[0104] The specific calculation formula for the above process is as follows:

[0105] F1 1 =Concat(Z(F1), F2) Formula 4;

[0106] F1 2 =Concat(Z(F2), F3) Formula 5;

[0107] F1 3 =Concat(Z(F3), F4) Formula 6;

[0108] F2 1 =Concat(Z(F1 1 ), F1 2 ) Formula 7;

[0109] F2 2 =Concat(Z(F1 2 ), F1 3 ) Formula 8;

[0110] F3 1 =Concat(Z(F2 1 ), F2 2 ) Formula 9;

[0111] F=Conv 1×1 (F3 1 ) Formula 10;

[0112] Among them, Concat(·,·) represents the feature concatenation operation according to the channel dimension, Conv 1×1 (·) denotes a 1×1 convolution.

[0113] The step S3 comprises the following steps:

[0114] Step S31: The ResNet_v2_50 network that has been pre-trained and has the last layer removed is used as a composition feature extraction network, and the weights pre-trained on the ImageNet dataset are used as initial parameters;

[0115] Step S32: Input each batch of images of the composition quality evaluation training set obtained in step S1 into the composition feature extraction network in step S31 to obtain the output features of the last four stages of the ResNet_v2_50 network, and set the output feature of the i-th stage to be The dimension is

[0116] Step S33: Output features obtained in step S32 Input to the pyramid multi-scale feature fusion module designed in step S2 to obtain the final image composition feature F C , whose dimensions are Last feature F C Input to the fully connected layer to get the image composition quality score;

[0117] Step S34: Design the loss function of the image composition quality evaluation network. The specific calculation formula is as follows:

[0118]

[0119] Among them, m is the number of samples, x i is the true composition quality score of the i-th image, is the predicted composition quality score of the i-th image;

[0120] Step S35: Repeat the above steps S31 to S34 in batches until the loss value calculated in step S34 converges and stabilizes, save the network parameters, and complete the training process of the image composition quality evaluation network.

[0121] The step S4 specifically includes the following steps:

[0122] Step S41: first, the image composition quality evaluation network designed in step S3 is used as the image composition feature extraction sub-network after removing the fully connected layer, and the network weights trained in step S3 are used as the parameters of the image composition feature extraction sub-network. This part of the parameters does not participate in the training of the image aesthetic quality evaluation network based on composition perception; then the pre-trained SwinTransformer network is used and the last layer is removed as the image aesthetic feature extraction sub-network, which uses the weights pre-trained on the ImageNet dataset as the initial parameters;

[0123] Step S42: Input each batch of images in the aesthetic training set in step S1 into the two sub-networks in step S41 in sequence; suppose the aesthetic image is passed through the image composition feature extraction sub-network to obtain the image composition feature F C , whose dimensions are And after the image aesthetic feature extraction sub-network, the output features of the four stages are obtained as follows:

[0124] Step S43: Output features obtained in step S42 Input to the pyramid multi-scale feature fusion module designed in step S2 to obtain the final image aesthetic feature F A , whose dimensions are

[0125] Step S44: Construct a cross encoder, which consists of multi-head cross attention, layer normalization and fully connected layers. Assume that the input features of the cross encoder are q, k and v, and the dimensions of q, k and v are all c×s. First, input them to the multi-head cross attention. The output of the multi-head cross attention is added to v, and the layer normalization is performed, which is recorded as The intermediate output feature r of the cross encoder is obtained and then input into two fully connected layers, denoted as MLP c (·), the outputs of the two fully connected layers are added to r, and then layer normalized, denoted as Finally, the output feature r′ is obtained, whose dimension is c×s;

[0126] The formula of the cross encoder Encoder is r′=Encoder(q, k, v), where Encoder(·,·,·) represents the calculation of the cross encoder. The specific calculation formula is as follows:

[0127]

[0128]

[0129] Among them, MHCA(·,·,·) represents multi-head cross attention, + represents matrix addition operation;

[0130] Step S45: First, the image composition feature F obtained in step S42 is C , whose dimensions are After a 1×1 convolution, the dimension is reduced according to the channel dimension to obtain the reduced-dimensional composition feature F′ C , whose dimensions are Then the feature F′ C and the image aesthetic feature F obtained in step S43 A , whose dimensions are After the Reshape operation, the dimensions are adjusted to obtain the feature F″ with a dimension of C×S. C and the feature F′ of dimension C×S A ;in, The specific calculation formula is as follows:

[0131] F′ C =Conv 1×1 (F C ) Formula 14;

[0132] F″ C =Reshape(F′ C ) Formula 15;

[0133] F′ A =Reshape(FA ) Formula XVI;

[0134] Among them, Conv 1×1 (·) represents 1×1 convolution, and Reshape(·) represents dimension adjustment operation;

[0135] Step S46: The feature F″ obtained in step S45 C and feature F′ A Input to the cross encoder constructed in step S44, F″ C As the input feature q of the cross encoder, F′ A As the input features k and v of the cross encoder, the image aesthetic features fused with the composition features are obtained, and then the image aesthetic features fused with the composition features are input into the fully connected layer to obtain the image aesthetic score distribution;

[0136] The number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set; when the score set is set to {1, 2, …, 10}, then N is 10;

[0137] Step S47: Design a loss function for the image aesthetic quality evaluation network based on composition perception. The specific formula is as follows:

[0138]

[0139] in, and i They represent the probability corresponding to the i-th value of the aesthetic score in the score distribution predicted by the image aesthetic quality evaluation network based on composition perception and the true distribution of the label, respectively. i corresponds to the aesthetic score value of 1, 2, ...N, and N is the number of score values ​​in the dataset;

[0140] Step S48: Repeat the above steps S41 to S47 in batches until the loss value calculated in step S47 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network based on composition perception.

[0141] Step S5 specifically includes the following steps:

[0142] Step S51: input the images in the aesthetic quality evaluation test set into the trained image aesthetic quality evaluation network model based on composition perception, and output the corresponding image aesthetic score distribution p;

[0143] Step S52: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score score; the calculation formula is as follows:

[0144]

[0145] in, Indicates a rating of s i The probability of i represents the i-th aesthetic score value, and N represents the number of scores.

[0146] In the image aesthetic quality evaluation data set and the image composition quality evaluation data set, the image composition quality evaluation value is proportional to the image aesthetic evaluation value.

Claims

1. A method for evaluating image aesthetic quality based on composition perception, characterized in that: The method uses an image aesthetic quality evaluation dataset and an image composition quality evaluation dataset formed by human subjective evaluation to guide image aesthetic quality based on the composition attributes of the image, and includes the following steps: Step S1: preprocessing the data in the image aesthetic quality evaluation data set and the image composition quality evaluation data set to form an aesthetic quality evaluation training set and an aesthetic quality evaluation test set, and a composition quality evaluation training set and a composition quality evaluation test set; Step S2: Design a pyramid-type multi-scale feature fusion module; Step S3: designing an image composition quality evaluation network, and using the designed network training to obtain an image composition quality evaluation model; Step S4: designing an image aesthetic quality evaluation network based on composition perception, and using the designed network to train to obtain an image aesthetic quality evaluation model based on composition perception; Step S5: input the images in the aesthetic quality evaluation test set into the trained image aesthetic quality evaluation model based on composition perception, output the corresponding image aesthetic score distribution, and finally calculate the average value of the aesthetic score distribution as the image aesthetic quality score; The step S3 comprises the following steps: Step S31: The ResNet_v2_50 network that has been pre-trained and has the last layer removed is used as a composition feature extraction network, and the weights pre-trained on the ImageNet dataset are used as initial parameters; Step S32: Input each batch of images of the composition quality evaluation training set obtained in step S1 into the composition feature extraction network in step S31 to obtain the output features of the last four stages of the ResNet_v2_50 network, and set the output feature of the i-th stage to be (i=1,2,3,4), the dimension is Step S33: Output features obtained in step S32 (i=1, 2, 3, 4) is input into the pyramid multi-scale feature fusion module designed in step S2 to obtain the final image composition feature F C , whose dimensions are Last feature F C Input to the fully connected layer to get the image composition quality score; Step S34: Design the loss function of the image composition quality evaluation network. The specific calculation formula is as follows: Among them, m is the number of samples, x i is the true composition quality score of the i-th image, is the predicted composition quality score of the i-th image; Step S35: repeat the above steps S31 to S34 in batches until the loss value calculated in step S34 converges and tends to be stable, save the network parameters, and complete the training process of the image composition quality evaluation network; The step S4 specifically includes the following steps: Step S41: first, the image composition quality evaluation network designed in step S3 is used as the image composition feature extraction sub-network after removing the fully connected layer, and the network weights trained in step S3 are used as the parameters of the image composition feature extraction sub-network. This part of the parameters does not participate in the training of the image aesthetic quality evaluation network based on composition perception; then the pre-trained SwinTransformer network is used and the last layer is removed as the image aesthetic feature extraction sub-network, which uses the weights pre-trained on the ImageNet dataset as the initial parameters; Step S42: Input each batch of images in the aesthetic training set in step S1 into the two sub-networks in step S41 in sequence; suppose the aesthetic image is passed through the image composition feature extraction sub-network to obtain the image composition feature F C , whose dimensions are And after the image aesthetic feature extraction sub-network, the output features of the four stages are obtained as follows: (i=1, 2, 3, 4), Step S43: Output features obtained in step S42 (i=1, 2, 3, 4) is input into the pyramid multi-scale feature fusion module designed in step S2 to obtain the final image aesthetic feature F A , whose dimensions are Step S44: Construct a cross encoder Encoder, which consists of multi-head cross attention, layer normalization and fully connected layers. Assume that the input features of the cross encoder are q, k and v, and the dimensions of q, k and v are all c×s. First, input it to the multi-head cross attention, and the output of the multi-head cross attention is added to v, and the layer is normalized, which is recorded as The intermediate output feature r of the cross encoder is obtained and then input into two fully connected layers, denoted as MLP c (·), the outputs of the two fully connected layers are added to r, and then layer normalized, denoted as Finally, the output feature r′ is obtained, whose dimension is c×s; The formula of the cross encoder Encoder is r′=Encoder(q, k, v), where Encoder(·, ·, ·) represents the calculation of the cross encoder. The specific calculation formula is as follows: Among them, MHCA(·,·,·) represents multi-head cross attention, + represents matrix addition operation; Step S45: First, the image composition feature F obtained in step S42 is C , whose dimensions are After a 1×1 convolution, the dimension is reduced according to the channel dimension to obtain the reduced-dimensional composition feature F′ C , whose dimensions are Then the feature F′ C and the image aesthetic feature F obtained in step S43 A , whose dimensions are After the Reshape operation, the dimensions are adjusted to obtain the feature F″ with a dimension of C×S. C and the feature F′ of dimension C×S A ;in, The specific calculation formula is as follows: F′ C =Conv 1×1 (F C ) Formula 14; F″ C = Reshape(F′ C ) Official 15; F′ A = Reshape(F A ) Official 16; Among them, Conv 1×1 (·) represents 1×1 convolution, and Reshape(·) represents dimension adjustment operation; Step S46: The feature F″ obtained in step S45 C and feature F′ A Input to the cross encoder constructed in step S44, F″ C As the input feature q of the cross encoder, F′ A As the input features k and v of the cross encoder, the image aesthetic features fused with the composition features are obtained, and then the image aesthetic features fused with the composition features are input into the fully connected layer to obtain the image aesthetic score distribution; The number of categories output by the fully connected layer is N, where N is the number of scores in the aesthetic score set; when the score set is set to {1, 2, …, 10}, then N is 10; Step S47: Design a loss function for the image aesthetic quality evaluation network based on composition perception. The specific formula is as follows: in, and i They represent the probability corresponding to the i-th value of the aesthetic score in the score distribution predicted by the image aesthetic quality evaluation network based on composition perception and the true distribution of the label, respectively. i corresponds to the aesthetic score value of 1, 2, ...N, and N is the number of score values ​​in the dataset; Step S48: Repeat the above steps S41 to S47 in batches until the loss value calculated in step S47 converges and stabilizes, save the network parameters, and complete the training process of the image aesthetic quality evaluation network based on composition perception.

2. The method for evaluating image aesthetic quality based on composition perception according to claim 1, characterized in that: The step S1 specifically includes the following steps: Step S11: Pairing images in the image aesthetic quality assessment dataset and their corresponding labels, and images in the image composition quality assessment dataset and their corresponding labels; Step S12: dividing the images in the image aesthetic quality assessment dataset and the image composition quality assessment dataset into an aesthetic quality assessment training set and an aesthetic quality assessment test set, and a composition quality assessment training set and a composition quality assessment test set according to a required ratio; Step S13: scaling the images in the aesthetic quality assessment training set and the aesthetic quality assessment test set, and the composition quality assessment training set and the composition quality assessment test set to a fixed size H×W respectively; Step S14: performing random horizontal flipping operations on the images in the aesthetic quality evaluation training set and the composition quality evaluation training set respectively, for training set data enhancement; Step S15: normalizing the images in the aesthetic quality evaluation training set and the aesthetic quality evaluation test set, and the composition quality evaluation training set and the composition quality evaluation test set respectively.

3. The method for evaluating image aesthetic quality based on composition perception according to claim 1, characterized in that: The step S2 specifically includes the following steps: Step S21: Design a pyramid-type multi-scale feature fusion module to change the structure Z of the image feature dimension. The structure Z consists of 1 1×1 convolution, 1 3×3 convolution and 1 1×1 convolution. The input image feature X has a dimension of C. x ×H x ×W x ,C x , H x and W x are the number of channels, height and width of the image feature X respectively; feature X reduces the feature width and height and increases the number of channels through structure Z. The calculation formula is as follows: f=w1(X)+b1 Formula 1; f′=w2(f)+b2 Formula 2; f″=w3(f′)+b3 Formula 3; Among them, w1 and b1 are the weight and bias of the first 1×1 convolutional layer, and feature f is the feature after the first 1×1 convolutional layer, and its dimension is w2 and b2 are the weights and biases of the 3×3 convolutional layer, and feature f′ is the feature after the 3×3 convolutional layer, and its dimension is w3 and b3 are the weights and biases of the second 1×1 convolutional layer. The feature f″ is the feature after the second 1×1 convolutional layer, and its dimension is Step S22: The pyramid multi-scale feature fusion module consists of the six structures Z designed in step S21, six feature concatenation operations based on channel dimensions, and one 1×1 convolution. If the i-th input feature of the pyramid multi-scale feature fusion module is F i (i=1, 2, 3, 4), dimension is C i ×H i ×W i , then C i+1 =2C i ; First, feature F1 is concatenated with feature F2 according to the channel dimension after passing through structure Z to obtain feature F1 1 , whose dimension is 2C2×H2×W2. After passing through structure Z, feature F2 is concatenated with feature F3 according to the channel dimension to obtain feature F1. 2 , whose dimensions are 2C3×H3×W3, After passing through structure Z, feature F3 is concatenated with feature F4 according to the channel dimension to obtain feature F1 3 , whose dimensions are 2C4×H4×W4; Secondly, feature F1 1 After structure Z and feature F1 2 Concatenate according to the channel dimension to get feature F2 1 , whose dimension is 4C3×H3×W3, feature F1 2 After structure Z and feature F1 3 Concatenate according to the channel dimension to get feature F2 2 , whose dimensions are 4C4×H4×W4; Again, feature F2 1 After structure Z and feature F2 2 Concatenate according to the channel dimension to get feature F3 1 , whose dimensions are 8C4×H4×W4; Finally, feature F3 with dimension 8C4×H4×W4 1 After a 1×1 convolution, the dimension is reduced according to the channel dimension, and the final image fusion feature F is obtained after the dimension reduction, and its dimension is C4×H4×W4; The specific calculation formula for the above process is as follows: F1 1 =Concat(Z(F1), F2) Formula 4; F1 2 =Concat(Z(F2), F3) Formula 5; F1 3 =Concat(Z(F3), F4) Formula 6; F2 1 =Concat(Z(F1 1 ), F1 2 ) Formula 7; F2 2 =Concat(Z(F1 2 ), F1 3 ) Formula 8; F1 3 =Concat(Z(F2 1 ), F2 2 ) Formula 9; F=Conv 1×1 (F3 1 ) Formula 10; Among them, Concat(·,·) represents the feature concatenation operation according to the channel dimension, Conv 1×1 (·) denotes a 1×1 convolution.

4. The method for evaluating image aesthetic quality based on composition perception according to claim 1, characterized in that: Step S5 specifically The steps include: Step S51: input the images in the aesthetic quality evaluation test set into the trained image aesthetic quality evaluation network model based on composition perception, and output the corresponding image aesthetic score distribution p; Step S52: Calculate the average value of the image aesthetic score distribution p to obtain the image aesthetic quality score score; the calculation formula is as follows: in, Indicates a rating of s i The probability of i represents the i-th aesthetic score value, and N represents the number of scores.

5. The method for evaluating image aesthetic quality based on composition perception according to claim 1, characterized in that: In the image aesthetic quality evaluation data set and the image composition quality evaluation data set, the image composition quality evaluation value is proportional to the image aesthetic evaluation value.

Citation Information

Patent Citations

  • Picture aesthetics description modeling and description method and system based on aesthetics attribute retrieval

    CN113610128A

  • Image aesthetic quality evaluation method fused with multi-modal attention mechanism

    CN113657380A