Image Evaluation Method and System Based on Meta-Learning Reweighted Network Pseudo-Label Training
By adopting a meta-learning-based reweighted network pseudo-label training method in computer vision aesthetic evaluation, the image aesthetic evaluation problem caused by uneven distribution of data sets is solved, and a more robust and effective image quality evaluation is achieved.
Patent Information
- Application Number
- CN202111387152.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-22
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2041-11-22
AI Technical Summary
The existing computer vision aesthetic evaluation method is difficult to effectively evaluate the aesthetic quality of the image when the training data set is unevenly distributed, resulting in the model scoring concentrated in the middle segment, and lacks effective processing for high-score and low-score annotations.
The reweighted network pseudo-label training method based on meta-learning is adopted, and the parameters of the reweighted network are adjusted through meta-learning, combined with pseudo-label grouping regression training, and the regression subnet parameters are adjusted to realize the weighted average output of image quality evaluation.
This method solves the problem of imbalance in the training data set to a certain extent, provides a more robust image quality evaluation method, and can effectively screen the quality of social images.
Smart Images

Figure CN114049500B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of machine vision, and particularly to an image quality evaluation method and system based on meta-learning reweighted network pseudo-label training. Background Art
[0002] Aesthetic research is one of the hot areas in recent computer vision research. With the continuous development of social networks, more and more people hope to record and share the beautiful moments in their lives through photos. Aesthetic evaluation is highly abstract. In the academic community, end-to-end supervised training of image aesthetic evaluation is mainly carried out through deep neural network models. Therefore, the training results highly depend on the quality of the labeled data set.
[0003] In computational visual aesthetics, although existing general aesthetic benchmark data sets have fine numerical labels, the number of most data sets does not exceed 20,000, and there are serious distribution imbalance problems. The proportion of high-score and low-score labeled data is very low, resulting in the scores of deep learning models trained by supervised learning being concentrated in the middle segment. Currently, most methods in this field still improve image feature extraction, and there is no method improved for the distribution imbalance of the training data set. Summary of the Invention
[0004] To solve the above technical problems, the present invention provides an image quality evaluation method and system based on meta-learning reweighted network pseudo-label training.
[0005] The technical solution of the present invention is as follows: An image quality evaluation method based on meta-learning reweighted network pseudo-label training, including:
[0006] Step S1: Obtain a picture data set as a sample set, extract a meta data set and a training full data set from the sample set, construct a backbone feature extraction network, adjust the size and padding of the images in the sample set, and input them into the first convolutional layer of the EfficientNet-B4 network for adaptive feature pre-extraction, and perform feature extraction through the remaining network structure of EfficientNet-B4 to obtain high-dimensional feature maps of the meta data and the training full data;
[0007] Step S2: Input the high-dimensional feature maps into a ten-classification sub-network and a regression sub-network respectively, perform dimensionality reduction feature extraction through multiple fully connected layers, and obtain the image classification and image quality score of the input image respectively; wherein, the ten-classification sub-network and the regression sub-network include: an efficient channel attention module and a fully connected network;
[0008] Step S3: Calculate the loss of the ten-classification result obtained from the meta-dataset through cross-entropy and input it into the meta-learning reweighting network to adjust the parameters of the meta-learning reweighting network, obtaining an adjusted meta-learning reweighting network; Obtain a weighted loss of the ten-classification result obtained from the training full dataset through the adjusted meta-learning reweighting network, and adjust the parameters of the backbone feature extraction network and the ten-classification sub-network; wherein, the meta-learning reweighting network includes: a three-layer perceptron;
[0009] Step S4: Assign pseudo-labels to the training full dataset based on a binary classifier, perform grouped regression training on the sample set according to the pseudo-labels, adjust the parameters of the regression sub-network, and take the weighted average of multiple comprehensive results obtained from different grouped trainings to obtain an evaluation output.
[0010] Compared with the prior art, the present invention has the following advantages:
[0011] The present invention discloses an image quality evaluation method based on pseudo-label training of a meta-learning reweighting network, providing an image quality evaluation method with better robustness, capable of solving to a certain extent the problem caused by the imbalance of computer vision training datasets, and providing an effective screening for the current popular social image sharing. Description of the Drawings
[0012] Figure 1 It is a flowchart of an image quality evaluation method based on pseudo-label training of a meta-learning reweighting network in an embodiment of the present invention;
[0013] Figure 2 It is a structural schematic diagram of a meta-learning reweighting network in an embodiment of the present invention;
[0014] Figure 3 It is a structural block diagram of an image quality evaluation system based on pseudo-label training of a meta-learning reweighting network in an embodiment of the present invention. Detailed Embodiments
[0015] The present invention provides an image quality evaluation method based on pseudo-label training of a meta-learning reweighting network, which can solve to a certain extent the problem caused by the imbalance of computer vision training datasets and provide an efficient image quality evaluation for the current popular social image sharing.
[0016] For a better understanding of the present invention, the terms used in the following embodiments are explained:
[0017] Meta-learning: Using past knowledge and experience to guide the learning of new tasks, having the ability to learn to learn. In the embodiments of the present invention, this idea is used to learn the weights of different training errors.
[0018] Pseudo-label: Based on the technology of labeled data to give an approximate label. In the embodiments of the present invention, a binary classifier is used to perform binary pseudo-label annotation on the existing data set, and the data set is grouped and trained through the annotation.
[0019] In order to make the objectives, technical solutions and advantages of the present invention clearer, the following further elaborates on the present invention through specific embodiments and in conjunction with the accompanying drawings.
[0020] Embodiment 1
[0021] As Figure 1 shown, an image quality evaluation method based on meta-learning reweighted network pseudo-label training provided by the embodiments of the present invention includes the following steps:
[0022] Step S1: Obtain a picture data set as a sample set, extract a meta data set and a training full data set from the sample set, construct a backbone feature extraction network, adjust the size and padding of the images in the sample set, and input them into the first convolutional layer of the EfficientNet-B4 network for adaptive feature pre-extraction, and perform feature extraction through the remaining network structure of EfficientNet-B4 to obtain high-dimensional feature maps of the meta data and the training full data;
[0023] Step S2: Input the high-dimensional feature maps into a ten-classification sub-network and a regression sub-network respectively, perform dimensionality reduction feature extraction through multiple fully connected layers, and obtain the image classification and image quality score of the input image respectively; among them, the ten-classification sub-network and the regression sub-network include: an efficient channel attention module and a fully connected network;
[0024] Step S3: Calculate the loss of the ten-classification result obtained from the meta data set through cross-entropy and input it into the meta-learning reweighted network to adjust the parameters of the meta-learning reweighted network to obtain an adjusted meta-learning reweighted network; obtain the weighted loss of the ten-classification result obtained from the training full data set through the adjusted meta-learning reweighted network, and adjust the parameters of the backbone feature extraction network and the ten-classification sub-network; among them, the meta-learning reweighted network includes: a three-layer perceptron;
[0025] Step S4: Assign pseudo-labels to the training full data set based on the binary classifier, perform grouped regression training on the sample set according to the pseudo-labels, adjust the parameters of the regression sub-network, and take the weighted average of the multiple comprehensive results obtained from different grouped trainings to obtain the evaluation output.
[0026] In one embodiment, the above-mentioned step S1: Obtain a picture data set as a sample set, extract a meta data set and a training full data set from the sample set, construct a backbone feature extraction network, adjust the size and padding of the images in the sample set, and input them into the first convolutional layer of the EfficientNet-B4 network for adaptive feature pre-extraction, and perform feature extraction through the remaining network structure of EfficientNet-B4 to obtain high-dimensional feature maps of the meta data and the training full data, specifically including:
[0027] Step S11: Select a publicly available or self-built picture data set as the sample set; extract samples with accurate annotations and uniform data distribution from the sample set as the meta data set, and the remaining samples as the training full data set;
[0028] In this step, select or self-build a picture data set as the sample set from common publicly available picture data sets, such as AVA, AADB, PCCD, etc. Before image preprocessing, extract a small amount of high-quality data from the sample set as the meta data set. The number of the meta data set is much smaller than the total data set, and its characteristics are accurate annotations and uniform data distribution. In the embodiment of the present invention, according to the continuous evaluation annotations, they are divided into 10 segments on average, 200 images are extracted from each segment, and a total of 2000 images are screened as the meta data set, and the remaining part is used as the training full data set.
[0029] Step S12: Scale the images in the sample set proportionally, convert them into RGB three channels, place them in a canvas of a preset size, fit the left and upper sides of the canvas to the input images, and fill the redundant parts with 0;
[0030] In the embodiment of the present invention, scale the input image proportionally according to the aspect ratio of the length and width so that the long side is 800 and the short side is n, record the width side as w and the height side as h; and convert the image into RGB three channels, place it in a 3*800*800 canvas, fit the left and upper sides of the canvas to the original image, and fill the redundant parts with 0;
[0031] After preprocessing the canvas, make the offset mean and standard deviation of its RGB color channels within preset thresholds respectively; use the convolutional kernel of the first convolutional layer of the EfficientNet-B4 network to perform convolutional processing on the canvas, and obtain a high-dimensional feature map through the remaining structure of EfficientNet-B4.
[0032] Perform horizontal random flipping, tensor transformation and standard deviation offset on the canvas, and the offset mean and standard deviation of its RGB color channels are (0.485, 0.456, 0.406) and (0.229, 0.224, 0.225) respectively;
[0033] Preprocess the features of the pixel matrix 3*800*800 of the canvas after standardized offset. Use the convolution kernel of the first convolutional layer of the Efficientnet-B4 network to perform convolution processing on the features to obtain a feature map of 48*400*400. Then, perform position marking on this feature map, select the effective features in the range of 48*[w / 2]*[h / 2] of the feature map, and use two-dimensional adaptive pooling operation to pool it into an initial feature map of 48*190*190; perform feature extraction on this initial feature map through the remaining structure of EfficientNet-B4, and finally obtain a high-dimensional feature map of 1792*11*11.
[0034] In one embodiment, the above step S2: Input the high-dimensional feature map into the ten-classification sub-network and the regression sub-network respectively, and perform dimensionality reduction feature extraction through multiple fully connected layers to obtain the image classification and image quality score of the input image respectively; among them, the ten-classification sub-network and the regression sub-network include: an efficient channel attention module and a fully connected network, specifically including:
[0035] Step S21: Input the high-dimensional feature map into the ten-classification sub-network and the regression sub-network respectively. The ten-classification sub-network and the regression sub-network include: an efficient channel attention module and a fully connected network;
[0036] As Figure 2 shown, the inputs of the ten-classification sub-network and the regression sub-network are high-dimensional feature maps of 1792*11*11 respectively. Each sub-network contains an efficient channel attention module and a fully connected network with three fully connected layers;
[0037] Step S22: The efficient channel attention module performs global average pooling on each channel, and then captures local cross-channel attention information through one-dimensional convolution with a convolution kernel size of k and the interaction of each channel and its k neighbors, where k represents the area of local cross-channel attention and represents the number of neighbor nodes participating in the channel attention prediction. Multiply the local cross-channel attention information with the high-dimensional feature map to obtain an efficient channel attention feature map;
[0038] In the embodiment of the present invention, the value of k is taken as 7, and an efficient channel attention feature map of 1792*11*11 is obtained through the efficient channel attention module.
[0039] Step S23: The fully connected layer includes three fully connected layers. Among them, the first two fully connected layers in the fully connected layers of the ten-classification sub-network and the regression sub-network respectively perform high-dimensional feature extraction on the efficient channel attention feature map; the output of the third fully connected layer of the ten-classification sub-network is a 10-dimensional feature, obtaining a category annotation of ten equal parts of the image continuous score label, and the output of the third fully connected layer of the regression sub-network is a 1-dimensional feature, obtaining the predicted image continuous score output.
[0040] In this step, global average pooling is performed on the high-efficiency channel attention feature map to obtain 1792-dimensional features. The first two fully connected layers in the ten-classification sub-network and the regression sub-network respectively perform high-dimensional feature extraction on them, obtaining 896-dimensional features and 448-dimensional features respectively. The third fully connected layer sets different nodes according to different output results. The output of the third fully connected layer of the ten-classification sub-network is 10-dimensional features, obtaining the category annotation of the ten equal parts of the continuous score label of the image, and the output of the third fully connected layer of the regression sub-network is 1-dimensional features, obtaining the predicted continuous score output of the image.
[0041] In one embodiment, in the above step S3: The ten-classification result obtained from the meta-dataset is used to calculate the loss through cross-entropy and input into the meta-learning reweighting network to adjust the parameters of the meta-learning reweighting network, obtaining the adjusted meta-learning reweighting network; the ten-classification result obtained from training the full dataset is used to obtain the weighted loss through the adjusted meta-learning reweighting network, and the parameters of the backbone feature extraction network and the ten-classification sub-network are adjusted; among them, the meta-learning reweighting network includes: a three-layer perceptron, specifically including:
[0042] Step S31: Input the meta-dataset into the backbone feature extraction network and the ten-classification sub-network, input the obtained ten-classification cross-entropy loss into the meta-learning reweighting network, calculate the weights using its perceptron, and adjust the parameters of the meta-learning reweighting network with the weighted gradient to obtain the adjusted meta-learning reweighting network;
[0043] Step S32: Output the ten-classification result through the ten-classification sub-network for the high-dimensional feature map of the training full dataset, calculate the obtained loss value and then input it into the meta-learning reweighting network, and adjust the parameters of the backbone feature extraction network and the ten-classification sub-network with the weighted loss value to obtain the adjusted backbone feature extraction network and ten-classification sub-network.
[0044] In one embodiment, in the above step S4: Assign pseudo-labels to the training full dataset based on the binary classifier, perform grouped regression training on the sample set according to the pseudo-labels, adjust the parameters of the regression sub-network, and take the weighted average of the multiple comprehensive results obtained from different grouped trainings to obtain the evaluation output, specifically including:
[0045] Step S41: Block the parameters of the backbone feature extraction network, the ten-classification sub-network and the meta-learning reweighting network, open the parameters of the regression sub-network, and perform regression training using the training full dataset to obtain the regression model of the training full dataset;
[0046] Step S42: Perform binary classification training on the training full dataset, re-divide the training full dataset through the binary classification pseudo-labels obtained from the training to obtain the regression model of the pseudo-label dataset;
[0047] Similar to the training of the ten-classification model, use a binary classifier to perform binary classification training on the entire training dataset, so that the binary classifier assigns pseudo-labels to high-quality and low-quality images in the training dataset. According to different labels, the entire training dataset is divided into a high-quality pseudo-label dataset and a low-quality pseudo-label dataset. Repeat the above training process of the regression model to obtain a high-quality regression model, a low-quality regression model, and a regression model for the entire training dataset through the high-quality pseudo-label dataset, the low-quality pseudo-label dataset, and the entire training dataset.
[0048] Step S43: According to the output result Score1 of the pseudo-label regression model and the output result Score2 of the regression model for the entire training dataset, the final score Score = (Score1 + Score2) / 2.
[0049] Give the test image a binary classification pseudo-label through the binary classifier, and then according to the output result of the regression model obtained from the corresponding pseudo-label dataset. If the test image is labeled with a low-quality pseudo-label, output the result Score1 through the low-quality regression model, otherwise output the result Score1 through the high-quality regression model. Then output the result Score2 through the regression model for the entire training dataset, and the final score result takes the weighted average of the two, Score = (Score1 + Score2) / 2.
[0050] The present invention discloses an image quality evaluation method based on meta-learning reweighted network pseudo-label training, provides an image quality evaluation method with good robustness, can solve the problem caused by the imbalance of the computer vision training dataset to a certain extent, and provides an effective screening for the current popular social image sharing.
[0051] Embodiment 2
[0052] As Figure 3 shown, the embodiment of the present invention provides an image quality evaluation system based on meta-learning reweighted network pseudo-label training, including the following modules:
[0053] The high-dimensional feature map acquisition module 51 is used to obtain a picture dataset as a sample set, extract a meta-dataset and an entire training dataset from the sample set, construct a backbone feature extraction network, adjust the size and fill the images in the sample set, and input them into the first convolutional layer of the EfficientNet-B4 network for adaptive feature pre-extraction, and perform feature extraction through the remaining network structure of the EfficientNet-B4 to obtain high-dimensional feature maps of the meta-data and the entire training data;
[0054] An image classification and scoring module 52 is configured to input the high-dimensional feature map into a ten-classification sub-network and a regression sub-network respectively, perform dimensionality reduction feature extraction through multiple fully-connected layers, and obtain the image classification and image quality score of the input image respectively. The ten-classification sub-network and the regression sub-network include: an efficient channel attention module and a fully-connected network.
[0055] A meta-learning reweighting network module 53 calculates the loss through cross-entropy for the ten-classification results obtained from the meta-dataset and inputs it into the meta-learning reweighting network to adjust the parameters of the meta-learning reweighting network, obtaining an adjusted meta-learning reweighting network. The ten-classification results obtained from training the entire dataset are input into the adjusted meta-learning reweighting network to obtain a weighted loss, and the parameters of the backbone feature extraction network and the ten-classification sub-network are adjusted. The meta-learning reweighting network includes: a three-layer perceptron.
[0056] A pseudo-label training and evaluation module 54 assigns pseudo-labels to the entire training dataset based on a binary classifier, performs grouped regression training on the sample set according to the pseudo-labels, adjusts the parameters of the regression sub-network, and takes the weighted average of multiple comprehensive results obtained from different grouped trainings to obtain an evaluation output.
[0057] The above embodiments are provided only for the purpose of describing the present invention and are not intended to limit the scope of the present invention. The scope of the present invention is defined by the appended claims. All equivalent substitutions and modifications made without departing from the spirit and principles of the present invention shall be covered within the scope of the present invention.
Claims
1. An image quality evaluation method based on meta - learning re - weighted network pseudo - label training, characterized in that, it includes: Step S1: Obtain a picture data set as a sample set, extract a meta - data set and a training full - data set from the sample set, construct a backbone feature extraction network, adjust the size and padding of the images in the sample set, and input them into the first convolutional layer of the EfficientNet - B4 network for adaptive feature pre - extraction. Then, perform feature extraction through the remaining network structure of EfficientNet - B4 to obtain high - dimensional feature maps of the meta - data and the training full - data; Step S2: Input the high - dimensional feature maps into a ten - classification sub - network and a regression sub - network respectively, and perform dimensionality - reduction feature extraction through multiple fully - connected layers to obtain the image classification and the image quality score of the input image respectively. Among them, the ten - classification sub - network and the regression sub - network include: an efficient channel attention module and a fully - connected network, specifically including: Step S21: Input the high - dimensional feature maps into a ten - classification sub - network and a regression sub - network respectively. The ten - classification sub - network and the regression sub - network include: an efficient channel attention module and a fully - connected network; Step S22: The efficient channel attention module performs global average pooling on each channel, and then captures local cross - channel attention information through one - dimensional convolution with a convolution kernel size of k to interact with each channel and its k neighbors, where k represents the area of local cross - channel attention and represents the number of neighbor nodes participating in the channel attention prediction. Multiply the local cross - channel attention information with the high - dimensional feature map to obtain an efficient channel attention feature map; Step S23: The fully - connected layer includes three fully - connected layers. Among them, the first two fully - connected layers in the fully - connected layers of the ten - classification sub - network and the regression sub - network respectively perform high - dimensional feature extraction on the efficient channel attention feature map; the output of the third fully - connected layer of the ten - classification sub - network is a 10 - dimensional feature, obtaining a category annotation of the ten - equal division of the continuous image score label, and the output of the third fully - connected layer of the regression sub - network is a 1 - dimensional feature, obtaining the predicted continuous image score output; Step S3: Calculate the loss through cross - entropy for the ten - classification results obtained from the meta - data set and input it into the meta - learning re - weighted network to adjust the parameters of the meta - learning re - weighted network, obtaining an adjusted meta - learning re - weighted network; obtain a weighted loss for the ten - classification results obtained from the training full - data set through the adjusted meta - learning re - weighted network, and adjust the parameters of the backbone feature extraction network and the ten - classification sub - network; among them, the meta - learning re - weighted network includes: a three - layer perceptron; Step S4: Assign pseudo - labels to the training full - data set based on a binary classifier, perform grouped regression training on the sample set according to the pseudo - labels, adjust the parameters of the regression sub - network, and take the weighted average of multiple comprehensive results obtained from different grouped trainings to obtain an evaluation output.
2. The image quality evaluation method based on meta - learning re - weighted network pseudo - label training according to claim 1, characterized in that, Step S1: Obtain an image dataset as a sample set, extract a meta-dataset and a training full-dataset from the sample set, construct a backbone feature extraction network, resize and pad the images in the sample set, and input them into the first convolutional layer of the EfficientNet-B4 network for adaptive feature pre-extraction, and perform feature extraction through the remaining network structure of EfficientNet-B4 to obtain high-dimensional feature maps of the meta-data and the training full-data, specifically including: Step S11: Select a publicly available or self-built image evaluation dataset as the sample set; extract samples with accurate annotations and uniform data distribution from the sample set as the meta-dataset, and the remaining samples as the training full-dataset; Step S12: Scale the images in the sample set proportionally, convert them into RGB three channels, place them in a canvas of a preset size, align the left and upper sides of the canvas with the input image, and fill the extra parts with 0; Step S13: After preprocessing the canvas, make the offset mean and standard deviation of its RGB color channels within preset thresholds respectively; use the convolution kernel of the first convolutional layer of the EfficientNet-B4 network to perform convolution processing on the canvas, and obtain high-dimensional feature maps of the meta-data and the training full-data through the remaining structure of EfficientNet-B4.
3. The image quality evaluation method based on meta-learning reweighted network pseudo-label training according to claim 1, characterized in that Step S3: Calculate the loss through cross-entropy for the ten-classification results obtained from the meta-dataset and input it into the meta-learning reweighted network to adjust the parameters of the meta-learning reweighted network, and obtain an adjusted meta-learning reweighted network; obtain a weighted loss for the ten-classification results obtained from the training full-dataset through the adjusted meta-learning reweighted network, and adjust the parameters of the backbone feature extraction network and the ten-classification sub-network; wherein, the meta-learning reweighted network includes: a three-layer perceptron, specifically including: Step S31: Input the meta-dataset into the backbone feature extraction network and the ten-classification sub-network, input the obtained ten-classification cross-entropy loss into the meta-learning reweighted network, calculate the weights using its perceptron, and adjust the parameters of the meta-learning reweighted network with the weighted gradient to obtain an adjusted meta-learning reweighted network; Step S32: Output the ten-classification results for the high-dimensional feature maps of the training full-data through the ten-classification sub-network, calculate the obtained loss value and input it into the meta-learning reweighted network, and adjust the parameters of the backbone feature extraction network and the ten-classification sub-network with the weighted loss value to obtain an adjusted backbone feature extraction network and ten-classification sub-network.
4. The image quality evaluation method based on meta-learning reweighted network pseudo-label training according to claim 1, characterized in that Step S4: Assign pseudo-labels to the training full dataset based on a binary classifier, perform grouped regression training on the sample set according to the pseudo-labels, adjust the parameters of the regression sub-network, and take the weighted average of multiple comprehensive results obtained from different grouped trainings to obtain an evaluation output, which specifically includes: Step S41: Block the parameters of the backbone feature extraction network, the ten-class sub-network, and the meta-learning reweighting network, open the parameters of the regression sub-network, and use the training full dataset for regression training to obtain a regression model for the training full dataset; Step S42: Perform binary classification training on the training full dataset, re-partition the training full dataset through the binary classification pseudo-labels obtained from the training to obtain a regression model for the pseudo-label dataset; Step S43: According to the output result Score1 of the regression model of the pseudo-label dataset and the output result Score2 of the regression model of the training full dataset, the final score Score = (Score1 + Score2) / 2.
5. An image quality evaluation system based on pseudo-label training of a meta-learning reweighting network Characterized in that It includes the following modules: A high-dimensional feature map acquisition module, configured to obtain a picture dataset as a sample set, extract a meta-dataset and a training full dataset from the sample set, construct a backbone feature extraction network, perform size adjustment and padding on the images in the sample set, and input them into the first convolutional layer of the EfficientNet-B4 network for adaptive feature pre-extraction, and perform feature extraction through the remaining network structure of the EfficientNet-B4 to obtain high-dimensional feature maps of the meta-data and the training full data; An image classification and scoring module, configured to respectively input the high-dimensional feature maps into a ten-class sub-network and a regression sub-network, perform dimensionality reduction feature extraction through multiple fully-connected layers, and respectively obtain the image classification and the image quality score of the input image; wherein, the ten-class sub-network and the regression sub-network include: an efficient channel attention module and a fully-connected network, specifically including: Step S21: Respectively input the high-dimensional feature maps into a ten-class sub-network and a regression sub-network, and the ten-class sub-network and the regression sub-network include: an efficient channel attention module and a fully-connected network; Step S22: The efficient channel attention module performs global average pooling on each channel, and then captures local cross-channel attention information through one-dimensional convolution with a convolution kernel size of k for the interaction between each channel and its k neighbors, where k represents the area of local cross-channel attention and represents the number of neighbor nodes participating in the prediction of the channel attention. Multiply the local cross-channel attention information with the high-dimensional feature map to obtain an efficient channel attention feature map; Step S23: The fully connected layer includes three fully connected layers. Among them, the first two fully connected layers in the fully connected layers of the ten-classification sub-network and the regression sub-network respectively perform high-dimensional feature extraction on the efficient channel attention feature map; the output of the third fully connected layer of the ten-classification sub-network is a 10-dimensional feature, obtaining the class annotation of the ten equal parts of the continuous image score label, and the output of the third fully connected layer of the regression sub-network is a 1-dimensional feature, obtaining the predicted continuous image score output. The meta-learning reweighting network module is used to calculate the loss of the ten-classification result obtained from the meta-dataset through cross-entropy and input it into the meta-learning reweighting network to adjust the parameters of the meta-learning reweighting network, obtaining the adjusted meta-learning reweighting network; obtaining the weighted loss of the ten-classification result obtained from the training full dataset through the adjusted meta-learning reweighting network, and adjusting the parameters of the backbone feature extraction network and the ten-classification sub-network; among them, the meta-learning reweighting network includes: a three-layer perceptron. The pseudo-label training evaluation module is used to assign pseudo-labels to the training full dataset based on the binary classifier, perform grouped regression training on the sample set according to the pseudo-labels, adjust the parameters of the regression sub-network, and take the weighted average of multiple comprehensive results obtained from different grouped trainings to obtain the evaluation output.
Citation Information
Patent Citations
Attention mechanism-based image aesthetics quality evaluation method
CN110473164A
Image score label prediction method based on deep convolutional neural network
CN111340123A