Multi-branch fused image aesthetic quality evaluation method and system
Through the multi-branch fusion image aesthetic quality evaluation method, combined with the main branch network, technical quality branch network and theme classification branch network, the problem of single evaluation dimensions in the existing technology is solved, and the comprehensive evaluation of image aesthetic quality and theme classification are achieved, and the performance and interpretability of evaluation are improved.
Patent Information
- Application Number
- CN202510219163.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-26
- Publication Date
- 2025-06-13
AI Technical Summary
The prior art has a single evaluation dimension in image aesthetic quality evaluation, making it difficult to fully understand and evaluate the aesthetic semantic characteristics of images.
The multi-branch fusion image aesthetic quality evaluation method is adopted, and the main branch network, technical quality branch network and theme classification branch network are combined to integrate the composition, style, theme and other aesthetic attribute information of the image to adaptively extract semantic information.
The performance of the image aesthetic quality evaluation method is improved, the aesthetic evaluation results are output through the image aesthetic score distribution, and the subject classification results of the image are provided, making the evaluation results more interpretable.
Smart Images

Figure CN120147656A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical fields of image processing and computer vision, and particularly relates to an image aesthetic quality evaluation method and system with multi-branch fusion. Background Art
[0002] Today, with the rapid development of multimedia technology and artificial intelligence, images have become an indispensable way for humans to obtain information due to their visual intuitiveness, emotional transmission, and cultural mediation. And the quality of images at the aesthetic level is also the focus of current attention. Image aesthetic quality evaluation quantifies the visual attractiveness of images through algorithms, understands and predicts the subjective feelings of humans towards the beauty of images, assists humans in quickly identifying and enhancing image beauty, and is applied to image retrieval, photo enhancement, and recommendation systems, and has important application values in industries such as photography, design, and medicine. The advantage of image aesthetic quality evaluation is that the model can be directly related to the visual experience of humans, while the main difficulty comes from the subjectivity of beauty, so it is still a challenging and promising proposition.
[0003] Currently, a large number of advanced studies on general image aesthetic quality evaluation have emerged. Most models extract image features based on the CNN or Transformer architecture and train the model using aesthetic subjective scoring data. However, understanding the aesthetic semantic features of images is a complex process. Humans' judgments of different images will be based on different aesthetic rules, and the change measure of this rule is generally called the image aesthetic attribute. Past algorithms usually design models for a specific aesthetic attribute such as lighting, color, composition, and style. Summary of the Invention
[0004] In order to overcome the deficiencies of the prior art, the purpose of the present invention is to provide an image aesthetic quality evaluation method and system with multi-branch fusion to solve the problem of single evaluation dimension in the process of most models extracting image aesthetic features.
[0005] To achieve the above purpose, the present invention provides the following technical solutions:
[0006] An image aesthetic quality evaluation method with multi-branch fusion, comprising the following steps:
[0007] Step S1, preprocess the data in the aesthetic image dataset, classify and preprocess it according to nine common image sizes, so as to perform batch training subsequently and keep most of the image composition information undamaged;
[0008] Step S2, design a main branch network with multi-level feature output, and adaptively fuse rich multi-level semantic features for different images by using an ordered weighted average operator;
[0009] Step S3: Design a technical quality branch network, and through identifying the technical distortion of the image and performing pre-training, learn the aesthetic attribute features such as the composition and style of the image;
[0010] Step S4: Design a theme classification branch network, and use a vision-language large model to output the theme classification label and theme attribute features of the image;
[0011] Step S5: Fuse the features of the main branch, technical quality branch, and theme classification branch, and input them into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the evaluation result of the image aesthetic quality.
[0012] Preferably, the specific implementation steps of step S1 are as follows:
[0013] Step S11: Select five common aspect ratios that can widely represent most professional photography images, art works, mobile phone photos, etc., and consider horizontal and vertical formats. Divide the image sizes into nine categories, and set the corresponding resolutions for each size;
[0014] Step S12: Classify the images in the dataset according to the category closest to their sizes, and scale them to the corresponding resolutions to obtain nine image subsets with different aspect ratios, so as to meet the requirements of batch training and maintain the image composition information;
[0015] Step S13: Divide the preprocessed images into a training set and a test set according to a preset ratio.
[0016] Preferably, the specific implementation steps of step S2 are as follows:
[0017] Step S21: Use the image classification network ResNet50 as the feature extraction module, extract feature vectors after four Conv Block convolutional modules, and perform dimensionality reduction operations such as pooling to obtain four 1×1×256-dimensional feature vectors;
[0018] Step S22: Use an ordered weighted average operator to process the feature vectors obtained in step S21, sort each corresponding value of the four vectors in descending order and assign weights to obtain a 256-dimensional aggregated feature vector; the formula of the ordered weighted average operator is:
[0019]
[0020] where, a i is the i-th value after sorting in descending order, ω i is the corresponding weight, and is a learnable parameter during the training phase;
[0021] Step S23: Input the training set that has gone through step S13 into the network model in step S21 to obtain the image aesthetic feature vector X sem .
[0022] Preferably, the specific implementation steps of step S3 are as follows:
[0023] Step S31: Use Mobilenet-V2 as the basic network for the technical quality evaluation task;
[0024] Step S32: Perform distortion processing on the unlabeled high-definition images with high aesthetic quality. Each distortion d corresponds to a set of distortion intensity values V, and an image set I = {I ν , ν∈V} is obtained, where I 0 = I. The distortion set D includes 5 types of technical distortions, namely JPEG compression, defocus blur, pixelation, Gaussian noise, and impulse noise, and their corresponding 6 types of distortion intensities, 6 types of style distortions, namely brightness, contrast, saturation, exposure, color temperature, and hue, and their corresponding 10 types of distortion intensities, and 6 types of composition distortions, namely rotation, horizontal cropping, vertical cropping, left diagonal cropping, right diagonal cropping, and image aspect ratio compression, and their corresponding 10 types of distortion intensities. Thus, 150 distorted images corresponding to the original images with high aesthetic quality are obtained;
[0025] Step S33: Randomly select image pairs from the distorted image set and input them into the technical quality network in step S31 to perform the distortion intensity ranking and distortion intensity value prediction tasks;
[0026] The loss function for distortion intensity ranking is as follows:
[0027]
[0028] where f rank represents the distortion intensity ranking model, and m represents the minimum difference value that can be distinguished by the model for the distortion intensities of the distorted images I i and I j , which is set to 0.2.
[0029] The loss function for distortion intensity value prediction is as follows:
[0030]
[0031] where f reg represents the distortion intensity value regression model, and norm(i) is the normalized intensity value, which normalizes all predicted positive intensity values to the range [0,1] and all predicted negative intensity values to [-1,0].
[0032] The loss function of the multi-task network based on distortion intensity ranking and regression is as follows:
[0033]
[0034] Step S34: Repeat step S33 until the loss function converges, and freeze the technical quality network parameters;
[0035] Step S35: Modify the last layer of the technical quality network, add a BatchNorm layer, a Relu layer, and an adaptive pooling layer to obtain a technical quality feature vector;
[0036] Step S36: Input the training set into the modified technical quality branch and output a 128-dimensional technical quality feature vector X tech 。
[0037] Preferably, the specific implementation steps of step S4 are as follows:
[0038] Step S41: Select classification labels covering common themes in the image, convert the classification labels into sentence form, and input them into the text encoder module of the vision-language multimodal large model CLIP to obtain category vectors;
[0039] Step S42: Input the training set images into the image encoder module of CLIP to obtain image features;
[0040] Step S43: Calculate the cosine similarity between the category vector and the image features, and obtain a prediction probability vector through the softmax function as the theme classification feature vector X theme ,and determine the theme classification of the image.
[0041] Preferably, the specific implementation steps of step S5 are as follows:
[0042] Step S51: Concatenate the aesthetic feature vector, technical quality feature vector, and theme classification feature vector obtained in steps S23, S36, and S43, and perform dimensionality reduction processing through a Relu layer, a 1×1 convolutional layer, and a softmax layer to obtain a 10-dimensional feature vector;
[0043] Step S52: According to the loss function of the image aesthetic score distribution prediction network, use backpropagation and the Adam optimizer to update the network parameters. The learning rate is adjusted cosine cyclically using the CosineAnnealingLR strategy. The loss function of the image aesthetic score distribution prediction network is as follows:
[0044]
[0045] Among them, CDF represents the sum of cumulative probabilities N represents the total number of images, and the true and predicted probability mass distribution functions are p and Set r = 2 to measure the Euclidean distance between CDFs;
[0046] The loss function is used to measure the Euclidean distance between the predicted and true aesthetic score distributions;
[0047] Step S53. Repeat Step S51 to Step S52 until the loss value converges, save the network parameters, and the process expression of the entire aesthetic model is:
[0048]
[0049] where F aes represents the entire aesthetic evaluation model of multi-branch fusion, and θ aes represents the parameters of the entire multi-branch aesthetic prediction model. represents concat feature concatenation;
[0050] Step S54. Input the test set images into the trained model, and output the aesthetic score distribution and theme classification of the images.
[0051] An image aesthetic quality evaluation system with multi-branch fusion attributes, including:
[0052] A memory for storing data;
[0053] A processor for executing program instructions;
[0054] And computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the multi-branch fusion-based image aesthetic quality evaluation method described in any one of claims 1-6 is implemented.
[0055] Preferably, an input / output interface for receiving image data to be evaluated and outputting the image aesthetic quality evaluation result.
[0056] Compared with the prior art, the beneficial effects of the present invention are:
[0057] The present invention can effectively fuse aesthetic attribute information such as the composition, style, and theme of images, adaptively extract semantic information for different images, and improve the performance of the image aesthetic quality evaluation method. The aesthetic evaluation result is output in the form of an image aesthetic score distribution, and at the same time, the theme classification result of the image is output, making the aesthetic evaluation result of the image more interpretable. Brief Description of the Drawings
[0058] Figure 1 is one of the implementation flowcharts of the method of the present invention;
[0059] Figure 2 is another implementation flowchart of the method of the present invention;
[0060] Figure 3 is the flowchart of the main aesthetic branch in the present invention;
[0061] Figure 4It is a multi-task flow chart based on distortion intensity ranking and regression in the pre-training of the technical quality branch in the present invention;
[0062] Figure 5 It is the flow chart of the technical quality branch in the present invention;
[0063] Figure 6 It is the flow chart of the theme classification branch in the present invention. Specific embodiments
[0064] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0065] Embodiment 1:
[0066] Please refer to Figures 1 to 6 As shown, the multi-branch fusion image aesthetic quality evaluation method includes the following steps:
[0067] Step S1: Preprocess the data in the aesthetic image dataset, classify and preprocess it according to nine common image sizes, so as to perform batch training subsequently and keep most of the image composition information undamaged;
[0068] Step S2: Design a main branch network with multi-level feature output, and adaptively fuse rich multi-level semantic features for different images by using an ordered weighted average operator;
[0069] Step S3: Design a technical quality branch network, and pre-train by identifying the technical distortion of the image to learn the aesthetic attribute features such as the composition and style of the image;
[0070] Step S4: Design a theme classification branch network, and use a vision-language large model to output the theme classification label and theme attribute features of the image;
[0071] Step S5: Fuse the features of the main branch, technical quality branch, and theme classification branch, and input them into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the image aesthetic quality evaluation result.
[0072] The specific implementation steps of Step S1 are as follows:
[0073] Step S11: Select five common aspect ratios that can widely represent most professional photography images, art works, mobile phone photos, etc., consider horizontal and vertical formats, divide the image sizes into nine categories, and set the corresponding resolutions for each size;
[0074] Step S12: Classify the images in the dataset according to the category closest to their size, and scale them to the corresponding resolution to obtain nine subsets of images with different aspect ratios, so as to meet the requirements of batch training and maintain the image composition information;
[0075] Step S13: Divide the preprocessed images into a training set and a test set according to a preset ratio.
[0076] As can be seen from the above, by preprocessing the data in the aesthetic image dataset and classifying and preprocessing it according to nine common image sizes, this step first selects five common aspect ratios that can widely represent most professional photography images, art works, mobile phone photos, etc., and considers horizontal and vertical formats, divides the image sizes into nine categories in detail, and sets the corresponding resolutions for each size. Subsequently, the images in the dataset are classified according to the category closest to their size and scaled to the corresponding resolution, obtaining nine subsets of images with different aspect ratios. These subsets not only meet the requirements of batch training but also effectively maintain the composition information of the images. Finally, the preprocessed images are scientifically divided into a training set and a test set, providing a high-quality data basis for subsequent steps. Through detailed image size classification and preprocessing, not only the efficiency of data processing is improved, but also the integrity of the composition information of the images during the scaling process is ensured, providing more accurate and effective data support for subsequent model training.
[0077] Embodiment 2:
[0078] Reference Figures 1 to 6 As shown, the multi-branch fusion image aesthetic quality evaluation method includes the following steps:
[0079] Step S1: Preprocess the data in the aesthetic image dataset and classify and preprocess it according to nine common image sizes for subsequent batch training and to keep most of the image composition information undamaged;
[0080] Step S2: Design a main branch network with multi-level feature output, and adaptively fuse rich multi-level semantic features for different images using an ordered weighted average operator;
[0081] Step S3: Design a technical quality branch network to learn the aesthetic attribute features such as the composition and style of the image by identifying the technical distortion of the image and performing pre-training;
[0082] Step S4: Design a theme classification branch network to output the theme classification label and theme attribute features of the image using a vision-language large model;
[0083] Step S5: Integrate the features of the main branch, the technical quality branch, and the subject classification branch, and input them into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the evaluation result of the image aesthetic quality.
[0084] The specific implementation steps of Step S2 are as follows:
[0085] Step S21: Use the image classification network ResNet50 as the feature extraction module to extract feature vectors after four Conv Block convolutional modules, and perform dimensionality reduction operations such as pooling to obtain four 1×1×256-dimensional feature vectors.
[0086] Step S22: Process the feature vectors obtained in Step S21 using an ordered weighted average operator, sort the corresponding values of the four vectors in descending order and assign weights to obtain a 256-dimensional aggregated feature vector. The formula for the ordered weighted average operator is:
[0087]
[0088] where, a i is the i-th value after sorting in descending order, ω i is the corresponding weight, and it is a learnable parameter during the training phase.
[0089] Step S23: Input the training set that has gone through Step S13 into the network model in Step S21 to obtain the image aesthetic feature vector X sem .
[0090] As can be seen from the above, by designing a main branch network with multi-level feature output, this network uses the image classification network ResNet50 as the feature extraction module to extract feature vectors after four Conv Block convolutional modules, and through dimensionality reduction operations such as pooling, four 1×1×256-dimensional feature vectors are obtained. Subsequently, an ordered weighted average operator is used to process these feature vectors, sort the corresponding values of the four vectors in descending order and assign weights, and finally a 256-dimensional aggregated feature vector is obtained. This step realizes the effective extraction and integration of multi-level semantic features of the image. Through the application of the ordered weighted average operator, the network can adaptively fuse rich multi-level semantic features for different images, improving the accuracy and robustness of feature extraction. This adaptive fusion method helps the model better capture the aesthetic information in the image and provides strong support for subsequent aesthetic quality evaluation.
[0091] Example 3:
[0092] Refer to Figures 1 to 6 As shown, the multi-branch fusion image aesthetic quality evaluation method includes the following steps:
[0093] Step S1: Preprocess the data in the aesthetic image dataset, classify and preprocess it according to nine common image sizes for subsequent batch training and ensure that most of the image composition information is not damaged;
[0094] Step S2: Design a main branch network with multi-level feature output, and use the ordered weighted average operator to adaptively fuse rich multi-level semantic features for different images;
[0095] Step S3: Design a technical quality branch network, learn the aesthetic attribute features such as the composition and style of the image by identifying the technical distortion of the image and performing pre-training;
[0096] Step S4: Design a theme classification branch network, and use the vision-language large model to output the theme classification label and theme attribute features of the image;
[0097] Step S5: Fuse the features of the main branch, technical quality branch, and theme classification branch, and input them into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the evaluation result of the image aesthetic quality.
[0098] The specific implementation steps of Step S3 are as follows:
[0099] Step S31: Use Mobilenet-V2 as the basic network for the technical quality evaluation task;
[0100] Step S32: Perform distortion processing on the unlabeled high-definition images with high aesthetic quality. Each distortion d corresponds to a set of distortion intensity values V, and an image set I = {I ν , ν ∈ V} is obtained, where I 0 = I. The distortion set D includes 5 types of technical distortions including JPEG compression, defocus blur, pixelation, Gaussian noise, and impulse noise and their corresponding 6 types of distortion intensities, 6 types of style distortions including brightness, contrast, saturation, exposure, color temperature, and hue and their corresponding 10 types of distortion intensities, and 6 types of composition distortions including rotation, horizontal cropping, vertical cropping, left diagonal cropping, right diagonal cropping, and image aspect ratio compression and their corresponding 10 types of distortion intensities. Thus, 150 distorted images corresponding to the high-aesthetic-quality original images are obtained;
[0101] Step S33: Randomly select image pairs from the distorted image set, input them into the technical quality network in Step S31, and perform the distortion intensity ranking and distortion intensity value prediction tasks;
[0102] The loss function for distortion intensity ranking is as follows:
[0103]
[0104] Among them, f rankDenote the distortion intensity ranking model as, and denote the distorted image as I i and I j The minimum difference value that can be distinguished by the model for the distortion intensity of is set to 0.2.
[0105] The loss function for predicting the distortion intensity value is as follows:
[0106]
[0107] Among them, f reg Denotes the distortion intensity value regression model, and norm(i) is the normalized intensity value, so that all predicted positive intensity values are normalized to the range [0,1] and all predicted negative intensity values are normalized to [-1,0].
[0108] The loss function of the multi-task network based on distortion intensity ranking and regression is as follows:
[0109]
[0110] Step S34: Repeat step S33 until the loss function converges, and freeze the technical quality network parameters;
[0111] Step S35: Modify the last layer of the technical quality network, add a BatchNorm layer, a Relu layer, and an adaptive pooling layer to obtain the technical quality feature vector;
[0112] Step S36: Input the training set into the modified technical quality branch, and output the 128-dimensional technical quality feature vector X tech .
[0113] As can be seen from the above, by designing the technical quality branch network, which is based on Mobilenet-V2 and performs various distortion processes on unlabeled high-definition images with high aesthetic quality, an image set containing various distortion types and distortion intensities is obtained. Subsequently, image pairs are randomly selected from the distorted image set and input into the technical quality network for training. By the distortion intensity ranking and distortion intensity value prediction tasks, the network's ability to identify technical distortions is improved. Finally, the network structure is modified and the technical quality feature vector is output. Through multi-task learning and the processing of various distortion types, the technical quality branch network can learn aesthetic attribute features such as the composition and style of the image, and can effectively identify technical distortions in the image, enabling the model to more comprehensively consider the technical quality factors of the image when evaluating the aesthetic quality of the image, improving the accuracy and reliability of the evaluation.
[0114] Example 4:
[0115] Refer to Figures 1 to 6 As shown, the multi-branch fusion image aesthetic quality evaluation method includes the following steps:
[0116] Step S1: Preprocess the data in the aesthetic image dataset and classify and preprocess it according to nine common image sizes for subsequent batch training and to ensure that most of the image composition information is not damaged;
[0117] Step S2: Design a main branch network with multi-level feature output, and use an ordered weighted average operator to adaptively fuse rich multi-level semantic features for different images;
[0118] Step S3: Design a technical quality branch network, and learn the aesthetic attribute features such as the composition and style of the image by identifying the technical distortion of the image and performing pre-training;
[0119] Step S4: Design a theme classification branch network, and use a vision-language large model to output the theme classification label and theme attribute features of the image;
[0120] Step S5: Fuse the features of the main branch, technical quality branch, and theme classification branch, and input them into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the evaluation result of the image aesthetic quality.
[0121] The specific implementation steps of Step S4 are as follows:
[0122] Step S41: Select classification labels covering common themes in the image, convert the classification labels into sentence form, and input them into the text encoder module of the vision-language multi-modal large model CLIP to obtain category vectors;
[0123] Step S42: Input the training set images into the image encoder module of CLIP to obtain image features;
[0124] Step S43: Calculate the cosine similarity between the category vector and the image feature, and obtain the prediction probability vector through the softmax function as the theme classification feature vector X theme , and determine the theme classification of the image.
[0125] As can be seen from the above, by designing a theme classification branch network, this network uses the text encoder and image encoder modules of the vision-language multi-modal large model CLIP to calculate the cosine similarity between the category vector and the image feature, and obtains the prediction probability vector through the softmax function as the theme classification feature vector, realizing the effective recognition of the image theme classification. By introducing the vision-language multi-modal large model CLIP, the theme classification branch network can accurately identify the theme classification and theme attribute features of the image, enabling the model to consider the theme factors of the image when evaluating the image aesthetic quality, and further improving the comprehensiveness and accuracy of the evaluation.
[0126] Example Five:
[0127] Reference Figures 1 to 6As shown in the figure, the multi-branch fusion image aesthetic quality evaluation method includes the following steps:
[0128] Step S1: Preprocess the data in the aesthetic image dataset and classify and preprocess it according to nine common image sizes for subsequent batch training and to ensure that most of the image composition information is not damaged;
[0129] Step S2: Design a main branch network with multi-level feature output and use an ordered weighted average operator to adaptively fuse rich multi-level semantic features for different images;
[0130] Step S3: Design a technical quality branch network to learn aesthetic attribute features such as the composition and style of the image by identifying the technical distortion of the image and performing pre-training;
[0131] Step S4: Design a theme classification branch network and use a vision-language large model to output the theme classification label and theme attribute features of the image;
[0132] Step S5: Fuse the features of the main branch, technical quality branch, and theme classification branch and input them into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the evaluation result of the image aesthetic quality.
[0133] The specific implementation steps of Step S5 are as follows:
[0134] Step S51: Concatenate the aesthetic feature vector, technical quality feature vector, and theme classification feature vector obtained in Step S23, Step S36, and Step S43, and perform dimensionality reduction processing through a Relu layer, a 1×1 convolutional layer, and a softmax layer to obtain a 10-dimensional feature vector;
[0135] Step S52: According to the loss function of the image aesthetic score distribution prediction network, use backpropagation and the Adam optimizer to update the network parameters, and the learning rate is adjusted cosine-periodically using the CosineAnnealingLR strategy. The loss function of the image aesthetic score distribution prediction network is as follows:
[0136]
[0137] Among them, CDF represents the sum of cumulative probabilities N represents the total number of images, and the true and predicted probability mass distribution functions are p and Set r = 2 to measure the Euclidean distance between CDFs;
[0138] The loss function is used to measure the Euclidean distance between the predicted and true aesthetic score distributions;
[0139] Step S53: Repeat steps S51 to S52 until the loss value converges, and save the network parameters. The process expression of the entire aesthetic model is as follows:
[0140]
[0141] where F aes represents the entire multi-branch fusion aesthetic evaluation model, and θ aes represents the parameters of the entire multi-branch aesthetic prediction model. denotes concat feature concatenation;
[0142] Step S54: Input the test set images into the trained model, and output the image aesthetic score distribution and theme classification.
[0143] As can be seen from the above, by fusing the features of the main branch, technical quality branch, and theme classification branch and inputting them into the aesthetic score distribution prediction module for prediction, in this step, the feature vectors obtained from each branch are first concatenated, and dimensionality reduction processing is performed through the Relu layer, 1×1 convolutional layer, and softmax layer to obtain a 10-dimensional feature vector. Then, according to the loss function of the image aesthetic score distribution prediction network, the network parameters are updated using backpropagation and the Adam optimizer until the loss value converges. Finally, the test set images are input into the trained model to output the image aesthetic score distribution and theme classification. Through the fusion of multi-branch features and the application of the aesthetic score distribution prediction module, step S5 realizes a comprehensive evaluation of the image aesthetic quality. This evaluation method not only considers the technical quality and theme factors of the image but also integrates multi-level semantic features, making the evaluation results more accurate and reliable. At the same time, through the optimization of the loss function and the adjustment of model parameters, the accuracy of the evaluation is further improved.
[0144] An image aesthetic quality evaluation system with multi-branch fusion, characterized by comprising:
[0145] A memory for storing data;
[0146] A processor for executing program instructions;
[0147] And computer program instructions stored on the memory and executable by the processor. When the processor runs the computer program instructions, the image aesthetic quality evaluation method with multi-branch fusion as described in any one of claims 1-6 is implemented.
[0148] An input / output interface for receiving image data to be evaluated and outputting the image aesthetic quality evaluation result.
[0149] As described above, by equipping with a memory, the storage function of data is realized, and information such as image data, algorithm parameters, and historical evaluation results can be saved. At the same time, necessary data support is provided for the processor to ensure the continuity and traceability of the evaluation method. By equipping with a processor, the execution of program instructions is realized, and complex image aesthetic quality evaluation algorithms can be run, achieving the beneficial effects of efficiently and accurately evaluating the image aesthetic quality, improving the real-time performance and accuracy of the evaluation. By storing computer program instructions, the automated execution of the image aesthetic quality evaluation method with multi-branch fusion is realized, reducing human intervention and improving the objectivity and consistency of the evaluation. By adding input and output interfaces, the interaction function between the system and the external environment is realized, and the image data input by the user can be conveniently received, and the evaluation result can be output to the user, reducing the usage threshold of the system and improving the user's satisfaction and experience.
[0150] Although the embodiments of the present invention have been shown and described, for those of ordinary skill in the art, it can be understood that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present invention. The scope of the present invention is defined by the appended claims and their equivalents.
Claims
1. A multi-branch fusion image aesthetic quality evaluation method, characterized in that: The steps include: Step S1, preprocessing the data in the aesthetic image dataset, and performing classification preprocessing according to nine common image sizes, so as to facilitate subsequent batch training and keep most of the image composition information intact; Step S2, designing a main branch network with multi-level feature output, and using an ordered weighted average operator to adaptively fuse rich multi-level semantic features for different images; Step S3, designing a technical quality branch network to learn the aesthetic attribute characteristics of the image, such as composition and style, by identifying the technical distortion of the image and performing pre-training; Step S4: design a topic classification branch network, and use the visual language model to output the topic classification labels and topic attribute features of the image; Step S5: The features of the main branch, the technical quality branch, and the subject classification branch are integrated and input into the aesthetic score distribution prediction module to obtain the aesthetic score distribution of the image as the image aesthetic quality evaluation result.
2. The multi-branch fusion image aesthetic quality evaluation method according to claim 1, characterized in that: The specific implementation steps of step S1 are as follows: Step S11, selecting five common aspect ratios that can widely represent most professional photographic images, artworks, mobile phone photos, etc., and considering the horizontal and vertical frames, dividing the image sizes into nine categories, and setting the resolution corresponding to each size; Step S12: Classify the images in the data set according to the category closest to their size and scale them to the corresponding resolution to obtain nine image subsets with different aspect ratios to meet the batch training requirements and maintain the image composition information; Step S13: Divide the preprocessed images into a training set and a test set according to a preset ratio.
3. The multi-branch fusion image aesthetic quality evaluation method according to claim 1, characterized in that: The specific implementation steps of step S2 are as follows: Step S21, using the image classification network ResNet50 as a feature extraction module, extracting feature vectors from four Conv Block convolution modules, and performing dimensionality reduction operations such as pooling to obtain four 1×1×256-dimensional feature vectors; Step S22: Use the ordered weighted average operator to process the feature vector obtained in step S21, arrange each corresponding value of the four vectors in descending order and assign weights to obtain a 256-dimensional aggregate feature vector; the ordered weighted average operator formula is: Among them, a i is the i-th value after descending order, ω i is the corresponding weight and is a learnable parameter in the training phase; Step S23: Input the training set obtained in step S13 into the network model in step S21 to obtain the image aesthetic feature vector X sem .
4. The multi-branch fusion image aesthetic quality evaluation method according to claim 1, characterized in that: The specific implementation steps of step S3 are as follows: Step S31, using Moilenet-V2 as the basic network for the technical quality evaluation task; Step S32: Distort the unlabeled high-definition image with high aesthetic quality, and each distortion d corresponds to a set of distortion intensity values V, so as to obtain an image set I=I consisting of an original image and a distorted image. ν ,ν∈V, I0=I, the distortion set D includes 5 technical distortions and 6 corresponding distortion intensities, including JPEG compression, defocus blur, pixelation, Gaussian noise and impulse noise, 6 style distortions and 10 corresponding distortion intensities, including brightness, contrast, saturation, exposure, color temperature and hue, 6 composition distortions and 10 corresponding distortion intensities, including rotation, horizontal cropping, vertical cropping, left diagonal cropping, right diagonal cropping and image aspect ratio compression, thus obtaining 150 distorted images corresponding to the original image with high aesthetic quality; Step S33, randomly selecting an image pair from the distorted image set, inputting the image pair into the technical quality network of step S31, and performing distortion intensity sorting and distortion intensity value prediction tasks; The loss function for distortion intensity ranking is as follows: Among them, f rank represents the distortion intensity sorting model, and m represents the distorted image I i and I j The minimum difference value of the distortion intensity that can be distinguished by the model is set to 0.2; The loss function for distortion intensity value prediction is as follows: Among them, f reg represents the distorted intensity value regression model, normi is the normalized intensity value, so that all predicted positive intensity values are normalized to the range [0,1] and all predicted negative intensity values are normalized to [-1,0]; The loss function of the multi-task network based on distortion intensity ranking and regression is as follows: Step S34, repeating step S33 until the loss function converges, freezing the technical quality network parameters; Step S35, modify the last layer of the technical quality network, add a BatchNorm layer, a Relu layer, and an adaptive pooling layer, and obtain a technical quality feature vector; Step S36: Input the training set into the modified technical quality branch and output a 128-dimensional technical quality feature vector X tech .
5. The multi-branch fusion image aesthetic quality evaluation method according to claim 1, characterized in that: The specific implementation steps of step S4 are as follows: Step S41, select classification labels covering common topics in the image, convert the classification labels into sentence form, input them into the text encoder module of the visual language multimodal large model CLIP, and obtain the category vector; Step S42, input the training set image into the image encoder module of CLIP to obtain image features; Step S43: Calculate the cosine similarity between the category vector and the image feature, and obtain the predicted probability vector through the softmax function as the topic classification feature vector X theme , and determine the subject classification of the image.
6. The multi-branch fusion image aesthetic quality evaluation method according to claim 1, characterized in that: The specific implementation steps of step S5 are as follows: Step S51, concatenate the aesthetic feature vector, technical quality feature vector, and topic classification feature vector obtained in step S23, step S36, and step S43, and reduce the dimension into a 10-dimensional feature vector through a Relu layer, a 1×1 convolution layer, and a softmax layer; Step S52: According to the loss function of the image aesthetic score distribution prediction network, back propagation and Adam optimizer are used to update the network parameters, and the learning rate is adjusted by cosine periodicity using the CosineAnnealingLR strategy. The loss function of the image aesthetic score distribution prediction network is as follows: Among them, CDF represents the sum of cumulative probabilities N represents the total number of images, and the probability mass distribution functions of the real and predicted images are p and Setting r = 2 is used to measure the Euclidean distance between CDFs; The loss function is used to measure the Euclidean distance between the predicted and true aesthetic rating distributions; Step S53, repeat step S51 to step S52 until the loss value converges, save the network parameters, and the process expression of the entire aesthetic model is: Among them, F aes represents the entire multi-branch fusion aesthetic evaluation model, θ aes Represents the parameters of the entire multi-branch aesthetic prediction model, Indicates concat feature concatenation; Step S54: Input the test set images into the trained model and output the image aesthetic score distribution and topic classification.
7. The multi-branch fusion image aesthetic quality evaluation system is characterized by: include: A memory for storing data; A processor for executing program instructions; And computer program instructions stored in the memory and capable of being executed by the processor, when the processor executes the computer program instructions, the multi-branch fusion image aesthetic quality evaluation method according to any one of claims 1-6 is implemented.
8. The multi-branch fusion image aesthetic quality evaluation system according to claim 7, characterized in that: Also includes: The input and output interface is used to receive the image data to be evaluated and output the image aesthetic quality evaluation results.
Citation Information
Cited By
Image aesthetic evaluation system and method, electronic equipment and storage medium
CN121259355A