Image evaluation method, device, electronic device and storage medium

By dividing image aesthetic attributes into multiple groups and using multi-scale feature extraction and dynamic weighted loss function to train the model, the problem of insufficient accuracy of image aesthetic evaluation in the prior art is solved, and higher accuracy and interpretability of aesthetic quality prediction are achieved.

CN115035083BActive Publication Date: 2025-08-26GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210729978.6
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-24
Publication Date
2025-08-26
Estimated Expiration
2042-06-24

AI Technical Summary

Technical Problem

Existing machine learning or deep learning-based image aesthetic evaluation methods fail to effectively consider different aesthetic attributes and the different cognitive stages of overall aesthetic quality in aesthetic perception, resulting in insufficient accuracy in image aesthetic quality prediction.

Method used

The aesthetic attributes of the image are divided into multiple aesthetic attribute groups, and the trained aesthetic attribute prediction model predicts the aesthetic quality score of the image based on the aesthetic characteristics of multiple aesthetic attribute groups, reflecting the stage of aesthetic perception, and using multi-scale feature extraction and dynamic weighted loss function to train the model.

Benefits of technology

It improves the accuracy of image aesthetic quality prediction, is more in line with human aesthetic process, and enhances the interpretability and accuracy of image aesthetic evaluation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115035083B_ABST
    Figure CN115035083B_ABST
Patent Text Reader

Abstract

The present application discloses an image evaluation method, apparatus, electronic device, and storage medium. The method includes: obtaining an image to be evaluated; inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, obtaining aesthetic features for each of multiple aesthetic attribute groups corresponding to the image to be evaluated, output by the multi-scale feature extraction module; inputting the aesthetic features of each aesthetic attribute group into a prediction module of the aesthetic attribute prediction model, and obtaining an aesthetic quality score for the image to be evaluated, output by the prediction module. The aesthetic attribute prediction model is configured to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated. This improves the accuracy of image aesthetic quality prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of image processing, and specifically relates to an image evaluation method, device, electronic device, and storage medium. Background Art

[0002] Image aesthetic evaluation requires humans to assess the beauty of an image from an artistic perspective. Accurately judging images from this aesthetic perspective requires extensive training. Therefore, subjective image aesthetic evaluation is often abstract and difficult to teach, hindering its practical application in real-time systems. The rapid development of machine learning in recent years has significantly facilitated the development of reproducible, objective image aesthetic evaluation methods. However, the accuracy of these image evaluation methods remains to be improved. Summary of the Invention

[0003] In view of the above problems, the present application proposes an image evaluation method, device, electronic device and storage medium to improve the above problems.

[0004] In a first aspect, an embodiment of the present application provides an image evaluation method, the method comprising: obtaining an image to be evaluated; inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and obtaining aesthetic features of each aesthetic attribute group in a plurality of aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; inputting the aesthetic features of each aesthetic attribute group into a prediction module of the aesthetic attribute prediction model, and obtaining an aesthetic quality score of the image to be evaluated output by the prediction module, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated.

[0005] In a second aspect, an embodiment of the present application provides an image evaluation device, comprising: an image acquisition unit for acquiring an image to be evaluated; a feature acquisition unit for inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and obtaining the aesthetic features of each aesthetic attribute group in a plurality of aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; a score acquisition unit for inputting the aesthetic features of each aesthetic attribute group into a prediction module of the aesthetic attribute prediction model, and obtaining the aesthetic quality score of the image to be evaluated output by the prediction module, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated.

[0006] In a third aspect, an embodiment of the present application provides an electronic device comprising one or more processors and a memory; one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs are configured to execute the above-mentioned method.

[0007] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium, in which program code is stored, wherein the above method is executed when the program code is run.

[0008] The embodiment of the present application provides an image evaluation method, device, electronic device and storage medium. First, an image to be evaluated is obtained, and then the image to be evaluated is input into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and the aesthetic features of each aesthetic attribute group in the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module are obtained, and then the aesthetic features of each aesthetic attribute group are input into the prediction module of the aesthetic attribute prediction model, and the aesthetic quality score of the image to be evaluated output by the prediction module is obtained, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated. Through the above method, according to the multi-level process of aesthetic cognition, the aesthetic attributes of the image are divided into multiple aesthetic attribute groups, and then the aesthetic quality score of the image is predicted based on the aesthetic features corresponding to each of the multiple aesthetic attribute groups through the trained aesthetic attribute prediction model, which reflects the stage of aesthetic perception, is more in line with the human aesthetic process, and improves the accuracy of image aesthetic quality prediction. BRIEF DESCRIPTION OF THE DRAWINGS

[0009] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the following briefly introduces the drawings required for use in the description of the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative work.

[0010] Figure 1 A schematic diagram of an application scenario of an image evaluation method proposed in an embodiment of the present application is shown;

[0011] Figure 2 A schematic diagram of an application scenario of an image evaluation method proposed in an embodiment of the present application is shown;

[0012] Figure 3 A flowchart of an image evaluation method proposed in one embodiment of the present application is shown;

[0013] Figure 4A flowchart of an image evaluation method proposed in another embodiment of the present application is shown;

[0014] Figure 5 A flowchart of step S220 in another embodiment of the present application is shown;

[0015] Figure 6 A flowchart of an image evaluation method proposed in another embodiment of the present application is shown;

[0016] Figure 7 A schematic diagram of the structure of an aesthetic attribute prediction model in another embodiment of the present application is shown;

[0017] Figure 8 A flowchart of an image evaluation method in another embodiment of the present application is shown;

[0018] Figure 9 The following is a structural block diagram of an image evaluation device proposed in an embodiment of the present application;

[0019] Figure 10 The following is a structural block diagram of an image evaluation device proposed in an embodiment of the present application;

[0020] Figure 11 A structural block diagram of an electronic device or server for executing an image evaluation method according to an embodiment of the present application is shown in real time in the present application;

[0021] Figure 12 The present invention shows a storage unit in real time for storing or carrying program codes for implementing the image evaluation method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0022] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0023] With the rapid development of mobile internet and the widespread adoption of smartphones, the amount of visual content, such as images and videos, is increasing rapidly. The perception and understanding of this visual content has become a multidisciplinary research topic, encompassing computer vision, computational photography, and human psychology. Image aesthetics assessment is a recent hot topic within computer vision perception and understanding. Image aesthetics aims to use computer systems to simulate the way humans perform aesthetic perception, computation, and evaluation of images. Humans appreciate images by making aesthetic decisions based on visual stimuli. Therefore, using computers to mimic this human ability presents challenges across multiple interdisciplinary fields, including image processing, computer vision, and psychology. Image aesthetics reflects the human visual pursuit and yearning for beauty, and therefore, visual aesthetics assessment holds significant significance in fields such as photography, videography, advertising design, and artwork production. In recent years, this field has attracted considerable research interest.

[0024] Image aesthetic evaluation requires humans to assess the beauty of images from an artistic perspective. Accurately judging images from this aesthetic perspective requires extensive training. Therefore, subjective image aesthetic evaluation is often abstract and difficult to teach, making it unsuitable for real-time systems. The rapid development of machine learning in recent years has significantly facilitated the development of reproducible, objective image aesthetic evaluation methods. Machine learning, particularly deep learning systems, can efficiently and accurately mimic human thought processes. Therefore, utilizing machine learning or deep learning methods for image aesthetic evaluation is an important research topic.

[0025] However, the inventors found in their research on relevant image evaluation methods that the current image aesthetic evaluation methods based on machine learning or deep learning mainly place all aesthetic attributes and overall aesthetic quality in the same learning stage and predict them simultaneously, without taking into account that different aesthetic attributes and overall aesthetic quality are at different cognitive stages in aesthetic perception. As a result, when predicting the aesthetic quality of an image, the accuracy of the prediction needs to be improved.

[0026] Therefore, the inventors proposed the image evaluation method, device, electronic device and storage medium in the present application. First, the image to be evaluated is obtained, and then the image to be evaluated is input into the multi-scale feature extraction module of the trained aesthetic attribute prediction model, and the aesthetic features of each aesthetic attribute group in the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module are obtained, and then the aesthetic features of each aesthetic attribute group are input into the prediction module of the aesthetic attribute prediction model to obtain the aesthetic quality score of the image to be evaluated output by the prediction module, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated. Through the above method, the aesthetic attributes of the image are divided into multiple aesthetic attribute groups according to the multi-level process of aesthetic cognition, and the aesthetic quality score of the image is predicted by the trained aesthetic attribute prediction model, which reflects the stage of aesthetic perception, is more in line with the human aesthetic process, and improves the accuracy of image aesthetic quality prediction.

[0027] In the embodiment of the present application, the image evaluation method provided can be executed by an electronic device. In this manner, all steps in the image evaluation method provided by the embodiment of the present application can be executed by the electronic device. For example, Figure 1 As shown, the processor of the electronic device 100 executes the steps of acquiring an image to be evaluated; inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and acquiring the aesthetic features of each aesthetic attribute group in a plurality of aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; and inputting the aesthetic features of each aesthetic attribute group into a prediction module of the aesthetic attribute prediction model, and acquiring the aesthetic quality score of the image to be evaluated output by the prediction module.

[0028] Furthermore, the image evaluation method provided in the embodiment of the present application can also be executed by a server (cloud). Correspondingly, in this method executed by the server, the electronic device can obtain the image to be evaluated and synchronously send the image to be evaluated to the server. The server then inputs the image to be evaluated into the multi-scale feature extraction module of the trained aesthetic attribute prediction model in real time, obtains the aesthetic features of each aesthetic attribute group in the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; inputs the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model, and obtains the aesthetic quality score of the image to be evaluated output by the prediction module.

[0029] Alternatively, the electronic device and the server may collaborate to perform the process. In this collaborative process, some steps of the image evaluation method provided in the embodiment of the present application are performed by the electronic device, while other steps are performed by the server.

[0030] For example, Figure 2 As shown, the electronic device 100 can execute the image evaluation method including: obtaining an image to be evaluated, and then the server 200 executes inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and obtaining the aesthetic features of each aesthetic attribute group in the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; inputting the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model, and obtaining the aesthetic quality score of the image to be evaluated output by the prediction module.

[0031] It should be noted that in this method of collaborative execution by the electronic device and the server, the steps respectively executed by the electronic device and the server are not limited to the methods introduced in the above examples. In actual applications, the steps respectively executed by the electronic device and the server can be dynamically adjusted according to actual conditions.

[0032] It should be noted that the electronic device 100 can be used for Figure 1 and Figure 2 In addition to the smartphone shown in FIG, it can also be a car device, a wearable device, a tablet computer, a laptop computer, a smart speaker, etc. The server 120 can be an independent physical server, or a server cluster or a distributed system composed of multiple physical servers.

[0033] The embodiments of the present application will be described in detail below with reference to the accompanying drawings.

[0034] See also Figure 3 , an image evaluation method provided in the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0035] Step S110: Acquire an image to be evaluated.

[0036] In the embodiments of the present application, the image to be evaluated can be any image that is acquired. Optionally, the image to be evaluated can be an image acquired in real time by an image acquisition device, or can be an image acquired from a storage area (e.g., a cloud server) that pre-stores many images based on an image acquisition request. Of course, the image to be evaluated can also be an image that is searched from a website, and this is not specifically limited here.

[0037] For example, as one approach, if the images to be evaluated are pre-stored in a cloud server, a large number of images need to be stored in the cloud server first. Then, when an image to be evaluated is needed, an image acquisition request can be sent to the cloud server. Upon receiving the image acquisition request, the cloud server returns the image corresponding to the image acquisition request as the image to be evaluated. When storing the images, an image identifier can be assigned to each image, and a corresponding relationship between the image identifier and the image can be established. When an image acquisition request is sent to the cloud server, the image acquisition request can include the specified image identifier. Upon receiving the image acquisition request carrying the specified image identifier, the cloud server can return the image corresponding to the specified image identifier as the image to be evaluated. Optionally, when establishing the corresponding relationship between the image identifier and the image, an application scenario identifier can be pre-assigned to each image. When establishing the corresponding relationship, a corresponding relationship between the application scenario identifier, image identifier, and image can be established. When acquiring the corresponding image to be evaluated based on the application scenario, the corresponding application scenario identifier can be searched first, followed by the corresponding image identifier, thereby finding a specific image in a specific application scenario as the image to be evaluated.

[0038] As another approach, when the image to be evaluated is captured in real time, if a designated application is detected to be launched, an image acquisition device may begin capturing the image in real time as the image to be evaluated. The designated application may be an application with image processing capabilities, and the image acquisition device may be a camera, a webcam, or other device with image acquisition capabilities. For example, if a designated application is detected to be launched in an electronic device, the image acquisition device in the electronic device may capture the image in real time as the image to be evaluated.

[0039] Step S120: inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and obtaining aesthetic features of each of the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module.

[0040] In this embodiment of the present application, each image to be evaluated corresponds to multiple aesthetic attribute groups, each of which includes at least one aesthetic attribute. Aesthetic attributes are attributes used to describe an image, and may include elements of balance, content, color harmony, depth of field, lighting, objects, the rule of thirds, color vividness, repetitiveness, symmetry, motion blur, and so on. Aesthetic features are characteristics of the image that correspond to the aesthetic attributes; that is, one aesthetic attribute corresponds to one aesthetic feature.

[0041] As one approach, after obtaining an image to be evaluated, the image can be input into a multi-scale feature extraction module of a pre-trained aesthetic attribute prediction model. The multi-scale feature extraction module then extracts aesthetic features from the image to be evaluated. When the multi-scale feature extraction module extracts aesthetic features from the image to be evaluated, it can output a set of aesthetic features corresponding to each of the multiple aesthetic attribute groups corresponding to the image to be evaluated.

[0042] Step S130: inputting the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model, and obtaining the aesthetic quality score of the image to be evaluated output by the prediction module.

[0043] The aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated and the aesthetic attribute scores corresponding to the aesthetic attributes included in each aesthetic attribute group based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated.

[0044] In this embodiment of the present application, each aesthetic attribute of the image being evaluated can be associated with an aesthetic attribute score. In other words, the number of aesthetic attribute scores is equal to the number of aesthetic attributes associated with the image being evaluated. The aesthetic quality score of the image being evaluated is the overall aesthetic score of the image being evaluated, and each image being evaluated is assigned an aesthetic quality score.

[0045] An image evaluation method provided in an embodiment of the present application divides the aesthetic attributes of an image into multiple aesthetic attribute groups according to the multi-level process of aesthetic cognition, and then predicts the aesthetic quality score of the image based on the aesthetic features corresponding to each of the multiple aesthetic attribute groups through a trained aesthetic attribute prediction model. This reflects the stage-by-stage nature of aesthetic perception, is more in line with the human aesthetic process, and improves the accuracy of image aesthetic quality prediction.

[0046] See also Figure 4 , an image evaluation method provided in the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0047] Step S210: Acquire an aesthetic attribute dataset, where the aesthetic attribute dataset includes multiple images, aesthetic attribute scores corresponding to each image, and aesthetic quality scores corresponding to each image.

[0048] In an embodiment of the present application, the aesthetic attribute scores corresponding to each image are obtained by averaging multiple first scores corresponding to the aesthetic attributes of each image; the aesthetic quality score corresponding to each image is obtained by averaging multiple second scores corresponding to each image; wherein, the first scores are scores of different evaluators on the aesthetic attributes of the image, and the second scores are scores of different evaluators on the aesthetic quality of the image.

[0049] As a method, each image in the aesthetic attribute dataset is evaluated simultaneously by multiple evaluators. Different evaluators may have different evaluations of each aesthetic attribute of the image, and different evaluators may also have different evaluations of the aesthetic quality of the image. Therefore, when determining the aesthetic attribute score corresponding to each image and the aesthetic quality score corresponding to each image, the average of the scores of multiple evaluators can be calculated.

[0050] For example, if the aesthetic attributes include eight types, such as balance elements, content, color harmony, depth of field, light, object, rule of thirds, and color vividness, then for image I in the aesthetic attribute dataset, a , the corresponding evaluator b has an opinion on image I a The first score of the eight aesthetic attributes is And for image I a The second rating is Where a = 1, 2, ..., n, n is the number of images included in the aesthetic attribute dataset, b = 1, 2, ..., m, m is the number of evaluators who evaluated each image.

[0051] For the convenience of calculation, in the embodiment of the present application, the multiple first scores corresponding to the image can be normalized, and the multiple second scores corresponding to the image can be normalized. Optionally, in the embodiment of the present application, the multiple first scores corresponding to the image can be normalized to between [-1, 1], and the multiple second scores corresponding to the image can be normalized to between [0, 1]. For example, as described above, the image I a The 8 first scores corresponding to the 8 aesthetic attributes Normalize to [-1,1], and transform image I a The corresponding second score Normalized to [0,1].

[0052] After normalizing the first score corresponding to each image, the average of multiple first scores corresponding to a certain aesthetic attribute of each image is used as the aesthetic attribute score of the image. The calculation formula can be as follows: Where i represents the i-th aesthetic attribute of the image, j represents the j-th evaluator, represents the aesthetic attribute score of the i-th aesthetic attribute of the image.

[0053] After normalizing the second score corresponding to each image, the average of multiple second scores corresponding to each image is used as the aesthetic quality score of the image. The calculation formula can be as follows: Among them, S a is the aesthetic quality score of image j for the jth evaluator.

[0054] Optionally, obtaining the aesthetic attribute dataset may include the following steps: obtaining the average of multiple first scores corresponding to each aesthetic attribute of each image as the aesthetic attribute score corresponding to each image, wherein the first score is the score of each aesthetic attribute of the image by different evaluators; obtaining the average of multiple second scores corresponding to each image as the aesthetic quality score corresponding to each image, wherein the second score is the score of the aesthetic quality of the image by different evaluators; and obtaining the aesthetic attribute dataset based on each image, the aesthetic attribute score corresponding to each image, and the aesthetic quality score corresponding to each image.

[0055] Step S220: Based on the aesthetic attribute data set and the preset loss function, the aesthetic attribute prediction model to be trained is trained until a training end condition is met, thereby obtaining the trained aesthetic attribute prediction model.

[0056] In the embodiment of the present application, in order to make the aesthetic attribute score and aesthetic quality score output by the aesthetic attribute prediction model consistent with the actual aesthetic attribute score and aesthetic quality score, a dynamically weighted preset loss function can be used. total =L A +λL S ,in, In the above formula, g represents the aesthetic attribute group, N is the number of aesthetic attributes, and w i Represents the weight of the i-th aesthetic attribute of the image, which is a dynamically changing learnable parameter, y i and are the true and predicted scores of the i-th aesthetic attribute of the image, y and are the true and predicted results of the aesthetic quality scores of the images, L total is the overall loss, L A is the loss of aesthetic attributes, L S is the aesthetic quality loss, λ is a dynamically changing learnable parameter, and L g is the aesthetic attribute group loss.

[0057] The training end condition may be that the loss value of a preset loss function reaches a preset value, or that the number of iterations of the model reaches a preset number of iterations, which is not specifically limited here.

[0058] As a way, such as Figure 5 As shown, step S220 may further include the following steps:

[0059] Step S221: scaling the size of the image in the aesthetic attribute dataset to a preset size to obtain a scaled aesthetic attribute dataset.

[0060] In this embodiment of the present application, since the input size of the aesthetic attribute prediction model to be trained is fixed, it is necessary to perform a scaling operation on images of different sizes. This scaling operation is used to scale all images in the aesthetic attribute dataset to a preset size. Optionally, the images in the aesthetic attribute dataset are scaled to a size of 299*299.

[0061] Step S222: performing a random vertical flip operation with a preset probability on the image in the scaled aesthetic attribute dataset to obtain an enhanced aesthetic attribute dataset.

[0062] In the embodiment of the present application, the preset probability is a pre-set probability for performing a vertical flip operation on the scaled image. Optionally, the preset probability may be 0.5. A random vertical flip operation with a probability of 0.5 is performed on the scaled image to enhance the image.

[0063] Step S223: normalizing the pixel values ​​of the image in the enhanced aesthetic attribute dataset to obtain a normalized aesthetic attribute dataset.

[0064] In the embodiment of the present application, for the convenience of calculation, the pixel values ​​of the images in the enhanced aesthetic attribute dataset may be normalized to between [0, 1].

[0065] Step S224: Based on the normalized aesthetic attribute data set and the preset loss function, the aesthetic attribute prediction model to be trained is trained until a training end condition is met, thereby obtaining the trained aesthetic attribute prediction model.

[0066] After obtaining a normalized aesthetic attribute dataset using the aforementioned method, the images in the normalized aesthetic attribute dataset are input into the aesthetic attribute prediction model to be trained. The Adam optimization algorithm is then used to iteratively optimize the preset loss function by applying the images in the normalized aesthetic attribute dataset until the calculated loss value no longer decreases, resulting in a trained aesthetic attribute prediction model. For any input image, the trained aesthetic attribute prediction model can determine the corresponding aesthetic quality score for that input image, as well as the aesthetic attribute scores for each aesthetic attribute.

[0067] Step S230: Acquire the image to be evaluated.

[0068] Step S240: inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and obtaining aesthetic features of each of the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module.

[0069] Step S250: inputting the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model, and obtaining the aesthetic quality score of the image to be evaluated output by the prediction module.

[0070] The present application provides an image evaluation method in which an aesthetic attribute prediction model can predict the aesthetic attribute scores of various aesthetic attributes of an image, providing a rich set of aesthetic elements and significantly enhancing the interpretability of image aesthetics. Furthermore, a dynamically weighted loss function is used to train the pre-trained aesthetic attribute prediction model, enabling the trained model to predict both the aesthetic attribute scores and the aesthetic quality score of the image. This avoids the errors introduced by using prior knowledge to determine the aesthetic attribute loss weights, adaptively reconciles the relationship between various aesthetic attributes and aesthetic quality, facilitates rapid model convergence, and saves model training time.

[0071] See also Figure 6 , an image evaluation method provided in the embodiment of the present application is applied to Figure 1 or Figure 2 The electronic device or server shown, the method includes:

[0072] Step S310: Acquire an aesthetic attribute dataset, where the aesthetic attribute dataset includes multiple images, aesthetic attribute scores corresponding to each image, and aesthetic quality scores corresponding to each image.

[0073] Step S320: Obtaining the Spearman correlation coefficient between each aesthetic attribute score and the aesthetic quality score corresponding to each image in the aesthetic attribute dataset.

[0074] In the embodiment of the present application, aesthetic experience is a multi-level information processing process, and different aesthetic attributes play a role in different stages of aesthetic perception. Therefore, in order to more effectively predict the aesthetic attributes of an image, the various aesthetic attributes of the image should first be grouped, and hierarchical predictions should be performed on this basis. The aesthetic attribute dataset in the embodiment of the present application is an existing AADB dataset, which contains annotations of 11 aesthetic attributes. The 11 aesthetic attributes are: balance elements, content, color harmony, depth of field, light, objects, rule of thirds, color vividness, repeatability, symmetry, and motion blur. Among them, the labels of repeatability, symmetry, and motion blur are almost zero, so these three aesthetic attributes are not used in the embodiment of the present application.

[0075] According to the psychological model of human aesthetic perception, content attributes are at a higher cognitive stage in the aesthetic process, so they are considered as a high-level independent aesthetic attribute group. The remaining seven aesthetic attributes can be grouped according to their importance relative to aesthetic quality. Specifically, a statistical study can be conducted based on the Spearman Rank Order Correlation Coefficient (SROCC) between the aesthetic quality scores of the images included in the AADB dataset and the scores of each aesthetic attribute. The Spearman correlation coefficients between the aesthetic attribute scores and the aesthetic quality scores of the images in the AADB dataset are shown in Table 1:

[0076] Table 1

[0077]

[0078] Step S330: Based on the Spearman correlation coefficient, the aesthetic attributes corresponding to the image are divided into multiple aesthetic attribute groups, wherein the correlation between each aesthetic attribute score and the aesthetic quality score is different in different aesthetic attribute groups.

[0079] In the embodiment of the present application, the aesthetic attributes corresponding to the image can be divided into three aesthetic attribute groups based on the Spearman correlation coefficient. Specifically, as can be seen from Table 1 above, the four aesthetic attributes such as balance elements, rule of thirds, depth of field, and object can be divided into one aesthetic attribute group, and the correlation between these four aesthetic attributes and the aesthetic quality score is relatively low. At the same time, color vividness, color harmony, and light can be divided into another aesthetic attribute group, and the aesthetic attribute scores of the aesthetic attributes in this aesthetic attribute group have a moderate correlation with the aesthetic quality score. In addition, it can be seen from Table 1 that the score of the content attribute has the highest correlation with the aesthetic quality score, so it is reasonable to arrange it in the highest-level aesthetic attribute group.

[0080] Of course, the aesthetic attributes included in the three aesthetic attribute groups may also be different from the above groupings. For example, balance elements, the rule of thirds, and depth of field may be grouped into one aesthetic attribute group, objects, color vividness, color harmony, and light may be grouped into one aesthetic attribute group, and content may be grouped into another aesthetic attribute group.

[0081] Optionally, the aesthetic attributes corresponding to the image may be divided into more or fewer than three aesthetic attribute groups based on the Spearman correlation coefficient, which is not specifically limited herein. For example, the aesthetic attributes corresponding to the image may be divided into four aesthetic attribute groups based on the Spearman correlation coefficient. Specifically, balance elements and the rule of thirds may be divided into one aesthetic attribute group, depth of field may be divided into one aesthetic attribute group, objects, color vividness, color harmony, and light may be divided into one aesthetic attribute group, and content may be divided into one aesthetic attribute group.

[0082] Step S340: Based on the aesthetic attribute data set and the preset loss function, the aesthetic attribute prediction model to be trained is trained until a training end condition is met, thereby obtaining the trained aesthetic attribute prediction model.

[0083] Step S350: Acquire the image to be evaluated.

[0084] Step S360: inputting the image to be evaluated into the multiple residual block groups, and obtaining multiple groups of aesthetic features corresponding to the image to be evaluated output by the multiple residual block groups, wherein each residual block group outputs a group of aesthetic features, and each aesthetic attribute group corresponds to a group of aesthetic features.

[0085] In an embodiment of the present application, the multi-scale feature extraction module includes multiple residual block groups, the prediction module includes multiple first fully connected layers, a second fully connected layer, a first output layer connected to the multiple first fully connected layers, and a second output layer connected to the second fully connected layer, wherein each residual block group corresponds to a first fully connected layer.

[0086] In an embodiment of the present application, the aesthetic attribute prediction model uses a ResNet50 network model as its base network model. The aesthetic attribute prediction model includes 16 residual blocks, each of which is used to extract aesthetic features from an image. The output of each residual block serves as the input to the next residual block. Optionally, the 16 residual blocks can be divided into multiple residual block groups, with the number of residual block groups being the same as the number of aesthetic attribute groups. Each residual block group is used to extract a set of aesthetic features corresponding to a set of aesthetic attribute groups for an image, and the aesthetic features output by each residual block in each residual block group are used as a set of aesthetic features.

[0087] As a way, if the aesthetic attributes of the image are divided into three aesthetic attribute groups based on the Spearman correlation coefficient, including the first aesthetic attribute group, the second aesthetic attribute group, and the third aesthetic attribute group; then the 16 residual blocks included in the aesthetic attribute prediction model can be divided into three residual block groups, the 1st to 10th residual blocks are divided into one residual block group, the 11th to 13th residual blocks are divided into one residual block group, and the 14th to 16th residual blocks are divided into one residual block group.

[0088] Among them, the aesthetic features of the images extracted by the 1st to 10th residual blocks are the aesthetic features corresponding to the first aesthetic attribute group, the aesthetic features of the images extracted by the 11th to 13th residual blocks are the aesthetic features corresponding to the second aesthetic attribute group, and the aesthetic features of the images extracted by the 14th to 16th residual blocks are the aesthetic features corresponding to the third aesthetic attribute group.

[0089] The multiple first fully connected layers of the aesthetic attribute prediction model can be respectively a first fully connected layer 1, a first fully connected layer 2 and a first fully connected layer 3, wherein the first fully connected layer 1 can be composed of 5888 nodes, and the first fully connected layer 1 is connected after the 1st to 10th residual blocks; the first fully connected layer 2 can be composed of 3072 nodes, and the first fully connected layer 2 is connected after the 11th to 13th residual blocks; the first fully connected layer 3 can be composed of 6144 nodes, and the first fully connected layer 3 is connected after the 14th to 16th residual blocks; the second fully connected layer of the aesthetic attribute feature model can be composed of 15104 nodes, and the second fully connected layer is connected after the 1st to 16th residual blocks.

[0090] Optionally, the aesthetic attribute prediction model further includes a first output layer and a second output layer. The first output layer is connected after the first fully connected layer 1, the first fully connected layer 2, and the first fully connected layer 3. To ensure that the predicted scores are between [-1, 1], the Tanh activation function can be used as the activation function of the first output layer. The second output layer is connected after the second fully connected layer. To ensure that the predicted scores are between [0, 1], the Sigmoid activation function can be used as the activation function of the second output layer. The first output layer is used to output the aesthetic attribute scores of the image, and the second output layer is used to output the aesthetic quality score of the image. The second fully connected layer and multiple first fully connected layers are each connected after a BN layer and a Dropout layer.

[0091] Optionally, the first fully connected layer 1, the first fully connected layer 2, the first fully connected layer 3 and the second fully connected layer may also be composed of other numbers of nodes; and the division of the 16 residual blocks may also be other division methods, which are not specifically limited here.

[0092] Optional, such as Figure 7 As shown in , each residual block is also connected to a global average pooling layer to compress the features of the image extracted by each residual block. Figure 7As can be seen from the figure, the aesthetic features of the images extracted by different residual block groups correspond to different aesthetic attribute groups. Attribute group 1 is the first aesthetic attribute group, attribute group 2 is the second aesthetic attribute group, and attribute group 3 is the third aesthetic attribute group.

[0093] Step S370: Input the multiple sets of aesthetic features into their respective corresponding first fully connected layers to obtain output data of each first fully connected layer; splice the multiple sets of aesthetic features and input them into the second fully connected layer to obtain output data of the second fully connected layer.

[0094] In this embodiment of the present application, after obtaining the aesthetic features corresponding to each aesthetic attribute group, the aesthetic features corresponding to each aesthetic attribute group are input into the corresponding first fully connected layer to obtain the output data of each first fully connected layer. When predicting the aesthetic quality score of the image to be evaluated, all aesthetic features are required. Therefore, the aesthetic features extracted from each residual block group are concatenated and integrated and input into the second fully connected layer to obtain the output data of the second fully connected layer.

[0095] Step S380: Inputting the output data of each first fully connected layer into the first output layer, and obtaining the aesthetic attribute scores corresponding to the aesthetic attributes included in each aesthetic attribute group corresponding to the image to be evaluated output by the first output layer.

[0096] In an embodiment of the present application, the output data of each first fully connected layer is used to predict the aesthetic attribute score of each aesthetic attribute of the image to be evaluated.

[0097] Step S390: Input the output data of the second fully connected layer into the second output layer to obtain the aesthetic quality score of the image to be evaluated output by the second output layer.

[0098] In an embodiment of the present application, the output data of the second fully connected layer is used to predict the aesthetic quality score of the image to be evaluated.

[0099] Exemplarily, the flowchart of the image evaluation method can be as follows: Figure 8 As shown, embodiments of the present application may include an image aesthetic attribute grouping module, a deep network multi-scale feature extraction module, and an aesthetic attribute and aesthetic quality prediction module. The image aesthetic attribute grouping module combines aesthetic psychology knowledge, photography rules, and statistical laws of the data set to divide the image's aesthetic attributes into different aesthetic attribute groups. The deep network multi-scale feature extraction module performs global average pooling and concatenation of features from each layer of the deep network to obtain multi-scale features of the image, and groups the multi-scale features to obtain multiple aesthetic feature groups. The aesthetic attribute and aesthetic quality prediction module uses the multiple aesthetic feature groups obtained by the deep network multi-scale feature extraction module to predict the image's aesthetic attributes and aesthetic quality.

[0100] This method significantly improves the prediction accuracy of the aesthetic attribute prediction model. As shown in Table 2, the performance of different methods is measured using the Spearman correlation coefficient. The Spearman correlation coefficient quantitatively measures the rank correlation between the predicted aesthetic attribute results and the true results. The larger the Spearman correlation coefficient value, the better the prediction performance of the method. Table 2 shows the performance comparison of four aesthetic attribute evaluation methods.

[0101] Table 2

[0102] Aesthetic attributes Kong Method Malu method Abdenebaui method Method of the present invention Balance Elements 0.220 0.186 0.179 0.263 Color harmony 0.471 0.475 0.48 0.556 content 0.508 0.584 0.572 0.592 Depth of Field 0.479 0.495 0.496 0.534 light 0.443 0.399 0.401 0.524 Object 0.602 0.666 0.652 0.616 Rule of Thirds 0.225 0.178 0.232 0.235 Color vividness 0.648 0.681 0.685 0.718

[0103] As shown in Table 2, the prediction performance of the aesthetic attribute prediction model in this method for the eight aesthetic attributes of images is higher than that of other methods except for the object attribute, which shows that this method has high accuracy in predicting the aesthetic attributes of images.

[0104] As shown in Table 3, Table 3 gives the performance comparison of six aesthetic quality prediction methods.

[0105] Table 3

[0106]

[0107]

[0108] As shown in Table 3, the aesthetic attribute prediction model in this method has a higher prediction performance for the aesthetic quality of images than other methods, which shows that this method has a high accuracy in predicting the aesthetic quality of images.

[0109] An image evaluation method provided in embodiments of the present application extracts aesthetic features from an image using multiple residual block groups. These features are then divided into groups corresponding to aesthetic attribute groups to predict the scores of the individual aesthetic attributes of the image being evaluated. The method also integrates the aesthetic features of the image being evaluated, extracted from all residual block groups, to predict the aesthetic quality score of the image being evaluated. Leveraging the multi-scale information extracted by a deep network, this method effectively reflects the different cognitive stages of aesthetic attributes and aesthetic quality in aesthetic prediction, demonstrating high feasibility.

[0110] See also Figure 9 , an embodiment of the present application provides an image evaluation device 400, the device 400 comprising:

[0111] The image acquisition unit 410 is used to acquire an image to be evaluated.

[0112] The feature acquisition unit 420 is used to input the image to be evaluated into the multi-scale feature extraction module of the trained aesthetic attribute prediction model, and obtain the aesthetic features of each aesthetic attribute group in the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module.

[0113] Optionally, the multi-scale feature extraction module includes multiple residual block groups, the prediction module includes multiple first fully connected layers, a second fully connected layer, a first output layer connected to the multiple first fully connected layers, and a second output layer connected to the second fully connected layer, wherein each residual block group corresponds to a first fully connected layer; the feature acquisition unit 420 is specifically used to input the image to be evaluated into the multiple residual block groups, and obtain multiple groups of aesthetic features corresponding to the image to be evaluated output by the multiple residual block groups, wherein each residual block group outputs a group of aesthetic features, and each aesthetic attribute group corresponds to a group of aesthetic features.

[0114] The score acquisition unit 430 is used to input the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model, and obtain the aesthetic quality score of the image to be evaluated output by the prediction module, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated.

[0115] Optionally, the score acquisition unit 430 is specifically used to input the multiple groups of aesthetic features into their respective corresponding first fully connected layers, and obtain the output data of each first fully connected layer; input the output data of each first fully connected layer into the first output layer, and obtain the aesthetic attribute scores corresponding to the aesthetic attributes included in each aesthetic attribute group corresponding to the image to be evaluated output by the first output layer; splice the multiple groups of aesthetic features and input them into the second fully connected layer, and obtain the output data of the second fully connected layer; input the output data of the second fully connected layer into the second output layer, and obtain the aesthetic quality score of the image to be evaluated output by the second output layer.

[0116] See also Figure 10 As shown, the apparatus 400 may further include:

[0117] The model training unit 440 is used to obtain an aesthetic attribute dataset, wherein the aesthetic attribute dataset includes multiple images, aesthetic attribute scores corresponding to each image, and aesthetic quality scores corresponding to each image; based on the aesthetic attribute dataset and a preset loss function, the aesthetic attribute prediction model to be trained is trained until the training end condition is met, thereby obtaining the trained aesthetic attribute prediction model.

[0118] Among them, the aesthetic attribute scores corresponding to each image are obtained by averaging multiple first scores corresponding to each aesthetic attribute of each image; the aesthetic quality score corresponding to each image is obtained by averaging multiple second scores corresponding to each image; among them, the first scores are scores of different evaluators on the aesthetic attributes of the image, and the second scores are scores of different evaluators on the aesthetic quality of the image.

[0119] Optionally, the model training unit 440 is further used to obtain the Spearman correlation coefficient between each aesthetic attribute score and the aesthetic quality score corresponding to each image in the aesthetic attribute data set; based on the Spearman correlation coefficient, the aesthetic attributes corresponding to the image are divided into multiple aesthetic attribute groups, wherein in different aesthetic attribute groups, the correlation between each aesthetic attribute score and the aesthetic quality score is different.

[0120] Optionally, the model training unit 450 is specifically used to scale the size of the image in the aesthetic attribute data set to a preset size to obtain a scaled aesthetic attribute data set; enhance the image in the scaled aesthetic attribute data set with a random vertical flip operation with a preset probability to obtain an enhanced aesthetic attribute data set; normalize the pixel values ​​of the image in the enhanced aesthetic attribute data set to obtain a normalized aesthetic attribute data set; train the aesthetic attribute prediction model to be trained based on the normalized aesthetic attribute data set and the preset loss function until the training end condition is met to obtain the trained aesthetic attribute prediction model.

[0121] It should be noted that the device embodiment in this application corresponds to the aforementioned method embodiment. The specific principles in the device embodiment can be found in the contents of the aforementioned method embodiment and will not be repeated here.

[0122] The following will be combined Figure 11 An electronic device or server provided in this application is described.

[0123] See also Figure 11 Based on the above-mentioned image evaluation method and apparatus, the present application also provides another electronic device or server 800 that can execute the above-mentioned image evaluation method. The electronic device or server 800 includes one or more (only one is shown in the figure) processors 802, a memory 804, and a network module 806 that are coupled to each other. The memory 804 stores a program that can execute the content of the above-mentioned embodiments, and the processor 802 can execute the program stored in the memory 804.

[0124] The processor 802 may include one or more processing cores. The processor 802 utilizes various interfaces and circuits to connect various components within the electronic device or server 800. It executes instructions, programs, code sets, or instruction sets stored in the memory 804 and accesses data stored in the memory 804 to perform various functions and process data within the electronic device or server 800. Optionally, the processor 802 may be implemented using at least one of the following hardware forms: a digital signal processing (DSP), a field-programmable gate array (FPGA), or a programmable logic array (PLA). The processor 802 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing display content; and the modem handles wireless communications. It is understood that the modem may not be integrated into the processor 802 and may be implemented separately via a communication chip.

[0125] The memory 804 may include a random access memory (RAM) or a read-only memory (ROM). The memory 804 may be used to store instructions, programs, codes, code sets, or instruction sets. The memory 804 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for implementing at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the following various method embodiments, etc. The data storage area may also store data (such as a phone book, audio and video data, chat history data), etc., created by the electronic device or server 800 during use.

[0126] The network module 806 is used to receive and transmit electromagnetic waves, realize the mutual conversion between electromagnetic waves and electrical signals, and thus communicate with a communication network or other devices, such as communicating with an audio playback device. The network module 806 may include various existing circuit components for performing these functions, such as an antenna, a radio frequency transceiver, a digital signal processor, an encryption / decryption chip, a subscriber identity module (SIM) card, a memory, etc. The network module 806 can communicate with various networks such as the Internet, an enterprise intranet, a wireless network, or communicate with other devices via a wireless network. The above-mentioned wireless network may include a cellular telephone network, a wireless local area network, or a metropolitan area network. For example, the network module 806 can exchange information with a base station.

[0127] Please refer to Figure 12 , which shows a block diagram of a computer-readable storage medium provided in an embodiment of the present application. The computer-readable storage medium 900 stores program code, which can be called by a processor to execute the method described in the above method embodiment.

[0128] The computer-readable storage medium 900 can be an electronic memory such as a flash memory, an EEPROM (Electrically Erasable Programmable Read-Only Memory), an EPROM, a hard disk, or a ROM. Alternatively, the computer-readable storage medium 900 includes a non-transitory computer-readable storage medium. The computer-readable storage medium 900 has storage space for program code 910 for executing any of the method steps in the above method. These program codes can be read from or written to one or more computer program products. The program code 910 can be compressed, for example, in a suitable form.

[0129] The present application provides an image evaluation method, device, electronic device, and storage medium. The method first obtains an image to be evaluated, then inputs the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, obtains the aesthetic features of each of the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module, then inputs the aesthetic features of each aesthetic attribute group into a prediction module of the aesthetic attribute prediction model, and obtains the aesthetic quality score of the image to be evaluated output by the prediction module. The aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated. Through the above method, the aesthetic attributes of an image are divided into multiple aesthetic attribute groups according to the multi-level process of aesthetic cognition, and the aesthetic quality score of the image is predicted using the trained aesthetic attribute prediction model. This method reflects the staged nature of aesthetic perception, is more consistent with the human aesthetic process, and improves the accuracy of image aesthetic quality prediction.

[0130] The embodiments of the present invention are described above in conjunction with the accompanying drawings, but the present invention is not limited to the above-mentioned specific implementation methods. The above-mentioned specific implementation methods are merely illustrative and not restrictive. Under the guidance of the present invention, ordinary technicians in this field can also make many forms without departing from the scope of protection of the present invention and the claims, all of which are protected by the present invention.

Claims

1. An image evaluation method, characterized in that: The method comprises: Acquire an aesthetic attribute dataset, wherein the aesthetic attribute dataset includes a plurality of images, aesthetic attribute scores corresponding to each image, and an aesthetic quality score corresponding to each image; Based on the aesthetic attribute data set and the preset loss function, the aesthetic attribute prediction model to be trained is trained until the training end condition is met to obtain the trained aesthetic attribute prediction model. The preset loss function is ,in, , , , in the above formula, represents the aesthetic attribute group, N is the number of aesthetic attributes, represents the weight of the i-th aesthetic attribute of the image, and is a dynamically changing learnable parameter. and are the true and predicted scores of the i-th aesthetic attribute of the image, and are the true and predicted results of the aesthetic quality scores of the images, For the overall loss, Loss of aesthetic properties, For loss of aesthetic quality, is a dynamically changing learnable parameter, is an aesthetic attribute group loss, wherein the aesthetic attribute group is divided based on the Spearman correlation coefficient between the aesthetic attribute score and the aesthetic quality score; Get the image to be evaluated; Inputting the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and obtaining aesthetic features of each of the multiple aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; The aesthetic features of each aesthetic attribute group are input into the prediction module of the aesthetic attribute prediction model, and the aesthetic quality score of the image to be evaluated output by the prediction module is obtained, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated.

2. The method according to claim 1, characterized in that The aesthetic attribute scores corresponding to each image are obtained by averaging multiple first scores corresponding to each aesthetic attribute of each image; the aesthetic quality score corresponding to each image is obtained by averaging multiple second scores corresponding to each image; wherein, the first scores are scores of different evaluators on the aesthetic attributes of the image, and the second scores are scores of different evaluators on the aesthetic quality of the image.

3. The method according to claim 1, characterized in that The method further includes training the aesthetic attribute prediction model to be trained based on the aesthetic attribute data set and the preset loss function until a training end condition is satisfied, and obtaining the trained aesthetic attribute prediction model. Obtaining the Spearman correlation coefficient between each aesthetic attribute score and the aesthetic quality score corresponding to each image in the aesthetic attribute dataset; Based on the Spearman correlation coefficient, the aesthetic attributes corresponding to the image are divided into a plurality of aesthetic attribute groups, wherein the correlation between the aesthetic attribute scores and the aesthetic quality score is different in different aesthetic attribute groups.

4. The method according to claim 3, characterized in that The multi-scale feature extraction module includes a plurality of residual block groups, the prediction module includes a plurality of first fully connected layers, a second fully connected layer, a first output layer connected to the plurality of first fully connected layers, and a second output layer connected to the second fully connected layer, wherein each residual block group corresponds to a first fully connected layer; inputting the image to be evaluated into the multi-scale feature extraction module of the trained aesthetic attribute prediction model, and obtaining aesthetic features of each of the plurality of aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module, comprises: Inputting the image to be evaluated into the multiple residual block groups, and obtaining multiple groups of aesthetic features corresponding to the image to be evaluated output by the multiple residual block groups, wherein each residual block group outputs a group of aesthetic features, and each aesthetic attribute group corresponds to a group of aesthetic features; Inputting the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model and obtaining the aesthetic quality score of the image to be evaluated output by the prediction module includes: Inputting the multiple groups of aesthetic features into their respective corresponding first fully connected layers to obtain output data of each first fully connected layer; Inputting the output data of each first fully connected layer into the first output layer, and obtaining the aesthetic attribute scores corresponding to the aesthetic attributes included in each aesthetic attribute group corresponding to the image to be evaluated output by the first output layer; splicing the multiple sets of aesthetic features and inputting them into the second fully connected layer to obtain output data of the second fully connected layer; The output data of the second fully connected layer is input into the second output layer to obtain the aesthetic quality score of the image to be evaluated output by the second output layer.

5. The method according to claim 1, characterized in that The method of training the aesthetic attribute prediction model to be trained based on the aesthetic attribute dataset and the preset loss function until a training end condition is met to obtain the trained aesthetic attribute prediction model includes: scaling the size of the images in the aesthetic attribute dataset to a preset size to obtain a scaled aesthetic attribute dataset; performing an enhancement process on the image in the scaled aesthetic attribute dataset by performing a random vertical flip operation with a preset probability to obtain an enhanced aesthetic attribute dataset; Normalizing the pixel values ​​of the images in the enhanced aesthetic attribute dataset to obtain a normalized aesthetic attribute dataset; Based on the normalized aesthetic attribute data set and the preset loss function, the aesthetic attribute prediction model to be trained is trained until a training end condition is met, thereby obtaining the trained aesthetic attribute prediction model.

6. The method according to any one of claims 1 to 5, characterized in that: The aesthetic attributes include balancing elements, color harmony, content, depth of field, lighting, objects, the rule of thirds, and color vibrancy.

7. An image evaluation device, characterized in that: The device comprises: The model training unit is used to obtain an aesthetic attribute data set, wherein the aesthetic attribute data set includes multiple images, aesthetic attribute scores corresponding to each image, and aesthetic quality scores corresponding to each image; based on the aesthetic attribute data set and a preset loss function, the aesthetic attribute prediction model to be trained is trained until the training end condition is met, thereby obtaining the trained aesthetic attribute prediction model, wherein the preset loss function is ,in, , , , in the above formula, represents the aesthetic attribute group, N is the number of aesthetic attributes, represents the weight of the i-th aesthetic attribute of the image, and is a dynamically changing learnable parameter. and are the true and predicted scores of the i-th aesthetic attribute of the image, and are the true and predicted results of the aesthetic quality scores of the images, For the overall loss, Loss of aesthetic properties, For loss of aesthetic quality, is a dynamically changing learnable parameter, is an aesthetic attribute group loss, wherein the aesthetic attribute group is divided based on the Spearman correlation coefficient between the aesthetic attribute score and the aesthetic quality score; An image acquisition unit, configured to acquire an image to be evaluated; a feature acquisition unit, configured to input the image to be evaluated into a multi-scale feature extraction module of a trained aesthetic attribute prediction model, and acquire aesthetic features of each of the plurality of aesthetic attribute groups corresponding to the image to be evaluated output by the multi-scale feature extraction module; A score acquisition unit is used to input the aesthetic features of each aesthetic attribute group into the prediction module of the aesthetic attribute prediction model, and obtain the aesthetic quality score of the image to be evaluated output by the prediction module, wherein the aesthetic attribute prediction model is used to predict the aesthetic quality score of the image to be evaluated based on the aesthetic features of each aesthetic attribute group corresponding to the image to be evaluated.

8. An electronic device, characterized in that: The method comprises one or more processors; one or more programs are stored in a memory and configured to execute the method according to any one of claims 1 to 6 by the one or more processors.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores program code, wherein when the program code is executed by a processor, the method according to any one of claims 1 to 6 is executed.

Citation Information

Patent Citations

  • Image processing method and device, computer equipment and computer readable storage medium

    CN112102304A

  • Multi-dimensional aesthetic quality evaluation method and device for mobile game images and medium

    CN113902723A