Method, apparatus, storage medium, and computer device for selecting an image

By obtaining image candidate sets, determining categories, obtaining aesthetic quality models, evaluating aesthetic quality scores and selecting target images, the problem of insufficient evaluation of aesthetic quality in traditional image selection methods is solved, and the aesthetics and selection accuracy of images are improved.

CN111062930BActive Publication Date: 2025-06-13TENCENT TECHNOLOGY (SHENZHEN) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN201911323354.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-12-20
Publication Date
2025-06-13
Estimated Expiration
2040-08-23

AI Technical Summary

Technical Problem

Traditional image selection methods mainly evaluate the degree of distortion of images, resulting in insufficient evaluation of image aesthetic quality, which in turn affects the aesthetics of the selected image.

Method used

By obtaining the image candidate set of business content, determining the category it belongs to, analytical quality models are obtained based on the category, evaluating the aesthetic quality scores of each image in the image candidate set, and selecting the target image based on these scores.

Benefits of technology

The aesthetics of the selected images are improved, and by providing different categories of aesthetic quality evaluation standards, the fine-grained distinction of image aesthetic quality is achieved, and the accuracy of selection is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN111062930B_ABST
    Figure CN111062930B_ABST
Patent Text Reader

Abstract

The present application relates to a method, apparatus, storage medium, and computer device for image selection. The method includes: obtaining an image candidate set of business content; determining the category to which the business content belongs, and obtaining a corresponding aesthetic quality model according to the category to which the business content belongs; obtaining the aesthetic quality scores of each candidate image in the image candidate set according to the aesthetic quality model; and selecting a target image according to the aesthetic quality scores of each candidate image in the image candidate set. The present application improves the aesthetic degree of the selected image.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of computer technologies, and particularly to a method, apparatus, storage medium, and computer device for selecting images. Background Art

[0002] With the rapid popularization of intelligent devices such as cameras, video cameras, and smart phones, visual content data such as images and videos has been increasing day by day. How to find content with high aesthetic quality from a large amount of data has become a difficult problem. In recent years, with the rapid development of technologies such as computer vision algorithms and pattern recognition, people hope that computers can simulate human perception and understanding of beauty, automatically evaluate the "aesthetic feeling" of images, and select images with high aesthetic quality from the evaluation results.

[0003] However, the "aesthetic feeling" of an image is an abstract concept. In traditional image selection methods, generally, image features are manually designed, and a classifier is trained based on the extracted image features. The classifier classifies images into two categories: high quality and low quality, and selects images from the high-quality images. This image selection method mainly evaluates the distortion degree of images (such as distortion caused by poor imaging conditions, lossy compression, and distortion caused by channel attenuation during image transmission), and less evaluates the aesthetic quality of images (such as composition, color, depth of field, etc.), resulting in insufficient aesthetic degree of the selected images. Summary of the Invention

[0004] Based on this, it is necessary to provide a method, apparatus, storage medium, and computer device for selecting images to address the problem of insufficient aesthetic degree of the images selected by traditional image selection methods.

[0005] An image selection method, the method comprising:

[0006] Obtaining an image candidate set of service content;

[0007] Determining the category to which the service content belongs, and obtaining a corresponding aesthetic quality model according to the category to which the service content belongs;

[0008] Obtaining the aesthetic quality scores of each candidate image in the image candidate set according to the aesthetic quality model;

[0009] Selecting a target image according to the aesthetic quality scores of each candidate image in the image candidate set.

[0010] An image selection apparatus, the apparatus comprising:

[0011] An obtaining module, configured to obtain an image candidate set of service content;

[0012] A determination module, configured to determine the category to which the service content belongs, and obtain a corresponding aesthetic quality model according to the category to which the service content belongs;

[0013] The obtaining module is further configured to obtain the aesthetic quality scores of the candidate images in the image candidate set according to the aesthetic quality model;

[0014] A selection module, configured to select a target image according to the aesthetic quality scores of the candidate images in the image candidate set.

[0015] A storage medium, on which computer-executable instructions are stored. When the computer-executable instructions are executed by a processor, the processor is caused to execute the steps of the method for selecting an image.

[0016] A computer device, including a memory and a processor. Computer-readable instructions are stored in the memory. When the computer-readable instructions are executed by the processor, the processor is caused to execute the steps of the method for selecting an image.

[0017] For the above method, device, storage medium and computer device for selecting an image, an image candidate set of service content is obtained, the category to which the service content belongs is determined, a corresponding aesthetic quality model is obtained according to the category to which the service content belongs, the aesthetic quality scores of the candidate images in the image candidate set are obtained according to the aesthetic quality model, and a target image is selected according to the aesthetic quality scores of the candidate images in the image candidate set. In this method for selecting an image, the aesthetic quality of the image is evaluated through the aesthetic quality model, the evaluation of the image in terms of aesthetic quality is strengthened, thereby improving the aesthetic degree of the selected image; moreover, different aesthetic quality evaluation criteria are provided for different categories of images, realizing a fine-grained distinction of the aesthetic quality of the images, and improving the accuracy of image selection. Description of the Drawings

[0018] Figure 1 It is the internal structure diagram of a terminal for implementing the method for selecting an image in an embodiment;

[0019] Figure 2 It is the flowchart of the method for selecting an image in an embodiment;

[0020] Figure 3 It is the partial structure block diagram of the image selection model in an embodiment;

[0021] Figure 4 It is the flowchart block diagram of the training of the image selection model in an embodiment;

[0022] Figure 5 It is the structure block diagram of the image selection model in an embodiment;

[0023] Figure 6 Partial structural block diagram of the image selection model in another embodiment;

[0024] Figure 7 Flow schematic diagram of the image selection method in another embodiment;

[0025] Figure 8 Effect of the image selection method in one embodiment;

[0026] Figure 9 Structural block diagram of the image selection system in one embodiment;

[0027] Figure 10 Structural block diagram of the image selection device in one embodiment;

[0028] Figure 11 Structural block diagram of the image selection device in another embodiment;

[0029] Figure 12 Internal structure diagram of a computer device in one embodiment. Detailed implementation

[0030] In order to make the objectives, technical solutions and advantages of this application clearer, the following further elaborates on this application in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain this application and are not used to limit this application.

[0031] Figure 1 Internal structure schematic diagram of a terminal in one embodiment. As Figure 1 shown, the terminal includes a processor, a non-volatile storage medium, an internal memory, a network interface, a display screen, and an input device connected through a system bus. Among them, the non-volatile storage medium of the terminal stores an operating system and can also store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can implement an image selection method. The processor is used to provide computing and control capabilities to support the operation of the entire terminal. The internal memory can also store computer-readable instructions. When the computer-readable instructions are executed by the processor, the processor can execute an image selection method. The network interface is used for network communication with a server or other terminals. The display screen of the terminal can be a liquid crystal display screen or an electronic ink display screen, etc. The input device can be a touch layer covering the display screen, or a button, trackball, or touchpad provided on the terminal housing, or an external keyboard, touchpad, or mouse, etc.

[0032] Those skilled in the art can understand, Figure 1The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the terminal to which the solution of this application is applied. The specific terminal may include more or fewer components than those shown in the figure, or combine some components, or have different component arrangements.

[0033] As Figure 2 shown, in one embodiment, a method for selecting an image is provided. Referring to Figure 2 , in this embodiment, this method is mainly illustrated by taking the application of this method to the above Figure 1 terminal as an example. The method for selecting an image specifically includes the following steps:

[0034] S202, obtain an image candidate set of the service content.

[0035] Among them, the service content may be video content, graphic and text content, etc. that involve images.

[0036] The type of service content may be PGC (Professional Generated Content, professional production content), which refers to the content produced by traditional broadcasters in the manner of TV programs and adjusted according to the dissemination characteristics of the Internet; it may also be UGC (User Generated Content, user-generated content), which refers to the content created by users and displayed through Internet platforms or provided to other users; it may also be PUGC (Professional UserGenerated Content, professional user-generated content), which refers to the content that combines UGC and PGC; it may also be MCN (Multi-Channel Network, multi-channel network), which refers to the content that combines PGC content and continuously outputs with the support of capital.

[0037] Among them, the image candidate set refers to the set of candidate images for selecting the target image, and the target image may be the cover, attached drawing, etc. of the service content.

[0038] Taking the target image as the cover as an example, if the service content is video content, the candidate images may be the images obtained by frame extraction of the video, and the set of images obtained by frame extraction is used as the image candidate set; if the service content is graphic and text content, the candidate images may be the images in the graphic and text content, and the set of images in the graphic and text content is used as the image candidate set. Taking the target image as the attached drawing in the graphic and text content as an example, the candidate images may be local images, and an image candidate set is generated according to the local images corresponding to the received image selection instruction.

[0039] S204, determine the category to which the service content belongs, and obtain the corresponding aesthetic quality model according to the category to which the service content belongs.

[0040] Among them, the category refers to the field involved in the business content. The category can be subdivided into multiple levels, such as the first-level category, the second-level category, and the third-level category, and the classification of the first-level category, the second-level category, and the third-level category is refined in turn. For example, the first-level category can be technology, games, live broadcasts, finance, sports, entertainment, real estate, fashion, education, etc.; for the first-level category of news, the second-level category can be domestic news and international news; for the second-level category of domestic news, the third-level category can be economic news, legal news, technology news, sports news, social news, etc. The category in this embodiment can be one of the above first-level categories, second-level categories, or third-level categories.

[0041] The category to which the business content belongs can be determined according to the meta-information of the business content. The meta-information can include attribute information and marking information. The attribute information is used to characterize the attributes of the business content itself, such as: the size, format, title, author, publisher, publishing time, whether it is original, shooting tags (equipment, location, time, etc. of the shooting), (video) bit rate, link to the cover image (in the case of a user uploading a cover image), the category to which the business content belongs (this category is the category automatically recognized by the system or the category received as input), etc.; the marking information is used to characterize the marking of the business content by manual review, such as: the category to which the business content belongs (this category is the category marked manually), etc. The category of the business content can be determined preferentially according to the marking information. If there is no marking information for the business content, the category of the business content can be determined according to the attribute information.

[0042] Among them, the aesthetic quality model is used to evaluate the aesthetic quality of each candidate image in the image candidate set. The aesthetic quality model can score each candidate image, and the score is used to represent the aesthetic quality evaluation of the candidate image.

[0043] Since the average level of the aesthetic quality of images in different categories is different, there can be different aesthetic quality evaluation criteria for images in different categories. Generally speaking, the user's aesthetic quality standard for the live broadcast category is lower than that for the news category. This is because the average level of the aesthetic quality of news category images is higher than that of live broadcast category images. For the same image, if it is in the live broadcast category, the user's score for this image may be 9 points, and if it is in the news category, the user's score for this image may be 5 points. Therefore, different aesthetic quality evaluation criteria can be set for images in different categories, so as to achieve a fine-grained distinction of the aesthetic quality of images.

[0044] Different aesthetic quality models can be set for images of different categories. Specifically, when training the aesthetic quality model, the sample images used for training carry the labeled categories and the labeled scores. The aesthetic quality model is trained using sample images of different categories, so that the scores predicted by the aesthetic quality model continuously approach the labeled scores of the sample images, thereby obtaining the aesthetic quality model parameters corresponding to different categories, that is, the aesthetic quality models corresponding to different categories. Different categories are associated with their corresponding aesthetic quality models and stored.

[0045] The aesthetic quality model can be a convolutional neural network model. Since the image stores the value of each pixel in the computer, the full connection of the traditional neural network may lead to an explosive growth of parameters. The computer cannot update and calculate so many parameters. The weight sharing, local perception and downsampling methods of the convolutional neural network model can effectively solve the problem of too many parameters, so that the parameter scale of the convolutional neural network model is reduced to a level acceptable to the computer, and the use of multiple convolutional layers can make the image features extracted by the convolutional neural network model richer.

[0046] After inputting the feature map obtained by the previous convolution layer, it is processed by a variety of convolution kernels - using convolution kernels of various sizes such as 4*1, 3*3, 5*5, and a maximum pooling layer. Using convolution kernels of different sizes means extracting different aesthetic features for receptive fields of different sizes, and finally splicing them together means the fusion of aesthetic features of different scales. Although this can increase the aesthetic features of the extracted image, the existence of the 5*5 convolution kernel will lead to an increase in network parameters, which in turn increases the complexity of network calculations. In order not to reduce the richness of the extracted aesthetic features, you can add Inception network modules, VGG Net network modules, etc. to the convolutional neural network model. Figure 3 As shown in the figure, taking the Inception network module as an example, a 1*1 convolution operation can be added before the 3*3 and 5*5 convolution operations to reduce the parameter scale of the convolutional neural network model.

[0047] If the scale of the aesthetic features of the image extracted by the previous layer of the network is 192 * 28 * 28, that is, 192 feature maps of size 28 * 28, and if it is directly processed by 32 convolutional kernels of size 5 * 5 with a stride of 1, then the scale of the resulting feature maps is 32 * 24 * 24, that is, 32 feature maps of size 24 * 24. The total number of parameters required for this convolution step is 192 * 32 * 5 * 5 + 32, that is, 153,632 parameters. If a 1 * 1 convolutional kernel is added before the 5 * 5 convolutional kernel processing, that is, after first processing the feature maps with 32 convolutional kernels of size 1 * 1, the feature maps are first changed to a scale of 32 * 28 * 28. The number of parameters required for this step is 192 * 32 * 1 * 1 + 32, that is, 6,176 parameters. Then, after processing with 32 convolutional kernels of size 5 * 5, the original feature maps are finally transformed into feature maps with a scale of 32 * 24 * 24. The number of parameters required for this step is 32 * 32 * 5 * 5 + 32, that is, 25,632 parameters. The total number of parameters required for the latter method is 6,176 + 25,632 = 31,808. Compared with the number of parameters of the previous method, which is 153,632, the number of parameters in the convolutional neural network model is greatly reduced. In this way, the aesthetic features extracted by the convolutional neural network model are not reduced, and the parameter scale of the convolutional neural network model can be reduced.

[0048] S206. Obtain the aesthetic quality scores of each candidate image in the image candidate set according to the aesthetic quality model.

[0049] Among them, the aesthetic quality score is used to represent the aesthetic quality evaluation of the candidate image by the aesthetic quality model. Multiple preset scores (such as 1 - 10) can be preset in advance. The candidate image is input into the aesthetic quality model to obtain the probability values of the candidate image corresponding to each preset score, and the preset score corresponding to the maximum probability value is selected as the aesthetic quality score of the candidate image.

[0050] S208. Select a target image according to the aesthetic quality scores of each candidate image in the image candidate set.

[0051] Each candidate image can be sorted according to its aesthetic quality score, and the target image is selected according to the sorting result. For example, a preset number of candidate images are selected as the target images in descending order of the aesthetic quality score. The preset number can be set according to the actual application scenario.

[0052] The image selection method provided in this embodiment obtains a candidate set of images of business content, determines the category to which the business content belongs, obtains a corresponding aesthetic quality model according to the category to which the business content belongs, obtains the aesthetic quality score of each candidate image in the image candidate set according to the aesthetic quality model, and selects a target image according to the aesthetic quality score of each candidate image in the image candidate set. This image selection method evaluates the aesthetic quality of the image through the aesthetic quality model, strengthens the evaluation of the aesthetic quality of the image, thereby improving the aesthetics of the selected image; and provides different aesthetic quality evaluation standards for images of different categories, realizes fine-grained distinction of the aesthetic quality of the image, and improves the accuracy of image selection.

[0053] In one embodiment, the training method of the aesthetic quality model includes: obtaining a sample image and annotation information of the sample image, wherein the annotation information of the sample image includes an annotation category of the sample image and an annotation score of the sample image; and training the pre-trained model according to the sample image and the annotation information of the sample image to obtain the aesthetic quality model.

[0054] The aesthetic quality model is trained based on a pre-trained model, which is a convolutional neural network model trained based on an image dataset, and the image dataset may be an ImageNet image dataset, an AVA (The Aesthetic Visual Analysis dataset) dataset, a PN (Photo Net) dataset, a CUHK-PQ (The CUHK-PhotoQuality) dataset, etc. The pre-trained model may be an Inception-v1, Inception-v2, Inception-v3, Inception-v4 model, etc., which is trained based on the ImageNet image dataset.

[0055] The source of the sample image may be sample business content, which may come from various business websites, such as QQ Kandian, WeChat Kanyikan, Jinri Toutiao, and Yidian Zixun, etc. If the sample business content is video content, the video may be framed, and the image obtained by frame extraction may be annotated to obtain the sample image; if the sample business content is graphic content, the image in the graphic content may be obtained, and the obtained image may be annotated to obtain the sample image.

[0056] The annotation information refers to the content of the annotation of the sample image. The annotation information may include the annotation category and the annotation score. In one embodiment, the input annotation information is received and used as the annotation information of the sample image, that is, the annotation information may be generated by the annotation personnel.

[0057] The annotation category refers to the field involved in the sample image, which can be determined with reference to the field involved in the sample business content. The annotation category can be subdivided into multiple levels, such as the first-level annotation category, the second-level annotation category, and the third-level annotation category. The classification of the first-level annotation category, the second-level annotation category, and the third-level annotation category is refined in sequence. The first-level annotation category can be technology, games, live broadcasts, finance, sports, entertainment, real estate, fashion, education, etc.; for the first-level annotation category of news, the second-level annotation category can be domestic news and international news; for the second-level annotation category of domestic news, the third-level annotation category can be economic news, legal news, technology news, sports news, social news, etc. The annotation category in this embodiment can be one of the above first-level annotation categories, second-level annotation categories, or third-level annotation categories.

[0058] Multiple preset scores (such as 1-10) can be preset in advance, and the sample image is scored based on the consideration of the index parameters of the sample image. Among them, the index parameter refers to the factor that affects the annotation score of the sample image. Optionally, the index parameter includes at least one of aesthetics information, key object information, user attention information, and relevance information.

[0059] The aesthetics information is used to characterize the aesthetic quality of the sample image, and the aesthetics information may include color features and composition features. The color features may include brightness, contrast, saturation, color temperature, hue, color components, color coordination, etc.; the composition features may include the rule of thirds, symmetrical composition, framed composition, central composition, leading line composition, diagonal composition, triangular composition, balanced composition, etc.

[0060] The key object information is used to characterize the proportion of the key area in the sample image. The key area can be a face area, a motion area, a significant area, etc. The significant area refers to the area related to the theme of the sample business content. The key object information may include the number and aggregated size of the face area, the number and aggregated size of the motion area, and the number and aggregated size of the significant area.

[0061] The user attention information is used to characterize the factors that attract the user's attention in the sample image, such as moving objects. The motion characteristics of the available frames can be used to characterize the attention level, and the user attention information is extracted according to the optical flow result. The motion information may include motion statistical characteristics, boundary jitter characteristics, camera jitter characteristics, motion entropy characteristics, and key target motion characteristics. The motion statistical characteristics are used to describe the motion intensity of the moving object, the boundary jitter characteristics are used to describe the motion of the moving object relative to the camera, the camera jitter characteristics are used to describe the absolute camera motion in the motion direction, the motion entropy characteristics are used to describe the change in the motion direction of the moving object, and the key target motion characteristics are used to describe the motion characteristics of the key target.

[0062] The relevance information is used to characterize the relevance between the sample image and the theme of the sample service content.

[0063] As Figure 4 shown, the training process of the aesthetic quality model is as follows: divide the sample images into a training set and a test set, use the training set to train the parameters of the aesthetic quality model, and use the test set to test the evaluation effect of the aesthetic quality model.

[0064] The method for selecting images provided in this embodiment uses the sample images and the annotation information of the sample images to train the aesthetic quality model, improves the evaluation ability of the aesthetic quality model for the aesthetic quality of images, and realizes the fine-grained distinction of the aesthetic quality of images by the aesthetic quality model.

[0065] In one embodiment, the determination method of the sample images includes: obtaining at least two annotation scores for each original image; selecting the original images whose difference between the at least two annotation scores is within a preset range as the sample images.

[0066] Among them, the original images can be determined in the sample service content. That is, if the sample service content is video content, the video can be frame-extracted, and the extracted images can be annotated to obtain the original images; if the sample service content is text-image content, the images in the text-image content are obtained and the obtained images are annotated to obtain the original images. Each original image can be annotated by at least two annotators. Select the sample images from the original images. That is, when the difference between the annotation scores of at least two annotators is within the preset range, the original image is used as the sample image. The preset range can be set according to actual applications. For example, the preset range can be [1, 2].

[0067] When different annotators annotate the same sample image, different annotators may have different cognitions of the annotation categories and annotation scores. If the cognitive deviation of the annotators from the annotation scores is too large, it is not conducive to the learning of the aesthetic quality model. Therefore, the annotation scores of different annotators for the same sample image are controlled within a certain range.

[0068] By obtaining at least two annotation scores for each sample image, the distribution information of the annotation scores of the sample image can be obtained, and then the probability value of the annotation score relative to each preset score can be obtained. The probability value can be used to train the aesthetic quality model.

[0069] The method for selecting images provided in this embodiment obtains the distribution information of the annotation scores of the sample images, and then obtains the probability value of the annotation score relative to each preset score. The probability value is used to train the aesthetic quality model, thereby improving the training accuracy of the aesthetic quality model.

[0070] In one embodiment, training the pre-trained model according to the sample image and the annotation information of the sample image to obtain the aesthetic quality model includes: inputting the sample image and the annotation information of the sample image into the pre-trained model to obtain the predicted score probability distribution information of the sample image; updating the parameters of the pre-trained model according to the difference between the predicted score probability distribution information and the reference score probability distribution information to obtain the aesthetic quality model, where the reference score probability distribution information is generated according to the annotation scores of the sample image.

[0071] Among them, the predicted score probability distribution information refers to the probability value of a sample image predicted by the aesthetic quality model relative to each preset score for a sample image; the reference score probability distribution information refers to the probability value of a sample image calculated according to at least two annotation scores relative to each preset score for a sample image.

[0072] Each sample image includes at least two annotation scores, and the reference score probability distribution information of the sample image can be determined according to at least two annotation scores. For example, if the preset scores are 1-10 and the annotation scores are 9 and 10, then the reference score probability distribution information is: the probability values of scores 1-8 are 0, and the probability values of 9 and 10 are both 50%.

[0073] For this sample image, the aesthetic quality model outputs the predicted score probability distribution information, and the predicted score probability distribution information can be an N-dimensional vector, and each dimension represents the probability value of the image relative to each preset score, where N refers to the number of preset scores.

[0074] According to the difference between the predicted score probability distribution information and the reference score probability distribution information, calculate the loss information, and update the parameters of the aesthetic quality model according to the loss information.

[0075] The method for selecting images provided in this embodiment trains the aesthetic quality model according to the difference between the predicted score probability distribution information and the reference score probability distribution information, improving the training accuracy of the aesthetic quality model.

[0076] In one embodiment, the predicted score probability distribution information includes first predicted score probability distribution information, second predicted score probability distribution information, and third predicted score probability distribution information. The aesthetic quality model includes a first loss function, a second loss function, and a third loss function. The first loss function, the second loss function, and the third loss function are at different levels in the aesthetic quality model. The first loss function is used to calculate the difference between the first predicted score probability distribution information and the reference score probability distribution information. The second loss function is used to calculate the difference between the second predicted score probability distribution information and the reference score probability distribution information. The third loss function is used to calculate the difference between the third predicted score probability distribution information and the reference score probability distribution information;

[0077] Updating the parameters of the pre-trained model according to the difference between the predicted score probability distribution information and the reference score probability distribution information to obtain the aesthetic quality model includes: calculating first loss information according to the first loss function, calculating second loss information according to the second loss function, and calculating third loss information according to the third loss function; updating the parameters of the pre-trained model according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model.

[0078] The activation function and the loss function can be set at different levels in the aesthetic quality model. The specific setting levels and the number of settings can be set according to the actual application. Each activation function can adopt the Sigmoid function, and each loss function can adopt the softmax function.

[0079] Among them, the first predicted score probability distribution information, the second predicted score probability distribution information, and the third predicted score probability distribution information are calculated by activation functions (first activation function, second activation function, third activation function) at different levels in the aesthetic quality model; the first loss function, the second loss function, and the third loss function are loss functions at different levels in the aesthetic quality model. The first loss function is used to calculate the difference between the first predicted score probability distribution information and the reference score probability distribution information to obtain first loss information; the second loss function is used to calculate the difference between the second predicted score probability distribution information and the reference score probability distribution information to obtain second loss information; the third loss function is used to calculate the difference between the third predicted score probability distribution information and the reference score probability distribution information to obtain third loss information.

[0080] Specifically, corresponding weights can be set for the first loss information, the second loss information, and the third loss information, and the total loss information can be obtained based on the weighted sum of the first loss information, the second loss information, and the third loss information, and the parameters of the pre-trained model can be updated using the total loss information. It is also possible to use the first loss information, the second loss information, and the third loss information to update the parameters on the paths where the first loss function, the second loss function, and the third loss function are located, respectively.

[0081] As Figure 5 shown, the three loss functions softmax0, softmax1, and softmax2 are at different levels in the aesthetic quality model, and the three loss functions obtain the first loss information, the second loss information, and the third loss information respectively. Taking the Figure 5 softmax0 branch in Figure 6 as an example, as

[0082] shown, the training process of the aesthetic quality model is described as follows: First, the first three layers perform convolution operations on the sample image, using convolution kernels of 7*7, 3*3, and 1*1 to process the sample image respectively, and an LRN layer is also added in the middle, mainly to accelerate the convergence speed of the network. From the fourth layer to the ninth layer, Inception modules are used respectively. First, the fourth layer uses different 1*1 convolution kernels to perform convolution processing on the sample image. The fifth layer is preceded by 3*3 and 5*5 convolution kernels respectively to further process the sample image, and the processing results obtained in the fifth layer are combined as the first part of the features of the sample image obtained by the network. The sixth layer and the seventh layer repeat the structure of the fourth layer and the fifth layer above to obtain the second part of the features of the sample image. The eighth layer and the ninth layer repeat the structure of the fourth layer and the fifth layer above again, and at the same time, a convolution operation with a 1*1 convolution kernel of the tenth layer is added after the ninth layer to obtain the third part of the features of the sample image. In the eleventh layer, a fully connected layer is used to fuse all the features of the sample image together to form a feature vector. The twelfth layer is also a fully connected layer, and the probability values of the sample image belonging to each preset score are obtained through the first activation function. Finally, the first loss information is calculated using the first loss function.

[0082] The three loss functions can adopt different loss calculation methods, such as Cross Entropy, JS (Jensen-Shannon divergence), KL (Kullback-Leibler divergence), EMD (empirical mode decomposition), and Euclidean distance, etc. In one embodiment, the first loss function can use JS divergence to measure the difference between the predicted score probability distribution information and the reference score probability distribution information, the second loss function can use EMD (empirical mode decomposition) to measure the difference between the predicted score probability distribution information and the reference score probability distribution information, and the third loss function can use Euclidean distance to measure the difference between the predicted score probability distribution information and the reference score probability distribution information.

[0083] Taking JS divergence as an example:

[0084] JS(Pr, Pg) = KL(Pr||Pg) + KL(Pg||Pm)

[0085] where Pm = (Pr + Pg) / 2

[0086]

[0087] where Pr is the predicted score probability distribution information, and Pg is the reference score probability distribution information.

[0088] Taking EMD as an example:

[0089] W(Pr, Pg) = infγ∈π(Pr,Pg) E(x,y)~y[||x - y||]

[0090] where Pr is the predicted score probability distribution information, Pg is the reference score probability distribution information, and W(Pr, Pg) is all possible joint distributions of Pr and Pg combined. For each possible joint distribution γ, samples x and y can be sampled from it (x,y)~γ, the distance ||x - y|| of this pair of samples can be calculated, and the expected value E(x,y)~γ[||x - y||] of the sample pair distance under this joint distribution γ can be calculated. Among all possible joint distributions, the lower bound obtained for this expected value is the EMD distance.

[0091] The method for selecting images provided in this embodiment sets loss functions at different levels of the aesthetic quality model, and updates the parameters of the aesthetic quality model according to the loss information calculated by the loss functions, improving the training accuracy of the aesthetic quality model.

[0092] In one embodiment, updating the parameters of the pre-trained model according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model includes: updating the parameters on the paths where the first loss function, the second loss function, and the third loss function are located according to the first loss information, the second loss information, and the third loss information respectively to obtain the aesthetic quality model.

[0093] Specifically, the parameters on the path where the first loss function is located are updated using the first loss information, the parameters on the path where the second loss function is located are updated using the second loss information, and the parameters on the path where the third loss function is located are updated using the third loss function.

[0094] The method for selecting images provided in this embodiment sets loss functions at different levels of the aesthetic quality model, and updates the parameters on the paths where the loss functions are located according to the loss information calculated by the loss functions, improving the accuracy of training the aesthetic quality model.

[0095] In one embodiment, the method further includes: periodically obtaining new sample images and the annotation information of the new sample images, and training the pre-trained model according to the new sample images and the annotation information of the new sample images to obtain the aesthetic quality model.

[0096] Among them, during the application process of the aesthetic quality model, the sample image library for storing sample images is continuously updated, and the new sample images refer to the sample images newly added to the sample image library.

[0097] Based on the business content feedback by the user, after the aesthetics of the business content is manually reviewed and found to be low, the images in the business content are obtained, marked, and the marked images are added to the sample image library. The new sample images and the annotation information of the new sample images can be periodically obtained to optimize and update the aesthetic quality model.

[0098] The method for selecting images provided in this embodiment periodically optimizes the aesthetic quality model to ensure the accuracy of prediction of the aesthetic quality model.

[0099] In one embodiment, obtaining the image candidate set of the business content includes: if the business content is a video, extracting video frames from the business content according to the frame extraction rule to obtain the image candidate set, and the frame extraction rule includes at least one of a motion rule, a user attention rule, a sparse coding rule, a sparse reconstruction rule, a key point rule, and a portrait rule.

[0100] In this embodiment, video scene transition frames can be extracted. The video scene transition frames can be identified by the degree of difference between the current frame and the previous frame, and this degree of difference can be quantified by at least one feature such as brightness, color, and edges. Taking brightness as an example, if the difference in brightness between the current frame and the previous frame is greater than the brightness threshold, then the current frame is a video scene transition frame; if the difference in brightness between the current frame and the previous frame is less than or equal to the brightness threshold, then the current frame is not a video scene transition frame.

[0101] After extracting the video scene transition frames, candidate images are determined in combination with the frame extraction rules. Among them, the frame extraction rules may include at least one of the following: motion rule, user attention rule, sparse coding rule, sparse reconstruction rule, key point rule, and portrait rule. The motion rule means: selecting frames with a motion amplitude greater than a preset amplitude; the user attention rule means: selecting frames with features that attract the user's attention, such as frames with motion features; the sparse coding rule means: selecting frames with a sparse coding amplitude greater than a preset amplitude; the sparse reconstruction rule means: selecting frames with a sparse reconstruction quantity less than a predetermined quantity; the key point rule means: selecting frames with a number of key points greater than a preset quantity, where key points refer to points related to the theme of the business content; the portrait rule means: selecting frames with a portrait.

[0102] The method for selecting images provided in this embodiment improves the aesthetic quality of candidate images.

[0103] In one embodiment, the selecting the target image according to the aesthetic quality scores of the candidate images in the image candidate set includes: sorting the candidate images in the image candidate set according to the aesthetic quality scores of the candidate images; selecting a preset number of candidate images as the target image according to the sorting result.

[0104] The candidate images can be sorted according to the aesthetic quality scores of the candidate images, and the target image is selected according to the sorting result. For example, a preset number of candidate images are selected as the target image in descending order of the aesthetic quality scores. The preset number can be set according to the actual application scenario.

[0105] The method for selecting images provided in this embodiment selects the target image according to the aesthetic quality score, improving the beauty of the selected image.

[0106] As Figure 7 shown, in a specific embodiment, the method for selecting images includes the following steps:

[0107] S702, obtaining an image candidate set of the business content;

[0108] S704, determining the category to which the business content belongs, and obtaining a corresponding aesthetic quality model according to the category to which the business content belongs;

[0109] S706, obtain the aesthetic quality scores of each candidate image in the image candidate set according to the aesthetic quality model;

[0110] S1008, select a cover according to the aesthetic quality scores of each candidate image in the image candidate set.

[0111] In a specific embodiment, the method for selecting an image can be used to select a cover for business content. In the traditional Internet industry, business content is often pushed in the form of an information stream, such as QQ Browser, WeChat Moments, Toutiao, and Yidian Zixun. When users browse the information stream, the first things they notice are the title, cover, and author of the business content. As the first impression of the business content for users, the cover directly affects the click-through rate of the business content.

[0112] As Figure 8 shown, Figure 8 On the left is the cover selected by the traditional method, Figure 8 and on the right is the cover selected by this embodiment. Through the method for selecting an image in this embodiment, a cover with high aesthetic quality and more focused theme can be selected.

[0113] The method for selecting an image provided in this embodiment can select a cover with high aesthetic quality and more focused theme, thereby improving the click-through rate of business content.

[0114] In one embodiment, as Figure 9 shown, the image selection system includes: a content production module, an uplink and downlink content interface module, a content information storage module, a content warehousing module, a video frame extraction and graphic parsing module, an aesthetic quality model, an aesthetic quality scoring module, a sample image module, a content distribution module, a content distribution outlet module, a content consumption module, a feedback module, and an audit module. The main functions of each module are as follows:

[0115] Among them, the content production module is used to: receive the business content uploaded by the producer of the business content. The business content can be PGC (Professional Generated Content), UGC (User Generated Content), PUGC (Professional User Generated Content), and MCN (Multi-Channel Network), etc.

[0116] The up and down content interface module is used for: obtaining the service content sent by the content production module and the meta information of the service content. The meta information may include attribute information and marking information. The attribute information is used to characterize the attributes of the service content itself, and the marking information is used to characterize the marking of the service content by manual review. Sending the meta information of the service content to the content information storage module.

[0117] The content information storage module is used for: storing the meta information of the service content.

[0118] The content warehousing module is used for: being responsible for the scheduling of the entire image selection system. The content warehousing module obtains the service content through the up and down content interface module, and obtains the meta information of the service content through the content information storage module; calls the video frame extraction and graphic and text parsing module to process the service content to obtain the image candidate set of the service content; according to the category to which the service content belongs, calls the corresponding aesthetic quality model of the category to score each candidate image, so as to determine the target image of the service content; sends the service content, the meta information of the service content, and the target image to the content distribution module for the content distribution exit module to call the content in the content distribution module and output it to the content consumption module.

[0119] The video frame extraction and graphic and text parsing module is used for: performing frame extraction on the video to obtain the image candidate set of the video, or extracting the images in the graphic and text content to obtain the candidate set of the graphic and text content.

[0120] The aesthetic quality model is used for: evaluating the aesthetic quality of the candidate images; regularly obtaining the images with marking information in the sample image module for optimization training.

[0121] The aesthetic quality scoring module is used for: servitizing the aesthetic quality model to construct a service that can be called on the link.

[0122] The sample image module is used for: storing the images with marking information obtained from the content information storage module and the images with marking information sent by the review module. The images stored in the sample image module are used to train the aesthetic quality model.

[0123] The content distribution module is used for: obtaining and storing the service content, the meta information of the service content, and the target image from the content warehousing module.

[0124] The content distribution exit module is used for: providing an information stream to the content consumption module according to the content stored in the content distribution module.

[0125] The content consumption module is used for: displaying the information stream, and the content consumption module is provided with a feedback module.

[0126] A feedback module, configured to: receive the input feedback information, and obtain the service content corresponding to the feedback information, where the feedback information is used to indicate that the aesthetic degree of the service content is relatively low.

[0127] An audit module, configured to: receive the service content sent by the feedback module, conduct a review on the service content, after confirming that the aesthetic degree of the service content is relatively low, obtain the image in the service content, mark the image, and add the marked image to the sample image module as the basis for subsequent iterative optimization of the aesthetic quality model.

[0128] Figure 2 and Figure 7 is a schematic flowchart of a method for selecting an image in an embodiment. It should be understood that although Figure 2 and Figure 7 the steps in the flowchart are shown in sequence according to the indication of the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless otherwise clearly stated in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, Figure 2 and Figure 7 at least a part of the steps in

[0129] such as Figure 10 shown, in an embodiment, an image selection device 1000 is provided, including: an acquisition module 1002, a determination module 1004, and a selection module 1006.

[0130] The acquisition module 1002 is configured to acquire an image candidate set of the service content;

[0131] The determination module 1004 is configured to determine the category to which the service content belongs, and obtain the corresponding aesthetic quality model according to the category to which the service content belongs;

[0132] The acquisition module 1002 is further configured to obtain the aesthetic quality scores of the respective candidate images in the image candidate set according to the aesthetic quality model;

[0133] The selection module 1006 is configured to select a target image according to the aesthetic quality scores of the respective candidate images in the image candidate set.

[0134] The above-mentioned image selection device 1000 obtains a candidate set of images of service content, determines the category to which the service content belongs, obtains a corresponding aesthetic quality model according to the category to which the service content belongs, obtains the aesthetic quality scores of each candidate image in the image candidate set according to the aesthetic quality model, and selects a target image according to the aesthetic quality scores of each candidate image in the image candidate set. This image selection method evaluates the aesthetic quality of images through an aesthetic quality model, strengthens the evaluation of images in terms of aesthetic quality, thereby improving the aesthetic degree of the selected images; moreover, different aesthetic quality evaluation criteria are provided for images of different categories, realizing fine-grained differentiation of the aesthetic quality of images and improving the accuracy of image selection.

[0135] In one embodiment, as Figure 11 shown, the image selection device 1000 further includes a training module 1008. The obtaining module 1002 is further configured to: obtain a sample image and annotation information of the sample image, where the annotation information of the sample image includes the annotation category of the sample image and the annotation score of the sample image; the training module 1008 is configured to: train the pre-trained model according to the sample image and the annotation information of the sample image to obtain the aesthetic quality model. The annotation score of the sample image is determined according to at least one of index parameters, and the index parameters include at least one of aesthetic degree information, key object information, user attention information, and relevance information; the aesthetic degree information includes color features and composition features; the key object information includes the number and aggregation size of face regions, the number and aggregation size of motion regions, and the number and aggregation size of salient regions; the user attention information includes motion statistical features, boundary jitter features, camera jitter features, motion entropy features, and key target motion features; the relevance information includes the relevance between the sample image and the theme of the sample service content.

[0136] In one embodiment, the obtaining module 1002 is further configured to: obtain at least two annotation scores for each original image; select the original image whose difference between the at least two annotation scores is within a preset range as the sample image.

[0137] In one embodiment, the training module 1008 is further configured to: input the sample image and the annotation information of the sample image into the pre-trained model to obtain the predicted score probability distribution information of the sample image; update the parameters of the pre-trained model according to the difference between the predicted score probability distribution information and the reference score probability distribution information to obtain the aesthetic quality model, where the reference score probability distribution information is generated according to the annotation score of the sample image.

[0138] In one embodiment, the predicted score probability distribution information includes first predicted score probability distribution information, second predicted score probability distribution information, and third predicted score probability distribution information. The aesthetic quality model includes a first loss function, a second loss function, and a third loss function. The levels of the first loss function, the second loss function, and the third loss function in the aesthetic quality model are different. The first loss function is used to calculate the difference between the first predicted score probability distribution information and the reference score probability distribution information. The second loss function is used to calculate the difference between the second predicted score probability distribution information and the reference score probability distribution information. The third loss function is used to calculate the difference between the third predicted score probability distribution information and the reference score probability distribution information. The training module 1008 is further configured to: calculate first loss information according to the first loss function, calculate second loss information according to the second loss function, and calculate third loss information according to the third loss function; update the parameters of the pre-trained model according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model.

[0139] In one embodiment, the training module 1008 is further configured to: update the parameters on the paths where the first loss function, the second loss function, and the third loss function are located respectively according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model.

[0140] In one embodiment, the acquisition module 1002 is further configured to: regularly acquire new sample images and the annotation information of the new sample images; the training module 1008 is further configured to: train the pre-trained model according to the new sample images and the annotation information of the new sample images to obtain the aesthetic quality model.

[0141] In one embodiment, the acquisition module 1002 is further configured to: if the service content is a video, extract video frames from the service content according to a frame extraction rule to obtain the image candidate set, where the frame extraction rule includes at least one of a motion rule, a user attention rule, a sparse coding rule, a sparse reconstruction rule, a key point rule, and a portrait rule.

[0142] In one embodiment, the selection module 1006 is further configured to: sort the candidate images in the image candidate set according to the aesthetic quality scores of the candidate images; select a preset number of candidate images as the target images according to the sorting result.

[0143] Figure 12The internal structure diagram of a computer device in an embodiment is shown. The computer device may specifically be the Figure 1 terminal in. As Figure 12 shown, the computer device includes a processor, a memory, and a network interface connected through a system bus. Among them, the memory includes a non-volatile storage medium and an internal memory. The non-volatile storage medium of the computer device stores an operating system and may also store a computer program. When the computer program is executed by the processor, the processor can implement the method for selecting an image. The internal memory may also store a computer program. When the computer program is executed by the processor, the processor can execute the method for selecting an image.

[0144] Those skilled in the art can understand that Figure 12 the structure shown in is only a block diagram of some structures related to the solution of the present application, and does not constitute a limitation on the computer device to which the solution of the present application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.

[0145] In an embodiment, the image selection device provided by the present application may be implemented in the form of a computer program, and the computer program can run on a computer device such as Figure 12 shown. The memory of the computer device may store each program module that makes up the image selection device. For example, Figure 10 the acquisition module 1002, the determination module 1004, and the selection module 1006 shown. The computer program composed of each program module enables the processor to execute the steps in the image selection method of each embodiment of the present application described in this specification.

[0146] In an embodiment, a computer device is provided, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the above-mentioned image selection method. Here, the steps of the image selection method may be the steps in the image selection method of the above-mentioned various embodiments.

[0147] In an embodiment, a storage medium is provided, storing a computer program. When the computer program is executed by the processor, the processor is caused to execute the steps of the above-mentioned image selection method. Here, the steps of the image selection method may be the steps in the image selection method of the above-mentioned various embodiments.

[0148] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The program can be stored in a non-volatile computer-readable storage medium. When the program is executed, it can include the processes of the embodiments of the above methods. Among them, any reference to a memory, storage, database, or other medium used in the embodiments provided in the present application can include non-volatile and / or volatile memories. Non-volatile memories can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memories can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link (Sync Hour link) DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and Rambus dynamic RAM (RDRAM), etc.

[0149] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope recorded in this specification.

[0150] The above-described embodiments only represent several implementation manners of the present application. The description is relatively specific and detailed, but it should not be construed as a limitation on the patent scope of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the patent of the present application should be subject to the appended claims.

Claims

1. A method for selecting an image, characterized in that, the method includes: Obtaining an image candidate set of business content; Determining the category to which the business content belongs, obtaining a corresponding aesthetic quality model according to the category to which the business content belongs, where the category refers to the field involved in the business content, and the category to which the business content belongs is determined according to the meta-information of the business content. The meta-information includes at least one of attribute information and marking information. The marking information is used to represent the marking of the business content by manual review. For images of different categories, there are different aesthetic quality evaluation criteria. The aesthetic quality standard of the live broadcast category is lower than that of the news category. For the same image, the scores in the live broadcast category and the news category are different; Obtaining the aesthetic quality scores of each candidate image in the image candidate set according to the aesthetic quality model; the training of the aesthetic quality model uses sample images and the annotation information of the sample images. The annotation information includes the annotation category of the sample image and the annotation score of the sample image. Based on the consideration of the index parameters of the sample image, the sample image is annotated with a score. The index parameters include at least one of key object information and relevance information. The key object information is used to represent the proportion of the key area in the sample image. The key area includes the area related to the theme of the sample business content. The relevance information includes the relevance between the sample image and the theme of the sample business content; Selecting a target image according to the aesthetic quality scores of each candidate image in the image candidate set. The target image includes the cover or attached drawing of the business content.

2. The method according to claim 1, characterized in that, the aesthetic quality model is trained based on a pre-trained model, and the pre-trained model is a convolutional neural network model trained according to an image data set.

3. The method according to claim 2, characterized in that, the training method of the aesthetic quality model includes: Obtaining sample images and the annotation information of the sample images; Training the pre-trained model according to the sample images and the annotation information of the sample images to obtain the aesthetic quality model.

4. The method according to claim 3, characterized in that, the annotation score of the sample image is determined according to index parameters, and the index parameters further include at least one of aesthetics information and user attention information.

5. The method according to claim 4, characterized in that, the aesthetics information includes color features and composition features; The key object information includes the number and aggregated size of face regions, the number and aggregated size of motion regions, and the number and aggregated size of salient regions; the user attention information includes motion statistical features, boundary jitter features, camera jitter features, motion entropy features, and key target motion features.

6. The method according to claim 3, characterized in that, the determination method of the sample image includes: Obtaining at least two annotation scores of each original image; Selecting the original image whose difference between the at least two annotation scores is within a preset range as the sample image.

7. The method according to claim 3, characterized in that, Training the pre-trained model based on the sample image and the annotation information of the sample image to obtain the aesthetic quality model includes: Inputting the sample image and the annotation information of the sample image into the pre-trained model to obtain the predicted score probability distribution information of the sample image; Updating the parameters of the pre-trained model according to the difference between the predicted score probability distribution information and the reference score probability distribution information to obtain the aesthetic quality model, where the reference score probability distribution information is generated according to the annotation scores of the sample image.

8. The method according to claim 7, wherein, The predicted score probability distribution information includes first predicted score probability distribution information, second predicted score probability distribution information, and third predicted score probability distribution information. The aesthetic quality model includes a first loss function, a second loss function, and a third loss function. The first loss function, the second loss function, and the third loss function are at different levels in the aesthetic quality model. The first loss function is used to calculate the difference between the first predicted score probability distribution information and the reference score probability distribution information. The second loss function is used to calculate the difference between the second predicted score probability distribution information and the reference score probability distribution information. The third loss function is used to calculate the difference between the third predicted score probability distribution information and the reference score probability distribution information; The updating the parameters of the pre-trained model according to the difference between the predicted score probability distribution information and the reference score probability distribution information to obtain the aesthetic quality model includes: Calculating first loss information according to the first loss function, calculating second loss information according to the second loss function, and calculating third loss information according to the third loss function; Updating the parameters of the pre-trained model according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model.

9. The method according to claim 8, wherein, The first loss function uses the divergence method to measure the difference between the first predicted score probability distribution information and the reference score probability distribution information; the second loss function uses the empirical mode decomposition method to measure the difference between the second predicted score probability distribution information and the reference score probability distribution information; the third loss function uses the Euclidean distance to measure the difference between the third predicted score probability distribution information and the reference score probability distribution information.

10. The method according to claim 8, wherein, The updating the parameters of the pre-trained model according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model includes: Respectively updating the parameters on the paths where the first loss function, the second loss function, and the third loss function are located according to the first loss information, the second loss information, and the third loss information to obtain the aesthetic quality model.

11. The method according to claim 3, wherein, the method further includes: periodically obtaining newly added sample images and annotation information of the newly added sample images; training the aesthetic quality model according to the newly added sample images and the annotation information of the newly added sample images.

12. The method according to claim 1, wherein, the obtaining of the image candidate set of the service content includes: if the service content is a video, extracting video frames from the service content according to a frame extraction rule to obtain the image candidate set, and the frame extraction rule includes at least one of a motion rule, a user attention rule, a sparse coding rule, a sparse reconstruction rule, a key point rule, and a portrait rule; or the selecting of the target image according to the aesthetic quality scores of the candidate images in the image candidate set includes: sorting the candidate images according to the aesthetic quality scores of the candidate images in the image candidate set; and selecting a preset number of candidate images as the target image according to the sorting result.

13. An image selection device, wherein, the device includes: an obtaining module, configured to obtain an image candidate set of service content; a determining module, configured to determine the category to which the service content belongs, obtain a corresponding aesthetic quality model according to the category to which the service content belongs, the category refers to the field involved in the service content, the category to which the service content belongs is determined according to the meta information of the service content, and the meta information includes at least one of attribute information and marking information, and the marking information is used to represent the marking of the service content by manual review. For images of different categories, there are different aesthetic quality evaluation criteria. The aesthetic quality standard for the live broadcast category is lower than that for the news category. For the same image, the scores are different under the live broadcast category and under the news category; the obtaining module is further configured to obtain the aesthetic quality scores of the candidate images in the image candidate set according to the aesthetic quality model. The aesthetic quality model is trained using sample images and annotation information of the sample images. The annotation information includes the annotation category and the annotation score of the sample image. Based on the consideration of the index parameters of the sample image, the sample image is annotated with a score. The index parameters include at least one of key object information and relevance information. The key object information is used to represent the proportion of the key area in the sample image, and the key area includes the area related to the theme of the sample service content. The relevance information includes the relevance of the sample image to the theme of the sample service content; a selecting module, configured to select a target image according to the aesthetic quality scores of the candidate images in the image candidate set, and the target image includes the cover or the attached drawing of the service content.

14. A computer device, including a memory and a processor. When the computer program stored in the memory is executed by the processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 12.

15. A storage medium, wherein, The computer-executable instructions are stored on the storage medium, and when the computer-executable instructions are executed by a processor, the processor is caused to execute the steps of the method according to any one of claims 1 to 12.

Citation Information

Patent Citations

  • Automatic image grading method and device

    CN107153838A

  • Neural network model training method, electronic device and storage medium

    CN108009638A

  • Aesthetic image quality prediction system and method based on depth drift-diffusion method

    CN109583500A