Face image attraction prediction method and electronic equipment
By introducing face prior features and multimodal coding technology based on face ID in the face image attraction prediction model, the shortcomings of existing methods in understanding complex contexts and subtle differences are solved, and more accurate and robust face attraction prediction is achieved.
Patent Information
- Application Number
- CN202510011727.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-03
- Publication Date
- 2025-06-03
AI Technical Summary
Existing face image attraction prediction methods are insufficient in understanding complex contexts and subtle differences, and the CNN-based methods are limited to local features, limiting long-term modeling capabilities.
By introducing face prior features based on face ID, and combining multimodal coding and cross-modal fusion technology, a face image attraction prediction model is constructed. The model acquires multi-scale features through deep learning networks, fuses face prior features, uses text encoder and image encoder for multimodal encoding, and optimizes text encoding vectors through cross-modal fusion, and finally calculates the similarity between image and text encoding vectors to predict attractiveness.
It improves the accuracy of face image attraction prediction, enhances the ability to capture individual features of faces, improves the robustness of the algorithm and adapts to complex lighting, posture and expression changes.
Smart Images

Figure CN120088826A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of information processing technologies, and particularly to a method for predicting the attractiveness of a face image and an electronic device. Background Art
[0002] The Facial Attractiveness Prediction / Facial Beauty Prediction (FAP / FBP) algorithm aims to predict the quality of a face image in terms of its attractiveness to people through an objective algorithm, simulating the subjective scoring results of humans. This algorithm can be used to judge the attractiveness level of faces in individual video contents (such as live streaming rooms or certain short video clips), helping to evaluate the relevant face image quality and beautification degree in business, and so on.
[0003] To measure the attractiveness of a face image, early face attractiveness prediction methods mainly relied on manually designed features, such as geometric features, traditional image descriptors, global descriptors, etc. However, since these manually designed features are difficult to comprehensively present the characteristics of face attractiveness, these methods perform poorly.
[0004] With the progress of deep learning, FAP / FBP solutions based on CNN (Convolutional Neural Networks) perform better. Inspired by psychological research, specific CNN architectures are designed, and the application of label distribution learning achieves better performance. Some methods also combine prior knowledge including facial key points, parsing maps, etc., and some methods integrate person attributes related to attractiveness. However, these CNN-based methods mainly focus on local features, limiting their ability to perform long-term modeling.
[0005] To solve the above problems, some methods also propose implementation solutions combining Transformer and CNN to utilize global and local features related to face attractiveness. However, these methods still lack in understanding complex contexts and subtle differences in tasks, so there is still a large room for improvement in the accuracy of face image attractiveness prediction. Summary of the Invention
[0006] This application provides a method for predicting the attractiveness of a face image and an electronic device, which can improve the accuracy of face image attractiveness prediction.
[0007] This application provides the following solutions:
[0008] A method for predicting the attractiveness of a face image, comprising:
[0009] Determine a face image to be predicted;
[0010] In the process of obtaining the face image features through a deep learning network, introduce the face prior features based on the face ID provided by the face recognition network to obtain the fused face image features;
[0011] Construct multiple text contents by dividing into multiple attraction levels, and use a text encoder to encode each of the multiple text contents respectively to obtain multiple text encoding vectors, where the multiple text contents are respectively used to textually describe the attractiveness of the face image through the multiple attraction levels;
[0012] Use the fused face image features to perform cross-modal optimization processing on the text encoding vectors to obtain multiple optimized text encoding vectors;
[0013] Use an image encoder to perform image encoding on the face image to obtain an image encoding vector;
[0014] By calculating the similarity between the image encoding vector and each of the multiple optimized text encoding vectors respectively, output the probabilities that the face image belongs to the multiple attraction levels respectively, for determining the attractiveness prediction result of the face image.
[0015] Among them, the face image features obtained through the deep learning network are multi-scale features, and the multi-scale features are generated by processing the features of different receptive fields output by different stages of the deep learning network.
[0016] Among them, the face prior features based on the face ID provided by the pre-trained face recognition network are used to represent the features that distinguish the face image from other images.
[0017] Among them, the text encoder and the image encoder are obtained after fine-tuning the parameters on the basis of a multi-modal pre-trained neural network, and the multi-modal pre-trained neural network is pre-trained in a multi-modal data environment, which includes an image encoder for processing images and a text encoder for processing text.
[0018] Among them, it further includes:
[0019] Obtain a first prediction score for face attractiveness prediction according to the fused face image features;
[0020] Obtain a second prediction score for face attractiveness prediction according to the image encoding vector and the multiple text encoding vectors before optimization;
[0021] Obtain a third prediction score for face attractiveness prediction according to the image encoding vector and the multiple optimized text encoding vectors;
[0022] By fusing the first prediction score, the second prediction score, and the third prediction score, an attraction prediction result of the face image is obtained.
[0023] Among them, the face image to be predicted includes:
[0024] Performing image sampling from the live stream;
[0025] Determine whether there is face image content in the collected image, and determine whether the size, clarity, facial pose, and occlusion of the face image content meet the conditions. If so, determine the collected image as the face image to be predicted.
[0026] Among them, the performing image sampling from the live stream includes:
[0027] By performing multiple samplings on the live stream of the same live room at multiple different time points to obtain multiple face images to be predicted, so as to evaluate the change of the face image attraction in the live room according to the attraction prediction results of multiple face images, or evaluate the image processing effect of the facial beautification algorithm and / or image enhancement algorithm used in the live room.
[0028] Among them, by evaluating the face image attractions corresponding to multiple live rooms in the live system, the overall face image attraction of the live system is evaluated.
[0029] A method for training a face image attraction prediction model, the prediction model includes a face prior feature introduction module, a multi-modal encoding module, and a cross-modal fusion module; the face prior feature introduction module includes a deep learning network and a pre-trained face recognition network for introducing face prior features based on face IDs into the face images extracted by the deep learning network; the multi-modal encoding module includes a text encoder and an image encoder, and the text encoder is used for text encoding according to multiple text contents constructed based on multiple divided attraction levels; the method includes:
[0030] Obtain a training data set, where the training data set includes multiple face images and corresponding attraction annotation information;
[0031] In the first training stage, parallel training is performed on the face prior feature introduction module and the multi-modal encoding module; among them, regarding the face prior feature introduction module, while keeping the parameters of the pre-trained face recognition network unchanged, the deep learning network is trained using the training data.
[0032] In the second training stage, with the parameters of the face prior feature introduction module and the multi-modal encoding module fixed, the cross-modal fusion module is trained;
[0033] Specifically, in the second training stage, by inputting the face images in the training data into the face prior feature introduction module trained in the first stage, face image features fused with face prior features based on face IDs are obtained. The face images are input into the image encoder trained in the first stage to obtain image encoding vectors, and the multiple text contents are input into the text encoder trained in the first stage to obtain multiple text vectors. Through the cross-modal fusion module, the multiple text vectors are optimized using the face image features fused with face prior features based on face IDs to obtain optimized multiple text encoding vectors. By calculating the similarity between the image encoding vectors and the optimized multiple text encoding vectors respectively, the probabilities of the face images belonging to the multiple attractiveness levels are obtained. The probabilities of the multiple attractiveness levels are weighted and summed with the level scores corresponding to the multiple attractiveness levels to obtain an attractiveness prediction score for the face image, so as to optimize the parameters of the cross-modal fusion module using a preset loss function to complete the training of the cross-modal fusion module.
[0034] Among them, a decision fusion module is further included in the prediction model;
[0035] When training the face prior feature introduction module in the first training stage, face image features of the face images in the training data are obtained through a deep learning network, face prior features based on face IDs of the face images are obtained through a pre-trained face recognition network, the face image features are fused with the face prior features to obtain fused face image features, a first prediction score corresponding to the face image is obtained according to the fused face image features, and the parameters of the deep learning network are optimized using a preset loss function to complete the training of the face prior feature introduction module;
[0036] When training the multi-modal encoding module in the first training stage, use the image encoder to perform image encoding on the face image to obtain an image encoding vector; construct multiple pieces of text content by dividing into multiple attractiveness levels, and use the text encoder to encode each piece of the text content respectively to obtain multiple text encoding vectors; calculate the similarity between the image encoding vector and the multiple text encoding vectors respectively to obtain the probabilities that the face image belongs to the multiple attractiveness levels; and perform weighted summation on the probabilities of the multiple attractiveness levels and the level scores corresponding to the multiple attractiveness levels to obtain a second prediction score, and use a preset loss function to optimize the parameters of the multi-modal encoding module to complete the training of the multi-modal encoding module;
[0037] In the second training stage, use the attractiveness prediction score output by the cross-modal fusion module as the third prediction score, and fuse the first prediction score, the second prediction score, and the third prediction score through the decision fusion module to obtain an attractiveness prediction result, and use a preset loss function to optimize the parameters of the cross-modal fusion module and the decision fusion module to complete the training of the cross-modal fusion module and the decision fusion module.
[0038] A computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the foregoing are implemented.
[0039] An electronic device, comprising:
[0040] One or more processors; and
[0041] A memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the foregoing are executed.
[0042] A computer program product, comprising computer programs / computer-executable instructions, and when the computer programs / computer-executable instructions are executed by a processor in an electronic device, the steps of the method described in any one of the foregoing are implemented.
[0043] According to the specific embodiments provided in the present application, the present application discloses the following technical effects:
[0044] Through the embodiments of the present application, when predicting the attractiveness of a face image, prior face features based on the face ID can be introduced, and multiple pieces of text content constructed according to multiple divided attractiveness levels can be used to achieve cross-modal fusion between text features and image features. Among them, the introduction of prior face features based on the face ID can more accurately and comprehensively capture the individual features of the face, and improve the robustness of the algorithm to various complex lighting, face postures, and expression changes. By further optimizing the extracted text features using image features containing rich face semantic information and performing cross-modal semantic fusion, the semantic information of the features can be made richer, so as to more deeply model the personalized differences between different faces and improve the accuracy of the face attractiveness prediction results.
[0045] Of course, when implementing any product of the present application, it is not necessarily required to achieve all the above-mentioned advantages simultaneously. BRIEF DESCRIPTION OF THE DRAWINGS
[0046] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following will briefly introduce the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present application. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0047] Figure 1 is a schematic diagram of the algorithm architecture provided by the embodiments of the present application;
[0048] Figure 2 is a flowchart of the prediction method provided by the embodiments of the present application;
[0049] Figure 3 is a flowchart of the model training method provided by the embodiments of the present application;
[0050] Figure 4 is a schematic diagram of the electronic device provided by the embodiments of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0051] The following will clearly and completely describe the technical solutions in the embodiments of the present application with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only some embodiments of the present application, rather than all embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art belong to the scope of protection of the present application.
[0052] First of all, it should be noted that the inventors of this application found in the process of implementing this application that the existing FAP / FBP algorithms are still insufficient in dealing with complex contexts and subtle differences in understanding tasks, and still have very limited generalization and adaptation capabilities in more challenging scenarios. The main reasons are as follows:
[0053] First, the existing FAP / FBP algorithms have all made certain developments in deep learning modeling, including using architectures such as CNN and Transformer. However, they only train and learn the model based on the image data itself, that is, they use single-modal modeling techniques that rely on simple splicing of visual features, without combining data from more modalities.
[0054] Since they do not train and learn through cross-modal data, that is, they cannot extract relevant features from multiple information sources, it further leads to insufficient work in dealing with complex contexts in understanding tasks.
[0055] Second, in terms of face prior features, the existing FAP / FBP methods introduce richer facial representation information, including information such as relying on face key points, artificially designed geometric features, and person attributes as face prior features into the model. However, these features have limitations in representing complex facial information, can only represent some features related to face attractiveness, and cannot dynamically model the subtle differences between different faces, so there is still room for improvement in facial feature expression.
[0056] In addition, in terms of training data, the existing FAP / FBP methods lack training data corresponding to some special scenarios. For example, the demand for face attractiveness evaluation in the live broadcast scenario is relatively high. However, due to many uncontrollable situations such as lighting, facial expressions, camera angles, and background noise in the live broadcast scenario, if the facial data of other scenarios is directly used and transferred to the live broadcast scenario, it may cause the model to overfit the facial training data of other scenarios, limiting the generalization ability of the model.
[0057] Based on the above situation, in the embodiments of the present application, first, in view of the deficiencies of existing algorithms in understanding complex contexts in tasks and reflecting subtle differences between different faces, corresponding technical solutions are provided. In this solution, a face image attractiveness prediction model is provided. In this model, when first extracting features from a face image through a deep learning model, face prior features based on face IDs can be introduced. Among them, the deep learning model used in the embodiments of the present application can be a model based on architectures such as Transformer. When using such a model for face image feature extraction, usually, features can be extracted from aspects such as image texture, but usually, the subtle differences between different faces cannot be expressed. Therefore, face prior features can also be introduced to enhance the face image features in reflecting the subtle differences between different faces.
[0058] Among them, in the embodiments of the present application, instead of using some manually designed features as face prior features, a face ID-based method can be adopted to obtain more comprehensive prior features about the face. To achieve this purpose, it can be realized by means of a pre-trained face recognition network. Among them, the so-called pre-trained face recognition network here can refer to a deep learning model that has already been trained using a large amount of face datasets. Since a sufficient amount of face data is used, it can be considered that identity information such as face IDs is already included. And after being trained on a large amount of face data, such a face recognition network has a unique perception of a person's eyebrows, eyes, etc., and can learn how to judge the unique features that distinguish a face image from other face images, and so on. Therefore, in the embodiments of the present application, the above capabilities of the above-mentioned pre-trained face recognition network can be utilized to obtain face prior features based on face IDs. Specifically, the face image is input into the above-mentioned face recognition network, and the features of the face image output by the face recognition network can be used as the available face prior features based on face IDs in the embodiments of the present application. Then, the face prior features based on face IDs can be introduced into the face image features extracted by the deep learning model used in the embodiments of the present application to enhance the face image features. Among them, in a preferred manner, the face image features extracted by the deep learning model used in the embodiments of the present application can be multi-scale features, that is, they can be generated by splicing and other processing of features with different receptive fields output at different stages of the deep learning network to obtain a better prediction effect.
[0059] In addition to introducing face prior features based on face ID into face image features, the embodiments of the present application can also perform multi-modal feature fusion. To achieve this goal, the embodiments of the present application can first construct some text data to achieve cross-modal fusion with image data, and then obtain richer features. Among them, the specific text content can be generated according to multiple pre-divided attractiveness levels, and each attractiveness level can correspond to a piece of text content. Then, the above text content can be encoded separately by a text encoder to obtain multiple text vectors (embeddings). For example, if there are a total of 5 pieces of text content, 5 text vectors can be obtained. In addition, the face image features that have been fused with the face prior features based on face ID can also be used to optimize the text vectors to obtain optimized text vectors. These optimized multiple text vectors can be used to calculate the similarity with the image encoding vectors obtained by encoding the face image by an image encoder, and the probability that the face image belongs to the above various attractiveness levels can be obtained. This probability can be further used to determine the attractiveness prediction result of the face image. For example, the attractiveness level corresponding to the highest predicted probability value can be used as the qualitative attractiveness prediction result. Or, in another way, the probability corresponding to the above various attractiveness levels can be weighted and summed with the level scores corresponding to multiple attractiveness levels to obtain an attractiveness prediction score for the face image, that is, a quantitative attractiveness prediction result, and so on.
[0060] In addition, in another way, for the features obtained by fusing face image features and face prior features based on face ID, a prediction score (which can be called the first prediction score) for the attractiveness of the face image can also be obtained by performing dimensionality reduction processing on it. The second prediction score for face attractiveness prediction can also be obtained according to the image encoding vector of the face image and the multiple text encoding vectors before optimization. The attractiveness prediction score obtained according to the image encoding vector and the optimized multiple text encoding vectors can be used as the third prediction score. Then, the attractiveness prediction result of the face image can be obtained by fusing the first prediction score, the second prediction score, and the third prediction score.
[0061] It can be seen that the face image attractiveness prediction model provided by the embodiments of the present application has the following capabilities: First, it can introduce face prior features based on face ID into face image features. Second, it has multi-modal encoding capabilities and can encode data in multiple modalities such as images and texts. Moreover, it also needs to have cross-modal fusion capabilities and can perform cross-modal fusion of image features and text features to obtain richer features.
[0062] To achieve the above purpose, specifically in terms of the model structure, such as Figure 1As shown, a specific prediction model can be divided into multiple modules. For example, it can specifically include a face prior feature introduction module (corresponding to the ability to introduce face prior features based on face IDs into face image features), a multi-modal encoding module, and a cross-modal fusion module. Additionally, in an optional manner, a decision fusion module can also be included. Specifically, when training the model, a phased approach can be adopted.
[0063] First, in the first stage, the face prior feature introduction module and the multi-modal encoding module can be trained. Among them, the training of these two modules can be carried out in parallel. For the face prior feature introduction module, specifically, a deep learning network can be used as the backbone network, and additionally, a pre-trained face recognition network can be used to obtain face prior features based on face IDs. Specifically, the deep learning network can be a structure such as Swin Transformer (a model based on the Transformer architecture, especially suitable for computer vision tasks. It combines the self-attention mechanism of Transformer and the local perception characteristics of CNN (Convolutional Neural Network)). Its role is to extract face image features for the input face image. Of course, there may be situations where the subtle differences between faces in the extracted features are not expressed clearly enough. Therefore, face prior features based on face IDs can be introduced to enhance the face image features. In a preferred manner, the face image features extracted by the Swin Transformer deep learning network can be multi-scale features, that is, generated by processing the features with different receptive fields output at different stages of the deep learning network. After obtaining the face image features and face prior features, a cross-attention module can be used to perform feature fusion between the face image features and the face prior features to obtain optimized face image features.
[0064] Specifically, when performing the above-mentioned feature fusion based on the cross-attention module, the face image features can be used as the Query parameter of the cross-attention module, and the face prior features can be used as the Key and Value parameters of the cross-attention module (the face prior features can be mapped to two different vectors, serving as Key and Value respectively), and then the cross-attention module performs feature fusion processing, thereby realizing the enhancement of the face image features by the face prior features.
[0065] After completing the above feature fusion, modules such as average pooling and multi-layer perceptron can be used to perform dimensionality reduction on the fused features. For example, assume that the number of training samples in the same batch is B, and the features obtained after feature fusion are B×N×C dimensions. Through average pooling, N can be reduced to 1, and through a multi-layer perceptron, C can be reduced to 1, thus obtaining features of B×1. That is, each training sample corresponds to a value, and this value can be used as the predicted face attractiveness score corresponding to the specific training sample. Of course, this predicted score is only obtained by prediction after introducing the face prior feature based on the face ID. In the embodiments of the present application, this score is not the final prediction result. However, since the face prior feature introduction module needs to be trained separately, the value obtained above can be temporarily used as a predicted score (which can be called the first predicted score) so that by optimizing the preset loss function, this predicted score gradually approaches the labeled attractiveness score, thereby completing the training of the face prior feature introduction module. Among them, the loss function used here can be the L1 (Mean Absolute Error) loss function, etc. In a preferred implementation, this first predicted score can also be fused with the predicted scores obtained by other methods to obtain the final prediction result.
[0066] It should be noted that during the above process of training the face prior feature introduction module, for the pre-trained face recognition network, its parameters can be kept fixed, and only the part of the deep learning network such as the aforementioned Swin Transformer structure needs to be trained.
[0067] The first training phase also involves training the multi-modal encoding module. In the embodiments of the present application, the specific multi-modal encoding module may include an image encoder and a text encoder. The image encoder and the text encoder can be trained with training samples to enable them to extract features related to human face attractiveness prediction. In a preferred implementation, the specific text encoder and image encoder can be obtained by fine-tuning the parameters based on a multi-modal pre-trained neural network. The multi-modal pre-trained neural network is pre-trained in a multi-modal data environment. For example, it can be a CLIP (Contrastive Language-Image Pre-training) model, etc. Since it itself exactly includes an image encoder for processing images and a text encoder for processing text, and this pre-trained neural network is also trained in a multi-modal data environment and has the ability of multi-modal encoding, therefore, based on these capabilities, fine-tuning its parameters according to the training samples adopted in the embodiments of the present application can obtain the trained image encoder and text encoder. By this way of fine-tuning based on the existing pre-trained neural network, the powerful representation ability of the pre-trained model on massive data can be utilized, and at the same time, by fusing the features of the image and text modalities, the semantic information is greatly enriched, contributing to the improvement of the model performance.
[0068] Among them, specifically when training the multi-modal encoding module, the image encoder can receive a face image to generate an image encoding vector (Embeddings), and the text encoder can receive text to generate a text encoding vector. Among them, the text in the embodiments of the present application can be constructed according to specific scenarios. Specifically, when constructing text data, the face attractiveness can first be divided into multiple levels. For example, it can specifically be 5 levels (а∈{"bad","poor","fair","good","perfect"}). Additionally, a text template can be provided to combine into specific text content. For example, the text template can be "a photo of a person with {а} attractiveness (a face photo with a certain attractiveness)". In this way, 5 pieces of text content can be obtained, which are respectively: "a photo of a person with bad attractiveness", "a photo of a person with poor attractiveness", "a photo of a person with fair attractiveness", "a photo of a person with good attractiveness", "a photo of a person with perfect attractiveness". Then, the text encoder can encode each of the above pieces of text content respectively to obtain multiple text encoding vectors.
[0069] To implement the training of the above multi-modal encoding module, after obtaining the image encoding vector and multiple text encoding vectors, the similarity regression method can be used to calculate the similarity between the image encoding vector and each text encoding vector respectively, to express the probability that the specific face image (data in the training sample) predicted by this module belongs to each attractiveness level. After that, the probability corresponding to the above various attractiveness levels can be weighted and summed with the level scores corresponding to multiple attractiveness levels to obtain the attractiveness prediction score of the face image predicted by this multi-modal encoding module. This score can be called the second prediction score. After obtaining this second prediction score, the preset loss function can be optimized to make this second prediction score gradually approach the labeled attractiveness score, thereby completing the training of the multi-modal encoding module. Among them, the loss function used here can include the L1 loss function, the rank loss function (Fidelity Loss, Two-direction Ranking Loss), and so on. In a preferred implementation manner, this second prediction score can also be fused with the prediction scores obtained by other methods to obtain the final prediction result.
[0070] After the above-mentioned first-stage training is completed, the parameters of the face prior feature introduction module and the multi-modal encoding module can be fixed. However, these two trained modules will participate in the training process of the cross-modal fusion module in the second stage. Specifically, first, the face images in the training samples can be input into the face prior feature introduction module trained in the first stage to obtain face image features fused with face prior features based on face IDs. This feature can be a multi-scale feature. Additionally, the face images and multiple pieces of text content assembled in the embodiments of the present application can be input into the multi-modal encoding module trained in the first stage to obtain an image encoding vector and multiple text encoding vectors. Then, through the cross-modal fusion module, in a cross-attention manner, the face image features fused with face prior features based on face IDs can be further fused with the multiple text encoding vectors. When performing cross-attention fusion here, the text encoding vectors can be used as the Query parameter of the cross-attention module, and the face image features fused with face prior features based on face IDs can be used as the Key and Value parameters of the cross-attention module, thereby optimizing or enhancing the text encoding vectors through the face image features fused with face prior features based on face IDs, and adaptively representing the semantic information of each attraction level related to the current face.
[0071] After obtaining the optimized text encoding vectors, the similarity between each of the optimized text encoding vectors and the image encoding vector extracted by the aforementioned multi-modal encoding module can be calculated through a similarity regression method, so as to obtain the probability that the specifically input face image belongs to each attraction level. Then, the probability corresponding to each level can be weighted and summed with the level scores of each attraction level to obtain a value, which can be used as the face attractiveness prediction score predicted by the above cross-modal fusion module, specifically referred to as the third prediction score. Then, a preset loss function can be used to optimize the parameters of the cross-modal fusion module to complete the training of the cross-modal fusion module.
[0072] Since the above cross-modal fusion module has already introduced the face prior features based on the face ID and achieved cross-modal feature fusion, the third prediction score generated by it can have a good prediction effect and can be directly output as the final prediction result. Alternatively, in another more preferred implementation, a decision fusion module can be further added. Through this decision fusion module, the aforementioned first prediction score, second prediction score, and third prediction score can be further fused. For example, they can be fused through two linear layers or other means to obtain a fused score, which can be output as the final prediction result to further improve the prediction effect. Of course, the parameters of this decision fusion module can also participate in the training. Specifically, during the aforementioned second training stage, it can be trained together with the cross-modal fusion module. That is to say, after the cross-modal fusion module outputs the third prediction score, the decision fusion module can first fuse the first prediction score, second prediction score, and third prediction score to obtain a fused score, and then use the preset loss function to optimize the parameters of the cross-modal fusion module and the decision fusion module so that the fused score gradually approaches the labeled value, thereby completing the training of the cross-modal fusion module and the decision fusion module and ending the entire training process.
[0073] After completing the training of each module of the aforementioned prediction model, it can be deployed to a specific online environment to be used for predicting the attractiveness of the captured face images in specific application scenarios.
[0074] The capabilities, composition structures, and training methods of the prediction module provided in the embodiments of the present application are introduced above. Next, the acquisition method of specific training samples will be introduced. It should be noted here that the method for predicting face attractiveness provided in the embodiments of the present application can be unrestricted by application scenarios. For any image containing face image content in any scenario, the solution provided in the embodiments of the present application can be used for attractiveness testing. However, if it is necessary to use it in special application scenarios such as live broadcasts and short videos, it is recommended to use training samples closer to the application scenario as much as possible to complete the training process of the specific model, so as to improve the prediction effect of the specific model in the specific application scenario. Taking the live broadcast scenario as an example, the acquisition method of specific training sample data will be introduced below.
[0075] Step 1: Sample at a constant rate of one frame per second from the live stream to obtain multiple images. Among them, in the scenario of collecting training samples, sampling can be performed separately from the live streams corresponding to multiple live rooms to obtain more training sample data. Among them, sampling can be preferentially performed from the live streams corresponding to some live rooms with relatively high playback volumes and having a relatively large impact on the face attractiveness of the entire live broadcast system.
[0076] Step 2: Each sampled frame image is first passed through a face detection algorithm to detect whether there is a face in the image. If there is no face, the frame image can be skipped; if there is a face, the coordinate information of the face Bouding Box (anchor box, used to represent the true position of the target in the image) can be returned; if there are multiple face images, all face images can be retained; then the detected faces can be cropped according to the Bouding Box and input into the next step.
[0077] Step 3: Detect the face size. If the short side size of the cropped face image is less than a preset threshold (e.g., 90 pixels), skip the image.
[0078] Step 4: Clarity detection. Calculate the Laplacian variance of the Y channel of the image. If the variance is greater than a preset value 2 (e.g., 10), it indicates that the face clarity is poor, and the image is skipped.
[0079] Step 5: Face pose detection. Through a face pose estimation algorithm, estimate the deflection degrees of the pitch angle and yaw angle of the face, and retain the face images with the deflection degree of the pitch angle less than the first threshold (e.g., 20) and the deflection degree of the yaw angle less than the second threshold (e.g., 15).
[0080] Step 6: Face occlusion detection. Calculate the effective face ratio through a face segmentation algorithm. The denominator of this effective face ratio is the area of the Bouding Box, and the numerator is the face skin area segmented by the segmentation algorithm; a more accurate denominator can also be obtained through the face key points at the facial edge. Skip the face images with an effective face ratio lower than a certain threshold (e.g., 0.6).
[0081] Step 7: For the remaining face images after the above 6 steps of processing, duplicate faces and other abnormal faces can also be removed manually.
[0082] After filtering through the above steps, if there are no face images, the live broadcast room can be skipped or re-sampled. If there are face images, manual annotation can be performed to obtain a subjective quality score (MOS, Mean Opinion Score, an average score obtained by having a group of evaluators rate the training samples, aiming to evaluate the subjective experience of users), and a high-quality training dataset in the live broadcast scenario can be sorted out. The data included in this dataset are: multiple face images and their corresponding MOS scores respectively. The training samples in this training dataset can be used to train the aforementioned prediction model provided by the embodiments of the present application.
[0083] In the above manner, since after the images are collected from the live stream, low-quality faces such as unclear faces, faces that are too small in size, faces with large deflection angles, and faces with large occlusions can be filtered out, the face training data for the live scene is efficiently obtained for subsequent model training. It can also be used as a preprocessing process for the input of the face attractiveness prediction model on the live platform. Since the above-mentioned low-quality faces are filtered out, this preprocessing process can enhance the reliability and robustness of the online deployment algorithm to a certain extent.
[0084] Of course, after the model training is specifically completed, in the stage of specifically using the model for inference in the live scene, data can also be collected from the live stream in a similar manner as described above for the specific inference process. Different from the process of collecting training samples, in the inference stage, sampling can be performed at different time periods during the live broadcast of a specific live room, and there is no need to perform duplicate removal processing. After that, the attractiveness of multiple face images obtained by sampling can be evaluated respectively, so as to evaluate whether the state of the anchor in a specific live room has changed at different time periods before and after a live broadcast. Or, the attractiveness evaluation results of faces at different time periods can also be averaged to obtain the average face attractiveness evaluation result of the entire live broadcast, and so on.
[0085] The following details the specific implementation solutions provided by the embodiments of the present application.
[0086] Embodiment 1
[0087] First, Embodiment 1 provides a method for predicting the attractiveness of a face image. Refer to Figure 2 , this method may include:
[0088] S201: Determine the face image to be predicted.
[0089] The specific face image to be detected can be obtained by sampling from a specific application scenario. For example, assuming that it is necessary to predict the attractiveness of a face image in a certain live room, image sampling can be performed from the live stream of this live room, and then it can be determined whether there is face image content in the collected image, and whether the clarity, size, facial pose, and occlusion situation of the face image content meet the conditions. If so, the collected image is determined as the face image to be predicted.
[0090] Among them, in the process of predicting the facial attractiveness of a specific live streaming room, since the live streaming duration of a live streaming room may last for a relatively long time, during the live streaming process of the same anchor, their state may change. For example, after a long period of live streaming, the anchor may be relatively tired, which may also affect the attractiveness of their face, and so on. Therefore, specifically when predicting the facial attractiveness of a live streaming room, multiple samples can be taken at multiple different time points from the live stream of the same live streaming room to obtain multiple facial images to be predicted, so as to evaluate the change of the facial image attractiveness of the live streaming room according to the attractiveness prediction results of multiple facial images. Or, the image processing effect of the facial beautification algorithm and / or image enhancement algorithm used in the live streaming room can also be evaluated by taking the average of the attractiveness prediction results at different time points, and so on. Among them, regarding the image enhancement algorithm, it mainly refers to image quality enhancement, which is a very important function in live streaming and short videos. Generally speaking, the enhancement of image quality also includes the enhancement of the image quality of the face part. However, for facial attractiveness, the enhancement of image quality may result in different situations such as improvement or deterioration. Therefore, the facial attractiveness prediction result provided by the embodiments of the present application can also be used as auxiliary information to help determine whether the relevant image quality enhancement function is positive.
[0091] In addition, there are usually multiple live streaming rooms in a live streaming platform or live streaming system that are live streaming simultaneously. Therefore, the attractiveness of the corresponding facial images can also be evaluated respectively to evaluate the overall facial image attractiveness of the live streaming system. That is to say, it can be used to evaluate the facial attractiveness level of the entire live streaming business market. For example, the impact of the launch of a certain core beauty algorithm on the market, and the comparison of the overall appearance level of the live streaming market of competitors.
[0092] S202: In the process of obtaining the facial image features through the deep learning network, introduce the facial prior features based on the facial ID provided by the face recognition network to obtain the fused facial image features.
[0093] After obtaining the face image, first, the face image features can be obtained through a pre-trained deep learning network. Here, the deep learning model can be the deep learning model used in the face prior feature introduction module in the prediction model provided in the embodiments of the present application. This deep learning model has been pre-trained and has the ability to extract face features for the attractiveness prediction task. In a preferred implementation, this deep learning model can obtain multi-scale features of the face image. The multi-scale features can be generated by processing features with different receptive fields output at different stages of the deep learning network, such as splicing. Additionally, since the face prior feature introduction module also includes a face recognition network, the face model can also be input into this face recognition network, and this face recognition network can output face prior features based on the face ID of the face image. The face prior features can express features such as the subtle differences between the current face image and other face images. Then, through methods such as cross-attention, the above multi-scale features can be fused with the face prior features based on the face ID to obtain multi-scale features fused with the face prior features based on the face ID, thereby realizing the enhancement or optimization of the foregoing multi-scale features through the face prior features based on the face ID.
[0094] S203: Construct multiple text contents by dividing into multiple attractiveness levels, and use a text encoder to encode each of the multiple text contents respectively to obtain multiple text encoding vectors. The multiple text contents are respectively used to textually describe the attractiveness of the face image through the multiple attractiveness levels.
[0095] The above steps achieve the introduction of the face prior features based on the face ID. In step S203, it can prepare for multi-modal fusion. Specifically, to achieve multi-modal fusion, the embodiments of the present application can first construct text contents. Specifically, multiple text contents can be constructed by dividing into multiple attractiveness levels. These text contents are respectively used to textually describe the attractiveness of the face image through multiple attractiveness levels. Among them, the number of text contents is the same as the number of divided attractiveness levels, and these text contents can be shared when predicting the attractiveness of different face images.
[0096] Then, a text encoder can be used to encode each of the multiple text contents respectively to obtain multiple text encoding vectors. Here, the text encoder can be the text encoder in the multi-modal encoding module in the prediction model provided in the embodiments of the present application. This text encoder can be trained on a specific dataset (for example, it can be training sample data sampled from a live stream and preferably selected for a live broadcast scenario, etc.), so that it has the ability to perform text encoding for the attractiveness prediction task.
[0097] S204: Use the fused facial image features to perform cross-modal optimization processing on the text encoding vectors to obtain multiple optimized text encoding vectors.
[0098] After obtaining the facial image features fused with the facial prior features based on the facial ID and multiple text encoding vectors, the two can be fused through cross-attention or other means to enhance or optimize the text encoding vectors by using the facial image features fused with the facial prior features based on the facial ID, thereby obtaining optimized text encoding vectors. In this way, cross-modal feature fusion can be achieved, and richer features can be obtained. Specifically, when implemented, this cross-modal fusion operation can be completed by the cross-modal fusion module in the prediction model provided by the embodiments of the present application. The cross-modal fusion module has also been pre-trained on a specific data set to enable it to perform cross-modal fusion for the attractiveness prediction task.
[0099] S205: Use an image encoder to perform image encoding on the facial image to obtain an image encoding vector.
[0100] In addition, an image encoder can also be used to perform image encoding on the facial image to obtain an image encoding vector. The image encoder here can be the image encoder in the multi-modal encoding module of the prediction model provided by the embodiments of the present application. The image encoder has also been trained on a specific data set to enable it to perform image encoding for the attractiveness prediction task.
[0101] S206: By calculating the similarity between the image encoding vector and the multiple optimized text encoding vectors respectively, output the probabilities that the facial image belongs to the multiple attractiveness levels respectively, so as to be used to determine the attractiveness prediction result of the facial image.
[0102] After introducing the facial prior features based on the facial ID and performing cross-modal feature fusion, multiple optimized text encoding vectors are obtained. In addition, an image encoding vector of the facial image is obtained through an image encoder. Therefore, by calculating the similarity between the image encoding vector and the multiple optimized text encoding vectors respectively, the probabilities that the facial image belongs to the multiple attractiveness levels can be output, and this probability information can be used to determine the attractiveness prediction result of the facial image.
[0103] For example, in one way, the attraction level corresponding to the predicted highest probability can be directly output as the attraction prediction result. Alternatively, the probabilities corresponding to multiple attraction levels can be weighted and summed with the level scores corresponding to multiple preset attraction levels (for example, the level scores corresponding to 5 attraction levels are 1, 2, 3, 4, 5, etc.), obtaining a value, and then this value can be output as the final attraction score.
[0104] Or, in another way, after obtaining the face image features that incorporate the face prior features based on the face ID, a score can also be calculated based on these features. For example, after dimensionality reduction processing through average pooling, multi-layer perceptron, etc., a score can be obtained, and this score can be used as the prediction result of the face prior feature introduction module alone, called the first prediction score. In the multi-modal encoding module, before optimizing the text encoding vector, the probabilities that the face images separately predicted by this multi-modal encoding module belong to multiple attraction levels can also be obtained by calculating the similarity between the image encoding vector and the text encoding vector before optimization. Then, by weighted summing the probabilities corresponding to these multiple attraction levels with the level scores corresponding to multiple preset attraction levels, a score can also be calculated, and this score can be used as the prediction result separately predicted by this multi-modal encoding module, called the second prediction score. Correspondingly, after the text encoding vector is optimized by the cross-modal fusion module, the score obtained after calculating the similarity with the image encoding vector and weighted summing with the level scores can be used as the prediction result separately predicted by the cross-modal fusion module, called the third prediction score. After that, the above first prediction score, second prediction score, and third prediction score can be fused to obtain the final prediction score for output. The process of fusing multiple prediction scores can be completed by the decision fusion module in the prediction model provided by the embodiments of the present application, and this decision fusion module can also be pre-trained on a specific data set to enable it to have the ability to fuse scores for the attraction prediction task.
[0105] In summary, through the embodiments of the present application, when predicting the attractiveness of a face image, face prior features based on face IDs can be introduced, and multiple text contents constructed according to multiple divided attractiveness levels can also be used to achieve cross-modal fusion between text features and image features. Among them, the introduction of face prior features based on face IDs can more accurately and comprehensively capture the individual features of the face, and improve the robustness of the algorithm to various complex illuminations, face postures, and expression changes. By further optimizing the extracted text features with image features containing rich face semantic information and performing cross-modal semantic fusion, the semantic information of the features can be made richer, so as to more deeply model the personalized differences between different faces and improve the accuracy of the face attractiveness prediction results.
[0106] Embodiment 2
[0107] This Embodiment 2 provides separate protection for the specific training method of the prediction model. Specifically, this Embodiment 2 provides a method for training a face image attractiveness prediction model, where the prediction model includes a face prior feature introduction module, a multi-modal encoding module, and a cross-modal fusion module; the face prior feature introduction module includes a deep learning network and a pre-trained face recognition network for introducing face prior features based on face IDs into the face images extracted by the deep learning network; the multi-modal encoding module includes a text encoder and an image encoder, and the text encoder is used to perform text encoding on multiple text contents constructed according to multiple divided attractiveness levels; see Figure 3 , the method includes:
[0108] S301: Obtain a training data set, where the training data set includes multiple face images and corresponding attractiveness annotation information;
[0109] S302: In the first training stage, perform parallel training on the face prior feature introduction module and the multi-modal encoding module; among them, regarding the face prior feature introduction module, with the parameters of the pre-trained face recognition network fixed, use the training data to train the deep learning network;
[0110] S303: In the second training stage, with the parameters of the face prior feature introduction module and the multi-modal encoding module fixed, train the cross-modal fusion module;
[0111] Among them, in the second training stage, by inputting the face images in the training data into the face prior feature introduction module trained in the first stage, face image features fused with face prior features based on face IDs are obtained. The face images are input into the image encoder trained in the first stage to obtain image encoding vectors. The multiple text contents are input into the text encoder trained in the first stage to obtain multiple text vectors. And through the cross-modal fusion module, the multiple text vectors are optimized by using the face image features fused with face prior features based on face IDs to obtain optimized multiple text encoding vectors. By calculating the similarity between the image encoding vectors and the optimized multiple text encoding vectors respectively, as the probabilities that the face images belong to the multiple attractiveness levels, the probabilities of the multiple attractiveness levels are weighted and summed with the level scores corresponding to the multiple attractiveness levels to obtain an attractiveness prediction score for the face image, so as to optimize the parameters of the cross-modal fusion module by using a preset loss function to complete the training of the cross-modal fusion module.
[0112] Among them, in an optional manner, the prediction model may further include a decision fusion module. In this case, when training the face prior feature introduction module in the first training stage, face image features of the face images in the training data are obtained through a deep learning network, face prior features based on face IDs of the face images are obtained through a pre-trained face recognition network, the face image features are fused with the face prior features to obtain fused face image features, a first prediction score corresponding to the face image is obtained according to the fused face image features, and the parameters of the deep learning network are optimized by using a preset loss function and the attractiveness annotation information to complete the training of the face prior feature introduction module.
[0113] When training the multi-modal encoding module in the first training stage, the face images are image-encoded by using an image encoder to obtain image encoding vectors. Multiple text contents are constructed by dividing into multiple attractiveness levels, and the multiple text contents are respectively encoded by using a text encoder to obtain multiple text encoding vectors. The similarity between the image encoding vectors and the multiple text encoding vectors is calculated respectively to obtain the probabilities that the face images belong to the multiple attractiveness levels. And the probabilities of the multiple attractiveness levels are weighted and summed with the level scores corresponding to the multiple attractiveness levels to obtain a second prediction score, and the parameters of the multi-modal encoding module are optimized by using a preset loss function and the attractiveness annotation information to complete the training of the multi-modal encoding module.
[0114] In the second training stage, the attraction prediction score output by the cross-modal fusion module can be used as the third prediction score. The decision fusion module fuses the first prediction score, the second prediction score, and the third prediction score to obtain an attraction prediction result, and uses a preset loss function and the attraction annotation information to optimize the parameters of the cross-modal fusion module and the decision fusion module, so as to complete the training of the cross-modal fusion module and the decision fusion module.
[0115] Regarding the parts not described in detail in the second embodiment, reference may be made to the description in the first embodiment and other parts of this specification, which will not be elaborated here.
[0116] It should be noted that the embodiments of the present application may involve the use of user data. In actual applications, user-specific personal data can be used in the solutions described herein within the scope permitted by applicable laws and regulations (for example, with the user's explicit consent, giving the user a practical notice, etc.) and in compliance with the requirements of applicable laws and regulations of the country where it is located.
[0117] Corresponding to the first embodiment, the embodiment of the present application further provides a face image attraction prediction device, which may include:
[0118] A face image determination unit, configured to determine a face image to be predicted;
[0119] A face prior feature introduction unit, configured to introduce a face prior feature based on a face ID provided by a face recognition network during the process of obtaining the face image features through a deep learning network, so as to obtain fused face image features;
[0120] A text encoding unit, configured to construct multiple text contents by dividing multiple attraction levels, and use a text encoder to encode the multiple text contents respectively to obtain multiple text encoding vectors, where the multiple text contents are respectively used to textually describe the face image attraction through the multiple attraction levels;
[0121] A text encoding vector optimization unit, configured to perform cross-modal optimization processing on the text encoding vectors by using the fused face image features to obtain optimized multiple text encoding vectors;
[0122] An image encoding unit, configured to perform image encoding on the face image by using an image encoder to obtain an image encoding vector;
[0123] A similarity calculation unit, configured to output the probabilities that the face image belongs to the multiple attraction levels respectively by calculating the similarities between the image encoding vector and the optimized multiple text encoding vectors respectively, so as to be used to determine the attraction prediction result of the face image.
[0124] Among them, in a preferred implementation, the face image features obtained through the deep learning network are multi-scale features, and the multi-scale features are generated by processing the features of different receptive fields output by different stages of the deep learning network.
[0125] Specifically, the face prior features based on the face ID provided by the pre-trained face recognition network are used to express the features that distinguish the face image from other images.
[0126] The text encoder and the image encoder are obtained after parameter fine-tuning training on the basis of a multi-modal pre-trained neural network. The multi-modal pre-trained neural network is pre-trained in a multi-modal data environment, which includes an image encoder for processing images and a text encoder for processing text.
[0127] In addition, the device may further include:
[0128] A first prediction score determination unit, configured to obtain a first prediction score for face attractiveness prediction according to the fused face image features;
[0129] A second prediction score determination unit, configured to obtain a second prediction score for face attractiveness prediction according to the image encoding vector and multiple text encoding vectors before optimization;
[0130] A third prediction score determination unit, configured to obtain a third prediction score for face attractiveness prediction according to the image encoding vector and the optimized multiple text encoding vectors;
[0131] A score fusion processing unit, configured to obtain the attractiveness prediction result of the face image by fusing the first prediction score, the second prediction score, and the third prediction score.
[0132] Specifically, the face image determination unit may specifically be configured to:
[0133] Perform image sampling from the live stream; determine whether there is face image content in the captured image, and determine whether the size, clarity, facial pose, and occlusion condition of the face image content meet the conditions. If so, determine the captured image as the face image to be predicted.
[0134] Specifically, when sampling images from a live stream, multiple samplings can be performed on the live stream of the same live room at multiple different time points to obtain multiple face images to be predicted, so as to evaluate the change of the face image attractiveness of the live room according to the prediction results of the attractiveness of multiple face images, or evaluate the image processing effects of the facial beautification algorithm and / or image enhancement algorithm used in the live room.
[0135] In addition, the attractiveness of face images corresponding to multiple live rooms in the live system can also be evaluated to evaluate the overall attractiveness of face images in the live system.
[0136] In addition, an embodiment of the present application also provides a computer-readable storage medium, on which a computer program is stored, and when the program is executed by a processor, the steps of the method described in any one of the foregoing method embodiments are implemented.
[0137] And an electronic device, including:
[0138] One or more processors; and
[0139] A memory associated with the one or more processors, the memory is used to store program instructions, and when the program instructions are read and executed by the one or more processors, the steps of the method described in any one of the foregoing method embodiments are executed.
[0140] A computer program product, including a computer program / computer executable instructions, and when the computer program / computer executable instructions are executed by a processor in an electronic device, the steps of the method described in the foregoing method embodiments are implemented.
[0141] Wherein, Figure 4 An exemplary architecture of the electronic device is shown, which may specifically include a processor 410, a video display adapter 411, a disk drive 412, an input / output interface 413, a network interface 414, and a memory 420. The above-mentioned processor 410, video display adapter 411, disk drive 412, input / output interface 413, network interface 414, and the memory 420 can be communicatively connected through a communication bus 430.
[0142] Wherein, the processor 410 can be implemented in a general-purpose CPU (Central Processing Unit, processor), a microprocessor, an application-specific integrated circuit (Application Specific Integrated Circuit, ASIC), or one or more integrated circuits, etc., and is used to execute relevant programs to implement the technical solution provided by the present application.
[0143] The memory 420 can be implemented in the form of ROM (Read Only Memory), RAM (Random Access Memory), static storage devices, dynamic storage devices, etc. The memory 420 can store an operating system 421 for controlling the operation of the electronic device 400, and a Basic Input / Output System (BIOS) for controlling the low-level operations of the electronic device 400. Additionally, it can also store a web browser 423, a data storage management system 424, an attraction prediction system 425, and so on. The above-mentioned attraction prediction system 425 can be the application program that specifically implements the operations of the foregoing steps in the embodiments of the present application. In summary, when implementing the technical solution provided by the present application through software or firmware, the relevant program codes are stored in the memory 420 and are called and executed by the processor 410.
[0144] The input / output interface 413 is used to connect to the input / output module to achieve information input and output. The input / output module can be configured as a component in the device (not shown in the figure) or externally connected to the device to provide corresponding functions. Among them, the input devices can include a keyboard, a mouse, a touch screen, a microphone, various sensors, etc., and the output devices can include a display, a speaker, a vibrator, an indicator light, etc.
[0145] The network interface 414 is used to connect to a communication module (not shown in the figure) to achieve communication and interaction between this device and other devices. Among them, the communication module can achieve communication through wired means (such as USB, network cable, etc.) or through wireless means (such as mobile network, WIFI, Bluetooth, etc.).
[0146] The bus 430 includes a path for transmitting information between various components of the device (such as the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, and the memory 420).
[0147] It should be noted that although the above device only shows the processor 410, the video display adapter 411, the disk drive 412, the input / output interface 413, the network interface 414, the memory 420, the bus 430, etc., in the specific implementation process, this device may also include other components necessary for normal operation. In addition, those skilled in the art can understand that the above device may also only include the components necessary to implement the solution of the present application and do not necessarily include all the components shown in the figure.
[0148] From the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software plus a necessary general hardware platform. Based on such an understanding, the technical solution of the present application, in essence or the part that makes a contribution to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in various embodiments or some parts of the embodiments of the present application.
[0149] Each embodiment in this specification is described in a progressive manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system or system embodiment, since it is basically similar to the method embodiment, the description is relatively simple. For the relevant parts, reference can be made to the partial description of the method embodiment. The systems and system embodiments described above are only illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of this embodiment. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0150] The above has introduced in detail the face image attractiveness prediction method and electronic device provided by the present application. Specific examples are used in this article to elaborate on the principle and implementation manner of the present application. The description of the above embodiments is only used to help understand the method and its core idea of the present application; at the same time, for those of ordinary skill in the art, there will be changes in the specific implementation manner and application scope according to the idea of the present application. In summary, the content of this specification should not be construed as a limitation to the present application.
Claims
1. A method for predicting attractiveness of a facial image, characterized in that: include: Determine a face image to be predicted; In the process of acquiring the facial image features through the deep learning network, a priori facial features based on the face ID provided by the face recognition network are introduced to obtain fused facial image features; Constructing a plurality of text contents by dividing the plurality of attractiveness levels, and encoding the plurality of text contents respectively by using a text encoder to obtain a plurality of text encoding vectors, wherein the plurality of text contents are respectively used to perform a text description of the attractiveness of the face image according to the plurality of attractiveness levels; Using the fused facial image features to perform cross-modal optimization processing on the text encoding vector to obtain multiple optimized text encoding vectors; Using an image encoder to perform image encoding on the face image to obtain an image encoding vector; By respectively calculating the similarities between the image encoding vector and the optimized multiple text encoding vectors, the probability that the facial image belongs to the multiple attractiveness levels is output to determine the attractiveness prediction result of the facial image.
2. The method according to claim 1, characterized in that The facial image features obtained through the deep learning network are multi-scale features, and the multi-scale features are generated by processing features of different receptive fields output at different stages of the deep learning network.
3. The method according to claim 1, characterized in that The face prior features based on face ID provided by the pre-trained face recognition network are used to express the features that distinguish the face image from other images.
4. The method according to claim 1, characterized in that: The text encoder and the image encoder are obtained after parameter fine-tuning training based on a multimodal pre-trained neural network. The multimodal pre-trained neural network is pre-trained in a multimodal data environment, and includes an image encoder for processing images and a text encoder for processing text.
5. The method according to claim 1, characterized in that: Also includes: Acquire a first prediction score for facial attractiveness prediction based on the fused facial image features; Obtaining a second prediction score for facial attractiveness prediction based on the image coding vector and a plurality of text coding vectors before optimization; Obtaining a third prediction score for facial attractiveness prediction according to the image coding vector and the optimized multiple text coding vectors; By fusing the first prediction score, the second prediction score and the third prediction score, an attractiveness prediction result of the facial image is obtained.
6. The method according to claim 1, characterized in that The face image to be predicted includes: Sampling images from live streams; Determine whether there is facial image content in the collected image, and determine whether the size, clarity, facial posture, and occlusion of the facial image content meet the conditions. If so, determine the collected image as the facial image to be predicted.
7. The method according to claim 6, characterized in that The image sampling from the live stream includes: By sampling the live stream of the same live broadcast room multiple times at multiple different time points, multiple facial images to be predicted are obtained, so as to evaluate the changes in the attractiveness of the facial images in the live broadcast room based on the attractiveness prediction results of the multiple facial images, or to evaluate the image processing effects of the facial beautification algorithm and / or image enhancement algorithm used in the live broadcast room.
8. The method according to claim 7, characterized in that By evaluating the attractiveness of facial images corresponding to multiple live broadcast rooms in the live broadcast system, the attractiveness of facial images of the entire live broadcast system can be evaluated.
9. A method for training a facial image attractiveness prediction model, characterized in that: The prediction model includes a face prior feature introduction module, a multimodal encoding module and a cross-modal fusion module; the face prior feature introduction module includes a deep learning network and a pre-trained face recognition network, so as to introduce face prior features based on face ID into the face image extracted by the deep learning network; The multimodal encoding module includes a text encoder and an image encoder, and the text encoder is used to perform text encoding on a plurality of text contents constructed according to the divided plurality of attractiveness levels; the method includes: Acquire a training data set, wherein the training data set includes a plurality of face images and corresponding attractiveness annotation information; In the first training stage, the face prior feature introduction module and the multimodal encoding module are trained in parallel; wherein, with respect to the face prior feature introduction module, the deep learning network is trained using the training data while the parameters of the pre-trained face recognition network are fixed; In the second training stage, the cross-modal fusion module is trained while the parameters of the face prior feature introduction module and the multimodal encoding module are fixed; Among them, in the second training stage, by inputting the face image in the training data into the face prior feature introduction module trained in the first stage, the face image feature that integrates the face prior feature based on the face ID is obtained, the face image is input into the image encoder trained in the first stage to obtain an image coding vector, the multiple text contents are input into the text encoder trained in the first stage to obtain multiple text vectors, and the multiple text vectors are optimized by the cross-modal fusion module using the face image feature that integrates the face prior feature based on the face ID to obtain multiple optimized text coding vectors; by respectively calculating the similarity between the image coding vector and the optimized multiple text coding vectors as the probability that the face image belongs to the multiple attractiveness levels respectively, the probabilities of the multiple attractiveness levels and the level scores corresponding to the multiple attractiveness levels are weightedly summed to obtain the attractiveness prediction score of the face image, so as to optimize the parameters of the cross-modal fusion module using a preset loss function to complete the training of the cross-modal fusion module.
10. The method according to claim 9, characterized in that The prediction model also includes a decision fusion module; When training the face prior feature introduction module in the first training stage, facial image features of the face images in the training data are obtained through a deep learning network, and facial prior features of the face images based on face ID are obtained through a pre-trained face recognition network, and fused face image features are obtained by fusing the face image features with the face prior features, and a first prediction score corresponding to the face image is obtained according to the fused face image features, and the parameters of the deep learning network are optimized using a preset loss function to complete the training of the face prior feature introduction module; When the multimodal encoding module is trained in the first training stage, the face image is encoded by using an image encoder to obtain an image encoding vector; multiple text contents are constructed by dividing them into multiple attractiveness levels, and the multiple text contents are respectively encoded by using a text encoder to obtain multiple text encoding vectors; the similarities between the image encoding vector and the multiple text encoding vectors are respectively calculated to obtain the probability that the face image belongs to the multiple attractiveness levels; and the probabilities of the multiple attractiveness levels are weightedly summed with the level scores corresponding to the multiple attractiveness levels to obtain a second prediction score, and the parameters of the multimodal encoding module are optimized using a preset loss function to complete the training of the multimodal encoding module; In the second training stage, the attractiveness prediction score output by the cross-modal fusion module is used as the third prediction score, and the first prediction score, the second prediction score and the third prediction score are fused through the decision fusion module to obtain an attractiveness prediction result, and the parameters of the cross-modal fusion module and the decision fusion module are optimized using a preset loss function to complete the training of the cross-modal fusion module and the decision fusion module.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the steps of the method described in any one of claims 1 to 8 are implemented.
12. An electronic device, characterized in that: include: one or more processors; as well as A memory associated with the one or more processors, the memory being used to store program instructions, wherein the program instructions, when read and executed by the one or more processors, execute the steps of the method described in any one of claims 1 to 8.
13. A computer program product comprising a computer program / computer executable instructions, characterized in that: When the computer program / computer executable instructions are executed by a processor in an electronic device, the steps of the method according to any one of claims 1 to 8 are implemented.