Face quality evaluation model training method and face image quality evaluation method
By using a multi-task quality evaluation model, combining face key points, occlusion and angle tags in the fusion dataset, the problems of slow face quality evaluation and high memory consumption in the existing technology are solved, and fast and accurate facial image quality evaluation is achieved.
Patent Information
- Application Number
- CN202311478719.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-07
- Publication Date
- 2025-05-09
AI Technical Summary
The existing face quality evaluation methods require multiple models and multi-stage detection, resulting in slow speed, large memory consumption, and large calculations, making it impossible to effectively evaluate the multi-faceted information of face images.
By obtaining the fusion data set, including multiple face sample images, each image has face key point label, face occlusion label and face angle label, and is trained using the multi-task quality evaluation model. The model includes face key point branches, face occlusion branches and face angle branches, which can output these information at the same time.
It realizes fast and effective face quality evaluation based on a single model, reduces device memory consumption and calculation volume, and improves the accuracy and speed of face image quality evaluation.
Smart Images

Figure CN119963936A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of image processing, and more specifically to a training method for a face quality assessment model and a quality assessment method for face images. Background Art
[0002] Currently, face recognition technology is being used more and more widely. In the face recognition process, the person being recognized often has various natural movements from entering the screen to the completion of recognition (such as lowering the head, side face, motion blur, etc.), or the face is blocked (such as wearing sunglasses, hats, masks, etc.), or the face image quality is too low due to problems such as low face resolution and unqualified face aspect ratio, which is not suitable for face comparison and recognition. Using low-quality face images for recognition can easily lead to misidentification or non-recognition by the face recognition system. Therefore, it is crucial for the face recognition system to filter out unqualified face images from the video frame sequence and retain qualified face images.
[0003] The machine learning models currently used to evaluate the quality of facial images are all targeted at specific tasks and cannot output multi-faceted information based on a single model. Therefore, current facial quality evaluation methods often adopt a multi-stage detection method and require the configuration of multiple models, which reduces the quality evaluation speed and requires larger device memory consumption and computing power. Summary of the invention
[0004] A series of simplified concepts are introduced in the Summary of the Invention, which will be further described in detail in the Detailed Description of the Invention. The Summary of the Invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the scope of protection of the claimed technical solution.
[0005] According to one aspect of the present invention, a method for training a face quality assessment model is provided, the method comprising:
[0006] Acquire a fused data set, wherein the fused data set includes a plurality of face sample images, each of the face sample images has a face key point label, a face occlusion label, and a face angle label, and at least one of the face key point label, the face occlusion label, and the face angle label is a true label;
[0007] The multi-task quality assessment model to be trained is trained according to the fusion data set to obtain a trained multi-task quality assessment model, wherein the multi-task quality assessment model includes a face key point branch, a face occlusion branch and a face angle branch, wherein the face key point branch is used to output face key point information, the face occlusion branch is used to output face occlusion information, and the face angle branch is used to output face angle information.
[0008] In one embodiment, at least one of the face key point label, the face occlusion label and the face angle label is a default label;
[0009] Each of the face sample images also has a face key point loss weight corresponding to the face key point label, a face occlusion loss weight corresponding to the face occlusion label, and a face angle loss weight corresponding to the face angle label;
[0010] Among the face key point loss weights, the face occlusion loss weights and the face angle loss weights, the value of the loss weight corresponding to the default label is 0, and the value of the loss weight corresponding to the true label is not 0.
[0011] In one embodiment, the training of the multi-task quality assessment model to be trained according to the fusion data set to obtain a trained multi-task quality assessment model includes:
[0012] Inputting the face sample image into the multi-task quality evaluation model to be trained, and obtaining the face key point prediction results output by the face key point branch of the multi-task quality evaluation model to be trained, the face occlusion prediction results output by the face occlusion branch, and the face angle prediction results output by the face angle branch;
[0013] Obtaining the face key point loss according to the face key point prediction result, the face key point label and the face key point loss weight;
[0014] Obtaining the face occlusion loss according to the face occlusion prediction result, the face occlusion label and the face occlusion loss weight;
[0015] Obtaining a face angle loss according to the face angle prediction result, the face angle label and the face angle loss weight;
[0016] The face key point loss, the face occlusion loss and the face angle loss are summed to obtain a summary loss, and the model parameters of the multi-task quality evaluation model are optimized according to the summary loss to obtain a trained multi-task quality evaluation model.
[0017] In one embodiment, the facial key point prediction results, the facial key point labels, and the facial key point loss weights correspond to a plurality of facial key points respectively;
[0018] The step of obtaining the face key point loss according to the face key point prediction result, the face key point label and the face key point loss weight comprises:
[0019] Obtaining sub-face key point losses corresponding to different face key points according to the face key point prediction results and the face key point labels corresponding to different face key points;
[0020] The sub-face key point losses corresponding to multiple face key points are weightedly summed according to the face key point loss weights corresponding to multiple face key points to obtain the face key point loss.
[0021] In one embodiment, the face occlusion prediction result, the face occlusion label and the face occlusion loss weight correspond to a plurality of face parts respectively;
[0022] The obtaining of the face occlusion loss according to the face occlusion prediction result, the face occlusion label and the face occlusion loss weight includes:
[0023] Obtaining sub-face occlusion losses corresponding to different face parts according to the face occlusion prediction results and the face occlusion labels corresponding to different face parts;
[0024] The sub-face occlusion losses corresponding to multiple face parts are weightedly summed according to the face occlusion loss weights corresponding to multiple face parts to obtain the face occlusion loss.
[0025] In one embodiment, the face angle prediction result, the face angle label and the face angle loss weight correspond to multiple face angles respectively;
[0026] The obtaining of the face angle loss according to the face angle prediction result, the face angle label and the face angle loss weight comprises:
[0027] Obtaining sub-face angle losses corresponding to different face angles according to the face angle prediction results and the face angle labels corresponding to different face angles;
[0028] The sub-face angle losses corresponding to multiple face angles are weightedly summed according to the face angle loss weights corresponding to multiple face angles to obtain the face angle loss.
[0029] In one embodiment, before inputting the sample face image into the multi-task quality assessment model to be trained, the method further includes preprocessing the sample face image, wherein the preprocessing includes at least one of the following:
[0030] Scaling the face sample image to a preset size;
[0031] Performing random color changes on the face sample image;
[0032] The face sample image is normalized.
[0033] A second aspect of an embodiment of the present invention provides a method for evaluating the quality of a face image, the method comprising:
[0034] Get face image;
[0035] Inputting the face image into a trained multi-task quality assessment model, wherein the multi-task quality assessment model includes a face key point branch, a face occlusion branch, and a face angle branch;
[0036] Obtaining facial key point information output by the facial key point branch, facial occlusion information output by the facial occlusion branch, and facial angle information output by the facial angle branch;
[0037] A quality evaluation result of the face image is obtained based on the face key point information, the face occlusion information and the face angle information.
[0038] In one embodiment, the facial key point information includes coordinates of multiple facial key points.
[0039] In one embodiment, obtaining the quality evaluation result of the face image based on the face key point information, the face occlusion information and the face angle information includes:
[0040] Obtaining facial key point distribution information according to the coordinates of the plurality of facial key points;
[0041] The quality evaluation result is obtained according to the facial key point distribution information.
[0042] In one embodiment, the face occlusion information includes occlusion probabilities of multiple face parts.
[0043] In one embodiment, obtaining the quality evaluation result of the face image based on the face key point information, the face occlusion information and the face angle information includes:
[0044] Comparing the occlusion probabilities of the multiple face parts with the corresponding preset probability thresholds respectively to determine an occluded face part among the multiple face parts, wherein the occlusion probability of the occluded face part is greater than or equal to the corresponding preset probability threshold;
[0045] The quality evaluation result is obtained according to the blocked facial part.
[0046] In one embodiment, the multiple facial parts include the following: left eye, right eye, nose, mouth, left face, right face, forehead.
[0047] In one embodiment, the face angle information includes at least one of the following: a face yaw angle, a face pitch angle, and a face roll angle.
[0048] In one embodiment, obtaining the quality evaluation result of the face image based on the face key point information, the face occlusion information and the face angle information includes:
[0049] The face yaw angle, the face pitch angle and the face roll angle are respectively compared with the corresponding preset angle thresholds. When at least one of the face yaw angle, the face pitch angle and the face roll angle exceeds the corresponding preset angle threshold, it is determined that the face angle does not meet the preset requirements.
[0050] In one embodiment, the multi-task quality assessment model is trained based on a fusion data set, which includes multiple face sample images, each of which has a face key point label, a face occlusion label and a face angle label, and at least one of the face key point label, the face occlusion label and the face angle label is a true label.
[0051] In one embodiment, before inputting the face image into the multi-task quality assessment model, the method further includes preprocessing the face image, wherein the preprocessing includes at least one of the following:
[0052] Scaling the face image to a preset size;
[0053] The face image is normalized.
[0054] In one embodiment, the facial image includes a visible light facial image or an infrared facial image.
[0055] A third aspect of an embodiment of the present invention provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program executed by the processor, and the computer program executes the method described above when executed by the processor.
[0056] A fourth aspect of an embodiment of the present invention provides a computer-readable medium, on which a computer program is stored, and the computer program executes the method as described above when running.
[0057] The multi-task quality evaluation model trained according to the training method of the face quality evaluation model in the embodiment of the present invention can simultaneously output face key point information, face occlusion information and face angle information, and can reduce device memory consumption and calculation amount; the face image quality evaluation method according to the embodiment of the present invention evaluates the quality of the face image based on the multi-task quality evaluation model, and can obtain more accurate and comprehensive face quality evaluation results. BRIEF DESCRIPTION OF THE DRAWINGS
[0058] The above and other purposes, features and advantages of the present invention will become more apparent by describing the embodiments of the present invention in more detail in conjunction with the accompanying drawings. The accompanying drawings are used to provide a further understanding of the embodiments of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings, the same reference numerals generally represent the same components or steps.
[0059] Figure 1 A schematic block diagram showing an example electronic device for implementing a training method for a face quality assessment model and a method for assessing the quality of a face image according to an embodiment of the present invention;
[0060] Figure 2 A schematic flow chart showing a method for training a face quality assessment model according to an embodiment of the present invention;
[0061] Figure 3 A structural diagram of a deep convolutional network according to an embodiment of the present invention is shown;
[0062] Figure 4 A schematic flow chart showing a method for evaluating the quality of a face image according to an embodiment of the present invention;
[0063] Figure 5 A schematic block diagram of an electronic device according to an embodiment of the present invention is shown. DETAILED DESCRIPTION
[0064] In order to make the purpose, technical scheme and advantages of the present invention more obvious, the exemplary embodiments according to the present invention will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments of the present invention, and it should be understood that the present invention is not limited to the exemplary embodiments described herein. Based on the embodiments of the present invention described in the present invention, all other embodiments obtained by those skilled in the art without creative work should fall within the protection scope of the present invention.
[0065] First, refer to Figure 1 An example electronic device 100 for implementing the training method of the face quality assessment model and the quality assessment method of the face image according to the embodiment of the present invention is described.
[0066] like Figure 1 As shown, the electronic device 100 includes one or more processors 102, one or more storage devices 104, an input device 106, an output device 108, and an image sensor 110, and these components are interconnected through a bus system 112 and / or other forms of connection mechanisms (not shown). It should be noted that Figure 1 The components and structures of the electronic device 100 shown are merely exemplary and non-limiting. The electronic device may also have other components and structures as required.
[0067] The processor 102 may be a central processing unit (CPU) or other forms of processing units having data processing capabilities and / or instruction execution capabilities, and may control other components in the electronic device 100 to perform desired functions.
[0068] The storage device 104 may include one or more computer program products, and the computer program product may include various forms of computer-readable storage media, such as volatile memory and / or non-volatile memory. The volatile memory may include, for example, random access memory (RAM) and / or cache memory (cache), etc. The non-volatile memory may include, for example, read-only memory (ROM), hard disk, flash memory, etc. One or more computer program instructions may be stored on the computer-readable storage medium, and the processor 102 may run the program instructions to implement the client functions (implemented by the processor) and / or other desired functions in the embodiments of the present invention described below. Various applications and various data, such as various data used and / or generated by the application, may also be stored in the computer-readable storage medium.
[0069] The input device 106 may be a device used by a user to input instructions, and may include one or more of a keyboard, a mouse, a microphone, a touch screen, and the like.
[0070] The output device 108 may output various information (eg, images or sounds) to the outside (eg, a user), and may include one or more of a display, a speaker, and the like.
[0071] The image sensor 110 can capture images desired by a user (eg, photos, videos, etc.), and store the captured images in the storage device 104 for use by other components.
[0072] Exemplarily, an example electronic device for implementing the training method of a face quality assessment model and the quality assessment method of a face image according to an embodiment of the present invention may be implemented as a smartphone, a tablet computer, or the like.
[0073] Next, we will refer to Figure 2 The training method 200 of the face quality assessment model according to an embodiment of the present invention is described. Figure 2 As shown, the training method 200 of the face quality assessment model according to the embodiment of the present invention includes the following steps:
[0074] In step S210, a fused data set is obtained, wherein the fused data set includes a plurality of face sample images, each of the face sample images has a face key point label, a face occlusion label, and a face angle label, and at least one of the face key point label, the face occlusion label, and the face angle label is a true label;
[0075] In step S220, the multi-task quality assessment model to be trained is trained according to the fusion data set to obtain a trained multi-task quality assessment model, wherein the multi-task quality assessment model includes a face key point branch, a face occlusion branch and a face angle branch, wherein the face key point branch is used to output face key point information, the face occlusion branch is used to output face occlusion information, and the face angle branch is used to output face angle information.
[0076] The training method 200 of the face quality assessment model of the embodiment of the present invention is mainly used to train a multi-task quality assessment model that can simultaneously output face key point information, face occlusion information, and face angle information, so that accurate and comprehensive face quality assessment results can be obtained based on the multi-task quality assessment model. In the model training process, it is necessary to construct training samples and design loss functions, and the existing training data is often for a single-task model, lacking training data with face key point labels, face occlusion labels, and face angle labels. And labeling new training data with face key point labels, face occlusion labels, and face angle labels requires a huge workload.
[0077] In order to obtain enough training data and reduce the workload required for annotation, when constructing a fused dataset, the face key point dataset, face occlusion dataset and face angle dataset can be obtained respectively, and the face sample images in the above three datasets can be integrated together to obtain a fused dataset. Among them, the face sample images in the face key point dataset have real face key point labels, the face sample images in the face occlusion dataset have real face occlusion labels, and the face sample images in the face angle dataset have real face angle labels. After integrating the above face sample images into a fused dataset, each face sample image in the fused dataset has a face key point label, a face occlusion label and a face angle label, at least one of which is a real label, and the rest of the labels are default labels.
[0078] For example, if a face sample image comes from a face occlusion dataset, the face occlusion label of the face sample image is a true label, which can represent the true face occlusion information therein, while the face key point label and the face angle label are default labels, which are only default values, for example, the face key point label and the face angle label are default values 0. In some embodiments, each face sample image in the fused dataset may also have two true labels and one default label.
[0079] In addition to the face key point labels, face occlusion labels and face angle labels, when constructing the fusion dataset, it is also necessary to set the loss weight for each face sample image. Specifically, each face sample image has a face key point loss weight corresponding to the face key point label, a face occlusion loss weight corresponding to the face occlusion label and a face angle loss weight corresponding to the face angle label. Since the default label cannot represent the true information of the face sample image, the value of the loss weight corresponding to the default label is 0 in the face key point loss weight, face occlusion loss weight and face angle loss weight to avoid the negative impact of the default label on the model training process. On the contrary, the real label can represent the real information of the face sample image, so the value of the loss weight corresponding to the real label is not 0, for example, the value of the loss weight corresponding to the real label can be set to 1. For example, if the face occlusion label is the real label, the face key point label and the face angle label are the default labels, then the value of the face occlusion loss weight is not 0, and the values of the face key point loss weight and the face angle loss weight are 0.
[0080] After arranging the three different data sets, each face sample image has face key point labels, face key point loss weights, face occlusion labels, face occlusion loss weights, face angle labels, and face angle loss weights. In this way, the face key point dataset, face occlusion dataset, and face angle dataset can be fused to obtain a fused dataset, and the model training is performed based on the face sample images in the fused dataset. When adding a face sample image, it is only necessary to ensure that there is a real label in the face sample image. For example, it is only necessary to mark the real face occlusion label in the face sample image without marking the face angle label and face key point label.
[0081] Exemplarily, in order to enable the trained multi-task quality assessment model to output facial key point prediction results corresponding to multiple facial key points, the facial sample data may include facial key point labels and facial key point loss weights corresponding to multiple facial key points. For example, the facial key point labels may include the horizontal coordinates and vertical coordinates of facial key points such as left eye, right eye, nose, left corner of mouth, right corner of mouth in the facial sample image. Among them, the facial key points may be 5 points on the face, 68 points on the face, 106 points on the face, etc.
[0082] In order to enable the trained multi-task quality assessment model to output comprehensive face occlusion prediction results corresponding to multiple face parts, the face sample data may include face occlusion labels and face occlusion loss weights corresponding to multiple face parts. For example, face occlusion labels may include left eye occlusion, right eye occlusion, nose occlusion, mouth occlusion, left face occlusion, right face occlusion, forehead occlusion, etc.
[0083] In order to enable the trained multi-task quality evaluation model to output face angle prediction results corresponding to face angles in multiple dimensions, the face sample data may include face angle labels and face angle loss weights corresponding to multiple face angles. The multiple face angles may include face yaw angle (Yaw), face pitch angle (Pitch) and face roll angle (Roll).
[0084] Exemplarily, before training the multi-task quality evaluation model based on face sample images, the face sample images may also be preprocessed, and the preprocessing includes at least one of the following: scaling the face sample images to a preset size; performing random color changes on the face sample images; and normalizing the face sample images.
[0085] The size of the preset size can be arbitrarily designed. For example, the size of the face sample image can be uniformly scaled to 96×96 pixels. The random color change of the face sample image includes, for example, modifying the brightness and contrast of the face sample image, converting the color image to a grayscale image, etc. Random color change can improve the generalization ability of the multi-task quality evaluation model.
[0086] Normalization is to scale the data of the face sample image to a preset range. Before data scaling, the face sample image can also be subtracted from the mean to remove the statistical mean of the corresponding dimension of the data to eliminate the common part and highlight the characteristics and differences between individuals. The mean and scaling factor (scaling factor is non-zero) can be arbitrarily set. For example, the face sample image data can be normalized to the range of [-1.0, 1.0].
[0087] In some embodiments, the face sample images included in the fused data set include infrared face images and visible light face images (e.g., RGB images), so that the trained multi-task quality assessment model can perform quality assessment on both visible light face images and infrared face images. In another embodiment, the face sample images included in the fused data set may also include only visible light face images, and by performing random color changes on the face sample images, the trained multi-task quality assessment model may also be applicable to infrared face images.
[0088] After the fusion data set is constructed, the multi-task quality evaluation model to be trained can be trained according to the fusion data set to obtain a trained multi-task quality evaluation model. The multi-task quality evaluation model of the embodiment of the present invention includes a face key point branch, a face occlusion branch and a face angle branch, wherein the face key point branch is used to output face key point information, the face occlusion branch is used to output face occlusion information, and the face angle branch is used to output face angle information. After the training is completed, face key point information, face occlusion information and face angle information can be obtained simultaneously based on a single multi-task quality evaluation model, which can effectively reduce device memory consumption and calculation amount, and improve the quality evaluation speed of face images.
[0089] Specifically, a face sample image is input into the multi-task quality evaluation model to be trained, and the face key point prediction results output by the face key point branch of the multi-task quality evaluation model to be trained, the face occlusion prediction results output by the face occlusion branch, and the face angle prediction results output by the face angle branch are obtained. The face key point loss can be obtained according to the face key point prediction results, face key point labels, and face key point loss weights; the face occlusion loss can be obtained according to the face occlusion prediction results, face occlusion labels, and face occlusion loss weights; the face angle loss can be obtained according to the face angle prediction results, face angle labels, and face angle loss weights.
[0090] Next, the face key point loss, face occlusion loss and face angle loss are summed to obtain the summary loss, and the model parameters of the multi-task quality evaluation model are optimized according to the summary loss to obtain a trained multi-task quality evaluation model. Among them, if the face key point label of the face sample image is the true label, the face key point loss weight is not 0, so the face key point loss is also not 0; conversely, if the face key point label of the face sample image is the default label instead of the true label, the face key point loss weight is 0, so the face key point loss is also 0. Similarly, if the face angle label is the true label, the face angle loss is not 0, otherwise the face angle loss is 0; if the face occlusion label is the true label, the face occlusion loss is not 0, otherwise the face occlusion loss is 0. In other words, if a label of the face sample image is the default label instead of the true label, the loss value corresponding to the label is 0, which will not affect the model parameters.
[0091] Exemplarily, the face key point prediction results, face key point labels and face key point loss weights correspond to multiple face key points respectively. When calculating the face key point loss, the sub-face key point losses corresponding to different face key points can be obtained according to the face key point prediction results and face key point labels corresponding to different face key points, and the sub-face key point losses corresponding to the multiple face key points are weighted and summed according to the face key point loss weights corresponding to the multiple face key points to obtain the total face key point loss.
[0092] For example, assuming that there are Np facial key points in total, then obtain Np facial key point labels and Np facial key point loss weights, and obtain the multi-task quality evaluation model output Np facial key point prediction results, and the facial key point prediction results and facial key point labels may include the coordinates of the facial key points. According to the loss function, the difference between the facial key point label and the facial key point prediction result of each facial key point is calculated to obtain the sub-facial key point loss corresponding to each facial key point. The loss function includes but is not limited to any one of L1 Loss, L2 Loss, and smooth L1 Loss. According to the facial key point loss weights corresponding to the Np facial key points, the sub-facial key point loss results are weighted and summed to obtain the total facial key point loss.
[0093] Exemplarily, the face occlusion prediction results, face occlusion labels and face occlusion loss weights correspond to multiple face parts respectively; when calculating the face occlusion loss, the sub-face occlusion losses corresponding to different face parts can be obtained according to the face occlusion prediction results and face occlusion labels corresponding to different face parts; the sub-face occlusion losses corresponding to the multiple face parts are weightedly summed according to the face occlusion loss weights corresponding to the multiple face parts to obtain the total face occlusion loss.
[0094] For example, assuming that there are Nm face parts in total, then Nm face occlusion labels and Nm face occlusion loss weights are obtained, and the multi-task quality evaluation model outputs Nm face occlusion prediction results. The face occlusion prediction result may include the occlusion probability of the face part, or the occlusion confidence of the face part; the face occlusion label may include whether the face part is occluded. When the face part is occluded, its occlusion probability is 1, and when the face part is not occluded, its occlusion probability is 0. The difference between the face occlusion label and the face occlusion prediction result of each face occlusion is calculated according to the loss function to obtain the sub-face occlusion loss corresponding to each face occlusion. The loss function includes but is not limited to any one of BCE Loss, Focal Loss, and softmax Loss. According to the face occlusion loss weights corresponding to the Nm face occlusions, the sub-face occlusion loss results are weighted summed to obtain the total face occlusion loss.
[0095] Exemplarily, the face angle prediction results, face angle labels and face angle loss weights correspond to multiple face angles respectively; when calculating the face angle loss, the sub-face angle losses corresponding to different face angles can be obtained according to the face angle prediction results and face angle labels corresponding to different face angles; according to the face angle loss weights corresponding to the multiple face angles, the sub-face angle losses corresponding to the multiple face angles are weightedly summed to obtain the face angle loss.
[0096] For example, assuming that there are Na face angles in total, then obtain Na face angle labels and Na face angle loss weights, and obtain Na face angle prediction results output by the multi-task quality evaluation model. Calculate the difference between the face angle label and the face angle prediction result of each face angle according to the loss function to obtain the sub-face angle loss corresponding to each face angle. The loss function includes but is not limited to any one of L1 Loss, L2 Loss, and smooth L1 Loss. According to the face angle loss weight corresponding to Na face angle, perform weighted summation on the sub-face angle loss results to obtain the total face angle loss.
[0097] Finally, the face occlusion loss, face angle loss and face key point loss are summed to obtain the summary loss, and the summary loss is used to calculate the gradient and update the model parameters. The above steps are repeated until the summary loss converges, and a trained multi-task quality evaluation model is obtained.
[0098] The multi-task quality evaluation model of the embodiment of the present invention adopts a deep convolutional network structure, which can reduce the amount of calculation while meeting the prediction accuracy, and facilitate the deployment of terminal devices. The deep convolutional network mainly includes an input layer (Input layer), a convolution layer (CONV layer), a pooling layer (Pooling layer) and a fully connected layer (FC layer).
[0099] Figure 3An exemplary deep convolutional network structure is shown. The input data is normalized four-channel image data. 3×3, S=1 represents 3×3 convolution, stride 1, and ReLU activation function; MP3×3, S=2 represents 3×3 maximum pooling, stride 2, and GP represents global pooling operation; output1 is the output of the face occlusion branch, specifically 1×1×7 values, and the sigmoid function is used to limit the output to the [0,1] interval, corresponding to the occlusion probability values of the seven parts of the face; output2 is the output of the face angle branch, specifically 1×1×3 values, limited to the [-pi / 2,pi / 2] interval, corresponding to the face yaw angle (Yaw), face pitch angle (Pitch) and face roll angle (Roll); output3 is the output of the face key point branch, specifically 1×1×10 values, and the sigmoid function is used to limit the output to the [0,1] interval, corresponding to the ratio values of the horizontal and vertical coordinates of the five key points of the face (left eye, right eye, nose, left corner of the mouth, right corner of the mouth) in the face image.
[0100] Based on the above description, the training method 200 of the face quality assessment model of the embodiment of the present invention trains a multi-task quality assessment model that can simultaneously output face key point information, face occlusion information and face angle information, and has the characteristics of high efficiency, strong applicability and high accuracy, and is suitable for deployment in terminal devices with limited computing.
[0101] The above exemplary steps of the training method of the face quality assessment model according to the embodiment of the present invention are described in detail. Figure 4 A method 400 for evaluating the quality of a face image according to an embodiment of the present invention is described. Figure 4 As shown, the method comprises the following steps:
[0102] In step S410, a face image is acquired;
[0103] In step S420, the face image is input into a trained multi-task quality assessment model, where the multi-task quality assessment model includes a face key point branch, a face occlusion branch, and a face angle branch;
[0104] In step S430, the facial key point information output by the facial key point branch, the facial occlusion information output by the facial occlusion branch, and the facial angle information output by the facial angle branch are obtained;
[0105] In step S440, a quality evaluation result of the face image is obtained based on the face key point information, the face occlusion information and the face angle information.
[0106] Exemplarily, before inputting the face image into the multi-task quality assessment model, the face image is first preprocessed, and the preprocessing includes at least one of the following: scaling the face image to a preset size; normalizing the face image. The preset size can be arbitrarily designed, for example, scaling the width and height to 96×96 pixels respectively. Normalization includes mean subtraction and data scaling, wherein the mean and scaling factor (scaling factor is a non-zero value) can be arbitrarily taken, for example, the data of the face image can be normalized to the interval [-1.0, 1.0].
[0107] After preprocessing is completed, the multi-task quality evaluation model can be input into the trained multi-task quality evaluation model. The multi-task quality evaluation model of the embodiment of the present invention can simultaneously predict facial key point information, facial angle information and facial occlusion information based on a single model. There is no need to configure multiple models, which can effectively reduce device memory consumption and calculation amount, and improve the speed of face quality evaluation, thereby improving the speed of face recognition.
[0108] Exemplarily, the facial key point information includes the coordinates of Np facial key points. The coordinates of multiple facial key points may include the horizontal coordinates and vertical coordinates of the left eye, right eye, nose, left corner of mouth, right corner of mouth, etc. in the face image. Among them, the facial key points can be 5 points on the face, 68 points on the face, 106 points on the face, etc., which can be configured according to actual needs.
[0109] After obtaining the coordinates of multiple facial key points, facial key point distribution information can be obtained based on the coordinates of the multiple facial key points, and a quality evaluation result can be obtained based on the facial key point distribution information. For example, facial key point distribution information such as left and right pupil distance and face width-to-height ratio can be obtained based on the coordinates of five facial key points, and the facial key point distribution information can be compared with a preset threshold to determine whether the facial key point distribution information meets the facial image quality requirements.
[0110] Exemplarily, the face occlusion information includes the occlusion probabilities of Nm face parts. Exemplarily, the multiple face parts include the following: left eye, right eye, nose, mouth, left face, right face, forehead; the face occlusion information includes the left eye occlusion probability, the right eye occlusion probability, the nose occlusion probability, the mouth occlusion probability, the left face occlusion probability, the right face occlusion probability, the forehead occlusion probability, etc. By obtaining the occlusion probabilities of multiple face parts, it can be applied to usage scenarios with different levels of demand for face occlusion.
[0111] After obtaining the occlusion probabilities of Nm face parts, the occlusion probabilities of the Nm face parts can be compared with the corresponding preset probability thresholds to determine the occluded face parts among the multiple face parts, wherein the occlusion probability of the occluded face parts is greater than or equal to the corresponding preset probability threshold. If the occlusion probability of the kth (k∈[0,...,Nm-1]) face part is greater than the kth threshold, it is determined that the corresponding face part is occluded.
[0112] For example, if the probability of the left eye being blocked is greater than the first preset probability threshold, it means that the left eye is blocked; if the probability of the mouth being blocked is less than the second preset probability threshold, it means that the mouth is not blocked. The preset probability thresholds corresponding to different facial parts are the same or different. Furthermore, relevant quality evaluation results can be obtained according to the blocked facial parts. For example, if the left and right eyes are blocked parts, it can be judged that the target is wearing sunglasses; if the forehead is blocked parts, it can be judged that the target is wearing a hat; if the nose and mouth are blocked parts, it can be judged that the target is wearing a mask, etc.
[0113] Exemplarily, the face angle information includes Na face angle information, which may specifically include at least one of the following: face yaw angle (Yaw), face pitch angle (Pitch), and face roll angle (Roll). After obtaining the face yaw angle, face pitch angle, and face roll angle, the face yaw angle, face pitch angle, and face roll angle are respectively compared with corresponding preset angle thresholds. Exemplarily, when at least one of the face yaw angle, face pitch angle, and face roll angle exceeds the corresponding preset angle threshold, it can be determined that the face angle does not meet the preset requirement.
[0114] The multi-task quality evaluation model of the embodiment of the present invention is obtained by training based on a fusion data set, which includes multiple face sample images, each face sample image has a face key point label, a face occlusion label and a face angle label, and at least one of the face key point label, the face occlusion label and the face angle label is a real label; further, at least one of the face key point label, the face occlusion label and the face angle label is a default label, and the loss weight corresponding to the default label is 0, thereby reducing the workload required for data annotation. During the model training process, the face sample images can be randomly changed in color, or visible light images and infrared images can be used for training, so that the face images input to the multi-task quality evaluation model can include both visible light face images and infrared face images. The training method of the multi-task quality evaluation model can be specifically referred to above and will not be repeated here.
[0115] To sum up, the facial image quality assessment method 400 of the embodiment of the present invention evaluates the quality of facial images based on a multi-task quality assessment model, and can obtain facial key point information, facial occlusion information and facial angle information based on a single model, thereby obtaining more accurate and comprehensive facial quality assessment results. At the same time, it can reduce device memory consumption and calculation amount, and improve the quality assessment speed of facial images.
[0116] Exemplarily, the training method of the face quality assessment model and the quality assessment method of the face image according to the embodiments of the present invention can be implemented in a device, apparatus or system having a memory and a processor.
[0117] Exemplarily, the training method of the face quality assessment model and the quality assessment method of the face image according to the embodiment of the present invention can be deployed at a personal terminal, such as a smart phone, a tablet computer, a personal computer, etc. Alternatively, the training method of the face quality assessment model and the quality assessment method of the face image according to the embodiment of the present invention can also be deployed on the server side (or the cloud). Alternatively, the training method of the face quality assessment model and the quality assessment method of the face image according to the embodiment of the present invention can be distributedly deployed on the server side (or the cloud) and the personal terminal, and can also be distributedly deployed at different personal terminals.
[0118] Below, refer to Figure 5 An electronic device according to an embodiment of the present invention is described. Figure 5 A schematic block diagram of an electronic device 500 according to an embodiment of the present invention is shown. The electronic device 500 according to the embodiment of the present invention includes a memory 510 and a processor 520. Among them, the memory 510 stores program codes for implementing the corresponding steps in the training method of the face quality assessment model or the quality assessment method of the face image according to the embodiment of the present invention. The processor 520 is used to run the program code stored in the memory 510 to execute the corresponding steps of the training method of the face quality assessment model or the quality assessment method of the face image according to the embodiment of the present invention. The training method of the face quality assessment model and the quality assessment method of the face image according to the embodiment of the present invention can be specifically referred to above and will not be repeated here.
[0119] In addition, according to an embodiment of the present application, a storage medium is also provided, on which program instructions are stored, and when the program instructions are executed by a computer or a processor, the corresponding steps of the training method of the face quality assessment model or the quality assessment method of the face image of the embodiment of the present application are executed. The storage medium may include, for example, a memory card of a smart phone, a storage component of a tablet computer, a hard disk of a personal computer, a read-only memory (ROM), an erasable programmable read-only memory (EPROM), a portable compact disk read-only memory (CD-ROM), a USB memory, or any combination of the above storage media. The computer-readable storage medium may be any combination of one or more computer-readable storage media.
[0120] According to an embodiment of the present application, a computer program product is also provided, including a computer program / instruction, which, when executed by a processor, implements the above-mentioned training method of the face quality assessment model or the quality assessment method of the face image.
[0121] The electronic device, computer-readable medium and computer program product according to the embodiments of the present invention are used to implement the training method of the face quality assessment model or the quality assessment method of the face image according to the embodiments of the present invention, and therefore also have similar advantages.
[0122] Although example embodiments have been described herein with reference to the accompanying drawings, it should be understood that the above example embodiments are merely exemplary and are not intended to limit the scope of the present invention thereto. Various changes and modifications may be made therein by one of ordinary skill in the art without departing from the scope and spirit of the present invention. All such changes and modifications are intended to be included within the scope of the present invention as required by the appended claims.
[0123] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered to be beyond the scope of the present invention.
[0124] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are only schematic, for example, the division of the units is only a logical function division, and there may be other division methods in actual implementation, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed.
[0125] In the description provided herein, a large number of specific details are described. However, it is understood that embodiments of the present invention can be practiced without these specific details. In some instances, well-known methods, structures and techniques are not shown in detail so as not to obscure the understanding of this description.
[0126] Similarly, it should be understood that in order to streamline the present invention and help understand one or more of the various inventive aspects, in the description of the exemplary embodiments of the present invention, the various features of the present invention are sometimes grouped together into a single embodiment, figure, or description thereof. However, the method of the present invention should not be interpreted as reflecting the following intention: the claimed invention requires more features than the features explicitly stated in each claim. More specifically, as reflected in the corresponding claims, the inventive point is that the corresponding technical problem can be solved with less than all the features of a single disclosed embodiment. Therefore, the claims following the specific embodiment are hereby expressly incorporated into the specific embodiment, wherein each claim itself serves as a separate embodiment of the present invention.
[0127] It will be understood by those skilled in the art that, except for mutually exclusive features, all features disclosed in this specification (including the accompanying claims, abstracts and drawings) and all processes or units of any method or device disclosed in this specification may be combined in any combination. Unless otherwise expressly stated, each feature disclosed in this specification (including the accompanying claims, abstracts and drawings) may be replaced by an alternative feature that provides the same, equivalent or similar purpose.
[0128] In addition, those skilled in the art will appreciate that, although some embodiments described herein include certain features included in other embodiments but not other features, the combination of features of different embodiments is meant to be within the scope of the present invention and form different embodiments. For example, in the claims, any one of the claimed embodiments may be used in any combination.
[0129] The various component embodiments of the present invention can be implemented in hardware, or in software modules running on one or more processors, or in a combination thereof. It should be understood by those skilled in the art that a microprocessor or other suitable processor can be used in practice to implement some or all of the functions of some modules according to embodiments of the present invention. The present invention can also be implemented as a device program (e.g., computer program and computer program product) for executing part or all of the methods described herein. Such a program implementing the present invention can be stored on a computer-readable medium, or can have the form of one or more signals. Such a signal can be downloaded from an Internet website, or provided on a carrier signal, or provided in any other form.
[0130] It should be noted that the above embodiments illustrate the present invention rather than limit it, and that those skilled in the art may devise alternative embodiments without departing from the scope of the appended claims. In the claims, any reference symbol between brackets shall not be construed as a limitation on the claims. The word "comprising" does not exclude the presence of elements or steps not listed in the claims. The word "one" or "an" preceding an element does not exclude the presence of a plurality of such elements. The present invention may be implemented by means of hardware comprising a number of different elements and by means of a suitably programmed computer. In a unit claim enumerating a number of devices, several of these devices may be embodied by the same hardware item. The use of the words first, second, and third, etc., does not indicate any order. These words may be interpreted as names.
[0131] The above is only a specific embodiment of the present invention or an explanation of a specific embodiment. The protection scope of the present invention is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present invention, which should be included in the protection scope of the present invention. The protection scope of the present invention shall be based on the protection scope of the claims.
Claims
1. A training method for a face quality assessment model, characterized in that: The method comprises: Acquire a fused data set, wherein the fused data set includes a plurality of face sample images, each of the face sample images has a face key point label, a face occlusion label, and a face angle label, and at least one of the face key point label, the face occlusion label, and the face angle label is a true label; The multi-task quality assessment model to be trained is trained according to the fusion data set to obtain a trained multi-task quality assessment model, wherein the multi-task quality assessment model includes a face key point branch, a face occlusion branch and a face angle branch, wherein the face key point branch is used to output face key point information, the face occlusion branch is used to output face occlusion information, and the face angle branch is used to output face angle information.
2. The method according to claim 1, characterized in that: At least one of the face key point label, the face occlusion label and the face angle label is a default label; Each of the face sample images also has a face key point loss weight corresponding to the face key point label, a face occlusion loss weight corresponding to the face occlusion label, and a face angle loss weight corresponding to the face angle label; Among the face key point loss weights, the face occlusion loss weights and the face angle loss weights, the value of the loss weight corresponding to the default label is 0, and the value of the loss weight corresponding to the true label is not 0.
3. The method according to claim 2, characterized in that The step of training the multi-task quality assessment model to be trained according to the fusion data set to obtain a trained multi-task quality assessment model includes: Inputting the face sample image into the multi-task quality evaluation model to be trained, and obtaining the face key point prediction results output by the face key point branch of the multi-task quality evaluation model to be trained, the face occlusion prediction results output by the face occlusion branch, and the face angle prediction results output by the face angle branch; Obtaining the face key point loss according to the face key point prediction result, the face key point label and the face key point loss weight; Obtaining the face occlusion loss according to the face occlusion prediction result, the face occlusion label and the face occlusion loss weight; Obtaining a face angle loss according to the face angle prediction result, the face angle label and the face angle loss weight; The face key point loss, the face occlusion loss and the face angle loss are summed to obtain a summary loss, and the model parameters of the multi-task quality evaluation model are optimized according to the summary loss to obtain a trained multi-task quality evaluation model.
4. The method according to claim 3, characterized in that The facial key point prediction results, the facial key point labels and the facial key point loss weights correspond to a plurality of facial key points respectively; The step of obtaining the face key point loss according to the face key point prediction result, the face key point label and the face key point loss weight comprises: Obtaining sub-face key point losses corresponding to different face key points according to the face key point prediction results and the face key point labels corresponding to different face key points; The sub-face key point losses corresponding to multiple face key points are weightedly summed according to the face key point loss weights corresponding to multiple face key points to obtain the face key point loss.
5. The method according to claim 3, characterized in that: The face occlusion prediction result, the face occlusion label and the face occlusion loss weight respectively correspond to a plurality of face parts; The obtaining of the face occlusion loss according to the face occlusion prediction result, the face occlusion label and the face occlusion loss weight includes: Obtaining sub-face occlusion losses corresponding to different face parts according to the face occlusion prediction results and the face occlusion labels corresponding to different face parts; The sub-face occlusion losses corresponding to multiple face parts are weightedly summed according to the face occlusion loss weights corresponding to multiple face parts to obtain the face occlusion loss.
6. The method according to claim 3, characterized in that The face angle prediction result, the face angle label and the face angle loss weight respectively correspond to multiple face angles; The obtaining of the face angle loss according to the face angle prediction result, the face angle label and the face angle loss weight includes: Obtaining sub-face angle losses corresponding to different face angles according to the face angle prediction results and the face angle labels corresponding to different face angles; The sub-face angle losses corresponding to multiple face angles are weightedly summed according to the face angle loss weights corresponding to multiple face angles to obtain the face angle loss.
7. The method according to claim 1, characterized in that Before inputting the sample face image into the multi-task quality assessment model to be trained, the method further includes preprocessing the sample face image, wherein the preprocessing includes at least one of the following: Scaling the face sample image to a preset size; Performing random color changes on the face sample image; The face sample image is normalized.
8. A method for evaluating the quality of a face image, characterized in that: The method comprises: Get face image; Inputting the face image into a trained multi-task quality assessment model, wherein the multi-task quality assessment model includes a face key point branch, a face occlusion branch, and a face angle branch; Obtaining facial key point information output by the facial key point branch, facial occlusion information output by the facial occlusion branch, and facial angle information output by the facial angle branch; A quality evaluation result of the face image is obtained based on the face key point information, the face occlusion information and the face angle information.
9. The method according to claim 8, characterized in that The facial key point information includes coordinates of multiple facial key points.
10. The method according to claim 9, characterized in that The obtaining of the quality evaluation result of the face image based on the face key point information, the face occlusion information and the face angle information includes: Obtaining facial key point distribution information according to the coordinates of the plurality of facial key points; The quality evaluation result is obtained according to the facial key point distribution information.
11. The method according to claim 8, characterized in that The face occlusion information includes occlusion probabilities of multiple face parts.
12. The method according to claim 11, characterized in that The obtaining of the quality evaluation result of the face image based on the face key point information, the face occlusion information and the face angle information includes: Comparing the occlusion probabilities of the multiple face parts with the corresponding preset probability thresholds respectively to determine an occluded face part among the multiple face parts, wherein the occlusion probability of the occluded face part is greater than or equal to the corresponding preset probability threshold; The quality evaluation result is obtained according to the blocked facial part.
13. The method according to claim 11 or 12, characterized in that: The multiple facial parts include the following: left eye, right eye, nose, mouth, left face, right face, forehead.
14. The method according to claim 8, characterized in that The face angle information includes at least one of the following: a face yaw angle, a face pitch angle, and a face roll angle.
15. The method according to claim 14, characterized in that The obtaining of the quality evaluation result of the face image based on the face key point information, the face occlusion information and the face angle information includes: The face yaw angle, the face pitch angle and the face roll angle are respectively compared with the corresponding preset angle thresholds. When at least one of the face yaw angle, the face pitch angle and the face roll angle exceeds the corresponding preset angle threshold, it is determined that the face angle does not meet the preset requirements.
16. The method according to claim 8, characterized in that The multi-task quality evaluation model is trained based on a fusion data set, which includes multiple face sample images, each of which has a face key point label, a face occlusion label and a face angle label, and at least one of the face key point label, the face occlusion label and the face angle label is a true label.
17. The method according to claim 8, characterized in that Before inputting the face image into the multi-task quality assessment model, the method further includes preprocessing the face image, wherein the preprocessing includes at least one of the following: Scaling the facial image to a preset size; The face image is normalized.
18. The method according to claim 8, characterized in that The facial image includes a visible light facial image or an infrared facial image.
19. An electronic device, characterized in that: The electronic device comprises a memory and a processor, wherein a computer program executed by the processor is stored in the memory, and the computer program executes the method according to any one of claims 1 to 18 when executed by the processor.
20. A computer readable medium, characterized in that The computer readable medium stores a computer program, which executes the method according to any one of claims 1 to 18 when executed.