Training method for face feature extraction model, face feature extraction method and device

通过训练人脸特征提取模型,结合多种损失信息优化模型参数,解决了人脸角度变化对识别精度的影响,实现了高准确度的人脸特征提取和识别。

CN119810894BActive Publication Date: 2025-07-11MLKEY
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510287919.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-07-11
Estimated Expiration
2045-03-12

AI Technical Summary

Technical Problem

In the prior art, the face feature extraction algorithm is sensitive to changes in face angles, resulting in a decrease in face recognition accuracy.

Method used

By training the face feature extraction model, angle compensation and feature extraction methods are adopted, combining angle invariance loss information, measurement learning loss information and angle sensitivity loss information, model parameters are optimized to improve face recognition accuracy.

Benefits of technology

At different acquisition angles, the accuracy of facial feature extraction is improved and the accuracy of facial recognition is improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119810894B_ABST
    Figure CN119810894B_ABST
Patent Text Reader

Abstract

The present application relates to the field of image recognition technology, and discloses a training method for a face feature extraction model, a face feature extraction method, and a device. The training method includes: obtaining multiple groups of training face images; inputting the training face images into a preset model to be trained to perform angle compensation and feature extraction on the training face images, so as to obtain predicted face features; training the model to be trained according to the predicted face features and corresponding model loss information to obtain a face feature extraction model; the model loss information is obtained by fitting the angle invariance loss information, the metric learning loss information, and the angle sensitivity loss information of the predicted face features. The embodiments of the present application can perform angle compensation on the face to be recognized to improve the accuracy of face recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of image recognition technology, and in particular, to a training method for a face feature extraction model, a face feature extraction method, and a device. Background Art

[0002] Face recognition is a biometric identification technology based on human face feature information. Its basic principle includes three main steps: face detection, feature extraction, and matching recognition.

[0003] However, in the process of feature extraction, the face feature extraction algorithm is sensitive to the change of face angle. When the face in the face image is at a non-frontal angle, the accuracy of face feature extraction will decrease significantly, resulting in a decrease in face recognition accuracy. Summary of the Invention

[0004] The purpose of this application is to provide a training method for a face feature extraction model, a face feature extraction method, and a device, which can perform angle compensation on the face to be recognized to improve the face recognition accuracy.

[0005] An embodiment of this application provides a training method for a face feature extraction model, including:

[0006] Obtain multiple groups of training face images;

[0007] Input the training face images into a preset model to be trained to perform angle compensation and feature extraction on the training face images, and obtain predicted face features;

[0008] Train the model to be trained according to the predicted face features and the corresponding model loss information to obtain the face feature extraction model; the model loss information is obtained by fitting the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features.

[0009] In some embodiments, before the step of inputting the training face images into a preset model to be trained to perform angle compensation and feature extraction on the training face images and obtain predicted face features, it further includes:

[0010] Extract the roll angle feature of the face in the training face images;

[0011] Align the face in the training face images according to the roll angle feature to obtain the aligned training face images.

[0012] In some embodiments, the step of inputting the training face images into a preset model to be trained to perform angle compensation and feature extraction on the training face images and obtain predicted face features includes:

[0013] Extract facial features from the training facial images to obtain the original facial features;

[0014] Obtain the yaw angle feature and pitch angle feature of the face in the training facial image;

[0015] Perform angle compensation on the yaw angle feature and the pitch angle feature to obtain the angle compensation feature;

[0016] Fuse the original facial features and the angle compensation feature to obtain the predicted facial features.

[0017] In some embodiments, the performing angle compensation on the yaw angle feature and the pitch angle feature to obtain the angle compensation feature includes:

[0018] Concatenate the yaw angle feature and the pitch angle feature to obtain a first concatenated feature;

[0019] Perform a linear mapping and an activation operation on the first concatenated feature to obtain the angle compensation feature.

[0020] In some embodiments, the fusing the original facial features and the angle compensation feature to obtain the predicted facial features includes:

[0021] Concatenate the original facial features and the angle compensation feature to obtain a second concatenated feature;

[0022] Perform a linear mapping, normalization, and activation operation on the second concatenated feature to obtain the predicted facial features.

[0023] In some embodiments, the training the model to be trained according to the predicted facial features and the corresponding model loss information to obtain the facial feature extraction model includes:

[0024] Obtain the target facial image corresponding to the training facial image; the target facial image is a facial image obtained by collecting a face at a target acquisition angle;

[0025] Extract features from the target facial image to obtain target facial features;

[0026] Determine the model loss information according to the predicted facial features and the target facial features;

[0027] Iteratively adjust the network parameters of the model to be trained according to the model loss information until the training end condition is met, to obtain the facial feature extraction model.

[0028] In some embodiments, the calculation formula of the model loss information is:

[0029] ,

[0030] Among them, is the model loss information, is the angle invariance loss information, is the metric learning loss information, is the angle sensitivity loss information, 、 and are all loss weight coefficients;

[0031] The calculation formula of the angle invariance loss information is:

[0032] ,

[0033] Among them, N is the number of training face images, is the target face feature of the target face image corresponding to the i-th group of training face images, is the predicted face feature of the i-th group of training face images, and i is a positive integer;

[0034] The calculation formula of the metric learning loss information is:

[0035] ,

[0036] ,

[0037] ,

[0038] Among them, is the Arcface loss information of the target face image, is the Arcface loss information of the training face image, is the angle between the target face feature of the target face image corresponding to the i-th group of training face images and the weight vector of the j-th other category, is the angle between the target face feature of the target face image corresponding to the i-th group of training face images and the weight vector of the corresponding true category, is the angle between the predicted face feature of the i-th group of training face images and the weight vector of the j-th other category, is the angle between the predicted face feature of the i-th group of training face images and the weight vector of the corresponding true category, m is the angle margin, s is the scaling factor, n is the total number of categories, j is a positive integer, and e is the natural constant;

[0039] The calculation formula of the angle sensitivity loss information is:

[0040] ,

[0041] Among them, is the concatenated vector of the p-th feature element among N predicted face features, is the concatenated vector of the yaw angles of the faces in N training face images, is the concatenated vector of the pitch angles of the faces in N training face images, is and is the correlation coefficient between them, where P is the face feature dimension.

[0042] An embodiment of the present application also provides a face feature extraction method, including:

[0043] Obtain a face image to be extracted;

[0044] Input the face image to be extracted into a face feature extraction model to perform angle compensation and feature extraction on the face image to be extracted, and obtain face compensation features; the face feature extraction model is trained by the training method of the above-mentioned face feature extraction model.

[0045] An embodiment of the present application also provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned method is implemented.

[0046] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0047] The beneficial effects of the present application: By obtaining a face image to be extracted containing a face, inputting the face image to be extracted as the input of the trained face feature extraction model, and then obtaining the face compensation features corresponding to the face in the face image to be extracted output by the trained face feature extraction model. Among them, the face feature extraction model is trained according to the predicted face features and the corresponding model loss information. The predicted face features are obtained by performing angle compensation and feature extraction on the training face images, and the model loss information is fitted according to the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features. By training the model to be trained with the innovative model loss information and obtaining the required face feature extraction model, when performing face feature extraction on the face images to be extracted obtained from different acquisition angles, a relatively high-accuracy face feature extraction effect can be obtained. BRIEF DESCRIPTION OF THE DRAWINGS

[0048] Figure 1 is an application environment diagram of the training method of the face feature extraction model provided by the embodiment of the present application.

[0049] Figure 2It is a flowchart of a method for training a face feature extraction model provided by an embodiment of the present application.

[0050] Figure 3 It is a schematic diagram of the acquisition angle of a face in a training face image provided by an embodiment of the present application.

[0051] Figure 4 It is a flowchart of a face feature extraction method provided by an embodiment of the present application.

[0052] Figure 5 It is a schematic diagram of the hardware structure of an electronic device provided by an embodiment of the present application. Detailed implementation manners

[0053] In order to make the objectives, technical solutions and advantages of the present application clearer, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.

[0054] Figure 1 It is an application environment diagram of a method for training a face feature extraction model provided by an embodiment of the present application. Refer to Figure 1 , this face feature extraction method is applied to a training system of a face feature extraction model. The training system of this face feature extraction model includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 may specifically be a desktop terminal or a mobile terminal, and the mobile terminal may specifically be at least one of a mobile phone, a tablet computer, a notebook computer, etc. The server 120 may be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to upload multiple groups of training face images to the server 120. The server 120 is used to obtain multiple groups of training face images, input the training face images into a preset model to be trained, perform angle compensation and feature extraction on the training face images to obtain predicted face features, and train the model to be trained according to the predicted face features and corresponding model loss information to obtain a face feature extraction model, where the model loss information is obtained by fitting the angle invariance loss information, metric learning loss information and angle sensitivity loss information of the predicted face features.

[0055] In another embodiment, the above face feature extraction method may be directly applied to the terminal 110. The terminal 110 is used to obtain multiple groups of training face images, input the training face images into a preset model to be trained, perform angle compensation and feature extraction on the training face images to obtain predicted face features, and train the model to be trained according to the predicted face features and corresponding model loss information to obtain a face feature extraction model, where the model loss information is obtained by fitting the angle invariance loss information, metric learning loss information and angle sensitivity loss information of the predicted face features.

[0056] Refer to Figure 2 , in one embodiment, a method for training a face feature extraction model is provided. This method can be applied to both a terminal and a server. In this embodiment, it is exemplified by being applied to a terminal. The method for training the face feature extraction model includes but is not limited to steps S201 to S203.

[0057] Step S201: Obtain multiple sets of training face images.

[0058] The training face images refer to the images containing faces that are required for training the model. In order to train the model to be trained, it is necessary to use the training face images to train the model to be trained. The training face images respectively contain different faces and are obtained by collecting faces from different acquisition angles.

[0059] Refer to Figure 3 , the acquisition angles include the roll angle (Roll), yaw angle (Yaw), and pitch angle (Pitch) of the face in the training face image relative to the corresponding coordinate axes of the training face image.

[0060] Step S202: Input the training face images into a preset model to be trained to perform angle compensation and feature extraction on the training face images, and obtain predicted face features.

[0061] The predicted face features are the face features when the face in the training face image is collected at the target acquisition angle predicted by the model to be trained. By using the training face images as the input of the model to be trained, the model to be trained performs angle compensation and feature extraction on the training face images, and then obtains the predicted face features corresponding to the faces in the training face images output by the model to be trained, which is convenient for subsequently adjusting the network parameters in the model to be trained according to the predicted face features and the corresponding model loss information, so that the model to be trained is optimized in a more accurate direction.

[0062] Step S203: Train the model to be trained according to the predicted face features and the corresponding model loss information to obtain a face feature extraction model. Among them, the model loss information is obtained by fitting the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features.

[0063] After obtaining the predicted face features, calculate the model loss information based on the predicted face features and the corresponding target face features. Here, the target face features are the face features obtained by extracting features from the target face image, and the target face image is the face image obtained by capturing a face at the target acquisition angle. Calculating the model loss information is to determine the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features based on the predicted face features and the target face features, and fit the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features to obtain the model loss information. After obtaining the model loss information, adjust the network parameters in the face feature extraction model according to the model loss information, and continue to train the adjusted model to be trained by repeating the above steps until the training end condition is met, and obtain the face feature extraction model. The face feature extraction model is the model to be trained after training. In one embodiment, the training end condition is a threshold set in advance for measuring the deviation of the model loss information. When the calculated model loss information is less than the set threshold, it indicates that the model to be trained is trained.

[0064] The target acquisition angle is the acquisition angle facing the face directly. The roll angle, yaw angle, and pitch angle corresponding to the target acquisition angle are all 0°. That is, the target face image is the face image obtained by capturing a face facing the face directly.

[0065] The above training method of the face feature extraction model takes the training face image as the input of the model to be trained. The model to be trained performs angle compensation and feature extraction on the training face image to obtain the predicted face features. According to the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features, the corresponding model loss information is fitted, and then the model loss information is used to train the model to be trained. When the training end condition is met, the required face feature extraction model is obtained. By using the innovative model loss information to train the model to be trained and obtaining the required face feature extraction model, when performing face feature extraction on the face images to be extracted obtained from different acquisition angles, a relatively high-accuracy face feature extraction effect can be obtained.

[0066] In one embodiment, before inputting the training face image into the preset model to be trained to perform angle compensation and feature extraction on the training face image to obtain the predicted face features, it further includes: extracting the roll angle feature of the face in the training face image; aligning the face in the training face image according to the roll angle feature to obtain the aligned training face image.

[0067] To improve the image quality of the training face image, align the face in the training face image.

[0068] Align the face in the training face image according to the roll angle feature, which can be to extract the roll angle feature of the face in the training face image, compare the extracted roll angle feature of the face with the roll angle feature of the face in the target face image, and on the basis of keeping the yaw angle feature and pitch angle feature of the face in the face image unchanged, align the face in the training face image so that the roll angle feature of the face in the training face image is consistent with the roll angle feature of the face in the target face image, that is, the roll angle of the face in the training face image is the same as the roll angle of the face in the target face image. Extracting the roll angle feature of the face in the training face image can be to perform face key point detection on the training face image and determine the roll angle feature of the face in the training face image according to the face key point detection result.

[0069] In one embodiment, input the training face image into a preset model to be trained to perform angle compensation and feature extraction on the training face image to obtain predicted face features, specifically including: performing face feature extraction on the training face image to obtain original face features; obtaining the yaw angle feature and pitch angle feature of the face in the training face image; performing angle compensation on the yaw angle feature and pitch angle feature to obtain angle compensation features; and fusing the original face features and angle compensation features to obtain predicted face features.

[0070] The model to be trained can include a feature extraction network, an angle compensation network, and a feature fusion network. When inputting the training face image into the model to be trained, the yaw angle feature and pitch angle feature of the face in the training face image are input into the model to be trained together. Performing face feature extraction on the training face image is executed by the feature extraction network. Input the training face image into the feature extraction network, and the training face image performs face feature extraction in the feature extraction network by means of continuous convolution or straight-line transfer. Finally, the feature extraction network outputs the original face features obtained by performing face feature extraction on the training face image. Performing angle compensation on the yaw angle feature and pitch angle feature is executed by the angle compensation network. After obtaining the yaw angle feature and pitch angle feature of the face in the training face image, input the yaw angle feature and pitch angle feature into the angle compensation network. The angle compensation network splices the yaw angle feature and pitch angle feature and then performs feature mapping on the spliced angle feature according to the current network parameters, and finally outputs the angle compensation feature. Fusing the original face features and angle compensation features is executed by the feature fusion network. After obtaining the original face features and angle compensation features, input the original face features and angle compensation features into the feature fusion network. The feature fusion network splices the original face features and angle compensation features and then performs feature mapping on the spliced features according to the current network parameters, and finally outputs the predicted face features.

[0071] In one embodiment, angle compensation is performed on the yaw angle feature and the pitch angle feature to obtain an angle compensation feature, including: splicing the yaw angle feature and the pitch angle feature to obtain a first spliced feature; performing a linear mapping and an activation operation on the first spliced feature to obtain the angle compensation feature. Among them, the angle compensation network includes two fully connected layers and a ReLu activation layer. After splicing the yaw angle feature and the pitch angle feature into the first spliced feature, it is input into the angle compensation network, and passes through two fully connected layers and a ReLu activation layer in sequence to perform a linear mapping and an activation operation, and finally outputs the angle compensation feature.

[0072] In one embodiment, the original face feature and the angle compensation feature are fused to obtain a predicted face feature, including: splicing the original face feature and the angle compensation feature to obtain a second spliced feature; performing a linear mapping, normalization, and activation operation on the second spliced feature to obtain the predicted face feature. Among them, the feature fusion network includes two fully connected layers, a normalization layer, and a ReLu activation layer. After splicing the original face feature and the angle compensation feature into the second spliced feature, it is input into the feature fusion network, and passes through two fully connected layers, a normalization layer, and a ReLu activation layer in sequence to perform a linear mapping, normalization, and activation operation, and finally outputs the predicted face feature.

[0073] In one embodiment, the model to be trained is trained according to the predicted face feature and the corresponding model loss information to obtain a face feature extraction model, including: obtaining the target face image corresponding to the training face image; performing feature extraction on the target face image to obtain the target face feature; determining the model loss information according to the predicted face feature and the target face feature; iteratively adjusting the network parameters of the model to be trained according to the model loss information until the training end condition is met, and obtaining the face feature extraction model.

[0074] Before inputting the training face images into the model to be trained, it is possible to first obtain the target face images corresponding to the training face images, and then input the training face images, the target face images, and the yaw angle features and pitch angle features of the training face images into the model to be trained together. Feature extraction of the target face images can be performed by the feature extraction network of the model to be trained. The target face images are input into the feature extraction network, and the target face images are subjected to face feature extraction in the feature extraction network by means of continuous convolution or straight-line transfer. Finally, the feature extraction network outputs the target face features obtained by performing face feature extraction on the target face images. Iteratively adjusting the network parameters of the model to be trained according to the model loss information can be achieved by presetting a loss threshold range and a reset threshold as the training end conditions. When the model loss information is within the loss threshold range and the training reset times reach the reset threshold, the training is ended, and the model to be trained of the last iteration obtained is the face feature extraction model. When the training reset times do not reach the reset threshold, the network parameters of the model to be trained are adjusted according to the deviation degree of the model loss information from the loss threshold range, so that the model loss information gradually approaches the loss threshold range and finally falls within the loss threshold range during the iteration process. Then, the training face images and the target face images are obtained again, and the above steps are repeated to obtain new model loss information until the model loss information is within the loss threshold range. Repeat the above steps until the training reset times reach the reset threshold, and then end the training to obtain the face feature extraction model, which can make the obtained face feature extraction model adaptable to face feature extraction scenarios for various face images to be extracted.

[0075] In one embodiment, the calculation formula for the model loss information is:

[0076] ,

[0077] where, is the model loss information, is the angle invariance loss information, is the metric learning loss information, is the angle sensitivity loss information, , and are all loss weight coefficients;

[0078] The calculation formula for the angle invariance loss information is:

[0079] ,

[0080] where N is the number of training face images, is the target face feature of the target face image corresponding to the i-th group of training face images, is the predicted face feature of the i-th group of training face images, and i is a positive integer;

[0081] The calculation formula of the metric learning loss information is as follows:

[0082] ,

[0083] ,

[0084] ,

[0085] wherein, is the Arcface loss information of the target face image, is the Arcface loss information of the training face image, is the included angle between the target face feature of the target face image corresponding to the i-th group of training face images and the weight vector of the j-th other category, is the included angle between the target face feature of the target face image corresponding to the i-th group of training face images and the weight vector of the corresponding true category, is the included angle between the predicted face feature of the i-th group of training face images and the weight vector of the j-th other category, is the included angle between the predicted face feature of the i-th group of training face images and the weight vector of the corresponding true category, m is the angle interval, s is the scaling factor, n is the total number of categories, j is a positive integer, and e is the natural constant;

[0086] The calculation formula of the angle sensitivity loss information is as follows:

[0087] ,

[0088] wherein, is the concatenated vector of the p-th feature element among N predicted face features, is the concatenated vector of the yaw angles of the faces in N training face images, is the concatenated vector of the pitch angles of the faces in N training face images, is and the correlation coefficient between, P is the dimension of the face feature.

[0089] Referring to Figure 4 , in one embodiment, a face feature extraction method is provided. This method can be applied to both the terminal and the server. In this embodiment, taking the application to the terminal as an example, the face feature extraction method includes but is not limited to steps S401 to S402.

[0090] Step S401, obtaining the face image to be extracted.

[0091] The face image to be extracted refers to an image containing a human face from which face features are to be extracted. The face image to be extracted can be an image containing a human face obtained by the terminal through real-time shooting by calling the camera, or it can be a stored image containing a human face that has been acquired. The human face in the face image to be extracted can be at any angle. For example, it can be a side face or a front face, etc. The face image to be extracted can include one human face or multiple human faces. If it includes multiple human faces, the face features corresponding to each human face in the image need to be extracted subsequently.

[0092] Step S402: Input the face image to be extracted into the face feature extraction model to perform angle compensation and feature extraction on the face image to be extracted, and obtain face compensation features. This face feature extraction model is trained by the training method of the above-mentioned face feature extraction model.

[0093] In the above face feature extraction method, by obtaining the face image to be extracted containing a human face, using the face image to be extracted as the input of the trained face feature extraction model, and then obtaining the face compensation features corresponding to the human face in the face image to be extracted output by the trained face feature extraction model. The above face feature extraction model is trained according to the predicted face features and the corresponding model loss information. The predicted face features are obtained by performing angle compensation and feature extraction on the training face images, and the model loss information is fitted according to the angle invariance loss information, metric learning loss information, and angle sensitivity loss information of the predicted face features. By training the model to be trained with the innovative model loss information and obtaining the required face feature extraction model, when performing face feature extraction on the face images to be extracted obtained from different acquisition angles, a relatively high-accuracy face feature extraction effect can be obtained.

[0094] Refer to Figure 5 , in one embodiment, an electronic device 500 is provided. The electronic device 500 includes but is not limited to: at least one processing unit 510, at least one storage unit 520, a bus 530 connecting different system components (including the storage unit 520 and the processing unit 510), a display unit 540, etc.

[0095] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 510, so that the processing unit 510 executes the steps according to various exemplary embodiments of the present disclosure described in the method part of this specification.

[0096] The storage unit 520 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 5201 and / or a cache storage unit 5202, and may further include a read-only storage unit (ROM) 5203.

[0097] The storage unit 520 may also include a program / utility 5204 having a set (at least one) of program modules 5205. Such program modules 5205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment.

[0098] The bus 530 may represent one or more of several types of bus structures, including a memory bus or memory controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus structures.

[0099] The electronic device 500 may also communicate with one or more external devices 500' (such as a keyboard, a pointing device, a Bluetooth device, etc.), and may also communicate with one or more devices that enable a user to interact with the electronic device 500, and / or communicate with any device that enables the electronic device 500 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication may be carried out through an input / output (I / O) interface 550. In addition, the electronic device 500 may also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 560. The network adapter 560 may communicate with other modules of the electronic device 500 through the bus 530. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 500, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.

[0100] The embodiment of the present application also provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the above-mentioned method is implemented.

[0101] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described herein can be implemented by software, or can be implemented by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which may be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which may be a personal computer, a server, or a network device, etc.) to execute the above-mentioned method according to the embodiments of the present disclosure.

[0102] The program product may employ any combination of one or more readable media. The readable media may be a readable signal medium or a readable storage medium. A readable storage medium may be, for example but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the foregoing. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection having one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0103] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which the readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.

[0104] Those skilled in the art can understand that the above-described modules may be distributed in the device according to the description of the embodiments, or may be correspondingly changed and distributed in one or more devices that are uniquely different from the embodiments. The modules of the above embodiments may be combined into one module, or may be further split into multiple sub-modules.

[0105] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, settings, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.

Claims

1. A training method for a face feature extraction model, characterized in that Including: Obtain multiple groups of training face images; Input the training face images into a preset model to be trained, so as to perform angle compensation and feature extraction on the training face images, and obtain predicted face features; Train the model to be trained according to the predicted face features and the corresponding model loss information to obtain the face feature extraction model; the model loss information is obtained by fitting the angle invariance loss information, metric learning loss information and angle sensitivity loss information of the predicted face features; The calculation formula of the model loss information is: , Among them, is the model loss information, is the angle invariance loss information, is the metric learning loss information, is the angle sensitivity loss information, , and are all loss weight coefficients; The calculation formula of the angle invariance loss information is: , where N is the number of training face images, is the target face feature of the target face image corresponding to the i-th group of training face images, is the predicted face feature of the i-th group of training face images, and i is a positive integer; The calculation formula of the metric learning loss information is: , , , Among them, is the Arcface loss information of the target face image, is the Arcface loss information of the training face image, is the angle between the target face feature of the target face image corresponding to the i-th group of training face images and the weight vector of the j-th other category, is the angle between the target face feature of the target face image corresponding to the i-th group of training face images and the weight vector of the corresponding true category, is the angle between the predicted face feature of the i-th group of training face images and the weight vector of the j-th other category, is the angle between the predicted face feature of the i-th group of training face images and the weight vector of the corresponding true category, m is the angle interval, s is the scaling factor, n is the total number of categories, j is a positive integer, and e is the natural constant; The calculation formula of the angle sensitivity loss information is: , Among them, is the concatenated vector of the p-th feature element among N predicted face features, is the concatenated vector of the yaw angles of the faces in N training face images, is the concatenated vector of the pitch angles of the faces in N training face images, is and is the correlation coefficient between them, and P is the face feature dimension.

2. The training method of the face feature extraction model according to claim 1, characterized in that Before the step of inputting the training face images into a preset model to be trained, so as to perform angle compensation and feature extraction on the training face images and obtain predicted face features, it further includes: Extract the roll angle feature of the face in the training face image; Align the face in the training face image according to the roll angle feature to obtain an aligned training face image.

3. The training method of the face feature extraction model according to claim 1, characterized in that, The step of inputting the training face images into a preset model to be trained, so as to perform angle compensation and feature extraction on the training face images and obtain predicted face features includes: Perform face feature extraction on the training face images to obtain original face features; Obtain the yaw angle feature and pitch angle feature of the face in the training face image; Perform angle compensation on the yaw angle feature and the pitch angle feature to obtain an angle compensation feature; Fuse the original face features and the angle compensation feature to obtain the predicted face features.

4. The training method of the face feature extraction model according to claim 3, wherein The step of performing angle compensation on the yaw angle feature and the pitch angle feature to obtain an angle compensation feature includes: Concatenate the yaw angle feature and the pitch angle feature to obtain a first concatenated feature; Perform a linear mapping and an activation operation on the first concatenated feature to obtain the angle compensation feature.

5. The training method of the face feature extraction model according to claim 3, characterized in that The step of fusing the original face features and the angle compensation feature to obtain the predicted face features includes: Concatenate the original face features and the angle compensation feature to obtain a second concatenated feature; Perform a linear mapping, normalization and activation operation on the second concatenated feature to obtain the predicted face features.

6. The training method of the face feature extraction model according to claim 1, wherein The step of training the model to be trained according to the predicted face features and the corresponding model loss information to obtain the face feature extraction model includes: Obtain the target face image corresponding to the training face image; the target face image is a face image obtained by collecting a face at a target acquisition angle; Perform feature extraction on the target face image to obtain target face features; Determine the model loss information according to the predicted face features and the target face features; Iteratively adjust the network parameters of the model to be trained according to the model loss information until the training end condition is met, and obtain the face feature extraction model.

7. A face feature extraction method, characterized in that, Including: Obtain a face image to be extracted; Input the face image to be extracted into the face feature extraction model to perform angle compensation and feature extraction on the face image to be extracted, and obtain face compensation features; the face feature extraction model is trained by the training method of the face feature extraction model according to any one of claims 1 to 6.

8. An electronic device, characterized in that, The electronic device includes a memory and a processor, the memory stores a computer program, and when the processor executes the computer program, the method according to any one of claims 1 to 7 is implemented.

9. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Face recognition method and device, robot and storage medium

    CN109543633A

  • Training method of deep convolutional neural network for face recognition

    CN111985310A