Face recognition method and device
The face recognition method improves accuracy by extracting and fusing facial pose and facial features to enhance user recognition in complex environments with rapid pose changes.
Patent Information
- Application Number
- JP2024538337
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-12-31
- Filing Date
- 2022-07-26
- Publication Date
- 2025-11-27
- Estimated Expiration
- 2042-07-26
AI Technical Summary
Conventional face recognition technologies suffer from poor generalization and performance in complex environments and cross-working scenarios, particularly when face pose changes rapidly and dramatically, falling short of practical application requirements.
A face recognition method that extracts facial pose features and facial features from an image, generates target facial features and pose information, and determines a user recognition result based on predetermined conditions, including facial structure, edge, and angle features, along with skin color and texture, to improve accuracy.
Enhances the accuracy of face recognition by fusing facial pose and facial features, improving the precision of user recognition results in dynamic environments.
Smart Images

Figure 0007777229000002 
Figure 0007777229000003 
Figure 0007777229000004
Abstract
Description
[Technical Field]
[0001] The present disclosure relates to the field of data processing technology, and in particular to a face recognition method and apparatus. [Background technology]
[0002] Facial recognition systems are an emerging biometric technology, a high-precision technology currently being tackled in the international scientific and technological field, with broad potential for development. While traditional face detection technologies often achieve excellent performance for faces in ideal environments, they suffer from poor generalization and performance in some complex environments and cross-working scenarios. In particular, in real-world application scenarios, especially when the face pose changes rapidly and dramatically, the facial recognition accuracy of traditional methods still falls far short of practical application requirements, and further research and improvement are needed. Summary of the Invention
[0003] In view of this, the embodiments of the present disclosure provide a face recognition method, an apparatus, a computer device and a computer-readable storage medium to solve the problem that in the prior art, face recognition results are inaccurate when face pose changes rapidly and abruptly.
[0004] The first aspect of the disclosed embodiment is A step of acquiring a face image to be recognized; extracting facial pose features and facial features from a face image to be recognized; generating target facial features and facial pose information of a face image to be recognized based on the facial pose features and facial features; If the face pose information satisfies a predetermined condition, a step of identifying a user recognition result for the face image to be recognized based on the target face features of the face image to be recognized is provided.
[0005] A second aspect of the disclosed embodiment is an image acquisition unit for acquiring a face image to be recognized; a feature extraction unit for extracting facial pose features and facial features of a face image to be recognized; an information generating unit for generating target facial features, facial pose information, and corresponding score vectors of the facial pose information of the face image to be recognized according to the facial pose features and facial features; A face recognition device is provided, which includes a result identification unit for identifying a user recognition result for a face image to be recognized based on target facial features of the face image to be recognized when the face pose information and the score vector corresponding to the face pose information satisfy a predetermined condition.
[0006] A third aspect of an embodiment of the present disclosure provides a computing device comprising a memory, a processor, and a computer program stored in the memory and executable by the processor, the computing device implementing the steps of the method when the processor executes the computer program.
[0007] A fourth aspect of an embodiment of the present disclosure provides a computer-readable storage medium having a computer program stored thereon, the computer program implementing the steps of the above method when executed by a processor.
[0008] Beneficial advantages of the disclosed embodiment over the prior art: In the disclosed embodiment, a facial image to be recognized can be first obtained, then facial pose features and facial features of the facial image to be recognized can be extracted, and then target facial features and facial pose information of the facial image to be recognized are generated based on the facial pose features and facial features. If the facial pose information satisfies a predetermined condition, a user recognition result of the facial image to be recognized is determined based on the target facial features of the facial image to be recognized. In this embodiment, the facial pose features reflect spatial information of the face and include facial structure features, edge features, and angle features, and the facial features reflect features such as skin color and texture. Therefore, by fusing the facial pose features and facial features, detailed information of the facial pose features and facial features can be merged, and target facial features and facial pose information of the facial image to be recognized can be generated based on the fused features, thereby improving the accuracy of the identified target facial features and facial pose information and the accuracy of the user recognition result determined based on the target facial features and facial pose information. [Brief explanation of the drawings]
[0009] In order to more clearly explain the technical solutions in the embodiments of the present disclosure, the following briefly introduces drawings necessary for describing the embodiments or prior art. Obviously, the drawings described below are only some embodiments of the present disclosure, and those skilled in the art can obtain other drawings based on these drawings without the need for creative work. [Figure 1] FIG. 1 is a scenario schematic diagram of an application scenario of an embodiment of the present disclosure. [Figure 2] 1 is a flowchart of a face recognition method provided in an embodiment of the present disclosure. [Figure 3] FIG. 1 is a block diagram of a face recognition device provided in an embodiment of the present disclosure. [Figure 4] FIG. 1 is a schematic diagram of a computer device provided in an embodiment of the present disclosure. DETAILED DESCRIPTION OF THE INVENTION
[0010] In the following description, for purposes of explanation, not limitation, specific details, such as particular system structures and techniques, are provided to thoroughly understand the embodiments of the present disclosure. However, those skilled in the art should understand that the present disclosure can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, devices, circuits, and methods are omitted to avoid obscuring the description of the present disclosure with unnecessary details.
[0011] Hereinafter, a face recognition method and device according to an embodiment of the present disclosure will be described in detail with reference to the drawings.
[0012] In the prior art, conventional face detection technologies have good performance for faces in many ideal environments, but in some applications such as complex environments or cross-working scenarios, the generalization ability and performance to the environment are poor. In particular, in real application scenarios, especially when the face pose changes rapidly and dramatically, the face recognition accuracy of conventional methods is still far from the requirements of practical applications, and further research and improvement are needed.
[0013] To solve the above problems, the present invention provides a face recognition method, which first obtains a face image to be recognized, then extracts facial pose features and facial features from the face image to be recognized, and then generates target face features and facial pose information for the face image to be recognized based on the facial pose features and facial features. If the facial pose information satisfies a predetermined condition, a user recognition result for the face image to be recognized is determined based on the target face features of the face image to be recognized. In this embodiment, the facial pose features reflect spatial information of the face and include facial structure features, edge features, and angle features, and the facial features reflect features such as skin color and texture. Therefore, by combining the facial pose features and facial features to generate the target face features and facial pose information for the face image to be recognized, the accuracy of the determined target face features and facial pose information can be improved, and the accuracy of the user recognition result determined based on the target face features and facial pose information can be improved.
[0014] By way of example, an embodiment of the present invention may be applied to the application scenario shown in Figure 1. This scenario may comprise a terminal device 1 and a server 2.
[0015] The terminal device 1 may be hardware or software. If the terminal device 1 is hardware, it may be various electronic devices that have image collection functions and support communication with the server 2, including, but not limited to, smartphones, tablets, laptop computers, and desktop computers. If the terminal device 1 is software, it may be installed in, for example, the above-mentioned electronic devices. The terminal device 1 may be implemented as multiple software programs or software modules, or as a single software program or software module; the embodiments of the present disclosure are not limited thereto. The server 2 may be a server that provides various services, such as a back-end server that receives bills sent by terminal devices that establish communication connections with it. The back-end server can receive, analyze, and otherwise process the bills sent by the terminal devices to generate processing results. The server 2 may be a single server, a server cluster consisting of several servers, or a cloud computing service center; the embodiments of the present disclosure are not limited thereto.
[0016] The server 2 may be hardware or software. If the server 2 is hardware, it may be various electronic devices that provide various services to the terminal device 1. If the server 2 is software, it may be multiple software programs or software modules that provide various services to the terminal device 1, or it may be a single software program or software module that provides various services to the terminal device 1, and the embodiments disclosed herein are not limited to these.
[0017] The terminal device 1 and the server 2 may be communicatively connected via a network. The network may be a wired network connected using a coaxial cable, a twisted pair, or an optical fiber, or may be a wireless network that can realize interconnection of various communication devices without the need for wiring, such as Bluetooth, Near Field Communication (NFC), or Infrared, and the embodiments of the present disclosure are not limited thereto.
[0018] Specifically, a user can input a facial image to be recognized using terminal device 1, and terminal device 1 transmits the business data to be evaluated and the target evaluation method to server 2. Server 2 first extracts facial pose features and facial features of the facial image to be recognized, and then generates target facial features and facial pose information for the facial image to be recognized based on the facial pose features and facial features. If the facial pose information satisfies a predetermined condition, server 2 can identify a user recognition result for the facial image to be recognized based on the target facial features of the facial image to be recognized, and return the user recognition result for the facial image to terminal device 1, which then displays the user recognition result for the facial image to be recognized to the user. By fusing the facial pose features and facial features in this way, detailed information about the facial pose features and facial features can be merged, and target facial features and facial pose information for the facial image to be recognized can be generated based on the fused features. This improves the accuracy of the identified target facial features and facial pose information, further improving the accuracy of the user recognition result identified based on the target facial features and facial pose information.
[0019] It should be noted that the specific types, numbers and combinations of the terminal device 1, the server 2 and the network can be adjusted according to the actual needs of the application scenario, and the embodiments disclosed herein are not limited thereto.
[0020] It should be noted that the above application scenarios are merely exemplified for easy understanding of the present disclosure, and the embodiments of the present disclosure are not limited in this respect, but may be used in any application scenario.
[0021] Figure 2 is a flowchart of a face recognition method provided in an embodiment of the present disclosure. The face recognition method of Figure 2 may be performed by the terminal device or the server of Figure 1. As shown in Figure 2, the face recognition method includes the following steps:
[0022] S201: A face image to be recognized is acquired.
[0023] In this embodiment, the facial image to be recognized can be understood as an image that needs to be recognized. For example, the facial image to be recognized may be collected by a surveillance camera installed at a fixed location, collected by a mobile terminal device, or read from a storage device in which the image is previously stored.
[0024] S202: The facial pose features and facial features of the face image to be recognized are extracted.
[0025] After acquiring a target facial image, in order to accurately recognize a profile face in the target facial image, it is first necessary to extract facial pose characteristics and facial features from the target facial image, where the facial pose characteristics can reflect spatial information of the face, such as facial structure characteristics, edge characteristics, and angle characteristics, and the facial features can reflect semantic information about the face's skin color, texture, age, lighting, race, etc.
[0026] S203: Based on the facial pose features and facial features, target facial features and facial pose information of the face image to be recognized are generated.
[0027] In this embodiment, facial pose features and facial features are first fused, and facial feature information is supplemented using the facial pose features to enrich facial feature information. Then, target facial features and facial pose information for the facial image to be recognized are generated based on the facial pose features and facial features, thereby improving the accuracy of the identified target facial features and facial pose information. Here, the target facial features may be understood to include a feature vector of facial detail information (e.g., texture information, skin color information, etc.). The facial pose information may include the yaw angle, pitch angle, and roll angle of the face, and the facial pose information may reflect the steering angle and steering amplitude of the face.
[0028] S204: If the face pose information satisfies a predetermined condition, the user recognition result of the face image to be recognized is determined based on the target face features of the face image to be recognized.
[0029] If the face is turned to the side or the head is lowered too much, the pose angle of the facial image will also be large, thereby reducing the facial information available for recognition and making the face more vulnerable to attack (e.g., easier to decrypt using a photo of the user being impersonated). Therefore, in this embodiment, it is first necessary to determine whether the facial pose information satisfies certain conditions, which in one implementation are that the yaw angle is less than or equal to a certain yaw angle threshold, the pitch angle is less than or equal to a certain pitch angle threshold, and the roll angle is less than or equal to a certain roll angle threshold.
[0030] When the facial pose information satisfies a predetermined condition, the user recognition result of the facial image to be recognized can be determined based on the target facial features of the facial image to be recognized. For example, in this embodiment, several predetermined user information can be set in advance, and each predetermined user information has corresponding predetermined user facial features. In this embodiment, the similarity between the target facial features of the facial image to be recognized and each of the predetermined user facial features can be determined. For example, the vector distance can be measured using methods such as Euclidean distance and cosine distance, and the similarity is determined based on the vector distance. Then, the user information corresponding to the predetermined user facial feature with the greatest similarity can be used as the user recognition result of the facial image to be recognized.
[0031] Beneficial advantages of the disclosed embodiment over the prior art: In the disclosed embodiment, a facial image to be recognized can be first obtained, then facial pose features and facial features of the facial image to be recognized can be extracted, and then target facial features and facial pose information of the facial image to be recognized are generated based on the facial pose features and facial features. If the facial pose information satisfies a predetermined condition, a user recognition result of the facial image to be recognized is determined based on the target facial features of the facial image to be recognized. In this embodiment, the facial pose features reflect spatial information of the face and include facial structure features, edge features, and angle features, and the facial features reflect features such as skin color and texture. Therefore, by fusing the facial pose features and facial features, detailed information of the facial pose features and facial features can be merged, and target facial features and facial pose information of the facial image to be recognized can be generated based on the fused features, thereby improving the accuracy of the identified target facial features and facial pose information and the accuracy of the user recognition result determined based on the target facial features and facial pose information.
[0032] Next, a description will be given of an implementation of S202, that is, how to extract the facial pose features and facial features of the face image to be recognized. In this embodiment, the facial pose features may include N facial pose features, and the facial features may include N facial features, and S202 may include the following steps:
[0033] S202a: For the first facial feature, a facial image to be recognized is input to a first convolutional layer to acquire the first facial feature.
[0034] S202b: Regarding the first face pose feature, the face image to be recognized is input to the first residual block, and the first face pose feature is obtained.
[0035] S202c: For the pose feature of the ith face, input the pose feature of the (i-1)th face into the ith residual block to obtain the pose feature of the ith face.
[0036] S202d: For the i-th facial feature, input the i-1-th face pose feature and the i-1-th facial feature into the i-th convolutional layer to obtain the i-th facial feature, where i is greater than or equal to 2 and less than or equal to N, and N and i are positive integers.
[0037] In this embodiment, two models, a facial feature extraction model and a facial pose feature extraction model, may be provided. Here, the facial pose feature extraction model may include N residual blocks, including at least a first residual block, a second residual block, ..., an Nth residual block, where the N residual blocks are cascade-connected. In one implementation, the facial pose feature extraction model includes at least four residual blocks, each of which is composed of two, three, and two residual networks, respectively. Each residual network may have a network architecture including two convolutional layers (i.e., conv), two batch normalization layers (i.e., bn), and two hyperbolic tangent activation functions (i.e., tanh), with a specific connection structure of conv+bn+tanh+conv+bn+tanh. The output feature maps of these three residual blocks have 64, 128, and 256 channels, respectively. The reason for using the hyperbolic tangent activation function is that each feature calculation takes a value between (-1, 1) and can contribute to subsequent pose calculations. As can be seen, a face image to be recognized can be input into the first residual block to obtain the pose features of the first face, the pose features of the first face can be input into the second residual block to obtain the pose features of the second face, ... for the pose features of the i-th face, the pose features of the (i-1)th face can be input into the i-th residual block to obtain the pose features of the i-th face, where i is greater than or equal to 2 and less than or equal to N, and both N and i are positive integers.
[0038] Here, the facial feature extraction model may include N convolution layers, including at least a first convolution layer, a second convolution layer, ..., an Nth convolution layer, where these N convolution layers are cascaded. Here, the facial feature extraction model may be an IResNet-50 network, and each convolution layer may include one convolution operator with a convolution kernel size of 3x3 and 192 channels, and one convolution operator with a convolution kernel size of 1x1 and 128 channels. In this embodiment, for the first facial feature, the facial image to be recognized is input to the first convolutional layer to obtain the first facial feature, the pose feature of the first face (e.g., dimensions (28, 28, 64)) and the first facial feature (e.g., dimensions (28, 28, 128)) are input to the second convolutional layer (e.g., the first facial feature and the first facial feature are subjected to feature fusion, and the fused feature is input to the second convolutional layer) to obtain the second facial feature, ..., the pose feature and the i-1st facial feature of the (i-1)th face are input to the ith convolutional layer to obtain the ith facial feature, where i is greater than or equal to 2 and less than or equal to N, and both N and i are positive integers. It should be emphasized that the facial pose features are used as supplementary information and input into the convolutional layer together with the facial features to calculate the next facial feature. The reason is that the pose information takes spatial information into consideration more and can extract facial structure features, edge features, and angle features relatively completely. However, facial feature extraction needs to take into account various semantic information such as age, lighting, race, skin color, texture, etc., which results in certain deficiencies in spatial structure and semantic confusion. Therefore, when the facial pose features extracted by the facial pose feature extraction model are added to the facial feature extraction model, they effectively supplement the information processing of facial features.
[0039] It should be noted that the number of residual blocks in the facial pose feature extraction model is smaller than the number of convolutional layers in the facial feature extraction model, and the number of channels is smaller. This is because the information processed by the facial pose feature extraction model is relatively single semantic information, which does not require a large amount of calculation.
[0040] Next, an implementation of S203, i.e., how to generate target facial features and facial pose information of a face image to be recognized, will be described. In this embodiment, the step of "generating target facial features of a face image to be recognized based on facial pose features and facial features" in S203 may include the following steps:
[0041] The pose feature and Nth facial feature of the Nth face are input to the N+1th convolutional layer to obtain the target facial features of the face image to be recognized.
[0042] In this embodiment, the facial feature extraction model further includes an N+1th convolutional layer, and the Nth facial pose feature and the Nth facial feature can be input into the N+1th convolutional layer to obtain the target facial features of the facial image to be recognized.
[0043] In this embodiment, the step of "generating facial pose information of the face image to be recognized based on the facial pose features and facial features" in S203 may include the following steps.
[0044] Step a: Based on the facial pose features, generate a corresponding attention map of the facial pose features.
[0045] In this embodiment, a downsampling process is performed on all facial pose features to make the dimensions of all facial pose features the same. Then, each facial pose feature is input into an attention model to obtain a corresponding attention map for the facial pose feature. Here, the attention model includes a convolution operator with a convolution kernel size of 1x1 and a channel count of 1, and a sigmoid function. For example, if the dimension of one facial pose feature is (28,28,64), the facial pose feature is first subjected to two convolution operations with a convolution kernel size of 3x3, a step size of 2, and a channel count of 64 to reduce the dimension to (7,7,64), and then the 1x1 convolution operation and sigmoid operation are introduced to obtain the corresponding attention map as an attention map with a dimension of (7,7,1). Then, the three dimensions are reduced to one dimension to obtain the final attention map with a dimension of (7,7,1).
[0046] Step b: Based on the facial features, generate a corresponding attention map of the facial features.
[0047] In this embodiment, a downsampling process is performed on all facial features to make the dimensions of all facial features the same. Then, each facial feature is input into an attention model to obtain a corresponding attention map of the facial feature. Here, the attention model includes a convolution operator with a convolution kernel size of 1x1 and a channel count of 1, and a sigmoid function. For example, if the dimension of one facial feature is (28,28,64), the facial feature is first subjected to two convolution operations with a convolution kernel size of 3x3, a step size of 2, and a channel count of 64 to reduce the dimension to (7,7,64), and then the 1x1 convolution operation and the sigmoid operation are introduced to obtain the corresponding attention map as an attention map with a dimension of (7,7,1). Then, the three dimensions are reduced to one dimension to obtain the final attention map with a dimension of (7,7,1).
[0048] Step c: Obtain a full-space attention map based on the corresponding attention map of the face pose features, the corresponding attention map of the face features and the N face pose features.
[0049] In this embodiment, first, a second-stream attention map can be generated based on the corresponding attention map of the facial pose feature and the corresponding attention map of the facial feature. For example, based on the corresponding attention map of the i-th facial pose feature and the corresponding attention map of the i-th facial feature, the i-th second-stream attention map can be generated. Specifically, the i-th second-stream attention map D is calculated by the following formula: i can be calculated, and D i =reshape([A,B] T W1), where A is the corresponding attention map of facial features, B is the corresponding attention map of facial pose features, W1 is the learning parameter matrix, whose matrix dimensions are (7x7x2, qx7x7), and each row of W1 represents the association value between each point and other points in space. Since such associations can take various forms, we introduce q types of association learning, and reshape() is a function that converts a matrix into a matrix of a specific dimension.
[0050] Then, a dark and shallow attention map can be generated based on the corresponding attention maps of the facial features. A dark and shallow attention map can be generated based on the corresponding attention maps of all facial features. Specifically, the dark and shallow attention map E can be calculated by the following formula: E=reshape([A1,A2,...,A N ] T W2), where reshape() is a function that transforms a matrix into a matrix of a certain dimension, A is the corresponding attention map of the facial features, N is the number of facial features, W2 is the learning parameter matrix, and W 2の The matrix dimensions are (7*7*2, q'*q), and the parameter q' may be set to 4, where q' also plays the role of a partition, and the dimensions of E are (q', q).
[0051] Next, based on the two-stream attention map, the dark and shallow attention map, and the N facial pose features, a full-space attention map is obtained. For example, the full-space attention map R can be calculated by the following formula: R=(E T D1+E T D2+…+E T D N ) T P N where E is the dark and shallow attention map and D i is the i-th second-stream attention map, 1≦i≦N, and P N is the pose feature of the Nth face.
[0052] Step d: Based on the full-space attention map, generate face pose information for the face image to be recognized.
[0053] In this embodiment, the facial pose information of the facial image to be recognized and the score vector corresponding to the facial pose information can be generated based on the full space attention map. The score vector corresponding to the facial pose information is used to evaluate the richness of the overall facial information (identification information contained in the facial photo). As can be understood, the higher the score vector, the richer the overall facial information, and conversely, the lower the score vector, the less rich the overall facial information. Specifically, the facial pose information O of the facial image to be recognized can be calculated using the following formula: O=sigmoid((relu(R T W3)) T W4), where R is the full-space attention map, W3 and W4 are parameter matrices, and the dimensions of W3 and W4 are (256, 128) and (128, 1), respectively.
[0054] As can be seen, various types of full-space attention information are combined into the full-space attention map, and deep pose and shallow pose (i.e., dark and shallow attention map) and overall face feature information (i.e., Nth face pose feature) are also combined, so the facial pose information of the face image to be recognized generated based on the full-space attention map and the corresponding score vector of the facial pose information are more accurate.
[0055] Correspondingly, when the face pose information satisfies a predetermined condition, the step of identifying a user recognition result of the face image to be recognized based on the target face features of the face image to be recognized may include:
[0056] When the facial pose information and the score vector corresponding to the facial pose information satisfy a predetermined condition, a user recognition result for the facial image to be recognized is specified based on the target facial features of the facial image to be recognized.
[0057] Here, the predetermined conditions may be that the yaw angle is equal to or less than a predetermined yaw angle threshold, the pitch angle is equal to or less than a predetermined pitch angle threshold, the roll angle is equal to or less than a predetermined roll angle threshold, and the score vector corresponding to the facial pose information is equal to or greater than a predetermined score. Note that if the score vector corresponding to the facial pose information is smaller than a predetermined score, the face becomes more vulnerable to attack (for example, it becomes easier to decrypt using a photo of a user who is being impersonated), so the score vector corresponding to the facial pose information must be equal to or greater than a predetermined score.
[0058] It should be noted that the above embodiment can be applied to a face recognition model, where the training process of the face recognition model is described as follows.
[0059] In this embodiment, N class centers (one positive class center and N-1 negative class centers), that is, N vectors, may be introduced. These vectors and the face vector T (after normalization) are point-multiplied to obtain N values, which represent the similarity x between the current face and the N class centers. A common training method is to perform a softmax operation on these N x values and then calculate the cross-entropy. However, such a decision boundary is not accurate, and the training method is also inefficient. Therefore, in this embodiment, the similarity between T and the positive class center is x. + and x + needs to be subtracted from the decision value related to the pose,evaluation value. Subtracting a larger decision value means that the decision boundary of the facial feature becomes smaller.,Here, if we set the yaw angle as y, the pitch angle as p, the roll angle as r, the evaluation value as s, and the decision value as m, the general formula is, x + =x + -m, JPEG0007777229000001.jpg49170As above, f is a function, and m0 and m1 are hyperparameters, for example, m0=0.2, m1=0.2.
[0060] The above formula also includes a value i that determines which faces are currently given larger decision values and smaller decision boundaries. Generally, the larger m, the tighter the decision boundaries, the closer the facial features are to the class center, and the larger the gradients generated. When i is set to 0, the formula assigns larger decision values, smaller decision boundaries, and larger gradients to faces with very small yaw, pitch, and roll angles and very high evaluation values. As i gradually increases, the formula assigns larger decision values, smaller decision boundaries, and larger gradients to faces with relatively large yaw, pitch, and roll angles and relatively low evaluation values. That is, when i takes the value 0, the network primarily trains on forward-facing faces and provides gradients for forward-facing faces. As i gradually increases, the network primarily trains on profile faces and provides gradients for profile faces.
[0061] Technical solution for determining training for face recognition models: When the model starts training, set i to 0. After that, the model loss gradually decreases. If the model's accuracy gradually improves on the validation set, i can be gradually increased by 0.1 each time, such as from 0 to 0.1, from 0.1 to 0.2, etc. Continue to observe the accuracy on the validation set. If a decrease in accuracy is observed, decrease i by 0.1, and after a certain period of training, return i to its original value. Repeat this process three times to the original value. If the accuracy does not continue to improve, decrease i by 0.1 and terminate training after fitting. In this case, the resulting spatial distribution of faces will show an appropriate distribution, with small angles near the class center and large angles at the class edge.
[0062] Due to the reasonable spatial distribution, the inference process can obtain better feature representations of faces in large poses, and the three pose angles and evaluation values can be directly obtained, which makes face comparison more flexible and allows the comparison to be stopped if the pose angles of the two photos in face comparison are too large.
[0063] All the above-mentioned optional technical solutions may be arbitrarily combined to form optional embodiments of the present disclosure, and will not be described one by one here.
[0064] The following are apparatus embodiments of the present disclosure for carrying out the method embodiments of the present disclosure. For details not disclosed in the apparatus embodiments of the present disclosure, please refer to the method embodiments of the present disclosure.
[0065] 3 is a schematic diagram of a face recognition device provided in an embodiment of the present disclosure. As shown in FIG. 3, the face recognition device includes: an image acquisition unit 301 for acquiring a face image to be recognized; a feature extraction unit 302 for extracting facial pose features and facial features of the face image to be recognized; an information generating unit 303 for generating target facial features, facial pose information, and corresponding score vectors of the facial pose information of the face image to be recognized according to the facial pose features and facial features; and a result determination unit 304 for determining a user recognition result for the face image to be recognized based on the target face features of the face image to be recognized when the face pose information and the corresponding score vector of the face pose information satisfy a predetermined condition.
[0066] Preferably, the facial pose features comprise N facial pose features, and the facial features comprise N facial features, and the feature extraction unit 302 Regarding the first facial feature, a facial image to be recognized is input to a first convolutional layer to acquire the first facial feature; Regarding a first face pose feature, inputting a face image to be recognized into a first residual block to obtain the first face pose feature; For the pose feature of the ith face, input the pose feature of the (i-1)th face into the ith residual block to obtain the pose feature of the ith face; For the i-th facial feature, the i-1st face pose feature and the i-1st facial feature are input into the i-th convolutional layer to obtain the i-th facial feature, where i is greater than or equal to 2 and less than or equal to N, and N and i are positive integers.
[0067] Preferably, the information generating unit 303 comprises: The pose features and N face features of the Nth face are input to the N+1th convolutional layer to obtain the target face features of the face image to be recognized.
[0068] Preferably, the information generating unit 303 comprises: and generating a corresponding attention map of the facial pose features based on the facial pose features; for generating a corresponding attention map of the facial features based on the facial features; and obtaining a full-space attention map based on the corresponding attention map of the facial pose features, the corresponding attention map of the facial features, and the N facial pose features; This is to generate facial pose information of a face image to be recognized based on the full-space attention map.
[0069] Preferably, the information generating unit 303 specifically: generating a two-stream attention map based on the corresponding attention map of facial pose features and the corresponding attention map of facial features; for generating dark and shallow attention maps based on corresponding attention maps of facial features; Based on the two-stream attention map, the dark and shallow attention map and the N facial pose features, a full-space attention map is obtained.
[0070] Preferably, the information generating unit 303 specifically: generating facial pose information of a face image to be recognized and a score vector corresponding to the facial pose information based on the full-space attention map; Correspondingly, the result determination unit 304: When the facial pose information and the score vector corresponding to the facial pose information satisfy a predetermined condition, the user recognition result of the facial image to be recognized is specified based on the target facial features of the facial image to be recognized.
[0071] Preferably, the facial pose information includes a yaw angle, a pitch angle, and a roll angle, and the predetermined conditions are that the yaw angle is equal to or less than a predetermined yaw angle threshold, the pitch angle is equal to or less than a predetermined pitch angle threshold, the roll angle is equal to or less than a predetermined roll angle threshold, and the score vector corresponding to the facial pose information is equal to or greater than a predetermined score.
[0072] Preferably, the result determination unit 304 comprises: to identify the similarity between target facial features of a facial image to be recognized and each of the predetermined user facial features; This is to determine the user information corresponding to the predetermined user facial feature with the greatest similarity as the facial image user recognition result to be recognized.
[0073] The technical solution provided in the disclosed embodiments is a face recognition device including an image acquisition unit for acquiring a face image to be recognized, a feature extraction unit for extracting facial pose features and facial features of the face image to be recognized, an information generation unit for generating target face features, facial pose information, and a corresponding score vector of the facial pose information of the face image to be recognized based on the facial pose features and facial features, and a result determination unit for determining a user recognition result of the face image to be recognized based on the target face features of the face image to be recognized when the facial pose information and the corresponding score vector of the facial pose information satisfy a predetermined condition. In this embodiment, the facial pose features reflect the spatial information of the face and include facial structure features, edge features, and angle features, and the facial features reflect features such as facial skin color and texture. Therefore, by fusing the facial pose features and facial features, the detailed information of the facial pose features and facial features can be fused, and the target facial features and facial pose information of the face image to be recognized can be generated based on the fused features. Therefore, the accuracy of the identified target facial features and facial pose information can be improved, and the accuracy of the user recognition results identified based on the target facial features and facial pose information can be improved.
[0074] It should be understood that the magnitude of the numbers of each step in the above embodiments does not mean the order of execution, and the execution order of each process should be determined by its function and inherent logic, and does not arbitrarily limit the implementation process of the embodiments disclosed herein.
[0075] 4 is a schematic diagram of a computer device 4 provided in an embodiment of the present disclosure. As shown in FIG. 4, the computer device 4 of the embodiment includes a processor 401, a memory 402, and a computer program 403 stored in the memory 402 and executable by the processor 401. When the processor 401 executes the computer program 403, it implements the steps in each of the above method embodiments. Alternatively, when the processor 401 executes the computer program 403, it implements the functions of each module / unit in each of the above device embodiments.
[0076] For example, the computer program 403 may be divided into one or more modules / units, and one or more modules / units may be stored in the memory 402 and executed by the processor 401 to accomplish the present disclosure. One or more modules / units may be a series of computer program command sections capable of performing specific functions, and the command sections are for explaining the process of the computer program 403 being executed on the computer device 4.
[0077] The computing device 4 may be a computing device such as a desktop computer, a laptop computer, a palmtop computer, a cloud server, etc. The computing device 4 may include, but is not limited to, a processor 401 and a memory 402. As will be appreciated by those skilled in the art, FIG. 4 is merely an example of the computing device 4 and is not intended to limit the computing device 4, which may include more or fewer components than those shown, or may combine certain components or different components; for example, the computing device may include input / output devices, network access devices, buses, etc.
[0078] Processor 401 may be a central processing unit (CPU), or may be another general-purpose processor, a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic device, a discrete gate or transistor logic device, a discrete hardware component, etc. The general-purpose processor may be a microprocessor, any common processor, etc.
[0079] The memory 402 may be an internal storage unit of the computer device 4, such as a hard disk or memory of the computer device 4. The memory 402 may also be an external storage device of the computer device 4, such as a plug-in hard disk, a Smart Media Card (SMC), a Secure Digital (SD) card, a Flash Card, etc., that is installed in the computer device 4. Furthermore, the memory 402 may comprise both an internal storage unit of the computer device 4 and an external storage device. The memory 402 is for storing computer programs and other programs and data required by the computer device. The memory 402 may also be used to temporarily store data that has been output or that is to be output.
[0080] Those skilled in the art will clearly understand that, for convenience and brevity, only the division of the above functional units and modules has been used as an example. However, in actual applications, the above functions can be assigned to different functional units or modules as needed, i.e., all or part of the above-described functions can be achieved by dividing the internal structure of the device into different functional units or modules. The functional units and modules in the embodiments may be integrated into a single processing unit, each unit may exist physically independently, or two or more units may be integrated into a single unit. The integrated unit may be implemented in the form of hardware or a software functional unit. The specific names of the functional units and modules are provided solely for the purpose of distinguishing them from one another and do not limit the scope of protection of the present disclosure. For the specific operating processes of the units and modules in the above system, reference may be made to the corresponding processes in the above-described method embodiments, and further description thereof will be omitted.
[0081] In the above embodiments, the description of each embodiment has its own emphasis, and for the details or parts not described in an embodiment, reference can be made to the relevant descriptions of other embodiments.
[0082] Those skilled in the art can recognize that the units and algorithm steps of each example described in the embodiments disclosed herein can be realized by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed by hardware or software is determined by the specific application and design constraints of the technical solution. Those skilled in the art may use different methods to realize the described functions for each specific application, but such realization should not be considered beyond the scope of this disclosure.
[0083] It should be understood that the disclosed devices / computer devices and methods in the embodiments provided in this disclosure can be implemented in other ways. For example, the device / computer device embodiments described above are merely illustrative, and the division into modules or units is merely a logical division of functions. In actual implementation, other division methods may be used. For example, multiple units or components may be combined or integrated into other systems, or some features may be omitted or not implemented. Furthermore, the couplings or direct couplings or communication connections shown or discussed may be indirect couplings or communication connections through several interfaces, devices, or units, and may be electrical, mechanical, or other types.
[0084] Units described as separate components may or may not be physically separated, and components shown as units may or may not be physical units, i.e., located in one location or distributed across multiple network units, some or all of which may be selected according to actual needs to achieve the objectives of the solutions of this embodiment.
[0085] Note that the functional units in this disclosure may be integrated into one processing unit, each unit may exist physically independently, or two or more units may be integrated into one unit. The integrated unit may be realized in the form of hardware or in the form of a software functional unit.
[0086] The integrated module / unit may be realized in the form of a software functional unit and stored in a computer-readable storage medium when sold or used as an independent product. Based on this understanding, the present disclosure provides that the realization of all or part of the processes in the above-described method embodiments can be accomplished by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when the computer program is executed by a processor, the steps of the above-described method embodiments can be realized. The computer program may include computer program code, which may be in source code format, object code format, an executable file, or some intermediate format. The computer-readable storage medium may include any entity or device capable of carrying computer program code, such as a recording medium, a U-disk, a removable hard disk, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory (RAM), an electrical carrier wave signal, an electrical communication signal, and a software distribution medium. Furthermore, the content contained on a computer-readable storage medium may be increased or decreased as required by the legislation and patent practice of a jurisdiction. For example, in some jurisdictions, the legislation and patent practice may require that a computer-readable storage medium not include electrical carrier signals and telecommunications signals.
[0087] The above-mentioned embodiments are only for illustrating the technical solutions of the present disclosure, and are not intended to limit the same. Although the present disclosure has been described in detail with reference to the above-mentioned embodiments, those skilled in the art may still amend the technical solutions described in the above-mentioned embodiments or equivalently replace some technical features thereof, and it should be understood that such amendments or replacements will not deviate from the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present disclosure, and all of them should be included in the protection scope of the present disclosure.
Claims
1. A face recognition method, comprising: A step of acquiring a face image to be recognized; extracting facial pose features and facial features from the face image to be recognized; generating target facial features and facial pose information of the face image to be recognized based on the facial pose features and the facial features; When the face pose information satisfies a predetermined condition, specifying a user recognition result for the face image to be recognized based on target face features of the face image to be recognized; Including, The facial pose features include N facial pose features, and the facial features include N facial features, and the step of extracting the facial pose features and facial features of the face image to be recognized includes: Regarding a first facial feature, inputting the facial image to be recognized into a first convolutional layer to acquire the first facial feature; Regarding a first face pose feature, inputting the face image to be recognized into a first residual block to obtain a first face pose feature; For the pose feature of the i-th face, input the pose feature of the (i-1)-th face into the i-th residual block to obtain the pose feature of the i-th face; For the i-th facial feature, input the i-1-th face pose feature and the i-1-th facial feature into the i-th convolutional layer to obtain the i-th facial feature, where i is greater than or equal to 2 and less than or equal to N, and N and i are positive integers. A method characterized by:
2. generating target facial features and facial pose information of the face image to be recognized based on the facial pose features and the facial features, inputting the pose feature of the Nth face and the Nth face feature into the (N+1)th convolutional layer, and acquiring target face features of the face image to be recognized; 2. The method of claim 1 .
3. The step of generating facial pose information of the face image to be recognized based on the facial pose features and the facial features includes: generating a corresponding attention map of the facial pose features based on the facial pose features; generating a corresponding attention map of the facial features based on the facial features; obtaining a full space attention map based on the corresponding attention map of the facial pose features, the corresponding attention map of the facial features, and N facial pose features; generating facial pose information of the facial image to be recognized based on the full-space attention map.
2. The method of claim 1 .
4. The step of obtaining a full space attention map based on the corresponding attention map of the facial pose features, the corresponding attention map of the facial features, and N facial pose features includes: generating a two-stream attention map based on the corresponding attention map of the facial pose features and the corresponding attention map of the facial features; generating a dark and shallow attention map based on the corresponding attention map of the facial features; and obtaining a full-space attention map based on the two-stream attention map, the dark and shallow attention map, and the N facial pose features.
4. The method of claim 3.
5. The step of generating facial pose information of the facial image to be recognized based on the full-space attention map includes: generating facial pose information of the face image to be recognized and a score vector corresponding to the facial pose information based on the full-space attention map; Correspondingly, when the face pose information satisfies a predetermined condition, the step of identifying a user recognition result of the face image to be recognized based on target face features of the face image to be recognized includes: and when the facial pose information and the score vector corresponding to the facial pose information satisfy a predetermined condition, specifying a user recognition result for the facial image to be recognized based on target facial features of the facial image to be recognized.
4. The method of claim 3.
6. The facial pose information includes a yaw angle, a pitch angle, and a roll angle, and the predetermined condition is that the yaw angle is equal to or less than a predetermined yaw angle threshold, the pitch angle is equal to or less than a predetermined pitch angle threshold, the roll angle is equal to or less than a predetermined roll angle threshold, and the score vector corresponding to the facial pose information is equal to or greater than a predetermined score.
6. The method of claim 5.
7. A face recognition method, comprising: A step of acquiring a face image to be recognized; extracting facial pose features and facial features from the face image to be recognized; generating target facial features and facial pose information of the face image to be recognized based on the facial pose features and the facial features; When the face pose information satisfies a predetermined condition, specifying a user recognition result for the face image to be recognized based on target face features of the face image to be recognized; Including, The step of identifying a user recognition result of the face image to be recognized based on features of a target face of the face image to be recognized, Identifying a similarity between a target facial feature of the facial image to be recognized and each of the predetermined user facial features; and determining the user information corresponding to the predetermined user facial feature having the greatest similarity as a facial image user recognition result to be recognized. A method characterized by:
8. A face recognition device, The face recognition device includes an image acquisition unit for acquiring a face image to be recognized, and a feature extraction unit for extracting face pose features and face features of the face image to be recognized, an information generating unit for generating target facial features, facial pose information, and a score vector corresponding to the facial pose information of the face image to be recognized according to the facial pose features and the facial features; a result determination unit for determining a user recognition result of the face image to be recognized based on target face features of the face image to be recognized when the face pose information and the score vector corresponding to the face pose information satisfy a predetermined condition; The facial pose features comprise N facial pose features, and the facial features comprise N facial features, and the feature extraction unit Regarding a first facial feature, the facial image to be recognized is input to a first convolutional layer to acquire the first facial feature; Regarding a first face pose feature, input the face image to be recognized into a first residual block to obtain a first face pose feature; For the pose feature of the i-th face, input the pose feature of the (i-1)-th face into the i-th residual block to obtain the pose feature of the i-th face; For the i-th face feature, input the i-1-th face pose feature and the i-1-th face feature into the i-th convolutional layer to obtain the i-th face feature, where i is greater than or equal to 2 and less than or equal to N, and N and i are positive integers. A face recognition device characterized by:
9. 10. A computing device comprising a memory, a processor, and a computer program stored in said memory and executable by said processor, said computer program implementing the steps of the method of claim 1 when said processor executes said computer program.
1. A computer device characterized by:
10. A computer-readable storage medium having a computer program stored thereon, the computer program performing the steps of the method of claim 1 when executed by a processor. A computer-readable storage medium comprising:
Citation Information
Patent Citations
Face recognition method and device, computer equipment and storage medium
CN112001932A
Face authentication unit, face authentication method, and entrance / exit management device
JP2007148988A
Attention detecting device and attention detecting method
JP2017068815A
Driving condition monitoring method and device, driver monitoring system, and vehicle
JP2019536673A
Head wearing device and control device for the same
JP2021159594A