Face recognition method and device, equipment and storage medium
By extracting and fusing face angle features, using the multi-head attention mechanism and feature redirection network to generate face recognition results, the problem of face recognition algorithm being sensitive to angle changes is solved and the recognition accuracy is improved.
Patent Information
- Application Number
- CN202510734622.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-04
- Publication Date
- 2025-07-04
- Estimated Expiration
- 2045-06-04
AI Technical Summary
In the prior art, face recognition algorithms are sensitive to changes in face angles, resulting in a decrease in recognition accuracy.
By extracting the target face features and face angle features of the image to be identified, combining the target face features in the pre-registered face features, the multi-head attention mechanism and feature redirection network are used to generate face recognition results.
It improves the adaptability and accuracy of face recognition, and can accurately recognize in multiple face angle scenarios.
Smart Images

Figure CN120260102A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular, to a face recognition method, device, equipment, and storage medium. Background Art
[0002] Face recognition is a biometric identification technology that identifies a person's identity based on the person's facial feature information. Its basic principle includes three main steps: face detection, feature extraction, and matching recognition.
[0003] However, in the process of feature extraction, the face feature extraction algorithm is sensitive to changes in the face angle. When the face in the image is at a non-frontal angle, the accuracy of face feature extraction will decrease significantly, resulting in a decrease in face recognition accuracy. Summary of the Invention
[0004] The purpose of this application is to provide a face recognition method, device, equipment, and storage medium, which can adapt to face recognition scenarios with multiple face angles and improve face recognition accuracy.
[0005] An embodiment of this application provides a face recognition method, including: Obtain an image to be recognized; Extract the target face features in the image to be recognized to obtain the face features to be recognized; Extract the face angle features in the image to be recognized to obtain the face angle features to be recognized; Extract the target face features corresponding to the face angle features to be recognized from the pre-registered face features to obtain the face features to be matched; the pre-registered face features include the target face features in multiple groups of pre-registered images collected from the same face at multiple angles; Generate a face recognition result of the image to be recognized based on the similarity between the face features to be recognized and the face features to be matched.
[0006] In some embodiments, the face recognition method further includes: Obtain multiple groups of pre-registered images corresponding to several faces; Extract the target face features in the pre-registered images; Fuse the target face features in the pre-registered images to obtain the pre-registered face features.
[0007] In some embodiments, the fusing of the target face features in each of the pre-registered images includes: Stitch the target face features in each of the pre-registered images to obtain a first stitched feature; Perform non-linear activation processing on the first stitched feature to obtain a first activation feature; Perform multi-head attention mechanism processing on the first activation feature to obtain an attention weight feature; Superimpose the first activation feature and the attention weight feature to obtain the pre-registered face feature.
[0008] In some embodiments, the extracting of the face angle feature in the image to be recognized includes: Extract the yaw angle feature and the pitch angle feature of the face in the image to be recognized; Perform angle encoding processing on the yaw angle feature and the pitch angle feature to obtain the face angle feature to be recognized.
[0009] In some embodiments, the performing of the angle encoding processing on the yaw angle feature and the pitch angle feature includes: Concatenate the yaw angle feature and the pitch angle feature to obtain a second concatenated feature; Map the second concatenated feature to a preset feature space to obtain the face angle feature to be recognized.
[0010] In some embodiments, the extracting of the target face feature corresponding to the face angle feature to be recognized in the pre-registered face feature includes: According to the mapping relationship between the face angle feature to be recognized and the target face feature in the pre-registered face feature, perform feature redirection on the pre-registered face feature to extract the corresponding target face feature in the pre-registered face feature, and obtain the face feature to be matched.
[0011] In some embodiments, the performing of the feature redirection on the pre-registered face feature according to the mapping relationship between the face angle feature to be recognized and the target face feature in the pre-registered face feature includes: Concatenate the pre-registered face feature and the face angle feature to be recognized to obtain a third concatenated feature; Perform probabilistic activation processing on the third concatenated feature to obtain a second activation feature; Extract the corresponding target face feature in the pre-registered face feature according to the second activation feature to obtain the face feature to be matched.
[0012] An embodiment of the present application further provides a face recognition device, including: A first module, configured to obtain an image to be recognized; A second module, configured to extract the target face feature in the image to be recognized to obtain a face feature to be recognized; A third module, configured to extract the face angle feature in the image to be recognized to obtain a face angle feature to be recognized; The fourth module is configured to extract the target face feature corresponding to the face angle feature to be recognized from the pre-registered face features, so as to obtain the face feature to be matched; the pre-registered face features include the target face features in multiple groups of pre-registered images obtained by collecting the same face from multiple angles; The fifth module is configured to generate a face recognition result of the image to be recognized according to the similarity between the face feature to be recognized and the face feature to be matched.
[0013] An embodiment of the present application further provides an electronic device, which includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the above-mentioned face recognition method is implemented.
[0014] An embodiment of the present application further provides a computer-readable storage medium, which stores a computer program, and when the computer program is executed by a processor, the above-mentioned face recognition method is implemented.
[0015] The beneficial effects of the present application: By performing face feature extraction and face angle feature extraction on the image to be recognized, and extracting the corresponding target face feature in the pre-registered face features according to the face angle feature to be recognized obtained by extraction, the face feature to be matched is obtained, and then the face recognition result of the image to be recognized is generated according to the similarity between the face feature to be recognized and the face feature to be matched. Since the face feature to be matched extracted from the pre-registered face features is the target face feature with a face angle feature matching the face angle feature to be recognized, it can adapt to the face recognition scenarios of multiple face angles and improve the face recognition accuracy. Description of the Drawings
[0016] Figure 1 It is an application environment diagram of the face recognition method provided by an embodiment of the present application.
[0017] Figure 2 It is a flowchart of the face recognition method provided by an embodiment of the present application.
[0018] Figure 3 It is a schematic diagram of the face angle feature provided by an embodiment of the present application.
[0019] Figure 4 It is a flowchart of the method before step S201 provided by an embodiment of the present application.
[0020] Figure 5 It is a schematic structural diagram of the face recognition device provided by an embodiment of the present application.
[0021] Figure 6 It is a schematic hardware structure diagram of the electronic device provided by an embodiment of the present application. Detailed Embodiments
[0022] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0023] It should be noted that although the functional modules are divided in the device schematic diagram and the logical sequence is shown in the flowchart, in some cases, the steps shown can be executed in a different order from the module division in the device or the sequence in the flowchart. Terms such as "first" and "second" in the description, claims and drawings are used to distinguish similar objects and are not used to describe a specific order or sequence.
[0024] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.
[0025] The face recognition method provided by the embodiments of this application can be executed by a computer device, which can be a terminal device or a server. Among them, the terminal device includes but is not limited to mobile phones, computers, intelligent voice interaction devices, intelligent home appliances, vehicle-mounted terminals, aircraft, etc. The server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server.
[0026] In addition, the information, data and signals involved in the embodiments of this application are all authorized by the relevant objects or fully authorized by all parties, and the collection, use and processing of the relevant data all comply with the relevant laws, regulations and standards of the relevant countries and regions.
[0027] To facilitate the understanding of the face recognition method provided by the embodiments of this application, the application scenario of the face recognition method will be exemplarily introduced below taking the execution entity of the face recognition method as a server as an example.
[0028] Figure 1 is the application environment diagram of the face recognition method provided by the embodiments of this application. Refer to Figure 1, the face recognition method is applied to a face recognition system. The face recognition system includes a terminal 110 and a server 120. The terminal 110 and the server 120 are connected through a network. The terminal 110 can be a desktop terminal or a mobile terminal, and the mobile terminal can be at least one of a mobile phone, a tablet computer, a laptop computer, etc. The server 120 can be implemented by an independent server or a server cluster composed of multiple servers. The terminal 110 is used to send an image to be recognized to the server 120. The server 120 is used to obtain the image to be recognized, extract the target face features in the image to be recognized to obtain the face features to be recognized, extract the face angle features in the image to be recognized to obtain the face angle features to be recognized, extract the target face features corresponding to the face angle features to be recognized in the pre-registered face features to obtain the face features to be matched, and generate a face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched. Among them, the pre-registered face features include the target face features in multiple groups of pre-registered images obtained by collecting the same face from multiple angles.
[0029] It should be understood that Figure 1 The application scenarios shown are only examples. In actual applications, the face recognition method provided in the embodiments of the present application can also be applied to other scenarios. For example, the above face recognition method can be directly applied to the terminal 110. The terminal 110 is used to obtain the image to be recognized, extract the target face features in the image to be recognized to obtain the face features to be recognized, extract the face angle features in the image to be recognized to obtain the face angle features to be recognized, extract the target face features corresponding to the face angle features to be recognized in the pre-registered face features to obtain the face features to be matched, and generate a face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched. Among them, the pre-registered face features include the target face features in multiple groups of pre-registered images obtained by collecting the same face from multiple angles.
[0030] Figure 2 is a flowchart of a face recognition method provided in an embodiment of the present application. Refer to Figure 2 , in some embodiments, the method includes but is not limited to steps S201 to S205.
[0031] Step S201, obtain an image to be recognized.
[0032] The image to be recognized refers to an image containing a face that needs to be authenticated, and can be obtained by collecting through an image acquisition device or obtained from an image storage device.
[0033] Step S202, extract the target face features in the image to be recognized to obtain the face features to be recognized.
[0034] The face features to be recognized refer to the feature vectors of face features with identity distinctiveness extracted from the image to be recognized, such as the 512-dimensional feature vectors extracted by a deep convolutional network. The target face features refer to the high-level semantic feature vectors such as face key points, textures, or geometric structures extracted by a deep convolutional neural network. Specifically, models such as ResNet or ArcFace can be used to implement feature encoding.
[0035] After obtaining the image to be recognized, it is possible to use a pre-trained face feature extraction network to extract the target face features in the image to be recognized, so as to obtain the face features to be recognized. The face feature extraction network can be a network constructed based on the deep convolutional network architecture. The pre-trained face feature extraction network focuses on perceiving the face features in the image. By performing continuous convolution operations on the image to be recognized, the target face features in the image to be recognized are extracted layer by layer, and finally the face features to be recognized are output.
[0036] Step S203: Extract the face angle features in the image to be recognized to obtain the face angle features to be recognized.
[0037] The face angle features to be recognized refer to the feature vectors reflecting the orientation angles of the face in the image to be recognized, and can be obtained by the pose parameters of the face in the image to be recognized through a three-dimensional face key point detection algorithm. The face angle features to be recognized are used to locate the feature data corresponding to the corresponding perspective in the pre-registered feature library. For example, when it is detected that the face has a left deflection, the features with a left angle in the pre-registered library are preferentially selected for comparison. Refer to Figure 3 , and the orientation angles include the roll angle (Roll), yaw angle (Yaw), and pitch angle (Pitch) of the face relative to the corresponding coordinate axes of the image.
[0038] After obtaining the image to be recognized, it is possible to use a pre-trained angle feature extraction network to extract the face angle features in the image to be recognized, so as to obtain the face angle features to be recognized. The angle feature extraction network can be a network constructed based on the deep convolutional network architecture. The pre-trained angle feature extraction network focuses on perceiving the face angle features in the image. By performing continuous convolution operations on the image to be recognized, the target face angle features in the image to be recognized are extracted layer by layer, and finally the face angle features to be recognized are output.
[0039] Step S204: Extract the target face features corresponding to the face angle features to be recognized in the pre-registered face features to obtain the face features to be matched.
[0040] The pre-registered face features include the target face features in multiple groups of pre-registered images obtained by collecting the same face from multiple angles. It can be understood that the pre-registered face features refer to the feature set constructed by the target face features in multiple groups of pre-registered images obtained by pre-collecting the same face, such as the target face features in multiple groups of pre-registered images obtained by collecting the same face from the front, left, and right, etc. The pre-registered face features are collected and pre-stored in the memory space of the execution entity before face recognition of the image to be recognized. There are multiple groups of pre-registered face features pre-stored in the memory space of the execution entity, and different pre-registered face features include the pre-registered images obtained by collecting different faces. Exemplarily, each group of pre-registered face features includes the target face features in five groups of pre-registered images, and the orientation angles of the faces in each group of pre-registered images are α≈0° and β≈0°, α∈[30°, 60°] and β≈0°, α∈[-60°, -30°] and β≈0°, α≈0° and β∈[-30°, -10°], and α≈0° and β∈[10°, 30°] respectively, where α is the yaw angle of the face and β is the pitch angle of the face.
[0041] After obtaining the angle features of the face to be recognized, it can be to use a pre-trained feature redirection network to extract the target face features corresponding to the angle features of the face to be recognized in the pre-registered face features to obtain the face features to be matched. The feature redirection network can be a deep neural network based on the attention mechanism. The pre-trained feature redirection network focuses on perceiving the mapping relationship between the angle features of the face to be recognized and the target face features in the pre-registered face features, determines the target face features corresponding to the angle features of the face to be recognized in the pre-registered face features based on the weight calculation or value transformation in the attention weight operation, and extracts them, and finally outputs the face features to be matched.
[0042] Step S205, generate the face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched.
[0043] The similarity between the face features to be recognized and the face features to be matched can be obtained by calculating the cosine similarity between the face features to be recognized and the face features to be matched.
[0044] After obtaining the face features to be matched, calculate the similarity between the face features to be recognized and the face features to be matched, and compare the calculated similarity with a preset similarity threshold. If it is greater than the similarity threshold, it indicates that the face features to be recognized and the face features to be matched are successfully matched, and a face recognition result of successful face recognition is generated. If it is not greater than the similarity threshold, it indicates that the face features to be recognized and the face features to be matched fail to match. Then, match the face features to be recognized with the next set of face features to be matched until it is greater than the similarity threshold or all the face features to be matched are traversed. When it is still not greater than the similarity threshold after traversing all the face features to be matched, a face recognition result of failed face recognition is generated. Compared with the prior art, the face recognition method provided by the embodiment of the present application extracts face features and face angle features from the image to be recognized, extracts corresponding target face features in the pre-registered face features according to the extracted face angle features to be recognized, obtains the face features to be matched, and then generates a face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched. Since the face features to be matched extracted from the pre-registered face features are target face features with face angle features matching the face angle features to be recognized, it can adapt to face recognition scenarios with multiple face angles and improve the face recognition accuracy.
[0045] Figure 4 It is a flowchart of the method before step S201 provided by the embodiment of the present application. Refer to Figure 4 , in some embodiments, the method includes but is not limited to steps S401 to S403.
[0046] Step S401, obtain multiple groups of pre-registered images corresponding to several faces.
[0047] Step S402, extract the target face features in the pre-registered images.
[0048] Step S403, fuse the target face features in the pre-registered images to obtain pre-registered face features.
[0049] The multiple groups of pre-registered images refer to multiple images of the same face collected from different angles. For example, it can be image collection of the same face with yaw angles of ±30° and ±60° and pitch angles of ±15°, so as to cover the common face pose change range.
[0050] For each group of pre-registered images corresponding to a group of faces, after obtaining the multiple groups of pre-registered images corresponding to the same face, a pre-trained facial feature extraction network may be used to extract the target facial features in the pre-registered images to obtain the target facial features in the multiple groups of pre-registered images. The target facial features in the pre-registered images may be fused by a method combining splicing and nonlinear transformation to map the target facial features in the multiple groups of pre-registered images to a unified feature space to generate pre-registered facial features. Thus, by constructing a multi-angle feature fusion mechanism, a corresponding relationship between angle changes and feature space is established in the registration stage, so that the generated pre-registered facial features can adaptively match the facial features to be identified at different angles.
[0051] In a specific embodiment, the dimension of the target facial features in the pre-registered image is greater than the dimension of the facial features to be identified. The target facial features in the pre-registered image can be extracted by inputting the pre-registered image into a pre-trained facial feature extraction network, obtaining the feature vector output by the intermediate layer network of the facial feature extraction network without dimensionality reduction, and obtaining the target facial features in the pre-registered image. In this way, the target facial features in the pre-registered image with richer feature content can be obtained, and after the target facial features in the pre-registered image are integrated, the generated pre-registered facial features can adaptively match the facial features to be identified with multiple feature contents.
[0052] In some embodiments, step S403 specifically includes: splicing the target facial features in each pre-registered image to obtain a first spliced feature; performing linear encoding processing on the first spliced feature to obtain a first encoding feature; performing nonlinear activation processing on the first spliced feature to obtain a first activation feature; performing multi-head attention mechanism processing on the first activation feature to obtain an attention weight feature; superimposing the first activation feature and the attention weight feature to obtain a pre-registered facial feature.
[0053] Splicing refers to connecting multiple feature vectors in a preset order. Specifically, it can be achieved by connecting them in series along the feature dimension direction, and forming a unified representation by integrating features from different angles. Non-linear activation processing refers to the nonlinear transformation of linear features through activation functions. Specifically, it can be implemented by ReLU function or Sigmoid function, which is used to enhance the expressive power of features and explore potential correlations. Multi-head attention mechanism processing refers to splitting features into multiple subspaces and calculating attention weights separately. Specifically, it can be implemented by multiple parallel self-attention modules, and weight allocation is achieved by capturing the dynamic dependency between different feature vectors. Superposition refers to the addition of multiple feature vectors. Specifically, it can be implemented by residual connection structure, which is used to retain the original feature information and fuse the key features after attention screening.
[0054] When fusing the target face features of the pre-registration images of the same face in the pre-registration stage, first, the target face features in each pre-registration image are spliced to form a unified feature representation containing multi-angle information, that is, the first spliced feature. Subsequently, the first spliced feature is input into a pre-trained feature fusion network. In the feature fusion network, a non-linear activation process is applied to the first spliced feature, and the non-linear correlation between features is enhanced through an activation function to obtain the first activated feature. Then, a multi-head attention mechanism is used to perform parallel weight calculation on the first activated feature, analyze the importance of features from different angles in multiple subspaces, generate dynamic attention weights, and obtain the attention weight feature. Finally, the first activated feature and the attention weight feature are superimposed to obtain the pre-registered face feature. This makes the generated pre-registered face feature retain the details of the first activated feature and can also strengthen the contribution degree of key angle features.
[0055] In a specific embodiment, the feature fusion network is composed of two fully connected layers, a ReLU layer, a multi-head attention layer, and an Add layer. After the first spliced feature is input into the feature fusion network, the first fully connected layer maps the first spliced feature into the corresponding feature space and outputs the mapped feature to the ReLU layer. The mapped feature undergoes non-linear activation in the ReLU layer to obtain the first activated feature. Then, the first activated feature is input into the multi-head attention layer and the Add layer. The multi-head attention layer performs a multi-head attention mechanism process on the first activated feature to generate the attention weight feature, and the generated attention weight feature is input into the Add layer. Furthermore, the first activated feature and the attention weight feature are superimposed to generate the corresponding superimposed feature. Finally, the superimposed feature is mapped by the second fully connected layer to obtain the pre-registered face feature.
[0056] In some embodiments, step S203 specifically includes: extracting the yaw angle feature and the pitch angle feature of the face in the image to be recognized; performing angle encoding processing on the yaw angle feature and the pitch angle feature to obtain the face angle feature to be recognized.
[0057] The yaw angle feature refers to the horizontal direction pose parameter of the face rotating around the vertical axis. Specifically, it can be implemented by combining a face key point detection algorithm with a three-dimensional reconstruction model. The horizontal rotation angle is calculated by analyzing the coordinate offsets of the left and right eyes, the tip of the nose, and the corners of the mouth. The pitch angle feature refers to the vertical direction pose parameter of the face rotating around the horizontal axis. Specifically, it can be implemented by using a facial contour curvature analysis method. The pitch angle value is obtained by calculating the change rate of the vertical projection distance between the mandibular line and the midpoint of the eyebrows. Angle encoding processing refers to the process of converting discrete angle values into continuous feature vectors. Specifically, it can be implemented by using a multi-layer perceptron neural network. The angle values are mapped into a high-dimensional feature space to form encoded features with geometric preservation.
[0058] In the stage of extracting the angle features of the face to be recognized, the angle feature extraction network first locates the facial region in the image to be recognized, and then captures the horizontal rotation state and the vertical tilt state through a dual-branch feature extraction network. The yaw angle branch generates horizontal rotation parameters by tracking the symmetry changes between the bilateral cheekbones and the nose bridge, and the pitch angle branch outputs vertical tilt parameters by analyzing the projection ratio changes in the forehead and chin regions. The two types of angle parameters are input into an angle encoder for joint encoding. The encoding process can be to map the original angle values to the corresponding feature space using a non-linear transformation function to form a composite angle feature vector with rotational invariance. This vector is matched with the angle encoding in the pre-registered feature library through cosine similarity calculation to achieve cross-angle feature alignment.
[0059] In some embodiments, the angle encoding process for the yaw angle feature and the pitch angle feature specifically includes: splicing the yaw angle feature and the pitch angle feature to obtain a second spliced feature; mapping the second spliced feature to a preset feature space to obtain the angle feature of the face to be recognized.
[0060] When performing the angle encoding process on the yaw angle feature and the pitch angle feature, by vector splicing the yaw angle feature and the pitch angle feature, a composite feature expression containing three-dimensional space angle information is formed. This composite feature is transformed through the non-linear transformation of the preset feature space to convert the discrete angle parameters into continuous embedding vectors, making the features of different angles comparable in a unified space. An activation function is used in the feature space mapping process to enhance the discriminability of the angle features, enabling the accurate retrieval of the target face features corresponding to the pre-registered face features at the corresponding angles in the pre-registered feature library in the subsequent matching stage.
[0061] In a specific embodiment, the angle encoder is composed of two fully connected layers. After obtaining the yaw angle feature and the pitch angle feature of the face in the image to be recognized, the obtained yaw angle feature and pitch angle feature are spliced, and then continuous fully connected processing is performed on the spliced angle features to map the spliced angle features to the corresponding feature space to obtain the angle feature of the face to be recognized.
[0062] In some embodiments, step S204 specifically includes: performing feature redirection on the pre-registered face features according to the mapping relationship between the angle feature of the face to be recognized and the target face features in the pre-registered face features, so as to extract the corresponding target face features in the pre-registered face features to obtain the face features to be matched.
[0063] Feature redirection refers to dynamically screening the effective target face features corresponding to the current angle in the pre-registered face features according to the current angle features, which can be specifically implemented by combining an attention mechanism with a feature projection matrix, and completing feature alignment by establishing a non-linear mapping between the angle feature space and the registration feature space.
[0064] When there is an angular deviation in the face to be recognized, input the angular feature of the face to be recognized into the pre-trained feature redirection network to generate a corresponding feature space projection matrix. Perform a tensor operation on the target face feature in the pre-registered face feature and this projection matrix, and screen out the target face feature that best matches the current angle in the pre-registered face feature by calculating the feature correlation score to obtain the face feature to be matched.
[0065] In some embodiments, feature redirection of the pre-registered face feature according to the mapping relationship between the angular feature of the face to be recognized and the target face feature in the pre-registered face feature specifically includes: concatenating the pre-registered face feature and the angular feature of the face to be recognized to obtain a third concatenated feature; performing a probabilistic activation process on the third concatenated feature to obtain a second activation feature; extracting the corresponding target face feature in the pre-registered face feature according to the second activation feature to obtain the face feature to be matched.
[0066] First, concatenate the pre-registered face feature and the angular feature of the face to be recognized to form a composite feature containing spatial association information, obtaining a third concatenated feature. Subsequently, perform a non-linear transformation on the third concatenated feature through a probabilistic activation function to calculate the matching probability distribution between the target face feature in the pre-registered face feature and the current angular feature of the face to be recognized. Based on this probability distribution, perform a weighted sum on the target face feature in the pre-registered face feature, and finally determine the target face feature that best matches the current face angle, that is, the face feature to be matched.
[0067] In a specific embodiment, the feature redirection network is composed of two fully connected layers and a GELU layer. After inputting the angular feature of the face to be recognized and the pre-registered face feature into the feature redirection network, the first fully connected layer maps the angular feature of the face to be recognized and the pre-registered face feature to the corresponding feature space, performs a probability-driven non-linear activation in the GELU layer, generates an activation value affected by the statistical characteristics of the input distribution through Gaussian distribution weighting, and then inputs this activation value into the second fully connected layer to obtain a feature redirection indication result pointing to the corresponding target face feature in the pre-registered face feature. Extract the corresponding target face feature in the pre-registered face feature according to this feature redirection indication result to obtain the face feature to be matched.
[0068] In a specific embodiment, the above-mentioned face feature extraction network, feature fusion network, angular feature extraction network, angular encoder, and feature redirection network constitute a face recognition model. The loss function for training this face recognition model is as follows: , , , Wherein, is the total loss of the face recognition model, is the weight parameter, is the feature invariance loss, is the feature distance loss, is the i-th face feature to be matched of the current pre-registered face, is the target face feature of the i-th in the pre-registered face features, and n is the number of multi-angle faces in the pre-registered face library, is the face feature to be matched obtained by searching based on the face angle feature of the face to be recognized, is the face feature to be recognized, is the face feature to be matched of other pre-registered faces.
[0069] Please refer to Figure 5 , the embodiment of the present application also provides a face recognition device, which can implement the above face recognition method. The device includes: The first module 501 is used to obtain the image to be recognized; The second module 502 is used to extract the target face feature in the image to be recognized to obtain the face feature to be recognized; The third module 503 is used to extract the face angle feature in the image to be recognized to obtain the face angle feature to be recognized; The fourth module 504 is used to extract the target face feature corresponding to the face angle feature to be recognized in the pre-registered face features to obtain the face feature to be matched; the pre-registered face features include the target face features in multiple groups of pre-registered images collected from the same face at multiple angles; The fifth module 505 is used to generate the face recognition result of the image to be recognized according to the similarity between the face feature to be recognized and the face feature to be matched.
[0070] The specific implementation manner of this face recognition device is basically the same as the specific embodiment of the above face recognition method, and will not be elaborated here.
[0071] Figure 6 is a block diagram of an electronic device shown according to an exemplary embodiment.
[0072] Next, refer to Figure 6 to describe the electronic device 600 according to this embodiment of the present disclosure. Figure 6 The electronic device 600 shown is only an example and should not bring any limitation to the functions and usage scope of the embodiments of the present disclosure.
[0073] As Figure 6As shown, the electronic device 600 is presented in the form of a general-purpose computing device. The components of the electronic device 600 may include, but are not limited to: at least one processing unit 610, at least one storage unit 620, a bus 630 connecting different system components (including the storage unit 620 and the processing unit 610), a display unit 640, etc.
[0074] Among them, the storage unit stores program code, and the program code can be executed by the processing unit 610, so that the processing unit 610 executes the steps according to various exemplary embodiments of the present disclosure described in the above-mentioned face recognition method part of this specification.
[0075] The storage unit 620 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6201 and / or a cache storage unit 6202, and may further include a read-only storage unit (ROM) 6203.
[0076] The storage unit 620 may also include a program / utility 6204 having a set (at least one) of program modules 6205. Such program modules 6205 include, but are not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include the implementation of a network environment.
[0077] The bus 630 may represent one or more of several types of bus structures, including a storage unit bus or a storage unit controller, a peripheral bus, a graphics acceleration port, a processing unit, or a local bus using any bus structure in a variety of bus structures.
[0078] The electronic device 600 can also communicate with one or more external devices 600' (such as a keyboard, a pointing device, a Bluetooth device, etc.), and can also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (such as a router, a modem, etc.). Such communication can be carried out through an input / output (I / O) interface 650. And, the electronic device 600 can also communicate with one or more networks (such as a local area network (LAN), a wide area network (WAN), and / or a public network, such as the Internet) through a network adapter 660. The network adapter 660 can communicate with other modules of the electronic device 600 through the bus 630. It should be understood that although not shown in the figure, other hardware and / or software modules can be used in combination with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0079] An embodiment of the present application also provides a computer-readable storage medium storing a computer program, which when executed by a processor implements the above-mentioned face recognition method.
[0080] The face recognition method, device, equipment and storage medium provided by the embodiments of the present application extract face features and face angle features from the image to be recognized, extract corresponding target face features in the pre-registered face features according to the extracted face angle features of the face to be recognized to obtain the face features to be matched, and then generate the face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched. Since the face features to be matched extracted from the pre-registered face features are target face features with face angle features matching the face angle features of the face to be recognized, it can adapt to the face recognition scenarios of multiple face angles and improve the face recognition accuracy.
[0081] Through the description of the above embodiments, those skilled in the art can easily understand that the example embodiments described here can be implemented by software or by a combination of software and necessary hardware. Therefore, the technical solutions according to the embodiments of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on the network, including several instructions to enable a computing device (which can be a personal computer, a server, or a network device, etc.) to execute the above method according to the embodiments of the present disclosure.
[0082] The program product can adopt any combination of one or more readable media. The readable medium can be a readable signal medium or a readable storage medium. The readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples (a non-exhaustive list) of the readable storage medium include: an electrical connection with one or more wires, a portable disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above.
[0083] A computer-readable storage medium may include a data signal propagated in a baseband or as part of a carrier wave, in which a readable program code is carried. Such a propagated data signal may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the foregoing. The readable storage medium may also be any readable medium other than the computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. The program code contained on the readable storage medium may be transmitted by any appropriate medium, including but not limited to wireless, wired, optical fiber cable, RF, etc., or any suitable combination of the foregoing.
[0084] Those skilled in the art can understand that the above-mentioned modules can be distributed in the device according to the description of the embodiments, or can be correspondingly changed and be located in one or more devices that are only different from this embodiment. The modules of the above embodiments can be combined into one module, or can be further split into multiple sub-modules.
[0085] The exemplary embodiments of the present disclosure have been specifically shown and described above. It should be understood that the present disclosure is not limited to the detailed structures, settings, or implementation methods described herein; on the contrary, the present disclosure is intended to cover various modifications and equivalent settings included within the spirit and scope of the appended claims.
Claims
1. A face recognition method, characterized in that, including: obtaining an image to be recognized; extracting target face features in the image to be recognized to obtain face features to be recognized; extracting face angle features in the image to be recognized to obtain face angle features to be recognized; extracting target face features corresponding to the face angle features to be recognized from pre-registered face features to obtain face features to be matched; the pre-registered face features include target face features in multiple groups of pre-registered images collected from the same face at multiple angles; generating a face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched.
2. The face recognition method according to claim 1, wherein It further includes: obtaining multiple groups of pre-registered images corresponding to several faces; extracting target face features in the pre-registered images; fusing the target face features in the pre-registered images to obtain the pre-registered face features.
3. The face recognition method according to claim 2, wherein The fusing of the target face features in each of the pre-registered images includes: concatenating the target face features in each of the pre-registered images to obtain a first concatenated feature; performing non-linear activation processing on the first concatenated feature to obtain a first activated feature; performing a multi-head attention mechanism processing on the first activated feature to obtain an attention weight feature; superimposing the first activated feature and the attention weight feature to obtain the pre-registered face features.
4. The face recognition method according to claim 1, wherein The extracting of the face angle features in the image to be recognized includes: extracting the yaw angle feature and the pitch angle feature of the face in the image to be recognized; performing angle encoding processing on the yaw angle feature and the pitch angle feature to obtain the face angle features to be recognized.
5. The face recognition method according to claim 4, characterized in that The performing of the angle encoding processing on the yaw angle feature and the pitch angle feature includes: concatenating the yaw angle feature and the pitch angle feature to obtain a second concatenated feature; mapping the second concatenated feature to a preset feature space to obtain the face angle features to be recognized.
6. The face recognition method according to claim 1, wherein The extracting of the target face features corresponding to the face angle features to be recognized from the pre-registered face features includes: performing feature redirection on the pre-registered face features according to the mapping relationship between the face angle features to be recognized and the target face features in the pre-registered face features, so as to extract the corresponding target face features in the pre-registered face features to obtain the face features to be matched.
7. The face recognition method according to claim 6, wherein The performing of the feature redirection on the pre-registered face features according to the mapping relationship between the face angle features to be recognized and the target face features in the pre-registered face features includes: concatenating the pre-registered face features and the face angle features to be recognized to obtain a third concatenated feature; performing probabilistic activation processing on the third concatenated feature to obtain a second activated feature; extracting the corresponding target face features in the pre-registered face features according to the second activated feature to obtain the face features to be matched.
8. A face recognition device, characterized in that, including: a first module for obtaining an image to be recognized; a second module for extracting target face features in the image to be recognized to obtain face features to be recognized; a third module for extracting face angle features in the image to be recognized to obtain face angle features to be recognized; The fourth module is used to extract the target face features corresponding to the face angle features to be recognized in the pre-registered face features, so as to obtain the face features to be matched; the pre-registered face features include the target face features in multiple groups of pre-registered images obtained by collecting the same face from multiple angles. The fifth module is used to generate a face recognition result of the image to be recognized according to the similarity between the face features to be recognized and the face features to be matched.
9. An electronic device, characterized in that, The electronic device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the face recognition method according to any one of claims 1 to 7 is implemented.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the face recognition method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Face recognition method and device, robot and storage medium
CN109543633A
Face recognition method, device and system, computing equipment and storage medium
CN112016508A
Face recognition method and device applied to underground coal mine
CN119625809A
Face feature extraction model training method, face feature extraction method and device
CN119810894A