Face Recognition Method, Device, Terminal and Storage Medium

The face image is segmented and featured through the face analysis network and feature extraction network, and a five-feature feature descriptor is generated, which solves the problem of low face recognition accuracy in the prior art, and achieves higher accuracy and robust recognition.

CN114821724BActive Publication Date: 2025-07-08DUXIAOMAN TECH (BEIJING) CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202210455469.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-27
Publication Date
2025-07-08
Estimated Expiration
2042-04-27

AI Technical Summary

Technical Problem

In the prior art, the recognition accuracy of face recognition is low.

Method used

By receiving face images and using face analysis network model to divide them into multiple facial area images, each area is extracted and mapped network model processing is performed, and a five-feature feature descriptor is generated, and a preset threshold is used to determine whether the two face images are the same person.

Benefits of technology

It improves the accuracy of face recognition and the robustness of the model, and conforms to the face recognition process in the real world.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114821724B_ABST
    Figure CN114821724B_ABST
Patent Text Reader

Abstract

The present application discloses a face recognition method, device, terminal and storage medium. The method includes: receiving a first face image and a second face image; based on the first face image, the second face image and a face parsing network model, obtaining a plurality of facial region image pairs corresponding to the first face image and a plurality of facial region image pairs corresponding to the second face image; based on the plurality of facial region image pairs corresponding to the first face image, the plurality of facial region image pairs corresponding to the second face image, a feature extraction network model and a mapping network model, determining a facial feature descriptor corresponding to the first face image and a facial feature descriptor corresponding to the second face image; based on the facial feature descriptor corresponding to the first face image, the facial feature descriptor corresponding to the second face image and a preset threshold, determining that the portraits in the first face image and the second face image are of the same person. The present invention improves the accuracy of face recognition and the robustness of the model.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the technical fields of computer vision and deep learning. Specifically, it relates to a face recognition method, device, terminal, and storage medium. Background Technique

[0002] Face recognition is a biometric identification technology that identifies a person based on their facial feature information. A camera or webcam is used to collect images or video streams containing human faces, and the faces in the images are automatically detected and tracked, and then the detected faces are recognized.

[0003] Currently, for face recognition, it is necessary to determine in advance whether the face in the face image is occluded. When the face in the face image is not occluded, facial features are extracted from two face images respectively, and then the facial features are compared and recognized to determine whether the portraits in the two face images are the same person. When the face in the face image is occluded, the occlusion is removed and restored to obtain complete face features, and then the facial features are compared and recognized to determine whether the portraits in the two face images are the same person.

[0004] However, using the above method for face recognition has the problem of low recognition accuracy. Summary of the Invention

[0005] The main purpose of this application is to provide a face recognition method, device, terminal, and storage medium to solve the problem of low recognition accuracy in related technologies.

[0006] To achieve the above purpose, in the first aspect, this application provides a face recognition method, including:

[0007] Receiving a first face image and a second face image;

[0008] Based on the first face image, the second face image, and a face parsing network model, obtaining multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image;

[0009] Based on the multiple pairs of facial region images corresponding to the first face image, the multiple pairs of facial region images corresponding to the second face image, a feature extraction network model, and a mapping network model, determining a facial feature descriptor corresponding to the first face image and a facial feature descriptor corresponding to the second face image;

[0010] Based on the facial feature descriptor corresponding to the first face image, the facial feature descriptor corresponding to the second face image, and a preset threshold, determining that the portraits in the first face image and the second face image are the same person.

[0011] In a possible implementation manner, based on the first face image, the second face image, and the face parsing network model, multiple facial region image pairs corresponding to the first face image and multiple facial region image pairs corresponding to the second face image are obtained, including:

[0012] Input the first face image and the second face image into the face parsing network model respectively to obtain the semantic segmentation map corresponding to the first face image and the semantic segmentation map corresponding to the second face image;

[0013] Process the semantic segmentation map corresponding to the first face image and the semantic segmentation map corresponding to the second face image respectively to obtain multiple facial region image pairs corresponding to the first face image and multiple facial region image pairs corresponding to the second face image.

[0014] In a possible implementation manner, based on multiple facial region image pairs corresponding to the first face image, multiple facial region image pairs corresponding to the second face image, the feature extraction network model, and the mapping network model, the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image are determined, including:

[0015] Input multiple facial region image pairs corresponding to the first face image and multiple facial region image pairs corresponding to the second face image into the feature extraction network model respectively to obtain multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image;

[0016] Input multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image into the mapping network model respectively to obtain the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image.

[0017] In a possible implementation manner, the multiple facial region image pairs include a forehead region image pair, a left eye region image pair, a right eye region image pair, a nose region image pair, and a sub-nose region image pair; the feature extraction network model includes a forehead feature extraction network model, a left eye feature extraction network model, a right eye feature extraction network model, a nose feature extraction network model, and a sub-nose feature extraction network model; the multiple facial feature descriptors include a forehead feature descriptor, a left eye feature descriptor, a right eye feature descriptor, a nose feature descriptor, and a sub-nose feature descriptor;

[0018] Input multiple facial region image pairs corresponding to the first face image and multiple facial region image pairs corresponding to the second face image into the feature extraction network model respectively to obtain multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image, including:

[0019] Input the forehead region image pairs corresponding to the first face image and the forehead region image pairs corresponding to the second face image into the forehead feature extraction network model respectively, and output the forehead feature descriptor corresponding to the first face image and the forehead feature descriptor corresponding to the second face image;

[0020] Input the left eye region image pairs corresponding to the first face image and the left eye region image pairs corresponding to the second face image into the left eye feature extraction network model respectively, and output the left eye feature descriptor corresponding to the first face image and the left eye feature descriptor corresponding to the second face image;

[0021] Input the right eye region image pairs corresponding to the first face image and the right eye region image pairs corresponding to the second face image into the right eye feature extraction network model respectively, and output the right eye feature descriptor corresponding to the first face image and the right eye feature descriptor corresponding to the second face image;

[0022] Input the nose region image pairs corresponding to the first face image and the nose region image pairs corresponding to the second face image into the nose feature extraction network model respectively, and output the nose feature descriptor corresponding to the first face image and the nose feature descriptor corresponding to the second face image;

[0023] Input the sub - nose region image pairs corresponding to the first face image and the sub - nose region image pairs corresponding to the second face image into the sub - nose feature extraction network model respectively, and output the sub - nose feature descriptor corresponding to the first face image and the sub - nose feature descriptor corresponding to the second face image.

[0024] In a possible implementation, input the multiple facial feature descriptors corresponding to the first face image and the multiple facial feature descriptors corresponding to the second face image into the mapping network model respectively to obtain the facial feature descriptors corresponding to the five sense organs of the first face image and the facial feature descriptors corresponding to the five sense organs of the second face image, including:

[0025] Integrate the forehead feature descriptor, left eye feature descriptor, right eye feature descriptor, nose feature descriptor and sub - nose feature descriptor corresponding to the first face image and then input them into the mapping network model to obtain the facial feature descriptor corresponding to the five sense organs of the first face image;

[0026] Integrate the forehead feature descriptor, left eye feature descriptor, right eye feature descriptor, nose feature descriptor and sub - nose feature descriptor corresponding to the second face image and then input them into the mapping network model to obtain the facial feature descriptor corresponding to the five sense organs of the second face image.

[0027] In a possible implementation, the forehead region image pair includes a forehead region image and a mask image corresponding to the forehead region image. The left eye region image pair includes a left eye region image and a mask image corresponding to the left eye region image. The right eye region image pair includes a right eye region image and a mask image corresponding to the right eye region image. The nose region image pair includes a nose region image and a mask image corresponding to the nose region image. The sub-nose region image pair includes a sub-nose region image and a mask image corresponding to the sub-nose region image.

[0028] In a possible implementation, based on the facial feature descriptors corresponding to the first face image, the facial feature descriptors corresponding to the second face image, and a preset threshold, determining that the portraits in the first face image and the second face image are of the same person includes:

[0029] Calculating the similarity between the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image to obtain a similarity value;

[0030] Converting the similarity value into a similarity score and comparing the similarity score with the preset threshold;

[0031] When the similarity score reaches the preset threshold, the portraits in the first face image and the second face image are of the same person.

[0032] In a second aspect, an embodiment of the present invention provides a face recognition device, including:

[0033] An image receiving module, configured to receive a first face image and a second face image;

[0034] A face parsing module, configured to obtain a plurality of facial region image pairs corresponding to the first face image and a plurality of facial region image pairs corresponding to the second face image based on the first face image, the second face image, and a face parsing network model;

[0035] A facial feature determination module, configured to determine facial feature descriptors corresponding to the first face image and facial feature descriptors corresponding to the second face image based on the plurality of facial region image pairs corresponding to the first face image, the plurality of facial region image pairs corresponding to the second face image, a feature extraction network model, and a mapping network model;

[0036] A face comparison module, configured to determine that the portraits in the first face image and the second face image are of the same person based on the facial feature descriptors corresponding to the first face image, the facial feature descriptors corresponding to the second face image, and a preset threshold.

[0037] In a third aspect, an embodiment of the present invention provides a terminal, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the steps of any of the above face recognition methods are implemented.

[0038] In a fourth aspect, an embodiment of the present invention provides a computer-readable storage medium storing a computer program, and when the computer program is executed by a processor, the steps of any one of the above face recognition methods are implemented.

[0039] An embodiment of the present invention provides a face recognition method, device, terminal, and storage medium, including: receiving a first face image and a second face image, then based on the first face image, the second face image, and a face parsing network model, obtaining a plurality of facial region image pairs corresponding to the first face image and a plurality of facial region image pairs corresponding to the second face image, and then based on the plurality of facial region image pairs corresponding to the first face image, the plurality of facial region image pairs corresponding to the second face image, a feature extraction network model, and a mapping network model, determining a facial feature descriptor corresponding to the first face image and a facial feature descriptor corresponding to the second face image, and finally based on the facial feature descriptor corresponding to the first face image, the facial feature descriptor corresponding to the second face image, and a preset threshold, determining that the portraits in the first face image and the second face image are of the same person. The present invention uses network models to sequentially parse, extract features, and perform network mapping on face images to determine the facial feature descriptors corresponding to each face image, and by calculating the similarity between the facial feature descriptors corresponding to two face images, compares whether the portraits in the two face images are of the same person. This method not only well conforms to the face recognition process in the real world and has its rationality, but also improves the accuracy of face recognition and the robustness of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0040] The drawings constituting a part of this application are used to provide a further understanding of this application, making other features, objectives, and advantages of this application more obvious. The schematic embodiments and descriptions thereof in the drawings are used to explain this application and do not constitute an improper limitation of this application. In the drawings:

[0041] Figure 1 is a flowchart of the implementation of a face recognition method provided by an embodiment of the present invention;

[0042] Figure 2 is a flowchart of the implementation of face parsing provided by an embodiment of the present invention;

[0043] Figure 3 is a flowchart of the implementation of a face recognition method provided by another embodiment of the present invention;

[0044] Figure 4 is a schematic diagram of the interaction between a mask image and a corresponding region image provided by an embodiment of the present invention;

[0045] Figure 5It is a schematic structural diagram of a face recognition device provided by an embodiment of the present invention;

[0046] Figure 6 It is a schematic diagram of a terminal provided by an embodiment of the present invention. Specific embodiments

[0047] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Apparently, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.

[0048] The terms "first", "second", "third", "fourth", etc. (if any) in the specification and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and do not necessarily describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.

[0049] It should be understood that in various embodiments of the present invention, the magnitudes of the sequence numbers of the various processes do not mean the order of execution is prior or subsequent. The order of execution of the various processes should be determined according to their functions and internal logics, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0050] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily limit to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0051] It should be understood that in the present invention, "a plurality of" means two or more. "And / or" is only a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally represents an "or" relationship between the front and rear associated objects. "Including A, B, and C" and "including A, B, C" mean that all of A, B, and C are included. "Including A, B, or C" means including any one of A, B, and C. "Including A, B, and / or C" means including any one or any two or all three of A, B, and C.

[0052] It should be understood that in the present invention, "B corresponding to A", "B corresponding to A relatively", "A corresponding to B relatively" or "B corresponding to A relatively" means that B is associated with A, and B can be determined according to A. Determining B according to A does not mean determining B only according to A, and B can also be determined according to A and / or other information. The matching between A and B is that the similarity between A and B is greater than or equal to a preset threshold.

[0053] Depending on the context, as used herein, "if" can be interpreted as "when...", "while...", "in response to determining", or "in response to detecting".

[0054] The technical solution of the present invention will be described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.

[0055] To make the objectives, technical solutions, and advantages of the present invention clearer, the following will be described through specific embodiments with reference to the accompanying drawings.

[0056] In one embodiment, as Figure 1 shown, a face recognition method is provided, including the following steps:

[0057] Step S101: Receive a first face image and a second face image;

[0058] Step S102: Based on the first face image, the second face image, and the face parsing network model, obtain multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image.

[0059] The present invention uses a face parsing network model to parse the received first face image and second face image to determine multiple pairs of facial region images corresponding to each face image. Specifically, the first face image and the second face image are respectively input into the face parsing network model to obtain a semantic segmentation map corresponding to the first face image and a semantic segmentation map corresponding to the second face image, and then the semantic segmentation map corresponding to the first face image and the semantic segmentation map corresponding to the second face image are respectively processed to obtain multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image. Among them, the face parsing network model can be a classification model such as a neural network model, and the face parsing network model is obtained based on training a classification model such as a neural network model.

[0060] Combined with Figure 2Explain the steps for determining multiple pairs of facial region images corresponding to each face image. Taking the multiple facial regions including the forehead region, left eye region, right eye region, nose region, and sub-nose region as an example. Let the face image be the first face image. Inputting the first face image into the face parsing network model can obtain the semantic segmentation map corresponding to the first face image. In the semantic segmentation map, the forehead region, left eye region, right eye region, nose region, and sub-nose region are respectively marked in the form of a square. Among them, the forehead region (forehead) includes the facial region above the eyes, the left eye region (l_eye) includes the left eye and the upper eyebrow of the left eye, the right eye region (r_eye) includes the right eye and the upper eyebrow, the nose region (nose) includes the entire nose region in the face image, and the sub-nose region (mouth&chin) includes all regions of the mouth and chin, from the upper position to below the nose.

[0061] By further processing the above semantic segmentation map, pairs of forehead region images, left eye region images, right eye region images, nose region images, and sub-nose region images can be obtained. Among them, the pair of forehead region images includes the forehead region image and the mask image corresponding to the forehead region image. The pair of left eye region images includes the left eye region image and the mask image corresponding to the left eye region image. The pair of right eye region images includes the right eye region image and the mask image corresponding to the right eye region image. The pair of nose region images includes the nose region image and the mask image corresponding to the nose region image. The pair of sub-nose region images includes the sub-nose region image and the mask image corresponding to the sub-nose region image. Among them, in Figure 2 the mask image can be represented as a mask image.

[0062] In the above way, five pairs of facial region images corresponding to the first face image can be determined. The method for determining five pairs of facial region images corresponding to the second face image is the same as that of the first face image, which will not be elaborated here. In addition, the division of facial regions is not limited to Figure 2 as shown, and other methods can also be used for division.

[0063] Step S103: Based on the multiple pairs of facial region images corresponding to the first face image, the multiple pairs of facial region images corresponding to the second face image, the feature extraction network model, and the mapping network model, determine the five-feature descriptor corresponding to the first face image and the five-feature descriptor corresponding to the second face image.

[0064] After determining multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image through the above embodiments, the multiple pairs of facial region images corresponding to the first face image and the multiple pairs of facial region images corresponding to the second face image need to be input into the feature extraction network model respectively to obtain multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image.

[0065] Among them, the feature extraction network model needs to be extracted and trained, and the feature extraction network model includes multiple sub-network models. For example Figure 3 As shown, each sub-network model (Feature Extractor) in the multiple sub-network models corresponds to a facial region. Taking the left eye as an example of the facial region, the corresponding sub-network model is the left eye feature extraction network model.

[0066] When training the sub-network model, each sub-network model uses resNet18 as the backbone, outputs 512-dimensional features, and replaces the max-pooling of resNet18 with 3*3conv, stride = 2. There are two interactions between the mask image and the corresponding facial region image. For example Figure 4 As shown. Resize the mask image and the corresponding facial region image to the same size, then perform a multiplication operation (an attention operation) on the mask image and the corresponding facial region image, and input it into the sub-network model corresponding to this facial region. When downsampling to 1 / 2 in this sub-network model, perform another multiplication operation on the mask image and the corresponding facial region image. Through such two calculations, the influence of occlusion is eliminated. During training, random occlusion can be added to enhance the mask image and the corresponding facial region image, thereby improving the robustness and accuracy of the network. Among them, the loss function adopted by the sub-network model is arcface loss.

[0067] After the above training, multiple sub-network models (that is, the feature extraction network model) can be obtained. The following takes the feature extraction network model including the forehead feature extraction network model, the left eye feature extraction network model, the right eye feature extraction network model, the nose feature extraction network model, and the sub-nose feature extraction network model; multiple pairs of facial region images including the forehead region image pair, the left eye region image pair, the right eye region image pair, the nose region image pair, and the sub-nose region image pair; multiple facial feature descriptors including the forehead feature descriptor, the left eye feature descriptor, the right eye feature descriptor, the nose feature descriptor, and the sub-nose feature descriptor as an example to illustrate the process of determining multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image, which is specifically as follows:

[0068] The forehead region image pairs corresponding to the first face image and the forehead region image pairs corresponding to the second face image are respectively input into the forehead feature extraction network model to output the forehead feature descriptors corresponding to the first face image and the forehead feature descriptors corresponding to the second face image; the left eye region image pairs corresponding to the first face image and the left eye region image pairs corresponding to the second face image are respectively input into the left eye feature extraction network model to output the left eye feature descriptors corresponding to the first face image and the left eye feature descriptors corresponding to the second face image; the right eye region image pairs corresponding to the first face image and the right eye region image pairs corresponding to the second face image are respectively input into the right eye feature extraction network model to output the right eye feature descriptors corresponding to the first face image and the right eye feature descriptors corresponding to the second face image; the nose region image pairs corresponding to the first face image and the nose region image pairs corresponding to the second face image are respectively input into the nose feature extraction network model to output the nose feature descriptors corresponding to the first face image and the nose feature descriptors corresponding to the second face image; the sub-nose region image pairs corresponding to the first face image and the sub-nose region image pairs corresponding to the second face image are respectively input into the sub-nose feature extraction network model to output the sub-nose feature descriptors corresponding to the first face image and the sub-nose feature descriptors corresponding to the second face image. Among them, the above-mentioned feature descriptors are 512-dimensional features.

[0069] After obtaining multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image through the feature extraction network model, the multiple facial feature descriptors corresponding to the first face image and the multiple facial feature descriptors corresponding to the second face image need to be respectively input into the mapping network model to obtain the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image.

[0070] Among them, the facial feature descriptors are the total feature descriptors obtained by integrating the forehead feature descriptors, left eye feature descriptors, right eye feature descriptors, nose feature descriptors, and sub-nose feature descriptors. The mapping network model has three layers and is a fully connected mapping layer containing a 1280-node hidden layer, and is also determined after being trained using arcface loss.

[0071] After obtaining the trained mapping network model, the forehead feature descriptors, left eye feature descriptors, right eye feature descriptors, nose feature descriptors, and sub-nose feature descriptors corresponding to the first face image are integrated and input into the mapping network model to obtain the facial feature descriptors corresponding to the first face image, and the forehead feature descriptors, left eye feature descriptors, right eye feature descriptors, nose feature descriptors, and sub-nose feature descriptors corresponding to the second face image are integrated and input into the mapping network model to obtain the facial feature descriptors corresponding to the second face image. Among them, the facial feature descriptors are 512-dimensional new features.

[0072] Step S104: Determine that the human faces in the first face image and the second face image are the same person based on the facial feature descriptors corresponding to the first face image, the facial feature descriptors corresponding to the second face image, and a preset threshold.

[0073] After determining the facial feature descriptors corresponding to each face image, it is necessary to determine the similarity value between the two face images based on the facial feature descriptors corresponding to the two face images, and then judge whether the human faces in the first face image and the second face image are the same person based on the preset threshold and the similarity value. Specifically, first calculate the similarity between the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image to obtain a similarity value, then convert the similarity value into a similarity score, and compare the similarity score with the preset threshold. When the similarity score reaches the preset threshold, the human faces in the first face image and the second face image are the same person.

[0074] An embodiment of the present invention provides a face recognition method, including: receiving a first face image and a second face image, then obtaining multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image based on the first face image, the second face image, and a face parsing network model, and then determining the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image based on the multiple pairs of facial region images corresponding to the first face image, the multiple pairs of facial region images corresponding to the second face image, a feature extraction network model, and a mapping network model, and finally determining that the human faces in the first face image and the second face image are the same person based on the facial feature descriptors corresponding to the first face image, the facial feature descriptors corresponding to the second face image, and a preset threshold. The present invention uses network models to parse, extract features, and perform network mapping on face images in sequence to determine the facial feature descriptors corresponding to each face image, and compares whether the human faces in the two face images are the same person by calculating the similarity between the facial feature descriptors corresponding to the two face images. This method not only well fits the face recognition process in the real world and is reasonable, but also improves the accuracy of face recognition and the robustness of the model.

[0075] It should be understood that the magnitudes of the sequence numbers of the steps in the above embodiments do not mean the order of execution. The order of execution of each process should be determined according to its function and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0076] The following is an apparatus embodiment of the present invention. For details not described in detail herein, reference may be made to the corresponding method embodiments above.

[0077] Figure 5The figure shows a schematic structural diagram of a face recognition device provided by an embodiment of the present invention. For ease of description, only parts related to the embodiment of the present invention are shown. A face recognition device includes an image receiving module 51, a face parsing module 52, a facial feature determination module 53, and a face comparison module 54, as follows:

[0078] The image receiving module 51 is configured to receive a first face image and a second face image;

[0079] The face parsing module 52 is configured to obtain a plurality of facial region image pairs corresponding to the first face image and a plurality of facial region image pairs corresponding to the second face image based on the first face image, the second face image, and a face parsing network model;

[0080] The facial feature determination module 53 is configured to determine a facial feature descriptor corresponding to the first face image and a facial feature descriptor corresponding to the second face image based on the plurality of facial region image pairs corresponding to the first face image, the plurality of facial region image pairs corresponding to the second face image, a feature extraction network model, and a mapping network model;

[0081] The face comparison module 54 is configured to determine that the portraits in the first face image and the second face image are the same person based on the facial feature descriptor corresponding to the first face image, the facial feature descriptor corresponding to the second face image, and a preset threshold.

[0082] In a possible implementation manner, the face parsing module 52 includes:

[0083] A face parsing sub-module, configured to respectively input the first face image and the second face image into the face parsing network model to obtain a semantic segmentation map corresponding to the first face image and a semantic segmentation map corresponding to the second face image;

[0084] An image processing sub-module, configured to respectively process the semantic segmentation map corresponding to the first face image and the semantic segmentation map corresponding to the second face image to obtain a plurality of facial region image pairs corresponding to the first face image and a plurality of facial region image pairs corresponding to the second face image.

[0085] In a possible implementation manner, the facial feature determination module 53 includes:

[0086] A feature extraction sub-module, configured to respectively input the plurality of facial region image pairs corresponding to the first face image and the plurality of facial region image pairs corresponding to the second face image into the feature extraction network model to obtain a plurality of facial feature descriptors corresponding to the first face image and a plurality of facial feature descriptors corresponding to the second face image;

[0087] A network mapping sub-module, configured to input multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image into a mapping network model respectively, so as to obtain the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image.

[0088] In a possible implementation, the multiple facial region image pairs include a forehead region image pair, a left eye region image pair, a right eye region image pair, a nose region image pair, and a sub-nose region image pair; the feature extraction network model includes a forehead feature extraction network model, a left eye feature extraction network model, a right eye feature extraction network model, a nose feature extraction network model, and a sub-nose feature extraction network model; the multiple facial feature descriptors include a forehead feature descriptor, a left eye feature descriptor, a right eye feature descriptor, a nose feature descriptor, and a sub-nose feature descriptor;

[0089] The feature extraction sub-module includes:

[0090] A forehead feature extraction unit, configured to input the forehead region image pair corresponding to the first face image and the forehead region image pair corresponding to the second face image into the forehead feature extraction network model respectively, and output the forehead feature descriptor corresponding to the first face image and the forehead feature descriptor corresponding to the second face image;

[0091] A left eye feature extraction unit, configured to input the left eye region image pair corresponding to the first face image and the left eye region image pair corresponding to the second face image into the left eye feature extraction network model respectively, and output the left eye feature descriptor corresponding to the first face image and the left eye feature descriptor corresponding to the second face image;

[0092] A right eye feature extraction unit, configured to input the right eye region image pair corresponding to the first face image and the right eye region image pair corresponding to the second face image into the right eye feature extraction network model respectively, and output the right eye feature descriptor corresponding to the first face image and the right eye feature descriptor corresponding to the second face image;

[0093] A nose feature extraction unit, configured to input the nose region image pair corresponding to the first face image and the nose region image pair corresponding to the second face image into the nose feature extraction network model respectively, and output the nose feature descriptor corresponding to the first face image and the nose feature descriptor corresponding to the second face image;

[0094] A sub-nose feature extraction unit, configured to input the sub-nose region image pair corresponding to the first face image and the sub-nose region image pair corresponding to the second face image into the sub-nose feature extraction network model respectively, and output the sub-nose feature descriptor corresponding to the first face image and the sub-nose feature descriptor corresponding to the second face image.

[0095] In a possible implementation, the network mapping sub-module includes:

[0096] A first network mapping unit, configured to integrate the forehead feature descriptor, left-eye feature descriptor, right-eye feature descriptor, nose feature descriptor, and sub-nose feature descriptor corresponding to the first face image, and input the integrated result into a mapping network model to obtain the facial feature descriptor corresponding to the first face image;

[0097] A second network mapping unit, configured to integrate the forehead feature descriptor, left-eye feature descriptor, right-eye feature descriptor, nose feature descriptor, and sub-nose feature descriptor corresponding to the second face image, and input the integrated result into a mapping network model to obtain the facial feature descriptor corresponding to the second face image.

[0098] In a possible implementation, the forehead region image pair includes a forehead region image and a mask image corresponding to the forehead region image, the left-eye region image pair includes a left-eye region image and a mask image corresponding to the left-eye region image, the right-eye region image pair includes a right-eye region image and a mask image corresponding to the right-eye region image, the nose region image pair includes a nose region image and a mask image corresponding to the nose region image, and the sub-nose region image pair includes a sub-nose region image and a mask image corresponding to the sub-nose region image.

[0099] In a possible implementation, the face comparison module 54 includes:

[0100] A similarity calculation unit, configured to calculate the similarity between the facial feature descriptor corresponding to the first face image and the facial feature descriptor corresponding to the second face image to obtain a similarity value;

[0101] A score comparison unit, configured to convert the similarity value into a similarity score and compare the similarity score with a preset threshold;

[0102] A portrait determination unit, configured to determine that the portraits in the first face image and the second face image are the same person when the similarity score reaches the preset threshold.

[0103] Figure 6 It is a schematic diagram of a terminal provided by an embodiment of the present invention. As Figure 6 shown, the terminal 6 in this embodiment includes: a processor 61, a memory 62, and a computer program 63 stored in the memory 62 and executable on the processor 61. When the processor 61 executes the computer program 63, the steps in the above-mentioned embodiments of each face recognition method are implemented, such as Figure 1 the steps 101 to 104 shown. Alternatively, when the processor 61 executes the computer program 63, the functions of each module / unit in the above-mentioned embodiments of each face recognition device are implemented, such as Figure 5 the functions of the modules / units 51 to 54 shown.

[0104] The present invention also provides a readable storage medium storing a computer program, which is used to implement the face recognition method provided by the above various embodiments when executed by a processor.

[0105] Among them, the readable storage medium can be a computer storage medium or a communication medium. The communication medium includes any medium facilitating the transmission of a computer program from one place to another. The computer storage medium can be any available medium accessible by a general-purpose or special-purpose computer. For example, the readable storage medium is coupled to the processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can be located in an application specific integrated circuit (ASIC). Additionally, the ASIC can be located in a user device. Of course, the processor and the readable storage medium can also exist as discrete components in a communication device. The readable storage medium can be a read-only memory (ROM), a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0106] The present invention also provides a program product, which includes execution instructions stored in a readable storage medium. At least one processor of the device can read the execution instructions from the readable storage medium, and the execution of the execution instructions by at least one processor causes the device to implement the face recognition method provided by the above various embodiments.

[0107] In the embodiment of the above device, it should be understood that the processor can be a central processing unit (CPU for short), and can also be other general-purpose processors, digital signal processors (DSP for short), application specific integrated circuits (ASIC for short), etc. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor, etc. The steps of the method disclosed in combination with the present invention can be directly implemented by the execution of the hardware processor, or implemented by the combination of hardware and software modules in the processor.

[0108] The above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit it; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the various embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A face recognition method, characterized in that, Including: Receiving a first face image and a second face image; Based on the first face image, the second face image, and a face parsing network model, obtaining multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image; Based on the multiple pairs of facial region images corresponding to the first face image, the multiple pairs of facial region images corresponding to the second face image, a feature extraction network model, and a mapping network model, determining a five - sense feature descriptor corresponding to the first face image and a five - sense feature descriptor corresponding to the second face image; Wherein, the five - sense feature descriptor is a total feature descriptor obtained by integrating a forehead feature descriptor, a left - eye feature descriptor, a right - eye feature descriptor, a nose feature descriptor, and a sub - nose feature descriptor, and the dimension of the five - sense feature descriptor is the same as the dimensions of the forehead feature descriptor, the left - eye feature descriptor, the right - eye feature descriptor, the nose feature descriptor, and the sub - nose feature descriptor; Wherein, the feature extraction network model includes multiple sub - network models, and each sub - network model in the multiple sub - network models corresponds to a facial region; When training the sub - network model, each sub - network model uses resNet18 as the backbone and outputs features with a dimension of 512; There are two interactions between the mask image and the corresponding facial region image: Resizing the mask image and the corresponding facial region image to the same size, then performing a multiplication operation on the mask image and the corresponding facial region image, and then inputting the image after the multiplication operation into the sub - network model corresponding to the facial region. When the sub - network model downsamples to 1 / 2, performing another multiplication operation on the mask image and the corresponding facial region image. The loss function used by the sub - network model is arcface loss; Based on the five - sense feature descriptor corresponding to the first face image, the five - sense feature descriptor corresponding to the second face image, and a preset threshold, determining that the portraits in the first face image and the second face image are of the same person.

2. The face recognition method according to claim 1, wherein The step of obtaining multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image based on the first face image, the second face image, and a face parsing network model includes: Inputting the first face image and the second face image into the face parsing network model respectively to obtain a semantic segmentation map corresponding to the first face image and a semantic segmentation map corresponding to the second face image; Processing the semantic segmentation map corresponding to the first face image and the semantic segmentation map corresponding to the second face image respectively to obtain multiple pairs of facial region images corresponding to the first face image and multiple pairs of facial region images corresponding to the second face image.

3. The face recognition method according to claim 2, wherein The step of determining a five - sense feature descriptor corresponding to the first face image and a five - sense feature descriptor corresponding to the second face image based on the multiple pairs of facial region images corresponding to the first face image, the multiple pairs of facial region images corresponding to the second face image, a feature extraction network model, and a mapping network model includes: Input the multiple pairs of facial region images corresponding to the first face image and the multiple pairs of facial region images corresponding to the second face image into the feature extraction network model respectively, to obtain multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image; Input the multiple facial feature descriptors corresponding to the first face image and the multiple facial feature descriptors corresponding to the second face image into the mapping network model respectively, to obtain the facial feature descriptors of the facial features of the first face image and the facial feature descriptors of the facial features of the second face image.

4. The face recognition method according to claim 3, wherein The multiple pairs of facial region images include a forehead region image pair, a left eye region image pair, a right eye region image pair, a nose region image pair, and an image pair under the nose; the feature extraction network model includes a forehead feature extraction network model, a left eye feature extraction network model, a right eye feature extraction network model, a nose feature extraction network model, and an image pair under the nose feature extraction network model; the multiple facial feature descriptors include a forehead feature descriptor, a left eye feature descriptor, a right eye feature descriptor, a nose feature descriptor, and an image pair under the nose feature descriptor; The step of inputting the multiple pairs of facial region images corresponding to the first face image and the multiple pairs of facial region images corresponding to the second face image into the feature extraction network model respectively, to obtain multiple facial feature descriptors corresponding to the first face image and multiple facial feature descriptors corresponding to the second face image, includes: Input the forehead region image pair corresponding to the first face image and the forehead region image pair corresponding to the second face image into the forehead feature extraction network model respectively, and output the forehead feature descriptor corresponding to the first face image and the forehead feature descriptor corresponding to the second face image; Input the left eye region image pair corresponding to the first face image and the left eye region image pair corresponding to the second face image into the left eye feature extraction network model respectively, and output the left eye feature descriptor corresponding to the first face image and the left eye feature descriptor corresponding to the second face image; Input the right eye region image pair corresponding to the first face image and the right eye region image pair corresponding to the second face image into the right eye feature extraction network model respectively, and output the right eye feature descriptor corresponding to the first face image and the right eye feature descriptor corresponding to the second face image; Input the nose region image pair corresponding to the first face image and the nose region image pair corresponding to the second face image into the nose feature extraction network model respectively, and output the nose feature descriptor corresponding to the first face image and the nose feature descriptor corresponding to the second face image; Input the image pair under the nose corresponding to the first face image and the image pair under the nose corresponding to the second face image into the image pair under the nose feature extraction network model respectively, and output the image pair under the nose feature descriptor corresponding to the first face image and the image pair under the nose feature descriptor corresponding to the second face image.

5. The face recognition method according to claim 3, wherein Respectively inputting the multiple facial feature descriptors corresponding to the first face image and the multiple facial feature descriptors corresponding to the second face image into the mapping network model to obtain the facial feature descriptors corresponding to the facial features of the first face image and the facial feature descriptors corresponding to the facial features of the second face image includes: Integrating the forehead feature descriptor, left eye feature descriptor, right eye feature descriptor, nose feature descriptor, and sub-nose feature descriptor corresponding to the first face image and inputting them into the mapping network model to obtain the facial feature descriptors corresponding to the facial features of the first face image; Integrating the forehead feature descriptor, left eye feature descriptor, right eye feature descriptor, nose feature descriptor, and sub-nose feature descriptor corresponding to the second face image and inputting them into the mapping network model to obtain the facial feature descriptors corresponding to the facial features of the second face image.

6. The face recognition method according to claim 4 or 5, wherein The forehead region image pair includes a forehead region image and a mask image corresponding to the forehead region image. The left eye region image pair includes a left eye region image and a mask image corresponding to the left eye region image. The right eye region image pair includes a right eye region image and a mask image corresponding to the right eye region image. The nose region image pair includes a nose region image and a mask image corresponding to the nose region image. The sub-nose region image pair includes a sub-nose region image and a mask image corresponding to the sub-nose region image.

7. The face recognition method according to any one of claims 1-5, characterized in that, Based on the facial feature descriptors corresponding to the first face image, the facial feature descriptors corresponding to the second face image, and a preset threshold, determining that the portraits in the first face image and the second face image are of the same person includes: Calculating the similarity between the facial feature descriptors corresponding to the first face image and the facial feature descriptors corresponding to the second face image to obtain a similarity value; Converting the similarity value into a similarity score and comparing the similarity score with the preset threshold; When the similarity score reaches the preset threshold, the portraits in the first face image and the second face image are of the same person.

8. A face recognition device, characterized in that, Including: An image receiving module, configured to receive a first face image and a second face image; A face parsing module, configured to obtain multiple facial region image pairs corresponding to the first face image and multiple facial region image pairs corresponding to the second face image based on the first face image, the second face image, and a face parsing network model; A facial feature determination module, configured to determine the facial feature descriptors corresponding to the facial features of the first face image and the facial feature descriptors corresponding to the facial features of the second face image based on the multiple facial region image pairs corresponding to the first face image, the multiple facial region image pairs corresponding to the second face image, a feature extraction network model, and a mapping network model; Wherein, the facial feature descriptor is a total feature descriptor obtained by integrating the forehead feature descriptor, left eye feature descriptor, right eye feature descriptor, nose feature descriptor, and sub-nose feature descriptor, and the dimension of the facial feature descriptor is the same as the dimensions of the forehead feature descriptor, left eye feature descriptor, right eye feature descriptor, nose feature descriptor, and sub-nose feature descriptor. Among them, the feature extraction network model includes multiple sub-network models, and each sub-network model in the multiple sub-network models corresponds to a facial region; When training the sub-network model, each sub-network model uses resNet18 as the backbone and outputs 512-dimensional features; There are two interactions between the mask image and the corresponding facial region image: Resize the mask image and the corresponding facial region image to the same size, then perform a multiplication operation on the mask image and the corresponding facial region image, and then input the image after the multiplication operation into the sub-network model corresponding to the facial region. When the sub-network model downsamples to 1 / 2, perform another multiplication operation on the mask image and the corresponding facial region image. The loss function used by the sub-network model is arcface loss; A face comparison module, configured to determine that the portraits in the first face image and the second face image are the same person based on the facial feature descriptors corresponding to the first face image, the facial feature descriptors corresponding to the second face image, and a preset threshold.

9. A terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the face recognition method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the face recognition method according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Similarity discrimination system based on five sense organs of human face

    CN106980819A

  • Block-based shielded face recognition algorithm

    CN108805040A

  • Access control method and system based on artificial intelligence

    CN112150692A

  • Image processing method and device, electronic equipment and computer readable medium

    CN112418054A