Model Generation Method, Apparatus, and Storage Medium, and Face Recognition Method and Apparatus
By cropping and training the face images of masks, a model for spatial transformation is generated, which solves the problem of poor facial recognition effect under masks and achieves higher recognition accuracy.
Patent Information
- Application Number
- CN202011563248.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2020-12-25
- Publication Date
- 2025-06-13
- Estimated Expiration
- 2040-12-25
AI Technical Summary
In the prior art, when wearing a mask, it is difficult to accurately identify the key points of face positions and facial features, resulting in a reduced facial recognition effect and low accuracy.
By obtaining a large number of face images of mask-wearing faces, cropping and obtaining face images that are not blocked by masks as supervision information, and the sample pair is formed to train them to obtain a model for spatial transformation. This model can effectively extract and feature extraction of mask faces to improve recognition accuracy.
Effective recognition of faces wearing masks has been achieved, and the accuracy of face recognition has been improved, especially in terms of local feature enhancement, and the identification of key areas such as eyes and eyebrows has been strengthened.
Smart Images

Figure CN114693987B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the field of face recognition technology, and in particular, to a model generation method, apparatus, storage medium, and face recognition method and apparatus. Background Art
[0002] Wearing a mask belongs to a large-area face occlusion. In the field of face recognition, since the face recognition model mainly determines identity based on facial features of the face, when recognizing a face wearing a mask, the face recognition model cannot accurately detect the face position and locate the key points of the facial features, greatly reducing the recognition effect and having a low recognition accuracy. Recognizing a face wearing a mask has always been a recognized difficult problem. Summary of the Invention
[0003] To solve the above technical problems of high difficulty and inaccurate recognition in face recognition of masked faces, the present disclosure provides a model generation method, apparatus, storage medium, and face recognition method and apparatus.
[0004] In a first aspect, the present disclosure provides a model generation method, which includes:
[0005] Obtain a first sample set, where the first sample set includes m first masked face images;
[0006] Perform a first process on each first masked face image to obtain corresponding supervision information, where the supervision information is the face image in the corresponding first masked face image that is not blocked by the mask;
[0007] Form a first sample pair by combining each first masked face image with the corresponding supervision information to obtain a first training sample set including m first sample pairs;
[0008] Train an initial spatial transformation model using the first training sample set to obtain a trained spatial transformation model;
[0009] Among them, the trained spatial transformation model can be used for effective face extraction in mask face recognition,
[0010] The first process includes at least a cropping process.
[0011] Optionally, training an initial spatial transformation model using the first training sample set to obtain a trained spatial transformation model includes:
[0012] Obtain a current first sub-training sample set including n first masked face images, where n is less than m;
[0013] Input the n first masked face images in the current first sub-training sample set into the current spatial transformation model for spatial transformation to obtain n current first effective face images;
[0014] The current first sub-training sample set is: the first sub-training sample set in the first training sample set for training the current spatial transformation model;
[0015] The current spatial transformation model is: the spatial transformation model obtained after training with all the first sub-training sample sets before the current first sub-training sample set;
[0016] Input n current first valid face images into the trained first feature extraction network to extract the first features of each current first valid face image;
[0017] Input the n pieces of supervision information in the current first sub-training sample set into the trained first feature extraction network to extract the second features of each piece of supervision information;
[0018] Obtain the first cosine distance between each first feature and the corresponding second feature;
[0019] Obtain the first loss value according to the n first cosine distances;
[0020] Judge whether the first loss value is less than or equal to the first threshold;
[0021] Update the training parameters of the current spatial transformation model using the first loss value;
[0022] If the first loss value is greater than the first threshold, obtain the next first sub-training sample set as the current first sub-training sample set, and perform the operation of respectively inputting the n first mask face images in the current first sub-training sample set into the current spatial transformation model for spatial transformation to obtain n current first valid face images;
[0023] If the first loss value is less than or equal to the first threshold, end the training and use the updated spatial transformation model as the trained spatial transformation model.
[0024] Optionally, it is characterized in that obtaining the first loss value according to the n first cosine distances includes:
[0025] Calculate the first loss value according to the loss function and the n first cosine distances.
[0026] Optionally, both the initial spatial transformation model and the trained spatial transformation model include a convolutional layer, a first inverted residual layer, an average pooling layer, a second inverted residual layer, a global average pooling layer, and a fully connected layer connected in sequence.
[0027] In a second aspect, the present disclosure provides a mask face recognition method, and the face recognition method includes:
[0028] Obtain the mask face image to be recognized;
[0029] The face to be recognized in the face image of the mask to be recognized is partially blocked by the mask;
[0030] Use the trained mask face recognition model to perform effective face extraction and face feature extraction on the face image of the mask to be recognized, so as to obtain the effective face features of the face to be recognized in the face image of the mask to be recognized;
[0031] Among them, the trained mask face recognition model includes: a trained spatial transformation model for effective face extraction and a trained mask face feature extraction model for feature extraction;
[0032] The effective face features are the face features of the face in the face image of the mask to be recognized that are not blocked by the mask;
[0033] Calculate the similarity between the effective face features of the face to be recognized and the effective face features of each known face in the pre-stored known face set, so as to obtain corresponding multiple similarities;
[0034] Perform face recognition on the face to be recognized based on multiple similarities.
[0035] Optionally, the trained mask face recognition model further includes a trained mask face feature extraction model;
[0036] Before using the trained mask face recognition model to perform effective face extraction and face feature extraction on the face image of the mask to be recognized, so as to obtain the effective face features of the face to be recognized in the face image of the mask to be recognized, the method further includes:
[0037] Embed the trained spatial transformation model into the first layer of the initial mask face feature extraction model to obtain the mask face feature extraction model to be trained,
[0038] Obtain a second training sample set, where the second training sample set includes multiple second mask face images with marked labels,
[0039] Use the second training sample set to train the mask face feature extraction model to be trained to obtain the trained mask face feature extraction model.
[0040] Optionally, using the trained mask face recognition model to perform effective face extraction and face feature extraction on the face image of the mask to be recognized, so as to obtain the effective face features of the face to be recognized in the face image of the mask to be recognized, includes:
[0041] Input the face image of the mask to be recognized into the trained spatial transformation model for spatial transformation, so as to obtain an effective face image of the face to be recognized in the face image of the mask to be recognized;
[0042] Input the valid face image of the face to be recognized into the trained mask face feature extraction model to obtain the valid face features of the face to be recognized in the face image of the mask face to be recognized.
[0043] In a third aspect, the present disclosure provides a model generation device, which includes:
[0044] A first sample acquisition module, configured to acquire a first sample set, where the first sample set includes m first mask face images;
[0045] A first processing module, configured to perform a first processing on each first mask face image to obtain corresponding supervision information, where the supervision information is the face image not covered by the mask in the corresponding first mask face image;
[0046] A first training sample generation module, configured to form a first sample pair by combining each first mask face image with the corresponding supervision information, so as to obtain a first training sample set including m first sample pairs;
[0047] A first training module, configured to train an initial space transformation model by using the first training sample set to obtain a trained space transformation model;
[0048] Wherein, the first processing at least includes a cropping process.
[0049] In a fourth aspect, the present disclosure provides a mask face recognition device, which includes:
[0050] An image acquisition module, configured to acquire a face image of a mask face to be recognized;
[0051] The face to be recognized in the face image of the mask face to be recognized is partially covered by the mask;
[0052] A feature extraction module, configured to perform valid face extraction and face feature extraction on the face image of the mask face to be recognized by using the trained mask face recognition model, so as to obtain the valid face features of the face to be recognized in the face image of the mask face to be recognized;
[0053] Wherein, the trained mask face recognition model includes: a trained space transformation model for valid face extraction;
[0054] The valid face feature is the face feature of the face not covered by the mask in the face image of the mask face to be recognized;
[0055] A similarity acquisition module, configured to calculate the similarity between the valid face features of the face to be recognized and the valid face features of each known face in the pre-stored known face set, so as to obtain corresponding multiple similarities;
[0056] An identification module, configured to perform face recognition on the face to be recognized based on the multiple similarities.
[0057] Optionally, the trained spatial transformation model is obtained according to the model generation device described above.
[0058] In a fifth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor is caused to execute the steps of the model generation method as described in any one of the foregoing.
[0059] In a sixth aspect, the present disclosure provides a computer-readable storage medium having a computer program stored thereon. When the computer program is executed by a processor, the processor is caused to execute the steps of the mask face recognition method as described in any one of the foregoing.
[0060] In a seventh aspect, the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the processor executes the steps of the model generation method as described in any one of the foregoing.
[0061] In an eighth aspect, the present disclosure provides a computer device, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the program, the processor executes the steps of the mask face recognition method as described in any one of the foregoing.
[0062] The above technical solutions provided by the embodiments of the present disclosure have the following advantages compared with the prior art:
[0063] The present disclosure obtains a first sample set, where the first sample set includes a plurality of first mask face images; performs a first process on each first mask face image to obtain corresponding supervision information, where the supervision information is a face image of the part of the corresponding first mask face image that is not blocked by the mask; forms a first sample pair by combining each first mask face image with the corresponding supervision information to obtain a first training sample set including a plurality of first sample pairs; and trains an initial spatial transformation model using the first training sample set to obtain a trained spatial transformation model. The trained spatial transformation model performs a spatial transformation on the mask face to obtain an effective face image including key regions such as eyes and eyebrows, upgrades the original face recognition model in a targeted manner, increases the weight of the visible region of the face, designs corresponding strategies in terms of local feature enhancement, strengthens the recognition of key regions such as eyes and eyebrows, and improves the accuracy of face recognition under a mask. Description of the Drawings
[0064] The drawings herein are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with the present invention and, together with the specification, are used to explain the principles of the present invention.
[0065] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, for those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.
[0066] Figure 1 The flowchart of a model generation method provided by an embodiment of the present disclosure;
[0067] Figure 2 The flowchart of a mask face recognition method provided by an embodiment of the present disclosure;
[0068] Figure 3 The structural block diagram of a model generation device provided by an embodiment of the present disclosure;
[0069] Figure 4 The structural block diagram of a mask face recognition device provided by an embodiment of the present disclosure. Detailed implementation manners
[0070] To make the objectives, technical solutions, and advantages of the embodiments of the present disclosure clearer, the following will clearly and completely describe the technical solutions in the embodiments of the present disclosure with reference to the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, rather than all of them. Based on the embodiments of the present disclosure, all other embodiments obtained by those of ordinary skill in the art without creative efforts belong to the scope of protection of the present disclosure.
[0071] Figure 1 The flowchart of a model generation method provided by an embodiment of the present disclosure. Refer to Figure 1 , the model generation method includes the following steps:
[0072] S100A: Obtain a first sample set, where the first sample set includes a plurality of first mask face images.
[0073] Specifically, the first mask face image is a face image wearing a mask. The face part in each first mask face image is blocked by the mask, and parts such as eyes, eyebrows, and forehead are outside the mask and visible.
[0074] The first sample set includes a large number of first mask face images as samples for training the STN network model.
[0075] The acquisition method of the first mask face image can be:
[0076] A large number of first masked face images are obtained by adding masks to a large number of different normal face images (faces without masks). This method can reduce the shooting work of a large number of masked face images and quickly obtain masked face images.
[0077] Of course, a large number of first masked face images can also be obtained by shooting different masked faces with a camera device.
[0078] S200A: Perform a first process on each first masked face image to obtain corresponding supervision information, where the supervision information is the face image of the part of the corresponding first masked face image that is not blocked by the mask.
[0079] Specifically, the first process at least includes a cropping process. The purpose of the cropping process is to crop out the face image outside the mask and not blocked by the mask in the masked face image. The supervision information is the image obtained after the first masked face image has at least undergone a cropping process.
[0080] Optionally, the first process may further include a correction process, such as skew correction; it may also include a denoising process, such as noise interference processing, etc. The purpose of the correction process is to correct the skewed face image into a non-skewed face image and center the face image. The purpose of the denoising process is to remove interference factors such as the background in the cropped face image. Of course, the first process may also include an image enhancement process to make the image clearer and easier to recognize.
[0081] S300A: Combine each first masked face image with the corresponding supervision information to form a first sample pair, so as to obtain a first training sample set containing multiple first sample pairs.
[0082] Specifically, each first masked face image has corresponding supervision information, and a first masked face image and the corresponding supervision information form a first sample pair. All the first sample pairs corresponding to all the first masked face images form the first training sample set. Therefore, the first training sample set includes multiple first sample pairs.
[0083] S400A: Train the initial spatial transformation model with the first training sample set to obtain a trained spatial transformation model.
[0084] Specifically, the spatial transformation model can be an STN network model. The STN network model (Spatial Transformer Networks) is a spatial transformation network model. The STN network model allows the neural network to learn how to perform spatial transformations on the input image to improve the geometric invariance of the model. For example, the STN network model has functions such as cropping the region of interest, scaling, and correcting the orientation of the image.
[0085] The trained spatial transformation model of the present application can be applied to mask face recognition to perform spatial transformation on mask face images to obtain effective face images. The effective face image is the face area image outside the mask in the mask face image.
[0086] The present disclosure has made a targeted upgrade to the original face recognition algorithm model, increased the weight of the visible face area, and designed corresponding strategies in terms of local feature enhancement. For example, the recognition of key areas such as eyes and eyebrows has been strengthened, and the face recognition accuracy under wearing a mask has been improved.
[0087] In one embodiment, step S400A specifically includes the following steps:
[0088] S410: Obtain the current first sub-training sample set containing n first mask face images, where n is less than m.
[0089] In a specific embodiment, the first training sample set obtained in step S300 contains m first sample pairs. These m first sample pairs are grouped to obtain multiple groups of first sub-training sample sets, and each first sub-training sample set includes n first sample pairs. Among them, n is less than m, and m and n are positive integers greater than 0. The number n of first sample pairs included in each first sub-training sample set can be the same or different.
[0090] The number n of first sample pairs included in each first sub-training sample set depends on the number of images that the spatial transformation model can process in parallel.
[0091] Each first sub-training sample set is used to train the model to be trained once. All first sub-training sample sets train the model to be trained in sequence in turn. The current first sub-training sample set is one of the multiple groups of first sub-training sample sets.
[0092] In a specific embodiment, the current first sub-training sample set can also be: composed of n first sample pairs randomly selected or selected according to a preset rule from the first training sample set at the current training moment.
[0093] S420: Input the n first mask face images in the current first sub-training sample set into the current spatial transformation model for spatial transformation to obtain n current first effective face images.
[0094] Specifically, the current first sub-training sample set is: the first sub-training sample set in the first training sample set for training the current spatial transformation model.
[0095] The current spatial transformation model is: the spatial transformation model obtained after being trained by all the first sub-training sample sets before the current first sub-training sample set.
[0096] The current spatial transformation model to be trained for each set of first sub-training sample sets is a semi-finished spatial transformation model that has been trained by the previously used first sub-training sample sets.
[0097] The current first sub-training sample set contains n first sample pairs. Therefore, the current first sub-training sample set contains n first mask face images and corresponding n pieces of supervision information. Inputting these n first mask face images into the current spatial transformation model for spatial transformation can obtain corresponding n first effective face images. Among them, the spatial transformation includes at least cropping processing. The first effective face image is a face image obtained after at least cropping processing by the current spatial transformation model. Since the first effective face image is an effective face image obtained during the training stage, the first effective face image may be a partial face not blocked by the mask, or may not be a completely partial face not blocked by the mask.
[0098] The supervision information is the partial face not blocked by the mask obtained through determining cropping processing. The supervision information is the expected result output after each first mask face is input into the spatial transformation model.
[0099] The first effective face images correspond one-to-one with the supervision information.
[0100] S430: Input the n current first effective face images into the trained first feature extraction network to extract the first features of each current first effective face image.
[0101] S440: Input the n pieces of supervision information in the current first sub-training sample set into the trained first feature extraction network to extract the second features of each piece of supervision information.
[0102] Specifically, the trained first feature extraction network is a feature extraction network obtained through the prior art, which will not be elaborated in this application. Inputting the n current first effective face images output by the current spatial transformation model into the trained first feature extraction network can extract the first features of each current first effective face image. Inputting each piece of supervision information in the current first sub-training sample set into the trained first feature extraction network can extract the second features of each piece of supervision information.
[0103] S450: Obtain the first cosine distance between each first feature and the corresponding second feature.
[0104] Specifically, the first effective face images correspond one-to-one with the supervision information. By calculating the first feature of the first effective face image and the second feature of the corresponding supervision information, the first cosine distance between the two can be obtained. Each first sample pair corresponds to a first cosine distance. Therefore, the current first sub-training sample set corresponds to n first cosine distances.
[0105] S460: Obtain a first loss value based on n first cosine distances.
[0106] Specifically,
[0107] Calculate the first loss value according to the loss function and n first cosine distances.
[0108] The loss function can use the Insightface function.
[0109] S470: Update the training parameters of the current spatial transformation model using the first loss value.
[0110] Specifically, the first loss value is used to characterize the training effect of the current spatial transformation model. The smaller the first loss value, the closer the result output by the spatial transformation model is to the supervision information (expected result).
[0111] S480: Determine whether the first loss value is less than or equal to a first threshold.
[0112] S481: If the first loss value is greater than the first threshold, obtain the next first sub-training sample set as the current first sub-training sample set, and execute step S420.
[0113] S482: If the first loss value is less than or equal to the first threshold, end the training, and use the updated spatial transformation model or the current spatial transformation model as the trained spatial transformation model.
[0114] Specifically, if the first loss value is less than or equal to the first threshold, it means that the result output by the current spatial transformation model has reached the expected effect, and the training is completed. The finally obtained trained spatial transformation module can be the current spatial transformation model or the spatial transformation model obtained by updating the training parameters of the current spatial transformation model according to the first loss value.
[0115] In a specific embodiment, the next first sub-training sample set can be: non-repetitively selected from multiple first sub-training sample sets obtained by grouping, or composed of n first sample pairs randomly selected or selected according to a preset rule from the first training sample set at the next training moment.
[0116] In an embodiment, both the initial spatial transformation model and the trained spatial transformation model include a convolutional layer, a first inverted residual layer, an average pooling layer, a second inverted residual layer, a global average pooling layer, and a fully connected layer connected in sequence.
[0117] Specifically, the spatial transformation model includes: 1) a Localisation Network for generating spatial transformation parameters; 2) a Grid Genator for calculating the coordinate correspondence between the target feature map and the original feature map; and 3) a Sampler for performing pixel sampling on the original feature map according to the coordinate correspondence to output the spatially transformed image.
[0118] Among them, the Localisation Network in the spatial transformation model includes a convolutional layer, a first inverted residual layer, an average pooling layer, a second inverted residual layer, a global average pooling layer, and a fully connected layer connected in sequence.
[0119] The convolutional layer is located in the first layer of the Localisation Network, and the size and stride of the convolutional kernel can be set according to actual applications. The size of the convolutional kernel should preferably enable the network to have a sufficient receptive field. For example, a convolutional kernel with a size of 5×5 and a stride of 2 is used. This can reduce the size of the feature map while ensuring a sufficient receptive field, thereby reducing the computational complexity of subsequent steps.
[0120] The first inverted residual layer is located behind the convolutional layer in the Localisation Network and is used for the first feature extraction to obtain the first feature map.
[0121] The average pooling layer is located behind the first inverted residual layer in the Localisation Network and is used to further reduce the size of the first feature map.
[0122] The second inverted residual layer is located behind the average pooling layer in the Localisation Network and is used for the second feature extraction to obtain the second feature map.
[0123] The global average pooling layer is located behind the second inverted residual layer in the Localisation Network and is used for feature map fusion. That is, the first feature map and the second feature map are fused.
[0124] The fully connected layer is located behind the global average pooling layer in the Localisation Network and is used to output the spatial transformation parameters.
[0125] Based on extracting sufficient features to effectively generate spatial transformation parameters, the network fully considers the computational complexity of the modules and reduces the impact of the modules on the model efficiency on mobile devices.
[0126] Table 1 shows the parameter table of each layer of the Localisation Network in an embodiment. Referring to Table 1, an image with a feature vector of 112x112 is input to the convolutional layer, which uses a convolutional kernel with a size of 5x5 and a stride of 2, and the number of channels of the convolutional layer is 32. The convolutional layer outputs a first feature vector of 56*56*32.
[0127] The first feature vector of 56*56*32 is input to the first inverted residual layer. The number of channels of the first inverted residual layer is 64 and the stride is 1. The first inverted residual layer outputs a second feature vector of 56*56*64.
[0128] The second feature vector of 56*56*64 is transmitted as input to the average pooling layer. The number of channels of the average pooling layer is 64, and the stride is 2. The average pooling layer outputs a third feature vector of 28*28*64.
[0129] The third feature vector of 28*28*64 is transmitted as input to the second inverted residual layer. The number of channels of the second inverted residual layer is 128, and the stride is 1. The second inverted residual layer outputs a fourth feature vector of 28*28*128.
[0130] The fourth feature vector of 28*28*128 is transmitted as input to the global average pooling layer. The number of channels of the global average pooling layer is 128, and the stride is 2. The global average pooling layer outputs a one-dimensional feature vector with 128 elements.
[0131] The one-dimensional feature vector is transmitted as input to the fully connected layer. The stride of the fully connected layer is 1. The fully connected layer outputs 6 parameters. These 6 parameters are the spatial transformation parameters.
[0132] Table 1: Parameter table of each layer of the local network
[0133] Operating layer Channel Step size Output Convolution layer 5x5 32 2 56*56*32 First inverted residual layer 64 1 56*56*64 Average pooling layer 64 2 28*28*64 Second inverted residual layer 128 1 28*28*128 Global average pooling layer 128 2 128 Fully connected layer - 1 6
[0134] The above-obtained 6 spatial transformation parameters are input to the grid generator to calculate the coordinate correspondence between the target feature map and the original feature map. The sampler samples the pixels of the original feature map according to the coordinate correspondence obtained by the grid generator, and then outputs the image after spatial transformation.
[0135] The trained spatial transformation model obtained above can crop out the face part in the mask face image that is not blocked by the mask. This enables subsequent face recognition to focus on face features and facilitates face recognition.
[0136] Figure 2 It is a schematic flowchart of a mask face recognition method provided by an embodiment of the present disclosure; refer to Figure 2 , the mask face recognition method includes the following steps:
[0137] S100B: Obtain a mask face image to be recognized.
[0138] Specifically, the mask face image to be recognized is a face image in which the face to be recognized is partially blocked by a mask. The mask face image to be recognized can come from an external device and is transmitted by the external device to the trained mask face recognition model for face recognition.
[0139] S200B: Use the trained mask face recognition model to perform effective face extraction and face feature extraction on the mask face image to be recognized, so as to obtain the effective face features of the face to be recognized in the mask face image to be recognized.
[0140] Specifically, the trained mask face recognition model includes: a trained spatial transformation model for effective face extraction and a trained mask face feature extraction model for feature extraction. The trained spatial transformation model is obtained according to the model generation method of any one of the previous items.
[0141] The effective face feature is the face feature of the face in the face image of the mask to be recognized that is not blocked by the mask.
[0142] The trained mask face recognition model of the present disclosure can perform effective face extraction on the face image of the mask to be recognized, and then perform face feature extraction to obtain effective face features. Performing effective face extraction first can strengthen the recognition of key areas (such as eyes, eyebrows, forehead, etc.) in the partial face image outside the mask that is not blocked by the mask, remove the interference factors of the mask part, and improve the accuracy of face recognition with a mask.
[0143] S300B: Calculate the similarity between the effective face feature of the face to be recognized and the effective face feature of each known face in the pre-stored known face set to obtain a corresponding plurality of similarities.
[0144] Specifically, the known face set can be existing faces. For example, the known face set is a blacklist face set, which stores the face images of each person included in the blacklist and is used to detect whether there are blacklisted persons. The known face set can also be a company employee face set, which stores the face images of all employees of the company and is used to detect whether the visiting person is an employee of the company or an unauthorized outsider. Of course, the known face set is not limited to the above application scenarios, and the known face set of the present application is specifically set in advance according to the actual application scenario.
[0145] The effective face feature of each known face in the known face set can be obtained in advance according to the trained mask face recognition model and stored in the known face database.
[0146] S400B: Perform face recognition on the face to be recognized based on the plurality of similarities.
[0147] Specifically, face recognition usually includes two tasks: face verification and face identification. Face verification is to verify whether two faces in two face images belong to the same person. It is a binary classification problem, and the correct rate of random guessing is 50%. Face identification is to identify the identity of the face to be recognized from a group of faces. This is a multi-classification problem. In either task, a comparison between two face images is required.
[0148] The method for face recognition can be implemented using conventional techniques. The face is recognized by calculating the similarity of face features between two faces. That is to say, it is possible to determine whether two faces belong to the same person by calculating the similarity between the two faces.
[0149] If two faces belong to the same person, the similarity between them must be numerically large; on the contrary, if two faces do not belong to the same person, the value of their similarity must be small.
[0150] In one embodiment, the trained mask face recognition model further includes a trained mask face feature extraction model. Before step S200B, the mask face recognition method further includes the following steps:
[0151] S010B: Embed the trained spatial transformation model into the first layer of the initial mask face feature extraction model to obtain the mask face feature extraction model to be trained.
[0152] S020B: Obtain a second training sample set, which includes a plurality of second mask face images with labeled tags.
[0153] S030B: Use the second training sample set to train the mask face feature extraction model to be trained to obtain the trained mask face feature extraction model.
[0154] Specifically, the mask face feature extraction model is a model for extracting features of the effective face outside the mask. The trained spatial transformation model is used to perform spatial transformation processing such as cropping on the initially obtained mask face image to obtain an effective face image outside the mask.
[0155] Since the second mask face image of the present application includes an image of the mask part, which is an interference factor, therefore, the image of the mask part in the second mask face image can be cropped by the trained spatial transformation model, and the effective face image outside the mask is retained. Using the effective face image outside the mask as a sample for training the initial mask face feature extraction model can better train the mask face feature extraction model, so that the mask face feature extraction model only extracts the effective face features of the effective face part and does not extract the features of the mask part, which is beneficial for face recognition and reduces the calculation amount.
[0156] Of course, the training samples of the trained mask face feature extraction model of the present application may also not use the trained spatial transformation model to perform spatial transformation processing, but directly use the effective face sample image with the effective face cropped to train the mask face feature extraction model to be trained.
[0157] The mask face feature extraction model of the present disclosure can select the MobilefaceNet lightweight network as the basic architecture, which is more suitable for mobile application scenarios.
[0158] In one embodiment, step S200B specifically includes the following steps:
[0159] S210B: Input the mask face image to be recognized into the trained spatial transformation model for spatial transformation to obtain the effective face image of the face to be recognized in the mask face image to be recognized;
[0160] S220B: Input the effective face image of the face to be recognized into the trained mask face feature extraction model to obtain the effective face features of the face to be recognized in the mask face image to be recognized.
[0161] Specifically, the trained mask face recognition model in the present application includes a trained spatial transformation model and a trained mask face feature extraction model.
[0162] The trained spatial transformation model is used to perform spatial transformation processing such as cropping on the initially obtained mask face image to obtain an effective face image outside the mask. The trained mask face feature extraction model is used to extract face features from the effective face obtained by the trained spatial transformation model to obtain effective face features. The effective face features include partial face features such as eyes, eyebrows, and forehead.
[0163] It should be understood that the steps in the above respective processes do not necessarily need to be executed in a fixed order. Unless there is a clear description in this article, the execution of these steps has no strict order limit, and these steps can be executed in other orders. Moreover, at least a part of the steps in the above respective processes can include multiple sub-steps or multiple stages. These sub-steps or stages do not necessarily need to be executed at the same time, but can be executed at different times. The execution order of these sub-steps or stages does not necessarily need to be sequential, but can be executed alternately or in turn with at least a part of other steps or sub-steps or stages of other steps.
[0164] Figure 3 It is a structural block diagram of a model generation device provided in an embodiment of the present disclosure; the model generation device includes:
[0165] The first sample acquisition module 100A is used to acquire a first sample set, and the first sample set includes m first mask face images;
[0166] The first processing module 200A is used to perform a first process on each first mask face image to obtain corresponding supervision information, and the supervision information is the face image in the corresponding first mask face image that is not blocked by the mask;
[0167] The first training sample generation module 300A is configured to form a first sample pair by combining each first mask face image with the corresponding supervision information, so as to obtain a first training sample set including m first sample pairs;
[0168] The first training module 400A is configured to train the initial spatial transformation model by using the first training sample set to obtain a trained spatial transformation model;
[0169] Wherein, the first processing at least includes cropping processing.
[0170] In one embodiment, the first training module 400A specifically includes:
[0171] The sampling module 410A is configured to obtain a current first sub-training sample set including n first mask face images, where n is less than m;
[0172] The first spatial transformation module 420A is configured to input the n first mask face images in the current first sub-training sample set into the current spatial transformation model respectively for spatial transformation, so as to obtain n current first valid face images;
[0173] Wherein, the current first sub-training sample set is: the first sub-training sample set in the first training sample set for training the current spatial transformation model;
[0174] The current spatial transformation model is: the spatial transformation model obtained after being trained by all the first sub-training sample sets before the current first sub-training sample set;
[0175] The first feature extraction module 430A is configured to input the n current first valid face images into the trained first feature extraction network to extract the first features of each current first valid face image;
[0176] The second feature extraction module 440A is configured to input the n supervision information in the current first sub-training sample set into the trained first feature extraction network to extract the second features of each supervision information;
[0177] The first cosine distance obtaining module 450A is configured to obtain the first cosine distance between each first feature and the corresponding second feature;
[0178] The first loss value obtaining module 460A is configured to obtain a first loss value according to the n first cosine distances;
[0179] The first judgment module 470A is configured to judge whether the first loss value is less than or equal to a first threshold;
[0180] The updating module 480A is configured to update the training parameters of the current spatial transformation model by using the first loss value;
[0181] A loop module 481A, configured to, if a first loss value is greater than a first threshold, obtain a next first sub-training sample set as the current first sub-training sample set, and perform operations of respectively inputting n first mask face images in the current first sub-training sample set into the current spatial transformation model for spatial transformation to obtain n current first valid face images;
[0182] An end module 482A, configured to, if the first loss value is less than or equal to the first threshold, end the training, and use the updated spatial transformation model or the current spatial transformation model as the trained spatial transformation model.
[0183] In one embodiment, the first loss value obtaining module 460A is specifically configured to calculate the first loss value according to a loss function and n first cosine distances.
[0184] In one embodiment, both the initial spatial transformation model and the trained spatial transformation model include a convolutional layer, a first inverted residual layer, an average pooling layer, a second inverted residual layer, a global average pooling layer, and a fully connected layer that are connected in sequence.
[0185] Figure 4 The block diagram of a mask face recognition device provided by an embodiment of the present disclosure. Refer to Figure 4 , the mask face recognition device includes:
[0186] An image acquisition module 100B, configured to acquire a mask face image to be recognized;
[0187] The face to be recognized in the mask face image to be recognized is partially blocked by a mask;
[0188] A feature extraction module 200B, configured to use the trained mask face recognition model to perform effective face extraction and face feature extraction on the mask face image to be recognized to obtain effective face features of the face to be recognized in the mask face image to be recognized;
[0189] Wherein, the trained mask face recognition model includes: a trained spatial transformation model for effective face extraction and a trained mask face feature extraction model;
[0190] The effective face features are face features of the face in the mask face image to be recognized that are not blocked by the mask;
[0191] A similarity acquisition module 300B, configured to calculate the similarity between the effective face features of the face to be recognized and the effective face features of each known face in a pre-stored known face set to obtain corresponding multiple similarities;
[0192] An identification module 400B, configured to perform face recognition on the face to be recognized based on the multiple similarities.
[0193] In one embodiment, the trained spatial transformation model is obtained by the model generation device according to any one of the preceding items.
[0194] In one embodiment, the mask face recognition device further includes:
[0195] An embedding module 010B, configured to embed the trained spatial transformation model into the first layer of the initial mask face feature extraction model to obtain a mask face feature extraction model to be trained.
[0196] A first sample acquisition module 020B, configured to acquire a second training sample set, where the second training sample set includes p second mask face images with labeled tags.
[0197] A second training module 030B, configured to train the mask face feature extraction model to be trained by using the second training sample set to obtain a trained mask face feature extraction model.
[0198] In one embodiment, the feature extraction module 200B specifically includes:
[0199] A second spatial transformation module 210B, configured to input the mask face image to be recognized into the trained spatial transformation model for spatial transformation to obtain an effective face image of the face to be recognized in the mask face image to be recognized.
[0200] A sub-feature extraction module 220B, configured to input the effective face image of the face to be recognized into the trained mask face feature extraction model to obtain an effective face feature of the face to be recognized in the mask face image to be recognized.
[0201] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: acquiring a first sample set, where the first sample set includes m first mask face images; performing a first process on each first mask face image to obtain corresponding supervision information, where the supervision information is a face image in the corresponding first mask face image that is not blocked by the mask; forming a first sample pair by combining each first mask face image with the corresponding supervision information to obtain a first training sample set including m first sample pairs; training the initial spatial transformation model by using the first training sample set to obtain a trained spatial transformation model; where the first process includes at least a cropping process.
[0202] In one embodiment, when the computer program is executed by the processor, the steps of any one of the above model generation methods are further implemented.
[0203] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining a first sample set, where the first sample set includes m first masked face images; performing a first process on each first masked face image to obtain corresponding supervision information, where the supervision information is the face image in the corresponding first masked face image that is not blocked by the mask; forming a first sample pair by combining each first masked face image with the corresponding supervision information to obtain a first training sample set including m first sample pairs; training an initial spatial transformation model using the first training sample set to obtain a trained spatial transformation model; where the first process includes at least a cropping process.
[0204] In one embodiment, when the processor executes the computer program, the steps of any one of the above model generation methods are also implemented.
[0205] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented: obtaining a to-be-recognized masked face image; the to-be-recognized face in the to-be-recognized masked face image is partially blocked by a mask; using the trained masked face recognition model to perform effective face extraction and face feature extraction on the to-be-recognized masked face image to obtain the effective face features of the to-be-recognized face in the to-be-recognized masked face image; where the trained masked face recognition model includes: a trained spatial transformation model for effective face extraction, and the trained spatial transformation model is obtained according to any one of the previous model generation methods; the effective face features are the face features of the face in the to-be-recognized masked face image that is not blocked by the mask; calculating the similarity between the effective face features of the to-be-recognized face and the effective face features of each known face in the pre-stored known face set to obtain corresponding multiple similarities; performing face recognition on the to-be-recognized face based on the multiple similarities.
[0206] In one embodiment, when the computer program is executed by the processor, the steps of any one of the above masked face recognition methods are also implemented.
[0207] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored on the memory and executable on the processor. When the processor executes the computer program, the following steps are implemented: obtaining a face image of a mask to be recognized, where the face to be recognized in the face image of the mask to be recognized is partially blocked by the mask; using a trained face recognition model for masks to perform effective face extraction and face feature extraction on the face image of the mask to be recognized, so as to obtain effective face features of the face to be recognized in the face image of the mask to be recognized; wherein, the trained face recognition model for masks includes: a trained spatial transformation model for effective face extraction, and the trained spatial transformation model is obtained according to the model generation method of any one of the previous items; the effective face features are the face features of the face in the face image of the mask to be recognized that is not blocked by the mask; calculating the similarity between the effective face features of the face to be recognized and the effective face features of each known face in a pre-stored set of known faces, so as to obtain corresponding multiple similarities; and performing face recognition on the face to be recognized based on the multiple similarities.
[0208] In one embodiment, when the processor executes the computer program, it also implements each step of any one of the above face recognition methods for masks.
[0209] It should be noted that in this text, relational terms such as "first" and "second" are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or further includes elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including a..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.
[0210] The above are only specific embodiments of the present invention, enabling those skilled in the art to understand or implement the present invention. Various modifications to these embodiments will be obvious to those skilled in the art, and the general principles defined herein can be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention will not be limited to these embodiments shown herein, but rather to the widest scope consistent with the principles and novel features claimed herein.
Claims
1. A model generation method, characterized in that, the method includes: Obtain a first sample set, where the first sample set includes m first masked face images; Perform a first process on each of the first masked face images to obtain corresponding supervision information, where the supervision information is the face image in the corresponding first masked face image that is not blocked by the mask; Form a first sample pair by combining each of the first masked face images with the corresponding supervision information to obtain a first training sample set containing the m first sample pairs; Train an initial spatial transformation model using the first training sample set to obtain a trained spatial transformation model; wherein, the first process at least includes a cropping process; The step of training the initial spatial transformation model using the first training sample set to obtain a trained spatial transformation model includes: Obtain a current first sub-training sample set containing n first masked face images, where n is less than m; Input the n first masked face images into the current spatial transformation model respectively for spatial transformation to obtain n current first valid face images; The current first sub-training sample set is: the first sub-training sample set in the first training sample set that is used to train the current spatial transformation model; The current spatial transformation model is: the spatial transformation model obtained after being trained by all the first sub-training sample sets before the current first sub-training sample set; Input the n current first valid face images into a trained first feature extraction network to extract the first features of each of the current first valid face images; Input the n supervision information in the current first sub-training sample set into a trained first feature extraction network to extract the second features of each of the supervision information; Obtain the first cosine distance between each first feature and the corresponding second feature; Obtain a first loss value according to the n first cosine distances; Determine whether the first loss value is less than or equal to a first threshold; Update the training parameters of the current spatial transformation model using the first loss value; If the first loss value is greater than the first threshold, obtain the next first sub-training sample set as the current first sub-training sample set, and perform the step of inputting the n first masked face images in the current first sub-training sample set into the current spatial transformation model respectively for spatial transformation to obtain n current first valid face images; If the first loss value is less than or equal to the first threshold, end the training, and use the updated spatial transformation model or the current spatial transformation model as the trained spatial transformation model.
2. The model generation method according to claim 1, characterized in that, Both the initial spatial transformation model and the trained spatial transformation model include a convolutional layer, a first inverted residual layer, an average pooling layer, a second inverted residual layer, a global average pooling layer, and a fully connected layer connected in sequence.
3. A masked face recognition method, characterized in that, the masked face recognition method includes: Obtain a masked face image to be recognized; The face to be recognized in the masked face image to be recognized is partially blocked by a mask; Using the trained mask face recognition model to perform effective face extraction and face feature extraction on the mask face image to be recognized, so as to obtain the effective face features of the face to be recognized in the mask face image to be recognized; Wherein, the trained mask face recognition model includes: a trained spatial transformation model for effective face extraction and a trained mask face feature extraction model for feature extraction; The effective face features are the face features of the face in the mask face image to be recognized that are not blocked by the mask; Calculating the similarity between the effective face features of the face to be recognized and the effective face features of each known face in the pre-stored known face set to obtain a corresponding plurality of similarities; Performing face recognition on the face to be recognized based on the plurality of similarities; The trained spatial transformation model is obtained according to the model generation method described in any one of claims 1-2.
4. The face recognition method according to claim 3, characterized in that, Before using the trained mask face recognition model to perform effective face extraction and face feature extraction on the mask face image to be recognized to obtain the effective face features of the face to be recognized in the mask face image to be recognized, the method further includes: Embedding the trained spatial transformation model into the first layer of the initial mask face feature extraction model to obtain a mask face feature extraction model to be trained, Obtaining a second training sample set, the second training sample set including p second mask face images with labeled tags, Training the mask face feature extraction model to be trained by using the second training sample set to obtain a trained mask face feature extraction model.
5. The face recognition method according to claim 4, characterized in that, The using the trained mask face recognition model to perform effective face extraction and face feature extraction on the mask face image to be recognized to obtain the effective face features of the face to be recognized in the mask face image to be recognized includes: Inputting the mask face image to be recognized into the trained spatial transformation model for spatial transformation to obtain an effective face image of the face to be recognized in the mask face image to be recognized; Inputting the effective face image of the face to be recognized into the trained mask face feature extraction model to obtain the effective face features of the face to be recognized in the mask face image to be recognized.
6. A model generation device, characterized in that, The model generation device includes: A first sample acquisition module for acquiring a first sample set, the first sample set including m first mask face images; A first processing module for performing first processing on each of the first mask face images to obtain corresponding supervision information, the supervision information being the face image of the face in the corresponding first mask face image that is not blocked by the mask; A first training sample generation module for forming a first sample pair by combining each of the first mask face images with the corresponding supervision information to obtain a first training sample set including the m first sample pairs; The first training module is used to train the initial spatial transformation model with the first training sample set to obtain a trained spatial transformation model; wherein, the first processing at least includes cropping processing; The training of the initial spatial transformation model with the first training sample set to obtain a trained spatial transformation model includes: Obtain a current first sub-training sample set containing n first mask face images, where n is less than m; Input the n first mask face images into the current spatial transformation model respectively for spatial transformation to obtain n current first valid face images; The current first sub-training sample set is: the first sub-training sample set in the first training sample set for training the current spatial transformation model; The current spatial transformation model is: the spatial transformation model obtained after being trained by all the first sub-training sample sets before the current first sub-training sample set; Input the n current first valid face images into the trained first feature extraction network to extract the first features of each current first valid face image; Input the n supervision information in the current first sub-training sample set into the trained first feature extraction network to extract the second features of each supervision information; Obtain the first cosine distance between each first feature and the corresponding second feature; Obtain a first loss value according to the n first cosine distances; Judge whether the first loss value is less than or equal to a first threshold; Update the training parameters of the current spatial transformation model with the first loss value; If the first loss value is greater than the first threshold, obtain the next first sub-training sample set as the current first sub-training sample set, and execute the step of inputting the n first mask face images in the current first sub-training sample set into the current spatial transformation model respectively for spatial transformation to obtain n current first valid face images; If the first loss value is less than or equal to the first threshold, end the training, and use the updated spatial transformation model or the current spatial transformation model as the trained spatial transformation model.
7. A mask face recognition device characterized in that the face recognition device includes: An image acquisition module for acquiring a mask face image to be recognized; The face to be recognized in the mask face image to be recognized is partially blocked by a mask; A feature extraction module for effectively extracting a face and extracting face features from the mask face image to be recognized by using a trained mask face recognition model, so as to obtain the effective face features of the face to be recognized in the mask face image to be recognized; wherein, the trained mask face recognition model includes: a trained spatial transformation model for effective face extraction and a trained mask face feature extraction model for feature extraction; The effective face features are the face features of the face in the mask face image to be recognized that is not blocked by the mask; A similarity acquisition module for calculating the similarity between the effective face features of the face to be recognized and the effective face features of each known face in a pre-stored known face set to obtain corresponding multiple similarities; An identification module for performing face recognition on the face to be recognized based on the multiple similarities; The trained spatial transformation model is obtained according to the model generation method described in any one of claims 1-2.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the processor is caused to execute the steps of the model generation method described in any one of claims 1-2.
Citation Information
Patent Citations
Image registration method, computer equipment and storage medium
CN110599526A
Face recognition method and device, equipment and medium
CN112036266A