Image recognition method, model training method and device
By adjusting the similarity threshold and optimizing model parameters, the problem of wearing a mask affecting face recognition was solved, accurate user identification in the presence of a mask was achieved, and the accuracy of security verification was improved.
Patent Information
- Application Number
- CN202310533850.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-05-09
- Publication Date
- 2025-10-03
- Estimated Expiration
- 2043-05-09
AI Technical Summary
During the face recognition process, the similarity calculation between the user image wearing a mask and the template image without a mask failed, resulting in the inability to identify the user's identity.
Image features and mask probabilities are obtained through the pre-trained target face recognition model, the similarity threshold is adjusted to identify images wearing masks, feature extraction network and classification network are used for image processing, and the model parameters are optimized in combination with the loss function.
When wearing a mask, the accuracy of facial recognition is improved, ensuring the effectiveness of security verification.
Smart Images

Figure CN116631026B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of image processing technology, and in particular to an image recognition method, a model training method and a device. Background Art
[0002] To ensure the safety of the target location, facial recognition is required when the user enters the target location, and security verification is performed based on the facial recognition results. For example, when a user needs to enter a school, a facial image of the user (which can be called the facial image to be recognized) is collected, and facial recognition is performed on the facial image to be recognized based on multiple template facial images. The multiple template facial images are pre-collected facial images of people on campus. If the recognition result indicates that the facial image to be recognized belongs to a person on campus, the user is determined to have passed security verification and is allowed to enter the school; if the recognition result indicates that the facial image to be recognized belongs to a stranger, the user is determined to have failed security verification and is prohibited from entering the school.
[0003] In related technologies, when performing face recognition on a facial image to be recognized, the similarity between the facial image to be recognized and a template facial image is calculated. If the similarity between the facial image to be recognized and a template facial image exceeds a similarity threshold, the facial image to be recognized is determined to belong to the user to which the template facial image belongs. If the similarity between the facial image to be recognized and each template facial image is not greater than the similarity threshold, the facial image to be recognized is determined to belong to an unfamiliar user.
[0004] However, the user may wear a mask, and the face in the face image to be identified is wearing a mask, while the face in the template face image is not wearing a mask. Recognition based on the template face image without a mask will result in the inability to identify the user to whom the face image to be identified belongs. Summary of the Invention
[0005] The purpose of the embodiments of the present application is to provide an image recognition method, model training method, and device to enable identification of the user to whom the face in the face image to be identified belongs even when the face in the face image to be identified is wearing a mask, thereby improving the accuracy of face recognition. The specific technical solution is as follows:
[0006] In a first aspect of an embodiment of the present application, an image recognition method is provided, the method comprising:
[0007] Obtain the face image to be recognized;
[0008] Inputting the face image to be recognized into a pre-trained target face recognition model, obtaining the image features of the face image to be recognized output by the target face recognition model, and the probability that the face in the face image to be recognized is wearing a mask;
[0009] Calculating the similarity between the image features of the face image to be identified and the image features of a preset template face image to obtain the similarity between the face image to be identified and the template face image; wherein the template face image is a pre-collected face image of the target user;
[0010] If the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, reducing the preset similarity threshold to obtain an adjusted similarity threshold;
[0011] When the similarity between the facial image to be identified and the template facial image is greater than the adjusted similarity threshold, it is determined that the facial image to be identified belongs to the target user to which the template facial image belongs.
[0012] Optionally, obtaining the face image to be recognized includes:
[0013] Get the image to be processed;
[0014] Face detection is performed on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as the face image to be recognized.
[0015] Optionally, performing face detection on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as the face image to be recognized includes:
[0016] Inputting the image to be processed into a pre-trained target face detection model, obtaining candidate image regions containing facial images in the image to be processed output by the target face detection model, and the probability that each candidate image region contains a facial image;
[0017] Based on the area of each candidate image region and / or the probability that each candidate image region contains a facial image, a facial image to be recognized is determined from each candidate image region.
[0018] Optionally, determining the facial image to be identified from each candidate image region based on the area of each candidate image region and / or the probability that each candidate image region contains a facial image includes:
[0019] From each candidate image region, determine the candidate image region with the highest probability of containing a face image as the current image region to be matched;
[0020] Calculate the overlap between the current image area to be matched and other candidate image areas;
[0021] Determine a candidate image region whose overlap with the current image region to be matched is greater than a preset overlap threshold, and obtain a repeated image region of the current image region to be matched;
[0022] From the candidate image regions that have not been screened, excluding the repeated image regions, determine the candidate image region with the highest probability of containing a face image, and use it as the current image region to be matched. Then, return to the step of calculating the degree of overlap between the current image region to be matched and the other candidate image regions until the repeated image region of each candidate image region is determined.
[0023] From the candidate image regions excluding the repeated image regions, the candidate image region with the largest area is determined as the face image to be recognized.
[0024] Optionally, the target face recognition model includes: a feature extraction network and a first classification network;
[0025] Inputting the face image to be identified into a pre-trained target face recognition model, obtaining the image features of the face image to be identified output by the target face recognition model, and the probability that the face in the face image to be identified is wearing a mask, includes:
[0026] Inputting the face image to be recognized into a pre-trained target face recognition model, performing feature extraction on the face image to be recognized through the feature extraction network to obtain image features of the face image to be recognized;
[0027] Normalizing the image features of the face image to be identified through the first classification network to obtain the probability that the face in the face image to be identified is wearing a mask.
[0028] Optionally, if the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, after reducing the preset similarity threshold to obtain an adjusted similarity threshold, the method further includes:
[0029] When the similarity between the facial image to be identified and the template facial image is not greater than the adjusted similarity threshold, it is determined that the facial image to be identified belongs to an unfamiliar user.
[0030] Optionally, after inputting the facial image to be recognized into a pre-trained target facial recognition model and obtaining the image features of the facial image to be recognized output by the target facial recognition model, and the probability that the face in the facial image to be recognized is wearing a mask, the method further includes:
[0031] If the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is greater than the similarity threshold, it is determined that the face image to be identified belongs to the target user to which the template face image belongs;
[0032] If the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is not greater than the similarity threshold, it is determined that the face image to be identified belongs to an unfamiliar user.
[0033] In a second aspect of an embodiment of the present application, a model training method is provided, the method being used to generate a target face recognition model as described in any one of the first aspects above, the method comprising:
[0034] Obtaining a sample facial image and label information of the sample facial image; wherein the label information includes a first label indicating whether the face in the sample facial image is wearing a mask, and a second label indicating the sample user to whom the sample facial image belongs;
[0035] Inputting the sample face image into an initial face recognition model, obtaining the probability of the face in the sample face image wearing a mask and the probability that the sample face image belongs to each sample user, output by the initial face recognition model;
[0036] Calculating a first loss function value based on a first label indicating whether the face in the sample face image is wearing a mask and a probability that the face in the sample face image is wearing a mask;
[0037] calculating a second loss function value based on a second label representing the sample user to which the sample facial image belongs and a probability that the sample facial image belongs to each sample user;
[0038] Calculating a weighted sum of the first loss function value and the second loss function value to obtain a target loss function value;
[0039] The model parameters of the initial face recognition model are adjusted based on the target loss function value until a preset convergence condition is reached to obtain a trained target face recognition model.
[0040] Optionally, the initial face recognition model includes: a feature extraction network, a first classification network, and a second classification network;
[0041] Inputting the sample face image into the initial face recognition model to obtain the probability of the face in the sample face image output by the initial face recognition model wearing a mask, and the probability that the sample face image belongs to each sample user, includes:
[0042] Inputting the sample face image into an initial face recognition model, performing feature extraction on the sample face image through the feature extraction network, and obtaining image features of the sample face image;
[0043] Normalizing the image features of the sample face image using the first classification network to obtain a probability that the face in the sample face image is wearing a mask;
[0044] The image features of the sample face images are normalized by the second classification network to obtain the probability that the sample face images belong to each sample user.
[0045] In a third aspect of the embodiments of the present application, an image recognition device is provided, the device comprising:
[0046] A face image acquisition module for acquiring a face image to be identified, used for acquiring a face image to be identified;
[0047] A face recognition module is configured to input the face image to be recognized into a pre-trained target face recognition model, obtain the image features of the face image to be recognized output by the target face recognition model, and obtain the probability that the face in the face image to be recognized is wearing a mask;
[0048] a similarity determination module, configured to calculate the similarity between the image features of the face image to be identified and the image features of a preset template face image, thereby obtaining the similarity between the face image to be identified and the template face image; wherein the template face image is a pre-collected face image of the target user;
[0049] A similarity threshold adjustment module is configured to reduce a preset similarity threshold to obtain an adjusted similarity threshold if the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold;
[0050] The first recognition result determination module is used to determine that the facial image to be recognized belongs to the target user to which the template facial image belongs when the similarity between the facial image to be recognized and the template facial image is greater than the adjusted similarity threshold.
[0051] Optionally, the module for acquiring the face image to be identified is specifically used to acquire the image to be processed;
[0052] Face detection is performed on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as the face image to be recognized.
[0053] Optionally, the face image to be recognized acquisition module is specifically configured to input the image to be processed into a pre-trained target face detection model, obtain candidate image regions containing face images in the image to be processed output by the target face detection model, and obtain the probability that each candidate image region contains a face image;
[0054] Based on the area of each candidate image region and / or the probability that each candidate image region contains a facial image, a facial image to be recognized is determined from each candidate image region.
[0055] Optionally, the face image acquisition module to be identified is specifically configured to determine, from among the candidate image regions, an image region with the highest probability of containing a face image, as the current image region to be matched;
[0056] Calculate the overlap between the current image area to be matched and other candidate image areas;
[0057] Determine a candidate image region whose overlap with the current image region to be matched is greater than a preset overlap threshold, and obtain a repeated image region of the current image region to be matched;
[0058] From the candidate image regions that have not been screened, excluding the repeated image regions, determine the candidate image region with the highest probability of containing a face image, and use it as the current image region to be matched. Then, return to the step of calculating the degree of overlap between the current image region to be matched and the other candidate image regions until the repeated image region of each candidate image region is determined.
[0059] From the candidate image regions excluding the repeated image regions, the candidate image region with the largest area is determined as the face image to be recognized.
[0060] Optionally, the target face recognition model includes: a feature extraction network and a first classification network;
[0061] The face recognition module is specifically configured to input the face image to be recognized into a pre-trained target face recognition model, perform feature extraction on the face image to be recognized through the feature extraction network, and obtain image features of the face image to be recognized;
[0062] Normalizing the image features of the face image to be identified through the first classification network to obtain the probability that the face in the face image to be identified is wearing a mask.
[0063] Optionally, the device further includes:
[0064] The second recognition result determination module is used to execute in the similarity threshold adjustment module if the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, reduce the preset similarity threshold, and after obtaining the adjusted similarity threshold, execute when the similarity between the face image to be identified and the template face image is not greater than the adjusted similarity threshold, determine that the face image to be identified belongs to an unfamiliar user.
[0065] Optionally, the device further includes:
[0066] a third recognition result determination module, configured to, after the face recognition module inputs the face image to be recognized into a pre-trained target face recognition model to obtain image features of the face image to be recognized output by the target face recognition model and a probability that the face in the face image to be recognized is wearing a mask, determine that the face image to be recognized belongs to the target user to which the template face image belongs if the probability that the face in the face image to be recognized is wearing a mask is not greater than a probability threshold and the similarity between the face image to be recognized and the template face image is greater than the similarity threshold;
[0067] The fourth recognition result determination module is used to determine that the face image to be identified belongs to an unfamiliar user if the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is not greater than the similarity threshold.
[0068] In a fourth aspect of the embodiments of the present application, a model training device is provided, the device being used to generate the target face recognition model described in any one of the first aspects above, the device comprising:
[0069] A sample face image acquisition module, configured to acquire a sample face image and label information of the sample face image; wherein the label information includes a first label indicating whether the face in the sample face image is wearing a mask, and a second label indicating the sample user to whom the sample face image belongs;
[0070] A face recognition module is configured to input the sample face image into an initial face recognition model, and obtain the probability of the face in the sample face image wearing a mask, as output by the initial face recognition model, and the probability that the sample face image belongs to each sample user;
[0071] A first loss function value determination module is configured to calculate a first loss function value based on a first label indicating whether the face in the sample face image is wearing a mask and a probability that the face in the sample face image is wearing a mask;
[0072] a second loss function value determining module, configured to calculate a second loss function value based on a second label representing a sample user to which the sample facial image belongs and a probability that the sample facial image belongs to each sample user;
[0073] a target loss function value determination module, configured to calculate a weighted sum of the first loss function value and the second loss function value to obtain a target loss function value;
[0074] The training module is used to adjust the model parameters of the initial face recognition model based on the target loss function value until a preset convergence condition is reached to obtain a trained target face recognition model.
[0075] Optionally, the initial face recognition model includes: a feature extraction network, a first classification network, and a second classification network;
[0076] The face recognition module is specifically configured to input the sample face image into an initial face recognition model, perform feature extraction on the sample face image through the feature extraction network, and obtain image features of the sample face image;
[0077] Normalizing the image features of the sample face image using the first classification network to obtain a probability that the face in the sample face image is wearing a mask;
[0078] The image features of the sample face images are normalized by the second classification network to obtain the probability that the sample face images belong to each sample user.
[0079] An embodiment of the present application further provides an electronic device, including:
[0080] Memory for storing computer programs;
[0081] The processor is used to implement the image recognition method described in any of the first aspects above, or the model training method described in any of the second aspects above, when executing the program stored in the memory.
[0082] An embodiment of the present application also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements the image recognition method described in any of the first aspects above, or the model training method described in any of the second aspects above.
[0083] An embodiment of the present application also provides a computer program product comprising instructions, which, when executed on a computer, enables the computer to execute any of the image recognition methods described in the first aspect above, or any of the model training methods described in the second aspect above.
[0084] Beneficial effects of the embodiments of the present application:
[0085] An image recognition method provided in an embodiment of the present application can obtain a facial image to be recognized; input the facial image to be recognized into a pre-trained target facial recognition model to obtain image features of the facial image to be recognized output by the target facial recognition model, and the probability that the face in the facial image to be recognized wears a mask; calculate the similarity between the image features of the facial image to be recognized and the image features of a preset template facial image to obtain the similarity between the facial image to be recognized and the template facial image; the template facial image is a facial image of a target user collected in advance; if the probability that the face in the facial image to be recognized wears a mask is greater than a preset probability threshold, reduce the preset similarity threshold to obtain an adjusted similarity threshold; when the similarity between the facial image to be recognized and the template facial image is greater than the adjusted similarity threshold, determine that the facial image to be recognized belongs to the target user to which the template facial image belongs.
[0086] Based on the above processing, based on the pre-trained target face recognition model, the probability that the face in the face image to be identified is wearing a mask can be determined. When the probability that the face in the face image to be identified is wearing a mask is greater than the preset probability threshold, it indicates that the face in the face image to be identified is wearing a mask, and the similarity threshold is reduced to obtain an adjusted similarity threshold. Since the difference between the face image wearing a mask and the face image not wearing a mask is large, the calculated similarity between the face image wearing a mask and the face image not wearing a mask is also low, and the adjusted similarity threshold is small. When the similarity between the face image to be identified and the template face image is greater than the adjusted similarity threshold, it can be determined that the face image to be identified belongs to the target user to which the template face image belongs. That is, even when the face in the face image to be identified is wearing a mask, the user to whom the face image to be identified belongs can be identified, thereby improving the accuracy of face recognition.
[0087] Of course, it is not necessary to achieve all the advantages described above at the same time when implementing any product or method of the present application. BRIEF DESCRIPTION OF THE DRAWINGS
[0088] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other embodiments can also be obtained based on these drawings.
[0089] Figure 1 A flowchart of the first image recognition method provided in an embodiment of the present application;
[0090] Figure 2 A flowchart of a second image recognition method provided in an embodiment of the present application;
[0091] Figure 3 A flowchart of a third image recognition method provided in an embodiment of the present application;
[0092] Figure 4 A flowchart of a fourth image recognition method provided in an embodiment of the present application;
[0093] Figure 5 An image to be processed provided in an embodiment of the present application;
[0094] Figure 6 A flowchart of the first model training method provided in an embodiment of the present application;
[0095] Figure 7 A flowchart of the second model training method provided in an embodiment of the present application;
[0096] Figure 8 A flowchart of a fifth image recognition method provided in an embodiment of the present application;
[0097] Figure 9 A flowchart of a sixth image recognition method provided in an embodiment of the present application;
[0098] Figure 10 A flowchart of a seventh image recognition method provided in an embodiment of the present application;
[0099] Figure 11 A flowchart of an eighth image recognition method provided in an embodiment of the present application;
[0100] Figure 12 A structural diagram of an image recognition device provided in an embodiment of the present application;
[0101] Figure 13 A structural diagram of a model training device provided in an embodiment of the present application;
[0102] Figure 14 A structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0103] The following will be combined with the drawings in the embodiments of this application to clearly and completely describe the technical solutions in the embodiments of this application. Obviously, the embodiments described are only part of the embodiments of this application, not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field based on this application are within the scope of protection of this application.
[0104] In related technologies, the user may wear a mask, so the face in the face image to be identified is wearing a mask, while the face in the template face image is not wearing a mask. Recognition is performed based on the template face image without a mask, resulting in the inability to identify the user to whom the face image to be identified belongs.
[0105] To address the above-mentioned issues, an embodiment of the present application provides an image recognition method, which is applied to an electronic device, which may be a terminal, a server, or the like. The electronic device obtains a facial image to be recognized; inputs the facial image to be recognized into a pre-trained target facial recognition model to obtain image features of the facial image to be recognized output by the target facial recognition model, as well as the probability that the face in the facial image to be recognized is wearing a mask; calculates the similarity between the image features of the facial image to be recognized and the image features of a preset template facial image, and obtains the similarity between the facial image to be recognized and the template facial image; the template facial image is a pre-collected facial image of a target user; if the probability that the face in the facial image to be recognized is wearing a mask is greater than a preset probability threshold, reduces the preset similarity threshold to obtain an adjusted similarity threshold; when the similarity between the facial image to be recognized and the template facial image is greater than the adjusted similarity threshold, it is determined that the facial image to be recognized belongs to the target user to whom the template facial image belongs. Even if the face in the facial image to be recognized is wearing a mask, the user to whom the facial image to be recognized belongs can still be identified, thereby improving the accuracy of facial recognition. Subsequently, security verification is performed based on the recognition result of the facial image to be recognized.
[0106] In one application scenario, when a user needs to enter a target place, security verification is performed based on the recognition result of the facial image to be recognized to ensure the safety of the target place. For example, when a user enters a school, it is determined whether the user is a school staff member based on the recognition result of the facial image to be recognized; or, when a user enters a company, it is determined whether the user is a company employee based on the recognition result of the facial image to be recognized; or, when a user enters a residential community, it is determined whether the user is a residential resident based on the recognition result of the facial image to be recognized.
[0107] In another application scenario, when a user account needs to be logged in, security verification is performed based on the recognition result of the face image to be recognized to determine whether the user currently logging in is the user to which the user account belongs, thereby ensuring the security of the user account.
[0108] See also Figure 1 , Figure 1 This is a flowchart of an image recognition method provided in an embodiment of the present application. The method may include the following steps:
[0109] S101: Obtain a face image to be recognized.
[0110] S102: Input the face image to be recognized into a pre-trained target face recognition model to obtain the image features of the face image to be recognized output by the target face recognition model, and the probability that the face in the face image to be recognized is wearing a mask.
[0111] S103: Calculating the similarity between the image features of the face image to be recognized and the image features of the preset template face image to obtain the similarity between the face image to be recognized and the template face image.
[0112] The template face image is a pre-collected face image of the target user.
[0113] S104: If the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, reduce the preset similarity threshold to obtain an adjusted similarity threshold.
[0114] S105: When the similarity between the face image to be recognized and the template face image is greater than the adjusted similarity threshold, determine that the face image to be recognized belongs to the target user to which the template face image belongs.
[0115] Based on the image processing method provided in the embodiment of the present application, based on the pre-trained target face recognition model, the probability that the face in the face image to be identified is wearing a mask can be determined. When the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, it indicates that the face in the face image to be identified is wearing a mask, and the similarity threshold is reduced to obtain an adjusted similarity threshold. Since the difference between the face image wearing a mask and the face image not wearing a mask is large, the calculated similarity between the face image wearing a mask and the face image not wearing a mask is also low, and the adjusted similarity threshold is small. When the similarity between the face image to be identified and the template face image is greater than the adjusted similarity threshold, it can be determined that the face image to be identified belongs to the target user to which the template face image belongs. That is, even when the face in the face image to be identified is wearing a mask, the user to whom the face image to be identified belongs can be identified, thereby improving the accuracy of face recognition.
[0116] In step S101, the facial image to be recognized is the ROI (Region of Interest) containing the facial image extracted from the image to be processed. The image to be processed is an image of the user captured by an image acquisition device. For example, an image of the user captured by a camera at the entrance of a school when the user needs to enter. Alternatively, an image of the user captured by a mobile phone camera when logging into a user account on the mobile phone.
[0117] In some embodiments, Figure 1 Based on Figure 2 , step S101 may include the following steps:
[0118] S1011: Acquire the image to be processed.
[0119] S1012: Performing face detection on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as a face image to be recognized.
[0120] The target face detection model is a RetinaFace detection model, or the target face detection model is an MTCNN (Multi-Task Convolutional Neural Network) detection model.
[0121] The target face detection model is obtained by training the initial face detection model based on sample images and their label information. The label information for the sample images is the location of the image region containing the face image within the sample image. It should be noted that the sample images used to train the initial face detection model in this embodiment are from a public dataset.
[0122] After acquiring the image to be processed, since the image to be processed also contains other image areas besides the face image, the electronic device can perform face detection on the image to be processed based on a pre-trained target face detection model to obtain the image area containing the face image in the image to be processed (i.e., the alternative image area in subsequent embodiments).
[0123] The target face detection model outputs the location of the candidate image region in the image to be processed. For example, the candidate image region is the image region occupied by the minimum bounding rectangle of the facial image. The target face detection model outputs the coordinates of the vertices of the minimum bounding rectangle in the image to be processed, as well as the coordinates of the key points of the facial image in the image to be processed. Key points of the facial image may include: the center of the left pupil, the center of the right pupil, the tip of the nose, the left and right corners of the mouth, etc.
[0124] The electronic device can extract each candidate image region containing a facial image from the image to be processed according to the position of the candidate image region in the image to be processed, and obtain the facial image to be recognized.
[0125] Based on the above processing, the face image to be recognized can be extracted from the image to be processed, avoiding the influence of other image areas in the image to be processed except the face image on the recognition result, which can improve the accuracy of subsequent face recognition.
[0126] In some embodiments, the image to be processed may contain multiple facial images, and the target face detection model may output multiple candidate image regions. When there is only one facial image to be recognized, the electronic device may determine the candidate image region to which the facial image to be recognized belongs from the multiple candidate image regions as the facial image to be recognized.
[0127] Accordingly, in Figure 2 Based on Figure 3 , step S1012 may include the following steps:
[0128] S10121: Input the image to be processed into a pre-trained target face detection model to obtain candidate image regions containing facial images in the image to be processed output by the target face detection model, and the probability that each candidate image region contains a facial image.
[0129] S10122: Determine a facial image to be recognized from each candidate image region based on the area of each candidate image region and / or the probability that each candidate image region contains a facial image.
[0130] The target face detection model includes: a backbone feature extraction network, an enhanced feature extraction network, and a prediction network. When the target face detection model is a RetinaFace detection model, the backbone feature extraction network can be a MobileNet (lightweight convolutional neural network) network or a ResNet (residual) network; the enhanced feature extraction network includes FPN (Feature Pyramid Network) and SSH (Single Stage Headless, context detection module). FPN can realize the fusion of multi-scale information for face images of different scales, so that the target face detection model can better adapt to complex application scenarios and enhance the robustness of the target face detection model. SSH can expand the context information of the image area that needs to be detected in the image to be processed, so that the final output result of the target face detection model is more accurate.
[0131] The electronic device inputs the image to be processed into the target face detection model. The backbone feature extraction network extracts features from the image to be processed, obtaining a feature vector (referred to as a first feature vector) representing the image features of the image to be processed. The enhanced feature extraction network then fuses the first feature vector to obtain a feature vector (referred to as a second feature vector) representing the image features of the image to be processed. Furthermore, the prediction network maps the second feature vector to obtain candidate image regions containing facial images in the image to be processed, as well as the probability that each candidate image region contains a facial image.
[0132] In one implementation, the candidate image region with the largest area is likely to be the face image requiring facial recognition, and the electronic device may extract the candidate image region with the largest area from the image to be processed to obtain the face image to be recognized. Alternatively, the candidate image region with the highest probability of containing a face image is likely to be the face image requiring facial recognition, and the electronic device may extract the candidate image region with the highest probability of containing a face image from the image to be processed to obtain the face image to be recognized.
[0133] In another implementation, the electronic device can determine the face image to be identified from each candidate image area based on the NMS (Non-Maximum Suppression) algorithm, the area of each candidate image area and the probability that each candidate image area contains a face image.
[0134] Accordingly, in Figure 3 Based on Figure 4 , step S10122 may include the following steps:
[0135] S101221: Determine, from among the candidate image regions, a candidate image region with the highest probability of containing a face image as the current image region to be matched.
[0136] S101222: Calculate the degree of overlap between the current image region to be matched and other candidate image regions.
[0137] S101223: Determine a candidate image region whose degree of overlap with the current image region to be matched is greater than a preset overlap threshold, and obtain a repeated image region of the current image region to be matched.
[0138] S101224: From the candidate image regions that have not been screened and are excluding the repeated image regions, determine the candidate image region with the highest probability of containing a face image as the current image region to be matched, and return to the step of calculating the degree of overlap between the current image region to be matched and the other candidate image regions until the repeated image regions of each candidate image region are determined.
[0139] S101225: Determine the candidate image region with the largest area from the candidate image regions excluding the repeated image regions as the face image to be recognized.
[0140] The electronic device may select the first candidate image region, i.e., the candidate image region with the highest probability of containing a face image, in descending order of probability of containing a face image, to obtain the current image region to be matched. Then, the electronic device calculates the degree of overlap between the current image region to be matched and the other candidate image regions. For example, the electronic device may calculate the IOU (Intersection over Union) of the two candidate image regions based on the following formula (1) as the corresponding degree of overlap.
[0141]
[0142] IOU represents the degree of overlap of two candidate image regions, S1 represents the area of the overlapping region of the two candidate image regions, and S2 represents the area of the combined region of the two candidate image regions.
[0143] For example, see Figure 5 , Figure 5 The image to be processed shown includes a candidate image region 1 and a candidate image region 2. Figure 5 The area filled with horizontal lines is the overlapping area of candidate image area 1 and candidate image area 2. The area formed by the area filled with grid lines and the area filled with horizontal lines is the merged area of candidate image area 1 and candidate image area 2.
[0144] If the degree of overlap between the current image region to be matched and a candidate image region exceeds a preset overlap threshold, it indicates that the current image region to be matched and the candidate image region contain the same facial image. Therefore, facial recognition can be performed only on the current image region to be matched. Accordingly, the electronic device can identify candidate image regions whose degree of overlap with the current image region to be matched exceeds the preset overlap threshold, obtaining the overlapping image region of the current image region to be matched. Subsequently, the electronic device may not process the overlapping image region of the candidate image region.
[0145] Then, the electronic device selects the candidate image region with the highest probability of containing a facial image from the candidate image regions that are not currently screened, in descending order of the probability of containing a facial image, to obtain the current image region to be matched. The candidate image regions that are not currently screened do not contain the determined repeated image region. The electronic device calculates the degree of overlap between the current image region to be matched and the other candidate image regions, and so on, to determine the repeated image region of each candidate image region. If the other candidate image regions except the repeated image region contain different facial images, the electronic device can determine the candidate image region with the largest area from the candidate image regions except the repeated image region, to obtain the facial image to be identified. The facial image to be identified is the facial image that needs to be recognized.
[0146] Based on the above processing, candidate image regions with a high probability of containing facial images and a large area can be determined. The determined candidate image regions are the facial images that require facial recognition (i.e., the facial images to be recognized). Subsequent recognition of the facial images to be recognized can improve the accuracy of subsequent facial recognition. In addition, duplicate image regions with a large degree of overlap can be deleted to prevent redundant calculations and avoid the waste of resources caused by multiple detections of the same facial image, thereby improving the efficiency of facial recognition.
[0147] In some embodiments, before performing face detection on the image to be processed, the electronic device may further perform pre-processing on the image to be processed, for example, performing size normalization, border filling, and other pre-processing on the image to be processed.
[0148] For step S102 , the target face recognition model may be a VGG (Visual Geometry Group) model, or a MobileFaceNet (lightweight face recognition network) model, etc.
[0149] The target face recognition model is obtained by training the initial face recognition model based on the sample face images. Figure 6 , Figure 6 A flowchart of a model training method provided in an embodiment of the present application, which is used to generate a target face recognition model, may include the following steps:
[0150] S601: Obtain sample face images and label information of the sample face images.
[0151] The label information includes a first label indicating whether the face in the sample face image is wearing a mask, and a second label indicating the sample user to whom the sample face image belongs.
[0152] S602: Input the sample face image into the initial face recognition model to obtain the probability of the face wearing a mask in the sample face image output by the initial face recognition model, and the probability that the sample face image belongs to each sample user.
[0153] S603: Calculate a first loss function value based on a first label indicating whether the face in the sample face image wears a mask and a probability that the face in the sample face image wears a mask.
[0154] S604: Calculate a second loss function value based on a second label representing the sample user to which the sample face image belongs and a probability that the sample face image belongs to each sample user.
[0155] S605: Calculate the weighted sum of the first loss function value and the second loss function value to obtain the target loss function value.
[0156] S606: Adjust the model parameters of the initial face recognition model based on the target loss function value until a preset convergence condition is reached to obtain a trained target face recognition model.
[0157] Based on the model training method provided in the embodiment of the present application, a target face recognition model can be obtained. Subsequently, based on the pre-trained target face recognition model, the probability of the face in the face image to be recognized wearing a mask can be determined. When the probability of the face in the face image to be recognized wearing a mask is greater than a preset probability threshold, it indicates that the face in the face image to be recognized is wearing a mask, and the similarity threshold is reduced to obtain an adjusted similarity threshold. Since the difference between the face image wearing a mask and the face image not wearing a mask is large, the calculated similarity between the face image wearing a mask and the face image not wearing a mask is also low, and the adjusted similarity threshold is small. When the similarity between the face image to be recognized and the template face image is greater than the adjusted similarity threshold, it can be determined that the face image to be recognized belongs to the target user to which the template face image belongs. That is, when the face in the face image to be recognized wears a mask, the user to whom the face image to be recognized belongs can also be identified, thereby improving the accuracy of face recognition.
[0158] The electronic device can obtain a sample facial image and pre-labeled label information for the sample facial image. The first label is a value indicating whether the face in the sample facial image is wearing a mask. For example, if the face in the sample facial image is wearing a mask, the first label is 1; if the face in the sample facial image is not wearing a mask, the first label is 0.
[0159] The second label is a vector representing the sample user to which the sample face image belongs. For example, the sample users include: sample user 1, sample user 2, and sample user 3. Sample face image 1 belongs to sample user 1, sample face image 2 belongs to sample user 2, and sample face image 3 belongs to sample user 3. The second label of sample face image 1 is [1, 0, 0]; the second label of sample face image 2 is [0, 1, 0]; and the second label of sample face image 3 is [0, 0, 1].
[0160] It should be noted that the sample facial images in this embodiment come from a public dataset.
[0161] In some embodiments, the initial face recognition model includes: a feature extraction network, a first classification network, and a second classification network. Figure 6 Based on Figure 7 , step S602 may include the following steps:
[0162] S6021: Input the sample face image into the initial face recognition model, extract features of the sample face image through the feature extraction network, and obtain image features of the sample face image.
[0163] S6022: Normalize the image features of the sample face image through the first classification network to obtain the probability that the face in the sample face image is wearing a mask.
[0164] S6023: Normalize the image features of the sample face image through the second classification network to obtain the probability that the sample face image belongs to each sample user.
[0165] When the initial face recognition model is the MobileFaceNet model, the feature extraction network includes a BottleNecks layer and multiple convolutional layers with different parameters and structures, such as normal convolutional layers, depthwise convolutional layers, and point-by-point convolutional layers. The first and second classification networks can be two Softmax (normalization) layers with different parameters.
[0166] The electronic device inputs the sample facial image into the initial face recognition model and extracts features from the sample facial image through the feature extraction network to obtain a feature vector (which may be referred to as a third feature vector) representing the image features of the sample facial image. The third feature vector is then normalized by the first classification network to obtain the probability that the face in the sample facial image is wearing a mask. The electronic device may also normalize the third feature vector using the second classification network to obtain the probability that the sample facial image belongs to each sample user.
[0167] In some embodiments, because facial features are obscured when a mask is worn, the available facial features extracted by the initial face recognition model are reduced. Therefore, the number of network channels in the feature extraction network of the initial face recognition model can be increased. For example, the number of channels in each convolutional layer in the feature extraction network can be increased so that the feature extraction network outputs a 512-dimensional feature vector. Furthermore, the target face recognition model trained based on the initial face recognition model can learn more diverse features. Even when a facial image is largely obscured by a mask, it can maintain good recognition performance, thereby improving the robustness of the target face recognition model.
[0168] The electronic device calculates a first loss function value based on a first label indicating whether a face in the sample face image is wearing a mask and a probability that the face in the sample face image is wearing a mask. The first loss function value may represent a difference between the probability of the face wearing a mask determined based on the initial face recognition model and the first label indicating whether the face in the sample face image is wearing a mask. Exemplarily, the electronic device calculates the first loss function value based on the following formula (2).
[0169] loss Mask =crossEnt(Softmax(logits Mask ),labels Mask_ )(2)
[0170] loss Mask Indicates the first loss function value; labels Mask_ID Indicates the first label of the sample face image; logits Mask represents the image features of the sample face image; Softmax represents the normalization function; and crossEnt represents the cross-entropy loss function. The first classification network is used to handle binary classification tasks. This means that the output of the first classification network is limited to two cases: the face in the face image to be identified is wearing a mask, and the face in the face image to be identified is not wearing a mask. In other words, the first label is limited to 0 and 1.
[0171] The electronic device may also calculate a second loss function value based on a second label indicating the sample user to which the sample face image belongs and the probability that the sample face image belongs to each sample user. The second loss function value may represent the difference between the probability that the sample face image belongs to each sample user determined based on the initial face recognition model and the second label indicating the sample user to which the sample face image belongs. The second loss function value is a loss function value that emphasizes training on negative and positive problems, and has a better effect on the recognition of difficult sample face images such as face images of people wearing masks. Exemplarily, the electronic device may calculate the second loss function value based on the following formula (3).
[0172] loss NPCFace =crossEnt(Softmax(logits NPCFace ),labels Face_ )(3)
[0173] loss NPCFace Represents the second loss function value; labels Face_ Indicates the second label of the sample face image; Face_ID indicates the ID of the sample user to which the sample face image belongs; logits NPCFace Represents the image features of the sample face image; Softmax represents the normalization function; crossEnt represents the cross entropy loss function.
[0174] Then, the electronic device can calculate the weighted sum of the first loss function value and the second loss function value based on the following formula (4) to obtain the target loss function value.
[0175] Loss total =αloss NPCFace +βlossMask (4)
[0176] Loss total represents the target loss function value; α represents the first preset coefficient; loss NPCFace represents the second loss function value; β represents the second preset coefficient; loss Mask Represents the value of the first loss function. α and β both belong to [0, 1], and the sum of α and β is 1. For example, to ensure that the trained target face recognition model meets the performance requirements for both unmasked and masked face recognition, the weights of the first and second loss functions can be balanced, i.e., α = β = 0.5.
[0177] Furthermore, the electronic device can adjust the model parameters of the initial face recognition model based on the target loss function value until a preset convergence condition is reached, thereby obtaining a trained target face recognition model. When the preset convergence condition is reached, it indicates that the difference between the first label and the probability that the face in the sample face image is wearing a mask is small, that is, the probability of the face in the sample face image wearing a mask output by the initial face recognition model is substantially consistent with the first label of whether the face in the actual sample face image is wearing a mask, and the difference between the second label and the probability that the sample face image belongs to each sample user is also small, that is, the sample user to which the sample face image output by the initial face recognition model belongs is substantially consistent with the sample user to which the actual sample face image belongs.
[0178] The preset convergence condition can be set by technical personnel based on experience. For example, the preset convergence condition can be: the number of iterations of training the initial face recognition model reaches a preset number; or the preset convergence condition can also be: the target loss function value obtained by calculating a preset number of times in a row is less than a preset value.
[0179] Based on the above processing, the first classification network of the initial face recognition model is used to determine whether the face in the sample face image is wearing a mask, and the second classification network of the initial face recognition model is used to determine the sample user to which the sample face image belongs. Compared with the related art, the face images of people wearing masks and the face images of people not wearing masks are directly mixed for training, which makes it difficult for the face recognition model to converge, and the trained face recognition model cannot take into account the recognition of faces without masks and the recognition of faces wearing masks well, resulting in limited application scenarios of the face recognition model. In the embodiment of the present application, the first loss function value of the first classification network and the second loss function value of the second classification network are calculated respectively, and the initial face recognition model is trained in combination with the first loss function value and the second loss function value, so that the initial face recognition model can not only focus on the correct classification of the sample face image category without wearing a mask during the training process, but also focus on the correct classification of whether the sample face image is wearing a mask. Furthermore, the target loss function value is calculated based on the first loss function value of the first classification network and the second loss function value of the second classification network, and forward feedback is performed based on the target loss function value to adjust the model parameters of the initial face recognition model, so that the trained target face recognition model can meet the needs of face recognition without wearing a mask and face recognition with wearing a mask.
[0180] Moreover, compared with the related art in which the face recognition model has a complex network structure, resulting in the face recognition model finally trained being larger and difficult to implement in practice, in the embodiment of the present application, by adjusting the network structure of the initial face recognition model and optimizing the loss function of the initial face recognition, the initial face recognition model can converge faster and better, and the ability of the trained target face recognition model to extract features from facial images of people wearing masks is enhanced. The resource consumption and time consumption of face recognition based on the target face recognition model are relatively low, and the target face recognition model is more robust and has strong feasibility, which can improve the accuracy and practicality of subsequent face recognition.
[0181] In some embodiments, when the face image to be identified is identified based on the target face recognition model, the number of template face images may not be obtained, and the probability that the face image to be identified belongs to the user to which each template face image belongs may not be output. Therefore, after the training of the initial face recognition model is completed, the feature extraction network and the first classification network in the trained initial face recognition model can be obtained to obtain the target face recognition model. When the face image to be identified is identified based on the target face recognition model, the image features of the face image to be identified can be obtained only through the feature extraction network of the target face recognition model, and the probability of the face in the face image to be identified wearing a mask can be obtained through the first classification network. Subsequently, based on the image features of the face image to be identified and the probability that the face in the face image to be identified wearing a mask, the user to whom the face image to be identified belongs is determined.
[0182] In some embodiments, before inputting the facial image to be recognized into the target facial recognition model, the electronic device may further pre-process the facial image to be recognized, for example, performing size normalization and other pre-processing on the image to be processed.
[0183] In some embodiments, the target face recognition model includes: a feature extraction network and a first classification network. Figure 1 Based on Figure 8 , step S102 may include the following steps:
[0184] S1021: Input the face image to be recognized into a pre-trained target face recognition model, perform feature extraction on the face image to be recognized through a feature extraction network, and obtain image features of the face image to be recognized.
[0185] S1022: Normalizing the image features of the face image to be identified through the first classification network to obtain the probability that the face in the face image to be identified is wearing a mask.
[0186] After the electronic device inputs the facial image to be recognized into the pre-trained target facial recognition model, it can extract features from the facial image to be recognized through the feature extraction network to obtain a feature vector (which can be called a fourth feature vector) representing the image features of the facial image to be recognized. The fourth feature vector is then normalized by the first classification network to obtain the probability that the face in the facial image to be recognized is wearing a mask.
[0187] For step S103, the template face image is a pre-collected face image of the target user. The template face image can be one, for example, the template face image can be a face image of the user to which the currently logged-in user account belongs.
[0188] There may also be multiple template facial images. For example, the template facial images may include facial images of students and staff members of a school; or, the template facial images may include facial images of company employees; or, the template facial images may include facial images of residents in a residential complex.
[0189] In this embodiment, the template face image is obtained after obtaining authorization from the target user.
[0190] After acquiring the template facial image, the template facial image can be subjected to feature extraction to obtain a feature vector (which can be referred to as the fifth feature vector) representing the image features of the template facial image, and the fifth feature vector representing the image features of the template facial image can be stored in a preset database. After obtaining the fourth feature vector representing the image features of the facial image to be identified, the electronic device can obtain the fifth feature vector representing the image features of the template facial image from the preset database. Furthermore, the electronic device can calculate the similarity between the fourth feature vector and the fifth feature vector based on a similarity algorithm to obtain the similarity between the facial image to be identified and the template facial image. The similarity algorithm can be cosine similarity, Euclidean distance, etc.
[0191] Exemplarily, the electronic device can calculate the similarity between the fourth eigenvector representing the image features of the face image to be identified and the fifth eigenvector representing the image features of the template face image based on the following formula (5) to obtain the similarity between the face image to be identified and the template face image.
[0192]
[0193] cos similarity represents the similarity between the face image to be identified and the template face image; a is the fourth eigenvector representing the image features of the face image to be identified; b is the fifth eigenvector representing the image features of the template face image; ‖a‖ is the modulus of the fourth eigenvector; ‖b‖ is the modulus of the fifth eigenvector.
[0194] For steps S104 and S105, when the probability that the face in the face image to be identified is wearing a mask is greater than the preset probability threshold, it indicates that the face in the face image to be identified is wearing a mask. Since the facial features in the face image wearing a mask are blocked by a large area, the difference between the face image wearing a mask and the face image not wearing a mask is large. Even if the face image to be identified belongs to the target user to which the template face image belongs, the calculated similarity between the face image to be identified and the template face image will be low. Therefore, the electronic device can reduce the preset similarity threshold. Furthermore, when the similarity between the face image to be identified and the template face image is greater than the adjusted similarity threshold, the electronic device can determine that the face image to be identified belongs to the target user to which the template face image belongs.
[0195] Exemplarily, the electronic device determines whether the face in the face image to be identified is wearing a mask based on the following formula (6).
[0196]
[0197] Face Mask Indicates the state of the face in the face image to be recognized wearing a mask, Face Mask A value of 1 indicates that the face in the face image to be identified is wearing a mask. Mask 0 indicates that the face in the face image to be recognized is not wearing a mask; p represents the probability that the face in the face image to be recognized is wearing a mask; u represents the preset probability threshold. Among them, u belongs to (0, 1).
[0198] The electronic device adjusts the preset similarity threshold based on the following formula (7).
[0199]
[0200] θ2 represents the adjusted similarity threshold; θ1 represents the preset similarity threshold; γ represents the preset coefficient; where θ 12 , γ, and θ2 all belong to (0, 1). The above formula (7) indicates that when the face in the face image to be recognized is wearing a mask, the similarity threshold is reduced; when the face in the face image to be recognized is not wearing a mask, the similarity threshold is not adjusted.
[0201] In some embodiments, Figure 1 Based on Figure 9 After step S104, the method may further include the following steps:
[0202] S106: When the similarity between the face image to be identified and the template face image is not greater than the adjusted similarity threshold, it is determined that the face image to be identified belongs to an unfamiliar user.
[0203] The adjusted similarity threshold is smaller. If the similarity between the facial image to be identified and the template facial image is not greater than the adjusted similarity threshold, it indicates that the facial image to be identified does not belong to the target user to which the template facial image belongs. The electronic device can then determine that the facial image to be identified belongs to an unfamiliar user.
[0204] In some embodiments, Figure 1 Based on Figure 10 After step S103, the method may further include the following steps:
[0205] S107: If the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is greater than the similarity threshold, it is determined that the face image to be identified belongs to the target user to which the template face image belongs.
[0206] S108: If the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is not greater than the similarity threshold, it is determined that the face image to be identified belongs to an unfamiliar user.
[0207] When the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, it indicates that the face in the face image to be identified is not wearing a mask. When the face image to be identified belongs to the target user to which the template face image belongs, the calculated similarity between the face image to be identified and the template face image is high. The electronic device can directly determine whether the similarity between the face image to be identified and the template face image is greater than the similarity threshold.
[0208] If the similarity between the facial image to be recognized and the template facial image is greater than a similarity threshold, the facial image to be recognized is determined to belong to the target user to which the template facial image belongs. If the similarity between the facial image to be recognized and the template facial image is less than the similarity threshold, indicating that the facial image to be recognized does not belong to the target user to which the template facial image belongs, the electronic device determines that the facial image to be recognized belongs to an unfamiliar user.
[0209] Based on the above processing, when the face in the face image to be identified is not wearing a mask, the target user to whom the face image to be identified belongs can also be identified. That is, the image recognition method provided in the embodiment of the present application can perform face recognition when the user is wearing a mask, and can also take into account the recognition of faces not wearing masks. The application scenarios of the image recognition method are wider, which improves the scope of application of the image recognition method.
[0210] In some embodiments, when there are multiple template facial images, the electronic device may sort the template facial images in a specified order and sequentially calculate the similarity between the face image to be recognized and each template facial image. The specified order may be the chronological order in which the template facial images were collected.
[0211] If, after calculating the similarity between the facial image to be identified and a template facial image, it is determined that the facial image to be identified belongs to the target user to which the template facial image belongs based on the similarity between the facial image to be identified and the template facial image, the similarity between the facial image to be identified and the other template facial images is no longer calculated. If, after calculating the similarity between the facial image to be identified and each template facial image, the target user to which the facial image to be identified belongs cannot be determined, the facial image to be identified is determined to belong to an unfamiliar user.
[0212] In some embodiments, after obtaining the recognition result of the facial image to be recognized, the method may further include the following steps: when the facial image to be recognized belongs to the target user, determining that the facial image to be recognized has passed the security verification; when the facial image to be recognized belongs to an unfamiliar user, determining that the facial image to be recognized has not passed the security verification.
[0213] In one application scenario, the target users are people inside the target location. For example, if the target location is a school, the target users include students and staff of the school; or, if the target location is a community, the target users include residents and staff of the community.
[0214] When a user attempts to enter a target location, if the target user is identified in the facial image, indicating that the facial image is someone inside the target location, the facial image is deemed to have passed security verification, and the user is allowed to enter the target location. If the facial image is identified as an unfamiliar user, indicating that the facial image is not someone inside the target location, the facial image is deemed to have failed security verification. Subsequently, to ensure the security of the target location, the user is prohibited from entering the target location.
[0215] In another application scenario, when logging into a user account, if the facial image to be recognized is identified as belonging to the target user, indicating that the facial image to be recognized belongs to the user to whom the user account belongs, the facial image to be recognized is determined to have passed security verification, and the user account login is successful. If the facial image to be recognized is identified as belonging to an unfamiliar user, indicating that the facial image to be recognized does not belong to the user to whom the user account belongs, the facial image to be recognized is determined to have failed security verification. Subsequently, to ensure the security of the user account, the login to the user account is stopped.
[0216] See also Figure 11 , Figure 11 A flowchart of an image recognition method provided in an embodiment of the present application, the method may include the following steps:
[0217] S1101: Determine whether a face is detected.
[0218] In this step, the electronic device obtains the image to be processed and performs face detection on the image based on the target face detection model. If the image to be processed does not contain a face image, the electronic device may not perform any processing. If the image to be processed does contain a face image, the electronic device may extract the face image from the image to be processed to obtain the face image to be recognized.
[0219] S1102: After the image is normalized and pre-processed, it is input into the target face recognition model.
[0220] In this step, after obtaining the face image to be recognized, the electronic device can preprocess the face image to be recognized, for example, perform size normalization, border filling and other preprocessing, and input the processed face image to be recognized into the target face recognition model.
[0221] S1103: Determine whether the face is wearing a mask.
[0222] In this step, the target face recognition outputs the probability that the face in the face image to be recognized is wearing a mask. When the probability that the face in the face image to be recognized is wearing a mask is greater than the preset probability threshold, it can be determined that the face in the face image to be recognized is wearing a mask; when the probability that the face in the face image to be recognized is wearing a mask is not greater than the preset probability threshold, it can be determined that the face in the face image to be recognized is not wearing a mask.
[0223] S1104: Adjust the similarity threshold.
[0224] In this step, when the face in the face image to be identified is wearing a mask, the electronic device reduces the preset similarity threshold, that is, adjusts the similarity threshold to obtain the adjusted similarity threshold.
[0225] S1105: Similarity matching calculation.
[0226] In this step, the electronic device can extract image features of the face image to be recognized using the target face recognition model. The electronic device can then calculate the similarity between the image features of the face image to be recognized and the image features of a preset template face image to obtain the similarity between the face image to be recognized and the template face image.
[0227] S1106: Determine whether the similarity is greater than a threshold.
[0228] In this step, if the face in the face image to be recognized is wearing a mask, then it is determined whether the similarity between the face image to be recognized and the template face image is greater than the adjusted similarity threshold. If the face in the face image to be recognized is not wearing a mask, then it is determined whether the similarity between the face image to be recognized and the template face image is greater than the similarity threshold.
[0229] S1107: Identification failed, determined to be a stranger.
[0230] In this step, if the face in the facial image to be recognized is wearing a mask and the similarity between the facial image to be recognized and the template facial image is not greater than the adjusted similarity threshold, it can be determined that the facial image to be recognized belongs to an unfamiliar user. If the face in the facial image to be recognized is not wearing a mask and the similarity between the facial image to be recognized and the template facial image is not greater than the similarity threshold, it can be determined that the facial image to be recognized belongs to an unfamiliar user.
[0231] S1108: Identification is successful, and the matching ID (identity identifier) is displayed.
[0232] In this step, if the face in the facial image to be identified is wearing a mask, and the similarity between the facial image to be identified and the template facial image is greater than the adjusted similarity threshold, it can be determined that the facial image to be identified belongs to the target user to which the template facial image belongs. If the face in the facial image to be identified is not wearing a mask, and the similarity between the facial image to be identified and the template facial image is greater than the similarity threshold, it can be determined that the facial image to be identified belongs to the target user to which the template facial image belongs. The electronic device can then display the ID of the target user.
[0233] Based on the above processing, face recognition is performed when the user is wearing a mask, while also taking into account the recognition of faces not wearing masks. The application scenarios of the image recognition method are broader, and the scope of application of the image recognition method is improved. In addition, in the embodiment of the present application, the probability of the face in the face image to be recognized wearing a mask can also be determined, and the similarity threshold is adaptively adjusted based on the probability of the face in the face image to be recognized wearing a mask. This optimizes the comprehensive recognition performance of the face image to be recognized wearing a mask and the face image to be recognized not wearing a mask, thereby improving the accuracy of face recognition.
[0234] In the technical solution of this application, the operations involved in acquiring, storing, using, processing, transmitting, providing and disclosing the user's facial images are all carried out with the user's authorization.
[0235] It should be noted that the face recognition model in this embodiment is not a face recognition model for a specific user and cannot reflect the personal information of a specific user.
[0236] It should be noted that the two-dimensional face images in this embodiment come from a public dataset.
[0237] and Figure 1 Corresponding to the method embodiment, see Figure 12 , Figure 12 This is a structural diagram of an image recognition device provided in an embodiment of the present application, the device comprising:
[0238] The face image acquisition module 1201 is used to acquire the face image to be recognized;
[0239] The face recognition module 1202 is configured to input the face image to be recognized into a pre-trained target face recognition model, obtain the image features of the face image to be recognized output by the target face recognition model, and obtain the probability that the face in the face image to be recognized is wearing a mask;
[0240] A similarity determination module 1203 is configured to calculate the similarity between the image features of the face image to be identified and the image features of a preset template face image, thereby obtaining the similarity between the face image to be identified and the template face image; wherein the template face image is a pre-collected face image of the target user;
[0241] A similarity threshold adjustment module 1204 is configured to reduce a preset similarity threshold to obtain an adjusted similarity threshold if the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold;
[0242] The first recognition result determination module 1205 is configured to determine that the facial image to be recognized belongs to the target user to which the template facial image belongs when the similarity between the facial image to be recognized and the template facial image is greater than the adjusted similarity threshold.
[0243] Optionally, the to-be-recognized face image acquisition module 1201 is specifically configured to acquire an image to be processed;
[0244] Face detection is performed on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as the face image to be recognized.
[0245] Optionally, the to-be-recognized facial image acquisition module 1201 is specifically configured to input the to-be-processed image into a pre-trained target face detection model, obtain candidate image regions containing facial images in the to-be-processed image output by the target face detection model, and obtain the probability that each candidate image region contains a facial image;
[0246] Based on the area of each candidate image region and / or the probability that each candidate image region contains a facial image, a facial image to be recognized is determined from each candidate image region.
[0247] Optionally, the to-be-recognized face image acquisition module 1201 is specifically configured to determine, from among the candidate image regions, a candidate image region with the highest probability of containing a face image, as the current image region to be matched;
[0248] Calculate the overlap between the current image area to be matched and other candidate image areas;
[0249] Determine a candidate image region whose overlap with the current image region to be matched is greater than a preset overlap threshold, and obtain a repeated image region of the current image region to be matched;
[0250] From the candidate image regions that have not been screened, excluding the repeated image regions, determine the candidate image region with the highest probability of containing a face image, and use it as the current image region to be matched. Then, return to the step of calculating the degree of overlap between the current image region to be matched and the other candidate image regions until the repeated image region of each candidate image region is determined.
[0251] From the candidate image regions excluding the repeated image regions, the candidate image region with the largest area is determined as the face image to be recognized.
[0252] Optionally, the target face recognition model includes: a feature extraction network and a first classification network;
[0253] The face recognition module 1202 is specifically configured to input the face image to be recognized into a pre-trained target face recognition model, perform feature extraction on the face image to be recognized through the feature extraction network, and obtain image features of the face image to be recognized;
[0254] Normalizing the image features of the face image to be identified through the first classification network to obtain the probability that the face in the face image to be identified is wearing a mask.
[0255] Optionally, the device further includes:
[0256] The second recognition result determination module is used to execute in the similarity threshold adjustment module 1204 if the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, reduce the preset similarity threshold, and after obtaining the adjusted similarity threshold, execute when the similarity between the face image to be identified and the template face image is not greater than the adjusted similarity threshold, determine that the face image to be identified belongs to an unfamiliar user.
[0257] Optionally, the device further includes:
[0258] a third recognition result determination module for, after the face recognition module 1202 inputs the face image to be recognized into a pre-trained target face recognition model to obtain image features of the face image to be recognized output by the target face recognition model and the probability that the face in the face image to be recognized is wearing a mask, determining that the face image to be recognized belongs to the target user to which the template face image belongs if the probability that the face in the face image to be recognized is wearing a mask is not greater than a probability threshold and the similarity between the face image to be recognized and the template face image is greater than the similarity threshold;
[0259] The fourth recognition result determination module is used to determine that the face image to be identified belongs to an unfamiliar user if the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is not greater than the similarity threshold.
[0260] Based on the image processing device provided by the embodiment of the present application, based on the pre-trained target face recognition model, the probability of the face in the face image to be identified wearing a mask can be determined. When the probability of the face in the face image to be identified wearing a mask is greater than the preset probability threshold, it indicates that the face in the face image to be identified is wearing a mask, and the similarity threshold is reduced to obtain an adjusted similarity threshold. Since the difference between the face image wearing a mask and the face image not wearing a mask is large, the calculated similarity between the face image wearing a mask and the face image not wearing a mask is also low, and the adjusted similarity threshold is small. When the similarity between the face image to be identified and the template face image is greater than the adjusted similarity threshold, it can be determined that the face image to be identified belongs to the target user to which the template face image belongs. That is, even when the face in the face image to be identified is wearing a mask, the user to whom the face image to be identified belongs can be identified, thereby improving the accuracy of face recognition.
[0261] and Figure 6 Corresponding to the method embodiment, see Figure 13 , Figure 13 This is a structural diagram of a model training device provided in an embodiment of the present application, wherein the device is used to generate any of the target face recognition models described above, and the device includes:
[0262] The sample facial image acquisition module 1301 is configured to acquire a sample facial image and label information of the sample facial image; wherein the label information includes a first label indicating whether the face in the sample facial image is wearing a mask, and a second label indicating the sample user to whom the sample facial image belongs;
[0263] A face recognition module 1302 is configured to input the sample face image into an initial face recognition model, and obtain the probability of the face in the sample face image wearing a mask, as output by the initial face recognition model, and the probability that the sample face image belongs to each sample user;
[0264] A first loss function value determining module 1303 is configured to calculate a first loss function value based on a first label indicating whether the face in the sample face image is wearing a mask and a probability that the face in the sample face image is wearing a mask;
[0265] A second loss function value determining module 1304 is configured to calculate a second loss function value based on a second label representing the sample user to which the sample facial image belongs and a probability that the sample facial image belongs to each sample user;
[0266] A target loss function value determination module 1305 is configured to calculate a weighted sum of the first loss function value and the second loss function value to obtain a target loss function value;
[0267] The training module 1306 is used to adjust the model parameters of the initial face recognition model based on the target loss function value until a preset convergence condition is reached to obtain a trained target face recognition model.
[0268] Optionally, the initial face recognition model includes: a feature extraction network, a first classification network, and a second classification network;
[0269] The face recognition module 1302 is specifically configured to input the sample face image into an initial face recognition model, perform feature extraction on the sample face image through the feature extraction network, and obtain image features of the sample face image;
[0270] Normalizing the image features of the sample face image using the first classification network to obtain a probability that the face in the sample face image is wearing a mask;
[0271] The image features of the sample face images are normalized by the second classification network to obtain the probability that the sample face images belong to each sample user.
[0272] Based on the model training device provided in the embodiment of the present application, a target face recognition model can be obtained. Subsequently, based on the pre-trained target face recognition model, the probability of the face in the face image to be recognized wearing a mask can be determined. When the probability of the face in the face image to be recognized wearing a mask is greater than a preset probability threshold, it indicates that the face in the face image to be recognized is wearing a mask, and the similarity threshold is reduced to obtain an adjusted similarity threshold. Since the difference between the face image wearing a mask and the face image not wearing a mask is large, the calculated similarity between the face image wearing a mask and the face image not wearing a mask is also low, and the adjusted similarity threshold is small. When the similarity between the face image to be recognized and the template face image is greater than the adjusted similarity threshold, it can be determined that the face image to be recognized belongs to the target user to which the template face image belongs. That is, even when the face in the face image to be recognized is wearing a mask, the user to whom the face image to be recognized belongs can be identified, thereby improving the accuracy of face recognition.
[0273] The present application also provides an electronic device, such as Figure 14 Shown, including:
[0274] Memory 1401, used for storing computer programs;
[0275] The processor 1402 is used to implement the steps of any of the above-mentioned image recognition methods or any of the above-mentioned model training methods when executing the program stored in the memory 1401.
[0276] Furthermore, the electronic device may further include a communication bus and / or a communication interface, and the processor 1402, the communication interface, and the memory 1401 communicate with each other via the communication bus.
[0277] The communication bus mentioned in the electronic device mentioned above may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus. This communication bus can be divided into address buses, data buses, control buses, etc. For ease of illustration, only one thick line is used in the figure, but this does not mean that there is only one bus or only one type of bus. The communication interface is used for communication between the electronic device mentioned above and other devices.
[0278] The memory may include random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage. Alternatively, the memory may be at least one storage device located away from the processor.
[0279] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components.
[0280] In another embodiment provided in the present application, a computer-readable storage medium is also provided, which stores a computer program. When the computer program is executed by a processor, it implements the steps of any of the above-mentioned image recognition methods or the steps of any of the above-mentioned model training methods.
[0281] In another embodiment provided in the present application, a computer program product comprising instructions is also provided, which, when executed on a computer, enables the computer to execute any image recognition method in the above embodiments, or steps of any model training method in the above embodiments.
[0282] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center via a wired (e.g., coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or other storage media (e.g., a solid-state drive (SSD)).
[0283] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply the existence of any such actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device comprising a series of elements includes not only those elements, but also other elements not explicitly listed, or elements inherent to such process, method, article, or device. In the absence of further limitations, an element defined by the phrase "comprising a ..." does not exclude the presence of other identical elements in the process, method, article, or device comprising the element.
[0284] Each embodiment in this specification is described in a related manner. Similar portions between embodiments can be referenced to each other. Each embodiment focuses on the differences between other embodiments. In particular, the device, electronic device, computer-readable storage medium, and computer program product embodiments are generally similar to the method embodiments, so their descriptions are relatively simple. For related portions, reference can be made to the descriptions of the method embodiments.
[0285] The above description is only a preferred embodiment of the present application and is not intended to limit the scope of protection of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application are included in the scope of protection of the present application.
Claims
1. An image recognition method, characterized in that: The method comprises: Obtain the face image to be recognized; Inputting the face image to be recognized into a pre-trained target face recognition model, obtaining the image features of the face image to be recognized output by the target face recognition model, and the probability that the face in the face image to be recognized is wearing a mask; Calculating the similarity between the image features of the face image to be identified and the image features of a preset template face image to obtain the similarity between the face image to be identified and the template face image; wherein the template face image is a pre-collected face image of the target user; and the face in the template face image is not wearing a mask; If the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, reducing the preset similarity threshold to obtain an adjusted similarity threshold; When the similarity between the face image to be identified and the template face image is greater than the adjusted similarity threshold, determining that the face image to be identified belongs to the target user to which the template face image belongs; After inputting the face image to be recognized into a pre-trained target face recognition model and obtaining the image features of the face image to be recognized output by the target face recognition model, and the probability that the face in the face image to be recognized is wearing a mask, the method further includes: If the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is greater than the similarity threshold, it is determined that the face image to be identified belongs to the target user to whom the template face image belongs; If the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is not greater than the similarity threshold, it is determined that the face image to be identified belongs to an unfamiliar user.
2. The method according to claim 1, characterized in that The step of obtaining a face image to be recognized includes: Get the image to be processed; Face detection is performed on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as the face image to be recognized.
3. The method according to claim 2, characterized in that The step of performing face detection on the image to be processed based on a pre-trained target face detection model to obtain an image region containing a face image in the image to be processed as the face image to be recognized includes: Inputting the image to be processed into a pre-trained target face detection model, obtaining candidate image regions containing facial images in the image to be processed output by the target face detection model, and the probability that each candidate image region contains a facial image; Based on the area of each candidate image region and / or the probability that each candidate image region contains a facial image, a facial image to be recognized is determined from each candidate image region.
4. The method according to claim 3, characterized in that The determining of the face image to be recognized from each candidate image region based on the area of each candidate image region and / or the probability that each candidate image region contains a face image includes: From each candidate image region, determine the candidate image region with the highest probability of containing a face image as the current image region to be matched; Calculate the overlap between the current image area to be matched and other candidate image areas; Determine a candidate image region whose overlap with the current image region to be matched is greater than a preset overlap threshold, and obtain a repeated image region of the current image region to be matched; From the candidate image regions that have not been screened, excluding the repeated image regions, determine the candidate image region with the highest probability of containing a face image, and use it as the current image region to be matched. Then, return to the step of calculating the degree of overlap between the current image region to be matched and the other candidate image regions until the repeated image region of each candidate image region is determined. From the candidate image regions excluding the repeated image regions, the candidate image region with the largest area is determined as the face image to be recognized.
5. The method according to claim 1, wherein The target face recognition model includes: a feature extraction network and a first classification network; Inputting the face image to be identified into a pre-trained target face recognition model, obtaining the image features of the face image to be identified output by the target face recognition model, and the probability that the face in the face image to be identified is wearing a mask, includes: Inputting the face image to be recognized into a pre-trained target face recognition model, performing feature extraction on the face image to be recognized through the feature extraction network to obtain image features of the face image to be recognized; Normalizing the image features of the face image to be identified through the first classification network to obtain the probability that the face in the face image to be identified is wearing a mask.
6. The method according to claim 1, characterized in that If the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold, the preset similarity threshold is reduced to obtain an adjusted similarity threshold, and the method further includes: When the similarity between the facial image to be identified and the template facial image is not greater than the adjusted similarity threshold, it is determined that the facial image to be identified belongs to an unfamiliar user.
7. The method according to claim 1, characterized in that The target face recognition model is trained based on the following method: Obtaining a sample facial image and label information of the sample facial image; wherein the label information includes a first label indicating whether the face in the sample facial image is wearing a mask, and a second label indicating the sample user to whom the sample facial image belongs; Inputting the sample face image into an initial face recognition model, obtaining the probability of the face in the sample face image wearing a mask and the probability that the sample face image belongs to each sample user, output by the initial face recognition model; Calculating a first loss function value based on a first label indicating whether the face in the sample face image is wearing a mask and a probability that the face in the sample face image is wearing a mask; calculating a second loss function value based on a second label representing the sample user to which the sample facial image belongs and a probability that the sample facial image belongs to each sample user; Calculating a weighted sum of the first loss function value and the second loss function value to obtain a target loss function value; The model parameters of the initial face recognition model are adjusted based on the target loss function value until a preset convergence condition is reached to obtain a trained target face recognition model.
8. The method according to claim 7, characterized in that The initial face recognition model includes: a feature extraction network, a first classification network and a second classification network; Inputting the sample face image into the initial face recognition model to obtain the probability of the face in the sample face image output by the initial face recognition model wearing a mask, and the probability that the sample face image belongs to each sample user, includes: Inputting the sample face image into an initial face recognition model, performing feature extraction on the sample face image through the feature extraction network, and obtaining image features of the sample face image; Normalizing the image features of the sample face image using the first classification network to obtain a probability that the face in the sample face image is wearing a mask; The image features of the sample face images are normalized by the second classification network to obtain the probability that the sample face images belong to each sample user.
9. An image recognition device, characterized in that: The device comprises: A face image acquisition module for acquiring a face image to be identified, used for acquiring a face image to be identified; A face recognition module is configured to input the face image to be recognized into a pre-trained target face recognition model, obtain the image features of the face image to be recognized output by the target face recognition model, and obtain the probability that the face in the face image to be recognized is wearing a mask; a similarity determination module, configured to calculate the similarity between the image features of the face image to be identified and the image features of a preset template face image, thereby obtaining the similarity between the face image to be identified and the template face image; wherein the template face image is a pre-collected face image of the target user; and the face in the template face image is not wearing a mask; A similarity threshold adjustment module is configured to reduce a preset similarity threshold to obtain an adjusted similarity threshold if the probability that the face in the face image to be identified is wearing a mask is greater than a preset probability threshold; a first recognition result determination module, configured to determine that the facial image to be recognized belongs to the target user to which the template facial image belongs, when the similarity between the facial image to be recognized and the template facial image is greater than the adjusted similarity threshold; The device further comprises: a third recognition result determination module, configured to, after the face recognition module inputs the face image to be recognized into a pre-trained target face recognition model to obtain image features of the face image to be recognized output by the target face recognition model and a probability that the face in the face image to be recognized is wearing a mask, determine that the face image to be recognized belongs to the target user to which the template face image belongs if the probability that the face in the face image to be recognized is wearing a mask is not greater than a probability threshold and the similarity between the face image to be recognized and the template face image is greater than the similarity threshold; The fourth recognition result determination module is used to determine that the face image to be identified belongs to an unfamiliar user if the probability that the face in the face image to be identified is wearing a mask is not greater than the probability threshold, and the similarity between the face image to be identified and the template face image is not greater than the similarity threshold.
10. An electronic device, characterized in that: include: Memory for storing computer programs; A processor, configured to implement the method according to any one of claims 1 to 8 when executing a program stored in a memory.
11. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 8 is implemented.
Citation Information
Patent Citations
Face recognition method and device, computer equipment and storage medium
CN111523414A
AI-based face recognition verification management method, system and equipment and storage medium
CN115810214A