Mask face recognition method, device and equipment based on knowledge distillation and memory enhancement method and storage medium
By using knowledge distillation and memory enhancement methods in facial recognition, the improved MobileNetV4 model and memory enhancement module are built, which solves the problem of degradation of recognition accuracy of traditional facial recognition algorithms under mask occlusion, and achieves more efficient feature extraction and recognition accuracy.
Patent Information
- Application Number
- CN202510102727.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-01-22
AI Technical Summary
Traditional face recognition algorithms reduce recognition accuracy under mask occlusion, and training deep learning networks requires a large number of face images and their labels, which increases the workload and is difficult to generalize to other data sets.
Using a method based on knowledge distillation and memory enhancement, the training method and detection effect of the model are improved by building a face database, extracting features, building a knowledge distillation model, comparing face databases and memory enhancement.
It improves the facial recognition performance and detection speed under mask occlusion, reduces storage and computing costs, enhances the model's feature extraction ability and recognition accuracy, and is suitable for complex environments such as underground environments.
Smart Images

Figure CN120183010A_ABST
Abstract
Description
Technical Field
[0001] The present invention provides a mask face recognition method, device, equipment and storage medium based on a knowledge distillation and memory enhancement method, belonging to the technical field of face recognition. Background Technique
[0002] Face recognition technology is a biometric-based identity verification technology that confirms identity by analyzing and recognizing an individual's facial features. This technology relies on the combination of multiple technologies such as computer vision, image processing, machine learning, and artificial intelligence, and can identify a person's facial features from images or videos. Face recognition technology can generally be divided into a feature point-based recognition method and a deep learning-based detection method. The feature point-based method divides the human face into several key feature points and realizes face recognition by comparing the feature points; the deep learning-based method includes methods based on convolutional neural networks, etc. First, the face in the image is intercepted, and then the intercepted face image is processed to encode the feature vector.
[0003] Among them, the feature point-based method has certain limitations. If detection is carried out while wearing a mask and a safety helmet, it may not be possible to effectively extract the feature points on the human face, especially areas such as the chin contour and nose that are severely blocked, thereby reducing the recognition accuracy. Using the deep learning-based method for mask and safety helmet face recognition has a great advantage in recognition accuracy compared with the traditional feature point method. However, simply using the deep learning-based method may also result in false detection and missed detection in the case of mask occlusion, thus affecting the performance of the overall algorithm.
[0004] In addition, to train a deep learning network with high accuracy, a large number of face images and their corresponding labels are required, which increases the workload of relevant personnel to a certain extent. The trained network is also only applicable to the corresponding dataset and is difficult to generalize to other types of face datasets.
[0005] The underground environment is complex, with insufficient light, multiple layers, high humidity, and miners often wear safety helmets, masks and other equipment, which increases the recognition difficulty. Therefore, the face recognition method underground has higher requirements for the recognition algorithm. Summary of the Invention
[0006] In order to solve the problem that the recognition accuracy of traditional face recognition algorithms decreases when blocked by a mask, the present invention proposes a mask face recognition method, device, equipment and storage medium based on a knowledge distillation and memory enhancement method, which can improve the overall training method of the model and the subsequent detection effect.
[0007] The technical solution adopted by the present invention is: a mask face recognition method based on a knowledge distillation and memory enhancement method, including the following steps:
[0008] S1: Construct a face database: Add masks to the existing face image database, and store the face images with added masks in the face image database to obtain the face database;
[0009] S2: Face mask feature extraction: Convert each image in the face database into a feature vector through a feature extraction model and store it;
[0010] S3: Construct a knowledge distillation model: The knowledge distillation model includes a teacher model and a student model, and its specific implementation steps are as follows:
[0011] (1) Teacher model construction: Read the pre-trained InceptionResNetv1 model as the teacher model to guide the training of the student model;
[0012] (2) Student model construction: Construct an improved MobileNetv4 model as the student model. The improved MobileNetV4 model includes FusedIB blocks, MSIB blocks, and ExtraDW blocks composed of depthwise separable convolutions. The structure of the improved MobileNetV4 model is: The input passes through Conv2D, two FusedIB blocks, an ExtraDW block, four MSIB blocks, ConvNext, two ExtraDW blocks, four MSIB blocks, an average pooling layer, and two Conv2D in sequence;
[0013] Among them, the MSIB block first uses a 1×1 convolution to expand the number of channels of the feature map to twice the original number of channels, then divides the feature map into 4 parts, and sequentially passes through convolutions with a kernel size of 3×3, 5×5, 3×3, a diffusion coefficient of 2, 5×5, a diffusion coefficient of 2, and then restores the number of channels through a 1×1 convolution and performs information interaction between different channels;
[0014] (3) Teacher knowledge transfer: Use a loss function to align the output of the student model with the output of the teacher model, and transfer the knowledge of the teacher model to the student model in this way,
[0015] (4) Student model storage: Save the student model that has obtained knowledge, and use the saved student model as the feature extraction model;
[0016] S4: Face database comparison: Collect face data and input it into the feature extraction model, perform similarity measurement on the extracted features and the features in the feature database, and output the face information corresponding to the feature with the highest similarity;
[0017] S5: Face Memory Enhancement: Perform memory enhancement operations on the corresponding features in the feature database, and update and store the enhanced data in the feature database.
[0018] Further, the specific implementation steps of step S1 are as follows:
[0019] (1) Extract face images from the database: Read each image from the image database, and each image represents a face information.
[0020] (2) Face feature point conversion: For the extracted image, first use the MTCNN model to extract the face image in the image, and then use the dlib library to convert the face into 68 feature key points, and obtain the absolute coordinates of the feature key points relative to the face image.
[0021] (3) Mask selection: Based on the input parameters, select a specified mask image and cover it with different colors. If no parameters are specified, randomly select the mask type and color for addition.
[0022] (4) Mask addition: Obtain the absolute coordinates corresponding to the chin contour key point and the nose bridge key point, adjust the size and rotation angle of the mask image, and cover the adjusted mask image to the chin position of the face to obtain a face image with a mask added.
[0023] (5) Storage of face mask images: Save the obtained face mask images into the image database to obtain a face database containing normal faces and faces wearing masks.
[0024] Further, the specific implementation steps of step S2 are as follows:
[0025] (1) Extract the images in the face database constructed in step S1: According to the input parameters, read the face images / face mask images in the face database. If no parameters are specified, read the face mask images.
[0026] (2) Face image feature extraction: Input the extracted images into the feature extraction model, and the feature extraction model converts each image into a feature vector.
[0027] (3) Storage of face mask feature vectors: Save the converted feature vectors together with their corresponding face user information into the feature database.
[0028] Further, the specific implementation steps of step S4 are as follows:
[0029] (1) Feature database reading: According to the selected parameters, extract face features / face mask features and their corresponding face user information from the feature database. If no parameters are specified, read the face mask features.
[0030] (2) Real-time camera detection: Turn on the device camera and capture the camera screen image at regular intervals;
[0031] (3) Real-time feature extraction: Input the captured camera screen into the feature extraction model to obtain the feature vector corresponding to the real-time face;
[0032] (4) Feature comparison: Perform similarity measurement on the feature vectors corresponding to the real-time face and the read features in sequence, and determine the features to be added to the candidate list according to the threshold. If it is greater than the threshold, add it; otherwise, do not add it;
[0033] (5) Face information output: After comparing the face features and the read features, sort the obtained candidate list and output the face user information corresponding to the feature with the highest similarity.
[0034] Further, the specific implementation steps of step S5 are as follows:
[0035] (1) Memory feature collection: During the face data comparison in step S4, the detected results are transmitted to the memory feature list in real time through the keyboard. The transmitted memory feature list stores the detected face information and the corresponding real-time face features at this moment;
[0036] (2) Memory enhancement: Use the features in the memory feature list to update the features in the feature database in sequence;
[0037] (3) Updated feature storage: Store the updated features in the feature database.
[0038] Further, using the features in the memory feature list to update the features in the feature database in sequence is achieved by performing weighted averaging on the two features to enhance the features.
[0039] A mask face recognition device based on knowledge distillation and memory enhancement method, comprising:
[0040] Face mask addition module: Used to add masks to each image in the existing face image database to obtain a face database containing face images and face mask images;
[0041] Face mask feature extraction module: Used to encode the face image to obtain a feature vector;
[0042] Knowledge Distillation Model Learning Module: It is used to use the obtained student model as a feature extraction model according to the knowledge distillation model. The student model adopts the improved MobileNetV4 model, which includes FusedIB blocks, MSIB blocks, and ExtraDW blocks composed of depthwise separable convolutions. The structure of the improved MobileNetV4 model is as follows: the input passes through Conv2D, two FusedIB blocks, an ExtraDW block, four MSIB blocks, ConvNext, two ExtraDW blocks, four MSIB blocks, an average pooling layer, and two Conv2D in sequence;
[0043] Among them, the MSIB block first uses a 1×1 convolution to expand the number of channels of the feature map to twice the original number of channels, and then divides the feature map into 4 parts, and sequentially passes through convolutions with a kernel size of 3×3, 5×5, 3×3, a diffusion coefficient of 2, 5×5, and a diffusion coefficient of 2, and then restores the number of channels through a 1×1 convolution and performs information interaction between different channels;
[0044] Face Database Comparison Module: It is used to perform face recognition by measuring the similarity between real-time features and features in the feature database, and to achieve real-time feedback of recognition results through face bounding box selection and information annotation;
[0045] Face Memory Enhancement Module: It is used to save the face feature data when the recognition is correct and update the face data already stored in the feature database.
[0046] A computer device includes a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, it implements the steps of the mask face recognition method based on the knowledge distillation and memory enhancement method.
[0047] A computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, it implements the steps of the mask face recognition method based on the knowledge distillation and memory enhancement method.
[0048] A mask face recognition device based on the knowledge distillation and memory enhancement method includes a computer device and a camera running on a portable mobile device.
[0049] The beneficial effects of the present invention compared with the prior art are as follows: The face mask addition module adds a preselected mask to the face and stores it in the database, which improves the face recognition performance and detection speed under mask occlusion, and effectively avoids the problems of missed detection and false detection of ordinary face recognition algorithms for mask occlusion; the face mask feature extraction module represents the image as a 512-dimensional face feature vector, greatly reducing the storage and calculation costs; the knowledge distillation model learning module provides an unsupervised learning method, which can train any deep learning network to be used as a feature extraction module; different similarity measurement methods are implemented in the face database comparison module, which can cope with different application scenarios; the face memory enhancement module enhances the memory of the vector representation of the face features, improving the detection accuracy. This method is applicable to the face recognition scenario with mask occlusion. Description of the Drawings
[0050] The present invention will be further described below with reference to the accompanying drawings:
[0051] Figure 1 It is a flow chart of a face mask recognition method based on knowledge distillation and memory enhancement;
[0052] Figure 2 It is the structural diagram of the teacher network: InceptionResNetV1 model;
[0053] Figure 3 It is the structural diagram of the student network: the improved MobileNetV4 model;
[0054] Figure 4 It is the knowledge transfer flow chart of the teacher network of the knowledge distillation method;
[0055] Figure 5 It is the memory enhancement flow chart;
[0056] Figure 6 It is a schematic diagram of adding a face mask;
[0057] Figure 7 It is a schematic diagram of the result of real-time face recognition using the method of the present invention (under bright conditions);
[0058] Figure 8 It is a schematic diagram of the result of real-time face recognition using the method of the present invention (under dim conditions). Detailed Embodiment
[0059] As Figures 1 to 8 shown, the present invention provides a face mask recognition method based on the knowledge distillation and memory enhancement method, which can be used in scenarios with complex or dim environments such as underground mines, and can effectively recognize faces under mask occlusion. This method combines traditional face recognition algorithms and deep learning-based face recognition algorithms. The process of this method includes the following steps:
[0060] S1: Construct a face database, and the specific implementation steps are as follows:
[0061] (1) Extract face images from the database: Read each image from the image database, and each image represents a face information;
[0062] (2) Convert facial feature points: For the extracted images, first use the MTCNN model to extract the face images in the images, and then use the dlib library to convert the faces into 68 feature key points, and obtain the absolute coordinates of the feature key points relative to the face images;
[0063] (3) Select a mask: Based on the input parameters, select a specified mask image and cover it with different colors. If no parameters are specified, randomly select the mask type and color for addition;
[0064] (4) Add the mask: Obtain the absolute coordinates corresponding to the key points of the chin contour and the key points of the nose bridge, adjust the size and rotation angle of the mask image, and cover the adjusted mask image to the chin position of the face to obtain a face image with a mask added;
[0065] (5) Store the face mask images: Save the obtained face mask images into the image database to obtain a face database containing normal faces and faces wearing masks.
[0066] S2: Extract face mask features, and the specific implementation steps are as follows:
[0067] (1) Extract the images in the face database constructed in step S1: According to the input parameters, read the face images / face mask images in the face database. If no parameters are specified, read the face mask images;
[0068] (2) Extract face image features: Input the extracted images into the feature extraction model, and the feature extraction model converts each image into a 512-dimensional feature vector; among them, the feature extraction model adopts the student model in the knowledge distillation model;
[0069] (3) Store the face mask feature vectors: Save the converted 512-dimensional feature vectors together with their corresponding face user information into the feature (vector) database. The extracted features with a length of 512 reduce the data scale by nearly 90% compared with the original image dimensions, and the extracted features can effectively reflect face information. Greatly reduce the data volume in the face comparison stage.
[0070] S3: Construct a knowledge distillation model: The knowledge distillation model includes a teacher model and a student model, and the specific implementation steps are as follows:
[0071] (1) Construction of teacher model: Read the pre-trained InceptionResNetv1 model as the teacher model to guide the training of the student model. The InceptionResNetV1 model is trained on the vggface2 dataset and has strong feature extraction capabilities, which can be used for knowledge transfer;
[0072] As Figure 2 shown, the teacher model of the method of the present invention uses the InceptionResNetv1 model. The InceptionResNetv1 network fuses the Inception block and the ResNet block, improving the model's ability to extract features. Among them, the Inception block ensures the model's recognition ability for face images of different sizes; the ResNet block ensures the stable update of the model during training. The InceptionResNetV1 network is pre-trained on the vggface2 dataset. The vggface2 dataset has a total of more than 40,000 pictures, covering face images of different ages, nationalities, and shooting angles. The teacher network trained using the vggface2 dataset has strong feature extraction capabilities.
[0073] (2) Construction of student model: Construct an improved MobileNetv4 model as the student model. Since the student model has a small number of parameters and fast feature extraction speed, transferring the knowledge of the teacher model to the student model can enhance the student model's feature extraction ability;
[0074] As Figure 3 shown, the student model of the method of the present invention uses an improved MobileNetV4 feature network as the student model. It uses a more lightweight image processing module and builds the MobileNetV4 network structure through the FusedIB block, IB block, and ExtraDW block composed of depthwise separable convolution (DWConv), which can reduce the parameters of the model, thereby improving the speed of feature extraction and face detection of the model. On this basis, the present invention improves the MobileNetV4 model and replaces the IB block in the model with a Multi-Scale Inverted Bottleneck block (MSIB). Compared with the IB module, the MSIB module processes the feature map using multi-scale parallel convolutional layers.
[0075] The MSIB structure first uses 1×1 convolution to expand the number of channels of the feature map to twice the original number of channels. Then, the feature map is divided into four parts and sequentially passed through convolutions with kernel sizes of 3×3, 5×5, 3×3, diffusion coefficients of 2, 5×5, and diffusion coefficients of 2. Then, 1×1 convolution is used to restore the number of channels and perform information interaction between different channels. The proposed MSIB structure can better extract face information of different sizes and improve the detection accuracy of the model while ensuring the number of parameters.
[0076] Compared with the IB module that uses per-channel convolution in the middle layer, MSIB evenly divides the feature map into four parts in the channel direction, and each part is processed using convolutions with different kernel sizes and different diffusion coefficients, which can extract face information of different scales. It can improve the detection accuracy and ensure that the parameters will not increase significantly.
[0077] (3) Teacher knowledge transfer: Use the vggface2 dataset and input it into the teacher model and the student model. Use the MSE loss function to align the output of the student model with the output of the teacher model, and in this way, transfer the knowledge of the teacher model to the student model. This loss function can be expressed as:
[0078]
[0079] In the formula: f student represents the output result of the student model, f teacher is the output result of the teacher model, n is the number of samples, and x i is the input.
[0080] As Figure 4 shown in the knowledge distillation process of the method of the present invention, unlabeled data is simultaneously passed to the teacher model and the student model, and the output of the teacher model is used as the label to align the output of the student model with the output of the teacher model, and the parameters of the student model are updated through the MSE loss function. In this way, the knowledge of the teacher model is imparted to the student model, and the student model obtained by this method is used as a feature extraction model, which can take into account both the accuracy and speed of extraction.
[0081] (4) Student model storage: Save the student model that has obtained knowledge, and the saved student model can be used as a feature extraction model.
[0082] The teacher model is characterized by strong feature extraction ability but slow speed, and the student model is characterized by limited feature extraction ability but fast speed. Using the method of knowledge distillation can transfer the knowledge of the teacher model to the student model, effectively improve the feature extraction ability of the student model, and thus improve the detection accuracy and speed of the method.
[0083] Using the knowledge distillation method for model training can avoid annotating images, effectively reduce the workload, and enable model training for different types of datasets.
[0084] S4: Face database comparison: Turn on the real-time detection camera, capture the camera screen at regular intervals, input it into the feature extraction model, measure the similarity between the extracted features and the features in the feature database, and output the face information corresponding to the feature with the highest similarity. The specific implementation steps are as follows;
[0085] (1) Feature database reading: According to the selected parameters, extract face features / face mask features and their corresponding face user information from the feature database. If no parameters are specified, read the face mask features;
[0086] (2) Real-time camera detection: Turn on the device camera and capture the camera screen image at regular intervals;
[0087] (3) Real-time feature extraction: Input the captured camera screen into the feature extraction model to obtain the feature vector corresponding to the real-time face;
[0088] (4) Feature comparison: Measure the similarity between the real-time face features and the read features in sequence. Decide which features to add to the candidate list according to the threshold. If it is greater than the threshold, add it; otherwise, do not add it;
[0089] (5) Face information output: After comparing all the face features with the read features, sort the obtained candidate list and output the face user information corresponding to the feature with the highest similarity.
[0090] When measuring similarity, different methods can be used to calculate similarity, such as Euclidean distance, cosine similarity, Manhattan distance, etc. Different distance calculation methods can be applied to different scenarios.
[0091] S5: Face memory enhancement: Use the candidate list obtained in step S4 as memory features, perform memory enhancement operations on the corresponding features in the feature database, and update and store the enhanced data in the feature database. The specific implementation steps are as follows:
[0092] (1) Memory feature collection: During the face data comparison in step S4, the detected results can be transmitted to the memory feature list in real time through the keyboard. The transmitted memory feature list saves the detected face information and the corresponding real-time face features at that moment;
[0093] (2) Memory Enhancement: Use the features in the memory feature list to update the features in the feature database one by one. The specific implementation steps are to perform weighted averaging on the two features to achieve feature enhancement. The weights can be passed through parameters. If no parameters are passed, the weights are both 0.5. The update formula can be expressed as:
[0094]
[0095] In the formula: represents the updated feature of the i-th user, represents the memory feature of the i-th user, represents the feature of the i-th user before update, and α1 and α2 represent weight factors.
[0096] (3) Update Feature Storage: Store the updated features into the feature database.
[0097] As Figure 5 shown, the memory enhancement process of the method of the present invention is as follows. During the detection process, the recognized correct features and their user information are stored in the memory list through keyboard input. After the detection is completed, all the features in the memory list are matched with the corresponding users in the feature database, and the memory features and the features in the feature database are weighted averaged. Compared with adjusting the parameter space of the model, adjusting the vector space of the feature data has a smaller operation cost and feasibility. The adjusted feature vector is more adaptable to the face detection scenario in more complex situations.
[0098] As Figure 6 shown, the face mask adding process of the method of the present invention is as follows. First, extract the face image from the database, use the dlib library to identify the face data in the image, add the incoming mask image to the face, and return the face data with the mask added.
[0099] As Figure 7 shown, when using the student model for real-time detection, it can be found that the model can identify the registered faces under normal circumstances, frame the recognized faces and return the registration information; when the face head and chin parts are blocked, the model can also accurately identify the face information; specifically, when no face information appears in the camera, "No face detected" is marked in the return.
[0100] As Figure 8 shown, when using the student model for real-time detection and performing face recognition tests under dim conditions, it can be found that the model can accurately identify faces under relatively dim conditions, and can also better identify face information when wearing a mask and blocking the head area.
[0101] The present invention also provides a mask face recognition device based on the knowledge distillation and memory enhancement method, which includes a face mask adding module, a face mask feature extraction module, a knowledge distillation model learning module, a face database comparison module, and a face memory enhancement module. The face mask adding module realizes the addition of a mask to the face based on dlib. The face mask feature extraction module can encode the face image to obtain a feature vector with a smaller storage space. The knowledge distillation model learning module realizes model learning without labels, and can combine the advantages of the teacher model and the student model to obtain the most effective feature extraction model. The face database comparison module performs face recognition through the similarity measurement between the real-time features and the features in the feature database, and realizes real-time feedback of the recognition result through face bounding and information annotation. The face memory enhancement module can save the face feature data when the recognition is correct, update the face data stored in the feature database, and enhance the recognition accuracy of this face.
[0102] The present invention also provides a computer device, including a memory and a processor. A computer program is stored in the memory, and when the processor executes the computer program, the steps of the mask face recognition method based on the knowledge distillation and memory enhancement method are realized.
[0103] The present invention also provides a computer-readable storage medium, including a stored computer program. When the computer program runs, it controls the device where the computer-readable storage medium is located to execute the mask face recognition method based on the knowledge distillation and memory enhancement method of the present invention.
[0104] The present invention also provides a mask face recognition device based on the knowledge distillation and memory enhancement method, including a computer device and a camera running on a portable mobile device. In this embodiment, Raspberry Pi 5 is used as the computer device, and a PCE camera is used as the detection camera for the face recognition method. This method can perform real-time detection on a mobile device.
[0105] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them. Although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements for some or all of the technical features. However, such modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.
Claims
1. A mask face recognition method based on knowledge distillation and memory enhancement method, characterized by: The following steps are involved: S1: Building a face database: adding a mask to an existing face image database, and storing the face image with the mask added into the face image database to obtain a face database; S2: Face mask feature extraction: Each image in the face database is converted into a feature vector and stored through the feature extraction model; S3: Build a knowledge distillation model: The knowledge distillation model includes a teacher model and a student model. The specific implementation steps are as follows: (1) Teacher model construction: Read the pre-trained InceptionResNetv1 model as the teacher model to guide the training of the student model; (2) Student model construction: An improved MobileNetv4 model is constructed as the student model. The improved MobileNetV4 model includes a FusedIB block, an MSIB block, and an ExtraDW block composed of channel-separable convolutions. The structure of the improved MobileNetV4 model is as follows: the input passes through Conv2D, two FusedIB blocks, an ExtraDW block, four MSIB blocks, ConvNext, two ExtraDW blocks, four MSIB blocks, an average pooling layer, and two Conv2Ds in sequence; The MSIB block first uses 1×1 convolution to expand the number of channels of the feature map to twice the original number of channels, then divides the feature map into 4 parts, and sequentially passes through convolutions with kernel sizes of 3×3, 5×5, 3×3, diffusion coefficient of 2, 5×5, diffusion coefficient of 2, and then restores the number of channels through 1×1 convolution, and performs information interaction between different channels; (3) Teacher knowledge transfer: Use the loss function to align the output of the student model with the output of the teacher model. In this way, the knowledge of the teacher model is transferred to the student model. (4) Student model storage: The student model that has obtained the knowledge is saved and used as a feature extraction model; S4: Face database comparison: collect face data and input it into the feature extraction model, measure the similarity between the extracted features and the features in the feature database, and output the face information corresponding to the features with the highest similarity; S5: Face memory enhancement: Perform memory enhancement operation on the corresponding features in the feature database, and update and store the enhanced data in the feature database.
2. According to claim 1, a mask face recognition method based on knowledge distillation and memory enhancement method is characterized in that: The specific implementation steps of step S1 are as follows: (1) Extracting face images from the database: Read each image from the image database, each image represents a face; (2) Facial feature point conversion: For the extracted image, the MTCNN model is first used to extract the face image in the image, and then the dlib library is used to convert the face into 68 feature key points to obtain the absolute coordinates of the feature key points relative to the face image; (3) Mask selection: Based on the input parameters, the specified mask image is selected and covered with different colors. If no parameters are specified, the mask type and color are randomly selected for addition; (4) Mask addition: Obtain the absolute coordinates of the chin contour key points and the nose bridge key points, adjust the size and rotation angle of the mask image, and cover the adjusted mask image to the chin position of the face to obtain the face image with the mask added; (5) Face mask image storage: The obtained face mask images are stored in an image database, and a face database containing normal faces and faces wearing masks is obtained.
3. The mask face recognition method based on knowledge distillation and memory enhancement method according to claim 2 is characterized in that: The specific implementation steps of step S2 are as follows: (1) Extracting the image in the face database constructed in step S1: Reading the face image / face mask image in the face database according to the input parameters. If no parameters are specified, reading the face mask image; (2) Face image feature extraction: The extracted images are input into the feature extraction model, which converts each image into a feature vector; (3) Face mask feature vector storage: The converted feature vector and its corresponding face user information are saved in the feature database.
4. The mask face recognition method based on knowledge distillation and memory enhancement method according to claim 3 is characterized in that: The specific implementation steps of step S4 are as follows: (1) Feature database reading: extracting facial features / face mask features and their corresponding facial user information from the feature database according to the selected parameters. If no parameters are specified, the face mask features are read; (2) Real-time camera detection: Turn on the device camera and capture the camera image at regular intervals; (3) Real-time feature extraction: The captured camera image is input into the feature extraction model to obtain the feature vector corresponding to the real-time face; (4) Feature comparison: The features corresponding to the real-time face are measured in similarity with the read features in turn, and the features to be added to the candidate list are determined based on the threshold. If the similarity is greater than the threshold, the feature is added, otherwise it is not added; (5) Face information output: After the face features are compared with the read features, the candidate list is sorted and the face user information corresponding to the feature with the highest similarity is output.
5. The mask face recognition method based on knowledge distillation and memory enhancement method according to claim 4 is characterized in that: The specific implementation steps of step S5 are as follows: (1) Memory feature collection: In the face data comparison process of step S4, the detected result is transferred to the memory feature list in real time through the keyboard, and the transferred memory feature list saves the detected face information and the corresponding real-time face features at that moment; (2) Memory enhancement: Use the features in the memory feature list to update the features in the feature database in sequence; (3) Update feature storage: store the updated features in the feature database.
6. The mask face recognition method based on knowledge distillation and memory enhancement method according to claim 5 is characterized in that: Using the features in the memory feature list to update the features in the feature database in sequence is to achieve feature enhancement by taking a weighted average of the two features.
7. A mask face recognition device based on knowledge distillation and memory enhancement method, characterized in that: include: Face mask adding module: used to add a mask to each image in the existing face image database to obtain a face database containing face images and face mask images; Face mask feature extraction module: used to encode the face image and obtain the feature vector; Knowledge distillation model learning module: used to use the student model obtained according to the knowledge distillation model as a feature extraction model, where the student model adopts the improved MobileNetV4 model, including FusedIB blocks, MSIB blocks and ExtraDW blocks composed of channel-separable convolutions. The structure of the improved MobileNetV4 model is: the input passes through Conv2D, two FusedIB blocks, ExtraDW blocks, four MSIB blocks, ConvNext, two ExtraDW blocks, four MSIB blocks, average pooling layer, and two Conv2D in sequence; The MSIB block first uses 1×1 convolution to expand the number of channels of the feature map to twice the original number of channels, then divides the feature map into 4 parts, and sequentially passes through convolutions with kernel sizes of 3×3, 5×5, 3×3, diffusion coefficient of 2, 5×5, diffusion coefficient of 2, and then restores the number of channels through 1×1 convolution, and performs information interaction between different channels; Face database comparison module: used to perform face recognition by measuring the similarity between real-time features and features in the feature database, and to provide real-time feedback of recognition results through face selection and information annotation; Face memory enhancement module: used to save face feature data when the recognition is correct and update the face data stored in the feature database.
8. A computer device, characterized in that: It includes a memory and a processor, wherein a computer program is stored in the memory, and when the processor executes the computer program, the steps of the mask face recognition method based on the knowledge distillation and memory enhancement method as described in any one of claims 1 to 6 are implemented.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by the processor, the steps of the mask face recognition method based on the knowledge distillation and memory enhancement method as described in any one of claims 1 to 6 are implemented.
10. A mask face recognition device based on knowledge distillation and memory enhancement method, characterized in that: The invention comprises a computer device as claimed in claim 8 and a camera running on a portable mobile device.
Citation Information
Patent Citations
Knowledge distillation network-based mask face shielding recognition method and device, and equipment
CN113343898A
Attention mechanism-fused mask shielded face detection and recognition method
CN115497139A
Face recognition method and device based on rapid mask generation
CN116665279A
Method for recognizing face with mask based on deep learning
CN117558044A
Recognition method in state of wearing mask on human face
WO2023103372A1