Multi-face mask recognition system and method based on MTCNN algorithm and SVM

By optimizing the MTCNN algorithm and improving the FaceNet algorithm, combined with MobileNetV2 and SVM, the problems of low recognition rate and high computational complexity of traditional face recognition technology when wearing masks are solved, and high-precision and high-efficiency multi-face mask recognition are achieved.

CN120388405APending Publication Date: 2025-07-29GUANGZHOU UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510391232.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-28
Publication Date
2025-07-29

AI Technical Summary

Technical Problem

When faced with mask wearing, traditional facial recognition technology has a low recognition rate and high computational complexity, which cannot meet the needs of practical application scenarios.

Method used

By optimizing the MTCNN algorithm, the MobileNetV2 neural network was introduced for mask detection, and the FaceNet algorithm was improved, combined with the support vector machine SVM for feature extraction and classification, the non-maximum suppression algorithm was optimized, and loss functions such as ArcFace, CosFace, SphereFace and Center Loss were used to improve recognition accuracy and robustness.

Benefits of technology

It improves the accuracy and efficiency of multi-face mask recognition, and is suitable for a variety of scenarios such as public places and hospitals, and can effectively respond to the identity identification needs when wearing masks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120388405A_ABST
    Figure CN120388405A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of computer vision and artificial intelligence, and particularly discloses a multi-face mask recognition system and method based on an MTCNN algorithm and an SVM, and the method comprises the steps: firstly obtaining an image through a camera, and carrying out the zooming, graying and normalization preprocessing; then, face detection and alignment are carried out by using the optimized MTCNN algorithm; then, a MobileNetV2 network is adopted to carry out mask detection on the face image; then, based on an improved FaceNet algorithm, respectively extracting 128-dimensional feature vectors of the face wearing the mask and the face not wearing the mask; and finally, carrying out classification through the optimized SVM classifier and outputting a result. According to the invention, by optimizing the MTCNN, introducing the MobileNetV2, improving the FaceNet and optimizing the SVM, high-precision and high-efficiency multi-face mask recognition is realized, and the method is suitable for various scenes.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the fields of computer vision and artificial intelligence technology, and more specifically, to a multi-face mask recognition system and method based on the MTCNN algorithm and SVM. Background Art

[0002] In various environments including public places, hospitals, factories, etc., ensuring that people wear masks is one of the important measures to protect public health. However, traditional face recognition technology often has a low recognition rate when dealing with the situation of wearing masks, and cannot effectively meet the needs of actual application scenarios. In the existing research results, some scholars proposed to use the MTCNN algorithm to achieve face detection and alignment operations, and combine the support vector machine (SVM) to carry out mask recognition research. However, this technical solution still exposes problems such as insufficient recognition accuracy and high computational complexity in actual applications.

[0003] Therefore, a multi-face mask recognition system and method based on the MTCNN algorithm and SVM are provided. Summary of the Invention

[0004] To solve the above technical problems, this application is proposed.

[0005] Specifically, according to one aspect of this application, a multi-face mask recognition system based on the MTCNN algorithm and SVM is provided, which includes:

[0006] An image acquisition and preprocessing module, configured to acquire an input image through a camera, and perform scaling, grayscaling, and normalization processing on the input image;

[0007] A face detection and alignment module, configured to optimize the MTCNN algorithm, and perform face detection and alignment on the preprocessed input image based on the optimized MTCNN algorithm to obtain multiple aligned face images; wherein, the MTCNN algorithm includes a proposal network P-Net, a refinement network R-Net, and an output network O-Net;

[0008] A mask detection module, configured to use the MobileNetV2 neural network to perform mask detection on the aligned face images, and output a binary classification result; wherein, the binary classification result includes: faces wearing masks and faces not wearing masks;

[0009] A face feature extraction module, configured to improve the FaceNet algorithm, and respectively extract 128-dimensional feature vectors of faces wearing masks and faces not wearing masks based on the improved FaceNet algorithm;

[0010] A classification decision module, which is used to classify the extracted 128-dimensional feature vector by using a support vector machine (SVM) classifier and output the final recognition result.

[0011] Preferably, in the face detection and alignment module, the MTCNN algorithm is optimized, including: continuously scaling the preprocessed input image to generate multiple sub-images with different resolutions to form an image pyramid; using the proposal network (P-Net) to generate preliminary face candidate boxes at each scale of the image pyramid; reducing the number of convolutional kernels of the refinement network (R-Net) and the output network (O-Net), and using depthwise separable convolution to replace standard convolution in the refinement network (R-Net) and the output network (O-Net); the refinement network (R-Net) takes the output of the proposal network (P-Net) as input, extracts features and classifies them through a convolutional neural network to optimize the face candidate boxes; the output network (O-Net) locates the face candidate boxes and facial feature points through a convolutional neural network, combines with the Bounding Box Regression model, and adjusts the position and size of the candidate window to obtain the face confidence, the bounding box regression value and the feature point coordinates; by optimizing the non-maximum suppression algorithm, the candidate windows output by the output network (O-Net) are screened to filter out overlapping candidate boxes; corresponding loss functions are defined for the face confidence, the bounding box regression value and the feature point coordinates through cross-entropy and Euclidean distance, which are in turn: face classification loss bounding box regression value loss and face feature point coordinate loss The face classification loss bounding box regression value loss and face feature point coordinate loss are weighted and summed to form the total loss function L of the MTCNN algorithm MTCNN .

[0012] Among them, by optimizing the non-maximum suppression algorithm, the candidate windows output by the output network (O-Net) are screened to filter out overlapping candidate boxes, including: adopting the soft suppression strategy (Soft-NMS) to replace the traditional non-maximum suppression strategy (NMS), and the specific formula is expressed as:

[0013]

[0014] Among them, S i is the confidence of candidate box i, IoU(b i , b j ) is the overlap degree with the surrounding box, and σ is a hyperparameter for controlling the suppression intensity.

[0015] Among them, corresponding loss functions are defined for the face confidence, the bounding box regression value and the feature point coordinates through cross-entropy and Euclidean distance, which are in turn: face classification loss Bounding box regression loss and facial feature point coordinate loss include:

[0016] Face classification loss Use binary cross entropy as the loss function:

[0017]

[0018] Bounding box regression loss Use Euclidean distance as the loss function:

[0019]

[0020]

[0021] Specifically, according to one aspect of the present application, a multi-face mask recognition method based on the MTCNN algorithm and SVM is provided, which includes:

[0022] S1. Acquire an input image through a camera, and perform scaling, grayscale, and normalization processing on the input image;

[0023] S2. Optimizing the MTCNN algorithm, and performing face detection and alignment on the preprocessed input image based on the optimized MTCNN algorithm to obtain multiple aligned face images; wherein the MTCNN algorithm includes a proposal network P-Net, a refinement network R-Net, and an output network O-Net;

[0024] S3. Use the MobileNetV2 neural network to perform mask detection on the aligned face image and output a binary classification result; wherein the binary classification result includes: a face wearing a mask and a face not wearing a mask;

[0025] S4. Improve the FaceNet algorithm and extract 128-dimensional feature vectors of faces wearing masks and faces not wearing masks based on the improved FaceNet algorithm;

[0026] S5. Use the support vector machine (SVM) classifier to classify the extracted 128-dimensional feature vector and output the final recognition result.

[0027] Among them, in the step S4, the FaceNet algorithm is improved, including: using two independent FaceNet networks to extract features of masked faces and unmasked faces respectively; by combining the angular margin constraint of ArcFace, the cosine margin enhancement of CosFace, the hypersphere normalization and angular margin of SphereFace, and the intra-class compactness constraint of Center Loss, to replace the original Triplet Loss.

[0028] Compared with the prior art, a multi-face mask recognition system and method based on the MTCNN algorithm and SVM provided by the present application have the following technical effects:

[0029] 1), The present application optimizes the MTCNN algorithm (such as reducing the number of convolutional kernels in the R-Net and O-Net, using the Bounding Box Regression method to calculate the bounding box regression vector, etc.) and introduces the MobileNetV2 neural network to detect masks for the aligned face images, effectively improving the efficiency of face detection and reducing the computational complexity. At the same time, by optimizing the NMS algorithm, the overlapping face windows are removed, improving the detection accuracy.

[0030] 2), The present application improves the FaceNet algorithm, uses two independent FaceNet networks, and introduces loss functions such as ArcFace, CosFace, SphereFace, and Center Loss. At the same time, by grid search or cross-validation to adjust the regularization parameters and kernel function parameters of the SVM, feature normalization and parameter tuning of the classifier are performed, enhancing the discrimination ability and robustness of the model. Especially in the case of wearing masks, the recognition accuracy has been significantly improved.

[0031] 3), The present application is applicable to various scenarios, such as public places, hospitals, factories, etc., and can effectively meet the identity recognition requirements in the case of wearing masks. BRIEF DESCRIPTION OF THE DRAWINGS

[0032] By describing the embodiments of the present application in more detail in conjunction with the accompanying drawings, the above and other objects, features, and advantages of the present application will become more obvious. The accompanying drawings are used to provide a further understanding of the embodiments of the present application, and constitute a part of the specification. Together with the embodiments of the present application, they are used to explain the present application and do not constitute a limitation to the present application. In the accompanying drawings, the same reference numerals generally represent the same components or steps.

[0033] Figure 1 The system block diagram of the embodiment of the present application is illustrated.

[0034] Figure 2Illustrates the flowchart of the traditional FaceNet model architecture according to an embodiment of the present application.

[0035] Figure 3 Illustrates the flowchart of the method according to an embodiment of the present application. Detailed implementation manners

[0036] Next, embodiments according to the present application will be described in detail with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all embodiments of the present application. It should be understood that the present application is not limited by the example embodiments described herein.

[0037] Embodiment:

[0038] Figure 1 Illustrates the system block diagram according to an embodiment of the present application. As Figure 1 shown, the multi-face mask recognition system based on the MTCNN algorithm and SVM according to an embodiment of the present application includes: an image acquisition and preprocessing module, configured to acquire an input image through a camera and perform scaling, grayscale conversion, and normalization processing on the input image; a face detection and alignment module, configured to optimize the MTCNN algorithm and perform face detection and alignment on the preprocessed input image based on the optimized MTCNN algorithm to obtain a plurality of aligned face images; wherein, the MTCNN algorithm includes a proposal network P-Net, a refinement network R-Net, and an output network O-Net; a mask detection module, configured to perform mask detection on the aligned face images using a MobileNetV2 neural network and output a binary classification result; wherein, the binary classification result includes: a face wearing a mask and a face not wearing a mask; a face feature extraction module, configured to improve the FaceNet algorithm and respectively extract 128-dimensional feature vectors of the face wearing a mask and the face not wearing a mask based on the improved FaceNet algorithm; a classification decision module, configured to classify the extracted 128-dimensional feature vectors using a support vector machine SVM classifier and output a final recognition result.

[0039] In the embodiment of the present application, the image acquisition and preprocessing module is configured to acquire an input image through a camera and perform scaling, grayscale conversion, and normalization processing on the input image. Considering problems such as inconsistent resolution of the input images, large amount of color image data and limited help of color information for mask recognition, and significant differences in the distribution range of image pixel values under different lighting conditions, directly processing these images will significantly increase the computational complexity and reduce the recognition accuracy. Therefore, in order to improve the accuracy of image recognition, it is necessary to perform a series of preprocessing measures such as scaling, grayscale conversion, and normalization on the input images, so as to ensure that the images can meet the input requirements of subsequent face detection and mask recognition algorithms.

[0040] In the embodiment of the present application, the face detection and alignment module is used to optimize the MTCNN algorithm, and perform face detection and alignment on the preprocessed input image based on the optimized MTCNN algorithm to obtain multiple aligned face images. Among them, the MTCNN algorithm includes a proposal network P-Net, a refinement network R-Net, and an output network O-Net. It is not difficult to understand that the original MTCNN algorithm will have a reduced detection accuracy when facing a face wearing a mask, and its computational complexity is relatively high. In addition, accurately locating the face area is crucial, and the poses and angles of faces in the image are different. If the unoptimized MTCNN algorithm is directly used for face detection and alignment operations, it is very likely to result in insufficient recognition accuracy and poor real-time performance. Therefore, in the embodiment of the present application, in order to improve the overall recognition accuracy and real-time performance of the system, the MTCNN algorithm is optimized before face detection and alignment.

[0041] Among them, the specific steps for optimizing the MTCNN algorithm are as follows:

[0042] 1) Continuously scale the preprocessed input image to generate multiple sub-images with different resolutions, forming an image pyramid. That is, an image pyramid is generated through multi-scale scaling to cover images with different resolutions, which can effectively solve the problem of face size variation and ensure that small-sized faces can also be effectively detected.

[0043] 2) Use the proposal network P-Net to generate preliminary face candidate boxes at each scale of the image pyramid.

[0044] 3) Reduce the number of convolutional kernels of the refinement network R-Net and the output network O-Net (such as reducing from 32 channels to 16 channels) to reduce the computational complexity, and use depthwise separable convolutions to replace some standard convolutions in the refinement network R-Net and the output network O-Net. Among them, the depthwise separable convolution decomposes the standard convolution into depthwise convolution and pointwise convolution, which can significantly reduce the number of parameters and the amount of computation while maintaining the feature extraction ability.

[0045] 4) The refinement network R-Net takes the output of the proposal network P-Net as input, performs feature extraction and classification through a convolutional neural network to optimize the face candidate boxes. Among them, R-Net can filter out most non-face windows, retain the candidate windows that may contain faces, and learn the mapping relationship from the candidate windows to the real windows through the Bounding Box Regression model, and output the border regression vectors corresponding to each candidate window to further correct the position and size of the candidate windows.

[0046] 5) The output network O-Net locates the face candidate frame and facial feature points (such as eyes, nose, and mouth) through a convolutional neural network, and combines the Bounding Box Regression model to adjust the position and size of the candidate window to obtain the face confidence, bounding box regression value, and feature point coordinates.

[0047] 6) By optimizing the non-maximum suppression (NMS) algorithm, the candidate windows output by the output network O-Net are screened to filter out overlapping candidate frames instead of directly eliminating high-scoring frames. The specific implementation is: using the soft suppression strategy Soft-NMS to replace the traditional non-maximum suppression strategy NMS, the formula is expressed as:

[0048]

[0049] Among them, S i is the confidence of candidate box i, IoU(b i ,b j ) is the overlap with the surrounding box, and σ is a hyperparameter that controls the suppression strength.

[0050] 7) Define the corresponding face classification loss for face confidence, bounding box regression value and feature point coordinates through cross entropy and Euclidean distance Bounding box regression loss and facial feature point coordinate loss Specifically:

[0051] Face classification loss Use binary cross entropy as the loss function:

[0052]

[0053] Bounding box regression loss Use Euclidean distance as the loss function:

[0054]

[0055]

[0056] Among them, the bounding box regression loss function is combined with the Bounding Box Regression model to optimize the parameters of the regression model by minimizing the offset between the predicted bounding box and the true bounding box (such as the difference in center point coordinates, width, and height), thereby realizing the mapping learning from candidate windows to true windows.

[0057] 8) All parameters are optimized iteratively through the total loss function LMTCNN. Specifically, the face classification loss Bounding box regression loss and facial feature point coordinate loss Perform weighted summation to form the total loss function \(L\) of the MTCNN algorithm MTCNN , and the specific formula is expressed as:

[0058] \(L\) MTCNN =\(\alpha\) det \(L\) det +\(\alpha\) box \(L\) box +\(\alpha\) landmarks \(L\) landmarks

[0059] where \(\alpha\) det , \(\alpha\) box and \(\alpha\) landmarks are all weight coefficients.

[0060] In the embodiment of the present application, the mask detection module is used to perform mask detection on the aligned face image by using the MobileNetV2 neural network and output a binary classification result; wherein, the binary classification result includes: a face wearing a mask and a face not wearing a mask. It is not difficult to understand that MobileNetV2, as a lightweight and efficient deep learning model, can effectively solve these problems. By processing the aligned face image, MobileNetV2 uses its optimized network structure and efficient computing power to accurately identify whether a face wears a mask and outputs a binary classification result. This optimized detection method not only significantly improves the overall recognition accuracy of the system, but also greatly reduces the computational complexity, thus meeting the strict requirements for real-time performance and computational efficiency in practical applications.

[0061] In the embodiment of the present application, the face feature extraction module is used to improve the FaceNet algorithm and extract 128-dimensional feature vectors of the face wearing a mask and the face not wearing a mask based on the improved FaceNet algorithm. It should be understood that the FaceNet algorithm, as a deep learning-based face feature extraction method, can map a face into a high-dimensional space, so that the feature vectors of the same face are close in this space, while the feature vectors of different faces are far apart. This characteristic makes the FaceNet algorithm highly discriminative and robust in processing face features and can effectively handle the situation of local face occlusion (such as mask occlusion). By extracting 128-dimensional feature vectors, the system can more accurately identify and distinguish the faces wearing masks and not wearing masks, thereby improving the overall recognition accuracy and reliability of the system. Therefore, by using the FaceNet algorithm to extract feature vectors, the limitations of traditional technologies can be overcome, and the advantages of deep learning can be utilized to achieve more efficient and accurate face feature extraction, ultimately improving the performance of the multi-face mask recognition system.

[0062] Particularly, Figure 2The figure shows the traditional FaceNet model architecture flow chart of the embodiment of the present application. Figure 2 As shown in Figure 1, FaceNet uses triplet loss, which improves the model's ability to distinguish by minimizing the distance difference between positive and negative samples. However, triplet loss has some limitations, such as difficulty in sample selection and slow training speed. Therefore, in the embodiment of the present application, before using the FaceNet algorithm for feature extraction, the FaceNet algorithm is improved. The specific improvement steps are as follows:

[0063] 1) Use two independent FaceNet networks to extract features of faces wearing masks and faces not wearing masks respectively.

[0064] 2) By using the angular interval constraint of ArcFace, the cosine boundary enhancement of CosFace, and the hyperspherical normalization and angular interval of SphereFace, combined with the intra-class compactness constraint of Center Loss, the original triplet loss Triplet Loss is replaced.

[0065] In particular, ArcFace's angular margin constraint includes: introducing a fixed interval (angular margin) in the angular dimension of the feature space, forcing features of different categories to be distributed in different angular regions of the hypersphere. By converting the cosine similarity in the original Softmax loss into an angular form, the specific formula is expressed as:

[0066]

[0067] Among them, m is the angular interval, which significantly increases the separation between classes and is especially suitable for fine-grained classification in high-dimensional feature space.

[0068] CosFace's cosine boundary enhancement includes: adding the classification boundary cosinemargin in the cosine similarity space, and adjusting the weight and bias terms. The specific formula is expressed as follows:

[0069]

[0070] Among them, this method simplifies the calculation of angle intervals, but can still effectively expand the distance between classes while maintaining computational efficiency.

[0071] SphereFace's hyperspherical normalization and angular spacing include: hyperspherical normalization of the weights of the last layer of the network (i.e., constraining the weight vector to be a point on the unit sphere) and introducing multiple spacing constraints in the angular space. (The core of this is to amplify angular differences through an extremely deep network (such as a 20-layer convolutional layer)). The specific formula is expressed as:

[0072]

[0073] Among them, m is an integer multiple, which further strengthens the angular separation between categories.

[0074] The within-class compactness constraint of Center Loss includes: by minimizing the distance from the features of samples of the same class to their class centers, and the specific formula is expressed as:

[0075]

[0076] Among them, is the center vector of class y i . This loss is used in combination with Softmax / CosFace, significantly improving the compactness of within-class features and alleviating the feature noise in occlusion scenarios.

[0077] That is to say, the improved FaceNet algorithm directly defines geometric constraints in the loss function by adopting loss functions such as ArcFace and CosFace, eliminating the dependence on sample selection. At the same time, it integrates angular margin, cosine margin, hypersphere normalization and within-class compactness constraints to construct a multi-dimensional feature optimization framework. This not only avoids the sampling bias problem of triplet loss and accelerates convergence, but also makes the model pay more attention to the feature expression of local unoccluded regions in the mask occlusion scenario, suppresses the feature noise caused by mask occlusion, enhances the consistency of samples of the same class, and improves the generalization performance and classification robustness of the model in mask / non-mask mixed data. Thus, it effectively solves the feature confusion problem in the mask occlusion scenario and significantly improves the training efficiency and classification accuracy.

[0078] In the embodiment of the present application, the classification decision module is used to classify the extracted 128-dimensional feature vector by using a support vector machine (SVM) classifier and output the final recognition result. Specifically, the classification decision module includes: normalizing the extracted 128-dimensional feature vector, and the specific formula is expressed as:

[0079]

[0080] Among them, x is the original feature value, min and max are respectively the minimum and maximum values of all elements in the feature vector, and x * is the normalized feature value; adjusting the regularization parameter and kernel function parameter of the support vector machine (SVM) through grid search or cross-validation; inputting the normalized feature vector into the adjusted support vector machine (SVM) classifier and performing classification using the One-vs-One strategy; obtaining the final recognition result according to the output of the support vector machine (SVM) classifier.

[0081] In summary, the multi-face mask recognition system based on the MTCNN algorithm and SVM according to the embodiments of the present application is elucidated. First, it acquires an image through a camera and performs preprocessing such as scaling, grayscaling, and normalization. Then, it uses the optimized MTCNN algorithm for face detection and alignment. Next, it employs the MobileNetV2 network to detect masks in the face images. After that, it extracts 128-dimensional feature vectors of masked and unmasked faces respectively based on the improved FaceNet algorithm. Finally, it classifies through the optimized SVM classifier and outputs the result. The present application realizes high-precision and high-efficiency multi-face mask recognition by optimizing MTCNN, introducing MobileNetV2, improving FaceNet, and optimizing SVM, and is applicable to various scenarios.

[0082] Figure 3 The flowchart of the method according to the embodiments of the present application is illustrated. As Figure 3 shown, the specific implementation process of the multi-face mask recognition method based on the MTCNN algorithm and SVM according to the embodiments of the present application includes: S1. Acquire an input image through a camera and perform scaling, grayscaling, and normalization processing on the input image; S2. Optimize the MTCNN algorithm, and based on the optimized MTCNN algorithm, perform face detection and alignment on the preprocessed input image to obtain multiple aligned face images; wherein, the MTCNN algorithm includes a proposal network P-Net, a refinement network R-Net, and an output network O-Net; S3. Use the MobileNetV2 neural network to detect masks in the aligned face images and output a binary classification result; wherein, the binary classification result includes: masked faces and unmasked faces; S4. Improve the FaceNet algorithm, and based on the improved FaceNet algorithm, extract 128-dimensional feature vectors of masked faces and unmasked faces respectively; S5. Use a support vector machine SVM classifier to classify the extracted 128-dimensional feature vectors and output the final recognition result.

[0083] Here, those skilled in the art can understand that the specific functions and operations of each step in the above multi-face mask recognition method based on the MTCNN algorithm and SVM have been introduced in detail in the description of the multi-face mask recognition method system based on the MTCNN algorithm and SVM above, and therefore, its repeated description will be omitted. Figure 1 For those skilled in the art, it is obvious that the present invention is not limited to the details of the above exemplary embodiments, and without departing from the spirit or basic characteristics of the present invention, the present invention can be implemented in other specific forms.

[0084]

[0085] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the preferred embodiments, those of ordinary skill in the art should understand that the technical solutions of the present invention can be modified or equivalently replaced without departing from the spirit of the technical solutions of the present invention.

Claims

1. A multi-face mask recognition system based on the MTCNN algorithm and SVM, characterized in that, Including: An image acquisition and preprocessing module, which is used to acquire an input image through a camera and perform scaling, grayscale conversion, and normalization processing on the input image; A face detection and alignment module, which is used to optimize the MTCNN algorithm and perform face detection and alignment on the preprocessed input image based on the optimized MTCNN algorithm to obtain multiple aligned face images; wherein, the MTCNN algorithm includes a proposal network P-Net, a refinement network R-Net, and an output network O-Net; A mask detection module, which is used to perform mask detection on the aligned face images by using a MobileNetV2 neural network and output a binary classification result; wherein, the binary classification result includes: a face wearing a mask and a face not wearing a mask; A face feature extraction module, which is used to improve the FaceNet algorithm and extract 128-dimensional feature vectors of the face wearing a mask and the face not wearing a mask respectively based on the improved FaceNet algorithm; A classification decision module, which is used to classify the extracted 128-dimensional feature vectors by using a support vector machine SVM classifier and output a final recognition result; Wherein, in the face detection and alignment module, optimizing the MTCNN algorithm includes: Continuously scaling the preprocessed input image to generate multiple sub-images with different resolutions, forming an image pyramid; Using the proposal network P-Net to generate preliminary face candidate boxes at each scale of the image pyramid; Reducing the number of convolutional kernels of the refinement network R-Net and the output network O-Net, and using depthwise separable convolution to replace standard convolution in the refinement network R-Net and the output network O-Net; The refinement network R-Net takes the output of the proposal network P-Net as input, performs feature extraction and classification through a convolutional neural network to optimize the face candidate boxes; The output network O-Net locates the face candidate boxes and facial feature points through a convolutional neural network, combines the Bounding BoxRegression model, and adjusts the position and size of the candidate window to obtain the face confidence, the bounding box regression value, and the feature point coordinates; By optimizing the non-maximum suppression algorithm, screening the candidate windows output by the output network O-Net to filter out overlapping candidate boxes; Define corresponding loss functions for face confidence, bounding box regression values, and facial landmark coordinates through cross-entropy and Euclidean distance, which are in turn: face classification loss Bounding box regression value loss And facial landmark coordinate loss Face classification loss Bounding box regression value loss And face feature point coordinate loss Are weighted and summed to form the total loss function L of the MTCNN algorithm MTCNN .

2. The multi-face mask recognition system based on the MTCNN algorithm and SVM according to claim 1, wherein By optimizing the non-maximum suppression algorithm, screening the candidate windows output by the output network O-Net to filter out overlapping candidate boxes, including: adopting the soft suppression strategy Soft-NMS to replace the traditional non-maximum suppression strategy NMS, and the specific formula is expressed as: Among them, S i is the confidence of candidate box i, and IoU(b i , b j ) is the overlap degree with surrounding boxes, and σ is a hyperparameter that controls the suppression intensity.

3. The multi-face mask recognition system based on the MTCNN algorithm and SVM according to claim 2, wherein, Define corresponding loss functions for face confidence, bounding box regression values, and facial landmark coordinates through cross-entropy and Euclidean distance, which are, in turn: face classification loss bounding box regression value loss and facial landmark coordinate loss including: Face classification loss Use binary cross-entropy as the loss function: Bounding box regression value loss Use the Euclidean distance as the loss function: Facial feature point coordinate loss Use the Euclidean distance as the loss function:

4. The multi-face mask recognition system based on the MTCNN algorithm and SVM according to claim 3, characterized in that, Face classification loss Bounding box regression value loss And face feature point coordinate loss Are weighted and summed to form the total loss function L of the MTCNN algorithm MTCNN , and the specific formula is expressed as: L MTCNN = α det L det + α box L box + α landmarks L landmarks Among them, α det , α box and α landmarks are all weight coefficients.

5. The multi-face mask recognition system based on the MTCNN algorithm and SVM according to claim 1, characterized in that, In the face feature extraction module, improving the FaceNet algorithm includes: Using two independent FaceNet networks to perform feature extraction on the face wearing a mask and the face not wearing a mask respectively; By using the angular margin constraint of ArcFace, the cosine margin enhancement of CosFace, and the hypersphere normalization and angular margin of SphereFace, combined with the intra-class compactness constraint of Center Loss, to replace the original triplet loss Triplet Loss.

6. The multi-face mask recognition system based on the MTCNN algorithm and SVM according to claim 1, wherein The classification decision module includes: Normalize the extracted 128-dimensional feature vector; Adjust the regularization parameter and kernel function parameter of the support vector machine (SVM) through grid search or cross-validation; Input the normalized feature vector into the adjusted SVM classifier and perform classification using the One-vs-One strategy; Obtain the final recognition result according to the output of the SVM classifier.

7. The multi-face mask recognition system based on the MTCNN algorithm and SVM according to claim 6, characterized in that, Normalize the extracted 128-dimensional feature vector, and the specific formula is expressed as: where x is the original eigenvalue, min and max are respectively the minimum and maximum values of all elements in the eigenvector, and x * is the normalized eigenvalue.

8. A multi-face mask recognition method based on the MTCNN algorithm and SVM, characterized in that, Including: S1. Obtain the input image through the camera, and perform scaling, grayscale conversion, and normalization on the input image; S2. Optimize the MTCNN algorithm, and perform face detection and alignment on the preprocessed input image based on the optimized MTCNN algorithm to obtain multiple aligned face images; wherein, the MTCNN algorithm includes a proposal network (P-Net), a refinement network (R-Net), and an output network (O-Net); S3. Use the MobileNetV2 neural network to perform mask detection on the aligned face images and output a binary classification result; wherein, the binary classification result includes: faces wearing masks and faces not wearing masks; S4. Improve the FaceNet algorithm, and respectively extract the 128-dimensional feature vectors of faces wearing masks and faces not wearing masks based on the improved FaceNet algorithm; S5. Use an SVM classifier to classify the extracted 128-dimensional feature vectors and output the final recognition result; Wherein, in the S4, improving the FaceNet algorithm includes: Using two independent FaceNet networks to respectively extract features of faces wearing masks and faces not wearing masks; Replacing the original triplet loss (Triplet Loss) with the angular margin constraint of ArcFace, the cosine margin enhancement of CosFace, and the hypersphere normalization and angular margin of SphereFace, combined with the intra-class compactness constraint of Center Loss.