A face mask recognition method and system based on improved YOLOv5
By improving the network structure and strategy of the YOLOv5 algorithm model, the shallow feature extraction and occlusion error judgment problems of face mask detection in natural scenes are solved, real-time automatic detection in public areas is realized, and the accuracy and speed of detection are improved.
Patent Information
- Application Number
- CN202211664167.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-23
- Publication Date
- 2025-05-06
- Estimated Expiration
- 2042-12-23
AI Technical Summary
In the natural scenes, the detection of face masks in the prior art has problems such as insufficient extraction of shallow features, difficulty in detecting small targets in dense crowds, and misjudgment of occlusion, making it difficult to achieve real-time automatic detection.
The improved YOLOv5 face mask recognition method is adopted to improve the performance of the model in a natural environment through network structure improvements (such as network pruning and CBAM convolutional block attention module) and network strategy improvements (such as hidden layer pruning, convolution kernel pruning and CIoU target loss function).
The average accuracy and running speed of the YOLOv5 algorithm model are improved, the effect of intensive small object detection and the perception of the model are enhanced, and real-time automatic detection of face mask wearing in public areas is achieved.
Smart Images

Figure CN116311417B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of computer vision technology and artificial intelligence technology, and in particular to a method and system for recognizing a face mask. Background Art
[0002] At present, in addition to manual detection, face recognition of masks is most widely used in public places. Computer vision technology uses various imaging systems instead of visual organs as input means, and uses computers instead of the brain to complete the processing and interpretation of visual information. With the continuous development of computer vision technology, computers can recognize various faces and give feedback. Although the detection speed and accuracy of small targets in natural scenes have been greatly improved at this stage, there is insufficient extraction of shallow features of the target, and there are problems with small target detection and occlusion misjudgment in dense crowds in natural scenes. Therefore, it is urgent to optimize and improve the existing technology, improve the performance of the model in a natural environment, and provide a method and system that can detect the wearing of face masks in real time and automatically in public areas. Summary of the invention
[0003] The technical problem to be solved by the present invention is to provide a face mask recognition method and system based on an improved YOLOv5, which can detect the wearing status of face masks in real time and automatically in public areas.
[0004] The present invention adopts the following technical solutions to solve the above technical problems:
[0005] The face mask recognition method based on improved YOLOv5 proposed in the present invention includes three aspects: network structure improvement, network strategy improvement and building a face mask recognition system. The specific steps are as follows:
[0006] S1. Network structure improvement: Use network pruning method to reduce the network parameters and structural complexity of YOLOv5 algorithm model, and add CBAM convolution block attention module after CSP module;
[0007] The network pruning method is mainly used to improve the generalization performance of the network and avoid overfitting by reducing network parameters and structural complexity, thereby obtaining a lighter and more efficient YOLOv5 algorithm model.
[0008] S2. Network strategy improvement: Use hidden layer pruning and convolution kernel pruning to adjust the depth and width of the YOLOv5 algorithm model, cut redundant structures, and select CIoU target loss function to improve the YOLOv5 model bounding box loss function BBox Loss The accuracy is improved by increasing the overlap between the predicted box and the true box.
[0009] S3. Build a face mask recognition system: The YOLOv5 algorithm model improved by steps S1 and S2 is trained with the data set to obtain an optimized algorithm model, which is embedded in the face recognition system interaction interface to build a face mask recognition system that can be applied in real time.
[0010] Furthermore, the CBAM convolutional block attention module described in step S1 is a lightweight module, and its specific functions include: when a feature map is given, CBAM will derive the attention map in sequence along the two independent dimensions of channel and space, use CBAM to extract the attention area, and then convolve the attention map with the input feature map to perform adaptive feature refinement, so as to achieve more accurate positioning of the mask position information in the face area.
[0011] Furthermore, in the network strategy improvement described in step S2, the hidden layer pruning method adjusts the network depth by controlling the number of residual components in the CSP structure, and the convolution kernel pruning method adjusts the network width by controlling the number of convolution kernels in the Focus and CBL structures.
[0012] Furthermore, in the improved YOLOv5 algorithm model after steps S1 and S2, the Backbone network is divided into three layers: the first layer is the Focus module, which slices the image and completely extracts features; the second layer is the CBL module, which is mainly composed of the convolution layer, batch normalization layer (BN) and ReLU activation function, where the number of convolution kernels in the convolution layer determines the size of the CBL module output image; the third layer is the CSP1_X module. The CSP2_X modules on the Neck network are all CSP2_1.
[0013] Furthermore, the classification loss part in the CIoU objective loss function described in step S2 is used to measure the performance of the model in face mask recognition classification, and the bounding box regression loss part is used when the bounding boxes overlap; the CIoU objective loss function is selected for the YOLOv5 model bounding box loss function BBox Loss Improvements are made to improve regression accuracy and reduce prediction error. The improved bounding box loss function is as follows:
[0014]
[0015] Among them, S 2 It represents the number of grids on the three scale feature maps; n refers to the number of candidate boxes generated by each grid; I ij obj Represents whether the jth candidate box in the i-th grid can be responsible for the currently predicted target object. If the overlap of the candidate boxes is greater than the selected threshold, I ij obj =1, otherwise I ij obj=0 and the threshold is selected from a value between 0 and 1 according to the actual application conditions, and the value is retained to one decimal place at most; is the width of the marker border; Is the height of the marker border.
[0016] Further, the steps for building the face mask recognition system described in step S3 are as follows:
[0017] S301, collecting public and web crawler-based face mask datasets;
[0018] S302, use Python programming combined with a third-party library to convert the format of the data set;
[0019] S303, using the labeling tool Labelbox to label all the above samples by category;
[0020] S304, expanding the labeled sample data set, where the expansion methods include: horizontal flipping, small angle rotation, scaling, cropping, translation, and Fancy PCA;
[0021] S305, splitting the data set into two categories: the data set of faces wearing masks and the data set of faces not wearing masks, with 80% used as a training set and 20% used as a test set, so that the trained model is more realistic and reliable;
[0022] S306: Save the training model and the test model, and encapsulate them into an exe executable file using Python.
[0023] Furthermore, the present invention also proposes a face mask recognition system based on an improved YOLOv5, comprising:
[0024] The network structure improvement module is used to reduce the network parameters and structural complexity of the YOLOv5 algorithm model using the network pruning method, and add the CBAM convolution block attention module after the CSP module to improve the generalization performance of the network;
[0025] The network strategy improvement module is used to adjust the depth and width of the YOLOv5 algorithm model by using hidden layer pruning and convolution kernel pruning, respectively, to reduce redundant structures, and select the CIoU target loss function to improve the bounding box loss function BBox of the YOLOv5 algorithm model. Loss Improvements are made to increase the overlap between the predicted box and the true box;
[0026] Build a face mask recognition system module to organically integrate the trained algorithm into the face recognition system interaction interface and build a real-time face mask recognition system.
[0027] Furthermore, the steps to build the face mask recognition system module are as follows:
[0028] Step 1: Collect public and web crawler face mask datasets;
[0029] Step 2: Use Python programming combined with third-party libraries to convert the data set format;
[0030] Step 3: Use the labeling tool Labelbox to label all the above samples;
[0031] Step 4: Expand the labeled sample data set. The expansion methods include: horizontal flipping, small angle rotation, scaling, cropping, translation, and Fancy PCA.
[0032] Step 5: Split the data set into two categories: faces wearing masks and faces not wearing masks, with 80% used as the training set and 20% as the test set, so that the trained model is more realistic and reliable;
[0033] Step 6: Save the training model and test model, and use Python to encapsulate them into exe executable files.
[0034] Furthermore, the present invention also proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the improved YOLOv5 face mask recognition method described above when executing the computer program.
[0035] The present invention adopts the above technical solution, and compared with the prior art, its significant technical effects are as follows:
[0036] 1. A compression strategy is used for the main network structure of the YOLOv5 algorithm model, that is, a network pruning method is used to obtain a lighter and more efficient model; a hidden layer pruning method is proposed to adjust the network depth, and a convolution kernel pruning method is used to adjust the network width and a CIoU target loss function is introduced to enhance the detection effect and running speed of dense small targets and improve the model perception.
[0037] 2. The improved YOLOv5 face mask recognition algorithm has an average accuracy value of 5% higher than the original algorithm, faster convergence speed, and smaller loss function value; the optimal trained model is selected and loaded into the face recognition system interactive interface to build a real-time face mask recognition system. While the performance is improved, it can also be organically integrated into daily life needs, which reflects the scientificity, effectiveness and practicality of this method. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] Figure 1 This is the network structure of the improved YOLOv5 algorithm model.
[0039] Figure 2The main steps to implement the face mask recognition system.
[0040] Figure 3 A system for image detection for face mask recognition system.
[0041] Figure 4 A system for video detection of face mask recognition system. DETAILED DESCRIPTION
[0042] The present invention is further described in detail below in conjunction with the accompanying drawings and specific implementation methods.
[0043] To achieve the above purpose, the present invention improves the YOLOv5 algorithm model in three aspects, including: network structure improvement, network strategy improvement and building a face mask recognition system, which is applied to the recognition of face masks. The specific steps are as follows:
[0044] S1. Network structure improvement: Use network pruning method to reduce the network parameters and structural complexity of the YOLOv5 algorithm model, and add a CBAM convolution block attention module after the CSP module.
[0045] The network pruning method is mainly used to improve the generalization performance of the network, and avoid overfitting by reducing network parameters and structural complexity, thereby obtaining a lighter and more efficient YOLOv5 algorithm model. The CBAM convolutional block attention module is a lightweight module. When a feature map is given, CBAM will derive the attention map along the two independent dimensions of channel and space in sequence, use CBAM to extract the attention area, and then convolve the attention map with the input feature map for adaptive feature refinement, so as to achieve more accurate positioning of the mask position information in the face area.
[0046] S2. Network strategy improvement: Use hidden layer pruning and convolution kernel pruning to adjust the depth and width of the YOLOv5 algorithm model, cut redundant structures, and select CIoU target loss function to improve the YOLOv5 model bounding box loss function BBox Loss The accuracy of the bounding box is improved by increasing the overlap between the predicted box and the true box, improving the regression accuracy and reducing the prediction error. The improved bounding box loss function is as follows:
[0047]
[0048] Among them, S 2 It represents the number of grids on the three scale feature maps; n refers to the number of candidate boxes generated by each grid; I ij obj Represents whether the jth candidate box in the i-th grid can be responsible for the currently predicted target object. If the overlap of the candidate boxes is greater than the selected threshold, I ij obj =1, otherwise Iij obj =0 and the threshold is selected from a value between 0 and 1 according to the actual application conditions, and the value is retained to one decimal place at most; is the width of the marker border; Is the height of the marker border.
[0049] The improvement of network strategy is mainly divided into two aspects: First, the number of residual components in the CSP structure is controlled by the hidden layer pruning method, thereby changing the network depth. The network depth comparison between the improved YOLOv5 algorithm model and the original YOLOv5 algorithm model is shown in Table 1. As can be seen from Table 1, compared with other YOLOv5 algorithm models, the improved algorithm model has different CSP modules on the Backbone and Neck networks.
[0050] Table 1 Network depth of four YOLOv5 algorithm models
[0051]
[0052] The second is to change the network width by controlling the number of convolution kernels in the Focus and CBL structures through the convolution kernel pruning method, thereby improving the target detection speed and average accuracy. The comparison of the network width between the improved YOLOv5 algorithm model and the original YOLOv5 algorithm model is shown in Table 2. As can be seen from Table 2, compared with other YOLOv5 algorithm models, the improved algorithm model selects different numbers of convolution kernels in different module structures.
[0053] Table 2 Network width of four YOLOv5 algorithm models
[0054]
[0055] The classification loss part in the CIoU objective loss function is used to measure the performance of the model in face mask recognition and classification, and the bounding box regression loss part is used when the bounding boxes overlap.
[0056] The improved network structure of the YOLOv5 algorithm model is as follows Figure 1 As shown in the figure, it mainly targets the Backbone network and the Neck network. In the improved YOLOv5 algorithm model, the Backbone network is divided into three layers, namely:
[0057] The first layer is the Focus module, which slices the image and fully extracts features to retain more effective information. Its purpose is to reduce the amount of model calculation and speed up model training. The specific structure is as follows: First, the input image dataset (three-channel image, size 608×608×3) is divided into 4 slices (slice size 304×304×3) through slicing operation; secondly, the four slices are deeply connected using the Concat operation to generate a feature map (image size 304×304×12); then a convolution operation is performed using a convolution layer composed of 40 convolution kernels to generate a new feature map (image size 304×304×40); finally, the output result is generated through batch normalization (BN) and leaky (LeakyReLU) activation function, and output to the CBL module.
[0058] The second layer is the CBL module, which is the smallest component in the YOLO network structure and the main component of the Backbone and Neck networks. It is mainly composed of convolutional layers, batch normalization layers (BN) and ReLU activation functions. The number of convolutional kernels in the convolutional layer determines the size of the output image of the CBL module.
[0059] The third layer is the CSP1_X module. It contains two CBL blocks and X residual components (Resunit), which can better extract the deep features of the image. The operating principle of the CSP1_X module is to pass the initial input value to the two branches, perform corresponding convolution operations in the two branches respectively, and then continuously calculate and deeply connect the output feature maps of the two branches, and then perform batch normalization and omission ReLU activation function processing. Finally, the CBL module is convolved. The size of the output feature map is the same as the size of the original input feature map of the CSP1_X module. In order to more accurately locate the mask position information in the face area, the CBAM convolution block attention module is added. When a feature map is given, CBAM will derive the attention map along the two independent dimensions of channel and space in turn, use CBAM to extract the attention area, and then convolve the attention map with the input feature map for adaptive feature refinement.
[0060] The CSP2_X modules of the Neck network are all CSP2_1.
[0061] S3. Build a face mask recognition system: organically integrate the trained algorithm into the face recognition system interaction interface to build a real-time face mask recognition system.
[0062] The main steps for building a face mask recognition system are as follows: Figure 2As shown. First, the data set is obtained: the data set of faces wearing masks mainly comes from the FaceMaskDetection and web crawler collection of the AIZOOTe team. The data set contains images that can detect faces and determine whether they are wearing masks, as well as 7959 mask annotation data with open source code; the face images without masks mainly come from the WIDERFACE data set, and 6896 face images are annotated by the LabelMe software after screening; to ensure the diversity of the data set, the total amount of the two types of data reached 20078 images after processing by web crawler technology and data expansion algorithm; in the model training process, 12,000 training sets, 4,000 validation sets, and 4,000 test sets are selected in proportion. During the experiment, the data is in VOC2007 format, and the two types of data sets (mask, no_mask) are labeled by category using labelImg software. Then, PyTorch-GPU1.12 combined with OpenCV is used to import the collected data into the deep learning framework for training, and the optimal model is obtained. After adjusting the parameters, the optimal model is selected and saved as a .pt file. Then use Python language to write the face mask recognition system interface. Finally, encapsulate the entire system into an exe executable file and embed the built system into a computing host or a board with an operating system for operation.
[0063] Figure 3 and Figure 4 The facial mask recognition system that has been built is presented. The system integrates two subsystems, the image detection system and the video detection system. The corresponding functions can be completed according to the button prompts. The real-time monitoring function of the camera is efficient, real-time, accurate and multi-target.
[0064] When recognizing a face mask image, click the exe executable file to run the system, select a picture to be detected, click to upload the picture, and then click to start recognition. The system will frame the face. If the face is wearing a mask, it will be displayed as masked and the current recognition accuracy is 0.94.
[0065] When recognizing videos, it is divided into local video recognition and real-time recognition of camera video streams. Select real-time recognition of camera video streams to operate, click the exe executable file to run the system, and click real-time monitoring of cameras. At this stage, the face in front of the camera can be recognized by masks. The system will frame the face in front of the camera. If the face is not wearing a mask, it will display no_mask and the current recognition accuracy is 0.94; if the face is wearing a mask, it will display masked and the current recognition accuracy is 0.95.
[0066] The present invention also proposes a face mask recognition system based on an improved YOLOv5, comprising:
[0067] The network structure improvement module is used to reduce the network parameters and structural complexity of the YOLOv5 algorithm model using the network pruning method, and add the CBAM convolution block attention module after the CSP module to improve the generalization performance of the network;
[0068] The network strategy improvement module is used to adjust the depth and width of the YOLOv5 algorithm model by using hidden layer pruning and convolution kernel pruning, reduce redundant structures, and select the CIoU target loss function to adjust the YOLOv5 model bounding box loss function BBox Loss The accuracy of the bounding box is improved by increasing the overlap between the predicted box and the real box. The improved bounding box loss function is as follows:
[0069]
[0070] Among them, S 2 It represents the number of grids on the three scale feature maps; n refers to the number of candidate boxes generated by each grid; I ij obj Represents whether the jth candidate box in the i-th grid can be responsible for the currently predicted target object. If the overlap of the candidate boxes is greater than the selected threshold, I ij obj =1, otherwise I ij obj =0 and the threshold is selected from a value between 0 and 1 according to the actual application conditions, and the value is retained to one decimal place at most; is the width of the marker border; Is the height of the marker border.
[0071] Build a face mask recognition system module to organically integrate the trained algorithm into the face recognition system interaction interface and build a real-time face mask recognition system.
[0072] The steps to build the face mask recognition system module are as follows:
[0073] Step 1: Collect public and web crawler face mask datasets;
[0074] Step 2: Use Python programming combined with third-party libraries to convert the data set format;
[0075] Step 3: Use the labeling tool Labelbox to label all the above samples;
[0076] Step 4: Expand the labeled sample data set. The expansion methods include: horizontal flipping, small angle rotation, scaling, cropping, translation, and Fancy PCA.
[0077] Step 5: Split the data set into two categories: faces wearing masks and faces not wearing masks, with 80% used as the training set and 20% as the test set, so that the trained model is more realistic and reliable;
[0078] Step 6: Save the training model and test model, and use Python to encapsulate them into exe executable files.
[0079] In addition, the present invention proposes an electronic device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the improved YOLOv5 face mask recognition method described above when executing the computer program.
[0080] It should be noted that the description of the system and the electronic device in the embodiments of the present application is similar to the description of the above-mentioned method embodiments, and has similar beneficial effects as the method embodiments, so it is not elaborated here. For details, please refer to the method provided in the embodiments of the present invention.
[0081] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general-purpose computer, a special-purpose computer or other programmable data processing device, so that the program code, when executed by the processor or controller, implements the functions / operations specified in the flow chart and / or block diagram. The program code can be executed entirely on the machine, partially on the machine, partially on the machine as a stand-alone software package and partially on a remote machine, or entirely on a remote machine or server.
[0082] The above embodiments are only for illustrating the technical idea of the present invention, and cannot be used to limit the protection scope of the present invention. Any changes made on the basis of the technical solution in accordance with the technical idea proposed by the present invention shall fall within the protection scope of the present invention.
Claims
1. A face mask recognition method based on improved YOLOv5, characterized in that: The specific steps are as follows: S1. Network structure improvement: Use network pruning method to reduce the network parameters and structural complexity of YOLOv5 algorithm model, and add CBAM convolution block attention module after CSP module; S2. Improvement of network strategy: Use hidden layer pruning and convolution kernel pruning to adjust the depth and width of the YOLOv5 algorithm model respectively, reduce redundant structures, and select the CIoU target loss function; the classification loss part in the CIoU target loss function is used to measure the performance of the model in face mask recognition and classification, and the bounding box regression loss part is used when the bounding boxes overlap; select the CIoU target loss function for the bounding box loss function BBox of the YOLOv5 algorithm model Loss The improved bounding box loss function is as follows: Among them, S 2 is the number of grids on the three scale feature maps; n is the number of candidate boxes generated by each grid; Indicates whether the jth candidate box in the i-th grid can be responsible for the currently predicted target object. If the overlap of the candidate boxes is greater than the selected threshold, otherwise is the width of the marker border; is the height of the marker border; In the improved bounding box loss function, with the variable The relevant threshold value is selected from 0 to 1 according to the actual application conditions. If the value is not a boundary value, only one decimal place can be retained; S3, building a face mask recognition system: using the improved YOLOv5 algorithm model in steps S1 and S2, the optimized algorithm model is obtained after training with the data set, and the model is embedded in the face recognition system interaction interface to build a face mask recognition system for real-time applications; wherein S301, collecting public and web crawler-based face mask datasets; S302, use Python programming combined with a third-party library to convert the format of the data set; S303, using the labeling tool Labelbox to label the data set after the format conversion; S304, expanding the labeled sample data set; S305, splitting the data set of faces wearing masks and faces not wearing masks into two categories, with 80% used as a training set and 20% used as a test set; S306, saving the training model and the test model, encapsulating them into an exe executable file using python; embedding them into a computing host or a board with an operating system for operation; the face mask recognition system includes a picture detection system and a video detection system.
2. The face mask recognition method based on improved YOLOv5 according to claim 1, characterized in that: The CBAM convolutional block attention module described in step S1 is used to derive an attention map along two independent dimensions, channel and space, in sequence when a feature map is given, use CBAM to extract the attention area, and then convolute the attention map with the input feature map to perform adaptive feature refinement.
3. The face mask recognition method based on improved YOLOv5 according to claim 1, characterized in that: In the network strategy improvement described in step S2, the hidden layer pruning method adjusts the network depth by controlling the number of residual components in the CSP structure, and the convolution kernel pruning method adjusts the network width by controlling the number of convolution kernels in the Focus and CBL structures.
4. The face mask recognition method based on improved YOLOv5 according to claim 1, characterized in that: In the YOLOv5 algorithm model improved by steps S1 and S2, the Backbone network is divided into three layers: the first layer is the Focus module, which slices the image and fully extracts features; The second layer is the CBL module, which includes a convolutional layer, a batch normalization layer, and a ReLU activation function. The number of convolution kernels in the convolutional layer determines the size of the output image of the CBL module. The third layer is the CSP1_X module. The CSP2_X modules on the Neck network side are all CSP2_1.
5. A face mask recognition system based on improved YOLOv5, characterized in that: include: The network structure improvement module is used to reduce the network parameters and structural complexity of the YOLOv5 algorithm model using the network pruning method, and add the CBAM convolution block attention module after the CSP module to improve the generalization performance of the network; The network strategy improvement module is used to adjust the depth and width of the YOLOv5 algorithm model by using hidden layer pruning and convolution kernel pruning, respectively, to reduce redundant structures, and select the CIoU target loss function to improve the bounding box loss function BBox of the YOLOv5 algorithm model. Loss Improvements are made to increase the overlap between the predicted box and the true box; the improved bounding box loss function is as follows: Among them, S 2 It represents the number of grids on the three scale feature maps; n refers to the number of candidate boxes generated by each grid; Indicates whether the jth candidate box in the i-th grid can be responsible for the currently predicted target object. If the overlap of the candidate boxes is greater than the selected threshold, otherwise The threshold is selected from a value between 0 and 1 according to the actual application conditions, and the value is retained to one decimal place at most; is the width of the marker border; is the height of the marker border; Build a face mask recognition system module to organically integrate the trained algorithm into the face recognition system interaction interface and build a real-time face mask recognition system; the specific contents are: Step 1: Collect public and web crawler face mask datasets; Step 2: Use Python programming combined with third-party libraries to convert the data set format; Step 3: Use the labeling tool Labelbox to label the dataset after the above format conversion; Step 4: Expand the labeled sample data set. The expansion methods include: horizontal flipping, small angle rotation, scaling, cropping, translation, and Fancy PCA. Step 5: Split the dataset into two types: faces wearing masks and faces not wearing masks, with 80% as the training set and 20% as the test set; Step 6: Save the training model and test model, and use Python to encapsulate them into exe executable files.
6. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, the steps of the method according to any one of claims 1 to 4 are implemented.