Mask-wearing detection method, system and device based on improved YOLOV7 and storage medium

CN116363728BActive Publication Date: 2026-10-09SHANGHAI JINLING INFORMATION TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202310121534.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-02-15
Publication Date
2026-10-09
Estimated Expiration
2043-02-15

AI Technical Summary

Technical Problem

[0005]本发明的主要目的在于提供了一种基于改进YOLOV7的佩戴口罩检测方法、系统、设备及存储介质,旨在解决现有技术中佩戴口罩检测速度较慢、精度低、边缘端不易部署的技术问题

Benefits of technology

[0034]本发明通过获取视频帧画面,然后将所述视频帧画面输入至目标检测网络模型进行识别,并输出识别结果,其中,目标检测网络模型为通过构建自校正卷积模块SCNet,取代YOLOV7模型特征提取网络中的常规卷积模块,将SCNet模块原有的平均池化采样改用最大池化,并在此基础上添加CBAM注意力机制,形成改进的网络模型。相比于现有技术,本发明将SCNet模块原有的平均池化采样改用最大池化,提高特征图辨识度的同时,减少了无用信息的影响,并添加的CBAM注意力机制,在提高改进的网络模型精确性的同时,使其更适合在边缘端部署。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116363728B_ABST
    Figure CN116363728B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of computer vision, and discloses a mask wearing detection method, system and device based on an improved YOLOV7 and a storage medium, the method comprising the following steps: acquiring a video frame picture; inputting the video frame into a target detection network model for identification and outputting an identification result, wherein the target detection network model is formed by replacing a conventional convolution module in a feature extraction network of a YOLOV7 model with a self-correcting convolution module SCNet, replacing original average pooling sampling of the SCNet module with maximum pooling, adding a CBAM attention mechanism on the basis, and forming an improved network model. Compared with the prior art, the original average pooling sampling of the SCNet module is replaced with maximum pooling, the influence of useless information is reduced while the feature map recognition degree is improved, and the added CBAM attention mechanism makes the model more suitable for edge deployment while improving the model accuracy.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to a method, system, device, and storage medium for detecting mask wearing based on an improved YOLOv7. Background Technology

[0002] Currently, respiratory infectious diseases are extremely difficult to control and require collective prevention from all of humanity. Clinical findings show that wearing masks correctly in the public is one of the most effective ways to prevent respiratory infectious diseases. At present, most places use staff reminders to check whether masks are worn correctly when entering or exiting certain special locations, while a few places use intelligent detection equipment assisted by personnel for checks. Regardless of the method, it demonstrates that advising the public to wear masks correctly to prevent respiratory infectious diseases is a necessary yet difficult task to implement. Especially in high-traffic areas such as subway entrances and hospital entrances, the low efficiency of checks and the increased crowd density due to congestion not only make management difficult but also increase the efficiency of bacterial and viral transmission. Therefore, efficiently addressing the issue of whether testing personnel are wearing masks correctly is urgently needed.

[0003] In recent years, with the continuous development of video surveillance and deep learning, real-time monitoring and detection have brought significant improvements in the efficiency of addressing the issue of mask-wearing. Currently, existing mask-wearing detection algorithms have achieved some success, but they still frequently suffer from slow detection speeds and inaccurate accuracy, often requiring manual verification.

[0004] There is an urgent need for a mask-wearing detection method based on an improved YOLOv7 to solve the technical problems of slow detection speed, low accuracy, and difficulty in deployment at the edge of the device in the existing technology. Summary of the Invention

[0005] The main objective of this invention is to provide a mask-wearing detection method, system, device, and storage medium based on an improved YOLOV7, aiming to solve the technical problems of slow mask-wearing detection speed, low accuracy, and difficulty in deployment at the edge in the prior art.

[0006] To achieve the above objectives, the present invention provides a mask-wearing detection method based on an improved YOLOv7, the method comprising the following steps:

[0007] Acquire video frames;

[0008] The video frame is input into the object detection network model for recognition and the recognition result is output. The object detection network model is an improved network model that replaces the conventional convolutional module in the YOLOv7 model feature extraction network by constructing a self-correcting convolutional module SCNet. The original average pooling sampling of the SCNet module is replaced with max pooling, and a CBAM attention mechanism is added on this basis.

[0009] Optionally, before the step of acquiring video frame images, the method further includes:

[0010] Collect images of people wearing masks and not wearing masks, perform pre-classification, and obtain an image dataset;

[0011] Optimize the YOLOv7 model to obtain a trained model;

[0012] The image dataset is input into the training model for training, and the training results are obtained.

[0013] Based on the training results, a target detection network model is selected.

[0014] Optionally, the step of optimizing the YOLOv7 model to obtain a trained model includes:

[0015] A self-calibrating convolutional module SCNet was established to replace the conventional convolutional module in the feature extraction network of the YOLOV7 model, thus obtaining the first optimized model.

[0016] The original average pooling sampling of the SCNet module was replaced with max pooling to obtain the second optimized model;

[0017] Add the CBAM attention mechanism to the second optimized model to obtain the trained model.

[0018] Optionally, the step of collecting portrait images of people wearing masks and not wearing masks, performing pre-classification, and obtaining an image dataset includes:

[0019] Collect facial images of people wearing masks and not wearing masks, and perform pre-classification;

[0020] Based on the pre-classification, the portrait images are labeled with location and category labels to obtain an image dataset.

[0021] Optionally, the step of inputting the image dataset into the training model for training and obtaining training results includes:

[0022] The image dataset is used as training samples and input into the training model to obtain multiple trained models. The loss function of the model is calculated during training.

[0023] The training results are obtained based on the loss function.

[0024] Optionally, the step of selecting the target detection network based on the training results specifically includes:

[0025] Based on the training results, the optimal training model is selected as the object detection network.

[0026] Optionally, after the step of inputting the portrait image into the object detection network model for recognition and outputting the recognition result, the method further includes:

[0027] The identification result is sent to the receiving end of the terminal device;

[0028] The receiving end of the terminal device displays the identification result and determines whether to trigger an alarm based on the displayed content.

[0029] Furthermore, to achieve the above objectives, this invention also proposes a mask-wearing detection system based on an improved YOLOv7, the system comprising:

[0030] The acquisition module is used to acquire video frame images, which include facial images of people wearing masks and not wearing masks;

[0031] The recognition module is used to input the video frame images into the target detection network model for recognition and output the recognition results. The target detection network model is an improved network model formed by constructing a self-correcting convolutional module SCNet to replace the conventional convolutional module in the YOLOv7 model feature extraction network, replacing the original average pooling sampling of the SCNet module with max pooling, and adding the CBAM attention mechanism on this basis.

[0032] Furthermore, to achieve the above objectives, the present invention also proposes a mask-wearing detection device based on an improved YOLOV7, the device comprising: a memory, a processor, and a mask-wearing detection program based on an improved YOLOV7 stored in the memory and executable on the processor, the mask-wearing detection program based on an improved YOLOV7 being configured to implement the steps of the mask-wearing detection method based on an improved YOLOV7 as described above.

[0033] Furthermore, to achieve the above objectives, the present invention also proposes a storage medium storing a mask-wearing detection program based on an improved YOLOv7, wherein when the mask-wearing detection program based on the improved YOLOv7 is executed by a processor, the mask-wearing detection program based on the improved YOLOv7 implements the steps of the mask-wearing detection method based on the improved YOLOv7 described above.

[0034] This invention acquires video frames, inputs them into an object detection network model for recognition, and outputs the recognition results. The object detection network model replaces the conventional convolutional modules in the YOLOv7 model's feature extraction network with a self-calibrating convolutional module (SCNet). The original average pooling sampling of the SCNet module is replaced with max pooling, and a CBAM attention mechanism is added to form an improved network model. Compared to existing technologies, this invention improves feature map discriminability by replacing average pooling sampling with max pooling, while reducing the influence of useless information. The added CBAM attention mechanism further enhances the accuracy of the improved network model and makes it more suitable for deployment at edge environments. Attached Figure Description

[0035] Figure 1 This is a schematic diagram of the structure of a mask-wearing detection device based on an improved YOLOV7, which is part of the hardware operating environment of the embodiment of the present invention.

[0036] Figure 2 This is a flowchart illustrating the first embodiment of the mask-wearing detection method based on the improved YOLOv7 of the present invention.

[0037] Figure 3 This is a flowchart illustrating the second embodiment of the mask-wearing detection method based on the improved YOLOv7 of the present invention;

[0038] Figure 4 This is a flowchart illustrating the third embodiment of the mask-wearing detection method based on the improved YOLOv7 of the present invention.

[0039] Figure 5 This is a structural block diagram of the first embodiment of the mask-wearing detection system based on the improved YOLOV7 of the present invention.

[0040] The realization of the objective, functional features and advantages of the present invention will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0041] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.

[0042] Reference Figure 1 , Figure 1 This is a schematic diagram of the structure of a mask-wearing detection device based on an improved YOLOV7, which is part of the hardware operating environment of the embodiment of the present invention.

[0043] like Figure 1As shown, the mask-wearing detection device based on the improved YOLOv7 may include: a processor 1001, such as a central processing unit (CPU), a communication bus 1002, a user interface 1003, a network interface 1004, and a memory 1005. The communication bus 1002 is used to enable communication between these components. The user interface 1003 may include a display screen or an input unit such as a keyboard; optionally, the user interface 1003 may also include a standard wired interface or a wireless interface. The network interface 1004 may optionally include a standard wired interface or a wireless interface (such as a Wireless-Fidelity (Wi-Fi) interface). The memory 1005 may be high-speed random access memory (RAM) or stable non-volatile memory (NVM), such as a disk drive. The memory 1005 may also optionally be a storage device independent of the aforementioned processor 1001.

[0044] Those skilled in the art will understand that Figure 1 The structure shown does not constitute a limitation on mask-wearing detection devices based on the improved YOLOV7 and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0045] like Figure 1 As shown, the memory 1005, which serves as a storage medium, may include an operating system, a network communication module, a user interface module, and a mask-wearing detection program based on an improved YOLOv7.

[0046] exist Figure 1 In the mask-wearing detection device based on the improved YOLOv7 shown, the network interface 1004 is mainly used for data communication with the network server; the user interface 1003 is mainly used for data interaction with the user; the processor 1001 and memory 1005 in the mask-wearing detection device based on the improved YOLOv7 of the present invention can be set in the mask-wearing detection device based on the improved YOLOv7. The mask-wearing detection generation device based on the improved YOLOv7 calls the mask-wearing detection program based on the improved YOLOv7 stored in the memory 1005 through the processor 1001 and executes the mask-wearing detection method based on the improved YOLOv7 provided in the embodiment of the present invention.

[0047] This invention provides a mask-wearing detection method based on an improved YOLOv7, referring to... Figure 2 , Figure 2This is a flowchart illustrating the first embodiment of the mask-wearing detection method based on the improved YOLOV7 of the present invention.

[0048] In this embodiment, the mask-wearing detection method based on the improved YOLOv7 includes the following steps:

[0049] Step S10: Acquire video frame images;

[0050] It should be noted that the executing entity in this embodiment can be a computing service device with data processing and program execution functions, such as a tablet computer or personal computer, or an electronic device capable of performing the same or similar functions, such as the one described above. Figure 1 The illustrated example is a mask-wearing detection device based on the improved YOLOv7. The following description uses the mask-wearing detection device based on the improved YOLOv7 as an example to illustrate this embodiment and the embodiments described below.

[0051] It should be noted that videos are composed of still images, which are called frames. A video frame is a dynamic picture composed of a series of still images.

[0052] It is understood that the aforementioned video frames may be video frames containing images of people wearing masks and not wearing masks, which are obtained in real time through monitoring, or video frames containing images of people wearing masks and not wearing masks, which can be obtained through other means. This embodiment does not limit this.

[0053] Step S20: Input the video frame into the object detection network model for recognition and output the recognition result. The object detection network model is an improved network model formed by constructing a self-correcting convolutional module SCNet to replace the conventional convolutional module in the YOLOV7 model feature extraction network, replacing the original average pooling sampling of the SCNet module with max pooling, and adding the CBAM attention mechanism on this basis.

[0054] It should be noted that the above object detection network model is an improvement on the YOLOv7 model. The self-calibrating convolutional module SCNet replaces the conventional convolutional module in the YOLOv7 model's feature extraction network. The original average pooling sampling of the SCNet module is replaced with max pooling, and the CBAM attention mechanism is added on this basis to construct a network model for detecting mask wearing.

[0055] Understandably, the above output results can be either "not wearing a mask" or "wearing a mask".

[0056] It should be understood that the above video frames not only contain images of people, but also other images that do not contain people or images without human faces. By filtering the above video frames and removing images that do not contain people or images without human faces, images of people wearing masks and not wearing masks are obtained.

[0057] In the specific implementation, images of people wearing masks and those not wearing masks, selected from video frames, are input into the object detection network model for mask recognition, and then the recognition results of the object detection network model are output.

[0058] This embodiment acquires video frames, inputs them into an object detection network model for recognition, and outputs the recognition results. The object detection network model replaces the conventional convolutional modules in the YOLOv7 model's feature extraction network with a self-calibrating convolutional module (SCNet). The original average pooling sampling of the SCNet module is replaced with max pooling, and a CBAM attention mechanism is added to form an improved network model. Compared to existing technologies, this invention replaces the original average pooling sampling of the SCNet module with max pooling, improving feature map discriminability while reducing the influence of useless information. The added CBAM attention mechanism improves the accuracy of the improved network model and makes it more suitable for deployment at edge environments.

[0059] refer to Figure 3 , Figure 3 This is a flowchart illustrating the second embodiment of the mask-wearing detection method based on the improved YOLOV7 of the present invention.

[0060] Based on the first embodiment described above, in this embodiment, before step S10, the method further includes:

[0061] Step S01: Collect images of people wearing masks and not wearing masks, perform pre-classification, and obtain an image dataset;

[0062] It should be noted that, in order to effectively achieve the pre-classification of the above-mentioned portrait images and obtain the image dataset, step S01 includes:

[0063] Step S011: Collect facial images of people wearing masks and not wearing masks, and perform pre-classification;

[0064] Step S012: Based on the pre-classification, label the portrait image with location and category labels to obtain the image dataset.

[0065] It should be noted that the location label indicates the position of the face in the portrait image; the category label indicates whether the person in the portrait image is wearing a mask.

[0066] It is understood that the annotation of human portrait images can be done manually or through a specific algorithm, and this embodiment and the following embodiments do not limit this.

[0067] In practice, both images of people wearing masks and those not wearing masks are labeled with location and category tags.

[0068] Step S02: Optimize the YOLOv7 model to obtain the trained model;

[0069] It should be noted that, in order to effectively optimize the YOLOv7 model and obtain the trained model, step S02 includes:

[0070] Step S021: Establish a self-calibrating convolutional module SCNet to replace the conventional convolutional module in the YOLOv7 model feature extraction network and obtain the first optimized model;

[0071] Step 1: Divide the input feature map X of size C*H*W into two parts, X1 and X2, of size C / 2*H*W;

[0072] It is understandable that the input feature map X mentioned above refers to the image after preprocessing the portrait image, that is, the portrait image labeled with location labels and category labels;

[0073] It should be noted that C, H, and W are concepts in image channel conversion, where C, H, and W represent the number of channels, height, and width of the image, respectively.

[0074] Step 2: Divide the convolution kernel K with shape [C, C, Kh, Kw] into 4 parts, denoted as K1, K2, K3, K4, each with shape [C / 2, C / 2, Kh, Kw].

[0075] It should be understood that, in image processing, a convolution kernel is a weighted average of pixels in a small region of the input image, which is then used to produce each corresponding pixel in the output image. The weights are defined by a function called the convolution kernel.

[0076] It should be noted that C represents the number of channels in the image, Kh is the convolution kernel in the height direction, and Kw is the convolution kernel in the width direction; [C, C, Kh, Kw] is a whole, representing the dimension of the convolution kernel K.

[0077] Step S022: Replace the original average pooling sampling of the SCNet module with max pooling to obtain the second optimized model;

[0078] Step 3: Process the self-correction space. This process can be described as follows: the features obtained by downsampling X1 by a factor of 4 and then upsampling them, along with the features obtained by simultaneously convolving X1 with K3, are then corrected using the Sigmoid activation function to obtain feature Y1.

[0079] T1 = AvgPool r (X1)

[0080] X'1=Up(F2(T1))=Up(T1*K2)

[0081] Y1'=F3(X1)·σ(X1+X'1)

[0082] Y1=F4(Y1'=Y1'*K4

[0083] In the formula, T1 represents the result after downsampling X1 by a factor of 4 using average pooling; AvgPool r This indicates average pooling by 4 times; Up is upsampling; F represents convolution operation; σ represents the Sigmoid activation function; X and Y only represent intermediate values ​​in the calculation process.

[0084] The fourth step is to extract feature Y2 from feature X2 through K1 convolution, and then concatenate Y1 and Y2 to obtain feature Y.

[0085] Step S023: Add the CBAM attention mechanism to the second optimized model to obtain the trained model.

[0086] First, the feature map is compressed in the spatial dimension, and then two one-dimensional vectors are obtained through two pooling functions. Therefore, the spatial attention extraction mechanism Mc can be represented as:

[0087]

[0088] In the formula, AvgPool and MaxPool represent average pooling and max pooling, respectively; MLP is the same as W0 and W1, representing two layers of parameters in the multilayer perceptron model; and σ represents the sigmoid activation function.

[0089] The feature map is then compressed along the channel dimension, resulting in two one-dimensional vectors through two pooling functions. Therefore, the attention extraction mechanism Ms along the channel dimension can be represented as:

[0090]

[0091] In the formula, F is the input feature map, which is also the output Y of SCNet. The other F values ​​are feature maps after pooling. W represents the parameters of the two layers in the perceptron model. 7×7 This indicates that a 7×7 convolution operation is being performed.

[0092] Therefore, the CBAM attention mechanism mainly performs the following operations:

[0093]

[0094]

[0095] In the formula, CBAM represents the Kronecker product, Mc represents the attention extraction mechanism in the spatial dimension, and Ms represents the attention extraction mechanism in the channel dimension. Because CBAM is relatively lightweight, it is especially suitable for deployment at the edge.

[0096] Step S03: Input the image dataset into the training model for training and obtain the training results;

[0097] It should be noted that, in order to obtain highly accurate training results, step S03 includes:

[0098] Step S031: Input the image dataset as training samples into the training model to train it, obtain multiple trained models, and calculate the loss function of the model during training;

[0099] Step S032: Obtain the training results based on the loss function.

[0100] It should be noted that the loss function mentioned above is the mean squared error loss function.

[0101] Step S04: Select an object detection network model based on the training results.

[0102] In the specific implementation, the aforementioned image dataset with location and category labels is used as the training sample set and input into the improved YOLOv7 model (training model) for training, resulting in multiple trained models. A loss function is calculated during training. The optimal improved YOLOv7 model (training model) is selected as the object detection network model based on the results of the loss function across the multiple trained models.

[0103] This embodiment collects portrait images of people wearing masks and not wearing masks, performs pre-classification, and then labels the images with location and category tags based on the pre-classification to obtain an image dataset. A self-calibrating convolutional module SCNet is then established to replace the conventional convolutional module in the YOLOv7 model's feature extraction network, resulting in a first optimized model. The original average pooling sampling of the SCNet module is replaced with max pooling to obtain a second optimized model. A CBAM attention mechanism is added to the second optimized model to obtain a training model. The image dataset is then used as training samples and input into the training model for training, resulting in multiple trained models. The loss function of each model is calculated during training. The training results are obtained based on the loss function. Finally, the optimal training model is selected as the object detection network based on the training results. Because this invention establishes a self-calibrating convolutional module SCNet to replace the conventional convolutional module in the YOLOv7 model's feature extraction network, and improves the original average pooling sampling of the SCNet module by replacing it with max pooling to improve the discriminative power of the feature maps, and then adds a CBAM attention mechanism, it improves accuracy and is easier to deploy at edge environments. Compared to existing mask detection network models, this invention solves the problems of slow detection speed and low detection accuracy, enabling low-computing-power terminals to run mask detection algorithms, and improving the speed performance, accuracy, and ease of deployment of the algorithm on the terminal.

[0104] refer to Figure 4 , Figure 4 This is a flowchart illustrating the third embodiment of the mask-wearing detection method based on the improved YOLOV7 of the present invention.

[0105] Based on the above embodiments, in this embodiment, after step S20, the method further includes:

[0106] Step S30: Send the recognition result to the receiving end of the terminal device;

[0107] Understandably, the above identification results include those who are wearing masks and those who are not.

[0108] It should be noted that the aforementioned terminal devices may be computers, mobile phones, iPads, or other devices with data display functions, or other electronic devices with the same or similar functions. This embodiment does not impose any restrictions on this.

[0109] Step S40: The terminal device receiving end displays the identification result and determines whether to perform alarm processing based on the displayed content.

[0110] It should be explained that if the recognition result shows that a mask is being worn, no alarm will be triggered; if the recognition result shows that a mask is not being worn, an alarm will be triggered.

[0111] In the specific implementation, the terminal device displays the identification results of whether a mask is being worn, output by the target detection network model, and determines whether to trigger an alarm. If the identification result shows that a mask is being worn, no alarm is triggered; if the identification result shows that a mask is not being worn, an alarm is triggered.

[0112] This embodiment establishes a self-calibrating convolutional module SCNet to replace the conventional convolutional module in the YOLOv7 model's feature extraction network, obtaining a first optimized model; the original average pooling sampling of the SCNet module is replaced with max pooling to obtain a second optimized model; a CBAM attention mechanism is added to the second optimized model to obtain a training model; the image dataset is then used as training samples to input into the training model for training, obtaining multiple trained models, and the loss function of the model is calculated during training; the training result is obtained based on the loss function; finally, based on the training result, the optimal training model is selected as the object detection network; then, video frames are acquired; the video frames are input into the object detection network model for recognition, and the recognition result is output; then, the recognition result is sent to the receiving end of the terminal device, and the recognition result is displayed, and an alarm is triggered based on the displayed content. This invention replaces the conventional convolutional modules in the YOLOv7 model's feature extraction network with a self-calibrating convolutional module, SCNet. It also improves upon the original average pooling sampling of SCNet by using max pooling to enhance the discriminative power of feature maps. Furthermore, the addition of a CBAM attention mechanism further improves accuracy and makes the invention easier to deploy at edge computing environments. Compared to existing mask detection network models, this invention solves the problems of slow detection speed and low accuracy, enabling the execution of mask detection algorithms on low-computing-power terminals and improving the speed, accuracy, and deployment ease of the algorithm.

[0113] Furthermore, this embodiment of the invention also proposes a storage medium storing a mask-wearing detection program based on an improved YOLOv7. When the improved YOLOv7 mask-wearing detection program is executed by a processor, it implements the steps of the mask-wearing detection method based on the improved YOLOv7 described above.

[0114] Reference Figure 5 , Figure 5 This is a structural block diagram of the first embodiment of the mask-wearing detection system based on the improved YOLOV7 of the present invention.

[0115] like Figure 5 As shown, the mask-wearing detection system based on the improved YOLOv7 proposed in this embodiment of the invention includes: an acquisition module 501 and an identification module 502.

[0116] The acquisition module 501 is used to acquire video frame images;

[0117] The recognition module 502 is used to input the video frame into the target detection network model for recognition and output the recognition result. The target detection network model is an improved network model formed by constructing a self-correcting convolutional module SCNet to replace the conventional convolutional module in the YOLOV7 model feature extraction network, replacing the original average pooling sampling of the SCNet module with max pooling, and adding the CBAM attention mechanism on this basis.

[0118] This embodiment acquires video frames; inputs these video frames into an object detection network model for recognition; and outputs the recognition results. The object detection network model replaces the conventional convolutional modules in the YOLOv7 model's feature extraction network with a self-calibrating convolutional module (SCNet). The original average pooling sampling of the SCNet module is replaced with max pooling, and a CBAM attention mechanism is added to form an improved network model. Compared to existing technologies, this invention replaces the original average pooling sampling of the SCNet module with max pooling, improving feature map discriminability while reducing the influence of useless information. The added CBAM attention mechanism further enhances the accuracy of the improved network model, making it more suitable for deployment at edge environments.

[0119] Based on the first embodiment of the mask-wearing detection system based on the improved YOLOV7 of the present invention, a second embodiment of the mask-wearing detection system based on the improved YOLOV7 of the present invention is proposed.

[0120] In this embodiment, the acquisition module 501 is further configured to collect portrait images of people wearing masks and not wearing masks, perform pre-classification, and obtain an image dataset; optimize the YOLOv7 model to obtain a training model; input the image dataset into the training model for training and obtain training results; and select an object detection network model based on the training results.

[0121] The acquisition module 501 is further used to establish a self-calibrating convolutional module SCNet to replace the conventional convolutional module in the YOLOv7 model feature extraction network to obtain a first optimized model; to replace the original average pooling sampling of the SCNet module with max pooling to obtain a second optimized model; and to add a CBAM attention mechanism to the second optimized model to obtain a training model.

[0122] The acquisition module 501 is also used to collect portrait images of people wearing masks and not wearing masks, and perform pre-classification; based on the pre-classification, the portrait images are labeled with location labels and category labels to obtain an image dataset.

[0123] The acquisition module 501 is further configured to input the image dataset as training samples into the training model for training, obtain multiple trained models, and calculate the loss function of the model during training; and obtain the training result based on the loss function.

[0124] The acquisition module 501 is further configured to select the optimal training model as the object detection network based on the training results.

[0125] This embodiment collects portrait images of people wearing masks and not wearing masks, performs pre-classification, and then labels the images with location and category tags based on the pre-classification to obtain an image dataset. A self-calibrating convolutional module SCNet is then established to replace the conventional convolutional module in the YOLOv7 model's feature extraction network, resulting in a first optimized model. The original average pooling sampling of the SCNet module is replaced with max pooling to obtain a second optimized model. A CBAM attention mechanism is added to the second optimized model to obtain a training model. The image dataset is then used as training samples and input into the training model for training, resulting in multiple trained models. The loss function of each model is calculated during training. The training results are obtained based on the loss function. Finally, the optimal training model is selected as the object detection network based on the training results. Because this invention establishes a self-calibrating convolutional module SCNet to replace the conventional convolutional module in the YOLOv7 model's feature extraction network, and improves the original average pooling sampling of the SCNet module by replacing it with max pooling to improve the discriminative power of the feature maps, and then adds a CBAM attention mechanism, it improves accuracy and is easier to deploy at edge environments. Compared to existing mask detection network models, this invention solves the problems of slow detection speed and low detection accuracy, enabling low-computing-power terminals to run mask detection algorithms, and improving the speed performance, accuracy, and ease of deployment of the algorithm on the terminal.

[0126] Other embodiments or specific implementations of the mask-wearing detection system based on the improved YOLOV7 of the present invention can be referred to the above-described method embodiments, and will not be repeated here.

[0127] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or system that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or system. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or system that includes that element.

[0128] The sequence numbers of the above embodiments of the present invention are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.

[0129] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as read-only memory / random access memory, magnetic disk, optical disk) and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of the present invention.

[0130] The above are merely preferred embodiments of the present invention and do not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.

Claims

1. A mask-wearing detection method based on an improved YOLOv7, characterized in that, The method includes the following steps: Acquire video frames; The video frame is input into the target detection network model for recognition and the recognition result is output. The target detection network model is an improved network model that replaces the conventional convolutional module in the YOLOV7 model feature extraction network by constructing a self-correcting convolutional module SCNet. The original average pooling sampling of the SCNet module is replaced with max pooling, and the CBAM attention mechanism is added on this basis. Before the step of acquiring video frame images, the method further includes: Collect images of people wearing masks and not wearing masks, perform pre-classification, and obtain an image dataset; Optimize the YOLOv7 model to obtain a trained model; The image dataset is input into the training model for training, and the training results are obtained. Based on the training results, select the target detection network model; The steps for optimizing the YOLOv7 model to obtain a trained model include: A self-calibrating convolutional module SCNet was established to replace the conventional convolutional module in the feature extraction network of the YOLOV7 model, thus obtaining the first optimized model. The original average pooling sampling of the SCNet module was replaced with max pooling to obtain the second optimized model; Add the CBAM attention mechanism to the second optimized model to obtain the trained model; The steps of the self-calibrating convolutional module SCNet in processing the input feature map include: C H The input feature map X of W is decomposed into two C / 2 along the channel dimension. H Feature submap of W and Where C, H, and W represent the number of channels, height, and width of the image, respectively; The convolution kernel K with shape [C, C, Kh, Kw] is divided into four parts. , , and Each part has a shape of [C / 2, C / 2, Kh, Kw], where Kh is the convolution kernel in the height direction and Kw is the convolution kernel in the width direction; For feature subgraphs Perform self-correction space processing: Features obtained by downsampling by a factor of 4 followed by upsampling, and simultaneously go through The features obtained after convolution are then corrected using the Sigmoid activation function to obtain the final features. ; For feature subgraphs go through Convolution extraction to obtain features ,Will and The final output feature Y of the self-correcting convolutional module SCNet is obtained by concatenating the features.

2. The method as described in claim 1, characterized in that, The steps of collecting images of people wearing masks and not wearing masks, performing pre-classification, and obtaining an image dataset include: Collect facial images of people wearing masks and not wearing masks, and perform pre-classification; Based on the pre-classification, the portrait images are labeled with location and category labels to obtain an image dataset.

3. The method as described in claim 1, characterized in that, The step of inputting the image dataset into the training model for training to obtain the trained model includes: The image dataset is used as training samples and input into the training model to obtain multiple trained models. The loss function of the model is calculated during training. The training results are obtained based on the loss function.

4. The method as described in claim 1, characterized in that, The step of selecting the object detection network model based on the training results specifically includes: Based on the training results, the optimal training model is selected as the object detection network model.

5. The method according to any one of claims 1-4, characterized in that, After the step of inputting the video frame into the target detection network model for recognition and outputting the recognition result, the method includes: The identification result is sent to the receiving end of the terminal device; The receiving end of the terminal device displays the identification result and determines whether to trigger an alarm based on the displayed content.

6. A mask-wearing detection system based on an improved YOLOv7, characterized in that, The system includes: The acquisition module is used to acquire video frame images; The recognition module is used to input the video frame into the target detection network model for recognition and output the recognition result. The target detection network model is an improved network model formed by constructing a self-correcting convolutional module SCNet to replace the conventional convolutional module in the YOLOV7 model feature extraction network, replacing the original average pooling sampling of the SCNet module with max pooling, and adding the CBAM attention mechanism on this basis. The acquisition module is also used to collect portrait images of people wearing masks and not wearing masks, perform pre-classification, and obtain an image dataset; optimize the YOLOv7 model to obtain a training model; input the image dataset into the training model for training and obtain training results; and select an object detection network model based on the training results. The acquisition module is also used to establish a self-correcting convolutional module SCNet to replace the conventional convolutional module in the YOLOV7 model feature extraction network to obtain a first optimized model; to replace the original average pooling sampling of the SCNet module with max pooling to obtain a second optimized model; and to add a CBAM attention mechanism to the second optimized model to obtain a training model. The steps of the self-calibrating convolutional module SCNet in processing the input feature map include: converting C... H The input feature map X of W is decomposed into two C / 2 along the channel dimension. H Feature submap of W and Where C, H, and W represent the number of channels, height, and width of the image, respectively; the convolution kernel K with shape [C, C, Kh, Kw] is divided into four parts. , , and Each part has a shape of [C / 2, C / 2, Kh, Kw], where Kh is the convolution kernel in the height direction and Kw is the convolution kernel in the width direction; for the feature submap Perform self-correction space processing: Features obtained by downsampling by a factor of 4 followed by upsampling, and simultaneously go through The features obtained after convolution are then corrected using the Sigmoid activation function to obtain the final features. ; for feature subgraphs go through Convolution extraction to obtain features ,Will and The final output feature Y of the self-correcting convolutional module SCNet is obtained by concatenating the features.

7. A mask-wearing detection device based on an improved YOLOv7, characterized in that, The device includes: a memory, a processor, and a mask-wearing detection program based on improved YOLOV7 stored in the memory and executable on the processor, the mask-wearing detection program based on improved YOLOV7 being configured to implement the steps of the mask-wearing detection method based on improved YOLOV7 as described in any one of claims 1 to 5.

8. A storage medium, characterized in that, The storage medium stores a mask-wearing detection program based on the improved YOLOV7, which, when executed by a processor, implements the steps of the mask-wearing detection method based on the improved YOLOV7 as described in any one of claims 1 to 5.

Citation Information

Patent Citations

  • Improved yolov5-based mask face detection method

    CN115171183A