Mask wearing recognition method and device, terminal equipment and storage medium
By acquiring and processing images of key facial regions and inputting them into a recognition network, the accuracy of user mask-wearing recognition was solved, thus improving the effectiveness of respiratory disease prevention.
Patent Information
- Application Number
- CN202110241157.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-03-04
- Publication Date
- 2025-12-23
- Estimated Expiration
- 2041-03-04
AI Technical Summary
Existing technologies cannot accurately identify whether users are wearing masks correctly, leading to potential risks in the prevention of respiratory diseases.
By acquiring images of the nose, left corner of the mouth, and right corner of the mouth of the person to be identified, and inputting them into a trained recognition network, the system can determine whether the user is wearing a mask correctly.
It enables accurate identification of whether users are wearing masks correctly, thus improving the effectiveness of respiratory disease prevention.
Smart Images

Figure CN115035560B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application belongs to the technical field of image processing, and particularly relates to a mask wearing recognition method and device, a terminal device, and a storage medium. BACKGROUND
[0002] The new coronavirus pneumonia has made people realize the importance of wearing a mask for preventing respiratory diseases. Wearing a mask has become an important way of preventing respiratory diseases. Correctly wearing a mask can block harmful gases, odors, droplets, viruses and other substances. If a mask is not worn correctly, it will bring great risks to the prevention of respiratory diseases. Therefore, recognizing whether a user correctly wears a mask plays an important role in the prevention of respiratory diseases. SUMMARY
[0003] The present application provides a mask wearing recognition method and device, a terminal device, and a storage medium to recognize whether a user correctly wears a mask.
[0004] In a first aspect, the present application provides a mask wearing recognition method, which comprises the following steps.
[0005] Obtaining a nose image, a left corner of the mouth image and a right corner of the mouth image of a to-be-recognized face, wherein the nose image refers to an image of a region where a nose is located in the to-be-recognized face, the left corner of the mouth image refers to an image of a region where a left corner of the mouth is located in the to-be-recognized face, and the right corner of the mouth image refers to an image of a region where a right corner of the mouth is located in the to-be-recognized face.
[0006] Inputting the nose image, the left corner of the mouth image and the right corner of the mouth image into a trained recognition network to obtain a recognition result of whether the to-be-recognized face wears a mask.
[0007] In a second aspect, the present application provides a mask wearing recognition device, which comprises the following modules.
[0008] A first obtaining module is configured to obtain a nose image, a left corner of the mouth image and a right corner of the mouth image of a to-be-recognized face, wherein the nose image refers to an image of a region where a nose is located in the to-be-recognized face, the left corner of the mouth image refers to an image of a region where a left corner of the mouth is located in the to-be-recognized face, and the right corner of the mouth image refers to an image of a region where a right corner of the mouth is located in the to-be-recognized face.
[0009] An image recognition module is configured to input the nose image, the left corner of the mouth image and the right corner of the mouth image into a trained recognition network to obtain a recognition result of whether the to-be-recognized face wears a mask.
[0010] In a third aspect, an embodiment of the present application provides a terminal device, comprising a memory, a processor, and a computer program stored in the memory and capable of running on the processor, and the processor implements the steps of the mask wearing identification method according to the first aspect when running the computer program.
[0011] In a fourth aspect, an embodiment of the present application provides a computer readable storage medium, which stores a computer program, and the computer program implements the steps of the mask wearing identification method according to the first aspect when executed by a processor.
[0012] In a fifth aspect, an embodiment of the present application provides a computer program product, which, when running on a terminal device, causes the terminal device to perform the steps of the mask wearing identification method according to the first aspect.
[0013] As can be seen from the above, the present application can identify whether the to-be-identified face correctly wears a mask by acquiring a nose image, a left corner of mouth image and a right corner of mouth image of the to-be-identified face, and inputting the acquired images into a trained identification network, thereby obtaining an identification result of the to-be-identified face wearing a mask. BRIEF DESCRIPTION OF DRAWINGS
[0014] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings needed to be used in the embodiments or prior art description will be briefly introduced as follows. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can also be obtained by those skilled in the art without creative labor.
[0015] Figure 1 is an implementation flow diagram of the mask wearing identification method provided by the first embodiment of the present application;
[0016] Figure 2 is an implementation flow diagram of the mask wearing identification method provided by the second embodiment of the present application;
[0017] Figure 3a is an example diagram of image splicing; Figure 3b is another example diagram of image splicing; Figure 3c is a structure example diagram of an identification network;
[0018] Figure 4 is a structure diagram of the mask wearing identification device provided by the third embodiment of the present application;
[0019] Figure 5 is a structure diagram of the terminal device provided by the fourth embodiment of the present application. DETAILED DESCRIPTION
[0020] In the following description, for purposes of explanation and not limitation, specific details are set forth such as particular architectures, techniques, etc. in order to provide a thorough understanding of the embodiments of the present application. However, it will be apparent to those skilled in the art that the present application can be practiced in other embodiments that depart from these specific details. In other instances, detailed descriptions of well-known methods, devices, circuits, and
[0021] It is to be understood that the terminology "includes", "has", "holds", "contains" and / or "comprising", when used in this specification and in the following claims, indicates the presence of the described features, integers, steps, operations, elements, and / or components but does not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0022] It is also to be understood that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to be limiting. As used in this specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.
[0023] It will be further understood that the terms "and / or", "including", "comprising" when used in this specification and in the following claims, specify the presence of features, integers, steps, operations, elements, and / or components with at least one of one or more associated possibilities, and do not preclude the presence or addition of one or more other features, integers, steps, operations, elements, components, and / or groups thereof.
[0024] As used in this specification and the appended claims, the term "if' can, in some instances, be interpreted as meaning "when", or "once", or "in response to a determination" or "in response to detecting". Similarly, the phrase "if a determination" or "if detecting [described condition or event]" can, in some instances, be interpreted to mean "once a determination" or "in response to a determination" or "once detecting [described condition or event]" or "in response to detecting [described condition or event]".
[0025] In particular implementations, the terminal device described in the embodiments of the present application includes, but is not limited to, other portable devices such as mobile phones, laptop computers, or tablet computers having touch-sensitive surfaces (e.g., touch screen displays and / or touch pads). It will also be appreciated that, in some embodiments, the device is not a portable communication device, but is a desktop computer having a touch-sensitive surface (e.g., a touch screen display and / or touch pad).
[0026] In the following discussion, a terminal device including a display and a touch-sensitive surface is described. It will be appreciated, however, that a terminal device can include one or more other physical user-interface devices, such as a physical keyboard, a mouse, and / or a joystick.
[0027] The terminal device supports various applications, such as one or more of the following: a drawing application, a presentation application, a word processing application, a website creation application, a disk authoring application, a spreadsheet application, a game application, a telephone application, a video conferencing application, an e-mail application, an instant messaging application, a workout support application, a photo management application, a digital camera application, a digital camcorder application, a web browsing application, a digital music player application, and / or a digital video player application.
[0028] The various applications that can be executed on the terminal device can use at least one common physical user interface device, such as a touch-sensitive surface. One or more functions of the touch-sensitive surface and corresponding information displayed on the terminal can be adjusted and / or changed between applications and / or within respective applications. In this way, the common physical architecture (e.g., touch-sensitive surface) of the terminal can support various applications with a user interface that is intuitive and transparent to the user.
[0029] It should be understood that the size of the serial number of each step in the embodiments does not mean the order of execution, and the execution order of each process should be determined according to its function and inherent logic, and should not constitute any limitation on the implementation process of the embodiments of the application.
[0030] In order to illustrate the technical solutions described in the present application, the following will be described by specific embodiments.
[0031] Referring to Figure 1 is a schematic diagram of the implementation process of the mask wearing identification method provided by Embodiment One of the present application, which is applied to a terminal device, such as Figure 1 As shown in the figure, the mask wearing identification method can include the following steps:
[0032] Step 101: Obtain the nose image, left corner image and right corner image of the face to be identified.
[0033] Among them, the nose image refers to the image of the region where the nose of the face to be identified is located, the left corner image refers to the image of the region where the left corner of the face to be identified is located, and the right corner image refers to the image of the region where the right corner of the face to be identified is located.
[0034] It should be noted that the terminal device can obtain the nose image, left corner image and right corner image of the face to be identified from the to-be-identified image including the face to be identified, or obtain the nose image, left corner image and right corner image of the face to be identified from other devices, which is not limited here.
[0035] The terminal device obtains a nose image, a left corner of mouth image, and a right corner of mouth image of the to-be-identified face from the to-be-identified image. Specifically, the nose image, the left corner of mouth image, and the right corner of mouth image can be obtained by locating a region where the nose is located, a region where the left corner of mouth is located, and a region where the right corner of mouth is located in the to-be-identified image by using a face key point positioning algorithm. The nose image can be obtained by cutting the image of the region where the nose is located from the to-be-identified image. The left corner of mouth image can be obtained by cutting the image of the region where the left corner of mouth is located from the to-be-identified image. The right corner of mouth image can be obtained by cutting the image of the region where the right corner of mouth is located from the to-be-identified image. When cutting the image of the region where the nose is located from the to-be-identified image, an image of a preset size can be cut with the tip of the nose as the center. The image of the preset size is the nose image. When cutting the image of the region where the left corner of mouth is located from the to-be-identified image, an image of a preset size can be cut with the left corner of mouth as the center. The image of the preset size is the left corner of mouth image. When cutting the image of the region where the right corner of mouth is located from the to-be-identified image, an image of a preset size can be cut with the right corner of mouth as the center. The image of the preset size is the right corner of mouth image.
[0036] The to-be-identified image can be stored by the terminal device, can be collected by an image collection device integrated in the terminal device, or can be obtained from another device, which is not limited herein. The face key points include, but are not limited to, the nose, the left corner of mouth, the right corner of mouth, the center of the left eye, and the center of the right eye.
[0037] As an optional embodiment, before obtaining the nose image, the left corner of mouth image, and the right corner of mouth image from the to-be-identified image, the to-be-identified image can be subjected to face alignment processing to obtain a to-be-identified image after face alignment. The nose image, the left corner of mouth image, and the right corner of mouth image can be obtained from the to-be-identified image after face alignment. This can reduce the influence of the face angle on the recognition result and improve the accuracy of the recognition result.
[0038] The recognition result includes, but is not limited to, that the to-be-identified face correctly wears a mask or that the to-be-identified face does not correctly wear a mask.
[0039] The face alignment processing of the to-be-identified image can be specifically that the affine transformation parameters of the face key points in the to-be-identified image to a face template are calculated, and the to-be-identified image is transformed according to the affine transformation parameters. The transformed image is the to-be-identified image after face alignment. The face template can be obtained by enlarging a face key point template to a specified size, for example, 200*200 pixels.
[0040] In step 102, the nose image, the left corner of mouth image, and the right corner of mouth image are input into the trained recognition network to obtain the recognition result of the to-be-identified face wearing a mask.
[0041] Since the sizes of the nose image, the left corner image and the right corner image are smaller than the to-be-identified image, the nose image, the left corner image and the right corner image can be processed by the identification network with low complexity to accurately obtain the identification result of whether the to-be-identified face correctly wears the mask. The low complexity of the identification network can be understood as that the depth, width, model size and other parameters of the identification network are small.
[0042] The identification network can be a convolutional neural network with high identification accuracy.
[0043] In an optional embodiment, the nose image, the left corner image and the right corner image are input into the trained identification network to obtain the identification result of whether the to-be-identified face wears the mask.
[0044] The nose image, the left corner image and the right corner image are input into the trained identification network to obtain the probability that the to-be-identified face correctly wears the mask.
[0045] The identification result of whether the to-be-identified face wears the mask is determined according to the probability that the to-be-identified face correctly wears the mask.
[0046] Specifically, the terminal device can pre-set a probability threshold, compare the probability that the to-be-identified face correctly wears the mask with the probability threshold, and if the probability that the to-be-identified face correctly wears the mask is greater than the probability threshold, it is determined that the to-be-identified face correctly wears the mask, and if the probability that the to-be-identified face correctly wears the mask is less than or equal to the probability threshold, it is determined that the to-be-identified face does not correctly wear the mask. The correct wearing of the mask can mean that the nose, the left corner and the right corner are not exposed to the air when the mask is worn, and the incorrect wearing of the mask can mean that the mask is not worn or at least one of the three key points, i.e., the nose, the left corner and the right corner, is exposed to the air when the mask is worn.
[0047] When it is determined that the to-be-identified face does not correctly wear the mask, the terminal device can issue a reminder to remind the user to correctly wear the mask. The reminder can include but is not limited to voice reminder, message reminder, alarm reminder, etc.
[0048] If the to-be-identified face does not correctly wear the mask, the nose or the two corners of the to-be-identified face will be exposed to the air, so the nose image and the two corner images of the to-be-identified face can be used to accurately identify whether the to-be-identified face correctly wears the mask. In addition, the mask wearing identification method provided in the embodiment is simple and easy to deploy in the terminal device.
[0049] The embodiment of the application can identify whether the to-be-identified face correctly wears a mask by acquiring a nose image, a left corner of the mouth image and a right corner of the mouth image of the to-be-identified face and inputting the acquired images into a trained recognition network, thereby obtaining a recognition result of the to-be-identified face wearing a mask.
[0050] Referring to Figure 2 is a schematic diagram of an implementation process of a mask wearing recognition method provided in Embodiment Two of the application, and the mask wearing recognition method is applied to a terminal device, such as Figure 2 , the mask wearing recognition method can include the following steps:
[0051] In step 201, a nose image, a left corner of the mouth image and a right corner of the mouth image of a to-be-identified face are acquired.
[0052] This step is the same as step 101, and details can be referred to the related description of step 101, which will not be repeated here.
[0053] In step 202, the nose image, the left corner of the mouth image and the right corner of the mouth image are spliced into a single-channel image.
[0054] The splicing can mean that the nose image, the left corner of the mouth image and the right corner of the mouth image are spliced into a seamless single-channel image, as Figure 3a illustrated in FIG. 2b is an example diagram of image splicing.
[0055] It should be noted that when splicing the three images of the nose image, the left corner of the mouth image and the right corner of the mouth image, the splicing order of the three images can not be limited. As Figure 3b illustrated in FIG. 2c is another example diagram of image splicing, Figure 3b and Figure 3a are example diagrams of different splicing orders of the three images. Figure 3a 32 in FIG. 3b represents an image size, in pixels.
[0056] In an optional embodiment, before splicing the nose image, the left corner of the mouth image and the right corner of the mouth image into a single-channel image, the method further includes:
[0057] graying the nose image, the left corner of the mouth image and the right corner of the mouth image to obtain a nose gray image, a left corner of the mouth gray image and a right corner of the mouth gray image;
[0058] Splicing the nose image, the left corner of the mouth image and the right corner of the mouth image into a single-channel image includes:
[0059] Splicing the nose gray image, the left corner of the mouth gray image and the right corner of the mouth gray image into a single-channel image.
[0060] The nose grayscale image refers to an image obtained by performing grayscale processing on the nose image, the left corner of the mouth grayscale image refers to an image obtained by performing grayscale processing on the left corner of the mouth image, and the right corner of the mouth grayscale image refers to an image obtained by performing grayscale processing on the right corner of the mouth image.
[0061] When determining whether the face to be recognized correctly wears a mask through the recognition network, mainly the texture information of the nose image, the left corner of the mouth image and the right corner of the mouth image is used. In this embodiment, by performing grayscale processing on the nose image, the left corner of the mouth image and the right corner of the mouth image, the complexity of image processing can be reduced without affecting the texture information of the three images, and the robustness of the recognition network is enhanced.
[0062] After performing grayscale processing on the nose image, the left corner of the mouth image and the right corner of the mouth image respectively, the nose grayscale image, the left corner of the mouth grayscale image and the right corner of the mouth grayscale image are all single-channel images. The three single-channel images can be spliced into a single-channel image to reduce the complexity of image processing. For example, the size of the nose grayscale image, the left corner of the mouth grayscale image and the right corner of the mouth grayscale image is 32*32 pixels, and the nose grayscale image, the left corner of the mouth grayscale image and the right corner of the mouth grayscale image can be spliced into a single-channel image with a size of 32*96 pixels.
[0063] In step 203, the single-channel image is input into the trained recognition network to obtain the probability that the face to be recognized correctly wears a mask.
[0064] Before training the recognition network, the training sample image can be obtained first, and then the recognition network is trained based on the training sample image, so that the trained recognition network can accurately identify whether the face correctly wears a mask.
[0065] The terminal device can determine the training sample image from the collected M face images, M is an integer greater than 1, and the M face images include face images correctly wearing a mask and face images not correctly wearing a mask. In the process of collecting the M face images, factors such as age, gender, skin color, light, and scene are considered, so that the M face images are as diverse as possible.
[0066] Taking the i-th face image among the M face images mentioned above as an example, the i-th face image is any face image among the M face images mentioned above, where i is an integer greater than zero and less than or equal to M. The face key points of the i-th face image can be obtained through the face key point localization algorithm. The face of the i-th face image is aligned with the face template based on the face key points of the i-th face image to obtain the face-aligned image corresponding to the i-th face image. In the face-aligned image, a nose image of a preset size can be cropped with the tip of the nose as the center, a left corner of the mouth image of a preset size can be cropped with the left corner of the mouth as the center, and a right corner of the mouth image of a preset size can be cropped with the right corner of the mouth as the center. For ease of differentiation, the aforementioned nose image, left corner of the mouth image, and right corner of the mouth image can be referred to as the nose image, left corner of the mouth image, and right corner of the mouth image corresponding to the i-th face image. By performing grayscale processing on the nose image, left corner of the mouth image, and right corner of the mouth image corresponding to the i-th face image, we can obtain the corresponding grayscale nose image, left corner of the mouth image, and right corner of the mouth image. These three grayscale images are then stitched together to form a single-channel image, which is a sample image. To facilitate subsequent training of the recognition network, if the face image corresponding to the sample image is a face image where a mask is correctly worn, the label of the sample image can be set to 1; if the face image corresponding to the sample image is a face image where a mask is not correctly worn, the label of the sample image can be set to 0.
[0067] By processing each of the above M face images as shown in the i-th face image, we can obtain M sample images. These M sample images can be divided into training sample images and test sample images. For example, if we collect 20,000 face images, we can obtain 20,000 sample images after processing. We can use 16,000 sample images as training sample images and 4,000 sample images as test sample images.
[0068] Training sample images are used to train the recognition network. Test sample images are used to determine the probability threshold.
[0069] like Figure 3c The diagram shown is an example of the structure of a recognition network. Figure 3cIt can be seen that the recognition network comprises an input layer, a dimension transformation, three image processing layers, two fully connected layers and a loss function. Each image processing layer comprises, from top to bottom, a convolution layer with a step of 1, batch normalization, scale scaling, an activation layer and a convolution layer with a step of 2. In the input layer, n*32*96*1 indicates that n images are input, the image size is 32*96, and the number of channels of the image is 1. In the dimension transformation, n*32*32*3 indicates that n images are input, the image size is 32*32, and the number of channels of the image is 3. In the convolution layer, 16[3, 3] indicates 16 3*3 convolution kernels, 32[3, 3] indicates 32 3*3 convolution kernels, and 64[3, 3] indicates 64 3*3 convolution kernels. In the fully connected layer, 48 indicates that the output dimension of the fully connected layer is 48, and 2 indicates that the output dimension of the fully connected layer is 2. The dimension transformation of the image is performed to facilitate the image recognition of the recognition network. The convolution layer with a step of 2 connected after the activation layer is helpful for feature extraction.
[0070] As an optional embodiment, when training the recognition network, the learning rate can be set to 0.1, the batch size can be set to 64, the loss function is a cross-entropy loss function, and the optimizer is a Momentum optimizer.
[0071] In step 204, the recognition result of the to-be-recognized face wearing a mask is determined according to the probability of the to-be-recognized face correctly wearing a mask.
[0072] For related descriptions of step 204, please refer to embodiment one, which will not be repeated here.
[0073] In an optional embodiment, the probability threshold can be determined by N test sample images, N being an integer greater than zero.
[0074] Specifically, the N test sample images are input into the trained recognition network to obtain the respective probabilities of correctly wearing a mask of the N test sample images;
[0075] The probability threshold is determined according to the respective probabilities of correctly wearing a mask of the N test sample images.
[0076] The first probability and the second probability can be preset, and the first probability is less than the second probability. Starting from the first probability, a probability is taken as a reference probability every interval of the third probability until the reference probability is the second probability. After determining each reference probability, the probability of each corresponding correct wearing of the mask of the N test sample images is compared with the reference probability to determine that the face in the test sample image with a probability greater than the reference threshold value correctly wears the mask, and the face in the test sample image with a probability less than or equal to the reference threshold value does not correctly wear the mask. According to the above recognition result, the accuracy of correct wearing of the mask and the accuracy of incorrect wearing of the mask are determined. Since the accuracy of correct wearing of the mask and the accuracy of incorrect wearing of the mask are inversely related, in order to improve the recognition accuracy of the probability threshold value, the reference probability corresponding to the closest accuracy of correct wearing of the mask and the accuracy of incorrect wearing of the mask can be determined as the probability threshold value. For example, the first probability is 0.0, the second probability is 0.999, and the third probability is 0.001.
[0077] wherein the accuracy of correct wearing of the mask acc_right = a / S, S represents the number of images of correct wearing of the mask in the N test sample images, and a represents the number of images of correct recognition in the S test sample images of correct wearing of the mask. The accuracy of incorrect wearing of the mask acc_wrong = b / (N-S), N-S represents the number of images of incorrect wearing of the mask in the N test sample images, and b represents the number of images of correct recognition in the N-S test sample images of incorrect wearing of the mask.
[0078] For the recognition result of each test sample image, the recognition result of the test sample image can be compared with the label of the test sample image. If the recognition result of the test sample image matches the label of the test sample image, it is determined that the recognition result of the test sample image is correct. If the determination result of the test sample image does not match the label of the test sample image, it is determined that the recognition result of the test sample image is incorrect. For example, if the recognition result of a test sample image is correct wearing of the mask, and the label of the test sample image is 1, it is determined that the recognition result of the test sample image matches the label of the test sample image, because the label of the test sample image is 1 indicates that the face in the test sample image correctly wears the mask.
[0079] The embodiment of the present application can reduce the complexity of image processing and enhance the robustness of the recognition network without affecting the texture information of the image by splicing the nose image, the left corner image and the right corner image into a single-channel image and taking the single-channel image as the input of the recognition network to identify whether the face to be identified correctly wears the mask.
[0080] Referring to Figure 4Fig. 3 is a structural schematic diagram of a mask wearing recognition device provided in Embodiment Three of the present application. For ease of illustration, only parts related to the present application are shown.
[0081] The mask wearing recognition device comprises:
[0082] The first acquisition module 41 is configured to acquire a nose image, a left corner of mouth image and a right corner of mouth image of a to-be-recognized face. The nose image refers to an image of a region where a nose is located in the to-be-recognized face. The left corner of mouth image refers to an image of a region where a left corner of mouth is located in the to-be-recognized face. The right corner of mouth image refers to an image of a region where a right corner of mouth is located in the to-be-recognized face.
[0083] The image recognition module 42 is configured to input the nose image, the left corner of mouth image and the right corner of mouth image into a trained recognition network to obtain a recognition result of whether the to-be-recognized face wears a mask.
[0084] Optionally, the mask wearing recognition device further comprises:
[0085] The image splicing module is configured to splice the nose image, the left corner of mouth image and the right corner of mouth image into a single-channel image.
[0086] The image recognition module 42 is specifically configured to:
[0087] input the single-channel image into the trained recognition network.
[0088] Optionally, the mask wearing recognition device further comprises:
[0089] The grayscale processing module is configured to perform grayscale processing on the nose image, the left corner of mouth image and the right corner of mouth image respectively to obtain a nose grayscale image, a left corner of mouth grayscale image and a right corner of mouth grayscale image.
[0090] The image splicing module is specifically configured to:
[0091] splice the nose grayscale image, the left corner of mouth grayscale image and the right corner of mouth grayscale image into the single-channel image.
[0092] Optionally, the image recognition module 42 comprises:
[0093] The recognition unit is configured to input the nose image, the left corner of mouth image and the right corner of mouth image into a trained recognition network to obtain a probability that the to-be-recognized face correctly wears a mask.
[0094] The determination unit is configured to determine a recognition result of whether the to-be-recognized face wears a mask according to the probability that the to-be-recognized face correctly wears a mask.
[0095] The mask wearing recognition device further includes:
[0096] The threshold acquisition module is configured to acquire a probability threshold.
[0097] The determination unit is specifically configured to:
[0098] If the probability that the to-be-recognized face correctly wears a mask is greater than the probability threshold, it is determined that the recognition result is that the to-be-recognized face correctly wears a mask.
[0099] If the probability that the to-be-recognized face correctly wears a mask is less than or equal to the probability threshold, it is determined that the recognition result is that the to-be-recognized face does not correctly wear a mask.
[0100] Optionally, the threshold acquisition module is specifically configured to:
[0101] acquire N test sample images, where N is an integer greater than zero;
[0102] input the N test sample images into the trained recognition network to obtain a probability that each of the N test sample images correctly wears a mask;
[0103] determine the probability threshold according to the probability that each of the N test sample images correctly wears a mask.
[0104] Optionally, the mask wearing recognition device further includes:
[0105] a second acquisition module configured to acquire a to-be-recognized image including the to-be-recognized face;
[0106] The first acquisition module 41 is specifically configured to:
[0107] acquire the nose image, the left corner image, and the right corner image from the to-be-recognized image.
[0108] Optionally, the mask wearing recognition device further includes:
[0109] an image alignment module configured to perform face alignment processing on the to-be-recognized image to obtain a face-aligned to-be-recognized image;
[0110] The first acquisition module 41 is specifically configured to:
[0111] acquire the nose image, the left corner image, and the right corner image from the face-aligned to-be-recognized image.
[0112] The mask wearing recognition device provided by the embodiments of the present application can be applied in the foregoing method embodiments one and two, and details are described in the foregoing method embodiments one and two, which will not be described here again.
[0113] Figure 5 is a structural schematic diagram of the terminal device provided in Embodiment Four of the present application. As shown in the figure, the terminal device 5 of this embodiment comprises one or more processors 50 (only one processor is shown in the figure), a memory 51, and a computer program 52 stored in the memory 51 and executable on the processor 50. The processor 50 implements the steps in each of the above mask wearing identification method embodiments when executing the computer program 52. Figure 5
[0114] The terminal device 5 can be a mobile phone, a desktop computer, a notebook computer, a palm computer, a cloud server, and the like. The terminal device can include, but is not limited to, the processor 50 and the memory 51. Those skilled in the art can understand that the terminal device 5 shown in the figure is only an example and does not constitute a limitation on the terminal device 5, which can include more or fewer components than those shown in the figure, or combine certain components, or different components, for example, the terminal device can also include an input / output device, a network access device, a bus, and the like. Figure 5
[0115] The processor 50 can be a central processing unit (CPU), and can also be other general-purpose processors, digital signal processors (DSP), application specific integrated circuits (ASIC), field-programmable gate arrays (FPGA) or other programmable logic devices, discrete gates or transistor logic devices, discrete hardware components, and the like. The general-purpose processor can be a microprocessor or the processor can also be any conventional processor.
[0116] The memory 51 can be an internal storage unit of the terminal device 5, such as a hard disk or a memory of the terminal device 5. The memory 51 can also be an external storage device of the terminal device 5, such as a plug-in hard disk, a smart media card (SMC), a secure digital (SD) card, a flash card, and the like equipped on the terminal device 5. Further, the memory 51 can include both the internal storage unit and the external storage device of the terminal device 5. The memory 51 is used to store the computer program and other programs and data required by the terminal device. The memory 51 can also be used to temporarily store data that has been output or will be output.
[0117] Those skilled in the art can clearly understand that, for the convenience and brevity of description, only the above-mentioned division of each functional unit and module is exemplified, and in actual application, the above-mentioned functions can be completed by different functional units and modules according to needs, that is, the internal structure of the device is divided into different functional units or modules to complete all or part of the functions described above. Each functional unit and module in the embodiment can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The above-mentioned integrated unit can be realized in the form of hardware or software. In addition, the specific names of each functional unit and module are only for easy distinction, and do not limit the protection scope of the present application. The specific working process of the units and modules in the above system can refer to the corresponding process in the foregoing method embodiments, which will not be repeated here.
[0118] In the above embodiments, the description of each embodiment has its own emphasis, and the parts not described or recorded in detail in a certain embodiment can be referred to the related description of other embodiments.
[0119] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be realized in electronic hardware or a combination of computer software and electronic hardware. Whether the functions are executed in hardware or software depends on the specific application and design constraints of the technical solution. A person skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of the present application.
[0120] In the embodiments provided in the present application, it should be understood that the disclosed apparatus / terminal device and method can be implemented in other ways. For example, the above-described apparatus / terminal device embodiments are only schematic, and the division of the modules or units is only a logical function division, and there can be another division in actual implementation. For example, a plurality of units or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the displayed or discussed mutual coupling or direct coupling or communication connection can be indirect coupling or communication connection through some interface, device or unit, and can be electrical, mechanical or other forms.
[0121] The units described as separate components can or can not be physically separated, and the components shown as units can or can not be physical units, that is, they can be located in one place, or can be distributed on multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the embodiment.
[0122] In addition, each of the function units in each of the embodiments of the present application can be integrated in one processing unit, or each unit can be physically present separately, or two or more units can be integrated in one unit. The integrated unit can be realized in the form of hardware or in the form of a software function unit.
[0123] The integrated module / unit, if realized in the form of a software function unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on such understanding, all or part of the processes in the above-mentioned embodiment methods can also be completed by a computer program instructing related hardware, and the computer program can be stored in a computer-readable storage medium. The computer program can implement the steps of each method embodiment when executed by a processor. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or some intermediate forms. The computer-readable medium can include any entity or device capable of carrying the computer program code, recording medium, U disk, mobile hard disk, magnetic disk, optical disk, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the computer-readable medium can include appropriate contents according to the requirements of legislation and patent practice in the jurisdiction, for example, in some jurisdictions, according to legislation and patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0124] The above-mentioned embodiment methods can also be completed by a computer program product, which, when running on a terminal device, causes the terminal device to execute the steps of the above-mentioned method embodiments.
[0125] The above-mentioned embodiments are only used to illustrate the technical solutions of the present application, rather than limit them; although the present application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can modify the technical solutions recorded in the foregoing embodiments, or make equivalent replacements to some technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of the embodiments of the present application, and should be included in the protection scope of the present application.
Claims
1. A method for identifying mask wearing, characterized in that, The mask wearing recognition method includes: Acquire images of the nose, left corner of the mouth, and right corner of the mouth of the face to be identified. The nose image refers to the image of the area where the nose is located in the face to be identified, the left corner of the mouth image refers to the image of the area where the left corner of the mouth is located in the face to be identified, and the right corner of the mouth image refers to the image of the area where the right corner of the mouth is located in the face to be identified. The nose image, the left corner of the mouth image, and the right corner of the mouth image are input into the trained recognition network to obtain the recognition result of the face to be recognized wearing a mask; Before inputting the nose image, the left corner of the mouth image, and the right corner of the mouth image into the trained recognition network, the method further includes: The nose image, the left corner of the mouth image, and the right corner of the mouth image are stitched together to form a single-channel image; The step of inputting the nose image, the left corner of the mouth image, and the right corner of the mouth image into the trained recognition network includes: The single-channel image is input into the trained recognition network.
2. The mask wearing recognition method as described in claim 1, characterized in that, Before stitching the nose image, the left corner of the mouth image, and the right corner of the mouth image into a single-channel image, the process also includes: The nose image, the left corner of the mouth image, and the right corner of the mouth image are respectively processed into grayscale to obtain a grayscale image of the nose, a grayscale image of the left corner of the mouth, and a grayscale image of the right corner of the mouth; The step of stitching the nose image, the left corner of the mouth image, and the right corner of the mouth image into a single-channel image includes: The grayscale image of the nose, the grayscale image of the left corner of the mouth, and the grayscale image of the right corner of the mouth are stitched together to form the single-channel image.
3. The mask wearing recognition method as described in any one of claims 1 to 2, characterized in that, The step of inputting the nose image, the left corner of the mouth image, and the right corner of the mouth image into the trained recognition network to obtain the recognition result of the face to be recognized wearing a mask includes: The nose image, the left corner of the mouth image, and the right corner of the mouth image are input into the trained recognition network to obtain the probability that the face to be recognized is correctly wearing a mask; Based on the probability that the face to be identified is correctly wearing a mask, the identification result of the face to be identified wearing a mask is determined.
4. The mask wearing recognition method as described in claim 3, characterized in that, Before determining the recognition result of the mask-wearing of the face to be recognized based on the probability that the face to be recognized is correctly wearing a mask, the process also includes: Obtain the probability threshold; The step of determining the recognition result of the face being identified wearing a mask based on the probability that the face being identified is correctly wearing a mask includes: If the probability that the face to be identified is correctly wearing a mask is greater than the probability threshold, then the identification result is determined to be that the face to be identified is correctly wearing a mask. If the probability that the face to be identified is wearing a mask correctly is less than or equal to the probability threshold, then the identification result is determined to be that the face to be identified is not wearing a mask correctly.
5. The mask wearing recognition method as described in claim 4, characterized in that, The acquisition probability threshold includes: Obtain N test sample images, where N is an integer greater than zero; The N test sample images are input into the trained recognition network to obtain the probability of correctly wearing a mask for each of the N test sample images; The probability threshold is determined based on the probability of correctly wearing a mask corresponding to each of the N test sample images.
6. The mask wearing recognition method according to any one of claims 1 to 2, characterized in that, Before acquiring the nose image, left corner of the mouth image, and right corner of the mouth image of the face to be identified, the process also includes: Acquire an image of the person to be identified, including the face to be identified; The acquisition of the nose image, left corner of the mouth image, and right corner of the mouth image of the face to be identified includes: The nose image, the left corner of the mouth image, and the right corner of the mouth image are obtained from the image to be identified.
7. The mask wearing recognition method as described in claim 6, characterized in that, Before obtaining the nose image, the left corner of the mouth image, and the right corner of the mouth image from the image to be identified, the method further includes: The image to be identified is subjected to face alignment processing to obtain the face-aligned image to be identified; The step of obtaining the nose image, the left corner of the mouth image, and the right corner of the mouth image from the image to be identified includes: From the face-aligned image to be identified, obtain the nose image, the left corner of the mouth image, and the right corner of the mouth image.
8. A mask wearing recognition device, characterized in that, The mask wearing recognition device includes: The first acquisition module is used to acquire the nose image, the left corner of the mouth image, and the right corner of the mouth image of the face to be identified. The nose image refers to the image of the area where the nose is located in the face to be identified, the left corner of the mouth image refers to the image of the area where the left corner of the mouth is located in the face to be identified, and the right corner of the mouth image refers to the image of the area where the right corner of the mouth is located in the face to be identified. An image recognition module is used to input the nose image, the left corner of the mouth image, and the right corner of the mouth image into a trained recognition network to obtain the recognition result of the face to be recognized wearing a mask; The mask wearing recognition device also includes: An image stitching module is used to stitch the nose image, the left corner of the mouth image, and the right corner of the mouth image into a single-channel image; The image recognition module is specifically used for: The single-channel image is input into the trained recognition network.
9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the mask wearing recognition method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the mask wearing recognition method as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Mask wearing state recognition method and device and computer equipment
CN111444869A
Mask wearing detection method and device, storage medium and electronic equipment
CN111444887A