A yawp detection method, device and electronic equipment

By acquiring key points of the mouth area in the image, determining the aspect ratio of the mouth, and querying a list of ratio thresholds based on the shooting angle, the yawn detection result is determined by combining the mouth aspect ratio and the ratio threshold. This solves the problem of high false positive rate in existing technologies and achieves higher detection accuracy.

CN115457618BActive Publication Date: 2026-05-08BEIJING CO WHEELS TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
BEIJING CO WHEELS TECH CO LTD
Filing Date
2022-03-29
Publication Date
2026-05-08

AI Technical Summary

Technical Problem

In existing yawn detection technologies, the fixed threshold for yawn opening and closing leads to poor detection accuracy and a high false positive rate at certain angles.

Method used

By acquiring key points of the mouth area in the image, determining the aspect ratio of the mouth, and querying a list of ratio thresholds based on the shooting angle, the yawn detection result is determined by combining the mouth aspect ratio and the ratio threshold, thereby reducing the false positive rate.

Benefits of technology

It improves the accuracy of yawn detection and reduces the false positive rate.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115457618B_ABST
    Figure CN115457618B_ABST
Patent Text Reader

Abstract

The application provides a yawning detection method and device and electronic equipment. The method comprises the following steps: determining a to-be-processed image and a mouth region image in the to-be-processed image; detecting the mouth region image to obtain a detection result of the mouth region image; determining a mouth height-width ratio according to at least one mouth key point in the detection result; determining a shooting angle of the to-be-processed image; querying a ratio threshold list according to the shooting angle to obtain a ratio threshold corresponding to the shooting angle; and determining a yawning detection result of the to-be-processed image according to the mouth height-width ratio and the ratio threshold. Thus, the mouth key point of the mouth region image can be obtained, the mouth height-width ratio can be determined, the shooting angle of the to-be-processed image can be determined, the yawning detection result of the to-be-processed image can be determined according to the ratio threshold of the shooting angle and the mouth height-width ratio, the misjudgment rate can be reduced, and the accuracy of the yawning detection can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of deep learning technology, and in particular to a yawn detection method, apparatus, and electronic device. Background Technology

[0002] Currently, the yawn detection solution involves locating the mouth area of ​​a face using face detection and facial landmark detection algorithms, performing landmark analysis on the detection results of the mouth area based on the facial landmark model, determining the mouth opening degree based on the obtained landmarks, and judging whether the mouth is in an open or closed state based on the mouth opening degree and opening degree threshold.

[0003] In the above scheme, the opening and closing degree threshold is fixed, which leads to poor detection accuracy at certain angles. Summary of the Invention

[0004] The present invention aims to solve, to a certain extent, the technical problems in the related technologies.

[0005] Therefore, the first objective of this invention is to propose a yawn detection method. This method obtains key points of the mouth in an image of the mouth region, determines the aspect ratio of the mouth, and determines the shooting angle of the image to be processed. Based on the ratio threshold of the shooting angle and the aspect ratio of the mouth, the yawn detection result of the image to be processed is determined, thereby reducing the false positive rate and improving the accuracy of yawn detection.

[0006] The second objective of this invention is to provide a yawn detection device.

[0007] The third objective of this invention is to provide an electronic device.

[0008] A fourth objective of this invention is to provide a non-transitory computer-readable storage medium.

[0009] The fifth objective of this invention is to provide a computer program product.

[0010] To achieve the above objectives, a first aspect of the present invention provides a yawn detection method, which includes the following steps: determining an image to be processed and a mouth region image in the image to be processed; detecting the mouth region image to obtain a detection result of the mouth region image, wherein the detection result includes: at least one mouth key point in the mouth region image; determining the mouth aspect ratio based on the at least one mouth key point; determining the shooting angle of the image to be processed; querying a ratio threshold list based on the shooting angle to obtain a ratio threshold corresponding to the shooting angle; and determining the yawn detection result of the image to be processed based on the mouth aspect ratio and the ratio threshold.

[0011] According to an embodiment of the present invention, a yawn detection method first determines an image to be processed and an image of the mouth region within the image to be processed; the mouth region image is detected to obtain a detection result, wherein the detection result includes: at least one key point of the mouth in the mouth region image; the aspect ratio of the mouth is determined based on the at least one key point of the mouth; the shooting angle of the image to be processed is determined; a ratio threshold list is consulted based on the shooting angle to obtain a ratio threshold corresponding to the shooting angle; and the yawn detection result of the image to be processed is determined based on the aspect ratio of the mouth and the ratio threshold. Thus, this method reduces the false positive rate and improves the accuracy of yawn detection by obtaining key points of the mouth in the mouth region image, determining the aspect ratio of the mouth, determining the shooting angle of the image to be processed, and determining the yawn detection result of the image to be processed based on the ratio threshold of the shooting angle and the aspect ratio of the mouth.

[0012] In addition, the yawn detection method proposed in the first aspect of the present invention may also have the following additional technical features:

[0013] According to one embodiment of the present invention, determining the image to be processed and the mouth region image in the image to be processed includes: determining the image to be processed; performing face detection on the image to be processed to obtain a face region image in the image to be processed; performing facial landmark detection on the face region image to obtain at least one facial landmark in the face region image; and determining the mouth region image in the image to be processed based on the mouth landmark among the at least one facial landmark and the image to be processed.

[0014] According to an embodiment of the present invention, determining the mouth region image in the image to be processed based on the mouth key point among the at least one facial key point and the image to be processed includes: determining the position information of the mouth region in the image to be processed based on the mouth key point among the at least one facial key point; and cropping the image to be processed according to the position information to obtain the mouth region image.

[0015] According to one embodiment of the present invention, determining the mouth aspect ratio based on the at least one mouth key point includes: determining the mouth width based on the corner of the mouth key point among the at least one mouth key point; determining the mouth height based on the lip tip key point among the at least one mouth key point; and determining the ratio of the mouth width to the mouth height as the mouth aspect ratio.

[0016] According to one embodiment of the present invention, determining the shooting angle of the image to be processed includes: inputting the at least one mouth key point into a preset shooting angle prediction model to obtain the shooting angle output by the shooting angle prediction model.

[0017] According to one embodiment of the present invention, before querying the ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle, the method further includes: obtaining a greater than preset number of sample mouth region images in the open-mouth state, the shooting angle corresponding to the sample mouth region images, and the sample mouth aspect ratio corresponding to the sample mouth region images; for each shooting angle, obtaining at least one sample mouth region image corresponding to the shooting angle, and the sample mouth aspect ratio corresponding to the sample mouth region image; summing and averaging at least one of the sample mouth aspect ratios to obtain a processed mouth aspect ratio; and determining the ratio threshold corresponding to the shooting angle based on the processed mouth aspect ratio.

[0018] According to one embodiment of the present invention, determining the yawn detection result of the image to be processed based on the mouth aspect ratio and the ratio threshold includes: when the mouth aspect ratio is greater than or equal to the ratio threshold, determining that the yawn detection result indicates that a face in the image to be processed is yawning; and when the mouth aspect ratio is less than the ratio threshold, determining that the yawn detection result indicates that a face in the image to be processed is not yawning.

[0019] According to an embodiment of the present invention, the detection result of the mouth region image further includes: mouth state. The step of determining the yawn detection result of the image to be processed based on the mouth aspect ratio and the ratio threshold includes: when the mouth aspect ratio is greater than or equal to the ratio threshold and the mouth state is an open mouth state, determining that the yawn detection result indicates that the face in the image to be processed exhibits yawning behavior; when the mouth aspect ratio is less than the ratio threshold, or when the mouth state is a closed mouth state, determining that the yawn detection result indicates that the face in the image to be processed does not exhibit yawning behavior.

[0020] To achieve the above objectives, a second aspect of the present invention provides a yawn detection device, comprising: a first determining module for determining an image to be processed and a mouth region image in the image to be processed; a first acquiring module for detecting the mouth region image and acquiring a detection result of the mouth region image, wherein the detection result includes at least one mouth key point in the mouth region image; a second determining module for determining a mouth aspect ratio based on the at least one mouth key point; a third determining module for determining the shooting angle of the image to be processed; a second acquiring module for querying a ratio threshold list based on the shooting angle to acquire a ratio threshold corresponding to the shooting angle; and a fourth determining module for determining a yawn detection result of the image to be processed based on the mouth aspect ratio and the ratio threshold.

[0021] According to an embodiment of the present invention, a yawn detection device determines an image to be processed and a mouth region image within the image to be processed through a first determining module; a first acquiring module detects the mouth region image and acquires a detection result of the mouth region image, wherein the detection result includes at least one mouth key point in the mouth region image; a second determining module is used to determine the aspect ratio of the mouth based on the at least one mouth key point; a third determining module is used to determine the shooting angle of the image to be processed; the second acquiring module is used to query a ratio threshold list based on the shooting angle to acquire a ratio threshold corresponding to the shooting angle; and a fourth determining module is used to determine the yawn detection result of the image to be processed based on the mouth aspect ratio and the ratio threshold. Thus, the device determines the mouth aspect ratio by acquiring the mouth key points of the mouth region image and determining the shooting angle of the image to be processed; and determines the yawn detection result of the image to be processed based on the ratio threshold of the shooting angle and the mouth aspect ratio, thereby reducing the false positive rate and improving the accuracy of yawn detection.

[0022] In addition, the yawn detection device proposed in the second aspect embodiment of the present invention may also have the following additional technical features:

[0023] According to an embodiment of the present invention, the first determining module includes: a first determining unit, a first acquiring unit, a second acquiring unit, and a second determining unit; the first determining unit is used to determine an image to be processed; the first acquiring unit is used to perform face detection on the image to be processed to obtain a face region image in the image to be processed; the second acquiring unit is used to perform facial key point detection on the face region image to obtain at least one facial key point in the face region image; the second determining unit is used to determine a mouth region image in the image to be processed based on a mouth key point among the at least one facial key point and the image to be processed.

[0024] According to one embodiment of the present invention, the second determining unit is specifically configured to: determine the position information of the mouth region in the image to be processed based on the mouth key point among the at least one facial key point; and perform cropping processing on the image to be processed according to the position information to obtain the mouth region image.

[0025] According to one embodiment of the present invention, the second determining module is specifically used to: determine the mouth width based on the corner of the mouth key point among the at least one mouth key point; determine the mouth height based on the lip tip key point among the at least one mouth key point; and determine the ratio of the mouth width to the mouth height as the mouth height-to-width ratio.

[0026] According to one embodiment of the present invention, the third determining module is specifically used to input the at least one key point of the mouth into a preset shooting angle prediction model to obtain the shooting angle output by the shooting angle prediction model.

[0027] According to an embodiment of the present invention, the method further includes: a third acquisition module, a fourth acquisition module, a processing module, and a fifth determination module; the third acquisition module is used to acquire more than a preset number of sample mouth region images in the open-mouth state, the shooting angle corresponding to the sample mouth region images, and the sample mouth aspect ratio corresponding to the sample mouth region images; the fourth acquisition module is used to acquire at least one sample mouth region image corresponding to each shooting angle, and the sample mouth aspect ratio corresponding to the sample mouth region image; the processing module is used to sum and average at least one of the sample mouth aspect ratios to obtain a processed mouth aspect ratio; the fifth determination module is used to determine a ratio threshold corresponding to the shooting angle based on the processed mouth aspect ratio.

[0028] According to an embodiment of the present invention, the fourth determining module is specifically used to determine that the yawning detection result is that the face in the image to be processed has yawning behavior when the aspect ratio of the mouth is greater than or equal to the ratio threshold; and to determine that the yawning detection result is that the face in the image to be processed does not have yawning behavior when the aspect ratio of the mouth is less than the ratio threshold.

[0029] According to an embodiment of the present invention, the detection result of the mouth region image further includes: mouth state. The fourth determining module is specifically used to determine that the yawning detection result is that the face in the image to be processed has yawning behavior when the mouth aspect ratio is greater than or equal to the ratio threshold and the mouth state is an open mouth state; and to determine that the yawning detection result is that the face in the image to be processed does not have yawning behavior when the mouth aspect ratio is less than the ratio threshold, or when the mouth state is an open mouth state.

[0030] To achieve the above objectives, a third aspect of the present invention provides an electronic device comprising: a processor and a memory; wherein the processor runs a program corresponding to the executable program code stored in the memory, for implementing the yawn detection method of the first aspect embodiment.

[0031] To achieve the above objectives, a fourth aspect of the present invention provides a non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that the computer program, when executed by a processor, implements the yawn detection method of the first aspect embodiment.

[0032] To achieve the above objectives, a fifth aspect of the present invention provides a computer program product that, when an instruction processor in the computer program product is executed, performs the yawn detection method of the first aspect of the present invention.

[0033] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0034] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein:

[0035] Figure 1 This is a flowchart of a yawn detection method according to an embodiment of the present invention;

[0036] Figure 2 This is a flowchart of a yawn detection method according to an embodiment of the present invention;

[0037] Figure 3 This is a flowchart of a yawn detection method according to another embodiment of the present invention;

[0038] Figure 4 This is a schematic diagram of a yawn detection device according to an embodiment of the present invention;

[0039] Figure 5 This is a structural block diagram of an electronic device according to an embodiment of the present invention. Detailed Implementation

[0040] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0041] The following description, with reference to the accompanying drawings, describes an embodiment of the yawn detection method, apparatus, and electronic device of the present invention.

[0042] Figure 1 This is a flowchart of a yawn detection method according to an embodiment of the present invention.

[0043] It should be noted that the executing entity in this embodiment of the invention is a yawn detection device, which can be configured in an electronic device so that the electronic device can perform the function of yawn detection.

[0044] Among them, electronic devices can be any device with computing capabilities, such as personal computers (PCs), mobile terminals, servers, etc. Mobile terminals can be hardware devices with various operating systems, touch screens and / or displays, such as in-vehicle devices, mobile phones, tablets, personal digital assistants, wearable devices, etc.

[0045] like Figure 1 As shown, the yawn detection method of this invention includes the following steps:

[0046] S101, determine the image to be processed, and the mouth region image in the image to be processed.

[0047] For example, the image to be processed can be acquired through a camera. The mouth region image can be obtained by detecting facial landmarks in the face region image of the image to be processed. Based on the mouth landmarks in the face landmarks, the position information of the mouth region can be determined. The image to be processed can then be cropped according to the position information to obtain the mouth region image, thereby narrowing the scope and facilitating the relocation of mouth landmarks, thus improving the accuracy of yawn detection.

[0048] S102, Detect the mouth region image and obtain the detection result of the mouth region image, wherein the detection result includes: at least one mouth key point in the mouth region image.

[0049] For example, a multi-task convolutional neural network model can be used to detect mouth region images and obtain multiple detection results output by the model, namely the coordinate information of at least one key point of the mouth and the state of the mouth.

[0050] Understandably, a multi-task convolutional neural network model includes a proposal network (P-Net), a refine network (R-Net), and an output network (O-Net). Before using these networks, the mouth region image is preprocessed by scaling the images in the mouth region image to different sizes, forming an "image pyramid." Calculations are performed on images of each size to detect the mouth region image at different dimensions. Then, mouth region images of different sizes are input into the proposal network to obtain candidate mouth region images output by the proposal network. The candidate mouth region images are then input into the refine network to obtain accurate candidate mouth region images output by the refine network. Finally, the accurate candidate mouth region images are input into the output network to obtain at least one mouth keypoint output by the output network.

[0051] Among them, at least one key point of the mouth may include a key point of the corner of the mouth and a key point of the tip of the lips.

[0052] S103, determine the mouth aspect ratio based on at least one key mouth point.

[0053] In this step, the mouth width is determined based on the corner of the mouth key point among at least one mouth key point; the mouth height is determined based on the lip tip key point among at least one mouth key point; and the ratio of the mouth width to the mouth height is determined as the mouth height-width ratio.

[0054] The key points of the corners of the mouth can include the key points of the left and right corners of the mouth. The width of the mouth can be calculated based on the coordinates of the key points of the left and right corners of the mouth.

[0055] The key points of the lip tip can include the key points of the upper lip tip and the key points of the lower lip tip. The height of the mouth is calculated based on the coordinates of the key points of the upper lip tip and the key points of the lower lip tip.

[0056] Step 104: Determine the shooting angle of the image to be processed.

[0057] In this step, the yawn detection device can input at least one key point of the mouth into a preset shooting angle prediction model to obtain the shooting angle output by the shooting angle prediction model.

[0058] The shooting angle prediction model can be a standard 3D face model, which calculates the shooting angle based on at least one key point of the mouth.

[0059] In this step, the shooting angle is obtained based on the shooting angle prediction model. Different shooting angles correspond to different ratio thresholds. The height-to-width ratio of the mouth is determined by using at least one key point of the mouth, namely the corner key point and the tip key point. The ratio threshold corresponding to the shooting angle is selected from multiple ratio thresholds. The ratio threshold is compared with the height-to-width ratio of the mouth to determine the yawn detection result, thereby reducing the misjudgment rate of the mouth region state.

[0060] S105, query the ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle.

[0061] In some embodiments, prior to this step, the yawn detection device may obtain a list of ratio thresholds from other devices. The other devices obtain the list by statistically analyzing the aspect ratio of the user's mouth when yawning, summing and averaging multiple mouth aspect ratios to form the ratio threshold list.

[0062] In some embodiments, prior to this step, the yawn detection device may first acquire more than a preset number of sample mouth region images in the open-mouth state, the shooting angle corresponding to the sample mouth region images, and the sample mouth aspect ratio corresponding to the sample mouth region images; for each shooting angle, acquire at least one sample mouth region image corresponding to the shooting angle, and the sample mouth aspect ratio corresponding to the sample mouth region image; sum and average the at least one sample mouth aspect ratio to obtain the processed mouth aspect ratio; and determine the ratio threshold corresponding to the shooting angle based on the processed mouth aspect ratio.

[0063] In one example, the processed aspect ratio of the mouth can be directly used as the ratio threshold corresponding to the shooting angle.

[0064] In another example, a specified value can be added to the processed mouth aspect ratio to serve as the ratio threshold corresponding to the shooting angle. Directly using the processed mouth aspect ratio as the ratio threshold may result in some mouth aspect ratios slightly below the ratio threshold not being considered as yawning. Adding a specified value to the processed mouth aspect ratio as the ratio threshold corresponding to the shooting angle can further improve the accuracy of yawn detection.

[0065] S106. Determine the yawn detection result of the image to be processed based on the aspect ratio of the mouth and the ratio threshold.

[0066] In this step, the yawn detection device can determine that the yawn detection result is that the face in the image to be processed has yawning behavior when the aspect ratio of the mouth is greater than or equal to the ratio threshold; and determine that the yawn detection result is that the face in the image to be processed does not have yawning behavior when the aspect ratio of the mouth is less than the ratio threshold.

[0067] Therefore, the yawn detection method of this invention obtains key points of the mouth in the mouth region image, determines the aspect ratio of the mouth, and determines the shooting angle of the image to be processed. Based on the ratio threshold of the shooting angle and the aspect ratio of the mouth, the yawn detection result of the image to be processed is determined, thereby reducing the false positive rate and improving the accuracy of yawn detection.

[0068] Figure 2 This is a flowchart of a yawn detection method according to an embodiment of the present invention.

[0069] like Figure 2 As shown, the yawn detection method of this invention includes the following steps:

[0070] S201, Determine the image to be processed.

[0071] The image to be processed can be acquired from a camera. The image to be processed can include images with or without faces.

[0072] S202, perform face detection on the image to be processed to obtain the face region image in the image to be processed.

[0073] In this step, the image to be processed is input into the face detection model, the position and size of the face are marked, and the face region image output by the face detection model is obtained.

[0074] If face detection is performed on the image to be processed and no face region image is obtained, then the subsequent processing of the image to be processed will be stopped.

[0075] S203, Perform facial landmark detection on the face region image to obtain at least one facial landmark in the face region image.

[0076] In this step, facial key points are detected based on the face region image, such as eyes, nose tip, corner of mouth, eyebrows, and contour points of various facial components, to obtain a set of facial feature points, thereby determining at least one facial key point. This can improve the accuracy of the key points and thus improve the accuracy of the detection results.

[0077] S204, Based on the mouth key point in at least one facial key point and the image to be processed, determine the mouth region image in the image to be processed.

[0078] In this step, the positional information of the mouth region in the image to be processed is determined based on the mouth key point among at least one facial key point; the image to be processed is then cropped according to the positional information to obtain the mouth region image.

[0079] In this step, the location information of the mouth region in the image to be processed can be determined based on the key points of the mouth, thereby obtaining a more accurate mouth region image, thus reducing the false judgment rate and improving the accuracy of the detection results.

[0080] S205, Detect the mouth region image and obtain the detection result of the mouth region image, wherein the detection result includes: at least one mouth key point in the mouth region image.

[0081] S206, Determine the mouth aspect ratio based on at least one mouth key point.

[0082] S207, Determine the shooting angle of the image to be processed.

[0083] S208, Query the ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle.

[0084] S209. Determine the yawn detection result of the image to be processed based on the aspect ratio of the mouth and the ratio threshold.

[0085] It should be noted that the execution process of steps S205-S209 can be found in the content of steps S102-S106 above, and will not be repeated here.

[0086] Therefore, by acquiring key points of the mouth region image, the aspect ratio of the mouth can be determined; and the shooting angle of the image to be processed can be determined; based on the ratio threshold of the shooting angle and the aspect ratio of the mouth, the yawn detection result of the image to be processed can be determined, thereby reducing the false positive rate and improving the accuracy of yawn detection.

[0087] Figure 3 This is a flowchart of another yawn detection method according to the present invention.

[0088] like Figure 3 As shown, the yawning detection method of this invention includes:

[0089] S301, Determine the image to be processed, and the mouth region image within the image to be processed.

[0090] The image to be processed can be an image that includes a face or an image that does not. Images that do not include faces include, for example, landscape images or road condition images.

[0091] S302, Detect the mouth region image and obtain the detection result of the mouth region image, wherein the detection result includes: at least one mouth key point in the mouth region image.

[0092] In this step, the mouth region image can be input into a convolutional neural network to obtain the detection results output by the convolutional neural network. The detection results include at least one mouth key point and mouth state. The mouth key point can be used to calculate the mouth aspect ratio. The mouth aspect ratio is compared with the ratio threshold and combined with the mouth state to detect yawning behavior, thereby improving the accuracy of the detection results.

[0093] S303, determine the mouth aspect ratio based on at least one mouth key point.

[0094] S304, determine the shooting angle of the image to be processed.

[0095] S305: Query the ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle.

[0096] S306, when the aspect ratio of the mouth is greater than or equal to the ratio threshold and the mouth is in an open state, the yawn detection result is determined to be that the face in the image to be processed has yawning behavior.

[0097] In this step, the yawning behavior is judged based on the mouth aspect ratio and mouth state. That is, when the mouth aspect ratio is greater than or equal to the ratio threshold and the mouth state is open, it is determined that the face is yawning, thereby reducing the possibility of false judgment and improving the accuracy of yawn detection.

[0098] S307, when the aspect ratio of the mouth is less than the ratio threshold, or when the mouth is in a non-open state, the yawn detection result is determined to be that there is no yawning behavior in the face of the image to be processed.

[0099] The "non-open mouth" state can include both a normal mouth state and other states. For example, a normal mouth state can be either closed or slightly open. Other states can be caused by abnormalities in the mouth, such as obstruction.

[0100] It should be noted that the contents of steps S301-S305 are the same as those of steps S101-S105 above, and will not be repeated here.

[0101] In summary, the yawn detection method according to embodiments of the present invention first determines the image to be processed and the mouth region image within the image to be processed; it then detects the mouth region image to obtain the detection result, wherein the detection result includes: at least one mouth key point in the mouth region image; based on the at least one mouth key point, it determines the aspect ratio of the mouth; and it determines the shooting angle of the image to be processed; it queries a ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle; when the mouth aspect ratio is greater than or equal to the ratio threshold and the mouth is in an open-mouth state, it determines that the yawn detection result indicates that the face in the image to be processed exhibits yawning behavior; when the mouth aspect ratio is less than the ratio threshold, or when the mouth is in a closed-mouth state, it determines that the face in the image to be processed does not exhibit yawning behavior, thereby reducing the false positive rate and improving the accuracy of yawn detection.

[0102] Figure 4 This is a schematic diagram of a yawn detection device according to an embodiment of the present invention.

[0103] like Figure 4 As shown, the yawn detection device 400 of this embodiment includes: a first determining module 401, a first acquiring module 402, a second determining module 403, a third determining module 404, a second acquiring module 405, and a fourth determining module 406.

[0104] The first determining module 401 is used to determine the image to be processed, and the mouth region image in the image to be processed;

[0105] The first acquisition module 402 is used to detect the mouth region image and acquire the detection result of the mouth region image, wherein the detection result includes at least one mouth key point in the mouth region image;

[0106] The second determining module 403 is used to determine the aspect ratio of the mouth based on the at least one mouth key point;

[0107] The third determining module 404 is used to determine the shooting angle of the image to be processed;

[0108] The second acquisition module 405 is used to query the ratio threshold list according to the shooting angle to obtain the ratio threshold corresponding to the shooting angle.

[0109] The fourth determining module 406 is used to determine the yawn detection result of the image to be processed based on the aspect ratio of the mouth and the ratio threshold.

[0110] According to an embodiment of the present invention, the first determining module 401 includes: a first determining unit, a first acquiring unit, a second acquiring unit, and a second determining unit; the first determining unit is used to determine an image to be processed; the first acquiring unit is used to perform face detection on the image to be processed to obtain a face region image in the image to be processed; the second acquiring unit is used to perform facial key point detection on the face region image to obtain at least one facial key point in the face region image; the second determining unit is used to determine a mouth region image in the image to be processed based on a mouth key point among the at least one facial key point and the image to be processed.

[0111] According to one embodiment of the present invention, the second determining unit is specifically configured to: determine the position information of the mouth region in the image to be processed based on the mouth key point among the at least one facial key point; and perform cropping processing on the image to be processed according to the position information to obtain the mouth region image.

[0112] According to one embodiment of the present invention, the second determining module 403 is specifically used to: determine the mouth width based on the corner of the mouth key point among the at least one mouth key point; determine the mouth height based on the lip tip key point among the at least one mouth key point; and determine the ratio of the mouth width to the mouth height as the mouth height-to-width ratio.

[0113] According to one embodiment of the present invention, the third determining module 404 is specifically used to input the at least one mouth key point into a preset shooting angle prediction model to obtain the shooting angle output by the shooting angle prediction model.

[0114] According to an embodiment of the present invention, the yawn detection device further includes: a third acquisition module, a fourth acquisition module, a processing module, and a fifth determination module; the third acquisition module is used to acquire more than a preset number of sample mouth region images in the open-mouth state, the shooting angle corresponding to the sample mouth region images, and the sample mouth aspect ratio corresponding to the sample mouth region images; the fourth acquisition module is used to acquire at least one sample mouth region image corresponding to each shooting angle, and the sample mouth aspect ratio corresponding to the sample mouth region image; the processing module is used to sum and average at least one of the sample mouth aspect ratios to obtain a processed mouth aspect ratio; the fifth determination module is used to determine a ratio threshold corresponding to the shooting angle based on the processed mouth aspect ratio.

[0115] According to an embodiment of the present invention, the fourth determining module 406 is specifically configured to: determine that the yawning detection result is that the face in the image to be processed has yawning behavior when the aspect ratio of the mouth is greater than or equal to the ratio threshold; and determine that the yawning detection result is that the face in the image to be processed does not have yawning behavior when the aspect ratio of the mouth is less than the ratio threshold.

[0116] According to an embodiment of the present invention, the detection result of the mouth region image further includes: mouth state. The fourth determining module 406 is specifically configured to: determine that the yawning detection result is that the face in the image to be processed has yawning behavior when the mouth aspect ratio is greater than or equal to the ratio threshold and the mouth state is an open mouth state; and determine that the yawning detection result is that the face in the image to be processed does not have yawning behavior when the mouth aspect ratio is less than the ratio threshold, or when the mouth state is an open mouth state.

[0117] It should be noted that for details not disclosed in the yawn detection device of this invention, please refer to the details disclosed in the yawn detection method of this invention, which will not be repeated here.

[0118] According to an embodiment of the present invention, a yawn detection device determines an image to be processed and a mouth region image within the image to be processed through a first determining module; a first acquiring module detects the mouth region image and acquires a detection result of the mouth region image, wherein the detection result includes at least one mouth key point in the mouth region image; a second determining module determines the aspect ratio of the mouth based on the at least one mouth key point; a third determining module determines the shooting angle of the image to be processed; the second acquiring module queries a ratio threshold list based on the shooting angle to acquire a ratio threshold corresponding to the shooting angle; and a fourth determining module determines the yawn detection result of the image to be processed based on the mouth aspect ratio and the ratio threshold. Thus, the device determines the mouth aspect ratio by acquiring the mouth key points of the mouth region image and determining the shooting angle of the image to be processed; and determines the yawn detection result of the image to be processed based on the ratio threshold of the shooting angle and the mouth aspect ratio, thereby reducing the false positive rate and improving the accuracy of yawn detection.

[0119] Based on the above embodiments, the present invention also proposes an electronic device.

[0120] The electronic device of this invention includes a processor and a memory; wherein the processor runs a program corresponding to the executable program code by reading executable program code stored in the memory, so as to implement the above-mentioned yawn detection method.

[0121] Based on the above embodiments, the present invention also proposes a non-transitory computer-readable storage medium.

[0122] The non-transitory computer-readable storage medium of this invention stores a computer program that, when executed by a processor, implements the yawn detection method described above.

[0123] Based on the above embodiments, the present invention also proposes a computer program product.

[0124] The computer program product of this invention executes the yawn detection method described above when the instructions in the computer program product are executed by the processor.

[0125] Figure 5 This is a structural block diagram of an electronic device provided in an embodiment of the present invention. Figure 5 The electronic device shown is merely an example and should not be construed as limiting the functionality and scope of the embodiments of the present invention.

[0126] like Figure 5As shown, the electronic device includes a processor 11, which can perform various appropriate actions and processes according to a program stored in read-only memory (ROM) 12 or a program loaded from memory 16 into random access memory (RAM) 13. The RAM 13 also stores various programs and data required for the operation of the electronic device 10. The processor 11, ROM 12, and RAM 13 are interconnected via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0127] The following components are connected to I / O interface 15: memory 16 including hard disks, etc.; and communication section 17 including network interface cards such as LAN (Local Area Network) cards, modems, etc., which performs communication processing via a network such as the Internet; and driver 18 is also connected to I / O interface 15 as needed.

[0128] In particular, according to embodiments of the present invention, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of the present invention include a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via the communication section 17. When the computer program is executed by the processor 11, it performs the functions defined in the methods of the present invention.

[0129] In an exemplary embodiment, a storage medium including instructions is also provided, such as a memory 16 including instructions, which can be executed by a processor 11 of an electronic device 10 to perform the above-described method. Optionally, the storage medium may be a non-transitory computer-readable storage medium, such as a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0130] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0131] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0132] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as should be understood by those skilled in the art to which embodiments of the invention pertain.

[0133] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0134] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0135] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0136] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a single module, or each unit can exist physically separately, or two or more units can be integrated into a single module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0137] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for detecting yawning, characterized in that, Includes the following steps: Identify the image to be processed, and the mouth region image within the image to be processed; The mouth region image is detected to obtain the detection result of the mouth region image, wherein the detection result includes at least one mouth key point in the mouth region image; Determine the mouth aspect ratio based on at least one key mouth point; Determine the shooting angle of the image to be processed; Query the ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle; The yawn detection result of the image to be processed is determined based on the aspect ratio of the mouth and the aspect ratio threshold. Before querying the ratio threshold list based on the shooting angle to obtain the ratio threshold corresponding to the shooting angle, the method further includes: Acquire more than a preset number of sample mouth region images in the open mouth state, the shooting angle corresponding to the sample mouth region images, and the aspect ratio of the sample mouth corresponding to the sample mouth region images; For each shooting angle, acquire at least one sample mouth region image corresponding to the shooting angle, and the sample mouth aspect ratio value corresponding to the sample mouth region image; The mouth aspect ratio values ​​of at least one of the samples are summed and averaged to obtain the processed mouth aspect ratio value. Based on the processed aspect ratio of the mouth, the ratio threshold corresponding to the shooting angle is determined.

2. The method according to claim 1, characterized in that, The determination of the image to be processed, and the mouth region image within the image to be processed, includes: Identify the image to be processed; Perform face detection on the image to be processed to obtain face region images in the image to be processed; Facial landmark detection is performed on the face region image to obtain at least one facial landmark in the face region image; Based on the mouth key point in the at least one facial key point and the image to be processed, determine the mouth region image in the image to be processed.

3. The method according to claim 2, characterized in that, The step of determining the mouth region image in the image to be processed based on the mouth key points in the at least one facial key point and the image to be processed includes: Based on the mouth key point among the at least one facial key point, determine the position information of the mouth region in the image to be processed; The image to be processed is cropped according to the location information to obtain the mouth region image.

4. The method according to claim 1, characterized in that, Determining the mouth aspect ratio based on the at least one key mouth point includes: The mouth width is determined based on the corner of the mouth key point among the at least one mouth key point; The mouth height is determined based on the lip tip key point among the at least one mouth key point; The ratio of the mouth width to the mouth height is defined as the mouth height-to-width ratio.

5. The method according to claim 1, characterized in that, Determining the shooting angle of the image to be processed includes: The at least one key point of the mouth is input into a preset shooting angle prediction model to obtain the shooting angle output by the shooting angle prediction model.

6. The method according to claim 1, characterized in that, The step of determining the yawn detection result of the image to be processed based on the aspect ratio of the mouth and the threshold value of the ratio includes: When the aspect ratio of the mouth is greater than or equal to the ratio threshold, the yawn detection result is determined to be that the face in the image to be processed is yawning. When the aspect ratio of the mouth is less than the ratio threshold, the yawn detection result is determined to be that there is no yawning behavior in the face of the image to be processed.

7. The method according to claim 1, characterized in that, The detection result of the mouth region image also includes: mouth state. The step of determining the yawn detection result of the image to be processed based on the mouth aspect ratio and a threshold value includes: When the aspect ratio of the mouth is greater than or equal to the ratio threshold, and the mouth is in an open mouth state, the yawn detection result is determined to be that the face in the image to be processed is yawning. When the aspect ratio of the mouth is less than the ratio threshold, or when the mouth is in a non-open mouth state, the yawn detection result is determined to be that there is no yawning behavior in the face of the image to be processed.

8. A yawn detection device, characterized in that, include: The first determining module is used to determine the image to be processed, and the mouth region image in the image to be processed; The first acquisition module is used to detect the mouth region image and acquire the detection result of the mouth region image, wherein the detection result includes at least one mouth key point in the mouth region image; The second determining module is used to determine the mouth aspect ratio based on the at least one mouth key point; The third determining module is used to determine the shooting angle of the image to be processed; The second acquisition module is used to query the ratio threshold list according to the shooting angle to obtain the ratio threshold corresponding to the shooting angle. The fourth determining module is used to determine the yawn detection result of the image to be processed based on the aspect ratio of the mouth and the ratio threshold. The device further includes: a third acquisition module, a fourth acquisition module, a processing module, and a fifth determination module; The third acquisition module is used to acquire more than a preset number of sample mouth region images in the open mouth state, the shooting angle corresponding to the sample mouth region image, and the sample mouth aspect ratio corresponding to the sample mouth region image. The fourth acquisition module is used to acquire at least one sample mouth region image corresponding to each shooting angle, and the sample mouth aspect ratio value corresponding to the sample mouth region image for each shooting angle. The processing module is used to sum and average the height-to-width ratio values ​​of at least one of the sample mouths to obtain the processed mouth height-to-width ratio value. The fifth determining module is used to determine the ratio threshold corresponding to the shooting angle based on the processed mouth aspect ratio.

9. The apparatus according to claim 8, characterized in that, The first determining module includes: a first determining unit, a first acquiring unit, a second acquiring unit, and a second determining unit; The first determining unit is used to determine the image to be processed; The first acquisition unit is used to perform face detection on the image to be processed in order to obtain a face region image in the image to be processed. The second acquisition unit is used to perform facial key point detection on the face region image to obtain at least one facial key point in the face region image; The second determining unit is used to determine the mouth region image in the image to be processed based on the mouth key point in the at least one facial key point and the image to be processed.

10. The apparatus according to claim 8, characterized in that, The second determining module is specifically used for, The mouth width is determined based on the corner of the mouth key point among the at least one mouth key point; The mouth height is determined based on the lip tip key point among the at least one mouth key point; The ratio of the mouth width to the mouth height is defined as the mouth height-to-width ratio.

11. The apparatus according to claim 8, characterized in that, The fourth determining module is specifically used for, When the aspect ratio of the mouth is greater than or equal to the ratio threshold, the yawn detection result is determined to be that the face in the image to be processed is yawning. When the aspect ratio of the mouth is less than the ratio threshold, the yawn detection result is determined to be that there is no yawning behavior in the face of the image to be processed.

12. The apparatus according to claim 8, characterized in that, The detection results of the mouth region image also include: mouth state. The fourth determining module is specifically used for: When the aspect ratio of the mouth is greater than or equal to the ratio threshold, and the mouth is in an open mouth state, the yawn detection result is determined to be that the face in the image to be processed is yawning. When the aspect ratio of the mouth is less than the ratio threshold, or when the mouth is in a non-open mouth state, the yawn detection result is determined to be that there is no yawning behavior in the face of the image to be processed.

13. An electronic device, characterized in that, include: Processor and memory; The processor reads executable program code stored in the memory to run a program corresponding to the executable program code, so as to implement the yawn detection method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Fatigue detection method, device and system and computer readable storage medium

    CN109447025A

  • Fatigue driving detection method based on train cab scene

    CN112016429A