Fatigue state detection method, and device
By acquiring facial key points, correcting and temporally stitching facial images, and utilizing feature extraction and detection models, the problem of inaccuracy in existing fatigue detection algorithms is solved, achieving higher fatigue detection accuracy.
Patent Information
- Application Number
- PCT/CN2024/133439
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-04-23
- Filing Date
- 2024-11-21
- Publication Date
- 2025-10-30
AI Technical Summary
Existing fatigue detection algorithms struggle to accurately detect driver fatigue and are significantly affected by the stability of detection rules.
By acquiring the facial key points of the user to be detected, the pre-trained feature extraction model is used to extract features from the facial image sequence. After fusing spatiotemporal information, the feature is input into the pre-trained detection model for fatigue state detection, including facial key point correction and temporal stitching.
It improves the accuracy of fatigue detection by giving more consideration to the impact of facial expression features on fatigue status, thus enhancing the accuracy of the detection results.
Smart Images

Figure CN2024133439_30102025_PF_FP_ABST
Abstract
Description
Methods and apparatus for detecting fatigue state
[0001] This application claims priority to Chinese Patent Application No. 202410489671.2, filed on April 23, 2024, entitled “Method and Apparatus for Detecting Fatigue Condition”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of fatigue detection technology, and in particular to a method and apparatus for detecting fatigue state. Background Technology
[0003] With the continuous increase in the number of modern cars, not only has the transportation industry developed, but traffic accidents have also increased. Fatigue driving is a major cause of traffic accidents, making the effective supervision and prevention of driver fatigue of extremely important in practice.
[0004] In related technologies, the state information of a driver's eyes, mouth, and other facial features can be used to determine whether the driver is fatigued. This process relies on detection algorithms to define rules. If the rules change, the corresponding detection results will also be affected, making it difficult to accurately detect fatigue.
[0005] Application content
[0006] In view of this, this application provides a method and apparatus for detecting fatigue state, the main purpose of which is to solve the problem that fatigue detection algorithms in the prior art are difficult to accurately detect fatigue state.
[0007] According to a first aspect of this application, a method for detecting fatigue state is provided, the method comprising:
[0008] Obtain key facial features of the user to be detected;
[0009] Within a set time window, face images containing facial landmarks are sequentially stitched together to obtain a face image sequence.
[0010] The face image sequence is subjected to feature extraction using a pre-trained feature extraction model to obtain facial expression features that incorporate spatiotemporal information;
[0011] The facial expression features fused with spatiotemporal information are input into a pre-trained detection model to detect the fatigue state of the user under test within a set time period, and obtain fatigue detection results.
[0012] Furthermore, the acquisition of facial key points of the user to be detected specifically includes:
[0013] Acquire the image of the user to be detected output by the camera device;
[0014] The face image of the user to be detected is obtained by identifying the face location in the image of the user to be detected through a pre-trained face recognition model.
[0015] The face image is subjected to key point detection by a pre-trained key point detection model to obtain the key points of the face of the user to be detected.
[0016] Furthermore, before sequentially stitching together face images containing facial landmarks within a set time window to obtain a face image sequence, the method further includes:
[0017] The facial image of the user to be detected is corrected based on the facial key points to obtain a facial image containing a preset background area;
[0018] Accordingly, the step of sequentially stitching together facial images containing facial landmarks within a set time window to obtain a facial image sequence specifically includes:
[0019] Within a set time window, the face images containing a preset background area are sequentially stitched together to obtain a face image sequence.
[0020] Furthermore, the step of correcting the facial image of the user to be detected based on the facial key points to obtain a facial image including a preset background area specifically includes:
[0021] The facial key points are subjected to affine transformation by a preset average face vector, and the affine transformation matrix of the facial key points is calculated.
[0022] The affine transformation matrix of the facial key points is applied to the facial image of the user to be detected in order to correct the facial image of the user to be detected and obtain a facial image containing a preset background area.
[0023] Further, the step of applying the affine transformation matrix of the facial key points to the face image of the user to be detected, in order to correct the face image of the user to be detected and obtain a face image containing a preset background region, specifically includes:
[0024] Based on the affine transformation matrix of the facial key points, determine the matrix transformation parameters corresponding to different affine transformation types of the face image of the user to be detected;
[0025] The matrix transformation parameters of the different radiation transformation types are combined in any order and number of times. The combined matrix transformation parameters are then used to correct the face image of the user to be detected, resulting in a face image containing a preset background area.
[0026] Furthermore, the step of inputting the facial expression features fused with spatiotemporal information into a pre-trained detection model to detect the fatigue state of the user to be detected within a set time period and obtain fatigue detection results specifically includes:
[0027] The facial expression features fused with spatiotemporal information are processed to obtain discrete features of the facial image, the discrete features including features obtained by spatial or temporal segmentation of facial image features;
[0028] The discrete features of the face image are input into a pre-trained detection model. The detection model combines the discrete features of the face image spatially or temporally to obtain the deep features of the facial expression.
[0029] The fatigue state of the user under test is detected based on the depth features of the facial expression within a set time period, and the fatigue detection result is determined.
[0030] Furthermore, the step of detecting the fatigue state of the user under test within a set time period based on the depth features of the facial expression and determining the fatigue detection result specifically includes:
[0031] The deep features of the facial expressions are processed using flat pooling and fully connected methods to obtain the global features of the facial expressions.
[0032] Based on the preset fatigue expression, determine the fatigue score of the global features of the facial expression on different fatigue labels;
[0033] Based on the global features of the facial expressions and the fatigue scores on different fatigue tags, the fatigue state of the user to be tested is detected within a set time period, and the fatigue detection result is determined.
[0034] According to a second aspect of this application, a fatigue state detection device is provided, the device comprising:
[0035] The acquisition unit is used to acquire the facial key points of the user to be detected;
[0036] The stitching unit is used to sequentially stitch together face images containing facial key points within a set time window to obtain a face image sequence.
[0037] The extraction unit is used to extract features from the face image sequence using a pre-trained feature extraction model to obtain facial expression features that incorporate spatiotemporal information;
[0038] The detection unit is used to input the facial expression features fused with spatiotemporal information into a pre-trained detection model, so as to detect the fatigue state of the user to be detected within a set time period through the detection model and obtain fatigue detection results.
[0039] Furthermore, the acquisition unit is specifically used to acquire the image of the user to be detected output by the camera; identify the face position in the image of the user to be detected by a pre-trained face recognition model to obtain the face image of the user to be detected; and perform key point detection on the face image by a pre-trained key point detection model to obtain the face key points of the user to be detected.
[0040] Furthermore, the device also includes:
[0041] The correction unit is used to correct the face image of the user to be detected based on the face key points before the face image containing face key points is sequentially stitched together within a set time window to obtain a face image sequence, so as to obtain a face image containing a preset area background.
[0042] Accordingly, the stitching unit is specifically used to stitch the face images containing a preset area background in a time sequence within a set time window to obtain a face image sequence.
[0043] Furthermore, the correction unit includes:
[0044] The calculation module is used to perform an affine transformation on the facial key points using a preset average face vector, and to calculate the affine transformation matrix of the facial key points.
[0045] The correction module is used to apply the affine transformation matrix of the facial key points to the face image of the user to be detected, so as to correct the face image of the user to be detected and obtain a face image containing a preset background area.
[0046] Furthermore, the correction module is specifically used to determine the matrix transformation parameters corresponding to different radial transformation types of the face image of the user to be detected based on the affine transformation matrix of the facial key points; to combine the matrix transformation parameters of the different radial transformation types in any order and number of times, and to use the combined matrix transformation parameters to correct the face image of the user to be detected, thereby obtaining a face image containing a preset background area.
[0047] Furthermore, the detection unit includes:
[0048] The processing module is used to process the facial expression features fused with spatiotemporal information to obtain discrete features of the facial image, the discrete features including features obtained by spatial or temporal segmentation of facial image features;
[0049] The combination module is used to input the discrete features of the face image into a pre-trained detection model, and to combine the discrete features of the face image spatially or temporally through the detection model to obtain the deep features of the face expression.
[0050] The detection module is used to detect the fatigue state of the user to be detected within a set time period based on the depth features of the facial expression, and to determine the fatigue detection result.
[0051] Furthermore, the detection module is specifically used to process the deep features of the facial expression using average pooling and a fully connected method to obtain the global features of the facial expression; determine the fatigue score of the global features of the facial expression on different fatigue labels based on the preset fatigue expression; and detect the fatigue state of the user to be detected within a set time period based on the fatigue score of the global features of the facial expression on different fatigue labels to determine the fatigue detection result.
[0052] According to a third aspect of this application, a storage medium is provided that stores a computer program thereon, which, when executed by a processor, implements the above-described fatigue state detection method.
[0053] According to a fourth aspect of this application, a fatigue state detection device is provided, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the program to implement the above-described fatigue state detection method.
[0054] By employing the above technical solution, this application provides a fatigue state detection method and apparatus. Compared with the current method that relies on detection algorithms to define rules to determine whether a user is in a fatigue state, this application obtains the facial key points of the user to be detected, and based on the facial key points, splices facial images containing the facial key points in a temporal sequence within a set time window to obtain a facial image sequence. A pre-trained feature extraction model is used to extract features from the facial image sequence to obtain facial expression features that integrate spatiotemporal information. These spatiotemporal facial expression features are then input into a pre-trained detection model to detect the fatigue state of the user to be detected within a set time period, obtaining a fatigue detection result. The entire process uses facial expression features that integrate spatiotemporal information to detect the user's fatigue state within a set time period, giving greater consideration to the influence of facial expression features on fatigue state. This allows for the addition of facial expression features that aid in classification to the fatigue state detection process, improving the accuracy of the fatigue detection results.
[0055] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0056] The accompanying drawings, which are included to provide a further understanding of this application and form part of this application, illustrate exemplary embodiments and are used to explain this application, but do not constitute an undue limitation of this application. In the drawings:
[0057] Figure 1 is a flowchart illustrating a fatigue state detection method according to an embodiment of this application;
[0058] Figure 2 is a flowchart illustrating a specific implementation of step 101 in Figure 1;
[0059] Figure 3 is a flowchart illustrating a fatigue state detection method in another embodiment of this application;
[0060] Figure 4 is a schematic flowchart of a specific implementation of step 105 in Figure 3;
[0061] Figure 5 is a flowchart illustrating a specific implementation of step 104 in Figure 1.
[0062] Figure 6 is a schematic diagram of the structure of a fatigue state detection device according to an embodiment of this application;
[0063] Figure 7 is a schematic diagram of the device structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0064] The present application will be described in detail below with reference to the accompanying drawings and embodiments. It should be noted that, unless otherwise specified, the embodiments and features described in the embodiments of the present application can be combined with each other.
[0065] In related technologies, fatigue detection algorithms primarily determine whether a user is fatigued by defining the duration of time their eyes are closed and their mouth is open, as defined by rules in the algorithm. This process relies on the stability of the rules in the detection algorithm; if the rules change, the corresponding detection results will also have errors, making it impossible to accurately detect fatigue.
[0066] To address this issue, this embodiment provides a method for detecting fatigue state, as shown in Figure 2. This method can be applied to an in-vehicle server and includes the following steps:
[0067] 101. Obtain the facial key points of the user to be detected.
[0068] The facial key points of the user to be detected can include, but are not limited to, facial features and contours, such as eyebrows, eyes, nose, mouth, and facial outlines. These facial key points include both 2D and 3D key points. For 2D key points, the output is the x and y coordinates; for 3D key points, the output is the x, y, and z coordinates. Various types of mobile phones, laptops, cameras, and other devices with facial capture capabilities can be used for face acquisition. For example, video surveillance cameras and other devices with facial capture capabilities transmit the captured facial images and / or real-time video streams to an in-vehicle server. The in-vehicle server then processes the facial images and / or real-time video streams to obtain the facial key points of the user to be detected.
[0069] In the specific process of acquiring facial key points, face detection algorithms can be used to identify the position and size of faces in an image. Deep learning algorithms can be used for face detection. Then, the detected faces are aligned to reduce the impact of pose and scale variations on key point extraction. Common alignment methods include feature point-based alignment and geometric transformation-based alignment. Further, on the aligned face image, feature point localization algorithms are used to locate the key points. Common feature point localization methods include template matching, regression-based, and deep learning-based methods. Finally, the feature point localization results may contain some errors or inaccuracies. Smoothing and shape constraints can be used to improve the accuracy of facial key point extraction.
[0070] Specifically, face detection algorithms can use convolutional neural network (CNN) models. CNN models are a type of deep learning model that can effectively extract features from images. Here, CNN models can use frameworks such as VGGNet and ResNet, or be trained according to actual needs.
[0071] Furthermore, to improve the accuracy of facial landmark detection, different types of facial images can be added to the facial sample images during the training process of the facial detection algorithm using deep learning. This enriches the application scenarios covered by facial images. For example, facial images acquired under different lighting conditions, including strong light, low light, and backlight, can be used as facial sample images. Facial images taken from multiple angles, poses, and varying distances can also be used as facial sample images. Furthermore, facial images taken with different occlusions, including eyes, masks, hats, and hands covering the mouth, can also be used as sample facial images.
[0072] The execution subject of this application embodiment can be a fatigue state detection device. It can automatically detect and locate certain key facial points, such as eyes, nose, and mouth, from a given face image. The location information of these key points contains information in the face image, which is very important for tasks such as face recognition, expression analysis, and pose estimation.
[0073] 102. Within a set time window, perform time-series stitching of facial images containing facial key points to obtain a facial image sequence.
[0074] Typically, a face image containing facial landmarks can represent the set of feature points of the face frame corresponding to the user to be detected. Based on the face image containing facial landmarks, the geometric features of the face of the user to be detected can be described, and facial features such as the mouth, eyes, and nose can be located in the face image of the user to be detected.
[0075] Considering the continuity of facial image changes within a preset time period, in order to further characterize the continuously changing facial expression features in the facial images, facial images containing facial key points can be connected in chronological order within a set time period to form a continuous facial image sequence. Here, the connection process can arrange the features of each window or frame in chronological order, or it can splice the features of the entire sequence to obtain the facial image sequence.
[0076] Specifically, in the process of temporal stitching of facial landmark images, recurrent neural networks or long short-term memory networks can be used to capture temporal features in facial landmarks. By using temporal features, facial images containing facial landmarks are stitched together to obtain a facial image sequence.
[0077] 103. Use a pre-trained feature extraction model to extract features from the face image sequence to obtain facial expression features that incorporate spatiotemporal information.
[0078] In this embodiment, the pre-trained feature extraction model may include a neural network model obtained by combining convolutional neural networks, recurrent neural networks and / or long short-term memory networks. The specific network architecture and parameter settings can be adjusted and optimized according to the actual situation.
[0079] Specifically, after taking the face image sequence as input, a convolutional neural network (CNN) can be used to extract the static features of each image frame. By using convolutional and pooling layers in the CNN, spatial information in the face image can be captured. To capture the temporal information in the face image sequence, a recurrent neural network (RNN) and / or a long short-term memory (LSTM) network is used to process the features of each image frame. The RNN and / or LSTM network can pass information from previous frames to subsequent frames, thereby establishing temporal relationships. Finally, in the output of the RNN and / or LSTM network, the features of each time step can be fused to obtain the feature representation of the entire face image sequence. Here, pooling operations and other methods can be used for feature fusion, ultimately obtaining facial expression features that fuse spatiotemporal information.
[0080] 104. Input the facial expression features fused with spatiotemporal information into a pre-trained detection model, so as to detect the fatigue state of the user to be detected within a set time period through the detection model, and obtain fatigue detection results.
[0081] The pre-trained detection model can be a model trained using a self-attention neural network. Here, facial expression features with spatiotemporal information can be fully connected for mapping, and a mapping label is added at the beginning of the fully connected mapping. Considering that the time setting only includes the fatigue state and is not sensitive to position, it is free to choose whether to add a position label. The temporal mapping vector with the mapping label is further input into the self-attention neural network model for training, and the probability values of facial expressions on different fatigue state labels are output. The fatigue state label with the highest probability is used as the fatigue state classification result output by the model. The actual fatigue state label corresponding to the facial expression is compared with the fatigue state classification result output by the model. The loss value of the model is calculated by using a loss function. When the loss value does not meet the model's iteration stopping condition, the above model training process is repeated to adjust the model parameters until the loss value meets the model's iteration stopping condition. The model output that meets the iteration stopping condition is used as the detection model.
[0082] Specifically, in the fatigue state detection process, firstly, a fully connected linear mapping is performed on the facial expression features that are fused with spatiotemporal information to obtain discrete units of facial expression coefficients. Then, a pre-trained detection model is used to extract deep features from the discrete units of facial expression coefficients. Furthermore, the extracted deep features are spatiotemporally fused through a fully connected layer. Finally, the spatiotemporally fused features are classified according to predefined fatigue state labels. The fatigue detection result is determined based on the probability values of the spatiotemporally fused features distributed on different fatigue state labels.
[0083] The fatigue detection method provided in this application, compared with the current method that relies on detection algorithms to define rules to determine whether a user is fatigued, obtains the facial key points of the user to be detected, and splices facial images containing the facial key points in a time sequence within a set time window to obtain a facial image sequence. A pre-trained feature extraction model is used to extract features from the facial image sequence to obtain facial expression features that incorporate spatiotemporal information. These spatiotemporal facial expression features are then input into a pre-trained detection model to detect the fatigue state of the user to be detected within a set time period, resulting in a fatigue detection result. The entire process uses spatiotemporal facial expression features to detect the user's fatigue state within a set time period, giving greater consideration to the influence of facial expression features on fatigue state. This allows for the addition of facial expression features that aid in classification to the fatigue detection process, improving the accuracy of the fatigue detection results.
[0084] In practical applications, the accuracy and efficiency of facial landmark detection are crucial for many applications, such as face recognition, expression analysis, and face tracking. Specifically, in the above embodiment, as shown in Figure 2, step 101 includes the following steps:
[0085] 201. Obtain the image of the user to be detected output by the camera device.
[0086] 202. The face location is identified in the image of the user to be detected by a pre-trained face recognition model, and the face image of the user to be detected is obtained.
[0087] 203. Perform key point detection on the face image using a pre-trained key point detection model to obtain the facial key points of the user to be detected.
[0088] Facial expressions can be used as a basis for judging fatigue status. By performing facial recognition on the image of the user to be identified, the face can be accurately located and identified in the image or video. Furthermore, the fatigue status of the user can be detected based on the facial expressions observed, thereby obtaining fatigue detection results.
[0089] Considering the accuracy and reliability of the image to be detected in subsequent analysis and application, the image of the user to be detected output by the camera device can be preprocessed, including image scaling, grayscale conversion, and histogram equalization, to improve the subsequent processing effect. Specifically, in the face recognition process, a pre-trained face recognition model can be used to find image regions that may contain faces. By calculating local features in the image, such as edges and textures, it is determined whether a face exists. Then, the extracted features are input into a classifier for classification to determine whether the region is a face. The classifier can learn the features of a series of positive and negative samples to determine whether a new image region is a face. Then, based on the output of the classifier, the regions in the image that may contain faces are determined, and the face image of the user to be detected is obtained. Rectangular or elliptical boxes are usually used to represent the position and size of the face. After obtaining the face image of the user to be detected, a pre-trained keypoint detection model can be used to locate the key points of the face, and the key points of the face of the user to be detected are obtained.
[0090] Considering the angle issues of human faces in practical applications, it is necessary to correct the facial pose in the image. Furthermore, in the above embodiment, as shown in Figure 3, before step 102, the method also includes the following steps:
[0091] 105. Correct the face image of the user to be detected based on the facial key points to obtain a face image containing a preset background area.
[0092] In practical applications, the facial images of the users to be detected may have different problems due to various factors, such as different camera angles and different human movements, which may cause the filtered facial images to not meet the optimal state for feature extraction. Since most of the training set used by the model during training consists of frontal face images, the facial images of the users to be detected need to be corrected before image feature extraction.
[0093] Specifically, facial landmarks provide information about the facial framework and the location of facial features. By using facial landmarks, the coordinates of the face and facial features can be obtained. Furthermore, based on these coordinates, the facial landmarks are aligned to a specific position—a standard position—to achieve standard or better analysis and comparison results. Rigid transformations, including translation, rotation, and scaling, can be used to align the facial features corresponding to the facial landmarks to a standard position in the image.
[0094] Furthermore, after facial correction is completed, some blank or overlapping areas may still be present. To eliminate these problems, interpolation techniques can be used to fill in blank areas, while overlapping areas can be smoothed out through fusion operations.
[0095] Accordingly, in step 102, the face images containing the preset area background are spliced together in a time sequence within a set time window to obtain a face image sequence.
[0096] It should be noted that in some cases, rigid transformation may not completely solve the alignment problem, especially when the facial pose changes significantly. In such cases, a more flexible affine transformation is needed to further improve the alignment effect by adjusting the distance, angle, and proportion between facial key points. Specifically, in the above embodiment, as shown in Figure 4, step 105 includes the following steps:
[0097] 301. Perform an affine transformation on the facial key points using a preset average face vector, and calculate the affine transformation matrix of the facial key points.
[0098] 302. Apply the affine transformation matrix of the facial key points to the face image of the user to be detected, so as to correct the face image of the user to be detected and obtain a face image containing a preset background area.
[0099] The preset average face vector is calculated by averaging a large number of face images with standard poses. Specifically, different average face data can be constructed based on gender using face data stored in a database. After determining the preset average face vector, the collected facial key points are mapped and fused with the preset average face vector. This process is equivalent to an affine transformation. Specifically, the Euclidean distance method can be used to calculate the difference between the preset average face vector and the facial key points. The affine transformation matrix of the facial key points is then calculated using this difference. Here, the difference can be adjusted using a fixed template or a custom setting.
[0100] Specifically, in the process of applying the affine transformation matrix of facial key points to the face image of the user to be detected for correction, the matrix transformation parameters corresponding to different affine transformation types of the face image of the user to be detected can be determined based on the affine transformation matrix of facial key points. Then, the matrix transformation parameters of different affine transformation types are combined in any order and number of times. The combined matrix transformation parameters are used to correct the face image of the user to be detected, resulting in a face image containing a preset background area.
[0101] Specifically, in the above embodiment, as shown in FIG5, step 104 includes the following steps:
[0102] 401. Process the facial expression features that have been fused with spatiotemporal information to obtain discrete features of the facial image.
[0103] 402. Input the discrete features of the face image into a pre-trained detection model, and use the detection model to combine the discrete features of the face image spatially or temporally to obtain the deep features of the face expression.
[0104] 403. Detect the fatigue state of the user to be detected within a set time period based on the depth features of the facial expression, and determine the fatigue detection result.
[0105] Discrete features include features obtained by spatially or temporally segmenting facial image features.
[0106] It is understandable that facial expression features fused with spatiotemporal information possess temporal characteristics. To convert these features into discrete labels for easier computer processing and analysis, the continuous data can be transformed into discrete labels, each representing a specific meaning or symbol. This facilitates data representation and processing. In machine learning and deep learning tasks, discretizing data with temporal characteristics converts it into fixed-length feature vectors, which can then be used as input for model training or other analytical tasks.
[0107] Specifically, the fatigue state of the user under test is detected based on the deep features of facial expressions within a set time period. In the process of determining the fatigue detection result, average pooling and fully connected methods can be used to process the deep features of facial expressions to obtain global features of facial expressions. Then, based on preset fatigue expressions, the fatigue score of the global features of facial expressions on different fatigue labels is determined. Finally, the fatigue state of the user under test is detected based on the fatigue scores of the global features of facial expressions on different fatigue labels within a set time period to determine the fatigue detection result.
[0108] Furthermore, as a specific implementation of the method in Figures 1-5, this application embodiment provides a fatigue state detection device, as shown in Figure 6. The device includes: an acquisition unit 51, a splicing unit 52, an extraction unit 53, and a detection unit 54.
[0109] Acquisition unit 51 is used to acquire facial key points of the user to be detected;
[0110] The stitching unit 52 is used to stitch face images containing facial key points in a time sequence within a set time window to obtain a face image sequence.
[0111] Extraction unit 53 is used to extract features from the face image sequence using a pre-trained feature extraction model to obtain facial expression features that incorporate spatiotemporal information.
[0112] The detection unit 54 is used to input the facial expression features fused with spatiotemporal information into a pre-trained detection model, so as to detect the fatigue state of the user to be detected within a set time through the detection model and obtain the fatigue detection result.
[0113] The fatigue detection device provided in this application, compared with the current method that relies on detection algorithms to define rules to determine whether a user is in a fatigue state, obtains the facial key points of the user to be detected, and splices facial images containing the facial key points in a time sequence within a set time window to obtain a facial image sequence. A pre-trained feature extraction model is used to extract features from the facial image sequence to obtain facial expression features that integrate spatiotemporal information. These spatiotemporal facial expression features are then input into a pre-trained detection model to detect the user's fatigue state within a set time period, resulting in a fatigue detection result. The entire process uses spatiotemporal facial expression features to detect the user's fatigue state within a set time period, giving greater consideration to the influence of facial expression features on fatigue state. This allows for the addition of facial expression features that aid in classification to the fatigue state detection process, improving the accuracy of the fatigue detection results.
[0114] In a specific application scenario, the acquisition unit 51 is specifically used to acquire the image of the user to be detected output by the camera; identify the face position in the image of the user to be detected by a pre-trained face recognition model to obtain the face image of the user to be detected; and perform key point detection on the face image by a pre-trained key point detection model to obtain the face key points of the user to be detected.
[0115] In specific application scenarios, the device further includes:
[0116] The correction unit is used to correct the face image of the user to be detected based on the face key points before the face image containing face key points is sequentially stitched together within a set time window to obtain a face image sequence, so as to obtain a face image containing a preset area background.
[0117] Accordingly, the stitching unit 52 is specifically used to stitch the face images containing the background of the preset area in a time sequence within a set time window to obtain a face image sequence.
[0118] In specific application scenarios, the correction unit includes:
[0119] The calculation module is used to perform an affine transformation on the facial key points using a preset average face vector, and to calculate the affine transformation matrix of the facial key points.
[0120] The correction module is used to apply the affine transformation matrix of the facial key points to the face image of the user to be detected, so as to correct the face image of the user to be detected and obtain a face image containing a preset background area.
[0121] In specific application scenarios, the correction module is specifically used to determine the matrix transformation parameters corresponding to different radiative transformation types of the face image of the user to be detected based on the affine transformation matrix of the facial key points; to combine the matrix transformation parameters of the different radiative transformation types in any order and number of times, and to use the combined matrix transformation parameters to correct the face image of the user to be detected, thereby obtaining a face image containing a preset background area.
[0122] In specific application scenarios, the detection unit includes:
[0123] The processing module is used to process the facial expression features fused with spatiotemporal information to obtain discrete features of the facial image, the discrete features including features obtained by spatial or temporal segmentation of facial image features;
[0124] The combination module is used to input the discrete features of the face image into a pre-trained detection model, and to combine the discrete features of the face image spatially or temporally through the detection model to obtain the deep features of the face expression.
[0125] The detection module is used to detect the fatigue state of the user to be detected within a set time period based on the depth features of the facial expression, and to determine the fatigue detection result.
[0126] In specific application scenarios, the detection module is specifically used to process the deep features of the facial expression using average pooling and full connection methods to obtain the global features of the facial expression; determine the fatigue score of the global features of the facial expression on different fatigue labels based on the preset fatigue expression; and detect the fatigue state of the user to be detected within a set time period based on the fatigue score of the global features of the facial expression on different fatigue labels to determine the fatigue detection result.
[0127] Based on the methods shown in Figures 1-5, this application also provides a storage medium storing a computer program that, when executed by a processor, implements the fatigue state detection method shown in Figures 1-5.
[0128] Based on this understanding, the technical solution of this application can be embodied in the form of a software product. This software product can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive), and includes several instructions to cause a computer device (such as a personal computer, server, or network device) to execute the methods described in the various implementation scenarios of this application.
[0129] Based on the methods shown in Figures 1-5 and the virtual device embodiment shown in Figure 6, in order to achieve the above objectives, this application embodiment also provides a physical device for detecting fatigue state, which can be a computer, smartphone, tablet computer, smartwatch, server, or network device, etc. The physical device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to implement the fatigue state detection method shown in Figures 1-5.
[0130] Optionally, the physical device may also include a user interface, a network interface, a camera, radio frequency (RF) circuitry, sensors, audio circuitry, a Wi-Fi module, etc. The user interface may include a display screen, input units such as a keyboard, etc., and optional user interfaces may also include USB interfaces, card reader interfaces, etc. The network interface may optionally include standard wired interfaces, wireless interfaces (such as Wi-Fi interfaces), etc.
[0131] In an exemplary embodiment, referring to FIG7, the above-described physical device includes a communication bus, a processor, a memory, and a communication interface. It may also include an input / output interface and a display device. The various functional units can communicate with each other via the bus. The memory stores a computer program, and the processor executes the program stored in the memory to perform the fatigue state detection method in the above embodiment.
[0132] Those skilled in the art will understand that the physical device structure for detecting fatigue state provided in this embodiment does not constitute a limitation on the physical device, and may include more or fewer components, or combine certain components, or have different component arrangements.
[0133] The storage medium may also include an operating system and a network communication module. The operating system is a program that manages the hardware and software resources of the physical device for detecting the aforementioned fatigue state, supporting the operation of information processing programs and other software and / or programs. The network communication module is used to enable communication between the various components within the storage medium, as well as communication with other hardware and software in the information processing physical device.
[0134] Through the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms, or it can be implemented by hardware. By applying the technical solution of this application, compared with the existing methods, this application uses facial expression features that integrate spatiotemporal information to detect the user's fatigue state within a set time. It takes into greater consideration the influence of facial expression features on fatigue state, and can add facial expression features that help classify fatigue state to the fatigue state detection process, thereby improving the accuracy of fatigue detection results.
[0135] Those skilled in the art will understand that the accompanying drawings are merely schematic diagrams of a preferred embodiment, and the modules or processes shown in the drawings are not necessarily essential for implementing this application. Those skilled in the art will understand that the modules in the apparatus of the embodiment can be distributed within the apparatus of the embodiment as described, or can be modified to be located in one or more apparatuses different from this embodiment. The modules of the above-described embodiment can be combined into one module, or further divided into multiple sub-modules.
[0136] The serial numbers in this application are for descriptive purposes only and do not represent the superiority or inferiority of any particular implementation scenario. The above disclosures are merely a few specific implementation scenarios of this application; however, this application is not limited thereto, and any variations conceived by those skilled in the art should fall within the protection scope of this application.
Claims
1. A method for detecting fatigue state, wherein, include: Obtain facial key points of the user to be detected; Within a set time window, face images containing facial landmarks are sequentially stitched together to obtain a face image sequence. The face image sequence is subjected to feature extraction using a pre-trained feature extraction model to obtain facial expression features that incorporate spatiotemporal information; The facial expression features fused with spatiotemporal information are input into a pre-trained detection model to detect the fatigue state of the user under test within a set time period, and obtain fatigue detection results.
2. The method according to claim 1, wherein, The acquisition of facial key points of the user to be detected specifically includes: Acquire the image of the user to be detected output by the camera device; The face image of the user to be detected is obtained by identifying the face location in the image of the user to be detected through a pre-trained face recognition model. The face image is subjected to key point detection by a pre-trained key point detection model to obtain the key points of the face of the user to be detected.
3. The method according to claim 1, wherein, Before sequentially stitching together facial images containing facial landmarks within a set time window to obtain a facial image sequence, the method further includes: The facial image of the user to be detected is corrected based on the facial key points to obtain a facial image containing a preset background area; Accordingly, the step of sequentially stitching together facial images containing facial landmarks within a set time window to obtain a facial image sequence specifically includes: Within a set time window, the face images containing a preset background area are sequentially stitched together to obtain a face image sequence.
4. The method according to claim 3, wherein, The step of correcting the facial image of the user to be detected based on the facial key points to obtain a facial image containing a preset background area specifically includes: The facial key points are subjected to affine transformation by a preset average face vector, and the affine transformation matrix of the facial key points is calculated. The affine transformation matrix of the facial key points is applied to the facial image of the user to be detected in order to correct the facial image of the user to be detected and obtain a facial image containing a preset background area.
5. The method according to claim 4, wherein, The step of applying the affine transformation matrix of the facial key points to the face image of the user to be detected, in order to correct the face image of the user to be detected and obtain a face image containing a preset background region, specifically includes: Based on the affine transformation matrix of the facial key points, determine the matrix transformation parameters corresponding to different affine transformation types of the face image of the user to be detected; The matrix transformation parameters of the different radiation transformation types are combined in any order and number of times. The combined matrix transformation parameters are then used to correct the face image of the user to be detected, resulting in a face image containing a preset background area.
6. The method according to any one of claims 1-5, wherein, The step of inputting the facial expression features fused with spatiotemporal information into a pre-trained detection model to detect the fatigue state of the user under test within a set time period and obtain fatigue detection results specifically includes: The facial expression features fused with spatiotemporal information are processed to obtain discrete features of the facial image, the discrete features including features obtained by spatial or temporal segmentation of facial image features; The discrete features of the face image are input into a pre-trained detection model. The detection model combines the discrete features of the face image spatially or temporally to obtain the deep features of the facial expression. The fatigue state of the user under test is detected based on the depth features of the facial expression within a set time period, and the fatigue detection result is determined.
7. The method according to claim 6, wherein, The step of detecting the fatigue state of the user under test within a set time period based on the depth features of the facial expression and determining the fatigue detection result specifically includes: The deep features of the facial expressions are processed using flat pooling and fully connected methods to obtain the global features of the facial expressions. Based on the preset fatigue expression, determine the fatigue score of the global features of the facial expression on different fatigue labels; Based on the global features of the facial expressions and the fatigue scores on different fatigue tags, the fatigue state of the user to be tested is detected within a set time period, and the fatigue detection result is determined.
8. A fatigue state detection device, wherein, include: The acquisition unit is used to acquire the facial key points of the user to be detected; The stitching unit is used to sequentially stitch together face images containing facial key points within a set time window to obtain a face image sequence. The extraction unit is used to extract features from the face image sequence using a pre-trained feature extraction model to obtain facial expression features that incorporate spatiotemporal information; The detection unit is used to input the facial expression features fused with spatiotemporal information into a pre-trained detection model, so as to detect the fatigue state of the user to be detected within a set time through the detection model and obtain fatigue detection results.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein... When the processor executes the computer program, it implements the steps of the method according to any one of claims 1 to 7.
10. A computer storage medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 7.
Citation Information
Patent Citations
Electric vehicle automatic driving method based on artificial intelligent platform
CN108609019A
Face fatigue driving detection method and device based on space-time feature recognition
CN110210382A
Driver fatigue state rapid detection method based on deep learning
CN110674701A
Dynamic expression recognition method, device and system and storage medium
CN117765583A
Face landmark detection method and apparatus for drive state monitoring
KR101788211B1