Fatigue state detection method and device
By acquiring facial key points, correcting and temporally stitching facial images, and using feature extraction and detection models to detect driver fatigue, the problem of inaccurate fatigue detection in existing technologies is solved, achieving higher detection accuracy.
Patent Information
- Application Number
- CN202410489671.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-23
- Publication Date
- 2025-10-24
AI Technical Summary
Existing fatigue detection algorithms struggle to accurately detect driver fatigue and are significantly affected by the stability of detection rules.
By acquiring the facial key points of the user to be detected, the pre-trained feature extraction model is used to extract features from the facial image sequence. After fusing spatiotemporal information, the feature is input into the pre-trained detection model for fatigue state detection, including facial key point correction, temporal stitching, and deep feature analysis.
It improves the accuracy of fatigue detection by giving more consideration to the impact of facial expression features on fatigue status, thus enhancing the reliability of the detection results.
Smart Images

Figure CN120833595A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of fatigue detection, in particular to a fatigue state detection method and device. BACKGROUND
[0002] With the continuous increase of modern automobile ownership, not only promotes the continuous development of transportation industry, also makes traffic accidents grow. Fatigue driving is the main cause of traffic accidents, effective supervision and prevention of fatigue driving, has a very important practical significance.
[0003] In the related art, the state information of the driver's eyes, mouth and other features can be used to determine whether the driver is in a fatigue state. This process needs to rely on detection algorithms to define rules, and if the rules change, the corresponding detection results will also be affected, making it difficult to accurately detect the fatigue state. SUMMARY
[0004] Therefore, the present application provides a fatigue state detection method and device, which mainly aims to solve the problem that the fatigue detection algorithm in the prior art is difficult to accurately detect the fatigue state.
[0005] According to a first aspect of the present application, a fatigue state detection method is provided, which comprises:
[0006] Obtaining the face key points of the user to be detected;
[0007] Time series splicing the face image containing the face key points within a set time window to obtain a face image sequence;
[0008] Using a pre-trained feature extraction model to extract features from the face image sequence to obtain face expression features fused with spatial and temporal information;
[0009] Inputting the face expression features fused with spatial and temporal information into a pre-trained detection model to detect the fatigue state of the user to be detected within a set time through the detection model to obtain a fatigue detection result.
[0010] Further, the face key points of the user to be detected are obtained, specifically comprising:
[0011] Obtaining the user image to be detected output by the photographing device;
[0012] Identifying the face position in the user image to be detected through a pre-trained face recognition model to obtain the face image of the user to be detected;
[0013] Detecting the key points of the face image through a pre-trained key point detection model to obtain the face key points of the user to be detected.
[0014] Further, before the face image sequence is obtained by sequentially splicing the face images containing the face key points within the set time window, the method further comprises:
[0015] correcting the face image of the to-be-detected user according to the face key points to obtain a face image containing a preset region background;
[0016] Correspondingly, the face image sequence is obtained by sequentially splicing the face images containing the face key points within the set time window, specifically comprising:
[0017] sequentially splicing the face images containing the preset region background within the set time window to obtain the face image sequence.
[0018] Further, the face image containing the preset region background is obtained by correcting the face image of the to-be-detected user according to the face key points, specifically comprising:
[0019] affine transform the face key points by a preset average face vector to calculate an affine transformation matrix of the face key points;
[0020] apply the affine transformation matrix of the face key points to the face image of the to-be-detected user to correct the face image of the to-be-detected user, and obtain the face image containing the preset region background.
[0021] Further, the face image containing the preset region background is obtained by applying the affine transformation matrix of the face key points to the face image of the to-be-detected user to correct the face image of the to-be-detected user, specifically comprising:
[0022] determine the matrix change parameters of different radial change types corresponding to the face image of the to-be-detected user according to the affine transformation matrix of the face key points;
[0023] combine the matrix change parameters of different radial change types in any order number, and correct the face image of the to-be-detected user using the combined matrix change parameters to obtain the face image containing the preset region background.
[0024] Further, the face expression feature fused with the space-time information is input into a pre-trained detection model to detect the fatigue state of the to-be-detected user within the set time by the detection model to obtain a fatigue detection result, specifically comprising:
[0025] processing the face expression feature fused with the space-time information to obtain discrete features of the face image, the discrete features including features obtained by dividing the face image features in space or time;
[0026] inputting the discrete features of the face image into a pre-trained detection model, combining the discrete features of the face image in space or time by the detection model, and obtaining deep features of a face expression;
[0027] detecting a fatigue state of the to-be-detected user within a set time according to the deep features of the face expression, and determining a fatigue detection result.
[0028] Further, the detecting the fatigue state of the to-be-detected user within the set time according to the deep features of the face expression, and determining the fatigue detection result specifically includes:
[0029] processing the deep features of the face expression using a global average pooling and a full connection manner, and obtaining global features of the face expression;
[0030] determining fatigue scores of the global features of the face expression on different fatigue labels according to pre-set fatigue expressions;
[0031] detecting the fatigue state of the to-be-detected user within the set time according to the fatigue scores of the global features of the face expression on the different fatigue labels, and determining the fatigue detection result.
[0032] According to a second aspect of the present application, a fatigue state detection device is provided, which includes:
[0033] an acquisition unit configured to acquire face key points of a to-be-detected user;
[0034] a splicing unit configured to time sequence splice face images containing the face key points within a set time window, and obtain a face image sequence;
[0035] an extraction unit configured to extract features of the face image sequence by using a pre-trained feature extraction model, and obtain face expression features fused with space-time information;
[0036] a detection unit configured to input the face expression features fused with the space-time information into a pre-trained detection model, and detect a fatigue state of the to-be-detected user within a set time by the detection model, and obtain a fatigue detection result.
[0037] Further, the acquisition unit is specifically configured to acquire a to-be-detected user image output by a photographing device, recognize a face position in the to-be-detected user image by using a pre-trained face recognition model, and obtain a face image of the to-be-detected user, and detect key points of the face image by using a pre-trained key point detection model, and obtain face key points of the to-be-detected user.
[0038] Further, the device further includes:
[0039] The correction unit is configured to correct a face image of a to-be-detected user according to the face key points to obtain a face image containing a preset region background before the face image containing the face key points is temporally spliced in a set time window to obtain a face image sequence.
[0040] Correspondingly, the splicing unit is specifically configured to temporally splice the face image containing the preset region background in the set time window to obtain the face image sequence.
[0041] Further, the correction unit comprises:
[0042] The calculation module is configured to perform affine transformation on the face key points by using a preset average face vector to calculate an affine transformation matrix of the face key points.
[0043] The correction module is configured to apply the affine transformation matrix of the face key points to the face image of the to-be-detected user to correct the face image of the to-be-detected user to obtain the face image containing the preset region background.
[0044] Further, the correction module is specifically configured to determine matrix change parameters corresponding to different radial change types of the face image of the to-be-detected user according to the affine transformation matrix of the face key points, combine the matrix change parameters of the different radial change types in any order number, and correct the face image of the to-be-detected user by using the combined matrix change parameters to obtain the face image containing the preset region background.
[0045] Further, the detection unit comprises:
[0046] The processing module is configured to process the face expression feature fused with the space-time information to obtain discrete features of the face image, the discrete features including features obtained by dividing the face image features in space or time.
[0047] The combination module is configured to input the discrete features of the face image into a pre-trained detection model, combine the discrete features of the face image in space or time by using the detection model, and obtain deep features of a face expression.
[0048] The detection module is configured to detect a fatigue state of the to-be-detected user in a set time according to the deep features of the face expression to determine a fatigue detection result.
[0049] Further, the detection module is specifically configured to process the deep features of the facial expression using a flat pooling and a full connection manner to obtain global features of the facial expression; determine fatigue scores of the global features of the facial expression on different fatigue labels according to preset fatigue expressions; and detect fatigue states of the user to be detected within a set time according to the fatigue scores of the global features of the facial expression on the different fatigue labels to determine a fatigue detection result.
[0050] According to a third aspect of the present application, a storage medium is provided, which stores a computer program, and the program is executed by a processor to implement the fatigue state detection method.
[0051] According to a fourth aspect of the present application, a fatigue state detection device is provided, which comprises a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, and the processor executes the program to implement the fatigue state detection method.
[0052] By the above technical solution, compared with the current method of defining rules to determine whether a user is in a fatigue state by relying on a detection algorithm, the fatigue state detection method and device provided by the present application obtain facial key points of a user to be detected, perform time sequence splicing on facial images containing the facial key points within a set time window according to the facial key points to obtain a facial image sequence, perform feature extraction on the facial image sequence by using a pre-trained feature extraction model to obtain facial expression features fused with space-time information, input the facial expression features fused with space-time information into a pre-trained detection model to detect fatigue states of the user to be detected within the set time by using the detection model, and obtain a fatigue detection result. The fatigue states of the user within the set time are detected by using the facial expression features fused with space-time information, the influence of the facial expression features on the fatigue states is considered more, the facial expression features that are helpful for classification can be added to the fatigue state detection process, and the accuracy of the fatigue detection result is improved.
[0053] The above description is only a summary of the technical solutions of the present application. In order to enable one skilled in the art to better understand the technical means of the present application, the content of the specification can be implemented, and in order to enable the above and other purposes, features and advantages of the present application to be more obvious and easy to understand, the following specific embodiments of the present application are described. BRIEF DESCRIPTION OF DRAWINGS
[0054] The drawings described herein are used to provide further understanding of the present application, and form a part of the present application. The schematic embodiments of the present application and their descriptions are used to explain the present application, and do not constitute an improper limitation on the present application. In the drawings:
[0055] Figure 1is a flowchart of a fatigue state detection method in an embodiment of the present application;
[0056] Figure 2 is Figure 1 is a flowchart of a specific implementation of step 101 in the method;
[0057] Figure 3 is a flowchart of a fatigue state detection method in another embodiment of the present application;
[0058] Figure 4 is Figure 3 is a flowchart of a specific implementation of step 105 in the method;
[0059] Figure 5 is Figure 1 is a flowchart of a specific implementation of step 104 in the method
[0060] Figure 6 is a structural diagram of a fatigue state detection device in an embodiment of the present application;
[0061] Figure 7 is a structural diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION
[0062] The present application will be described in detail below with reference to the accompanying drawings and in conjunction with embodiments. It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict.
[0063] In related technologies, whether a user is in a fatigue state is mainly determined by the time length of closing eyes and opening mouth defined by rules in a fatigue detection algorithm. This process needs to rely on the stability of the rules in the detection algorithm. If the rules change, the corresponding detection result will also have errors, and the fatigue state cannot be accurately detected.
[0064] To solve this problem, the present embodiment provides a fatigue state detection method, as shown in Figure 2 The method can be applied to a vehicle-mounted server and includes the following steps:
[0065] 101. Obtain the face key points of a user to be detected.
[0066] The face key points of the to-be-detected user can include but are not limited to facial features such as eyebrows, eyes, nose, mouth, and facial contour. The face key points include 2D key points and 3D key points. For the 2D key points, the x and y coordinate information of the face key points is output. For the 3D key points, the x, y, and z coordinate information of the face key points is output. In the face collection device, various types of mobile phones, notebooks, cameras, and other devices with face collection functions can be used. For example, a video monitoring camera or other device with a face collection function collects the face image and / or real-time video stream and transmits the collected face image and / or real-time video stream to the vehicle-mounted server. The vehicle-mounted server performs image processing on the face image and / or real-time video stream to obtain the face key points of the to-be-detected user.
[0067] In the process of obtaining specific face key points, a face detection algorithm can be used to identify the position and size of the face in the image. The face detection algorithm can use a deep learning algorithm. Then, the detected face is aligned to reduce the influence of posture and scale changes on key point extraction. Common alignment methods include feature point-based alignment and geometric transformation-based alignment. Further, on the aligned face image, a feature point positioning algorithm is used to position the key points. Common feature point positioning methods include template matching, regression, deep learning, etc. Finally, for the results of feature point positioning, some errors or inaccuracies may exist. Smoothing, shape constraint, and other operations can be used to improve the accuracy of the face key points.
[0068] Specifically, the face detection algorithm can use a convolutional neural network model, which is a deep learning model that can effectively extract features from images. The convolutional neural network model can use VGGNet, ResNet, or other frameworks, or be trained according to actual needs.
[0069] Further, to improve the accuracy of face key point detection, different types of face images can be added to the face sample images during the training process of the face detection algorithm using deep learning, thereby enriching the application scenarios covered by the face images. For example, face images captured under different light conditions, including strong light, dim light, and backlight, can be used as face sample images. Face images captured at multiple angles, poses, and distances can also be used as face sample images. Face images captured with different obstructions, including eyes, masks, hats, and hands covering the mouth, can also be used as sample face images.
[0070] The execution subject of the embodiment of the present application can be a fatigue state detection device. Some specific facial key points, such as eyes, nose, and mouth, are automatically detected and located from a given facial image. The position information of these key points contains information in the facial image, which is very important for facial recognition, expression analysis, and posture estimation.
[0071] 102. Time sequence splicing the facial image containing the facial key points in a set time window to obtain a facial image sequence.
[0072] Generally, the facial image containing the facial key points can represent the feature point set of the corresponding face frame of the to-be-detected user. The geometric features of the face of the to-be-detected user can be depicted according to the facial image containing the facial key points, and the five organs such as the mouth, eyes, and nose can be located in the facial image of the to-be-detected user.
[0073] Considering the continuity of the facial image change in a preset time, in order to further represent the continuously changing expression features in the facial image, the facial image containing the facial key points can be connected in time sequence in a set time to form a continuous facial image sequence. Here, the connection process can arrange the features of each window or frame in time sequence, or can splice the features of the entire sequence to obtain the facial image sequence.
[0074] Specifically, in the process of time sequence splicing the facial key point image, a recurrent neural network or a long short-term memory network can be used to capture the time sequence features in the facial key points. The facial image containing the facial key points is spliced through the time sequence features to obtain the facial image sequence.
[0075] 103. Feature extraction is performed on the facial image sequence by using a pre-trained feature extraction model to obtain facial expression features fused with space-time information.
[0076] In the embodiment, the pre-trained feature extraction model can include a neural network model obtained by combining a convolutional neural network, a recurrent neural network, and / or a long short-term memory network. The specific network architecture and parameter settings can be adjusted and optimized according to actual conditions.
[0077] Specifically, after taking the face image sequence as input, a convolutional neural network can be used to extract the static features of each image frame. By using convolutional layers and pooling layers in the convolutional neural network, spatial information in the face image can be captured. In order to capture the time sequence information in the face image sequence, a recurrent neural network and / or a long short-term memory network is then used to process the features of each image frame. The recurrent neural network and / or the long short-term memory network can pass the information of the previous frame to the subsequent frame, thereby establishing a time sequence relationship. Finally, in the output of the recurrent neural network and / or the long short-term memory network, the features of each time step can be fused to obtain a feature representation of the entire face image sequence. Here, a pooling operation or the like can be used for feature fusion, and finally a face expression feature fused with spatial and temporal information is obtained.
[0078] 104. inputting the face expression feature fused with spatial and temporal information into a pre-trained detection model to detect the fatigue state of the user to be detected within a set time through the detection model, and obtaining a fatigue detection result.
[0079] The pre-trained detection model can be a model trained using a self-attention neural network. Here, the face expression feature fused with spatial and temporal information can be fully connected and mapped, and a mapping label can be added at the beginning of the fully connected mapping. Considering that the set time only has a fatigue state and is not sensitive to position, it can be freely selected whether to superimpose a position label. Further, the time sequence mapping vector with the mapping label is input into the self-attention neural network model for training, and the probability value of the face expression on different fatigue state labels is output. The fatigue state classification result output by the model is used as the fatigue state classification result output by the model. The actual fatigue state label corresponding to the face expression is compared with the fatigue state classification result output by the model, and the loss value of the model is calculated through a loss function. When the loss value does not meet the iteration stopping condition of the model, the above model training process is repeated to adjust the model parameters until the loss value meets the iteration stopping condition of the model. The model output that meets the iteration stopping condition is used as the detection model.
[0080] Specifically, in the fatigue state detection process, the face expression feature fused with spatial and temporal information is first linearly mapped by full connection to obtain discrete units of face expression coefficients. Then, the pre-trained detection model is used to extract deep features of the discrete units of face expression coefficients. Further, the extracted deep features are spatio-temporally fused through a fully connected layer. Finally, the spatio-temporally fused features are classified according to pre-defined fatigue state labels, and the fatigue detection result is determined according to the probability value of the spatio-temporally fused features distributed on different fatigue state labels.
[0081] Compared with the current method of defining rules to determine whether a user is in a fatigue state by relying on a detection algorithm, the fatigue state detection method provided in the embodiments of the present application obtains face key points of a user to be detected, performs time sequence splicing on face images containing the face key points within a set time window according to the face key points, obtains a face image sequence, performs feature extraction on the face image sequence by using a pre-trained feature extraction model, obtains face expression features fused with space-time information, inputs the face expression features fused with space-time information into a pre-trained detection model, and detects the fatigue state of the user to be detected within a set time by using the detection model to obtain a fatigue detection result. The fatigue state of the user within the set time is detected by using the face expression features fused with space-time information throughout the process, more consideration is given to the influence of face expression features on the fatigue state, face expression features that are helpful for classification can be added to the fatigue state detection process, and the accuracy of the fatigue detection result is improved.
[0082] In actual application scenarios, the accuracy and efficiency of face key point detection are crucial for many applications, such as face recognition, expression analysis, and face tracking. Specifically, in the above embodiments, as shown in step 101, the following steps are included: Figure 2
[0083] 201. Obtain a user image to be detected output by a photographing device.
[0084] 202. Identify a face position in the user image to be detected by using a pre-trained face recognition model, and obtain a face image of the user to be detected.
[0085] 203. Perform key point detection on the face image by using a pre-trained key point detection model, and obtain face key points of the user to be detected.
[0086] Face expression is used as a basis for judging a fatigue state. By performing face recognition on a user image to be recognized, the face can be accurately positioned and recognized in an image or a video. Further, the fatigue state of the user to be recognized is detected according to the recognized face expression, and a fatigue detection result is obtained.
[0087] In view of the accuracy and reliability of the to-be-detected image in subsequent analysis and application, the to-be-detected user image output by the photographing device can be preprocessed, including image scaling, grayscale, histogram equalization and the like, so as to improve the subsequent processing effect. Specifically, in the face recognition process, a pre-trained face recognition model can be used to find an image region that can contain a face, whether a face exists is determined by calculating local features such as edges and textures in the image, and then the extracted features are input into a classifier for classification to determine whether the region is a face. The classifier can determine whether a new image region is a face by learning the features of a series of positive and negative samples. Then, according to the output result of the classifier, a region in the image where a face can exist is determined, and a face image of the to-be-detected user is obtained, which is usually represented by a rectangular or elliptical frame to indicate the position and size of the face. After obtaining the face image of the to-be-detected user, a pre-trained key point detection model can be used to locate the key points of the face to obtain the face key points of the to-be-detected user.
[0088] In view of the angle problem of the face in actual application, the face posture in the image needs to be corrected. Further, in the above embodiment, as shown in Figure 3 Before step 102, the method further includes the following steps:
[0089] 105. Correct the face image of the to-be-detected user according to the face key points to obtain a face image containing a background of a preset region.
[0090] In actual application, due to various factors, the face image of the to-be-detected user collected can have different problems, for example, different camera angles and different human actions, so that the filtered face image cannot meet the best state for feature extraction. Since most of the training set used in the training process of the model is a front face image, the face image of the to-be-detected user needs to be corrected before image feature extraction.
[0091] Specifically, the face key points can feedback the positioning of the face frame and the features in the face. The positioning coordinate information of the face and the features can be obtained through the face key points. Further, the face key points are aligned to a specific position according to the positioning coordinate information of the face and the features. The specific position is a standard position, so as to achieve a standard or better analysis and comparison effect. When aligning the key points to the specific position, rigid transformation can be used, including translation, rotation and scaling operations, so that the position and size of the face features corresponding to the face key points in the image match the standard position.
[0092] Further, after completing the face correction, some blank areas or overlapping areas can also be generated. In order to eliminate these problems, interpolation technology can be used to fill the blank areas, and fusion operation can be used to smooth the overlapping areas.
[0093] Correspondingly, step 102, the face image containing the preset region background is time-series spliced in the set time window to obtain a face image sequence.
[0094] It should be noted that in some cases, rigid transformation may not completely solve the alignment problem, especially when the face pose changes greatly, at this time, a more flexible affine transformation is needed to further improve the alignment effect by adjusting the distance, angle and scale between the face key points. Specifically, in the above embodiment, as shown in Figure 4 Step 105 includes the following steps:
[0095] 301. Affine transform the face key points by the preset average face vector to calculate the affine transformation matrix of the face key points.
[0096] 302. Apply the affine transformation matrix of the face key points to the face image of the user to be detected to correct the face image of the user to be detected, and obtain a face image containing a preset region background.
[0097] The preset average face vector is a face vector calculated by averaging a large number of standard face images with different poses, and different average face data can be constructed according to different genders based on the face data stored in the database. After determining the preset average face vector, the collected face key points and the preset average vector are mapped and fused, which is equivalent to the process of affine transformation. Specifically, the Euclidean distance calculation method can be used to calculate the difference between the preset average face vector and the face key points, and the affine transformation matrix of the face key points is calculated by the difference. Here, the difference can be added to a fixed template or customized control.
[0098] Specifically, in the process of applying the affine transformation matrix of the face key points to the face image of the user to be detected to correct the face image of the user to be detected, the matrix change parameters of different radiation change types corresponding to the face image of the user to be detected can be determined according to the affine transformation matrix of the face key points, and then the matrix change parameters of different radiation change types are combined in any order number. The face image of the user to be detected is corrected using the combined matrix change parameters to obtain a face image containing a preset region background.
[0099] Specifically, in the above embodiment, as shown in Figure 5 Step 104 includes the following steps:
[0100] 401. Process the face expression features fused with the space-time information to obtain discrete features of the face image.
[0101] 402. Input the discrete features of the facial image into a pre-trained detection model, and use the detection model to spatially or temporally combine the discrete features of the facial image to obtain depth features of facial expressions.
[0102] 403 : Detect the fatigue state of the user to be detected within a set time based on the depth feature of the facial expression, and determine a fatigue detection result.
[0103] Among them, discrete features include features obtained by segmenting facial image features in space or time
[0104] It is understandable that facial expression features that incorporate spatiotemporal information have temporal characteristics. To convert these features into discrete tokens for easier computer processing and analysis, they can be processed to convert continuous data into discrete tokens, each representing a specific meaning or symbol. This allows for easier data representation and processing. In machine learning and deep learning tasks, discretizing data with temporal characteristics converts it into fixed-length feature vectors, which can be used as input for model training or other analytical tasks.
[0105] Specifically, when detecting the fatigue state of the user to be detected within a set time based on the deep features of facial expressions, and determining the fatigue detection results, the deep features of the facial expressions can be processed using average pooling and full connection methods to obtain global features of facial expressions. Then, based on the preset fatigue expressions, the fatigue scores of the global features of facial expressions on different fatigue labels are determined. Finally, based on the fatigue scores of the global features of facial expressions on different fatigue labels, the fatigue state of the user to be detected within the set time is detected to determine the fatigue detection results.
[0106] Further, as Figures 1-5 The specific implementation of the method, the embodiment of the present application provides a fatigue state detection device, such as Figure 6 As shown, the device includes: an acquisition unit 51, a splicing unit 52, an extraction unit 53, and a detection unit 54.
[0107] An acquisition unit 51 is used to acquire key points of the face of the user to be detected;
[0108] A splicing unit 52 is used to splice the facial images containing facial key points in a set time window to obtain a facial image sequence;
[0109] An extraction unit 53 is used to extract features from the facial image sequence using a pre-trained feature extraction model to obtain facial expression features that are integrated with temporal and spatial information;
[0110] The detection unit 54 is configured to input the facial expression feature fused with the space-time information into a pre-trained detection model to detect the fatigue state of the to-be-detected user within a set time through the detection model, and obtain a fatigue detection result.
[0111] Compared with the current way of defining rules to determine whether a user is in a fatigue state by means of a detection algorithm, the fatigue state detection device provided by the embodiment of the application obtains facial key points of a to-be-detected user, sequentially splices facial images containing the facial key points within a set time window according to the facial key points to obtain a facial image sequence, extracts features of the facial image sequence by using a pre-trained feature extraction model, obtains facial expression features fused with space-time information, inputs the facial expression features fused with the space-time information into a pre-trained detection model to detect the fatigue state of the to-be-detected user within the set time through the detection model, and obtains a fatigue detection result. The fatigue state of the user within the set time is detected by using the facial expression features fused with the space-time information, the influence of the facial expression features on the fatigue state is considered more, the facial expression features that are helpful to classification can be added to the fatigue state detection process, and the accuracy of the fatigue detection result is improved.
[0112] In a specific application scenario, the obtaining unit 51 is specifically configured to obtain a to-be-detected user image output by a photographing device, recognize a facial position in the to-be-detected user image by using a pre-trained face recognition model, and obtain a facial image of the to-be-detected user, and detect facial key points of the to-be-detected user by using a pre-trained key point detection model.
[0113] In a specific application scenario, the device further includes:
[0114] The correction unit is configured to correct the facial image of the to-be-detected user according to the facial key points before the facial image containing the facial key points is sequentially spliced within the set time window to obtain the facial image sequence, and obtain a facial image containing a background of a preset region.
[0115] Correspondingly, the splicing unit 52 is specifically configured to sequentially splice the facial image containing the background of the preset region within the set time window to obtain the facial image sequence.
[0116] In a specific application scenario, the correction unit includes:
[0117] The calculation module is configured to perform affine transformation on the facial key points by using a pre-set average face vector, and calculate an affine transformation matrix of the facial key points.
[0118] The correction module is configured to apply the affine transformation matrix of the face key point to a face image of a to-be-detected user to correct the face image of the to-be-detected user, and obtain a face image containing a preset region background.
[0119] In a specific application scenario, the correction module is specifically configured to determine matrix variation parameters of different radial variation types of the face image of the to-be-detected user according to the affine transformation matrix of the face key point, combine the matrix variation parameters of the different radial variation types in any order, and correct the face image of the to-be-detected user using the combined matrix variation parameters to obtain a face image containing a preset region background.
[0120] In a specific application scenario, the detection unit includes:
[0121] The processing module is configured to process the face expression feature fused with the space-time information to obtain discrete features of the face image, the discrete features including features obtained by dividing the face image features in space or time.
[0122] The combination module is configured to input the discrete features of the face image into a pre-trained detection model, combine the discrete features of the face image in space or time by using the detection model, and obtain deep features of the face expression.
[0123] The detection module is configured to detect the fatigue state of the to-be-detected user within a set time according to the deep features of the face expression, and determine a fatigue detection result.
[0124] In a specific application scenario, the detection module is specifically configured to process the deep features of the face expression using a flat pooling and a full connection method to obtain global features of the face expression, determine fatigue scores of the global features of the face expression on different fatigue labels according to a preset fatigue expression, and detect the fatigue state of the to-be-detected user within the set time according to the fatigue scores of the global features of the face expression on the different fatigue labels to determine the fatigue detection result.
[0125] Based on the above-mentioned method as shown in Figures 1-5 Accordingly, an embodiment of the present application also provides a storage medium having a computer program stored thereon, the program being executed by a processor to implement the above-mentioned fatigue state detection method as shown in Figures 1-5
[0126] Based on such understanding, the technical scheme of the present application can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a U disk, a mobile hard disk, etc.), and includes a plurality of instructions to make a computer device (which can be a personal computer, a server, or a network device, etc.) execute the method described in various implementation scenarios of the present application.
[0127] Based on the method as shown in Figures 1-5 , and Figure 6 the virtual device embodiment, in order to achieve the above-mentioned purpose, the embodiment of the present application also provides an entity device for detecting fatigue state, which can be a computer, a smart phone, a tablet computer, a smart watch, a server, or a network device, etc. The entity device includes a storage medium and a processor; the storage medium is used to store a computer program; the processor is used to execute the computer program to realize the fatigue state detection method as shown in Figures 1-5 .
[0128] Optionally, the entity device can also include a user interface, a network interface, a camera, a radio frequency (RF) circuit, a sensor, an audio circuit, a WI-FI module, etc. The user interface can include a display screen (Display), an input unit such as a keyboard (Keyboard), etc. The optional user interface can also include a USB interface, a card reader interface, etc. The network interface can optionally include a standard wired interface, a wireless interface (such as a WI-FI interface), etc.
[0129] In an exemplary embodiment, referring to Figure 7 , the above-mentioned entity device includes a communication bus, a processor, a memory, and a communication interface, and can also include an input / output interface and a display device, wherein the communication between various functional units can be completed through the bus. The memory stores a computer program, and the processor is used to execute the program stored on the memory to execute the fatigue state detection method in the above-mentioned embodiment.
[0130] Those skilled in the art can understand that the structure of the entity device for detecting fatigue state provided by the present embodiment does not constitute a limitation on the entity device, and can include more or fewer components, or combine certain components, or different component arrangements.
[0131] The storage medium can also include an operating system and a network communication module. The operating system is a program for managing the hardware and software resources of the above-mentioned entity device for detecting fatigue state, supporting the running of information processing programs and other software and / or programs. The network communication module is used to realize the communication between the components inside the storage medium, and the communication with other hardware and software in the information processing entity device.
[0132] Through the description of the above embodiments, those skilled in the art can clearly understand that the present application can be implemented by means of software with a necessary general hardware platform, or by hardware. By applying the technical solutions of the present application, compared with the existing manner, the present application uses the facial expression features fused with space-time information to detect the fatigue state of the user within a set time, more considers the influence of the facial expression features on the fatigue state, can add the facial expression features that are helpful to classification to the detection process of the fatigue state, and improves the accuracy of the fatigue detection result.
[0133] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred implementation scenario, and the modules or processes in the drawings are not necessarily essential for implementing the present application. Those skilled in the art can understand that the modules in the device in the implementation scenario can be distributed in the device in the implementation scenario according to the description of the implementation scenario, or can be changed and located in one or more devices different from the implementation scenario. The modules of the above implementation scenario can be combined as one module, or can be further split into multiple sub-modules.
[0134] The above application serial numbers are only for description, and do not represent the advantages and disadvantages of the implementation scenario. The above disclosure is only several specific implementation scenarios of the present application, but the present application is not limited thereto, and any changes that those skilled in the art can think of should fall within the protection scope of the present application.
Claims
1. A method of detecting a fatigue state, characterized by, The method comprises the following steps: obtaining facial key points of a user to be detected; sequentially splicing facial images containing facial key points within a set time window to obtain a facial image sequence; extracting features from the facial image sequence by using a pre-trained feature extraction model to obtain facial expression features fused with space-time information; inputting the facial expression features fused with space-time information into a pre-trained detection model to detect fatigue states of the user to be detected within a set time by using the detection model to obtain a fatigue detection result.
2. The method of claim 1, wherein, The method for obtaining facial key points of a user to be detected comprises the following steps: obtaining an image of a user to be detected output by a photographing device; identifying a facial position in the image of the user to be detected by using a pre-trained facial recognition model to obtain a facial image of the user to be detected; detecting key points of the facial image by using a pre-trained key point detection model to obtain facial key points of the user to be detected.
3. The method of claim 1, wherein, Before the step of sequentially splicing facial images containing facial key points within a set time window to obtain a facial image sequence, the method further comprises the following steps: correcting the facial image of the user to be detected according to the facial key points to obtain a facial image containing a background of a preset area; Correspondingly, the step of sequentially splicing facial images containing facial key points within a set time window to obtain a facial image sequence comprises the following step: sequentially splicing the facial image containing the background of the preset area within a set time window to obtain a facial image sequence.
4. The method of claim 3, wherein, The step of correcting the facial image of the user to be detected according to the facial key points to obtain a facial image containing a background of a preset area comprises the following steps: affine transforming the facial key points by using a pre-set average face vector to calculate an affine transformation matrix of the facial key points; applying the affine transformation matrix of the facial key points to the facial image of the user to be detected to correct the facial image of the user to be detected to obtain a facial image containing a background of a preset area.
5. The method of claim 4, wherein, The step of applying the affine transformation matrix of the facial key points to the facial image of the user to be detected to correct the facial image of the user to be detected to obtain a facial image containing a background of a preset area comprises the following steps: determining matrix variation parameters of different radial variation types of the facial image of the user to be detected according to the affine transformation matrix of the facial key points; combining the matrix variation parameters of different radial variation types in any order to correct the facial image of the user to be detected by using the combined matrix variation parameters to obtain a facial image containing a background of a preset area.
6. The method according to any one of claims 1-5, characterized in that, The step of inputting the facial expression features fused with space-time information into a pre-trained detection model to detect fatigue states of the user to be detected within a set time by using the detection model to obtain a fatigue detection result comprises the following steps: processing the facial expression features fused with space-time information to obtain discrete features of the facial image, wherein the discrete features comprise features obtained by dividing the facial image features in space or time. The discrete features of the face image are input into a pre-trained detection model, and the discrete features of the face image are combined in space or time by the detection model to obtain deep features of a face expression; According to the deep features of the face expression, the fatigue state of the user to be detected within a set time is detected to determine a fatigue detection result.
7. The method of claim 6, wherein, The fatigue state of the user to be detected within a set time is detected according to the deep features of the face expression to determine a fatigue detection result, specifically including: The deep features of the face expression are processed using a tie-pooling and a full connection method to obtain global features of the face expression; According to the fatigue expression set in advance, fatigue scores of the global features of the face expression on different fatigue labels are determined; According to the fatigue scores of the global features of the face expression on different fatigue labels, the fatigue state of the user to be detected within a set time is detected to determine a fatigue detection result.
8. A fatigue state detecting apparatus characterized by comprising: It includes: An acquisition unit is configured to acquire face key points of a user to be detected; A splicing unit is configured to sequentially splice face images containing face key points within a set time window to obtain a face image sequence; An extraction unit is configured to extract features of the face image sequence using a pre-trained feature extraction model to obtain face expression features fused with space-time information; A detection unit is configured to input the face expression features fused with space-time information into a pre-trained detection model to detect the fatigue state of the user to be detected within a set time by the detection model to obtain a fatigue detection result. 9.A computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the computer device is configured to perform the method according to any one of claims 1-8 when the computer program is executed by the processor. The processor executes the computer program to implement the steps of the method of any one of claims 1 to 7.
10. A computer storage medium having stored thereon a computer program, characterized in that The computer program is executed by the processor to implement the steps of the method of any one of claims 1 to 7.