Fatigue driving detection method and system combining face and upper body features of driver

Through the detection method of driver's facial and upper body characteristics, the multi-view camera and three-dimensional convolution module are used to solve the problem of low accuracy of existing fatigue driving detection methods, achieving higher accuracy and robustness.

CN119942503APending Publication Date: 2025-05-06SHANGHAI UNIV OF ENG SCI
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510010013.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-03
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

The current fatigue driving detection methods have low accuracy, poor generalization, low robustness, and poor generalization and accuracy of multi-feature detection methods.

Method used

Using a fatigue detection method combining the driver's facial and upper body features, video data is collected through a multi-view camera, preprocessing and feature extraction, and fatigue detection models of multiple continuous three-dimensional convolution modules are constructed.

Benefits of technology

It improves the accuracy and generalization of fatigue driving detection, enhances the robustness of the model, and makes the detection results more applicable to the production environment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119942503A_ABST
    Figure CN119942503A_ABST
Patent Text Reader

Abstract

The invention discloses a fatigue driving detection method and system combining face and upper body features of a driver, and the method comprises the steps: collecting and preprocessing a face video sequence and an upper body video sequence, and obtaining processed data; based on the processed data, extracting a facial feature vector and a feature vector of the upper body activity; and inputting the extracted facial feature vector and the feature vector of the upper body activity into a constructed fatigue detection model to complete the detection of fatigue driving. The invention designs a method for combining the upper body features and the facial features of the driver, and the upper body features only retain features related to fatigue actions of the driver. Meanwhile, the invention further designs a video classification method which only extracts part of features and is not large in parameter quantity and has the capability of being applied to a production environment.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of fatigue detection, and in particular to a fatigue driving detection method and system combining facial and upper body features of a driver. Background Art

[0002] Fatigue driving refers to the mental and physical dysfunction of the driver after long-term continuous driving. If the driver does not sleep well at night, even a short drive may lead to fatigue driving. Driving fatigue will affect the driver's concentration, observation, thinking judgment, and decision-making. As a common safety hazard, the potential risks of fatigue driving cannot be ignored. Drivers are prone to fatigue when driving for a long time, especially on highways or at night, which not only affects their reaction speed and judgment ability, but may also lead to serious traffic accidents. Therefore, the research and development of fatigue detection and warning technology in fatigue driving is particularly important.

[0003] At present, most fatigue driving detection methods build models based on a single feature, but driver fatigue is the result of the interaction of multiple factors, and these factors have complex relationships. This single detection method leads to low accuracy, poor generalization, and low robustness of the proposed fatigue driving detection method. Among the multi-feature detection methods, there are few methods that combine the driver's fatigue actions, and the fusion of multiple features is often only performed through simple weighted fusion, which has poor generalization and accuracy. Summary of the invention

[0004] The purpose of the present invention is to provide a fatigue detection method and system combining the driver's upper body movements and facial features, which are applied to multi-view cameras to improve the problems of existing fatigue detection methods such as single data source and low generalization.

[0005] To achieve the above object, the present invention provides a method for detecting fatigue driving by combining facial and upper body features of a driver, the steps comprising:

[0006] Collecting facial video sequences and upper body video sequences and performing preprocessing to obtain processed data;

[0007] extracting facial feature vectors and upper body activity feature vectors based on the processed data;

[0008] The extracted facial feature vectors and the feature vectors of upper body activities are input into the constructed fatigue detection model to complete the detection of fatigue driving; the constructed fatigue detection model includes: multiple continuous three-dimensional convolution modules, the first module consists of a 3D convolution layer, a batch normalization layer, a ReLU activation function and a 3D maximum pooling layer, wherein the convolution kernel size is 3×3×3, the number of input and output channels are 128 and 64 respectively, and the pooling operation uses a 2×2×2 kernel.

[0009] Preferably, the method for collecting and preprocessing a facial video sequence includes: collecting an original video file of the front of the driver through a camera at the dashboard position, and preprocessing the video using the OpenCV library; first, by converting the RGB video frame sequence into a grayscale image; then, using the cv2.CascadeClassifier function of OpenCV to perform face detection on the grayscale image, wherein the parameters are set to scaleFactor=1.1 and minNeighbors=4 to capture the driver's facial area.

[0010] Preferably, the method for collecting and preprocessing the upper body video sequence includes: collecting the original video file of the driver's side through a camera at the right A-pillar position, and using cv2.resize of the OpenCV library to adjust the image resolution to 224*224, and the shape of the video frame is adjusted to a batch of 1, a number of image channels of 3, a continuous number of frames of 16, and a resolution of 224*224; then using the Pose module of Mediapipe to detect the key points of the driver's upper body, setting the model complexity of the Mediapipe Pose module to 1, and setting the detection confidence parameter and the tracking confidence parameter to 0.3 to improve the detection efficiency; for each frame of the video image, using the process() method of the Pose module to detect the key points, and extracting 17 key points of the driver's upper body; for frames where no key points are detected, filling all zero coordinates to ensure the integrity of the data; finally, stacking the key points of each frame in chronological order to form a key point sequence output.

[0011] Preferably, the method for extracting facial feature vectors includes: receiving a preprocessed video frame sequence, decomposing the video frames into single-frame images along the time dimension, normalizing each frame and converting it into a standardized tensor sequence; then, inputting the standardized single-frame images one by one into a fine-tuned VGGFace model; and extracting output features by registering the forward propagation hook function of the penultimate fully connected layer of the model.

[0012] Preferably, the method for extracting the feature vector of the upper body activity includes: converting the extracted key point data into a Gaussian heat map related to the driver's activity; splitting the video frame tensor into single frames in the time dimension, and resizing each frame to 224×224 through torchvision.transforms, and then recombining the adjusted frames into a processed tensor, through model reasoning, feature map; combining the extracted Gaussian heat map and the feature map to extract the feature vector of the upper body activity:

[0013]

[0014] in, represents the processed body part feature map vector; c′ represents the channel index of the feature map; i represents the height direction feature map index; j represents the width direction feature map index; F c ′ represents the value of the c′th channel of the feature map at the pixel position (i, j); h′ represents the height of the feature map; w′ represents the width of the feature map; Represents the value of heatmap channel n at pixel position (i, j).

[0015] The present invention also provides a fatigue driving detection system combining the driver's facial and upper body features, the system is used to implement the above method, including: a collection module, an extraction module and a detection module;

[0016] The acquisition module is used to acquire facial video sequences and upper body video sequences and perform preprocessing to obtain processed data;

[0017] The extraction module is used to extract facial feature vectors and upper body activity feature vectors based on the processed data;

[0018] The detection module is used to input the extracted facial feature vectors and the feature vectors of upper body activities into the constructed fatigue detection model to complete the detection of fatigue driving; the constructed fatigue detection model includes: multiple continuous three-dimensional convolution modules, the first module consists of a 3D convolution layer, a batch normalization layer, a ReLU activation function and a 3D maximum pooling layer, wherein the convolution kernel size is 3×3×3, the number of input and output channels are 128 and 64 respectively, and the pooling operation uses a 2×2×2 kernel.

[0019] Compared with the prior art, the present invention has the following beneficial effects:

[0020] The present invention designs a method combining the driver's upper body features and facial features, and the upper body features only retain features related to the driver's fatigue movements. At the same time, the present invention also designs a video classification method, which only extracts some features and has a small number of parameters, and has the ability to be applied in a production environment. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] In order to more clearly illustrate the technical solution of the present invention, the following briefly introduces the drawings required for use in the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.

[0022] Figure 1 A schematic diagram of a method flow chart of an embodiment of the present invention;

[0023] Figure 2 A schematic diagram of the Vggface model structure of an embodiment of the present invention;

[0024] Figure 3 A schematic diagram of fatigue model detection according to an embodiment of the present invention;

[0025] Figure 4 The figure is a flowchart of a method according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0027] In order to make the above-mentioned objects, features and advantages of the present invention more obvious and easy to understand, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments.

[0028] Embodiment 1

[0029] like Figure 1 FIG. 1 is a schematic diagram of the method flow of this embodiment, and the steps include:

[0030] S1. Collect facial video sequences and upper body video sequences and perform preprocessing to obtain processed data.

[0031] In this embodiment, two cameras are arranged to obtain the driver's facial video and upper body video, wherein one camera is arranged at the vehicle dashboard and the other camera is fixed to the right A-pillar.

[0032] (1) Acquisition and processing of driver’s facial video

[0033] The original video file of the driver's front is collected by the camera at the dashboard position, and the video is preprocessed using the OpenCV library. First, the RGB video frame sequence is converted into a grayscale image. Then, the grayscale image is used for face detection using the cv2.CascadeClassifier function of OpenCV, where the parameters are set to scaleFactor=1.1 and minNeighbors=4 to capture the driver's facial area. Since the driver's face in the front video data is centered and has obvious features, a simple and lightweight cascade Haar face classifier is used to achieve efficient detection. The detected facial area is further adjusted to a resolution of 224×224 using the cv2.resize function to adapt to the subsequent model input requirements.

[0034] (2) Acquisition and processing of driver upper body video

[0035] The original video file of the driver's side is collected by the camera at the right A-pillar position, and the image resolution is adjusted to 224*224 using cv2.resize of the OpenCV library. The shape of the video frame is adjusted to (1, 3, 16, 224, 224), where 1 is the batch, 3 is the number of image channels, 16 is the number of consecutive frames, and 224*224 is the resolution. Next, the Pose module of Mediapipe is used to detect the key points of the driver's upper body. The model complexity of the Mediapipe Pose module is set to 1, and the detection confidence parameter and tracking confidence parameter are both set to 0.3 to improve the detection efficiency. For each frame of the video image, the process() method of the Pose module is used to detect the key points and extract the 17 key points of the driver's upper body. For frames where no key points are detected, fill all zero coordinates to ensure the integrity of the data. Finally, the key points of each frame are stacked in time order to form a key point sequence output with a dimension of (1, 16, 17, 2).

[0036] Through the above steps, the processed data is obtained.

[0037] S2. Based on the processed data, extract the facial feature vector and the feature vector of the upper body activity.

[0038] (1) Extraction of driver’s facial feature vector

[0039] In order to extract high-dimensional features related to fatigue, this embodiment uses the VGGFace model (such as Figure 2 Specifically, all existing weight parameters except the last three fully connected layers are frozen, and only the parameters of the last three fully connected layers are trained. The Adam optimizer is used in the fine-tuning process, and the learning rate is set to 0.0001. After training for 50 epochs, an optimized model suitable for fatigue feature extraction is obtained. The structural diagram of the model is shown in the appendix.

[0040] A preprocessed video frame sequence is received, and its shape is (1, 3, 16, 224, 224), where 16 represents the number of frames in the time dimension and 224×224 is the resolution of each frame. The video frame is decomposed into single-frame images along the time dimension, and each frame is normalized and converted into a standardized tensor sequence. Subsequently, the standardized single-frame images are input into the fine-tuned VGGFace model one by one. To obtain the deep feature output of the model, the output features of the layer are extracted by registering the forward propagation hook function (Forward Hook) of the penultimate fully connected layer of the model. The extracted single-frame features are one-dimensional vectors with a shape of (1, 4096). After stacking the features of all frames in time order, a feature sequence with a shape of (1, 16, 4096) is formed. In order to splice with the feature vector of the body part, the sequence is transformed in dimension, and the spatial resolution of the feature map is increased by upsampling. Specifically, the view function is used to reshape the features, and the interpolate function is used to adjust the spatial size to (112, 112). Finally, the generated video sequence feature shape is (1, 16, 64, 112, 112), which is used to characterize the high-dimensional features of facial video sequences in deep neural networks.

[0041] (2) Extraction of feature vectors of driver’s upper body movements

[0042] The key point data obtained in step S1 is converted into a Gaussian heat map related to the driver's activity. The key point data extracted from the video sequence is received in the form of a four-dimensional array with a shape of (1, 16, 17, 2), where 1 is the batch size, 16 is the number of time frames, 17 is the number of key points, and the two-dimensional coordinates (x, y) represent the position of each key point. Set the original image resolution ori_shape and the target heat map resolution heatmap_shape, and assume that both the input image and the output heat map are square. According to the ratio of the target heat map resolution to the original image resolution, calculate the coordinate scaling factor factor = heatmap_shape / ori_shape. Using this ratio, the normalized coordinates of the key points are mapped to the pixel coordinate system of the heat map. Initialize a zero matrix for each frame image and each key point in each batch to store the Gaussian heat map, with a shape of (B, F, J, heatmap_shape, heatmap_shape), and the meanings of the parameters are batch, number of time frames, number of key points, and heat map shape.

[0043] Next is the calculation method of Gaussian heat map. First, define the center point of Gaussian distribution (μ x , μ y ) and standard deviation σ = 1 are used to store the Gaussian heat map and construct a Gaussian distribution function with the center point as the core:

[0044]

[0045] Where G(x, y) represents the Gaussian distribution value at the point (x, y); μ x and μ y represents the coordinates of the center point of the Gaussian distribution; σ represents the standard deviation of the Gaussian distribution.

[0046] Next, the influence range of the Gaussian distribution is determined based on the center coordinates of the key points (center x, center y) and the size of the heat map. By calculating the effective boundary:

[0047] x0=max(0,center x-δ·σ), x1=min(W-1,center x+δ·σ),

[0048] y0=max(0,centery-δ·σ), y1=min(H-1,centery+δ·σ),

[0049] Among them, x0 and x1 represent the effective boundaries of the Gaussian distribution on the x-axis; y0 and y1 represent the effective boundaries of the Gaussian distribution on the y-axis; center x and center y represent the center coordinates of the key points; δ represents a constant used to limit the distribution range; W and H represent the width and height of the heat map.

[0050] Next, within the effective range of the Gaussian distribution, the square sum of the offset of each pixel is generated by gridding calculation, and the pixel value is calculated according to the Gaussian distribution formula. To avoid the distribution affecting meaningless distant pixels, the threshold th = ln (100) is set, and the Gaussian distribution is truncated so that the pixel values ​​beyond the threshold range are set to 0.

[0051] Finally, the heat map value is updated by comparing the maximum value of each element with the corresponding area of ​​the current heat map:

[0052] heatmap[y0:y1+1,x0:x1+1]=max(arr heatarr exp),

[0053] Among them, heatmap represents the final heat map; arr heat and arr exp represent the pixel values ​​of the area corresponding to the heat map and the pixel value array calculated according to the Gaussian distribution respectively; y0:y1+1 and x0:x1+1 represent the area range specified to be updated on the heat map.

[0054] A heat map representation of the key points of the driver's upper body is obtained, with a dimension of (1, 16, 17, 112, 112). The obtained action-related feature vector is combined with the heat map.

[0055] The original video data of the driver's side is used to highlight the image feature information most relevant to the driver's action, and the pre-trained R(2+1)D model is used to extract the feature vector. Specifically, first use pytorch's torch.hub.load to load the model. The corresponding model is r2plusld_3432_ig65m. The model uses a 34-layer convolutional network and is pre-trained based on the IG65M dataset. The names of the model layers are: stem, layer1, layer2, layer3, layer4, avgpool, fc. The model input is 32 frames of video and the output is the prediction results of 359 categories. A forward hook (ForwardHook) is registered in the layer1 layer of the model to extract the feature output of the intermediate layer during the forward propagation process. First, the video frame tensor is split into single frames in the time dimension, and each frame is resized to 224×224 through torchvision.transforms. Then the resized frames are reassembled into the processed tensor. Through model reasoning, the feature map of the layerl layer is obtained, which represents the shallow features of the input video frame at this layer. The feature map dimension is (1, 16, 64.112.112).

[0056] The extracted Gaussian heat map and feature map of driver activity are combined and calculated using the following formula:

[0057]

[0058] in, represents the processed body part feature map vector; c represents the channel index of the feature map; i represents the height direction feature map index; j represents the width direction feature map index; F c ′ represents the value of the c′th channel of the feature map at the pixel position (i, j); h′ represents the height of the feature map; w′ represents the width of the feature map; Represents the value of heatmap channel n at pixel position (i, j).

[0059] Finally, we get a feature vector representing the characteristics of the driver's body parts, which is highly correlated with the driver's actions and only contains the parts most relevant to the actions.

[0060] S3. Input the extracted facial feature vector and the feature vector of upper body activity into the constructed fatigue detection model to complete the detection of fatigue driving.

[0061] The facial and body feature vectors obtained by the feature extraction part are concatenated according to the channel dimension, and finally a complete feature sequence with a dimension of (1, 16, 128, 112, 112) is obtained. The concatenated data is input into the 3DCNN model. The calculation formula of 3D convolution is as follows:

[0062]

[0063] in, represents the output of the kth convolution kernel at time position t and spatial position (h, w), X is the input feature, w is the convolution kernel weight, b is the bias, and K t , K h , K w are the convolution kernel sizes in time and space dimensions respectively.

[0064] This embodiment designs a fatigue detection model based on a three-dimensional convolutional neural network (3DCNN) (such as Figure 3 As shown in the figure), it is used to extract deep spatiotemporal features from the spliced ​​spatiotemporal feature data and realize the final judgment of fatigue state. The model receives input data of dimension (1, 16, 128, 112, 112), including 16 frames of time series features, 128 channel features, and the height and width of each frame image are 112 pixels. The network contains multiple continuous three-dimensional convolution modules. The first module consists of a 3D convolution layer, a batch normalization layer, a ReLU activation function, and a 3D maximum pooling layer. The convolution kernel size is 3×3×3, the number of input and output channels are 128 and 64 respectively, and the pooling operation uses a 2×2×2 kernel; after the first module, the feature dimension is reduced to (1, 8, 64, 56, 56). The parameters of subsequent convolution modules are similar, and the network feature extraction capability is gradually deepened. The output dimension of the second module is (1, 4, 128, 28, 28), and the output dimension of the third module is (1, 2, 256, 14, 14). Finally, the spatial dimension is compressed through the global average pooling layer to form a one-dimensional vector, and three fully connected layers are connected to realize the classification and judgment of fatigue state. The network structure diagram is shown in the appendix. This network structure can effectively capture the changes in spatiotemporal features through hierarchical feature extraction and global information integration. The detailed process of this embodiment is as follows Figure 4 shown.

[0065] Embodiment 2

[0066] The present invention also provides a fatigue driving detection system that combines the driver's facial and upper body features, including: an acquisition module, an extraction module and a detection module; the acquisition module is used to acquire facial video sequences and upper body video sequences and perform preprocessing to obtain processed data; the extraction module is used to extract facial feature vectors and feature vectors of upper body activities based on the processed data; the detection module is used to input the extracted facial feature vectors and feature vectors of upper body activities into a constructed fatigue detection model to complete the detection of fatigue driving.

[0067] The following will explain in detail how the present invention solves technical problems in real life in conjunction with this embodiment.

[0068] Firstly, the facial video sequence and the upper body video sequence are collected by the collection module and preprocessed to obtain the processed data.

[0069] The acquisition module of this embodiment includes two cameras for acquiring facial video and upper body video of the driver, wherein one camera is arranged at the vehicle dashboard, and the other camera is fixed on the right A-pillar.

[0070] (1) Acquisition and processing of driver’s facial video

[0071] The original video file of the driver's front is collected by the camera at the dashboard position, and the video is preprocessed using the OpenCV library. First, the RGB video frame sequence is converted into a grayscale image. Then, the grayscale image is used for face detection using the cv2.CascadeClassifier function of OpenCV, where the parameters are set to scaleFactor=1.1 and minNeighbors=4 to capture the driver's facial area. Since the driver's face in the front video data is centered and has obvious features, a simple and lightweight cascade Haar face classifier is used to achieve efficient detection. The detected facial area is further adjusted to a resolution of 224×224 using the cv2.resize function to adapt to the subsequent model input requirements.

[0072] (2) Acquisition and processing of driver upper body video

[0073] The original video file of the driver's side is collected by the camera at the right A-pillar position, and the image resolution is adjusted to 224*224 using cv2.resize of the OpenCV library. The shape of the video frame is adjusted to (1, 3, 16, 224, 224), where 1 is the batch, 3 is the number of image channels, 16 is the number of consecutive frames, and 224*224 is the resolution. Next, the Pose module of Mediapipe is used to detect the key points of the driver's upper body. The model complexity of the Mediapipe Pose module is set to 1, and the detection confidence parameter and tracking confidence parameter are both set to 0.3 to improve the detection efficiency. For each frame of the video image, the process() method of the Pose module is used to detect the key points and extract the 17 key points of the driver's upper body. For frames where no key points are detected, fill all zero coordinates to ensure the integrity of the data. Finally, the key points of each frame are stacked in time order to form a key point sequence output with a dimension of (1, 16, 17, 2).

[0074] Through the above process, the processed data is obtained.

[0075] The extraction module then extracts the facial feature vector and the feature vector of the upper body activity based on the processed data.

[0076] (1) Extraction of driver’s facial feature vector

[0077] In order to extract high-dimensional features related to fatigue, this embodiment uses the VGGFace model (such as Figure 2 Specifically, all existing weight parameters except the last three fully connected layers are frozen, and only the parameters of the last three fully connected layers are trained. The Adam optimizer is used in the fine-tuning process, and the learning rate is set to 0.0001. After training for 50 epochs, an optimized model suitable for fatigue feature extraction is obtained. The structural diagram of the model is shown in the appendix.

[0078] A preprocessed video frame sequence is received, and its shape is (1, 3, 16, 224, 224), where 16 represents the number of frames in the time dimension and 224×224 is the resolution of each frame. The video frame is decomposed into single-frame images along the time dimension, and each frame is normalized and converted into a standardized tensor sequence. Subsequently, the standardized single-frame images are input into the fine-tuned VGGFace model one by one. To obtain the deep feature output of the model, the output features of the layer are extracted by registering the forward propagation hook function (Forward Hook) of the penultimate fully connected layer of the model. The extracted single-frame features are one-dimensional vectors with a shape of (1, 4096). After stacking the features of all frames in time order, a feature sequence with a shape of (1, 16, 4096) is formed. In order to splice with the feature vector of the body part, the sequence is transformed in dimension, and the spatial resolution of the feature map is increased by upsampling. Specifically, the view function is used to reshape the features, and the interpolate function is used to adjust the spatial size to (112, 112). Finally, the generated video sequence feature shape is (1, 16, 64, 112, 112), which is used to characterize the high-dimensional features of facial video sequences in deep neural networks.

[0079] (2) Extraction of feature vectors of driver’s upper body movements

[0080] The key point data obtained in step S1 is converted into a Gaussian heat map related to the driver's activity. The key point data extracted from the video sequence is received in the form of a four-dimensional array with a shape of (1, 16, 17, 2), where 1 is the batch size, 16 is the number of time frames, 17 is the number of key points, and the two-dimensional coordinates (x, y) represent the position of each key point. Set the original image resolution ori_shape and the target heat map resolution heatmap_shape, and assume that both the input image and the output heat map are square. According to the ratio of the target heat map resolution to the original image resolution, calculate the coordinate scaling factor factor = heatmap_shape / ori_shape. Using this ratio, the normalized coordinates of the key points are mapped to the pixel coordinate system of the heat map. Initialize a zero matrix for each frame image and each key point in each batch to store the Gaussian heat map, with a shape of (B, F, J, heatmap_shape, heatmap_shape), and the meanings of the parameters are batch, number of time frames, number of key points, and heat map shape.

[0081] Next is the calculation method of Gaussian heat map. First, define the center point of Gaussian distribution (μ x , μ y ) and standard deviation σ = 1 are used to store the Gaussian heat map and construct a Gaussian distribution function with the center point as the core:

[0082]

[0083] Where G(x, y) represents the Gaussian distribution value at the point (x, y); μ x and μ y represents the coordinates of the center point of the Gaussian distribution; σ represents the standard deviation of the Gaussian distribution.

[0084] Next, the influence range of the Gaussian distribution is determined based on the center coordinates of the key points (center x, center y) and the size of the heat map. By calculating the effective boundary:

[0085] x0=max(0,center x-δ·σ), x1=min(W-1,center x+δ·σ),

[0086] y0=max(0,center y-δ·σ), y1=min(H-1,center y+δ·σ),

[0087] Among them, x0 and x1 represent the effective boundaries of the Gaussian distribution on the x-axis; y0 and y1 represent the effective boundaries of the Gaussian distribution on the y-axis; center x and center y represent the center coordinates of the key points; δ represents a constant used to limit the distribution range; W and H represent the width and height of the heat map.

[0088] Next, within the effective range of the Gaussian distribution, the square sum of the offset of each pixel is generated by gridding calculation, and the pixel value is calculated according to the Gaussian distribution formula. To avoid the distribution affecting meaningless distant pixels, the threshold th = ln (100) is set, and the Gaussian distribution is truncated so that the pixel values ​​beyond the threshold range are set to 0.

[0089] Finally, the heat map value is updated by comparing the maximum value of each element with the corresponding area of ​​the current heat map:

[0090] heatmap[y0:y1+1,x0:x1+1]=max(arr heat, arr exp),

[0091] Among them, heatmap represents the final heat map; arr heat and arr exp represent the pixel values ​​of the area corresponding to the heat map and the pixel value array calculated according to the Gaussian distribution respectively; y0:y1+1 and x0:x1+1 represent the area range specified to be updated on the heat map.

[0092] A heat map representation of the key points of the driver's upper body is obtained, with a dimension of (1, 16, 17, 112, 112). The obtained action-related feature vector is combined with the heat map.

[0093] The original video data of the driver's side is used to highlight the image feature information most relevant to the driver's action, and the pre-trained R(2+1)D model is used to extract the feature vector. Specifically, first use pytorch's torch.hub.load to load the model. The corresponding model is r2plusld_34_32_ig65m. The model uses a 34-layer convolutional network and is pre-trained based on the IG65M dataset. The names of the model layers are: stem, layer1, layer2, layer3, layer4, avgpool, fc. The model input is 32 frames of video, and the output is the prediction results of 359 categories. A forward hook (ForwardHook) is registered in the layer1 layer of the model to extract the feature output of the intermediate layer during the forward propagation process. First, the video frame tensor is split into single frames in the time dimension, and each frame is resized to 224×224 through torchvision.transforms. Then the resized frames are reassembled into the processed tensor. Through model inference, the feature map of layer1 is obtained, which represents the shallow features of the input video frame at this layer. The feature map dimension is (1, 16, 64.112.112).

[0094] The extracted Gaussian heat map and feature map of driver activity are combined and calculated using the following formula:

[0095]

[0096] in, represents the processed body part feature map vector; c represents the channel index of the feature map; i represents the height direction feature map index; j represents the width direction feature map index; F c ′ represents the value of the c′th channel of the feature map at the pixel position (i, j); h′ represents the height of the feature map; w′ represents the width of the feature map; Represents the value of heatmap channel n at pixel position (i, j).

[0097] Finally, we get a feature vector representing the characteristics of the driver's body parts, which is highly correlated with the driver's actions and only contains the parts most relevant to the actions.

[0098] Finally, the detection module inputs the extracted facial feature vector and the feature vector of upper body activity into the constructed fatigue detection model to complete the detection of fatigue driving.

[0099] The facial and body feature vectors obtained by the feature extraction part are concatenated according to the channel dimension, and finally a complete feature sequence with a dimension of (1, 16, 128, 112, 112) is obtained. The concatenated data is input into the 3DCNN model. The calculation formula of 3D convolution is as follows:

[0100]

[0101] in, represents the output of the kth convolution kernel at time position t and spatial position (h, w), X is the input feature, w is the convolution kernel weight, b is the bias, and K t , K h , K w are the convolution kernel sizes in time and space dimensions respectively.

[0102] This embodiment designs a fatigue detection model based on a three-dimensional convolutional neural network (3DCNN) (such as Figure 3As shown in the figure), it is used to extract deep spatiotemporal features from the spliced ​​spatiotemporal feature data and realize the final judgment of fatigue state. The model receives input data of dimension (1, 16, 128, 112, 112), including 16 frames of time series features, 128 channel features, and the height and width of each frame image are 112 pixels. The network contains multiple continuous three-dimensional convolution modules. The first module consists of a 3D convolution layer, a batch normalization layer, a ReLU activation function, and a 3D maximum pooling layer. The convolution kernel size is 3×3×3, the number of input and output channels are 128 and 64 respectively, and the pooling operation uses a 2×2×2 kernel; after the first module, the feature dimension is reduced to (1, 8, 64, 56, 56). The parameters of subsequent convolution modules are similar, and the network feature extraction capability is gradually deepened. The output dimension of the second module is (1, 4, 128, 28, 28), and the output dimension of the third module is (1, 2, 256, 14, 14). Finally, the spatial dimension is compressed through the global average pooling layer to form a one-dimensional vector, and three fully connected layers are connected to realize the classification and judgment of fatigue state. The network structure diagram is shown in the appendix. This network structure can effectively capture the changes in spatiotemporal features through hierarchical feature extraction and global information integration. The detailed process of this embodiment is as follows Figure 4 shown.

[0103] The embodiments described above are only descriptions of the preferred embodiments of the present invention and are not intended to limit the scope of the present invention. Without departing from the design spirit of the present invention, various modifications and improvements made to the technical solutions of the present invention by ordinary technicians in this field should all fall within the protection scope determined by the claims of the present invention.

Claims

1. A method for detecting fatigue driving by combining the facial and upper body features of a driver, characterized in that the steps include: Collecting facial video sequences and upper body video sequences and performing preprocessing to obtain processed data; extracting facial feature vectors and upper body activity feature vectors based on the processed data; The extracted facial feature vectors and the feature vectors of upper body activities are input into the constructed fatigue detection model to complete the detection of fatigue driving; the constructed fatigue detection model includes: several continuous three-dimensional convolution modules, the first module consists of a 3D convolution layer, a batch normalization layer, a ReLU activation function and a 3D maximum pooling layer, wherein the convolution kernel size is 3×3×3, the number of input and output channels are 128 and 64 respectively, and the pooling operation uses a 2×2×2 kernel.

2. The fatigue driving detection method combining the driver's facial and upper body features according to claim 1 is characterized in that: The method for collecting and preprocessing facial video sequences includes: collecting the original video file of the front of the driver through a camera at the dashboard position, and preprocessing the video using the OpenCV library; first, by converting the RGB video frame sequence into a grayscale image; then, using the cv2.CascadeClassifier function of OpenCV to perform face detection on the grayscale image, where the parameters are set to scaleFactor=1.1 and minNeighbors=4 to capture the driver's facial area.

3. The fatigue driving detection method combining the driver's facial and upper body features according to claim 1 is characterized in that: The method for collecting and preprocessing the upper body video sequence includes: collecting the original video file of the driver's side through the camera at the right A-pillar position, and using the cv2.resize of the OpenCV library to adjust the image resolution to 224*224, and the shape of the video frame is adjusted to a batch of 1, an image channel number of 3, a continuous frame number of 16, and a resolution of 224*224; then using the Pose module of Mediapipe to detect the key points of the driver's upper body, setting the model complexity of the Mediapipe Pose module to 1, and setting the detection confidence parameter and the tracking confidence parameter to 0.3 to improve the detection efficiency; for each frame of the video image, using the process() method of the Pose module to detect the key points, and extracting 17 key points of the driver's upper body; for frames where no key points are detected, fill all zero coordinates to ensure the integrity of the data; finally, stacking the key points of each frame in chronological order to form a key point sequence output.

4. The fatigue driving detection method combining the driver's facial and upper body features according to claim 2 is characterized in that: The method for extracting facial feature vectors includes: receiving a preprocessed video frame sequence, decomposing the video frames into single-frame images along the time dimension, normalizing each frame and converting it into a standardized tensor sequence; then, inputting the standardized single-frame images one by one into a fine-tuned VGGFace model; and extracting output features by registering the forward propagation hook function of the penultimate fully connected layer of the model.

5. The fatigue driving detection method combining the driver's facial and upper body features according to claim 3 is characterized in that: The method for extracting the feature vector of upper body activity includes: converting the extracted key point data into a Gaussian heat map related to the driver's activity; splitting the video frame tensor into single frames in the time dimension, and resizing each frame to 224×224 through torchvision.transforms, and then recombining the resized frames into a processed tensor, through model reasoning, feature map; combining the extracted Gaussian heat map and feature map to extract the feature vector of upper body activity: in, represents the processed body part feature map vector; c′ represents the channel index of the feature map; i represents the height direction feature map index; j represents the width direction feature map index; F c′ represents the value of the c′th channel of the feature map at the pixel position (i, j); h′ represents the height of the feature map; w′ represents the width of the feature map; Represents the value of heatmap channel n at pixel position (i, j).

6. A fatigue driving detection system combining the driver's facial and upper body features, the system is used to implement the method according to any one of claims 1 to 5, characterized in that: include: Acquisition module, extraction module and detection module; The acquisition module is used to acquire facial video sequences and upper body video sequences and perform preprocessing to obtain processed data; The extraction module is used to extract facial feature vectors and upper body activity feature vectors based on the processed data; The detection module is used to input the extracted facial feature vector and the feature vector of the upper body activity into the constructed fatigue detection model to complete the detection of fatigue driving; the constructed fatigue detection model includes: a number of continuous three-dimensional convolution modules, the first module consists of a 3D convolution layer, a batch normalization layer, a ReLU activation function and a 3D maximum pooling layer, wherein the convolution kernel size is 3×3×3, the number of input and output channels are 128 and 64 respectively, and the pooling operation uses a 2×2×2 kernel.

Citation Information

Cited By

  • Non-contact athlete physiological status assessment method and system based on feature fusion and medium

    CN120616480A