A Fatigue Driving Behavior Recognition Method for 3D Point Cloud Image Data
The 3D point cloud image processing method using PVFTN and MSFLD for fatigue detection improves accuracy and efficiency in identifying fatigue driving states by leveraging adaptive thresholding and facial feature analysis.
Patent Information
- Application Number
- CN202310833987.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-07-10
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2043-07-10
AI Technical Summary
The existing fatigue driving detection methods have problems such as high equipment cost, unstable accuracy, great environmental impact, privacy violations and low detection efficiency, especially when wearing glasses or masks.
The point-element-voxel fusion conversion network PVFTN extracts the characteristics of 3D point cloud image data, and combines the multi-scale facial key point detector MSFLD and the adaptive threshold and statistical threshold to determine the driver's fatigue state.
It improves the accuracy and efficiency of fatigue driving detection, reduces interference to drivers, adapts to different environments and individual differences, and reduces equipment costs.
Smart Images

Figure CN116824558B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of image processing, and particularly to a method for recognizing fatigue driving behavior for 3D point cloud image data. Background Art
[0002] Fatigue driving refers to the phenomenon of psychological and physiological dysfunction of a driver during long-term driving due to excessive mental and physical fatigue. After the driver is fatigued, physiological functions, recognition, and control abilities decline, and they cannot respond to sudden accidents in a timely manner, seriously affecting safe driving. Accurately detecting the fatigue driving state of a driver and issuing a warning in a timely manner is crucial for ensuring traffic safety.
[0003] The detection of fatigue driving behavior has been widely studied in the past 20 years. Fatigue driving detection methods are divided into five categories according to input features: subjective reports, biological, physical, vehicle, and hybrid. The subjective report detection method is based on a questionnaire survey of the driver, and the fatigue level is obtained by analyzing and comparing the main symptoms of the respondents and the number of occurrences of various symptoms. The advantage of this detection method is that there is no invasion problem, and the disadvantage is that the process of measuring fatigue is not synchronized with the driving process and is easily affected by the emotions and physical conditions of the subjects. Detection methods based on driver biometric characteristics usually have high accuracy, but the detection process requires special equipment to measure biological signals, such as electroencephalogram (EEG), electrocardiogram (ECG), electrooculogram (EOG), and surface electromyogram (EMG), etc. The equipment cost is high, and the contact signal acquisition method causes certain interference to the driver, and the acceptance of the driver is low. Detection methods based on vehicle characteristics use indicators such as steering wheel movement and lane deviation for detection. It only needs to obtain vehicle information and will not cause any interference to the driver. However, this detection method is affected by road conditions and driver skills and has a low accuracy. Detection methods based on physical characteristics use image processing technology to monitor the fatigue state of the driver by detecting changes in individual characteristics such as the driver's eyes, mouth, head, and facial expressions. This detection method not only has high accuracy but also has the advantage of non-contact detection. However, this detection method is affected by objective factors such as light, occlusions (glasses, masks, etc.), and vehicle movement, and the detection accuracy varies greatly in different environments, with low robustness. The user's facial data needs to be collected during the detection process, involving user privacy issues. The fatigue detection method based on hybrid features is to fuse various features for fatigue detection.
[0004] In recent years, with the successful application of deep learning in computer vision, the fatigue driving detection method based on the deep learning of visual physical features has become the mainstream direction of current research. Its detection algorithms mainly include the method based on the Recurrent Neural Network (RNN), the method based on the Convolutional Neural Networks (CNN), and the fusion method of RNN and CNN. Generally, the sliding window plus classifier method is adopted. Due to the large amount of input video data, the system takes a long time and the detection efficiency is not high. Moreover, when the driver has a side face, wears glasses or wears a mask, the detection accuracy is not high. Summary of the Invention
[0005] In view of the above problems, the present invention proposes a fatigue driving behavior recognition method for 3D point cloud image data. First, collect the 3D point cloud image video frames containing the driver's facial information after frame processing as the images to be processed; then, use the Point-Voxel Fusion Transformer Network (PVFTN) as the 3D point cloud image data feature extraction model to process the images to be processed and obtain the facial region feature images of the driver. Among them, the Point-Voxel Fusion Transformer Network (PVFTN) includes three Point-Voxel Fusion (PVF) modules, and the PVF module includes a point-voxel branch and a voxel branch; then use the Multi-Scale Facial Landmark Detector (MSFLD) to process the facial region feature images to obtain the coordinates of 23 key points of each frame of the face region image; then calculate the aspect ratios of the left eye, right eye and mouth from the 23 key points of each frame of the face region image, and update the facial fatigue parameter matrix; finally, judge according to the facial fatigue parameter matrix by combining the statistical threshold and the adaptive threshold to obtain the judgment result of the driver's driving state, that is, the judgment of the driver's fatigue driving state in the motion state can be realized.
[0006] According to the first aspect of the present invention, there is provided a fatigue driving behavior recognition method for 3D point cloud image data, including:
[0007] Step 10, collect the 3D point cloud image video frames containing the driver's facial information after frame processing as the images to be processed.
[0008] Step 20, use the Point-Voxel Fusion Transformer Network (PVFTN) as the 3D point cloud image data feature extraction model to process the images to be processed and obtain the facial region feature images.
[0009] Among them, the image to be processed is input into the point-voxel fusion transformation network PVFTN, and is processed by the first PVF module, the second PVF module, and the third PVF module in sequence to obtain the first PVF feature map, the second PVF feature map, and the third PVF feature map. These three feature maps are finally processed by the MLP layer and aggregated to obtain the aggregated multi-scale feature as the facial region feature image output.
[0010] The PVF module includes a point-voxel branch and a voxel branch, which process the input feature F of the PVF module respectively. The point-voxel branch extracts the global feature F in the point-voxel domain. g , and the voxel branch extracts the local feature F in the voxel domain. l , and the output end of the PVF module fuses the global feature F g and the local feature F l , and obtains and outputs the PVF feature map F'.
[0011] Step 30, use the multi-scale facial key point detector MSFLD to process the facial region feature image to obtain the facial key point coordinates.
[0012] Step 40, update the facial fatigue parameter matrix according to the facial key point coordinates.
[0013] Step 50, make a judgment according to the facial fatigue parameter matrix by combining the statistical threshold and the adaptive threshold to obtain the judgment result of the driver's driving state.
[0014] Furthermore, in the fatigue driving behavior recognition method for 3D point cloud image data provided by the present invention, step 20 further includes: the point-voxel branch uses the self-attention mechanism of the embedded relative position representation RPR to process the input feature F to obtain the self-attention feature F in the point-voxel domain sa , and then is processed by the MLP layer and aggregated with the input feature F to obtain the global feature F g , and the calculation formula is:
[0015] F sa = softmax(QK T + B)V
[0016] F g = MLP(F sa ) + F
[0017] where Q, K, and V are the feature matrices generated by the input feature F according to the shared linear transformation, B is the relative position representation deviation matrix, and for any two points p = {x p , y p , z p} and q = {x q , y q , z q}, the relative position from point p to point q represents the deviation B pq = max(-β, min(β, D pq ))), D pq = (x p - x q ) + (y p - y q ) + (z p - z q ), where β is the preset maximum relative distance.
[0018] Furthermore, in the fatigue driving behavior recognition method for 3D point cloud image data provided by the present invention, step 20 further includes: the voxel branch processes the input feature F through a voxelization module, a feature aggregation module, and a de-voxelization module in sequence to obtain the voxel local feature F l . Among them, the voxelization module evenly divides the space of the input feature F in the point voxel domain into multiple volume regions in a non-overlapping manner. For each volume region, the mean value of all feature points in the volume region is calculated as the feature value of the volume region, thereby mapping the input feature F in the point voxel domain to the voxel domain feature O. The feature aggregation module adopts a window self-attention mechanism. First, the space of the voxel domain feature O is evenly divided into multiple window spaces in a non-overlapping manner. In each window space, the voxel domain feature O is processed through the self-attention mechanism to obtain the voxel domain self-attention feature O sa , and then processed through an MLP layer and aggregated with the voxel domain feature O to obtain the voxel domain local feature O g . The de-voxelization module assigns the feature value of each volume region in the voxel domain local feature O g to all feature points in the input feature F within the volume region, thereby obtaining the local feature F l .
[0019] Furthermore, in the fatigue driving behavior recognition method for 3D point cloud image data provided by the present invention, step 30 further includes: the facial region feature image is input into the multi-scale facial key point detector MSFLD, and is processed through the first convolutional layer, the depthwise separable convolutional layer, 4 inverted residual blocks, the second convolutional layer, the third convolutional layer, and 5 fully connected layers with different scales in sequence to obtain the facial key point coordinates.
[0020] Furthermore, in the fatigue driving behavior recognition method for 3D point cloud image data provided by the present invention, step 40 further includes: calculating the left eye aspect ratio EAR li , the right eye aspect ratio EAR ri , and the mouth aspect ratio MAR i , and combining with the time parameter t i of the corresponding current i-th frame 3D point cloud image video frame to update the facial fatigue parameter matrix:
[0021] Further, in the fatigue driving behavior recognition method for 3D point cloud image data provided by the present invention, step 50 further includes: the statistical threshold is the statistically obtained left eye aspect ratio threshold EAR based on data of different driving behavior types and different drivers in the dataset lst and the statistically obtained right eye aspect ratio threshold EAR rst and the statistically obtained mouth aspect ratio threshold MAR st . The adaptive threshold is the adaptively obtained left eye aspect ratio threshold EAR based on the data of the first p frames of the current test video lat and the adaptively obtained right eye aspect ratio threshold EAR rat and the adaptively obtained mouth aspect ratio threshold MAR at . When the data of the i-th frame simultaneously satisfies the condition EAR li <max(EAR lst , EAR lat ) or satisfies the condition EAR ri <max(EAR rst , EAR rat ) or satisfies the condition MAR i >min(MAR st , MAR at ), it is determined that the driving state of the driver is fatigued, otherwise it is normal
[0022] According to a second aspect of the present invention, there is provided a computer device, including: a memory for storing instructions; and a processor for calling the instructions stored in the memory to execute the fatigue driving behavior recognition method for 3D point cloud image data according to the first aspect
[0023] According to a third aspect of the present invention, there is provided a computer-readable storage medium, characterized in that it stores instructions which, when executed by a processor, execute the fatigue driving behavior recognition method for 3D point cloud image data according to the first aspect
[0024] Compared with the prior art, the above technical solution conceived by the present invention has at least the following beneficial effects
[0025] (1) Enter the 3D point cloud image for fatigue driving behavior recognition. The model for facial feature image extraction, the Point-Voxel Fusion Transformation Network (PVFTN), takes N points of the 3D point cloud image as input and feeds them into 3 stacked PVF blocks. The maximum pooling and repetition operators are used to extract effective global features representing the entire point cloud. And an MLP layer is used to aggregate multi-scale features. The PVF block consists of two branches. The upper branch is voxel-based and used to aggregate local features, while the lower branch is point-voxel-based and used to capture global features. By effectively fusing the two branches, complementary feature information can be provided, improving the detection accuracy of obtaining the driver's face area image.
[0026] (2) By replacing all the original bottleneck layers with inverted residual blocks and increasing the number of multi-scale fully connected layers on the basis of the original PFLD face key point detector, a face detector (MSFLD) with embedded multi-scale fully connected layers is constructed to increase the network depth, reduce the number of parameters, reduce the computational amount while ensuring the extraction of feature information, so as to keep the detection speed of the face detector unchanged and improve the positioning accuracy of key points.
[0027] (3) In the driving behavior decision-making, a driving behavior decision-making method combining an adaptive threshold and a statistical threshold is proposed. The adaptive threshold is obtained by calculating the eye aspect ratio and mouth aspect ratio of the first p frames of each test video and then taking the average. The adaptive threshold is dynamically changed, effectively solving the problem that the eye aspect ratio and mouth aspect ratio of different drivers are different. The statistical threshold is obtained by calculating the average of the eye aspect ratio and mouth aspect ratio of different driving behavior types and different drivers in the dataset, mainly solving the problem that the adaptive threshold may have a low eye aspect ratio threshold and a high mouth aspect ratio threshold when testing fatigue driving behavior videos.
[0028] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present invention. Description of the Drawings
[0029] The drawings here are incorporated into the specification and form a part of this specification, showing embodiments consistent with the present invention, and are used together with the specification to explain the principles of the present invention.
[0030] Figure 1 It is a schematic flowchart of a fatigue driving behavior recognition method for 3D point cloud image data shown according to an exemplary embodiment.
[0031] Figure 2 It is a schematic diagram of the architecture of the Point-Voxel Fusion Transformation Network (PVFTN) shown according to an exemplary embodiment.
[0032] Figure 3It is a schematic diagram of the architecture of the multi-scale face key point detector MSFLD model shown according to an exemplary embodiment.
[0033] Figure 4 It is a schematic diagram of the process of a driving behavior decision-making method that combines an adaptive threshold and a statistical threshold shown according to an exemplary embodiment. Detailed implementation manners
[0034] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention. In addition, the technical features involved in the various embodiments of the present invention described below can be combined with each other as long as they do not conflict with each other.
[0035] A fatigue driving behavior recognition method for 3D point cloud image data according to the present invention, as Figure 1 shown, includes:
[0036] Step 10, collecting the 3D point cloud image video frames containing the driver's face information after frame division as the image to be processed.
[0037] Step 20, using the point-voxel fusion transformation network PVFTN as the 3D point cloud image data feature extraction model to process the image to be processed and obtain the face region feature image.
[0038] Step 30, using the multi-scale face key point detector MSFLD to process the face region feature image and obtain the coordinates of 23 key points on the face.
[0039] Step 40, updating the face fatigue parameter matrix according to the face key point coordinates.
[0040] Step 50, making a judgment according to the face fatigue parameter matrix by combining a statistical threshold and an adaptive threshold to obtain the judgment result of the driver's driving state.
[0041] Among them, as Figure 2 shown, the image to be processed is input into the point-voxel fusion transformation network PVFTN and sequentially processed by the first PVF module, the second PVF module, and the third PVF module to obtain the first PVF feature map, the second PVF feature map, and the third PVF feature map respectively. These three feature maps are finally processed by the MLP layer and aggregated to obtain the aggregated multi-scale feature as the face region feature image output.
[0042] The PVF module includes a point-voxel branch and a voxel branch, which respectively process the input feature F of the PVF module. The point-voxel branch extracts the global feature F in the point-voxel domain g, the voxel branch extracts local feature F in the voxel domain l , the output end of the PVF module fuses the global feature F g and the local feature F l , to obtain and output the PVF feature map F'.
[0043] The Point-Voxel Fusion Transformation Network (PVFTN) takes N points of the 3D point cloud image as input and feeds them into 3 stacked PVF blocks, using max pooling and repeat operators to extract effective global features representing the entire point cloud. And an MLP layer is used to aggregate multi-scale features. The essence of the PVFTN model is to predict the final per-point segmentation score of the input point cloud. According to the segmentation score, the part label to which the point belongs is determined, and the part label to which the point belongs is determined as the part label category with the maximum segmentation score. The PVFTN model is obtained by training on a training data set with pre-calibrated part labels. When training the PVFTN model only for facial feature image extraction, the part labels of the training data only need two parts, namely the face and non-face.
[0044] The PVF block consists of two branches. The upper branch is voxel-based and is used to aggregate local features, while the lower branch is point-voxel-based and is used to capture global features. The voxel branch collects neighborhood information at a low resolution. However, in order to capture long-range dependencies, only the low-resolution voxel branch is limited. Therefore, we adopt the fusion transformation of voxel and point-voxel features at different scales.
[0045] The basic point-voxel-based method uses standard self-attention on the entire point cloud for global context aggregation. The standard self-attention calculation is permutation-invariant and is very suitable for processing irregular and unordered 3D points. However, the standard self-attention does not include relative position representation in its structure, and this ability is very important for 3D vision tasks. For example, the absolute coordinates of the same object may be completely different with a rigid transformation. Therefore, introducing relative position representation is usually more robust. Therefore, the present invention embeds RPR into the standard self-attention module to achieve a more stable feature extraction effect.
[0046] The point-voxel branch processes the input feature F using the self-attention mechanism embedded with the relative position representation RPR to obtain the self-attention feature F in the point-voxel domain sa , and then processes it through the MLP layer and aggregates it with the input feature F to obtain the global feature F g , and the calculation formula is:[[]]
[0047] F sa = softmax(QK T + B)V
[0048] F g = MLP(F sa)+F
[0049] Among them, Q, K, and V are feature matrices generated by the input feature F according to a shared linear transformation, and B is the relative position representation deviation matrix. For any two points p = {x p , y p , z p} and q = {x q , y q , z q} in the input feature F, the relative position representation deviation B pq = max(-β, min(β, D pq ))), D pq = (x p - x q ) + (y p - y q ) + (z p - z q ), and β is the preset maximum relative distance.
[0050] The voxel branch aims to effectively capture local information, which can bypass expensive sampling and neighbor point queries. Specifically, it includes three steps: voxelization, feature aggregation, and devoxelization. Voxelization maps the input point cloud to a set of new voxel features, and devoxelization converts the per-voxel features back to per-point features. Feature aggregation is performed using a standard transformer architecture on a regular 3D voxel grid, which can significantly improve accuracy. The standard transformer architecture performs global sparse attention SA, which results in quadratic complexity with respect to the number of voxels. In this regard, the window attention mechanism is used to calculate the sparse attention SA within a local 3D window.
[0051] The voxel branch processes the input feature F sequentially through a voxelization module, a feature aggregation module, and a devoxelization module to obtain the voxel local feature F l . Among them, the voxelization module evenly divides the space of the input feature F in the point domain into multiple volume regions in a non-overlapping manner. For each volume region, the mean value of all feature points in the volume region is calculated as the feature value of the volume region, thereby mapping the input feature F in the point domain to the voxel domain feature O. The feature aggregation module adopts the window self-attention mechanism. First, the space of the voxel domain feature O is evenly divided into multiple window spaces in a non-overlapping manner. The voxel domain feature O is processed through the self-attention mechanism within each window space to obtain the voxel domain self-attention feature O sa , and then it is processed through an MLP layer and aggregated with the voxel domain feature O to obtain the voxel domain local feature O g . The devoxelization module assigns the feature value of each volume region in the voxel domain local feature O g to all feature points in the input feature F within the volume region, thereby obtaining the local feature F l。
[0052] Step 30 further includes: As Figure 3 shown, the facial region feature image is input into the multi-scale facial key point detector MSFLD, and successively processed through the first convolutional layer, depthwise separable convolutional layer, 4 inverted residual blocks, second convolutional layer, third convolutional layer, and 5 different-scale fully connected layers to obtain the facial key point coordinates.
[0053] The multi-scale facial key point detector (MSFLD, Multi-scale Facial Landmark Detector) has the following innovations in the network structure compared with the PFLD (Practical Facial Landmark Detector) detector: (1) All the original bottleneck layers are replaced with inverted residual blocks, which can increase the network depth, reduce the number of parameters, reduce the computational amount while ensuring the extraction of feature information, and keep the speed of the facial detector unchanged; (2) 56×56 and 28×28 multi-scale fully connected layers are added, that is, the number of scales is increased from the original 3 to 5, which can accurately locate the facial key points and improve the positioning accuracy of the facial detector.
[0054] Step 40 further includes: calculating the left eye aspect ratio EAR li , right eye aspect ratio EAR ri and mouth aspect ratio MAR i , combining with the time parameter t i of the corresponding current i-th frame 3D point cloud image video frame, and updating the facial fatigue parameter matrix:
[0055] Assume that the width of the eye opening is ω, and the vertical distance between the upper and lower eyelids is h. According to the key points collected by the left and right eyes, the calculation formulas of ω and h for the left and right eyes are shown in formulas (1) and (2) respectively.
[0056]
[0057] In the formula: x6, x8, x9, x 11 are the abscissas of the key points of the left and right eyes respectively; y7, y 13 , y 10 , y 12 are the ordinates of the key points above and below the upper and lower eyelids of the left and right eyes of a person respectively. When the driver is dozing off, the opening and closing degree of the eyes changes, the vertical distance between the upper and lower eyelids decreases, and the width increases. In order to measure whether the eyes are in a fatigued state of dozing off, the left eye aspect ratio K left , right eye aspect ratio K rightThe calculation is as shown in Formulas (3) and (4).
[0058]
[0059] The width ω of the mouth mouth and the height h mouth The calculation is as shown in Formula (5).
[0060]
[0061] where: x18, x 20 are the abscissas of two key points on the left and right of the mouth; y 19 and y 21 are the ordinates of two key points above and below the mouth. When the driver yawns, the opening degree of the mouth is the largest, and the height of the mouth will increase and the width will decrease. To measure whether the mouth is in a fatigued state of yawning, the mouth aspect ratio
[0062] K mouth is introduced, and its calculation is as shown in Formula (6).
[0063]
[0064] Step 50 also includes: as Figure 4 shown, the statistical thresholds are the statistical left eye aspect ratio threshold EAR lst , the statistical right eye aspect ratio threshold EAR rst , and the statistical mouth aspect ratio threshold MAR st obtained based on the data of different driving behavior types and different drivers in the dataset. The adaptive thresholds are the adaptive left eye aspect ratio threshold EAR lat , the adaptive right eye aspect ratio threshold EAR rat , and the adaptive mouth aspect ratio threshold MAR at obtained based on the data of the first p frames of the current test video. When the data of the i-th frame simultaneously satisfies the condition EAR li <max(EAR lst , EAR lat ) and EAR ri <max(EAR rst , EAR rat ), or satisfies the condition MAR i >min(MAR st , MAR at ), it is determined that the driving state of the driver is fatigued, otherwise it is normal.
[0065] The adaptive threshold is obtained by calculating the eye aspect ratio and mouth aspect ratio of the first p frames of each test video and then taking the average. The adaptive threshold is dynamically changed, effectively solving the problem that the eye aspect ratio and mouth aspect ratio of different drivers are different. The statistical threshold is obtained by calculating the average of the eye aspect ratio and mouth aspect ratio of different driving behavior types and different drivers in the dataset, mainly solving the problem that when testing the video of drowsy driving behavior, the adaptive threshold may have a low eye aspect ratio threshold and a high mouth aspect ratio threshold. The fusion strategy of the adaptive threshold and the statistical threshold is as follows: When the test video is input into the model, the driving behavior type is unknown. If the input test video is a drowsy driving behavior, the adaptive threshold of the eyes is too low. In this case, it will lead to an incorrect judgment of the driving behavior. Therefore, the eye aspect ratio threshold takes the maximum value of the adaptive threshold and the statistical threshold, which can avoid misjudgment. Similarly, if the input test video is a yawn-driven behavior, the adaptive threshold of the mouth is too high, which will also lead to a misjudgment of the driving behavior. Therefore, the mouth aspect ratio threshold takes the minimum value of the adaptive threshold and the statistical threshold, which can avoid misjudgment. Using the driving behavior decision method combining the adaptive threshold and the statistical threshold to judge the driving behavior state can effectively improve the prediction accuracy.
[0066] Those skilled in the art will readily conceive of other embodiments of the present invention after considering the specification and practicing the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of the present invention that follow the general principles of the present invention and include common general knowledge or conventional technical means in the technical field not disclosed in this disclosure. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present invention are pointed out by the following claims.
[0067] It should be understood that the present invention is not limited to the exact structures already described and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present invention is only limited by the appended claims.
Claims
1. A fatigue driving behavior recognition method for 3D point cloud image data, characterized in that Including: Step 10: Collect the 3D point cloud image video frames containing the driver's facial information after frame division as the images to be processed; Step 20: Use the Point-Voxel Fusion Transformation Network (PVFTN) as the 3D point cloud image data feature extraction model to process the images to be processed and obtain the facial region feature images; Among them, the images to be processed are input into the PVFTN and successively processed through the first PVF module, the second PVF module, and the third PVF module to obtain the first PVF feature map, the second PVF feature map, and the third PVF feature map respectively. These three feature maps are finally processed through the MLP layer and aggregated to obtain the aggregated multi-scale features as the output of the facial region feature images; The PVF module contains a point-voxel branch and a voxel branch, which process the input feature F of the PVF module respectively. The point-voxel branch extracts the global feature F in the point-voxel domain g , and the voxel branch extracts the local feature F in the voxel domain l . At the output end of the PVF module, the global feature F g and the local feature F l are fused to obtain and output the PVF feature map F; Step 30: Use the Multi-Scale Facial Keypoint Detector (MSFLD) to process the facial region feature images and obtain the coordinates of 23 facial keypoints; Step 40: Update the facial fatigue parameter matrix according to the facial keypoint coordinates; Step 50: According to the facial fatigue parameter matrix, use a method combining statistical threshold and adaptive threshold to make a judgment and obtain the judgment result of the driver's driving state.
2. The fatigue driving behavior recognition method for 3D point cloud image data according to claim 1, characterized in that Step 20 further includes: The point element branch uses the self-attention mechanism of the embedded relative position representation RPR to process the input feature F to obtain the point element domain self-attention feature F sa , and then processes it through the MLP layer and aggregates it with the input feature F to obtain the global feature F g , and the calculation formula is: F sa = softmax(QK T + B)V F g = MLP(F sa ) + F Among them, Q, K, and V are feature matrices generated by the input feature F through a shared linear transformation, and B is the relative position representation deviation matrix. For any two points p = {x p , y p , z p}, q = {x q , y q , z q} in the input feature F, the relative position representation deviation B pq = max(-β, min(β, D pq )), D pq = (x p - x q ) + (y p - y q ) + (z p - z q ), and β is the preset maximum relative distance.
3. The fatigue driving behavior recognition method for 3D point cloud image data according to claim 2, wherein Step 20 further includes: The voxel branch processes the input feature F through a voxelization module, a feature aggregation module, and a devoxelization module in sequence to obtain the voxel local feature F l ; Among them, the voxelization module evenly divides the space of the input feature F in the point-voxel domain into multiple volume regions in a non-overlapping manner. For each volume region, calculate the mean value of all feature points in the volume region as the feature value of the volume region, so as to map the input feature F in the point-voxel domain to the voxel domain feature O; The feature aggregation module adopts the window self-attention mechanism. First, the space of the voxel domain feature O is evenly divided into multiple window spaces in a non-overlapping manner. The voxel domain feature O is processed through the self-attention mechanism within each window space to obtain the voxel domain self-attention feature O sa , and then it is processed through the MLP layer and aggregated with the voxel domain feature O to obtain the voxel domain local feature O g ; The voxelization removal module assigns the eigenvalue of each volume region in the voxel domain local feature O g to all the feature points in the input feature F within the volume region, thereby obtaining the local feature F l .
4. The fatigue driving behavior recognition method for 3D point cloud image data according to claim 1, wherein Step 30 further includes: The facial region feature images are input into the MSFLD and successively processed through the first convolutional layer, the depthwise separable convolutional layer, 4 inverted residual blocks, the second convolutional layer, the third convolutional layer, and 5 fully connected layers with different scales to obtain the coordinates of 23 facial keypoints.
5. The fatigue driving behavior recognition method for 3D point cloud image data according to claim 1, wherein Step 40 further includes: Calculate the left eye aspect ratio EAR based on the facial key point coordinates li , the right eye aspect ratio EAR ri and the mouth aspect ratio MAR i , combined with the time parameter t of the corresponding current i-th frame 3D point cloud image video frame i , update the facial fatigue parameter matrix:
6. The fatigue driving behavior recognition method for 3D point cloud image data according to claim 5, characterized in that Step 50 further includes: The statistical threshold is the statistically obtained left-eye aspect ratio threshold EAR based on data of different driving behavior types and different drivers in the dataset lst , the statistically obtained right-eye aspect ratio threshold EAR rst , and the statistically obtained mouth aspect ratio threshold MAR st ; The adaptive threshold is the adaptive left-eye aspect ratio threshold EAR obtained based on the first p-frame data of the current test video lat and the adaptive right-eye aspect ratio threshold EAR rat and the adaptive mouth aspect ratio threshold MAR at ; When the data of the i-th frame simultaneously meets the condition EAR li <max(EAR lst , EAR lat ), EAR ri <max(EAR rst , EAR rat ) or meets the condition MAR i >min(MAR st , MAR at ), it is determined that the driver's driving state is fatigued; otherwise, it is normal.
7. A computer device, characterized in that, Including: A memory for storing instructions; A processor for calling the instructions stored in the memory to execute the fatigue driving behavior recognition for 3D point cloud image data as described in any one of claims 1-6.
8. A computer-readable storage medium, characterized in that, Instructions are stored, and when the instructions are executed by the processor, the fatigue driving behavior recognition method for 3D point cloud image data as described in any one of claims 1-6 is executed.
Citation Information
Patent Citations
Multi-index fusion-based driver fatigue detection method
CN108875642A
Point cloud multi-task model training method and device and electronic equipment
CN115358413A