Driver monitoring system based on human body posture evaluation
By developing a driver monitoring system conducted by deep learning technology on a low-power embedded platform, the problems of insufficient computing power and difficulty in data acquisition in the existing system are solved, and efficient detection and early warning of driver fatigue and distraction are achieved, ensuring driver and traffic safety.
Patent Information
- Application Number
- CN202510278466.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-10
- Publication Date
- 2025-06-27
AI Technical Summary
Due to insufficient computing power, difficulty in collecting and labeling data, high accuracy requirements for algorithm models and complex system integration, existing driver monitoring systems are difficult to effectively monitor and prevent dangerous driving behaviors such as driver fatigue and distraction.
A driver monitoring system based on human posture evaluation was designed, using deep learning technology for facial feature extraction and multi-feature fusion, combining transfer learning to accelerate model convergence, and developed on the JetsonNano embedded platform with low power consumption, improving data acquisition efficiency through data augmentation and semi-automated data annotation tools, using the YOLOv5s model and Stacking method for behavior detection, and sending warning information to the driver through voice transmission function.
It effectively solves the problem of low computing power on the vehicle platform, improves the real-time and accuracy of the system, and can efficiently detect dangerous driving behaviors such as driver fatigue and distraction, prevent traffic accidents, and ensure driver safety.
Smart Images

Figure CN120220122A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a driver monitoring system based on human posture assessment. Background Art
[0002] The driver monitoring system (DMS) based on human posture assessment is a key technology for ensuring driving safety in the field of intelligent transportation. It uses advanced sensors and algorithms to monitor and analyze the driver's posture, behavior, and physiological state in real time to prevent traffic accidents. However, the existing driver monitoring systems have the following problems:
[0003] Software computing power bottleneck: Most mainstream in-vehicle computing platforms are mid- to low-end ARM CPUs / GPUs with limited computing power. This makes it difficult for algorithms such as face detection, key point detection, face recognition, gaze tracking, and gesture recognition to run efficiently and cannot meet the requirements of real-time driver monitoring.
[0004] Data collection and annotation dilemma: Computer vision algorithms have strict requirements for image quality. In actual driving scenarios, factors such as the signal-to-noise ratio of sound, differences in image contrast, occlusion and jitter of images, and brightness differences caused by light changes in different weather and time periods all require a large amount of data collection and annotation, consuming a large amount of human and time costs.
[0005] Algorithm model problem: Driver monitoring needs to detect various complex behaviors, which puts high requirements on the accuracy of the algorithm model. It is necessary to design appropriate deep learning models and multi-feature fusion strategies.
[0006] System integration problem: Driver monitoring needs to be seamlessly integrated with devices such as in-vehicle systems, mobile phones, and PCs. It is necessary to design a user-friendly interaction interface to ensure the stable and reliable operation of the system, which poses challenges to system integration and testing.
[0007] Therefore, a driver monitoring system based on human posture assessment is needed to solve problems related to data, algorithms, computing power, accuracy, and system integration. Summary of the Invention
[0008] Aiming at the deficiencies of the prior art, the present invention provides a driver monitoring system based on human posture assessment, which solves the problems raised in the above background art;
[0009] Technical solution: To solve the above technical problems, according to one aspect of the present invention, more specifically, a driver monitoring system based on human posture assessment includes a data acquisition end, a data processing end, a feature fusion end, and a display end;
[0010] The data acquisition module is used to collect the driver's facial video stream and other relevant physiological and behavioral data. The data acquisition module uses public datasets as the basis and expands the datasets through data augmentation techniques, and uses semi-automated data annotation tools to improve the annotation efficiency;
[0011] The fatigue detection module, based on face feature detection technology, determines whether there is a yawning behavior by extracting the features of the mouth area, uses an eye detection algorithm to locate the eye area and calculates the eye opening degree within the area, and determines whether the driver is in a fatigued state based on the EAR value and the PERCLOS value;
[0012] The distracted driving determination module uses the principle of human eye imaging to detect the driver's pupil points in real time to calculate the line-of-sight angle, combines the driver's normal driving line-of-sight angle value and the determination range of distraction to determine whether there is a distracted driving behavior, divides the ROI area for the face area, and learns the behavior features such as smoking and making phone calls through deep learning training within the ROI area, and uses hand detection and object detection for detection and determination;
[0013] The model module uses the YOLOv5s model for driver behavior detection, and performs model fusion through the Stacking method. Multiple fine-tuned single models are used as base learners, and XGBoost is used as a secondary learner to retrain the classification results of the base learners;
[0014] The hardware module uses the Jetson Nano embedded computing platform. The Jetson Nano embedded computing platform includes NVIDIA's Tegra X1 processor, which has four ARM Cortex-A57 CPU cores and one NVIDIA Maxwell GPU, as well as memory, storage, and various peripheral interfaces, and is used to provide computing support for the system;
[0015] The system software module consists of a video acquisition, fatigue feature extraction, distracted behavior detection, warning decision-making, and voice transmission module. It collects the driver's facial video stream through a camera device, extracts facial features, detects distracted behaviors, makes warning decisions based on the detection results, and sends warning messages to the driver through the voice transmission function;
[0016] The platform-side module includes a data acquisition end, a data processing end, a feature fusion end, and a display end. The data acquisition end is responsible for collecting various physiological and behavioral data of the driver. The data processing end processes the collected data. The feature fusion end fuses multiple features from the data processing end and uses machine learning algorithms for classification and recognition. The display end displays the monitoring results to the driver and can transmit them to the vehicle control system;
[0017] The data visualization module is used to display information such as the driver's facial expressions, eye movements, gaze trajectories, and attention distribution maps. By analyzing this information, it detects the driver's fatigue level and distraction situation, and displays corresponding alarms or prompt messages when the driver's state is abnormal.
[0018] Furthermore, the calculation formula for the EAR value is:
[0019]
[0020] The calculation formula for the PERCLOS value is:
[0021]
[0022] Furthermore, in the data acquisition module, three in-vehicle cameras located at different positions are used to collect data. The dataset contains two activity types: distraction activities and the gaze area of each participant, and each activity type has two groups of data: with and without appearance blocks.
[0023] Furthermore, during data preprocessing, the distracted driving determination module synthesizes all video files from a single participant into one file by using Python and FFmpeg.
[0024] Furthermore, during training, the model module uses loss according to the calculation principle of categorical crossentropy, and the formula is
[0025]
[0026] Based on transfer learning, the Fine-tune method is selected to lock the weight part levels of ImageNet and retrain other part levels. And during Fine-tune, Adam is first used as the main optimizer, and when it is optimized to meet the expected requirements, RMSprop is used for ultra-low learning rate tuning.
[0027] The beneficial effects of a driver monitoring system based on human posture assessment of the present invention are as follows:
[0028] (1) By using deep learning technology for facial feature extraction and multi-feature fusion, combined with transfer learning to accelerate model convergence, and developing on a low-power Jetson Nano embedded platform, the present invention effectively solves the problem of low computing power of the in-vehicle platform and improves the real-time performance and accuracy of the system.
[0029] (2) The present invention has high detection accuracy, can effectively detect dangerous driving behaviors such as driver fatigue and distraction, prevent traffic accidents, and ensure the safety of people's lives and property.
[0030] (3) The present invention is applicable to driver monitoring in the intelligent cockpit environment and can be widely used in multiple fields such as intelligent transportation systems and commercial vehicle fleet management, promoting the development of intelligent transportation and autonomous driving technologies. Description of the Drawings
[0031] The present invention will be further described in detail below with reference to the drawings and specific implementation methods.
[0032] Figure 1 Schematic structural diagram of the facial feature fatigue driving detection process in Embodiment 1 of the present invention;
[0033] Figure 2 Schematic diagram of the YOLOv5 network structure in Embodiment 1 of the present invention;
[0034] Figure 3 Schematic structural diagram of the Stacking training stage method in Embodiment 1 of the present invention;
[0035] Figure 4 Schematic flow diagram of Embodiment 1 of the present invention;
[0036] Figure 5 Schematic diagram of the final model accuracy in Embodiment 2 of the present invention;
[0037] Figure 6 Classification result diagram of LightGBM in Embodiment 2 of the present invention. Detailed Implementation Modes
[0038] The present invention will be described in detail below with reference to the drawings and embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other.
[0039] To make the technical solution of the present invention clearer, the present invention will be further described in detail below with reference to the drawings and specific embodiments.
[0040] Embodiment 1
[0041] Referring to Figure 1-4 , a driver monitoring system based on human posture assessment, including a data acquisition end, a data processing end, a feature fusion end, a display end, and a data acquisition module, which is used to collect the facial video stream of the driver and other relevant physiological and behavioral data. The data acquisition module uses a public data set as a basis and expands the data set through data augmentation technology, and improves the annotation efficiency by using a semi-automatic data annotation tool;
[0042] The fatigue detection module, based on face feature detection technology, determines whether there is a yawning behavior by extracting the features of the mouth area, uses an eye detection algorithm to locate the eye area of a person and calculates the eye opening degree within the area, and determines whether the driver is in a fatigued state based on the EAR value and the PERCLOS value;
[0043] The distracted driving determination module uses the principle of human eye imaging to detect the pupil points of the driver in real time to calculate the line-of-sight angle, combines the normal driving line-of-sight angle value of the driver and the determination range of distraction to determine whether there is a distracted driving behavior, divides the ROI area for the face area, and learns the behavior characteristics such as smoking and making phone calls through deep learning training within the ROI area, and uses hand detection and object detection for detection and determination;
[0044] The model module uses the YOLOv5s model for driver behavior detection, and performs model fusion through the Stacking method. Multiple fine-tuned single models are used as base learners, and XGBoost is used as a secondary learner to retrain the classification results of the base learners;
[0045] The hardware module uses the Jetson Nano embedded computing platform. The Jetson Nano embedded computing platform includes NVIDIA's Tegra X1 processor, which has four ARM Cortex-A57 CPU cores and one NVIDIA Maxwell GPU, as well as memory, storage, and various peripheral interfaces, and is used to provide computing support for the system;
[0046] The system software module consists of a video acquisition, fatigue feature extraction, distracted behavior detection, warning decision-making, and voice transmission module. It collects the driver's facial video stream through a camera device, extracts facial features, detects distracted behaviors, makes a warning decision based on the detection results, and sends a warning message to the driver through the voice transmission function;
[0047] The platform terminal module includes a data acquisition end, a data processing end, a feature fusion end, and a display end. The data acquisition end is responsible for collecting various physiological and behavioral data of the driver. The data processing end processes the collected data. The feature fusion end fuses multiple features from the data processing end and uses machine learning algorithms for classification and recognition. The display end displays the monitoring results to the driver and can transmit them to the vehicle control system;
[0048] The digital large-screen display module is used to display information such as the driver's facial expressions, eye activities, line-of-sight trajectories, and attention distribution maps. By analyzing this information, it detects the driver's fatigue level and distraction situation, and displays corresponding alarms or prompt messages when the driver's state is abnormal.
[0049] Preferably, the fatigue detection principle is analyzed as follows:
[0050] Through face feature detection technology, extract the features of the mouth area, and judge whether there is a yawning behavior according to the degree of mouth opening and closing. Locate the eye area through the eye detection algorithm, calculate the degree of eye opening and closing within the area, and judge whether there is a closed-eye behavior according to the threshold control.
[0051] After using the object detection algorithm model to classify and detect the face and eyes, it is necessary to further judge the driving state of the driver according to the classification results. This section will introduce some evaluation methods related to fatigue driving, including the PERCLOS criterion, EAR value, etc.
[0052] EAR value:
[0053] EAR is a conceptual model used to evaluate the eye state. When the human eye is in the open state, the value of EAR is a relatively constant value. When the human eye is in the closed state, the value of EAR will approach zero. Although individual differences will have a certain impact on the value of EAR, the change law of the EAR value when an individual changes from the open-eye state to the closed-eye state is the same. When the eyes close from the open state, the EAR value will rapidly drop from a relatively constant value to a zero value, and then quickly increase to the previous relatively constant EAR value after opening the eyes again. Therefore, the EAR value is often used to judge the opening and closing state of the eyes. Before calculating the EAR value using Equation (3-1), it is necessary to obtain the key points of the eye contour of the face.
[0054]
[0055] After obtaining the eye feature points, calculate the EAR value using the method shown in Equation (3-2). The numerator is to calculate the distance between the vertical eye landmarks, and the denominator is to calculate the distance between the horizontal eye landmarks. Through this method, it is possible to simply determine whether a person blinks based on the distance ratio between the eye landmarks.
[0056] PERCLOS value:
[0057] Fatigue is a dynamically changing process over time and is a description of the physiological state. Scholars and experts at Carnegie Mellon proposed a physical quantity - PERCLOS criterion for measuring fatigue by calculating the ratio of the number of frames in which the eyes are in the closed state per unit time through careful thinking and powerful experimental demonstrations. The specific implementation of PERCLOS is shown in Equation (3-2), where Nclose represents the number of frames of the closed eyes in unit time, and Ntotal represents the total number of frames in unit time.
[0058]
[0059] To calculate PERCLOS, it is necessary to determine whether the eyes are in a closed state. In addition to using a deep learning algorithm model to directly classify open and closed eyes, it is also possible to judge whether the eyes are in a closed state by calculating the proportion of the eyelid covering the pupil. Nowadays, the PERCOLS criterion has been widely recognized in the academic community, and many researchers use it as an effective indicator for detecting fatigue driving.
[0060] Figure 1 The framework of the fatigue driving detection process based on the driver's facial features is shown, summarizing the technologies and methods used in the key steps of the process.
[0061] Preferably, in the data acquisition module, three on-vehicle cameras located at different positions (on the dashboard, on the rearview mirror, and at the upper right window corner) are used to collect data. The dataset contains two activity types, namely distracted activities and the gaze area of each participant, and each activity type has two groups of data with or without appearance blocks (such as wearing a hat or sunglasses).
[0062] Preferably, when preprocessing the data, the distracted driving determination module synthesizes all video files from a single participant into one file by using Python and FFmpeg.
[0063] The convolutional neural network CNN is used as a whole to solve the problem of behavior detection. The adjacent pixels of the picture are changed into units with smaller length and width and deeper height, so as to extract local features. The features are mapped backward layer by layer to extract the final driving behavior features. Facing 10 different classification situations in the dataset, the calculation principle of categorical cross-entropy is used to calculate the loss (Formula 1), and the probabilities of each classification are summed up.
[0064]
[0065] On the basis of transfer learning, without locking the weights of some parts of the ImageNet model, Fine-tune is selected to lock the weight parts of ImageNet at the hierarchical level and retrain the other parts at the hierarchical level. The reason for doing this is that the characteristics of the dataset in this project are quite different from those of the ImageNet dataset. When training the data, if only the method of training the fully connected layer is adopted, it is very difficult to find the features during training. Therefore, it is necessary to advance to the level before the fully connected layer to train the model in order to ensure finding the features.
[0066] When performing fine-tuning, Adam is selected as the main optimizer. When the optimization reaches the expected requirements, RMSprop is used for fine-tuning with an ultra-low learning rate. After experiments, it is found that when performing single-model fine-tuning, using RMSprop results in better performance than Adam in terms of accuracy by 0.03 - 0.05 and in terms of LogLoss by 0.05 - 0.2 points. However, if RMSprop is used throughout, the convergence is relatively slow and it also reaches the accuracy achieved when starting to train with the Adam optimizer.
[0067] Example 2
[0068] System testing
[0069] 1. Data acquisition testing
[0070] Test the reliability of the acquisition device: By simulating different driving situations (such as different vehicle speeds, different road conditions, etc.), test the stability of the device and the accuracy of data acquisition in various situations. Test the accuracy of the acquisition device: Compare the acquisition device with a known accurate device to test the accuracy and error range of the acquisition device. Test the stability of the acquisition device: Conduct a long-term stability test on the acquisition device to ensure that the acquisition device does not have errors or instability during long-term use.
[0071] Feature extraction testing:
[0072] Comparison testing: Compare the features extracted by the system with known accurate features to test the accuracy and error range of the extraction algorithm. Robustness testing: Test the robustness of the algorithm under different driving behaviors, drivers, vehicle speeds, and road conditions. Time efficiency testing: Test the time required for the algorithm to extract features to ensure that the algorithm can operate in a real-time system.
[0073] Feature fusion testing:
[0074] Comparison testing: Compare the features fused by the system with known accurate features to test the accuracy and error range of the fusion algorithm. Cross-validation testing: Divide the dataset into a training set and a test set through the cross-validation method and test the performance of the algorithm on the test set. Robustness testing: Test the robustness of the algorithm under different driving behaviors, drivers, vehicle speeds, and road conditions.
[0075] 2. System effectiveness evaluation
[0076] When training the overall model, the accuracy rate is as shown in the figure and reaches over 98%. As Figure 5 shown.
[0077] In the simultaneous design model, LightGBM is used for classification, and the prediction results are as follows: The precision is 0.9738, indicating that 97.38% of the samples predicted as positive are truly positive classes. The recall is 0.9736, indicating that 97.36% of the actual positive samples are correctly predicted as positive classes. The F1-Score is 0.9735, which combines the precision and recall metrics to measure the comprehensive performance of the model. The accuracy is 0.9736, indicating the proportion of all samples that are correctly classified. These evaluation metrics reflect the performance of the classification model in the driver DMS safety monitoring and warning system.
[0078] As Figure 6 shown; Use the trained model to predict the test set (X_test) to obtain the prediction result y_pred. Calculate the prediction accuracy and add it to the accuracy list. By plotting the accuracy curve, it helps to analyze the influence of the n_estimators parameter on the random forest model, so as to select the optimal value of n_estimators.
[0079] The driver monitoring system (DMS) based on human pose estimation uses deep learning technology for facial feature extraction and multi-feature fusion. At the same time, transfer learning is used to accelerate model convergence, and it is developed on the low-power Jetson Nano embedded platform to solve the problem of low computing power on in-vehicle platforms. This innovation not only has technical content but also shows practical application value. The innovation points include depthwise separable convolution, transfer learning, embedded platform, and multi-feature fusion, and the effectiveness of the system is verified in experiments. The system is applicable to driver monitoring in the intelligent cockpit environment and has broad application prospects, which helps to prevent traffic accidents, protect people's lives and property safety, and promote the development of intelligent transportation and autonomous driving technologies
[0080] The above-described embodiments merely represent several implementation manners of the present invention. The description is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several deformations and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention shall be subject to the appended claims.
Claims
1. A driver monitoring system based on human posture assessment, comprising a data acquisition end, a data processing end, a feature fusion end, and a display end, characterized in that: A data collection module, used to collect the driver's facial video stream and other relevant physiological and behavioral data, the data collection module uses a public data set as a basis, expands the data set through data enhancement technology, and uses a semi-automatic data annotation tool to improve annotation efficiency; The fatigue detection module is based on facial feature detection technology. It extracts the features of the mouth area to determine whether there is yawning behavior. It uses the eye detection algorithm to locate the eye area and calculate the degree of eye opening in the area. It determines whether the driver is in a fatigued state based on the EAR value and PERCLOS value. The distracted driving judgment module uses the principle of human eye imaging to detect the driver's pupil point in real time to calculate the line of sight angle, and combines the driver's normal driving line of sight angle value and the distraction judgment range to determine whether there is distracted driving behavior. The face area is divided into ROI areas, and behavioral characteristics such as smoking and making phone calls are learned through deep learning training in the ROI area, and hand detection and object detection are used for detection and judgment; In the model module, the YOLOv5s model is used for driver behavior detection, and the model is integrated through the Stacking method. Multiple fine-tuned single models are used as base learners, and XGBoost is used as a secondary learner to retrain the classification results of the base learners. The hardware module uses the Jetson Nano embedded computing platform, which includes NVIDIA's Tegra X1 processor, four ARM Cortex-A57 CPU cores and an NVIDIA Maxwell GPU, as well as memory, storage, and various peripheral interfaces to provide computing support for the system; The system software module consists of video acquisition, fatigue feature extraction, distraction behavior detection, warning decision and voice transmission modules. It uses the camera to collect the driver's facial video stream, extract facial features, detect distraction behavior, make warning decisions based on the detection results, and send warning information to the driver through the voice transmission function; The platform module includes a data collection terminal, a data processing terminal, a feature fusion terminal and a display terminal. The data collection terminal is responsible for collecting various physiological and behavioral data of the driver. The data processing terminal processes the collected data. The feature fusion terminal fuses multiple features from the data processing terminal and uses a machine learning algorithm for classification and identification. The display terminal displays the monitoring results to the driver and can transmit them to the vehicle control system. The digital large-screen display module is used to display information such as the driver's facial expressions, eye movements, eye tracking and attention distribution. By analyzing this information, the driver's fatigue and distraction can be detected, and corresponding alarms or prompts can be displayed when the driver's status is abnormal.
2. A driver monitoring system based on human posture assessment according to claim 1, characterized in that: The EAR value calculation formula is: The PERCLOS value calculation formula is:
3. A driver monitoring system based on human posture assessment according to claim 1, characterized in that: In the data acquisition module, three vehicle-mounted cameras positioned at different locations are used to collect data. The data set includes two types of activities: distracting activities and each participant's gaze area, and each activity type has two sets of data with and without appearance blocks.
4. A driver monitoring system based on human posture assessment according to claim 1, characterized in that: During data preprocessing, the distracted driving determination module combines all video files from a single participant into one file by using Python and FFmpeg.
5. A driver monitoring system based on human posture assessment according to claim 1, characterized in that: When training, the model module uses the loss calculated based on the categorical crossentropy principle, and the formula is: On the basis of transfer learning, the Fine-tune method is selected to lock the weight part of ImageNet and retrain the other parts. In Fine-tune, Adam is first used as the main optimizer. When optimized to meet the expected requirements, RMSprop is used to tune the ultra-low learning rate.