Intelligent driver state monitoring control system and method based on multi-mode perception
The intelligent driver status monitoring system, which combines multimodal perception with infrared cameras and deep learning algorithms, analyzes the driver's status in real time, solving the problems of low accuracy and poor adaptability of existing systems. It achieves accurate monitoring and intelligent early warning of the driver's status, thereby improving driving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHERY COMMERCIAL VEHICLE (ANHUI) CO LTD
- Filing Date
- 2026-01-07
- Publication Date
- 2026-05-01
AI Technical Summary
Existing driver monitoring systems rely on manual observation or simple sensors, resulting in low accuracy, slow response, poor adaptability, and the ability to only monitor whether the driver's eyes are open, with low recognition accuracy. They are also unable to comprehensively monitor fatigue, distraction, and dangerous behavior.
The system employs a multimodal perception-based intelligent driver status monitoring system, which combines an infrared supplementary light camera, a control unit (MCU), and a warning unit. Through high-precision image acquisition and deep learning algorithms, it analyzes the driver's status in real time, identifies fatigue, distraction, and dangerous behaviors, and issues warnings through the vehicle's instruments or ECU.
It enables precise monitoring of driver fatigue, distraction, and dangerous behavior, provides intelligent early warnings, improves driving safety, and enhances the system's accuracy, reliability, and adaptability.
Smart Images

Figure CN121963156A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of active safety in automobiles, and in particular to an intelligent driver status monitoring system and method that integrates infrared visual perception and artificial intelligence algorithms. Background Technology
[0002] With the increasing number of cars on the road and the growing traffic congestion, driver fatigue, distraction, and dangerous driving behaviors have become one of the main causes of traffic accidents. Statistics show that globally, traffic accidents caused by driver fatigue or distraction account for more than 20% of all accidents annually. Therefore, developing an efficient and reliable driver condition monitoring system is crucial for improving road safety.
[0003] Traditional driver monitoring methods mainly rely on manual observation or simple sensor detection, which suffer from problems such as low accuracy, slow response, and poor adaptability.
[0004] The development of modern computer vision, artificial intelligence, and embedded systems has provided new technical means for DMS (Driver Monitoring System). However, existing technologies only consider driver fatigue and use only the method of whether the driver's eyes are open, resulting in low accuracy and incomplete monitoring of the driver's condition. Summary of the Invention
[0005] The purpose of this invention is to overcome the shortcomings of the prior art and provide an intelligent driver state monitoring and control system and method based on multimodal perception. Through high-precision image acquisition, deep learning algorithms and real-time data analysis, it can achieve accurate monitoring of driver fatigue, distraction and dangerous behavior.
[0006] To achieve the above objectives, the technical solution adopted by the present invention is as follows:
[0007] The intelligent driver status monitoring and control system based on multimodal perception includes an infrared supplementary light camera, a control unit (MCU), and an early warning unit. The infrared supplementary light camera is used to collect facial image data of the driver, and its output is connected to the control unit (MCU). The control unit (MCU) identifies the current driver status based on the collected facial data and drives the early warning unit to issue early warning information based on the identification result.
[0008] The control unit MCU has built-in fatigue monitoring subunit, distraction monitoring subunit, and dangerous action monitoring subunit, which are used to monitor driver fatigue, distraction, and dangerous actions respectively and identify the corresponding results.
[0009] The alarm unit includes an in-vehicle instrument panel, which issues voice, text, and graphic reminders; and / or the alarm unit includes an in-vehicle ECU, which promptly notifies the driver by opening the windows and activating the hazard lights.
[0010] The control system also includes a fatigue alarm control switch, which is used to control and disable the driver status alarm signal.
[0011] The intelligent driver status monitoring and control method based on multimodal perception includes collecting driver facial data through an infrared supplementary light camera and performing real-time analysis of the collected facial data. Based on the real-time analysis results of the driver status, an alarm reminder is issued. The driver status analysis results include driver fatigue, driver distraction, and driver dangerous actions. Based on the driver status analysis results, corresponding alarm reminders are issued.
[0012] Perclos data and yawn count data are obtained through facial recognition, and combined with the obtained vehicle speed to determine whether the current driver is fatigued and the level of fatigue; based on the fatigue status and fatigue level, the corresponding warning method is controlled to issue a warning to the driver.
[0013] The driver's fatigue state and corresponding fatigue level are determined by using pre-set Perclos value thresholds and yawn count thresholds; the Perclos value threshold and yawn count threshold are set based on the current vehicle speed, with the threshold decreasing as the speed increases.
[0014] Distraction monitoring includes identifying the driver's head position, turning angle, and gaze direction based on facial and head image data collected by cameras. When the head position deviates from the normal position or the head turning angle exceeds a set angle threshold and continues for a set time, it is determined that the driver is in a distracted state, and the control will issue a warning reminder corresponding to the distracted state.
[0015] During distraction monitoring, real-time vehicle data is acquired to determine whether the vehicle is in a turning state. When the vehicle is in a turning state, the analysis and monitoring function is automatically disabled until the turning state ends.
[0016] Dangerous action monitoring involves analyzing images captured by cameras to identify the driver's hand and facial movements, as well as the objects appearing in the image, to determine whether the driver is in a dangerous driving state, and issuing corresponding alarm signals based on the identification results.
[0017] The advantages of this invention are as follows: The intelligent driver state monitoring system based on multimodal perception achieves accurate monitoring of driver fatigue, distraction and dangerous behavior through high-precision image acquisition, deep learning algorithms and real-time data analysis, and provides intelligent early warning by combining vehicle CAN bus communication, thereby greatly improving driving safety. Attached Figure Description
[0018] The following is a brief explanation of the contents of each of the accompanying drawings and the markings in the drawings:
[0019] Figure 1 This is a schematic diagram of the overall architecture of the system of the present invention;
[0020] Figure 2 This is a block diagram of the camera hardware structure.
[0021] Figure 3 This is a block diagram of the controller hardware structure. Detailed Implementation
[0022] The specific embodiments of the present invention will be further described in detail below with reference to the accompanying drawings and the description of the preferred embodiments.
[0023] The intelligent driver status monitoring system based on multimodal perception in this embodiment achieves accurate monitoring of driver fatigue, distraction, and dangerous behavior through high-precision image acquisition, deep learning algorithms, and real-time data analysis. Combined with vehicle CAN bus communication, it provides intelligent early warning, thereby significantly improving driving safety.
[0024] This embodiment presents an intelligent driver state monitoring system based on multimodal perception. The system collects driver facial data using a high-performance infrared camera, and combines this data with advanced computer vision algorithms and vehicle CAN bus communication to achieve real-time monitoring and early warning of driver fatigue, distraction, and dangerous behaviors. The system features high precision, high reliability, and strong adaptability, and can be widely applied to passenger cars, commercial vehicles, and autonomous vehicles, primarily addressing safety hazards caused by fatigue, distraction, and dangerous behaviors during driving.
[0025] This embodiment of the intelligent driver state monitoring and control system based on multimodal perception includes an infrared supplementary light camera, a control unit (MCU), and a warning unit in its hardware component. The infrared supplementary light camera is used to collect facial image data of the driver, and its output is connected to the control unit (MCU). The control unit (MCU) identifies the current driver state based on the collected facial data and drives the warning unit to issue warning information based on the identification result. The control unit (MCU) incorporates a fatigue monitoring subunit, a distraction monitoring subunit, and a dangerous action monitoring subunit, which are used to respectively monitor driver fatigue, distraction, and dangerous actions, and identify the corresponding results.
[0026] Alarm control is implemented based on the results of fatigue monitoring, distraction monitoring, and dangerous action monitoring. Corresponding alarms are output through alarm units, including on-board instrument panels that issue voice, text, and graphic alerts; and / or on-board ECUs that promptly notify the driver by opening windows or activating hazard lights. Different alarms are output based on different monitoring results.
[0027] In a preferred embodiment, a fatigue alarm control switch is provided. The fatigue alarm control switch is used to control and turn off the driver status alarm signal. The fatigue alarm control switch is connected to the control unit MCU. It can turn off the current alarm and the monitoring system according to the user's actual needs. For example, in some working conditions where the driver is resting or does not need to be monitored, the fatigue alarm can be turned off manually.
[0028] This embodiment provides an intelligent driver state monitoring and control method based on multimodal perception. It includes acquiring driver facial data via an infrared supplemental lighting camera and performing real-time analysis of the acquired facial data. Based on the real-time analysis results of the driver's state, an alarm is issued. The driver's state analysis results include driver fatigue, driver distraction, and dangerous driver actions. Corresponding alarms are issued based on these results. Driver fatigue, driver distraction, and dangerous driver actions are identified, monitored, and alarmed using driver fatigue monitoring strategies, driver analysis monitoring strategies, and dangerous driver action monitoring strategies, respectively. The specific scheme is as follows:
[0029] Driver fatigue monitoring strategy: When monitoring the driver's condition, Perclos data and yawn count data are obtained through facial recognition, and combined with the vehicle speed to determine whether the driver is currently fatigued and the fatigue level. Based on the fatigue state and fatigue level, corresponding warning methods are used to issue warnings to the driver. Perclos data refers to the percentage of time the eyes are closed out of the total time. The driver's fatigue state and corresponding fatigue level are determined based on pre-set Perclos value thresholds and yawn count thresholds. These thresholds are set based on the current vehicle speed; the higher the speed, the lower the threshold. Perclos value and yawn count values per unit time can be identified through facial recognition and other methods. The Perclos value is compared with a set first threshold, and the yawn count value is compared with a set second threshold. If either or both exceed the set threshold, the driver is considered fatigued. First and second thresholds can be set for each fatigue level, and the current fatigue level is determined by comparing the driver's current fatigue level to the first and second thresholds corresponding to each fatigue level. In this embodiment, the values of the first threshold and the second threshold are set according to the vehicle speed. The higher the vehicle speed, the smaller the threshold. The reason for this is that the higher the vehicle speed, the greater the risk, and therefore the more stringent the classification and judgment of fatigue. Thus, it is necessary to reduce the threshold to improve the reliability and safety of fatigue monitoring.
[0030] The distraction detection strategy includes: identifying the driver's head position, turning angle, and gaze direction based on facial and head image data collected by cameras. If the head position deviates from the normal position or the head turning angle exceeds a set threshold and persists for a set time, the driver is determined to be distracted, and a warning alert corresponding to the distraction state is issued. During distraction detection, real-time vehicle data is acquired to determine if the vehicle is turning. When the vehicle is turning, the analysis and monitoring functions are automatically disabled until the turning state ends. The driver's distraction status, such as looking around, checking a phone, or chatting, can be assessed based on head position, head angle, and gaze direction. Pre-set ranges for head position, head angle, and gaze direction can be used to determine if the driver is distracted. If any of these conditions are met, the driver is considered distracted; otherwise, the driver is considered distracted. In this embodiment, the degree of distraction is determined by the degree of head position deviation, head angle deviation, gaze direction deviation, and duration. A greater deviation or longer duration results in a higher level of distraction. Different distraction levels are defined based on preset ranges for head position, head angle, and gaze direction.
[0031] The hazardous action monitoring strategy involves analyzing images captured by cameras to identify the driver's hand and facial movements, as well as the objects appearing in the scene, to determine whether the driver is in a dangerous driving state and to issue corresponding alarm signals based on the identification results. Through comprehensive analysis of the driver's hand and facial movements, as well as object characteristics, the system can detect whether the driver is smoking or making a phone call while driving. For example, when the system detects a hand gesture resembling picking up or lighting a cigarette, coupled with facial features resembling smoking, it immediately identifies it as smoking. Similarly, when the system detects the driver holding a phone close to their ear, accompanied by corresponding head movements, it identifies it as making a phone call. Once these hazardous actions are detected, and the vehicle speed exceeds 50 km / h, the system immediately issues a warning to remind the driver to correct the dangerous behavior and ensure driving safety.
[0032] In a preferred embodiment of this application, the system immediately detects whether the camera is obstructed after the vehicle is powered on. If camera obstruction is detected, the system will promptly notify the driver via voice, text, and graphic alerts on the instrument panel. The voice alert is broadcast twice consecutively, and the text is displayed on the screen for 6 seconds. This timely alert allows the driver to quickly notice any abnormalities in the camera and take timely measures to resolve the problem, ensuring that the system is always in normal working order.
[0033] In another preferred embodiment, the driving status of the vehicle is monitored. When the vehicle is in autonomous driving mode or high-speed driving mode (when the vehicle speed is greater than a set speed threshold), interactive tasks are periodically initiated at regular intervals. The interactive tasks include voice interaction tasks and behavioral interaction tasks.
[0034] The voice interaction task includes: after initiating the interaction task, issuing a text recognition task to the driver via the instrument panel or in-vehicle voice system. The text to be recognized is displayed on the instrument panel or in-vehicle display screen. After issuing the text recognition task instruction to the driver, the in-vehicle voice interaction system collects the driver's voice in real time and obtains the text read by the driver through voice recognition. The text obtained by voice recognition is compared with the text of the text task instruction. If the comparison result is the same, the voice interaction task ends; otherwise, the voice interaction task is repeated. If the voice interaction task cannot be ended after multiple attempts (reaching a set number of requirements), an alarm reminder is issued to the driver through the in-vehicle alarm unit, reminding the driver to pay attention to driving safety and take a break in time. In this embodiment, the in-vehicle voice system detects the volume of the driver's voice. Only when the volume is greater than a set threshold and the text obtained by voice recognition is the same as the text of the text task instruction is the voice interaction task considered to end. The purpose of this is to stimulate the driver by reading the text aloud, thereby eliminating some distraction and fatigue.
[0035] The behavioral interaction task includes the driver's mouth-opening task. The driver is given the mouth-opening task command through the instrument panel or the vehicle voice system. Then, the driver's facial data is collected by the vehicle camera. The system determines whether the driver has completed the mouth-opening task based on the facial data. If so, the mouth-opening task ends. Otherwise, the mouth-opening task is repeated. If the mouth-opening interaction task cannot be ended after multiple attempts (reaching the set number of times), an alarm reminder is issued to the driver through the vehicle alarm unit, reminding the driver to pay attention to driving safety and take a break in time.
[0036] In this embodiment, determining whether the driver has completed the mouth-opening task using facial data includes: issuing a mouth-opening command to the driver, determining whether the driver has opened their mouth using facial images and monitoring the size of the mouth opening. Only when the mouth is open and the size of the mouth opening is greater than a set threshold is the mouth-opening task considered complete; otherwise, the mouth-opening task is considered incomplete. The set threshold for the mouth opening size can be based on the maximum size that the driver can open their mouth, multiplied by a proportional coefficient K, where the proportional coefficient K is greater than 50%. That is, only when the driver opens their mouth more than 50% of their maximum mouth opening value can the mouth-opening requirement be met. The purpose of this is to stimulate the driver's mental alertness when the driver opens their mouth and opens it wide, thus avoiding fatigue.
[0037] By requiring drivers to read text aloud and open their mouths wide, driver fatigue and drowsiness are eliminated, providing a certain level of stimulation and reducing inattention caused by prolonged driving, thus improving driving safety. In this embodiment, each interactive task includes both voice and behavioral interaction tasks, using loud reading and mouth opening to invigorate the driver.
[0038] The hardware components and control strategies employed in the monitoring system of this embodiment include:
[0039] The image acquisition technology uses a high-performance infrared supplemental camera, whose technical parameters are as follows:
[0040] Wide Spectrum Adaptability: The high-performance infrared supplementary lighting camera equipped in this system has a wide operating wavelength range of 400-1100nm. This characteristic enables it to operate normally under various lighting conditions, whether it is strong direct sunlight during the day, weak light at night, or even the complex lighting environment in tunnels, it can clearly capture the driver's facial information. For example, when driving at night, ordinary cameras may not be able to obtain clear images due to insufficient light, but the camera of this invention can use infrared supplementary lighting to achieve clear imaging of the driver's face using infrared light, providing a reliable data foundation for subsequent analysis.
[0041] High frame rate and wide field of view: The lens focal length is set at 2.38±0.3mm, achieving a frame rate of 30FPS. The high frame rate ensures the system can acquire continuous image data in real time, providing smooth playback like a movie without any stuttering or information loss. The wide field of view (FOV) of 62.4±3° horizontally and 38.8±2° vertically fully covers the driver's face, from forehead to chin, and from left ear to right ear, capturing no crucial details. Even slight head movements during driving are fully captured by the camera, providing comprehensive information for accurate analysis of the driver's condition.
[0042] High resolution and pixel advantage: Boasting a high resolution of 1280*800 and a pixel size of 3.0μm, the acquired images are clear and sharp. The high resolution clearly reveals subtle facial features of the driver, such as bloodshot eyes and mouth movements; while the larger pixel size improves image sensitivity and signal-to-noise ratio, reduces image noise interference, and further enhances image quality. These advantages provide accurate data for subsequent image analysis and processing, enabling the system to more accurately identify the driver's expressions, movements, and eye states.
[0043] High Dynamic Range (HDR) Technology: A high dynamic range of 68dB is another highlight of this camera. In environments with strong contrast between bright and low light, such as a vehicle suddenly entering bright sunlight from a dark tunnel, ordinary cameras may produce images that are either too bright or too dark, resulting in the loss of some information. However, the camera of this invention, with its high dynamic range characteristics, can simultaneously capture details in both bright and dark areas, ensuring that all information in the image is clearly visible. This effectively avoids information loss due to changes in lighting, ensuring the stability and reliability of image acquisition.
[0044] Employing advanced computer vision and machine learning algorithms, the system performs real-time analysis of collected driver facial data. Its main functions include:
[0045] Fatigue Monitoring Algorithm: For fatigue monitoring, the system employs a self-developed judgment scheme based on vehicle speed, Perclos (the percentage of time the eyes are closed), and the number of yawns. It comprehensively judges the driver's fatigue state by accurately calculating key indicators such as eye closure degree, blink frequency, and pupil size, combined with the characteristics of yawning. For example, when the vehicle speed is greater than 50 km / h, if the percentage of blinks within 60 seconds reaches 13%, or if there are 2 yawns within 120 seconds, the system will immediately determine that the driver is in a state of mild fatigue (KSS = 7). If the percentage of blinks within 60 seconds is ≥15%, or if there are 3 yawns within 120 seconds, it is considered moderate fatigue (KSS = 8). And if there are a cumulative total of 2 eye closures within 60 seconds with each closure lasting ≥0.6 seconds, or continuous eye closures lasting ≥1.2 seconds, it is considered a state of severe fatigue (KSS = 9). This is because, under normal driving conditions, a driver's blinking frequency and yawning frequency are within a certain range. When these indicators exceed the normal range, it is likely that the driver has begun to feel fatigued. As fatigue deepens, the proportion of blinking frames and the number of yawns will further increase, allowing the system to accurately determine this and providing a scientific basis for taking timely warning measures.
[0046] Diverse warning methods and user settings: Once mild fatigue is detected, the system will promptly notify the driver through various means, including voice, text, and graphic alerts on the instrument panel, opening windows, and activating hazard lights. Voice alerts are repeated twice to attract the driver's auditory attention; text and graphics are displayed for 6 seconds visually, allowing the driver to quickly obtain fatigue information. Moderate fatigue is detected, and windows will automatically open; severe fatigue is detected, and hazard lights will automatically activate. Simultaneously, the driver can choose to turn the alarm on or off on the instrument panel, allowing for convenient settings based on actual needs. For example, in certain special circumstances, such as when the driver needs a short rest but does not want to be frequently disturbed, the alarm function can be temporarily turned off; while during normal driving, the alarm function can be turned on to ensure driving safety.
[0047] The fatigue alarm control switch is set to two types: instrument and large screen.
[0048] Instrument cluster: 1 (Master switch for driver detection assist function);
[0049] Large screens: 4
[0050] 1. Turn on the driver detection assistance function - master switch;
[0051] 2. Turn on the fatigue detection sub-switch;
[0052] 3. Turn on the distraction detection sub-switch;
[0053] 4. Turn on the hazard detection sub-switch.
[0054] Distraction Detection Algorithm: Based on head position, turning angle, and gaze direction, the distraction detection algorithm analyzes captured images to accurately identify whether the driver is engaging in distracted behaviors such as looking left and right, checking their phone, or chatting. When the vehicle speed is between 25km / h and 50km / h, if the driver's head posture deflection angle exceeds a "reasonable angle" (≥45° to the left or right, or ≥30° up or down) and lasts for more than a certain threshold of 4.5 seconds (calibrable), it is considered minor distraction. When the vehicle speed is greater than 50km / h, if the driver's head posture deflection angle exceeds a "reasonable angle" for 3.5 seconds (calibrable), it is considered severe distraction. For example, when the driver turns their head to look out the window or looks down at their phone for an extended period, the system can quickly detect these abnormal actions and, based on set time and angle thresholds, issue a timely distraction warning. Furthermore, when the user requests a turn, the system automatically disables the distraction detection function to avoid misjudgments caused by normal driving operations, improving the accuracy and reliability of the monitoring.
[0055] Dangerous Action Detection Algorithm: The system's algorithm can accurately identify dangerous behavior patterns such as smoking and making phone calls. Through comprehensive analysis of the driver's hand movements, facial expressions, and object features, the system can detect whether the driver is smoking or making a phone call while driving. For example, when the system detects a hand gesture resembling picking up or lighting a cigarette, coupled with facial features resembling smoking motions, it immediately identifies it as smoking. Similarly, when the system detects the driver holding a phone close to their ear with corresponding head movements, it identifies it as making a phone call. Once these dangerous actions are detected, and the vehicle speed exceeds 50 km / h, the system will immediately issue a warning, reminding the driver to correct the dangerous behavior and ensuring driving safety.
[0056] Timely obstruction detection: Upon vehicle power-on, the system immediately checks for camera obstruction. If obstruction is detected, the system will promptly notify the driver via voice, text, and graphic alerts on the instrument panel. The voice alert is repeated twice, and the text is displayed on the screen for 6 seconds. This timely notification allows the driver to quickly identify any camera malfunctions and take appropriate measures to resolve the issue, ensuring the system remains in normal working order.
[0057] Key components work together: such as Figure 1-3As shown, the camera hardware, as the front-end acquisition device of the entire system, integrates several key components that work together to ensure high-quality and stable image acquisition. The serializer uses the Maxim Integrated MAX96701, which performs high-speed serial transmission of data acquired by the image sensor. During data transmission, it ensures data stability and efficiency, preventing data loss or transmission errors, much like an efficient transport vehicle on a highway, quickly and accurately transporting large amounts of image data to the next stage. The CMOS sensor selected is the SmartSens SC133AT, which features high sensitivity and low noise, enabling it to capture high-quality images in various environments. Whether in bright daylight or dim light at night, it acts like a keen eye, clearly capturing every detail of the driver's face.
[0058] Infrared Illumination and Optical Focusing: Two Brightek SF3838F94CQ01 LEDs provide ample infrared illumination, ensuring clear imaging of the driver's face even at night or in low-light conditions. These two LEDs act like miniature suns, providing extra light to the camera in the dark, making the driver's face clearly visible even in low-light environments. The LED driver IC, MPQ7230, precisely controls the brightness and operating status of the LEDs, automatically adjusting the brightness according to changes in ambient light to achieve optimal illumination. The lens uses a YTCM006, boasting excellent optical performance and accurate focusing to ensure clear, distortion-free images, much like a professional camera lens, clearly capturing every detail of the driver's face.
[0059] Power Management and System Integration: The PMIC (Power Management Integrated Circuit) is the BV8001, responsible for managing the camera's power supply and ensuring stable operation of all components. It acts like a smart manager, rationally allocating power to ensure each component receives sufficient power while preventing equipment failures due to unstable power. The entire camera utilizes an 8-layer FR-4 PCB design, which improves system integration and reliability. The multi-layer PCB tightly integrates various components, reducing wiring complexity and signal interference, while enhancing device stability and durability. Furthermore, the camera boasts IP52 dust and water resistance, allowing it to withstand harsh in-vehicle environments such as dusty compartments and accidental spills, ensuring continuous normal operation.
[0060] Controller Hardware: With its powerful core processing capabilities, the controller hardware serves as the system's core processing unit, undertaking crucial tasks such as data processing, analysis and decision-making, and signal transmission. Its core SOC, the OAX4600, possesses powerful computing capabilities, enabling rapid processing of massive amounts of image data captured by the cameras. Like a supercomputer, it can analyze and process massive amounts of image data in a short time, extracting useful information such as the driver's facial expressions and movements. The W35N02J Memory provides stable data storage support for the system, ensuring reliable data preservation and rapid retrieval. Like a vast warehouse, it can categorize and store data generated during system operation, and quickly retrieve this data when needed, supporting the system's analysis and decision-making.
[0061] System Control and Communication Coordination: The S32K312 MCU is selected to be responsible for the overall control and coordination of the system, enabling the collaborative work of various functional modules. It acts like a commander, directing the various components in the system to work in an orderly manner, ensuring the efficient operation of the entire system. The MAX96714F Deserializer converts serial data into parallel data for easier subsequent processing. It acts like a translator, translating the "foreign language" of serial data into the parallel data "language" that the system can understand, allowing the data to be processed smoothly within the system. The TJA1043 CAN transceiver enables communication between the controller and the vehicle's CAN network, allowing the system to acquire vehicle information such as speed and gear position, and send the monitoring results to instruments and other devices. Through the CAN network, the system can interact with other components of the vehicle, achieving data sharing and collaborative work.
[0062] Stable power management and protection: The MPQ7920 PMIC is responsible for the controller's power management, ensuring stable system operation under different operating voltages. It automatically adjusts the power supply according to system needs, ensuring stable operation under various working conditions. Also using an 8-layer FR4 PCB design, the controller has IP52-level dust and water resistance, effectively guaranteeing system stability and durability. This design allows the controller to operate normally in harsh environments, such as humid vehicle compartments and dusty road conditions, improving system reliability and lifespan.
[0063] System Connection and Communication: Stable connections between components are achieved through carefully designed connectors. The DMS camera and controller use connectors from the brand "Dianlian". The electrical terminal of the camera-to-controller connector is model 818019638 (B-pin, white), and the wiring harness terminal is model 818A92001B (B-pin, white). The electrical terminal of the controller-to-camera connector is model 818026543, and the outgoing cable terminal is model 818026555. PIN1 and PIN2 of these connectors are defined as POC+ and POC-, respectively, for data transmission and power supply. They act like bridges, tightly connecting the camera and controller, ensuring stable data and power transmission, and enabling the two components to work collaboratively.
[0064] Information interaction with the vehicle network: The controller connects to other vehicle devices via HRS brand GT25H2-8DP-2.2H (10) (electrical end) and GT25-8DS-HU / R (outgoing end) connectors to achieve communication with the CAN network. Through the CAN network, the system can acquire vehicle information such as vehicle speed, gear position, and door / lock status, and send monitoring results such as driver fatigue and attention status to the instrument cluster (IPM) and the rearview mirror (RRM) to achieve information interaction and sharing. This allows the system to perform comprehensive analysis based on the actual operating status of the vehicle and the driver's status, providing more accurate early warning information. For example, when the system detects that the driver is fatigued and the vehicle speed is high, it can send a signal to the instrument cluster, causing the instrument cluster to issue a stronger warning prompt, reminding the driver to rest in time and ensuring driving safety.
[0065] This invention utilizes an intelligent driver condition monitoring system. Through advanced image acquisition technology, precise data analysis algorithms, a rational system architecture design, and a wealth of practical functions, it achieves real-time monitoring and accurate early warning of driver fatigue, distraction, and dangerous actions. This system boasts advantages such as high reliability, high precision, and high adaptability, effectively improving driving safety and reducing traffic accidents.
[0066] Implementing the monitoring system in this solution on a real vehicle includes the following steps:
[0067] I. System Installation and Initialization
[0068] 1. Install the camera above the dashboard directly in front of the driver, ensuring that its field of view completely covers the driver's face.
[0069] 2. Connect the camera to the controller and access the vehicle network via the CAN bus.
[0070] 3. After the system is powered on, it will automatically initialize, detect whether the camera is blocked, and complete the algorithm loading.
[0071] II. Work Process
[0072] 1. Image Acquisition: The camera captures real-time images of the driver's face at a frame rate of 30 FPS and transmits them to the controller via a serializer.
[0073] 2. Data processing: The controller preprocesses the image data (such as noise reduction and enhancement) and extracts key features (such as the state of the eyes and mouth).
[0074] 3. Status determination:
[0075] (1) Fatigue monitoring: Calculate indicators such as PERCLOS and blink frequency to determine the fatigue level.
[0076] (2) Distraction monitoring: Analyze head posture and gaze direction to determine the degree of distraction.
[0077] (3) Monitoring dangerous actions: identifying behaviors such as smoking and making phone calls.
[0078] 4. Alarm Output: Based on the judgment result, an alarm signal is sent to the instrument and large screen via the CAN bus to trigger voice, text or graphic reminders.
[0079] III. User Settings
[0080] Drivers can set alarm switches via the instrument panel or large screen: the instrument panel has one master switch (to turn the driver detection assistance function on / off). The large screen has four sub-switches (master switch, fatigue detection, distraction detection, and dangerous action detection), which can be used to turn each sub-function on and off.
[0081] Obviously, the specific implementation of this invention is not limited to the above-described methods. Any non-substantial improvements made using the inventive concept and technical solution of this invention are within the protection scope of this invention.
Claims
1. An intelligent driver state monitoring and control system based on multimodal perception, characterized in that: It includes an infrared supplementary light camera, a control unit (MCU), and a warning unit; wherein the infrared supplementary light camera is used to collect facial image data of the driver, and its output is connected to the control unit (MCU). The control unit (MCU) identifies the current state of the driver based on the collected facial data and drives the warning unit to issue a warning message based on the identification result.
2. The intelligent driver state monitoring and control system based on multimodal perception as described in claim 1, characterized in that: The control unit MCU has built-in fatigue monitoring subunit, distraction monitoring subunit, and dangerous action monitoring subunit, which are used to monitor driver fatigue, distraction, and dangerous actions respectively and identify the corresponding results.
3. The intelligent driver state monitoring and control system based on multimodal perception as described in claim 1 or 2, characterized in that: The alarm unit includes an in-vehicle instrument panel, which issues voice, text, and graphic alerts. The alarm unit may include an on-board ECU, which promptly notifies the driver by opening the windows and activating the hazard lights.
4. The intelligent driver state monitoring and control system based on multimodal perception as described in claim 1 or 2, characterized in that: The control system also includes a fatigue alarm control switch, which is used to control and disable the driver status alarm signal.
5. A method for intelligent driver state monitoring and control based on multimodal perception, characterized in that: The system uses an infrared supplemental lighting camera to collect facial data from the driver and performs real-time analysis on the collected facial data. Based on the real-time analysis results of the driver's status, the system will issue an alarm or reminder. The driver's status analysis results include driver fatigue, driver distraction, and driver dangerous actions. Based on the driver's status analysis results, the system will issue corresponding alarm or reminders.
6. The intelligent driver state monitoring and control method based on multimodal perception as described in claim 5, characterized in that: Perclos data and yawn count data are obtained through facial recognition, and combined with the obtained vehicle speed to determine whether the current driver is fatigued and the level of fatigue; based on the fatigue status and fatigue level, the corresponding warning method is controlled to issue a warning to the driver.
7. The intelligent driver state monitoring and control method based on multimodal perception as described in claim 6, characterized in that: The driver's fatigue state and corresponding fatigue level are determined by using pre-set Perclos value thresholds and yawn count thresholds; the Perclos value threshold and yawn count threshold are set based on the current vehicle speed, with the threshold decreasing as the speed increases.
8. The intelligent driver state monitoring and control method based on multimodal perception as described in claim 6, characterized in that: Distraction monitoring includes identifying the driver's head position, turning angle, and gaze direction based on facial and head image data collected by cameras. When the head position deviates from the normal position or the head turning angle exceeds a set angle threshold and continues for a set time, it is determined that the driver is in a distracted state, and the control will issue a warning reminder corresponding to the distracted state.
9. The intelligent driver state monitoring and control method based on multimodal perception as described in claim 8, characterized in that: During distraction monitoring, real-time vehicle data is acquired to determine whether the vehicle is in a turning state. When the vehicle is in a turning state, the analysis and monitoring function is automatically disabled until the turning state ends.
10. The intelligent driver state monitoring and control method based on multimodal perception as described in claim 6, characterized in that: Dangerous action monitoring involves analyzing images captured by cameras to identify the driver's hand and facial movements, as well as the objects appearing in the image, to determine whether the driver is in a dangerous driving state, and issuing corresponding alarm signals based on the identification results.