Monitoring system and method for preventing driver from fatigue and distraction

By combining infrared and visible light dual cameras and image fusion algorithms, the system monitors and calculates the driver's status in real time, solving the problems of low recognition accuracy and limited field of view in existing systems under complex lighting conditions. This enables local intelligent processing and proactive vehicle intervention, thereby improving driving safety.

CN121492955APending Publication Date: 2026-02-10SINO TRUK JINAN POWER CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511981326.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-02-10

AI Technical Summary

Technical Problem

Existing driver fatigue and distraction monitoring systems have low accuracy in complex lighting conditions, lack proactive intervention capabilities, have limited field of vision, and cannot work in conjunction with vehicle control systems, thus affecting driving safety.

Method used

It employs a combination of infrared and visible light dual cameras, integrated inside the vehicle's A-pillar. Through image fusion and feature extraction algorithms, it monitors the driver's status in real time and performs parallel calculations in conjunction with vehicle bus information to generate warnings and vehicle intervention commands. The integrated data processing module inside the A-pillar enables local intelligent processing.

Benefits of technology

Accurately identifying driver facial features under complex lighting conditions reduces data transmission delays, enables proactive vehicle intervention, improves the accuracy and real-time performance of the monitoring system, and reduces the risk of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121492955A_ABST
    Figure CN121492955A_ABST
Patent Text Reader

Abstract

The invention provides a monitoring system and method for preventing a driver from fatigue and distraction, and belongs to the technical field of automobile intelligent cabins. The system comprises an infrared camera, a visible light camera and a data processing module; the infrared camera and the visible light camera respectively collect an infrared image and a visible light image of the face of a driver; the data processing module fuses the multi-source image to extract a driver state feature vector, performs parallel calculation through a preset fatigue evaluation and distraction identification model in combination with steering wheel steering frequency information, and judges whether the driver is in a fatigue or distraction abnormal state; when it is judged that the driver is in the fatigue or distraction abnormal state, an early warning instruction and a vehicle intervention instruction are generated; the communication module sends the early warning instruction and the vehicle intervention instruction to a vehicle-mounted alarm system and a vehicle driving assistance system respectively. According to the invention, double cameras are fused to improve the monitoring precision, local processing enhances the real-time performance, and active intervention guarantees the driving safety.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of automotive intelligent cockpit technology, specifically relating to a monitoring system and method for preventing driver fatigue and distraction. Background Technology

[0002] DMS is short for Driver Monitoring System.

[0003] Driver Monitoring System (DMS) is an active safety technology that uses sensors, cameras, and other devices to monitor the driver's condition in real time, aiming to prevent accidents caused by abnormal driving behaviors such as fatigue and distraction. With the development of intelligent and connected vehicles, DMS has become a key component of Advanced Driver Assistance Systems (ADAS) and the safety redundancy of future autonomous driving cockpits.

[0004] However, the practical application and deployment of DMS (Digital Monitoring System) have the following limitations: First, the recognition accuracy is low. In complex lighting environments such as strong light, backlight, or sudden changes in tunnel lighting, a single infrared camera struggles to accurately capture key facial points (such as eyelids and corners of the mouth), leading to decreased accuracy in recognition during fatigue or distraction, and frequent missed or false positives. Second, the data processing method is unreasonable. Existing DMS systems typically function only as data acquisition devices, transmitting raw images to the vehicle's central control system for processing, consuming central control capacity, increasing system response latency and the risk of data loss, and affecting the real-time performance of monitoring. Third, there is a lack of proactive intervention capabilities. After detecting driver abnormalities, existing systems only alert the driver through audible and visual alarms, without linking with the vehicle's Advanced Driver Assistance Systems (ADAS), and cannot proactively intervene in the vehicle (such as deceleration or stabilizing the driving trajectory), making it difficult to effectively reduce safety risks during long-distance driving or when the driver's reaction is slow. Finally, the field of view is limited. Traditional DMS cameras have a narrow field of view, failing to cover the driver's upper body area, and are prone to feature loss due to changes in the driver's posture.

[0005] Therefore, there is an urgent need for an integrated monitoring solution that can adapt to complex environments, achieve local intelligent processing, and coordinate with vehicle control systems to fundamentally improve driving safety. Summary of the Invention

[0006] In a first aspect, embodiments of this application provide a monitoring system for preventing driver fatigue and distraction, including an image acquisition module, a data processing module, and a communication module; The image acquisition module includes an infrared camera and a visible light camera; Infrared cameras are used to capture infrared images of the driver's face, while visible light cameras are used to capture visible light images of the driver's face. The data processing module is used to receive infrared and visible light images of the driver's face and is configured to perform the following functions: The received infrared and visible light images of the driver's face are processed in real time, and the driver's state feature vector, including eye state and head posture, is output through image fusion and feature extraction algorithms. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is judged to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. When the driver is determined to be in an abnormal state of fatigue or distraction, a warning command and a vehicle intervention command are generated. A communication module is used to send the warning command to the vehicle alarm system to trigger an audible and visual warning; The vehicle intervention command is sent to the vehicle driving assistance system to trigger active control operations on at least one of the braking system, electronic stability system, or hazard warning lights.

[0007] Furthermore, the infrared camera and the visible light camera are integrated and packaged inside the vehicle's A-pillar, and the interior panel area of ​​the vehicle's A-pillar corresponding to the optical axis of the lens of the infrared camera and the visible light camera has an optical window that allows infrared light and visible light to pass through. The optical window is made of infrared-transmitting acrylic sheet or infrared-transmitting polycarbonate sheet that is integrally formed or embedded with the A-pillar interior panel.

[0008] Furthermore, the infrared camera uses a wide-angle lens, and the visible light camera is equipped with an angle adjustment mechanism; The wide-angle lens, in conjunction with the angle adjustment mechanism, enables the combined field of view of the image acquisition module to cover the driver's face, shoulders, and part of the upper limbs.

[0009] Secondly, embodiments of this application also provide a monitoring method for preventing driver fatigue and distraction, applied to the system described in the first aspect, comprising the following steps: S1. Simultaneously acquire infrared and visible light images of the driver's face using an infrared camera and a visible light camera; S2. Perform real-time fusion processing on the infrared image and the visible light image to extract the driver state feature vector, which includes eye state and head posture. S3. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is determined to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. S4. When it is determined that the driver is in an abnormal state of fatigue or distraction, generate a warning command and a vehicle intervention command, and send the warning command to the vehicle alarm system to trigger an audible and visual warning, and send the vehicle intervention command to the vehicle driving assistance system to trigger active control operation on at least one of the braking system, electronic stability system or hazard warning lights.

[0010] Furthermore, the specific steps of step S1 are as follows: S11. Detect ambient light intensity using a light sensor; If the ambient light intensity is higher than the upper limit threshold or backlighting is detected, the dynamic exposure adjustment function of the visible light camera will be activated. S12. Control the infrared camera and visible light camera installed in the A-pillar of the vehicle to simultaneously acquire infrared and visible light images of the driver's face at the same time; S13. The ambient light intensity, the acquired infrared image, and the visible light image are transmitted to the data processing module located inside the A-pillar via a wire harness.

[0011] Furthermore, the specific steps of step S2 are as follows: S21. Obtain ambient light intensity; If the ambient light intensity is below the lower threshold, the weight of the infrared image is set higher than that of the visible light image; If the ambient light intensity is higher than the upper limit threshold or backlighting is detected, the weight of the visible light image is set higher than that of the infrared image. S22. The received infrared image is enhanced by the CLAHE algorithm. The enhanced infrared image is then registered with the visible light image based on feature points, and then weighted and fused according to the set weights. S23. A lightweight multi-task convolutional neural network is used to process the face fusion image, locate the face bounding box, and after correcting the face pose based on affine transformation, output the coordinates of five key feature points: eyes, nose tip, left and right corners of mouth. S24. Based on the coordinates of key feature points, calculate eye feature parameters and head posture parameters; the eye feature parameters include the percentage of eyelid closure per unit time. blinking frequency and blink duration The head posture parameters include pitch angle. With deflection angle ; Eyelid closure percentage among eye feature parameters The calculation formula is as follows:

[0012] in, Let N be the duration of the i-th instance of eyelid closure within the statistical period T, and N be the number of closures within the statistical period. blink frequency and blink duration By analyzing the relative position changes of the reflected light spots from the iris and cornea in consecutive frames, a threshold segmentation method was used to extract them. The head pose parameters are as follows: A 3D face model is constructed based on the coordinates of five key feature points: the eyes, the tip of the nose, and the left and right corners of the mouth. The pitch angle θ and yaw angle φ of the head are calculated using the Perspective-n-Point algorithm, with the following formula:

[0013]

[0014] in,( , , () represents the coordinates of the nose tip. , , ( ) represents the coordinates of the center points of both eyes.

[0015] Furthermore, it also includes an adaptive feature extraction step for occlusion scenes: When the confidence level of feature points in the lower half of the face detected based on the face detection results is consistently below the threshold, it is determined to be a scenario where a mask is being worn. In scenarios where drivers wear masks, the motion amplitude of the driver's brow bone region in consecutive frames is used as an auxiliary fatigue feature; the motion amplitude is obtained by calculating the variance of the Euclidean distance between the coordinates of the key points of the brow bone in the frames. When infrared image analysis identifies a uniform low-temperature area in the eye region that exceeds a set area threshold and has a temperature below a set temperature threshold, it is determined to be a scene where sunglasses are being worn. In scenarios where sunglasses are worn, a continuous sequence of infrared images is input into a pre-trained temporal convolutional network to analyze the thermal imaging change patterns in the eyelid region and output alternative eye opening and closing probabilities.

[0016] Furthermore, the specific steps of step S3 are as follows: S31. Based on eye feature parameters and head posture parameters, and combined with steering wheel frequency information obtained in real time from the vehicle bus, calculate the driver fatigue score. The specific steps are as follows: S311. Based on the proportion of eyelid closure Calculate the normalized value of the percentage of eyelid closure ; S312. Based on blink frequency With blink duration Calculate the blinking abnormality coefficient B: Determine if the following conditions are met: blink duration The blink rate is greater than the first time threshold or the blink frequency is less than the first frequency threshold. If so, B=1; If not, B=0; S313. Head-based pitch angle And the duration of hold, to determine the head posture abnormality coefficient H: Determine the absolute value of the head's pitch angle Whether the pitch angle is greater than a preset first angle threshold and whether the holding time of the current pitch angle exceeds a second time threshold; If so, H=1; If not, H=0; S314. Calculate the steering wheel steering frequency anomaly coefficient S based on the steering wheel turning frequency: Determine if the steering wheel frequency is less than the second frequency threshold; If so, S=1; If not, S=0; S32. Calculate the fatigue score using the following formula. :

[0017] in, , , , Preset weighting coefficients; S33. Perform at least one of the following distraction detections in parallel: The system uses a YOLOv8 model to detect whether a pre-defined distracting target exists in the driver's field of vision, and simultaneously monitors the head yaw angle. When the confidence level of the distracted target detection is greater than or equal to the preset confidence threshold, and the absolute value of the head deflection angle is... When the duration of a state ≥ the preset second angle threshold is greater than the third time threshold, it is determined to be visual distraction; Based on the MediaPipe hand keypoint detection model, the coordinates of hand keypoints are obtained, and the distance between the hand and the face and the proportion of the occlusion area of ​​the hand on the face are calculated. When the distance between the hand and the face is less than the distance threshold and the occlusion area exceeds the preset proportion for a duration greater than the fourth time threshold, it is determined that the hand is centered. Based on the LSTM network model, the facial expression features represented by the sequence of facial key points are analyzed; when the distraction probability value output by the model is greater than or equal to the second confidence threshold and the duration of the state is greater than or equal to the third time threshold, it is determined to be cognitive distraction. S34. If fatigue score If the threshold is exceeded, it is determined to be a fatigue state; if any distraction detection is valid, it is determined to be the corresponding distraction state.

[0018] Furthermore, the YOLOv8s model in step S33 is obtained through the following steps: SS1. Collect raw images of the driver's cab scene by shooting real vehicles and crawling public network data, and manually annotate the collected raw images by adding bounding boxes and category labels to objects in the images that belong to preset categories, forming labeled image data; SS2. Randomly divide the labeled image data into training and validation subsets, normalize the size of the images in the training and validation subsets, and convert the labels into label format files that meet the input requirements of the YOLOv8s model; SS3. Load the weights of the YOLOv8s model pre-trained on a general dataset as the initial model, freeze the weights of the first N convolutional layers of the backbone network of the initial model, use the training subset as input, use stochastic gradient descent as the optimizer, and use a preset initial learning rate to iteratively train the unfrozen layers in the initial model to obtain an intermediate model. SS4. Use the validation subset to evaluate the intermediate model, calculate the mean accuracy, and when the mean accuracy exceeds the preset qualified threshold, solidify the corresponding intermediate model parameters into the final YOLOv8s model weight file and deploy it to the data processing module. The LSTM network model used for cognitive distraction identification in step S33 is obtained through the following steps: ST1. Collect facial video data of drivers under normal driving and cognitive distraction states, including inattentiveness and lack of response; extract the coordinates of facial key points in each video frame, generate a time sequence as training samples, and label the corresponding state labels. ST2. Using the time sequence as input and the corresponding distraction state as a supervision signal, a supervised training of an LSTM network model is performed until it can predict the distraction probability based on the input sequence. ST3. Deploy the trained LSTM network model in the data processing module, and during real-time monitoring, input the time sequence of facial key point coordinates from multiple consecutive frames into the LSTM model, and output a probability value representing the driver's state of cognitive distraction.

[0019] Furthermore, the specific steps of step S4 are as follows: S41. When the driver is in an abnormal state of fatigue or distraction, generate a first-level warning instruction; S42. If the abnormal state continues for more than the preset waiting time, a second-level vehicle intervention command is generated. S43. Send the warning command to the vehicle alarm system to trigger an audible and visual alarm to alert the driver; S44. Send a vehicle intervention command to the vehicle driver assistance system to trigger at least one of the following operations: The system controls the braking system to slightly reduce speed, enhances lane keeping assist torque through the electronic stability system, and automatically activates hazard warning lights.

[0020] As can be seen from the above technical solutions, this application has the following advantages: The monitoring system and method for preventing driver fatigue and distraction provided in this application effectively reduce the risk of traffic accidents caused by driver fatigue or distraction by accurately monitoring the driver's fatigue and distraction state, issuing timely warnings, and intervening in the vehicle, thus ensuring driving safety. It employs a dual-camera setup combined with image fusion and feature extraction algorithms, enabling accurate identification of the driver's facial features and state under complex lighting conditions, improving the accuracy and reliability of monitoring compared to a single-camera system. The data processing module is integrated into the vehicle's A-pillar for local intelligent processing, reducing reliance on the vehicle's central control system, lowering data transmission latency and packet loss risks, ensuring rapid feedback of monitoring results, and improving the system's real-time performance. The driver assistance system works in conjunction with the vehicle to detect abnormal driver behavior, issuing not only audible and visual warnings but also actively controlling the vehicle's braking system, electronic stability system, or hazard warning lights. This effectively reduces safety risks and overcomes the shortcomings of existing systems that rely solely on audible and visual alarms. This application is adaptable to various scenarios and features adaptive feature extraction for obstructed scenes, enabling it to handle situations where the driver is wearing a mask or sunglasses, thus expanding the system's applicability and allowing it to function normally in more complex scenarios. This application integrates and encapsulates hardware such as cameras within the vehicle's A-pillar using a concealed installation method. This avoids interfering with the driver's view, maintains the aesthetics of the cockpit interior, reduces external light interference with the camera, and improves the stability of image acquisition. Attached Figure Description

[0021] To more clearly illustrate the technical solution of this application, the accompanying drawings used in the description will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0022] Figure 1 This is a schematic diagram of the monitoring system for preventing driver fatigue and distraction according to the present invention.

[0023] Figure 2 This is a schematic diagram of the installation of the monitoring system for preventing driver fatigue and distraction according to the present invention.

[0024] Figure 3 This is a flowchart illustrating the monitoring method for preventing driver fatigue and distraction according to the present invention. Detailed Implementation

[0025] Various embodiments of this disclosure will be described more fully in the following detailed description of the specific steps of a monitoring system for preventing driver fatigue and distraction. This disclosure may have various embodiments, and adjustments and changes may be made therein. However, it should be understood that there is no intention to limit the various embodiments of this disclosure to the specific embodiments disclosed herein, but rather this disclosure should be understood to cover all adjustments, equivalents, and / or alternatives falling within the spirit and scope of the various embodiments of this disclosure.

[0026] This embodiment provides a monitoring system to prevent driver fatigue and distraction. Dual-camera fusion improves monitoring accuracy in complex environments, local intelligent processing reduces latency, and linkage with the driving assistance system enables proactive intervention to ensure driving safety.

[0027] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0028] Please see Figure 1 The diagram shown is a schematic of a monitoring system for preventing driver fatigue and distraction in a specific embodiment. The system includes an image acquisition module, a data processing module, and a communication module. The image acquisition module includes an infrared camera and a visible light camera; Infrared cameras are used to capture infrared images of the driver's face, while visible light cameras are used to capture visible light images of the driver's face. It should be noted that the infrared camera can clearly capture the driver's facial contours and eye movements in low light or no light environments, providing a reliable data source for monitoring at night or in low light conditions, and improving the system's monitoring capabilities under low light conditions. Visible light cameras can capture visible light images of the driver's face. In complex lighting scenarios such as strong light and backlight, clear facial details can be obtained through dynamic exposure adjustment, which makes up for the shortcomings of infrared cameras in strong light and enhances the system's adaptability under different lighting conditions. The data processing module is used to receive infrared and visible light images of the driver's face and is configured to perform the following functions: The received infrared and visible light images of the driver's face are processed in real time, and the driver's state feature vector, including eye state and head posture, is output through image fusion and feature extraction algorithms. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is judged to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. When the driver is determined to be in an abnormal state of fatigue or distraction, a warning command and a vehicle intervention command are generated. It should be noted that the data processing module is integrated inside the vehicle's A-pillar, enabling real-time processing and analysis of the collected image data. Through image fusion and feature extraction algorithms, it can accurately output the driver's state feature vector and combine it with vehicle bus information to perform fatigue assessment and distraction identification, reducing reliance on the vehicle's central control system, improving the system's real-time performance and reliability, and reducing latency and packet loss risks during data transmission. A communication module is used to send the warning command to the vehicle alarm system to trigger an audible and visual warning; The vehicle intervention command is sent to the vehicle driving assistance system to trigger active control operations on at least one of the braking system, electronic stability system, or hazard warning lights; It should be noted that the communication module sends warning commands and vehicle intervention commands to the corresponding systems, realizing effective linkage between the monitoring system and other vehicle systems. This ensures that warning and intervention operations can be executed in a timely and accurate manner, enhancing the system's active safety performance and improving driving safety.

[0029] This embodiment achieves real-time monitoring and accurate judgment of the driver's status, as well as corresponding early warning and vehicle intervention functions through the coordinated work of various modules, thus ensuring driving safety.

[0030] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process in this embodiment, another monitoring system for preventing driver fatigue and distraction is provided. This system includes an image acquisition module, a data processing module, and a communication module. The image acquisition module includes an infrared camera and a visible light camera. The infrared camera and the visible light camera are integrated and packaged inside the A-pillar of the vehicle, and optical windows that allow infrared light and visible light to pass through are opened in the interior panel area of ​​the A-pillar of the vehicle in front of the optical axis of the lens of the infrared camera and the visible light camera. The optical window is made of a material that allows infrared and visible light to pass through, ensuring that the camera can capture images normally. At the same time, it matches the A-pillar interior panel, so as not to affect the overall aesthetics of the vehicle interior, and also provides a certain degree of protection for the camera. Infrared cameras are used to capture infrared images of the driver's face, while visible light cameras are used to capture visible light images of the driver's face. For example, the infrared camera uses a near-infrared imaging module of model IMX390, equipped with a 120° wide-angle lens (focal length 2.8mm, aperture F1.6), supports 850nm band infrared light acquisition, and can output clear facial contour images with a resolution of 1920×1080 in low light environment of 0-1000lux, focusing on capturing key features such as eyelids and brow bones; The visible light camera uses the IMX586 high-definition module and is equipped with an electric angle adjustment mechanism (adjustment range ±15°). It supports dynamic exposure adjustment (exposure time 1ms-100ms adaptive). In strong light and backlight scenes above 10000 lux, it can automatically adjust the sensitivity (ISO 100-3200) to obtain facial detail images without overexposure or glare. The dual cameras capture data at a frame rate of 30fps. Through the cooperation of the wide-angle lens and the angle adjustment mechanism, the combined field of view reaches 70°×70°, which can completely cover the driver's face, shoulders and part of the upper limb area (the acquisition range covers 50cm horizontally and 70cm vertically), effectively avoiding feature loss due to slight changes in the driver's sitting posture. The data processing module is located inside the vehicle's A-pillar, receives infrared and visible light images of the driver's face, and is configured to perform the following functions: The received infrared and visible light images of the driver's face are processed in real time, and the driver's state feature vector, including eye state and head posture, is output through image fusion and feature extraction algorithms. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is judged to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. When the driver is determined to be in an abnormal state of fatigue or distraction, a warning command and a vehicle intervention command are generated. A communication module is used to send the warning command to the vehicle alarm system to trigger an audible and visual warning; The vehicle intervention command is sent to the vehicle driving assistance system to trigger active control operations on at least one of the braking system, electronic stability system, or hazard warning lights; The optical window is made of infrared-transmitting acrylic sheet or infrared-transmitting polycarbonate sheet that is integrally formed or embedded with the A-pillar interior panel. For example, an infrared-transmitting polycarbonate panel integrally formed with the A-pillar interior panel is used, with a thickness of 2mm, an infrared light transmittance of ≥92%, a visible light transmittance of ≥88%, and an anti-glare treatment on the surface (haze ≤0.5%), which ensures light penetration, avoids reflection interference with image acquisition, and is consistent with the interior style. The infrared camera uses a wide-angle lens, while the visible light camera has an angle adjustment mechanism; The wide-angle lens, in conjunction with the angle adjustment mechanism, enables the combined field of view of the image acquisition module to cover the driver's face, shoulders, and part of the upper limbs. Specifically, the total field of view of the infrared camera and the visible light camera is not less than 70°×70°; It also includes a BH1750 digital light sensor with a measurement range of 0-65535 lux and a measurement accuracy of ±20%, used to detect ambient light intensity in real time and provide a trigger signal for switching the working mode of the dual cameras; it also integrates a temperature sensor with a measurement range of -20℃ to 85℃, used to help determine the obstruction scene such as wearing sunglasses. For example, the data processing module uses a hardware motherboard (model RK3568) with an integrated SOC chip, equipped with a quad-core Cortex-A55 processor (1.8GHz), a Mali-G52 2EE GPU, 8GB eMMC storage chip and 2GB LPDDR4 memory, and supports local inference of lightweight deep learning models; the motherboard integrates an LVDS image interface (for connecting to dual cameras), a CAN FD bus interface (for connecting to the vehicle bus), and a USB 3.0 interface (for debugging and upgrading), which can directly receive raw image data collected by the camera and complete the entire process of feature extraction, model calculation and other processing locally, without relying on the vehicle central control system, reducing data transmission latency and packet loss risk; The communication module adopts a CAN FD bus communication module (communication rate up to 8Mbps, supports SAE J1939 protocol) to realize bidirectional data interaction with the vehicle alarm system and the vehicle driver assistance system (ADAS); it also integrates a Bluetooth 5.0 module to support short-range communication with mobile APP or diagnostic equipment for system parameter configuration, fault diagnosis and log export. like Figure 2 As shown, the hardware module of the monitoring system for preventing driver fatigue and distraction in this application is fixed inside the A-pillar by a customized bracket. The optical axis of the lens forms a 15° angle with the center of the driver's face, and the optical window is flush with the interior panel of the A-pillar without any protruding structure. After installation, it is connected to the vehicle's 12V power supply (power consumption ≤15W) through a wiring harness. The data processing module is connected to the vehicle network through the CAN bus to achieve linkage with other vehicle systems.

[0031] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0032] like Figure 3 As shown, the following are embodiments of the monitoring method for preventing driver fatigue and distraction provided in this disclosure. This method and the monitoring system for preventing driver fatigue and distraction in the above embodiments belong to the same inventive concept. For details not described in detail in the embodiments of the monitoring method for preventing driver fatigue and distraction, please refer to the embodiments of the monitoring system for preventing driver fatigue and distraction described above.

[0033] The method includes the following steps: S1. Simultaneously acquire infrared and visible light images of the driver's face using an infrared camera and a visible light camera; It should be noted that the simultaneous acquisition of infrared and visible light images can obtain information about the driver's face under different spectra, providing a data foundation for image fusion and status analysis. Infrared images perform well in low light or no light conditions, while visible light images provide clear details under normal lighting conditions. The combination of the two can effectively cope with various complex lighting scenarios such as strong light, backlight, and low light, improving environmental adaptability. Simultaneous acquisition ensures the temporal correspondence between the two types of images, which facilitates subsequent image registration and fusion processing, avoids errors caused by time differences, and improves monitoring accuracy. S2. Perform real-time fusion processing on the infrared image and the visible light image to extract the driver state feature vector, which includes eye state and head posture. It should be noted that by using image fusion algorithms, the advantages of infrared and visible light images are complemented to enhance the overall image quality and information content, and improve the recognition effect of facial features, especially in complex lighting or occlusion scenarios, it can accurately extract key features. Using algorithms such as multi-task convolutional neural networks, the bounding box of the face is accurately located and key feature points (such as the eyes, nose tip, and corners of the mouth) are extracted, providing accurate physiological feature parameters for subsequent fatigue and distraction state judgment. The fused image data can more comprehensively reflect the driver's state, reduce the misjudgment or omission that may be caused by a single image, and improve the accuracy of monitoring fatigue and distraction state. S3. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is determined to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. It should be noted that by combining the driver's physiological characteristics and driving behavior data, the driver's state is assessed from multiple dimensions, making the judgment more comprehensive and accurate, and avoiding misjudgments caused by single-dimensional data. Parallel computing is used to simultaneously perform fatigue assessment and distraction identification, improving processing efficiency and real-time performance, enabling rapid response to changes in the driver's state and timely issuance of warnings or intervention commands. Through pre-trained fatigue assessment and distraction identification models, it is possible to accurately determine whether the driver is in a state of fatigue or distraction, providing a basis for warning and intervention measures and reducing the risk of traffic accidents. S4. When it is determined that the driver is in an abnormal state of fatigue or distraction, generate a warning command and a vehicle intervention command, and send the warning command to the vehicle alarm system to trigger an audible and visual warning, and send the vehicle intervention command to the vehicle driving assistance system to trigger active control operation on at least one of the braking system, electronic stability system or hazard warning lights. It should be noted that when driver fatigue or distraction is detected, a warning command is immediately generated and an audible and visual alarm is triggered. This quickly reminds the driver to pay attention to their condition, adjust their behavior in time, and avoid dangerous driving behaviors caused by fatigue or distraction. In conjunction with the vehicle's driver assistance system, depending on the severity of the abnormal state, it actively triggers actions such as braking and deceleration, enhanced lane keeping assist, or activation of hazard warning lights, directly intervening in the vehicle's driving status, reducing the risk of accidents, and providing safety assurance. Based on the duration and severity of the abnormal state, warning and intervention commands are generated in stages, and system resources are rationally allocated. This not only promptly reminds the driver but also allows for decisive intervention when necessary, ensuring driving safety while avoiding excessive intervention that could affect normal driving.

[0034] This embodiment integrates infrared and visible light dual cameras inside the vehicle's A-pillar. Through image fusion and feature extraction algorithms, it improves facial recognition accuracy under complex lighting conditions. The local data processing module enables rapid analysis, reducing the burden on the central control system and providing strong real-time performance. Combined with the active intervention of the driver assistance system, it reduces the risk of fatigue and distraction, and improves driving safety.

[0035] Furthermore, as a refinement and extension of the specific implementation of the above embodiments, in order to fully illustrate the specific implementation process of this embodiment, another monitoring method for preventing driver fatigue and distraction is provided, which includes the following steps: S1. Simultaneously acquire infrared and visible light images of the driver's face using an infrared camera and a visible light camera; the specific steps of step S1 are as follows: S11. Detect ambient light intensity using a light sensor; If the ambient light intensity is higher than the upper limit threshold or backlighting is detected, the dynamic exposure adjustment function of the visible light camera will be activated. S12. Control the infrared camera and visible light camera installed in the A-pillar of the vehicle to simultaneously acquire infrared and visible light images of the driver's face at the same time; S13. The ambient light intensity, the acquired infrared image, and the visible light image are transmitted to the data processing module located inside the A-pillar via a wire harness; For example, the light sensor detects the ambient light intensity in real time, setting an upper limit threshold of 5000 lux and a lower limit threshold of 100 lux. When the ambient light intensity is greater than 5000 lux or backlighting is detected through image brightness gradient analysis (image contrast ratio > 8:1), the dynamic exposure adjustment function of the visible light camera is activated, shortening the exposure time to 1ms-10ms and reducing the sensitivity to below ISO 400 to avoid overexposure of facial details. When the ambient light intensity is less than 100 lux, the visible light camera stops dynamic adjustment, and the infrared camera automatically turns on the near-infrared fill light (3W power, illumination distance 1-2m) to ensure image acquisition effect in low-light environments.

[0036] The system controls dual cameras to simultaneously acquire infrared and visible light images at the same time, with a timestamp error of ≤1ms, ensuring the time consistency of the two types of images. After acquisition, the image data (in RAW10 format) and ambient light intensity data are transmitted to the data processing module in real time via LVDS cable, with a transmission delay of ≤20ms. S2. Perform real-time fusion processing on the infrared image and the visible light image to extract the driver's state feature vector, including eye state and head posture; the specific steps of step S2 are as follows: S21. Obtain ambient light intensity; If the ambient light intensity is below the lower threshold, the weight of the infrared image is set higher than that of the visible light image; If the ambient light intensity is higher than the upper limit threshold or backlighting is detected, the weight of the visible light image is set higher than that of the infrared image. S22. The received infrared image is enhanced by the CLAHE algorithm. The enhanced infrared image is then registered with the visible light image based on feature points, and then weighted and fused according to the set weights. S23. A lightweight multi-task convolutional neural network is used to process the face fusion image, locate the face bounding box, and after correcting the face pose based on affine transformation, output the coordinates of five key feature points: eyes, nose tip, left and right corners of mouth. The backbone network of the multi-task convolutional neural network is MobileNetV3 optimized by channel pruning, and its output layer connects the face detection branch and the facial landmark regression branch in parallel. S24. Based on the coordinates of key feature points, calculate eye feature parameters and head posture parameters; the eye feature parameters include the percentage of eyelid closure per unit time. blinking frequency and blink duration The head posture parameters include pitch angle. With deflection angle ; Eyelid closure percentage among eye feature parameters The calculation formula is as follows:

[0037] in, The duration of the i-th time within the statistical period T that is determined to be eyelid closure (e.g., closure degree ≥ 80%), and N is the number of closures within the statistical period; blink frequency and blink duration By analyzing the relative position changes of the reflected light spots from the iris and cornea in consecutive frames, a threshold segmentation method was used to extract them. The head pose parameters are as follows: A 3D face model is constructed based on the coordinates of five key feature points: the eyes, the tip of the nose, and the left and right corners of the mouth. The pitch angle θ and yaw angle φ of the head are calculated using the Perspective-n-Point algorithm, with the following formula:

[0038]

[0039] in,( , , () represents the coordinates of the nose tip. , , () represents the coordinates of the center points of both eyes; For example, after receiving ambient light intensity data, the data processing module automatically assigns image fusion weights: when the ambient light intensity is <100 lux, the weight of the infrared image is set to 0.7 and the weight of the visible light image is set to 0.3; when the ambient light intensity is >5000 lux or in backlight, the weight of the visible light image is set to 0.8 and the weight of the infrared image is set to 0.2; when the ambient light intensity is between 100 lux and 5000 lux, the weights of both images are set to 0.5, thus achieving complementary fusion under different lighting conditions. The CLAHE algorithm is used to enhance the contrast of infrared images, with a block size of 8×8 pixels and a contrast limit parameter of 2.0, improving the clarity of facial contours in low-light environments. Gaussian filtering (3×3 kernel, standard deviation 1.0) is used to denoise visible light images, removing noise interference from strong light conditions. Key feature points (at least 50 matching points) are extracted from the two images using the SIFT feature point detection algorithm. Mismatches are eliminated using the RANSAC algorithm (1000 iterations, interior point threshold of 2 pixels), completing image registration with a registration error ≤ 1 pixel. Subsequently, weighted fusion is performed according to preset weights to generate a fused image that combines contour clarity with rich detail. A multi-task convolutional neural network with channel-pruned optimization as the backbone is used to process the fused image. The network includes a face detection branch and a facial landmark regression branch. The face detection branch locates the face bounding box using the anchor box clustering algorithm (IOU threshold 0.5), and the landmark regression branch outputs the coordinates of five key feature points: eyes, nose tip, left and right corners of mouth, with a localization accuracy of ≤2 pixels. Affine transformation is used to correct the face pose, eliminating the influence of pitch and yaw angles (the corrected pose deviation is ≤3°), ensuring the accuracy of feature parameter calculation. Feature parameter calculation: Eye characteristic parameters include eyelid closure percentage, blink frequency, and blink duration. The eyelid closure percentage is calculated by dividing the total duration of eyelid closure (≥80% closure) within a statistical period T (default 60 seconds) by T, using the following formula: ,in Let N be the duration of the i-th blink, and N be the number of blinks within the statistical period. The blink frequency and blink duration were extracted by analyzing the relative position changes of the reflected light spots of the iris and cornea in consecutive frames using the Otsu threshold segmentation method. The statistical accuracy of the blink frequency was 0.1 blinks / minute, and the measurement accuracy of the blink duration was 10ms. Head pose parameters: A 3D face model is constructed based on the coordinates of five key feature points. The pitch angle θ and yaw angle φ are calculated using the Perspective-n-Point algorithm. The calculation formula is as follows: , , in,( , , () represents the coordinates of the nose tip. , , ( ) represents the coordinates of the center points of both eyes, and the angle calculation accuracy is ≤0.5°; S3. Based on the driver's state feature vector and combined with the steering wheel frequency information obtained from the vehicle bus, a pre-set fatigue assessment model and a distraction recognition model are used for parallel calculation to determine whether the driver is in an abnormal state of fatigue or distraction; the specific steps of step S3 are as follows: S31. Based on eye feature parameters and head posture parameters, and combined with steering wheel frequency information obtained in real time from the vehicle bus, calculate the driver fatigue score. The specific steps are as follows: S311. Based on the proportion of eyelid closure Calculate the normalized value of the percentage of eyelid closure ; S312. Based on blink frequency With blink duration Calculate the blinking abnormality coefficient B: Determine if the following conditions are met: blink duration The blinking frequency is greater than the first time threshold (e.g., 0.5s) or less than the first frequency threshold (e.g., 5 times / minute). If so, B=1; If not, B=0; S313. Head-based pitch angle And the duration of hold, to determine the head posture abnormality coefficient H: Determine the absolute value of the head's pitch angle Whether the pitch angle is greater than a preset first angle threshold (e.g., 20 degrees) and the holding time of the current pitch angle exceeds a second time threshold (e.g., 3 seconds). If so, H=1; If not, H=0; S314. Calculate the steering wheel steering frequency anomaly coefficient S based on the steering wheel turning frequency: Determine whether the steering wheel turning frequency is less than the second frequency threshold (e.g., 0.5 times / minute). If so, S=1; If not, S=0; S32. Calculate the fatigue score using the following formula. :

[0040] in, , , , Preset weighting coefficients; S33. Perform at least one of the following distraction detections in parallel: The system uses a YOLOv8 model to detect whether a pre-defined distracting target exists in the driver's field of vision, and simultaneously monitors the head yaw angle. When the confidence level of the distracted target detection is greater than or equal to the preset confidence threshold (e.g., 0.7), and the absolute value of the head deflection angle is... When the duration of a state ≥ a preset second angle threshold (e.g., 30 degrees) is greater than a third time threshold (e.g., 2 seconds), it is determined to be visual distraction; The YOLOv8s model in step S33 is obtained through the following steps: SS1. Collect raw images of the driver's cab scene by shooting real vehicles and crawling public network data, and manually annotate the collected raw images by adding bounding boxes and category labels to objects in the images that belong to preset categories, forming labeled image data; The preset categories include at least: mobile phones, tablets, paper documents, books, and food containers; SS2. Randomly divide the labeled image data into training and validation subsets, normalize the size of the images in the training and validation subsets, and convert the labels into label format files that meet the input requirements of the YOLOv8s model; SS3. Load the weights of the YOLOv8s model pre-trained on a general dataset as the initial model, freeze the weights of the first N convolutional layers of the backbone network of the initial model (where N is an integer greater than 100), use the training subset as input, use stochastic gradient descent as the optimizer, and use a preset initial learning rate (e.g., 1e-3) to iteratively train the unfrozen layers in the initial model to obtain an intermediate model. SS4. Use the validation subset to evaluate the intermediate model, calculate the mean accuracy, and when the mean accuracy exceeds the preset qualified threshold (e.g., 0.85), solidify the corresponding intermediate model parameters into the final YOLOv8s model weight file and deploy it to the data processing module. Based on the MediaPipe hand keypoint detection model, the coordinates of hand keypoints are obtained, and the distance between the hand and the face and the proportion of the face occluded by the hand are calculated. When the distance between the hand and the face is less than the distance threshold (e.g., 30cm) and the proportion of the occluded area exceeds the preset proportion (e.g., 15%) for a duration greater than the fourth time threshold (e.g., 1.5 seconds), it is determined that the hand is centered. Specifically, the MediaPipe Hands model's hand keypoint detection API is called to obtain the 3D coordinates of 21 keypoints on the hand; the Euclidean distance between the wrist keypoint and the nose tip keypoint is calculated, and the occlusion area ratio of the polygon formed by the hand keypoints to the facial bounding box is calculated; when the distance is <30cm and the occlusion area ratio is >15% for ≥1.5 seconds, it is determined to be the hand centering. The hand center detection function is implemented by calling the application programming interface of the MediaPipe Hands model. MediaPipe is an open-source cross-platform machine learning framework. Its Hands model provides an efficient and real-time 21-point hand keypoint detection model. This model has been fully pre-trained on a large-scale public hand dataset and has good generalization ability and robustness. It can accurately locate the three-dimensional coordinates (x, y, z) of 21 keypoints of the hand directly from a single frame of RGB image, where z represents the relative depth. This application uses the driver's facial region image frames captured by a visible light camera as input to the model, calls the MediaPipe Hands model for inference, obtains the hand key point coordinates output by the model, and then calculates two core metrics using a specific algorithm: Distance between hand and face: obtained by calculating the three-dimensional Euclidean distance between wrist key points (e.g., LANDMARK_WRIST) and facial reference points (e.g., the tip of the nose); The proportion of the area obscured by the hand on the face is obtained by connecting the key points of the hand to form an approximate polygon of the hand outline, and calculating the proportion of the overlapping area between the polygon and the pre-calibrated bounding box of the driver's face. The data obtained from the above calculations will be compared with a pre-set threshold to serve as a direct basis for determining whether the driver is engaging in distracted hand behaviors such as smoking or making phone calls. Based on the LSTM network model, facial expression features represented by the sequence of facial key points are analyzed. When the distraction probability value output by the model is greater than or equal to the second confidence threshold (e.g., 0.85) and the duration of this state is greater than or equal to the third time threshold (e.g., 3 seconds), it is determined to be cognitive distraction. The LSTM network model used for cognitive distraction identification is obtained through the following steps: ST1. Collect facial video data of drivers under normal driving and cognitive distraction states, including inattentiveness and lack of response; extract the coordinates of facial key points in each video frame, generate a time sequence as training samples, and label the corresponding state labels. ST2. Using the time sequence as input and the corresponding distraction state as a supervision signal, a supervised training of an LSTM network model is performed until it can predict the distraction probability based on the input sequence. ST3. Deploy the trained LSTM network model in the data processing module, and during real-time monitoring, input the time sequence of facial key point coordinates from multiple consecutive frames into the LSTM model, and output a probability value representing the driver's state of cognitive distraction. S34. If fatigue score If the threshold is exceeded, it is determined to be a fatigue state; if any distraction detection is valid, it is determined to be the corresponding distraction state. For example, fatigue assessment model: eyelid closure percentage normalization: when the eyelid closure percentage P≥20%, the normalized value Q=1; when P≤10%, Q=0; when 10%<P<20%, ​​Q=2P-0.2 (linear interpolation). Abnormal coefficient calculation: Blinking abnormal coefficient B: The first time threshold is set to 0.5s, and the first frequency threshold is set to 5 times / minute. When the blinking duration is >0.5s or the blinking frequency is <5 times / minute, B=1; otherwise, B=0. Head posture abnormality coefficient H: The first angle threshold is preset to 20° and the second time threshold is set to 3s. When the absolute value of the head pitch angle is >20° and the duration is >3s, H=1; otherwise, H=0. Steering wheel frequency anomaly coefficient S: Steering wheel data is obtained in real time from the vehicle CAN bus to calculate the number of steering turns per minute. The second frequency threshold is set to 0.5 times / minute. When the steering frequency is <0.5 times / minute, S=1; otherwise, S=0. Fatigue score calculation: preset weighting coefficients =0.4、 =0.3、 =0.2、 =0.1, according to the formula Calculate fatigue score, when When the value is ≥0.6, it is considered a state of fatigue; Distraction detection model: Visual distraction recognition: The YOLOv8s model is used, which is trained through the following steps: Data collection and annotation: 10,000 original images of the driver's cab scene are collected through real vehicle photography and public network data crawling, covering distracting targets such as mobile phones, tablets, paper documents, books, and food containers. Bounding boxes and category labels are manually annotated to form an annotated dataset. Data preprocessing: The dataset was divided into training and validation subsets in an 8:2 ratio, the images were normalized to 640×640 pixels, and the labels were converted to YOLO format (category ID, center x / y coordinates, width and height). Model training: Load the pre-trained YOLOv8s weights from the COCO dataset, freeze the first 120 convolutional layers of the backbone network, use stochastic gradient descent as the optimizer (initial learning rate 1e-3, momentum 0.9, weight decay 0.0005), iterate for 100 epochs, and validate every 10 epochs. Model deployment: When the mean average precision (mAP) of the validation set is ≥0.85, the model weights are fixed and deployed to the data processing module. During detection, when the detection confidence of a distracted target is ≥0.7 and the absolute value of the head deflection angle is ≥30° (preset second angle threshold) for a duration >2s (third time threshold), it is determined to be visual distraction. Hand centrifugation recognition: The MediaPipe Hands model is called to obtain the 3D coordinates of 21 key points on the hand. The Euclidean distance between the wrist key point (LANDMARK_WRIST) and the nose tip key point is calculated, as well as the proportion of the occlusion area of ​​the polygon formed by the hand key points on the facial bounding box. When the distance is <30cm (distance threshold) and the occlusion area proportion is >15% (preset proportion) for a duration >1.5s (fourth time threshold), it is determined as hand centrifugation (such as smoking, answering the phone, etc.). Cognitive Distraction Detection: An LSTM network model is used, trained and deployed through the following steps: Data collection: Facial video data of 50 drivers were collected under normal driving and cognitive distraction (distraction, lack of reaction, etc.). Each video was 3 minutes long. The coordinates of facial key points in each video frame were extracted to generate a time sequence (sequence length of 60 frames) as training samples, and the corresponding state labels (normal / distracted) were labeled. Model training: Using time series as input and state labels as supervision signals, supervised training of the LSTM network (128 hidden layer neurons, 50 iterations, batch size 32) is performed until the model achieves an accuracy of ≥90% on the validation set. Real-time inference: The coordinates of facial key points in 60 consecutive frames are used to form a temporal sequence and input into the LSTM model. The output is a cognitive distraction probability value. When the probability value is ≥0.85 (second confidence threshold) and the duration is ≥3s (third time threshold), it is judged as cognitive distraction. Parallel computing logic: The data processing module adopts a multi-threaded parallel computing architecture, with fatigue assessment and three types of distraction identification executed simultaneously, and the total computing latency ≤50ms to ensure real-time performance; when the fatigue score exceeds the preset threshold or any distraction identification is valid, it is determined to be the corresponding abnormal state; S4. When it is determined that the driver is in an abnormal state of fatigue or distraction, a warning command and a vehicle intervention command are generated. The warning command is sent to the vehicle alarm system to trigger an audible and visual warning, and the vehicle intervention command is sent to the vehicle driver assistance system to trigger active control operations on at least one of the braking system, electronic stability system, or hazard warning lights. The specific steps of step S4 are as follows: S41. When the driver is in an abnormal state of fatigue or distraction, generate a first-level warning instruction; S42. If the abnormal state continues for more than the preset waiting time, a second-level vehicle intervention command is generated. S43. Send the warning command to the vehicle alarm system to trigger an audible and visual alarm to alert the driver; S44. Send a vehicle intervention command to the vehicle driver assistance system to trigger at least one of the following operations: The system controls the braking system to slightly reduce speed, enhances lane keeping assist torque through the electronic stability system, and automatically activates hazard warning lights. For example, a tiered response mechanism: When the system first determines that the driver is fatigued or distracted, it immediately generates a first-level warning command with a warning delay of ≤100ms. If the abnormal state lasts for more than 5s (preset waiting time) and no effective response from the driver is detected (such as the head posture returning to normal or the hands leaving the face), a second-level vehicle intervention command is generated.

[0041] Warning Execution: The communication module sends the warning command to the vehicle alarm system via the CAN bus, triggering a multi-dimensional audio-visual warning: the in-vehicle speaker plays a voice prompt ("Please note that you are experiencing fatigue / distraction, please take a break"), with the volume adaptively adjusting to the vehicle speed (volume increases by 20% when the vehicle speed is >100km / h); the instrument panel displays a red warning icon accompanied by vibration (vibration frequency 5Hz, amplitude 0.5mm); a prominent warning pop-up appears on the central control screen, which disappears automatically after 3 seconds to avoid interfering with driving.

[0042] Intervention Execution: The vehicle intervention command is sent to the vehicle's Advanced Driver Assistance System (ADAS), triggering at least one of the following active control actions: Braking system: Controls the electro-hydraulic braking module for slight deceleration, with a deceleration range ≤5km / h and a smooth deceleration process (acceleration change rate ≤0.5m / s²). 2 This helps to avoid rear-end collisions caused by sudden braking.

[0043] Electronic Stability Program (ESP): Enhances lane keeping assist torque (torque increment ≤ 5N) (m) The power steering system adjusts and corrects the vehicle's trajectory to prevent it from deviating from the lane.

[0044] Hazard warning lights: Automatically turn on hazard lights at a flashing frequency of 1Hz to alert surrounding vehicles to take evasive action; once the driver returns to normal, the system automatically turns off the hazard warning lights and disengages.

[0045] In some embodiments, unlike the embodiments described above, an adaptive feature extraction step for occlusion scenes is also included: When the confidence level of feature points in the lower half of the face detected based on the face detection results is consistently below the threshold, it is determined to be a scenario where a mask is being worn. In scenarios where drivers wear masks, the motion amplitude of the driver's brow bone region in consecutive frames is used as an auxiliary fatigue feature; the motion amplitude is obtained by calculating the variance of the Euclidean distance between the coordinates of the key points of the brow bone in the frames. When infrared image analysis identifies a uniform low-temperature area in the eye region that exceeds a set area threshold and has a temperature below a set temperature threshold, it is determined to be a scene where sunglasses are being worn. In scenarios where sunglasses are worn, a continuous sequence of infrared images is input into a pre-trained temporal convolutional network to analyze the thermal imaging change patterns in the eyelid region and output alternative eye opening and closing probabilities. For example, adaptive processing of occlusion scenes: Mask-wearing scenario: When the confidence score of feature points in the lower half of the face (below the tip of the nose) is below 0.6 in 10 consecutive frames, it is determined to be a mask-wearing scenario. At this time, the system adds two key feature points in the brow bone region and calculates the variance of the Euclidean distance between the coordinates of these feature points in 30 consecutive frames. When the variance is <0.5, it is determined to be an abnormal brow bone movement amplitude, which is included in the evaluation as an auxiliary fatigue feature. Wearing sunglasses scenario: Through temperature analysis of infrared images, when the area around the eyes is greater than 200 pixels... 2 When a uniform low-temperature block (temperature < 30℃, converted from infrared image grayscale value) is detected, it is determined to be a scene of wearing sunglasses; 50 consecutive frames of infrared images are input into a pre-trained temporal convolutional network (TCN). This network analyzes the thermal imaging change pattern of the eyelid region and outputs the probability of eyelid opening and closing state (the threshold is set to 0.7, and if the probability is ≥ 0.7, it is determined to be open eyes, otherwise it is closed eyes), replacing the traditional eye feature parameters.

[0046] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A monitoring system for preventing driver fatigue and distraction, characterized in that, It includes an image acquisition module, a data processing module, and a communication module; The image acquisition module includes an infrared camera and a visible light camera; Infrared cameras are used to capture infrared images of the driver's face, while visible light cameras are used to capture visible light images of the driver's face. The data processing module is used to receive infrared and visible light images of the driver's face and is configured to perform the following functions: The received infrared and visible light images of the driver's face are processed in real time, and the driver's state feature vector, including eye state and head posture, is output through image fusion and feature extraction algorithms. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is judged to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. When the driver is determined to be in an abnormal state of fatigue or distraction, a warning command and a vehicle intervention command are generated. A communication module is used to send the warning command to the vehicle alarm system to trigger an audible and visual warning; The vehicle intervention command is sent to the vehicle driving assistance system to trigger active control operations on at least one of the braking system, electronic stability system, or hazard warning lights.

2. The monitoring system for preventing driver fatigue and distraction according to claim 1, characterized in that, The infrared camera and the visible light camera are integrated and packaged inside the vehicle's A-pillar, and the interior panel area of ​​the vehicle's A-pillar corresponding to the optical axis of the infrared camera and the visible light camera has an optical window that allows infrared light and visible light to pass through. The optical window is made of infrared-transmitting acrylic sheet or infrared-transmitting polycarbonate sheet that is integrally formed or embedded with the A-pillar interior panel.

3. The monitoring system for preventing driver fatigue and distraction according to claim 1, characterized in that, The infrared camera uses a wide-angle lens, while the visible light camera has an angle adjustment mechanism; The wide-angle lens, in conjunction with the angle adjustment mechanism, enables the combined field of view of the image acquisition module to cover the driver's face, shoulders, and part of the upper limbs.

4. A monitoring method for preventing driver fatigue and distraction, applied to the system as described in any one of claims 1-3, characterized in that, Includes the following steps: S1. Simultaneously acquire infrared and visible light images of the driver's face using an infrared camera and a visible light camera; S2. Perform real-time fusion processing on the infrared image and the visible light image to extract the driver state feature vector, which includes eye state and head posture. S3. Based on the driver's state feature vector, combined with the steering wheel frequency information obtained from the vehicle bus, the driver is determined to be in an abnormal state of fatigue or distraction by performing parallel calculations through a pre-set fatigue assessment model and a distraction recognition model. S4. When it is determined that the driver is in an abnormal state of fatigue or distraction, generate a warning command and a vehicle intervention command, and send the warning command to the vehicle alarm system to trigger an audible and visual warning, and send the vehicle intervention command to the vehicle driving assistance system to trigger active control operation on at least one of the braking system, electronic stability system or hazard warning lights.

5. The monitoring method for preventing driver fatigue and distraction according to claim 4, characterized in that, The specific steps of step S1 are as follows: S11. Detect ambient light intensity using a light sensor; If the ambient light intensity is higher than the upper limit threshold or backlighting is detected, the dynamic exposure adjustment function of the visible light camera will be activated. S12. Control the infrared camera and visible light camera installed in the A-pillar to simultaneously acquire infrared and visible light images of the driver's face at the same time; S13. The ambient light intensity, the acquired infrared image, and the visible light image are transmitted to the data processing module located inside the A-pillar via a wire harness.

6. The monitoring method for preventing driver fatigue and distraction according to claim 5, characterized in that, The specific steps of step S2 are as follows: S21. Obtain ambient light intensity; If the ambient light intensity is below the lower threshold, the weight of the infrared image is set higher than that of the visible light image; If the ambient light intensity is higher than the upper limit threshold or backlighting is detected, the weight of the visible light image is set higher than that of the infrared image. S22. The received infrared image is enhanced by the CLAHE algorithm. The enhanced infrared image is then registered with the visible light image based on feature points, and then weighted and fused according to the set weights. S23. A lightweight multi-task convolutional neural network is used to process the face fusion image, locate the face bounding box, and after correcting the face pose based on affine transformation, output the coordinates of five key feature points: eyes, nose tip, left and right corners of mouth. S24. Based on the coordinates of key feature points, calculate eye feature parameters and head posture parameters; the eye feature parameters include the percentage of eyelid closure per unit time. blinking frequency and blink duration The head posture parameters include pitch angle. With deflection angle ; Eyelid closure percentage among eye feature parameters The calculation formula is as follows: in, Let N be the duration of the i-th instance of eyelid closure within the statistical period T, and N be the number of closures within the statistical period. blink frequency and blink duration By analyzing the relative position changes of the reflected light spots from the iris and cornea in consecutive frames, a threshold segmentation method was used to extract them. The head pose parameters are as follows: A 3D face model is constructed based on the coordinates of five key feature points: the eyes, the tip of the nose, and the left and right corners of the mouth. The pitch angle θ and yaw angle φ of the head are calculated using the Perspective-n-Point algorithm, with the following formula: in,( , , () represents the coordinates of the nose tip. , , ( ) represents the coordinates of the center points of both eyes.

7. The monitoring method for preventing driver fatigue and distraction according to claim 6, characterized in that, It also includes an adaptive feature extraction step for occluded scenes: When the confidence level of feature points in the lower half of the face detected based on the face detection results is consistently below the threshold, it is determined to be a scenario where a mask is being worn. In scenarios where drivers wear masks, the motion amplitude of the driver's brow bone region in consecutive frames is used as an auxiliary fatigue feature; the motion amplitude is obtained by calculating the variance of the Euclidean distance between the coordinates of the key points of the brow bone in the frames. When infrared image analysis identifies a uniform low-temperature area in the eye region that exceeds a set area threshold and has a temperature below a set temperature threshold, it is determined to be a scene where sunglasses are being worn. In scenarios where sunglasses are worn, a continuous sequence of infrared images is input into a pre-trained temporal convolutional network to analyze the thermal imaging change patterns in the eyelid region and output alternative eye opening and closing probabilities.

8. The monitoring method for preventing driver fatigue and distraction according to claim 6, characterized in that, The specific steps of step S3 are as follows: S31. Based on eye feature parameters and head posture parameters, and combined with steering wheel frequency information obtained in real time from the vehicle bus, calculate the driver fatigue score. The specific steps are as follows: S311. Based on the proportion of eyelid closure Calculate the normalized value of the percentage of eyelid closure ; S312. Based on blink frequency With blink duration Calculate the blinking abnormality coefficient B: Determine if the following conditions are met: blink duration The blink rate is greater than the first time threshold or the blink frequency is less than the first frequency threshold. If so, B=1; If not, B=0; S313. Head-based pitch angle And the duration of hold, to determine the head posture abnormality coefficient H: Determine the absolute value of the head's pitch angle Whether the pitch angle is greater than a preset first angle threshold and whether the holding time of the current pitch angle exceeds a second time threshold; If so, H=1; If not, H=0; S314. Calculate the steering wheel steering frequency anomaly coefficient S based on the steering wheel turning frequency: Determine if the steering wheel frequency is less than the second frequency threshold; If so, S=1; If not, S=0; S32. Calculate the fatigue score using the following formula. : in, , , , Preset weighting coefficients; S33. Perform at least one of the following distraction detections in parallel: The system uses a YOLOv8 model to detect whether a pre-defined distracting target exists in the driver's field of vision, and simultaneously monitors the head yaw angle. When the confidence level of the distracted target detection is greater than or equal to the preset confidence threshold, and the absolute value of the head deflection angle is... When the duration of a state ≥ the preset second angle threshold is greater than the third time threshold, it is determined to be visual distraction; Based on the MediaPipe hand keypoint detection model, the coordinates of hand keypoints are obtained, and the distance between the hand and the face and the proportion of the occlusion area of ​​the hand on the face are calculated. When the distance between the hand and the face is less than the distance threshold and the occlusion area exceeds the preset proportion for a duration greater than the fourth time threshold, it is determined that the hand is centered. Based on the LSTM network model, the facial expression features represented by the sequence of facial key points are analyzed; when the distraction probability value output by the model is greater than or equal to the second confidence threshold and the duration of the state is greater than or equal to the third time threshold, it is determined to be cognitive distraction. S34. If fatigue score If the threshold is exceeded, it is determined to be a fatigue state; if any distraction detection is valid, it is determined to be the corresponding distraction state.

9. The monitoring method for preventing driver fatigue and distraction according to claim 8, characterized in that, The YOLOv8s model in step S33 is obtained through the following steps: SS1. Collect raw images of the driver's cab scene by shooting real vehicles and crawling public network data, and manually annotate the collected raw images by adding bounding boxes and category labels to objects in the images that belong to preset categories, forming labeled image data; SS2. Randomly divide the labeled image data into training and validation subsets, normalize the size of the images in the training and validation subsets, and convert the labels into label format files that meet the input requirements of the YOLOv8s model; SS3. Load the weights of the YOLOv8s model pre-trained on a general dataset as the initial model, freeze the weights of the first N convolutional layers of the backbone network of the initial model, use the training subset as input, use stochastic gradient descent as the optimizer, and use a preset initial learning rate to iteratively train the unfrozen layers in the initial model to obtain an intermediate model. SS4. Use the validation subset to evaluate the intermediate model, calculate the mean accuracy, and when the mean accuracy exceeds the preset qualified threshold, solidify the corresponding intermediate model parameters into the final YOLOv8s model weight file and deploy it to the data processing module. The LSTM network model used for cognitive distraction identification in step S33 is obtained through the following steps: ST1. Collect multiple facial video data of the driver under normal driving and cognitive distraction states, including inattentiveness and unresponsiveness; extract the coordinates of facial key points in each video frame, generate a time sequence as training samples, and label the corresponding state labels. ST2. Using the time sequence as input and the corresponding distraction state as a supervision signal, a supervised training of an LSTM network model is performed until it can predict the distraction probability based on the input sequence. ST3. Deploy the trained LSTM network model in the data processing module, and during real-time monitoring, input the time sequence of facial key point coordinates from multiple consecutive frames into the LSTM model, and output a probability value representing the driver's state of cognitive distraction.

10. The monitoring method for preventing driver fatigue and distraction according to claim 4, characterized in that, The specific steps of step S4 are as follows: S41. When the driver is in an abnormal state of fatigue or distraction, generate a first-level warning instruction; S42. If the abnormal state continues for more than the preset waiting time, a second-level vehicle intervention command is generated. S43. Send the warning command to the vehicle alarm system to trigger an audible and visual alarm to alert the driver; S44. Send a vehicle intervention command to the vehicle driver assistance system to trigger at least one of the following operations: The system controls the braking system to slightly reduce speed, enhances lane keeping assist torque through the electronic stability system, and automatically activates hazard warning lights.