Driver state analysis system based on facial expression and voice emotion

By collecting and integrating the driver's facial expressions and voice information, the problem of the existing technology being difficult to fully reflect the driver's status is solved, achieving more accurate status judgment and lower traffic accident risks.

CN120673382AInactive Publication Date: 2025-09-19GUIZHOU YIAN INFORMATION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510773983.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-19
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Most existing driver status monitoring technologies have limitations. A single sensor or indicator method is difficult to fully and accurately reflect the driver's true status, leading to misjudgment and increased risk of traffic accidents.

Method used

By simultaneously collecting the driver's facial expressions and voice information, a multimodal data fusion algorithm is used to perform comprehensive analysis, including data collection, preprocessing, feature extraction, classification and recognition, and state judgment.

Benefits of technology

It can more comprehensively and accurately reflect the driver's real state, improve the accuracy and timeliness of state judgment, and reduce the risk of traffic accidents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673382A_ABST
    Figure CN120673382A_ABST
Patent Text Reader

Abstract

The invention discloses a driver state analysis system based on facial expressions and voice emotions. The driver state analysis system comprises a data acquisition module, a data processing and analysis module and an early warning and feedback module. The data acquisition module is the basis of the whole system, is responsible for acquiring facial expression and voice information of a driver in real time, and comprises a facial expression acquisition sub-module and a voice emotion acquisition sub-module; the facial expression acquisition sub-module comprises facial expression acquisition hardware, the facial expression acquisition hardware adopts a high-definition camera as main equipment for facial expression acquisition, and the voice emotion acquisition sub-module comprises voice emotion acquisition hardware. The voice emotion acquisition hardware adopts a high-sensitivity microphone as voice emotion acquisition equipment; the data processing and analyzing module is the core of the whole system. According to the method, the fatigue degree of the driver can be identified more accurately, and misjudgment caused by limitation of single modal information is avoided.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent transportation and driver safety monitoring, and in particular to a driver status analysis system based on facial expression and voice emotion. Background Art

[0002] In today's society, cars have become an indispensable tool for daily commuting and cargo transportation. However, with the rapid increase in car ownership, road safety issues are becoming increasingly serious. Statistics show that a large number of traffic accidents are caused by poor driver performance, such as fatigue, distracted driving, and impulsive driving caused by emotional fluctuations. These poor driving behaviors significantly increase the risk of traffic accidents and pose a significant threat to people's lives and property.

[0003] However, existing technologies for driver state monitoring often have limitations in practical applications. Some methods based on single sensors or indicators, such as using cameras to monitor the driver's eye movements to determine fatigue, or inferring the driver's distraction based solely on vehicle parameters such as steering wheel angle and speed, often fail to fully and accurately reflect the driver's true state.

[0004] Taking eye movement fatigue monitoring as an example, judging fatigue solely by observing the degree of eye openness or closure through a camera is easily affected by individual differences in drivers (for example, some people are born with small eyes), lighting conditions (eye features are not obvious in strong or low light), and whether the driver wears glasses or sunglasses, resulting in inaccurate monitoring results. Similarly, judging distraction status based solely on vehicle parameters also has errors, because changes in vehicle parameters may be caused by a variety of reasons and are not necessarily related to driver distraction. At the same time, the driver's state is a complex psychological and physiological process, and single-modal information often cannot fully reflect their true state. For example, a driver may appear to be focused (with normal facial expression) on the surface, but actually be angry (with abnormal voice emotion). If only facial expression monitoring is relied upon, this potentially dangerous state will be overlooked. Summary of the Invention

[0005] The purpose of the present invention is to provide a driver status analysis system based on facial expressions and voice emotions to address the limitations of most existing driver status monitoring technologies mentioned in the above background technology. Some methods based on a single sensor or indicator, such as only monitoring the driver's eye movements through a camera to judge the degree of fatigue, or only inferring the driver's distraction status based on vehicle parameters such as steering wheel angle and speed, often cannot fully and accurately reflect the driver's actual status.

[0006] To achieve the above-mentioned purpose, the present invention provides the following technical solutions: including a data acquisition module, a data processing and analysis module, and an early warning and feedback module;

[0007] The data acquisition module is the basis of the entire system and is responsible for collecting the driver's facial expressions and voice information in real time. The module includes a facial expression acquisition submodule and a voice emotion acquisition submodule;

[0008] The facial expression acquisition submodule includes facial expression acquisition hardware, which uses a high-definition camera as the main device for facial expression acquisition; the voice emotion acquisition submodule includes voice emotion acquisition hardware, which uses a high-sensitivity microphone as the device for voice emotion acquisition;

[0009] The data processing and analysis module is the core of the entire system and is responsible for preprocessing, feature extraction, classification, recognition and comprehensive analysis of the collected facial expression and voice information. The module includes a facial expression analysis submodule, a voice emotion analysis submodule and a comprehensive analysis and state judgment submodule.

[0010] The facial expression analysis submodule includes image preprocessing, feature extraction, and expression classification and recognition;

[0011] The speech emotion analysis submodule includes speech preprocessing and emotion classification and recognition;

[0012] The comprehensive analysis and state judgment submodule includes multimodal data fusion and state judgment and quantitative evaluation;

[0013] Early warning and feedback module, including early warning submodule, feedback submodule and remote monitoring and notification;

[0014] Early warning submodule, including early warning mode and early warning threshold setting;

[0015] Feedback submodule, including feedback suggestion content and personalized feedback;

[0016] Remote monitoring and notification, including data transmission and remote notification.

[0017] Preferably, the high-definition camera has high resolution and autofocus function. In addition, the high-definition camera also has a light compensation function. The high-definition camera also integrates infrared technology. The high-definition camera is installed in the center above the windshield in front of the driver in the car, with an angle tilted downward by about 15-20 degrees. The facial image data collected by the high-definition camera is transmitted to the data processing and analysis module in real time via a data cable.

[0018] Preferably, the high-sensitivity microphone has a wide frequency response range, and the high-sensitivity microphone has a noise reduction function. The high-sensitivity microphone also has good directivity. The high-sensitivity microphone is installed in a suitable position near the driver's head in the car. The voice data collected by the high-sensitivity microphone is transmitted to the data processing and analysis module in real time through the audio line.

[0019] Preferably, the image preprocessing includes preprocessing the collected facial images to improve the accuracy of subsequent analysis, the preprocessing steps including image denoising and normalization, the image denoising adopting a median filter or a Gaussian filter algorithm to remove salt and pepper noise and Gaussian noise in the image, and the normalization processing including image size normalization and grayscale normalization, adjusting facial images of different sizes to a uniform size and adjusting the grayscale value of the image to a suitable range of 0-255 to facilitate subsequent feature extraction and classification recognition;

[0020] The feature extraction includes using a deep learning algorithm, a convolutional neural network (CNN), to extract features from the preprocessed facial image. In the present invention, a pre-trained CNN model is used as a feature extractor, and the facial image is input into the CNN model to extract high-dimensional feature vectors. These feature vectors contain rich information about the facial image, such as facial contour, facial features, and texture features.

[0021] The expression classification and recognition includes inputting the extracted facial feature vector into a classifier for expression classification and recognition. The classifier adopts a fully connected neural network FCNN in deep learning. In the present invention, in order to improve the accuracy of classification, a fully connected neural network in deep learning is adopted as a classifier. By training on a pre-established facial expression database, the classifier can accurately identify the emotional state represented by the driver's facial expression, such as happiness, sadness, anger, and fatigue, and quantify the intensity of the emotion.

[0022] Preferably, the speech preprocessing includes preprocessing the collected speech information, including speech segmentation and feature extraction steps. Speech segmentation is to divide the continuous speech signal into multiple short-time speech frames according to the time interval of 10-30 milliseconds to facilitate subsequent feature extraction. Feature extraction is to extract characteristic parameters that can reflect speech emotion from the speech frames.

[0023] The emotion classification and recognition includes inputting the extracted speech features into a classifier for emotion classification and recognition. The classifier can adopt a machine learning algorithm support vector machine (SVM). In the present invention, taking into account the temporal characteristics of the speech signal, a long short-term memory network (LSTM) is adopted as a classifier. By training on a pre-established speech emotion database, the LSTM classifier can accurately identify the emotional state contained in the driver's speech, such as calmness, excitement, anger, and anxiety, and evaluate the intensity of the emotion.

[0024] Preferably, the multimodal data fusion includes fusing the results obtained by the facial expression analysis submodule and the voice emotion analysis submodule, and using a multimodal data fusion algorithm to comprehensively consider the information reflected by the facial expression and voice emotion to judge the overall state of the driver. In the present invention, a late fusion method is adopted, that is, facial expression and voice emotion are first analyzed separately to obtain respective state judgment results, and then these results are fused to establish a weight distribution model, and corresponding weights are assigned to data of different modalities according to the importance of facial expression and voice emotion in judging the driver's state;

[0025] The state judgment and quantitative assessment include judging and quantitatively assessing the driver's state based on the comprehensive analysis results, and setting different state levels, such as normal, mild fatigue, moderate fatigue, severe fatigue, mild distraction, moderate distraction, severe distraction, and anger.

[0026] Preferably, the warning method includes issuing a warning signal in a timely manner when it is determined that the driver is in an abnormal state and reaches a certain warning threshold. The warning methods include but are not limited to sound alarms in the car, seat vibrations, and warning information displayed on the instrument panel. The sound alarm can be set with different alarm tones and frequencies according to different state levels;

[0027] The warning threshold setting includes adjusting the warning threshold according to different driver status levels and actual driving scenarios, comprehensively determining the warning threshold based on the driver's driving time, blinking frequency, and head posture factors. For angry driving status, the warning threshold is set based on the degree of rising tone and volume of the voice and the degree of anger in the facial expression. When these indicators exceed the threshold, the system will promptly issue an angry driving warning.

[0028] Preferably, the feedback suggestion content includes providing corresponding feedback suggestions to the driver at the same time as issuing the warning;

[0029] The personalized feedback includes the system personalizing warning thresholds and feedback suggestions based on the characteristics and driving habits of different drivers, recording the driver's historical status data and feedback effects, and continuously optimizing the personalized feedback plan based on this data;

[0030] Preferably, the data transmission includes the system transmitting the driver's status information to a remote monitoring center or the driver's mobile phone application in real time, and sending the status data obtained by the data processing and analysis module to a remote server through wireless communication technology. The remote server stores and manages the data and provides query and analysis functions.

[0031] Compared with the prior art, the present invention has the following beneficial effects:

[0032] By simultaneously collecting the driver's facial expressions and voice information and using a multimodal data fusion algorithm for comprehensive analysis, the present invention can more comprehensively and accurately reflect the driver's true condition. Compared with monitoring methods based on a single sensor or indicator, the accuracy of state judgment is greatly improved. For example, when judging fatigue status, combining facial expressions (such as eye condition and degree of facial muscle relaxation) and voice emotions (such as speech speed and tone changes) can more accurately identify the driver's fatigue level and avoid misjudgments caused by the limitations of single modal information.

[0033] The present invention also boasts efficient data processing and analysis capabilities, enabling real-time processing of collected facial expressions and voice information, and the issuance of timely warnings and feedback. Generally speaking, the entire process, from data collection to status assessment and warning issuance, can be controlled within a few hundred milliseconds, ensuring that drivers are alerted immediately and effectively avoiding traffic accidents caused by poor condition. BRIEF DESCRIPTION OF THE DRAWINGS

[0034] Figure 1 This is a schematic diagram of the overall process of a driver status analysis system based on facial expression and voice emotion of the present invention;

[0035] Figure 2 This is a flow chart of a data acquisition module of a driver status analysis system based on facial expression and voice emotion according to the present invention;

[0036] Figure 3 This is a schematic diagram of a data processing and analysis module of a driver status analysis system based on facial expression and voice emotion according to the present invention;

[0037] Figure 4 This is a schematic diagram of a warning and feedback module of a driver status analysis system based on facial expressions and voice emotions of the present invention. DETAILED DESCRIPTION

[0038] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.

[0039] See also Figure 1-4 The core purpose of the present invention is to provide a driver status analysis system based on facial expressions and voice emotions, which realizes real-time and accurate analysis of the driver's status by comprehensively utilizing facial expressions and voice emotion information; timely discovers the driver's fatigue, distraction, anger and other adverse conditions, provides timely warnings for the driver, reduces the risk of traffic accidents, improves the accuracy and reliability of driver status monitoring, overcomes the limitations of existing single sensor monitoring methods, provides personalized feedback suggestions for the driver, helps the driver adjust his status, improves driving safety, and can transmit the driver's status information in real time to the remote monitoring center or the mobile devices of relevant personnel to realize remote monitoring and management.

[0040] The driver status analysis system based on facial expressions and voice emotions mainly consists of three parts: data acquisition module, data processing and analysis module, and warning and feedback module. The modules cooperate with each other to complete the analysis and warning tasks of the driver status.

[0041] The data acquisition module is the foundation of the entire system and is responsible for collecting the driver's facial expressions and voice information in real time. This module includes a facial expression acquisition submodule and a voice emotion acquisition submodule;

[0042] The facial expression acquisition submodule includes facial expression acquisition hardware. The facial expression acquisition hardware uses a high-definition camera as the main equipment for facial expression acquisition. The camera has a high resolution (at least 1080P) and can clearly capture subtle changes in the driver's facial expressions. At the same time, the camera has an autofocus function, which can automatically adjust the focal length according to changes in the driver's facial position to ensure that the image is always clear; in addition, the camera also has a light compensation function, which can automatically adjust the exposure parameters under different lighting conditions (such as strong light during the day, weak light at night, and in tunnels) to ensure the stability of the quality of the collected facial images. In order to further improve the acquisition effect at night or in low-light environments, the camera also integrates infrared technology. The infrared camera can capture the thermal radiation information of the driver's face in the absence of visible light, and convert it into a visible facial image through a specific algorithm, thereby ensuring the continuity and stability of facial expression acquisition;

[0043] The installation location of the facial expression acquisition hardware includes the installation of a high-definition camera in the center of the windshield in front of the driver in the car, tilted downward by approximately 15-20 degrees to ensure that the driver's facial image can be fully captured while avoiding interference with the driver's line of sight. The installation height of the camera should be adjusted according to the adjustment range of the car seat, generally about 50-80 cm from the driver's face;

[0044] The data transmission of the facial expression acquisition hardware, including the real-time transmission of the collected facial image data to the data processing and analysis module via a data cable (such as USB3.0 or HDMI), should have good shielding performance to reduce external electromagnetic interference in order to ensure the stability and real-time performance of data transmission;

[0045] The voice emotion acquisition submodule includes voice emotion acquisition hardware. This hardware uses a high-sensitivity microphone as the voice emotion acquisition device. The microphone has a wide frequency response range (20Hz-20kHz) and can accurately capture the various frequency components of the driver's voice. The microphone also has a noise reduction function and uses advanced digital signal processing technology to effectively filter out ambient noise in the car (such as engine noise, wind noise, and passenger conversations), ensuring that the collected voice information is clear and accurate. In addition, the microphone should have good directionality, focusing on collecting voice from the driver's direction and reducing interference from other directions.

[0046] The installation location of the voice and emotion collection hardware, including the high-sensitivity microphone, should be installed in a suitable position near the driver's head in the car, such as on the ceiling near the driver's side, approximately 30-50 cm from the driver's mouth. The installation location should avoid being blocked by other objects, while also taking into account the driver's comfort and safety, and not affecting the driver's normal operation;

[0047] The data transmission of the voice emotion collection hardware includes the real-time transmission of the collected voice data to the data processing and analysis module through an audio cable (such as a 3.5mm audio cable or a digital audio interface). The audio cable should have good anti-interference capabilities to ensure the quality of the voice data.

[0048] The data processing and analysis module is the core of the entire system, responsible for preprocessing, feature extraction, classification and recognition, and comprehensive analysis of the collected facial expression and voice information. This module includes a facial expression analysis submodule, a voice emotion analysis submodule, and a comprehensive analysis and state judgment submodule.

[0049] The facial expression analysis submodule includes image preprocessing, feature extraction, and expression classification and recognition;

[0050] Image preprocessing involves preprocessing the collected facial images to improve the accuracy of subsequent analysis. The preprocessing steps include image denoising and normalization. Image denoising uses a median filter or Gaussian filter algorithm to remove salt and pepper noise and Gaussian noise from the image, making the image clearer. Normalization includes image size normalization and grayscale normalization, which resizes facial images of different sizes to a uniform size (e.g., 224×224 pixels) and adjusts the image's grayscale value to an appropriate range (e.g., 0-255) to facilitate subsequent feature extraction and classification recognition.

[0051] Feature extraction involves extracting features from preprocessed facial images using a deep learning algorithm (e.g., a convolutional neural network (CNN). CNN is a deep learning model specifically designed to process grid-structured data (e.g., images) and automatically learn features from images. In this invention, pretrained CNN models (VGG16 and ResNet) are used as feature extractors. Facial images are input into the CNN model to extract high-dimensional feature vectors, which contain rich information about the facial image, such as facial contours, facial features, and texture features.

[0052] Expression classification and recognition includes inputting the extracted facial feature vectors into a classifier for expression classification and recognition. The classifier may adopt a support vector machine (SVM) or a random forest machine learning algorithm, or a fully connected neural network (FCNN) in deep learning. In the present invention, in order to improve the accuracy of classification, a fully connected neural network in deep learning is adopted as a classifier. By training on a pre-established facial expression database, the classifier can accurately identify the emotional state represented by the driver's facial expression, such as happiness, sadness, anger, and fatigue, and quantify the intensity of the emotion. For example, for a fatigue expression, the degree of fatigue can be judged by analyzing features such as the degree of eye openness, blinking frequency, and the degree of relaxation of facial muscles, and quantified as mild fatigue, moderate fatigue, or severe fatigue.

[0053] Speech emotion analysis submodule, including speech preprocessing and emotion classification and recognition;

[0054] Speech preprocessing involves preprocessing the collected speech information, including speech segmentation, feature extraction, and other steps. Speech segmentation is to divide the continuous speech signal into multiple short-term speech frames according to a certain time interval (such as 10-30 milliseconds) to facilitate subsequent feature extraction. Feature extraction is to extract characteristic parameters that can reflect speech emotions from speech frames. Common speech features include acoustic features such as fundamental frequency (F0), energy, and formant, as well as prosodic features such as speech rate and pause. Fundamental frequency reflects the frequency of vocal cord vibration and is related to the pitch of speech; energy reflects the intensity of speech signals and is related to the loudness of speech; formant is an important acoustic characteristic of the vocal tract and is related to the timbre of speech; speech rate and pause reflect the rhythm and fluency of speech and are closely related to speech emotion.

[0055] Emotion classification and recognition, including inputting the extracted speech features into a classifier for emotion classification and recognition. The classifier can use a machine learning algorithm (such as a support vector machine (SVM)) or a deep learning algorithm (such as a recurrent neural network (RNN) and its variant LSTM). In the present invention, taking into account the temporal characteristics of the speech signal, a long short-term memory network (LSTM) is used as a classifier. LSTM is a special recurrent neural network that can effectively process long sequence data and remember long-term information. By training on a pre-established speech emotion database, the LSTM classifier can accurately identify the emotional state contained in the driver's voice, such as calmness, excitement, anger, and anxiety, and evaluate the intensity of the emotion. For example, for anger, the degree of anger can be judged by analyzing features such as rising voice tone, faster speaking speed, and increased volume, and quantified as mild anger, moderate anger, or severe anger.

[0056] Comprehensive analysis and status judgment submodule, including multimodal data fusion and status judgment and quantitative evaluation;

[0057] Multimodal data fusion includes fusing the results obtained by the facial expression analysis submodule and the voice emotion analysis submodule, and using a multimodal data fusion algorithm to comprehensively consider the information reflected by facial expressions and voice emotions to more accurately judge the driver's overall state. Commonly used multimodal data fusion methods include early fusion, mid-term fusion, and late fusion. In the present invention, a late fusion method is adopted, that is, facial expressions and voice emotions are analyzed separately to obtain their respective state judgment results, and then these results are fused. In specific implementation, a weight distribution model can be established to assign corresponding weights to data of different modalities based on the importance of facial expressions and voice emotions in judging the driver's state. For example, when judging fatigue status, facial expressions (such as eye status) may be more valuable than voice emotions, so the facial expression analysis results can be given a higher weight, while when judging anger status, voice emotions (such as tone and volume) may better reflect the driver's true emotions, so the voice emotion analysis results can be given a higher weight.

[0058] State judgment and quantitative assessment include judging and quantitatively evaluating the driver's state based on the comprehensive analysis results, and setting different state levels, such as normal, mild fatigue, moderate fatigue, severe fatigue, mild distraction, moderate distraction, severe distraction, and anger. For example, when the facial expression analysis results show that the driver blinks frequently and has a dull look, and the voice emotion analysis results show that the driver is depressed and speaks slowly, it is comprehensively judged that the driver may be in a fatigued driving state, and the driver is classified as mild fatigue, moderate fatigue, or severe fatigue according to the degree of fatigue. Similarly, when the facial expression shows anger (such as furrowed brows and downturned corners of the mouth) and the voice emotion analysis shows that the driver is emotionally excited and has a rising tone, it is judged that the driver may be in an angry driving state, and a quantitative assessment is performed based on the intensity of the anger.

[0059] Early warning and feedback module, including early warning submodule, feedback submodule and remote monitoring and notification;

[0060] Early warning submodule, including early warning mode and early warning threshold setting;

[0061] Warning methods include issuing warning signals in a timely manner when it is determined that the driver is in an abnormal state (such as fatigue, distraction, anger, etc.) and reaches a certain warning threshold. Warning methods include but are not limited to in-vehicle sound alarms, seat vibrations, and warning messages displayed on the instrument panel. The sound alarm can be set with different alarm tones and frequencies according to different state levels. For example, for mild fatigue, a relatively gentle prompt tone can be used; for severe fatigue or anger, a rapid and loud alarm tone is used to attract the driver's attention. Seat vibration can be set with different vibration modes and intensities to allow the driver to feel the warning more intuitively, such as using an intermittent vibration mode and adjusting the vibration intensity according to the state level. The warning message displayed on the instrument panel can clearly inform the driver of the current state and the measures to be taken, such as displaying prompts such as "You are in a fatigued driving state, please stop and rest as soon as possible" or "You are emotionally agitated, please remain calm when driving";

[0062] Warning threshold settings, including those adjusted according to different driver status levels and actual driving scenarios, should be implemented. For example, for fatigue driving, the warning threshold can be determined based on factors such as the driver's driving time, blinking frequency, and head posture. Generally speaking, when a driver's continuous driving time exceeds a certain length (e.g., 2-3 hours), and their blinking frequency decreases significantly, and their head frequently droops, the system should issue a fatigue warning. For angry driving, the warning threshold can be set based on factors such as the degree of voice tone and volume increase, and the degree of anger in facial expressions. When these indicators exceed a certain threshold, the system should promptly issue an angry driving warning.

[0063] Feedback submodule, including feedback suggestion content and personalized feedback;

[0064] Feedback suggestions include providing corresponding feedback suggestions to the driver at the same time as issuing a warning. For example, when the driver is judged to be in a state of fatigue, the feedback suggestions may include "please stop and rest in a safe place as soon as possible", "it is recommended to turn on the ventilation function in the car to stay awake", "you can listen to some refreshing music", etc. When the driver is judged to be distracted, the feedback suggestions may be "please concentrate on driving and avoid using distracting devices such as mobile phones", "please adjust the seat and rearview mirror to ensure a comfortable driving posture", "if you feel sleepy, you can drink some coffee or tea", and when the driver is judged to be in an angry state, the feedback suggestions may be "please stay calm, avoid emotional driving, try to take a deep breath to relax", "you can listen to some soothing music to relieve your emotions", "if your emotions are difficult to control, it is recommended to pull over and continue driving after your emotions have calmed down";

[0065] Personalized feedback includes the system's ability to personalize warning thresholds and feedback suggestions based on the characteristics and driving habits of different drivers. For example, for drivers who frequently drive long distances, the warning threshold for fatigue driving can be appropriately raised, while at the same time providing more detailed rest suggestions. For drivers who are easily distracted, monitoring and warning of distracted driving can be strengthened, and more targeted suggestions for avoiding distractions can be provided. In addition, the system can also record the driver's historical status data and feedback effects, and continuously optimize the personalized feedback plan based on this data to improve the effectiveness and pertinence of the feedback.

[0066] Remote monitoring and notification, including data transmission and remote notification;

[0067] Data transmission, including the system can also transmit the driver's status information in real time to a remote monitoring center or the driver's mobile phone application. Through wireless communication technology (such as 4G, 5G or Wi-Fi), the status data (including status level, quantitative evaluation results, etc.) obtained by the data processing and analysis module is sent to the remote server. The remote server stores and manages the data and provides query and analysis functions;

[0068] Remote notification, including remote monitoring center staff can view the driver's status information in real time through monitoring software. When the driver is found to be in an abnormal state, the driver or relevant personnel (such as family members, fleet managers) can be notified in time by phone, text message or application push. At the same time, the remote monitoring center can also provide remote guidance and suggestions based on the driver's status to help the driver adjust his status. For example, if the driver is in a fatigued driving state, the remote monitoring center can suggest the driver to stop and rest at the nearest service area and provide the location information of the service area.

[0069] By simultaneously collecting the driver's facial expressions and voice information and applying a multimodal data fusion algorithm for comprehensive analysis, this method can more comprehensively and accurately reflect the driver's true condition. Compared with monitoring methods based on a single sensor or indicator, this method significantly improves the accuracy of condition judgment. For example, when determining fatigue status, combining facial expressions (such as eye condition and degree of facial muscle relaxation) with voice emotion (such as speech rate and inflection) can more accurately identify the driver's fatigue level, avoiding misjudgments caused by the limitations of single modal information.

[0070] The system boasts efficient data processing and analysis capabilities, enabling real-time processing of collected facial expressions and voice information, and issuing timely warnings and feedback. Generally speaking, the entire process, from data collection to status assessment and warning issuance, is time-delayed within a few hundred milliseconds, ensuring drivers receive prompt alerts and effectively preventing accidents caused by poor performance.

[0071] The system is not significantly affected by environmental factors such as lighting conditions and in-vehicle noise, and can operate stably in a variety of complex driving environments. The facial expression acquisition submodule, through autofocus and light compensation, and with the assistance of infrared technology, ensures facial image quality in diverse lighting conditions, including strong light, low light, and at night. The voice emotion acquisition submodule, through noise reduction and excellent directionality, effectively filters out in-vehicle ambient noise, ensuring clear and accurate voice information.

[0072] The system can personalize warning thresholds and feedback suggestions based on the characteristics and driving habits of different drivers, improving the effectiveness and relevance of warnings and feedback to better meet their needs. For example, different warning thresholds can be set for drivers of different ages, genders, and driving experience; and specific driving suggestions can be provided for drivers who frequently navigate different road conditions.

[0073] At the same time, by transmitting driver status information to a remote monitoring center in real time, remote monitoring and management of the driver's status is achieved. Relevant personnel can promptly understand the driver's status and take appropriate measures, such as reminding the driver to rest and arranging rescue, thus improving the efficiency and level of road traffic safety management.

[0074] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.

[0075] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.

Claims

1. A driver status analysis system based on facial expression and voice emotion, characterized by: Including data acquisition module, data processing and analysis module and early warning and feedback module; The data acquisition module is the basis of the entire system and is responsible for collecting the driver's facial expressions and voice information in real time. The module includes a facial expression acquisition submodule and a voice emotion acquisition submodule; The facial expression acquisition submodule includes facial expression acquisition hardware, which uses a high-definition camera as the main device for facial expression acquisition; the voice emotion acquisition submodule includes voice emotion acquisition hardware, which uses a high-sensitivity microphone as the device for voice emotion acquisition; The data processing and analysis module is the core of the entire system and is responsible for preprocessing, feature extraction, classification, recognition and comprehensive analysis of the collected facial expression and voice information. The module includes a facial expression analysis submodule, a voice emotion analysis submodule and a comprehensive analysis and state judgment submodule. The facial expression analysis submodule includes image preprocessing, feature extraction, and expression classification and recognition; The speech emotion analysis submodule includes speech preprocessing and emotion classification and recognition; The comprehensive analysis and state judgment submodule includes multimodal data fusion and state judgment and quantitative evaluation; Early warning and feedback module, including early warning submodule, feedback submodule and remote monitoring and notification; Early warning submodule, including early warning mode and early warning threshold setting; Feedback submodule, including feedback suggestion content and personalized feedback; Remote monitoring and notification, including data transmission and remote notification.

2. The driver status analysis system based on facial expression and voice emotion according to claim 1, characterized in that: The high-definition camera has high resolution and autofocus function. In addition, the high-definition camera also has a light compensation function. The high-definition camera also integrates infrared technology. The high-definition camera is installed in the center position above the windshield in front of the driver in the car, with an angle downward of about 15-20 degrees. The facial image data collected by the high-definition camera is transmitted to the data processing and analysis module in real time via a data cable.

3. The driver status analysis system based on facial expression and voice emotion according to claim 2, characterized in that: The high-sensitivity microphone has a wide frequency response range, a noise reduction function, and good directivity. The high-sensitivity microphone is installed in a suitable position near the driver's head in the car. The voice data collected by the high-sensitivity microphone is transmitted to the data processing and analysis module in real time through the audio line.

4. The driver status analysis system based on facial expression and voice emotion according to claim 3, characterized in that: The image preprocessing includes preprocessing the collected facial images to improve the accuracy of subsequent analysis. The preprocessing steps include image denoising and normalization. The image denoising uses a median filter or a Gaussian filter algorithm to remove salt and pepper noise and Gaussian noise in the image. The normalization processing includes image size normalization and grayscale normalization. Facial images of different sizes are adjusted to a uniform size and the grayscale value of the image is adjusted to an appropriate range of 0-255 to facilitate subsequent feature extraction and classification recognition. The feature extraction includes using a deep learning algorithm, a convolutional neural network (CNN), to extract features from the preprocessed facial image. In the present invention, a pre-trained CNN model is used as a feature extractor, and the facial image is input into the CNN model to extract high-dimensional feature vectors. These feature vectors contain rich information about the facial image, such as facial contour, facial features, and texture features. The expression classification and recognition includes inputting the extracted facial feature vector into a classifier for expression classification and recognition. The classifier adopts a fully connected neural network FCNN in deep learning. In the present invention, in order to improve the accuracy of classification, a fully connected neural network in deep learning is adopted as a classifier. By training on a pre-established facial expression database, the classifier can accurately identify the emotional state represented by the driver's facial expression, such as happiness, sadness, anger, and fatigue, and quantify the intensity of the emotion.

5. The driver status analysis system based on facial expression and voice emotion according to claim 4, characterized in that: The speech preprocessing includes preprocessing the collected speech information, including speech segmentation and feature extraction steps. Speech segmentation is to divide the continuous speech signal into multiple short-term speech frames according to the time interval of 10-30 milliseconds to facilitate subsequent feature extraction. Feature extraction is to extract characteristic parameters that can reflect the speech emotion from the speech frames; The emotion classification and recognition includes inputting the extracted speech features into a classifier for emotion classification and recognition. The classifier can adopt a machine learning algorithm support vector machine (SVM). In the present invention, taking into account the temporal characteristics of the speech signal, a long short-term memory network (LSTM) is adopted as a classifier. By training on a pre-established speech emotion database, the LSTM classifier can accurately identify the emotional state contained in the driver's speech, such as calmness, excitement, anger, and anxiety, and evaluate the intensity of the emotion.

6. The driver status analysis system based on facial expression and voice emotion according to claim 5, characterized in that: The multimodal data fusion includes fusing the results obtained by the facial expression analysis submodule and the voice emotion analysis submodule, using a multimodal data fusion algorithm to comprehensively consider the information reflected by facial expressions and voice emotions to judge the overall state of the driver. In the present invention, a late fusion method is adopted, that is, facial expressions and voice emotions are first analyzed separately to obtain respective state judgment results, and then these results are fused to establish a weight distribution model. According to the importance of facial expressions and voice emotions in judging the driver's state, corresponding weights are assigned to data of different modes; The state judgment and quantitative assessment include judging and quantitatively assessing the driver's state based on the comprehensive analysis results, and setting different state levels, such as normal, mild fatigue, moderate fatigue, severe fatigue, mild distraction, moderate distraction, severe distraction, and anger.

7. The driver status analysis system based on facial expression and voice emotion according to claim 6, characterized in that: The warning method includes issuing a warning signal in a timely manner when it is determined that the driver is in an abnormal state and reaches a certain warning threshold. The warning methods include but are not limited to sound alarms in the car, seat vibrations, and warning information displayed on the instrument panel. The sound alarm can be set with different alarm tones and frequencies according to different status levels; The warning threshold setting includes adjusting the warning threshold according to different driver status levels and actual driving scenarios, comprehensively determining the warning threshold based on the driver's driving time, blinking frequency, and head posture factors. For angry driving status, the warning threshold is set based on the degree of rising tone and volume of the voice and the degree of anger in the facial expression. When these indicators exceed the threshold, the system will promptly issue an angry driving warning.

8. The driver status analysis system based on facial expression and voice emotion according to claim 7, characterized in that: The feedback suggestion content includes providing corresponding feedback suggestions to the driver at the same time as issuing the warning; The personalized feedback includes the system personalizing warning thresholds and feedback suggestions according to the characteristics and driving habits of different drivers, recording the driver's historical status data and feedback effects, and continuously optimizing the personalized feedback plan based on this data.

9. The driver status analysis system based on facial expression and voice emotion according to claim 8, characterized in that: The data transmission includes the system transmitting the driver's status information in real time to a remote monitoring center or the driver's mobile phone application, and sending the status data obtained by the data processing and analysis module to a remote server through wireless communication technology. The remote server stores and manages the data and provides query and analysis functions.