Driver fatigue wake-up device and method
By utilizing a large language model and a web voice chat module, combined with deep learning models and automatic speech recognition technology, the driver fatigue wake-up device addresses the safety hazards of driver fatigue and effectively wakes up drivers while improving safety.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- FUZHOU UNIV
- Filing Date
- 2024-09-30
- Publication Date
- 2026-06-05
Smart Images

Figure CN122157433A_ABST
Abstract
Description
[0001] This application is a divisional application of a patent application entitled "A Driver Fatigue Awakening Device", the original application was filed on September 30, 2024, application number 202411382995.2. Technical Field
[0002] This application relates to the field of voice interaction technology, and in particular to a driver fatigue wake-up device and method. Background Technology
[0003] Driver fatigue is a major cause of traffic accidents. Current technology only alerts and reports to a backend system for fatigued driving and poor driving habits. However, it lacks effective intervention methods for situations where drivers are fatigued and cannot easily stop to rest, and there are no specific measures to improve driver fatigue. If a driver needs a certain distance and time to reach a rest stop, continuing to drive while fatigued poses a significant safety hazard. Therefore, there is an urgent need for a driver fatigue awakening device and method. Summary of the Invention
[0004] The purpose of this application is to provide a driver fatigue awakening device and method, which can effectively intervene in situations where drivers are fatigued and it is inconvenient to stop and rest, thereby alleviating driver fatigue and improving driving safety.
[0005] To achieve the above objectives, this application provides the following solution.
[0006] This application provides a driver fatigue wake-up device, suitable for various vehicles or embedded in other fatigue monitoring systems, for long-distance transportation and professional driving occasions; the driver fatigue wake-up device includes an in-vehicle terminal and a Web voice chat room module; the in-vehicle terminal includes a command issuing module, an intelligent voice dialogue module, a speech synthesis module and an automatic speech recognition module; The instruction issuing module is used to issue voice interaction instructions when the driver is fatigued. The intelligent voice dialogue module is used to output dialogue text using a large language model after receiving voice interaction commands; the large language model records the driver's historical dialogues and selects questions or topics that are more suitable for the driver's cognitive level as the dialogue text based on the historical dialogues. The speech synthesis module is used to convert dialogue text into output speech using a trained deep learning model. The automatic speech recognition module is used to recognize the driver's response speech when the driver responds to the output speech, obtain the response text, and feed the response text back to the intelligent voice dialogue module; the intelligent voice dialogue module continues to output dialogue text based on the response text; The Web voice chat room module is used to conduct voice calls with friends on the list through the voice chat room to wake up drivers who have fallen into driving fatigue, and to keep drivers in a real-time environment of communication to help prevent them from falling into a state of fatigue driving; if the intelligent voice dialogue module fails to influence the driver, it automatically switches to the Web voice chat room module to wake up the driver through human-to-human interaction.
[0007] Optionally, the speech synthesis module uses a deep learning-based speech synthesis algorithm to play back the dialogue initiated by the large language model to the driver using a driver-customized prompt voice. The speech synthesis module includes the following three processing parts: Text analysis: Performing linguistic analysis on the input dialogue text to determine the sentence structure and the phoneme composition of each word; Speech synthesis: Extracting processed text from a speech synthesis library trained by deep learning, and converting linguistic descriptions into speech waveforms; Prosody processing: By processing the synthesized speech waveform, the output speech is made more coherent and natural, reducing driver discomfort.
[0008] Optionally, the deep learning-based speech synthesis algorithm employs at least one of the WaveNet model, MelGAN model, Tacotron model, FastSpeech model, and Deep Voice model, and supports drivers to customize the timbre of the prompt voice.
[0009] Optionally, the automatic speech recognition module includes a microphone, a feature extraction unit, and a text output unit; The microphone is used to collect the driver's response voice and convert the response voice into a digital signal; The feature extraction unit is used to extract features from the digital signal using the Mel frequency cepstral coefficient algorithm to obtain pronunciation feature data; The text output unit is used to compare the pronunciation feature data with the data in the sound model to determine the corresponding text content of the input and obtain the answer text; The sound model includes an acoustic model, a language model, and a pronunciation model; the acoustic model is obtained by training a deep neural network and is used to map MFCC features to phonemes or speech units; the language model is trained based on a large amount of text data and is used to predict text content based on the probability of word occurrence; the pronunciation model uses a finite state machine or a rule-based system to process the mapping relationship between phonemes and words.
[0010] Optionally, the feature extraction unit performs feature extraction on the digital signal using the Mel frequency cepstral coefficient algorithm, including: Audio capture: Recording sound signals through a microphone and converting them into digital signals; Preprocessing: The digital signal is preprocessed to obtain a preprocessed digital signal; the preprocessing includes removing DC offset and applying a pre-emphasis filter; Framing and windowing: The preprocessed digital signal is divided into multiple frames, and a window function is applied to each frame to reduce boundary effects, resulting in a windowed frame. Fourier Transform: Perform a Fast Fourier Transform on each windowed frame to convert the time-domain signal into a frequency-domain signal; Mel filter bank: A frequency domain signal is filtered through a set of Mel filters to obtain the frequency band of the Mel filter bank; the center frequencies of these filters are evenly distributed on the Mel scale. Logarithmic Transform: Perform a logarithmic transformation on the energy of each frequency band after passing through the Mel filter bank to obtain the logarithmic energy; Discrete Cosine Transform: Perform a discrete cosine transform on the logarithmic energy to obtain the MFCC coefficients as the pronunciation feature data.
[0011] Optionally, when comparing the pronunciation feature data with the data in the sound model, the text output unit matches the pronunciation feature data with the data feature templates in the acoustic model; the pronunciation feature data includes tone and vowel features, and the matching is indirectly completed through deep learning models and statistical models; when comparing similarity, the distance or similarity between the pronunciation feature data and the data feature templates stored in the acoustic model is calculated, and the phoneme or word most likely corresponding to the input speech is determined based on the calculation result; the Euclidean distance algorithm or cosine similarity algorithm is used when calculating the distance or similarity.
[0012] Optionally, the Web voice chat room module comprises five parts: Front-end user interface: Users join voice chat rooms through a web interface or application; Audio capture and transmission: The user's voice is captured through a microphone and converted into digital audio data, which is then sent to a server or directly transmitted to other users via the Internet using a real-time transmission protocol; Server processing: The server is responsible for coordinating voice data streams, managing user connections and permissions, and handling data synchronization needs; Audio reception and playback: The received audio data is decoded and played at the user end; Chat room management features: Users can create chat rooms, invite other users to join, or manage chat room settings; The voice chat room is built using WebSocket and uses WebRTC technology for audio data transmission.
[0013] Optionally, the driver fatigue wake-up device further includes a user terminal; the user terminal is used to enter the voice chat room and set the parameters of the voice chat room, and the user terminal is a mobile phone or tablet computer.
[0014] Optionally, the driver fatigue wake-up device further includes a communication module and a monitoring center; the communication module communicates with the server, monitoring center and mobile terminal through the operator's network; the server has API access function for network information services, and can provide information sources such as traffic information platform, meteorological information platform, AI language model provider or other network consulting platform through API access, so as to provide network information services for the vehicle terminal.
[0015] This application also provides a driver fatigue wake-up method, applied to the aforementioned driver fatigue wake-up device, the driver fatigue wake-up method comprising: When the fatigue monitoring system detects that the driver is fatigued, the command issuing module issues a voice interaction command. After receiving a voice interaction command, the intelligent voice dialogue module outputs dialogue text using a large language model. The large language model records the driver's historical dialogues and selects questions or topics that are more suitable for the driver's cognitive level as the dialogue text based on the historical dialogues. The speech synthesis module converts the dialogue text into output speech and performs prosodic processing on the speech waveform; When the driver responds to the output voice, the automatic speech recognition module recognizes the response voice to obtain the response text, and feeds the response text back to the intelligent voice dialogue module. The intelligent voice dialogue module continues to output dialogue text based on the response text. If the intelligent voice dialogue module fails to influence the driver, it automatically switches to the Web voice chat room module. The driver can then make voice calls with friends on the list through the Web voice chat room module to wake up the driver who has fallen into driving fatigue and to keep the driver in a real-time environment of communication to help prevent them from falling into a state of fatigue driving.
[0016] According to the specific embodiments provided in this application, the following technical effects are disclosed: This application provides a driver fatigue wake-up device and method. On the one hand, it records the driver's historical dialogues through a large language model and selects questions or topics that are more suitable for the driver's cognitive level as dialogue text based on the historical dialogues, making the questions for waking the driver from fatigue more intelligent and more likely to arouse the driver's interest. On the other hand, when the intelligent voice dialogue module fails to influence the driver, it automatically switches to the Web voice chat room module to wake the driver up through human-to-human interaction. This effectively intervenes in situations where it is inconvenient to stop and rest when the driver is fatigued, alleviating the driver's fatigue and improving driving safety. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a schematic diagram of a driver fatigue wake-up device provided in an embodiment of this application; Figure 2 This is a schematic diagram of the functional modules of a driver fatigue wake-up device provided in an embodiment of this application. Detailed Implementation
[0019] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0020] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0021] In one exemplary embodiment, such as Figure 1 As shown, this application provides a driver fatigue wake-up device, including an in-vehicle terminal. The in-vehicle terminal includes a command issuing module, an intelligent voice dialogue module, a speech synthesis module, and an automatic speech recognition module. The driver fatigue wake-up device provided in this application is suitable for various vehicles or can be embedded in other fatigue monitoring systems, and is particularly suitable for long-distance transportation and professional driving applications.
[0022] The command issuing module is used to issue voice interaction commands when the driver is fatigued. Fatigue can be identified through the driver's facial and hand features.
[0023] The intelligent voice dialogue module is used to output dialogue text using a large language model after receiving voice interaction commands; the dialogue text is text that conforms to the driver's style preferences.
[0024] When the fatigue monitoring system detects driver fatigue, it sends a signal, i.e., issuing a voice interaction command through the command issuing module. The intelligent voice dialogue module, based on the driver's style preferences, uses a large language model to generate dialogue text. This dialogue text consists of common-sense questions, which the driver answers, freeing them from monotonous driving. The intelligent voice dialogue module records the driver's historical conversations and selects questions or topics more relevant to the driver's cognitive level as dialogue text. This helps the driver maintain alertness and mental activity over the long term, effectively preventing driver fatigue and improving transportation efficiency.
[0025] Figure 2 Dialogue management in this context refers to maintaining and managing the dialogue process with users, determining the next steps and judgments made by the artificial intelligence. When a driver is in a state of deep fatigue, the large language model integrated into the system needs to ask the driver targeted questions based on the actual situation, helping the driver improve brain activity and recover from fatigue as quickly as possible. Simultaneously, the large language model records the driver's style preferences and asks questions that the driver is more interested in, helping the driver to develop a more effective thought process. It should be noted that the large language model uses a mature open API from a reputable large language model provider to ensure that the driver can engage in dialogue with the artificial intelligence.
[0026] The speech synthesis module is used to convert dialogue text into output speech using a trained deep learning model. The trained deep learning model is a model trained with sample dialogue text as input and sample output speech corresponding to the sample dialogue text as labels.
[0027] The speech synthesis module includes a text analysis unit, a speech synthesis unit, and a prosodic processing unit. The text analysis unit analyzes the dialogue text to obtain analysis results, including the structure of the dialogue text and the phoneme composition of each character. The speech synthesis unit converts the dialogue text into a speech waveform based on the analysis results. The prosodic processing unit performs prosodic processing on the speech waveform to obtain the output speech.
[0028] The speech synthesis module uses a deep learning-based speech synthesis algorithm to play back the dialogue initiated by the language model to the driver using a driver-customized prompt voice. The deep learning-based speech synthesis algorithm can employ models such as WaveNet, MelGAN, Tacotron, FastSpeech, and Deep Voice.
[0029] The speech synthesis module mainly includes the following three processing components: Text analysis: Performing linguistic analysis on the input text to determine the sentence structure and the phoneme composition of each word; Speech synthesis: Extracting processed text from a speech synthesis library trained by deep learning, and converting linguistic descriptions into speech waveforms; Prosody processing: By processing the synthesized speech waveform, the output speech is made more coherent and natural, reducing driver discomfort.
[0030] The speech synthesis module includes a speaker for playing the output speech. The speech synthesis module uses deep learning to extract the dialogue text from a trained speech synthesis library and synthesizes speech based on user-defined voice timbres, thereby broadcasting the output of the AI language model.
[0031] The automatic speech recognition module is used to recognize the driver's response speech when the driver responds to the output speech, obtain the response text, and feed the response text back to the intelligent voice dialogue module. The intelligent voice dialogue module uses a large language model to output dialogue text based on the response text.
[0032] Specifically, when the system detects that the driver is fatigued and the artificial intelligence intervenes (i.e., outputs the corresponding voice text to the driver through the intelligent voice dialogue module), it collects the driver's voice information, converts the sound signal into a digital signal, compares it with the data in the system's voice library after feature recognition, realizes the automatic recognition of the driver's response voice, converts it into text and feeds it back to the artificial intelligence, and then proceeds to the next step after the artificial intelligence makes a judgment.
[0033] The automatic speech recognition module includes a microphone used for recording and processing speech signals. The microphone captures the driver's responses and converts them into digital signals for subsequent computer processing.
[0034] The automatic speech recognition module includes a feature extraction unit and a text output unit. The feature extraction unit extracts features from the digital signal to obtain pronunciation feature data. The text output unit generates the response text based on the pronunciation feature data.
[0035] In terms of extracting features from digital signals to obtain pronunciation feature data, the feature extraction unit utilizes the Mel Frequency Cepstrum Coefficient (MFCC) algorithm to extract features from the digital signals and obtain pronunciation feature data, thereby facilitating accurate recognition of speech content. The pronunciation feature data includes features such as initials and finals.
[0036] The input to the MFCC algorithm is the digital signal obtained after the audio signal has been captured and converted by a microphone. Throughout the process, the original analog audio signal is converted into a discrete-time signal that can be processed by digital signal processing algorithms. The calculation process of the MFCC algorithm is as follows (1) to (7).
[0037] 1) Audio Acquisition: Recording sound signals through a microphone and converting them into digital signals. This typically involves sampling and quantization, converting continuous analog sound signals into discrete digital forms.
[0038] 2) Preprocessing: The digital signal is preprocessed to obtain a preprocessed digital signal. Preprocessing includes removing DC offset and applying a pre-emphasis filter to increase the energy of high-frequency components.
[0039] 3) Framing and Windowing: The preprocessed digital signal is divided into multiple frames, each typically 20-40 milliseconds long. A window function (such as a Hamming window) is then applied to each frame to reduce boundary effects, resulting in a windowed frame.
[0040] 4) Fourier Transform: Perform a Fast Fourier Transform (FFT) on each windowed frame to convert the time-domain signal of the windowed frame into a frequency-domain signal.
[0041] 5) Mel filter bank: The frequency domain signal is filtered through a set of Mel filters to simulate human auditory perception, resulting in the frequency band of the Mel filter bank. The center frequencies of these filters are uniformly distributed on the Mel scale.
[0042] 6) Logarithmic Transformation: Perform a logarithmic transformation on the energy of each frequency band after passing through the Mel filter bank to obtain the logarithmic energy.
[0043] 7) Discrete Cosine Transform (DCT): Performing a discrete cosine transform on the logarithmic energy yields MFCC coefficients, which are the speech feature data. Since the MFCC coefficients contain most of the speech information, usually only the first few coefficients are retained.
[0044] The text output unit is used for speech model comparison. Specifically, it compares and matches the extracted pronunciation feature data with the large sound model in the system. By comparing the similarity between the pronunciation feature data and the data in the sound model, the corresponding text content of the input is determined, and the answer text is obtained.
[0045] The sound model is a model or data structure used for speech recognition, which can be used to store and analyze speech features for matching with input speech features. The sound model can be an acoustic model, a language model, or a pronunciation model, etc.
[0046] The acoustic model is used to map sound signals (such as MFCC features) to possible phonemes or speech units. The acoustic model can be obtained by training a deep neural network (such as a convolutional neural network (CNN), a recurrent neural network (RNN), or a Transformer).
[0047] Language models are used to predict text content based on the probability of word occurrences. They are trained on large amounts of text data to improve the accuracy of speech recognition.
[0048] Pronunciation models are used to process the mapping relationship between phonemes and words. Pronunciation models can be finite state machines or rule-based systems.
[0049] The process of comparing the similarity between articulation feature data and data in the sound model is as follows. Articulation feature data (such as MFCC coefficients) is used for similarity comparison. This feature data includes tones and vowels. During feature extraction, the energy features of high-frequency components may contain tonal information. Tone features are identified by analyzing the frequency distribution and modulation patterns in the feature data. Vowel features are determined by analyzing the frequency distribution patterns and formants in the MFCC coefficients.
[0050] When comparing pronunciation feature data with data in the sound model, the pronunciation feature data extracted using the MFCC algorithm is matched with the data feature templates in the acoustic model. These templates are trained based on a large amount of speech data and are used to describe the MFCC feature distributions corresponding to various phonemes (including tones and vowels). The comparison process mainly involves matching MFCC features with feature templates in the sound model, rather than directly comparing them with specific "tones" or "vowels," but rather indirectly through deep learning models and statistical models.
[0051] Specifically, when comparing similarity, the distance or similarity between the pronunciation feature data and the data feature templates stored in the acoustic model is calculated, and the most likely phoneme or word corresponding to the input speech is determined based on the calculated result. In an exemplary embodiment of this application, the Euclidean distance algorithm can be used to calculate the distance between the pronunciation feature data and the data feature templates stored in the acoustic model, or the cosine similarity algorithm can be used to calculate the similarity between the pronunciation feature data and the data feature templates stored in the acoustic model.
[0052] Furthermore, the driver fatigue wake-up device also includes a Web voice chat room module. This module is used for voice calls with a list of friends via a voice chat room. The voice chat room is built using WebSocket. When driver fatigue is detected, the device proactively intervenes and prevents existing or potential driver fatigue through intelligent human-computer dialogue or by providing the driver with a customized real-time voice chat room, thereby increasing the driver's mental activity and preventing driver fatigue that could lead to traffic accidents.
[0053] like Figure 2 As shown, in addition to waking up drivers who are already experiencing driving fatigue, this application also establishes a Web voice chat room module, which helps prevent drivers from falling into a state of fatigued driving by enabling them to be in a real-time environment of communicating with people.
[0054] The Web voice chat room module mainly consists of the following five parts (1) to (5).
[0055] (1) Front-end user interface: Users join the voice chat room through the web interface or application. This interface includes voice call controls, user list, chat message window, etc. Users can choose to turn on their microphone, speaker, and other devices to make voice calls.
[0056] (2) Audio capture and transmission: The user's voice is captured through a microphone and converted into digital audio data. The audio data is then transmitted to a server via the Internet using a real-time transmission protocol (such as WebRTC) or directly to other users (peer-to-peer connection). This application uses WebRTC technology to build a voice chat room.
[0057] (3) Server processing: In a centralized system, the server can be responsible for coordinating voice data streams, managing user connections and permissions, and handling any data synchronization needs, such as voice activity indications, online status updates, etc.
[0058] For driver fatigue wake-up devices using peer-to-peer technology, the server's main role is to help users establish a connection (i.e., exchange signals) during the initial connection. Subsequent audio data is transmitted directly between users, reducing latency and bandwidth usage.
[0059] The driver fatigue wake-up device also includes a user terminal. The user terminal is used to access and configure the voice chat room's parameters. Users (drivers, relatives, or fleet managers) can configure the voice chat room at any time through the user terminal, which also provides a convenient way for users to join the voice chat room. The user terminal can be a mobile phone or tablet computer.
[0060] The driver fatigue wake-up device also includes a monitoring center. The monitoring center is a monitoring platform set up by the fleet or related industry management department. It is used to monitor and record driver and vehicle information in real time, and managers can direct and manage drivers through the monitoring center.
[0061] The driver fatigue wake-up device also includes a communication module, which can communicate with the server, monitoring center, and mobile terminal through the operator's network.
[0062] The server provides data connectivity and processing services for vehicle terminals, mobile devices, and monitoring centers. It features API access (Application Programming Interface) functionality for network information services, information exchange services between terminals, and chat functionality. Specifically, the API access allows information from traffic information platforms, weather information platforms, AI language model providers, or other online information platforms to be provided as information sources, thereby enabling network information services for vehicle terminals.
[0063] (4) Audio reception and playback: The received audio data is decoded and played at the user end. Echo cancellation and noise suppression methods are used to improve call quality.
[0064] (5) Chat room management function: Users can create chat rooms, invite other users to join, or manage chat room settings, such as access control, voice activity indicators, etc.
[0065] The main function of the web voice chat module is to keep drivers active and prevent them from falling into a state of deep fatigue when they need to drive long distances for extended periods. In cases where AI wake-up measures (through the intelligent voice dialogue module) fail to affect the driver, others can effectively supervise or even wake the driver in an emergency, thus significantly reducing the likelihood of traffic accidents caused by deep fatigue.
[0066] This application integrates a large language model, enabling more intelligent question selection for driver fatigue recovery, thus stimulating driver interest and alleviating fatigue. Simultaneously, this application provides a reminder sound recording function, employing a deep learning-based speech fusion algorithm that allows drivers to customize prompts, further aiding in overcoming fatigue. Furthermore, this application utilizes WebSocket to build a web chat room, enabling drivers to communicate in real-time with family, friends, or fleet managers via voice, preventing driver fatigue and ensuring driving safety.
[0067] This application discloses a system that utilizes AI language models and multimodal active intervention to prevent driver fatigue. By comprehensively applying computer technology, human-computer interaction, human-human interaction, and physical stimulation measures, it can effectively prevent driver fatigue behavior, intervene in deep fatigue states in a timely manner, and significantly improve the safety and efficiency of transportation.
[0068] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0069] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A driver fatigue wake-up device, characterized by, Suitable for various vehicles or embedded in other fatigue monitoring systems, for long-distance transportation and professional driving applications; the driver fatigue wake-up device includes an in-vehicle terminal and a Web voice chat room module; the in-vehicle terminal includes a command issuing module, an intelligent voice dialogue module, a voice synthesis module, and an automatic voice recognition module; The instruction issuing module is used to issue voice interaction instructions when the driver is fatigued. The intelligent voice dialogue module is used to output dialogue text using a large language model after receiving voice interaction commands. The large language model records the driver's historical dialogues and selects questions or topics that are more relevant to the driver's cognitive level as the dialogue text based on the historical dialogues. The speech synthesis module is used to convert dialogue text into output speech using a trained deep learning model. The automatic speech recognition module is used to recognize the driver's response speech when the driver responds to the output speech, obtain the response text, and feed the response text back to the intelligent voice dialogue module. The intelligent voice dialogue module continues to output dialogue text based on the response text; The Web voice chat room module is used to make voice calls with friends on the list through the voice chat room, so as to wake up drivers who have fallen into driving fatigue and keep drivers in a real-time environment of communication to help prevent them from falling into a state of fatigue driving. If the intelligent voice dialogue module fails to influence the driver, it will automatically switch to the Web voice chat room module for human-to-human interaction to wake up the driver.
2. The driver fatigue alerting device of claim 1, wherein The speech synthesis module uses a deep learning-based speech synthesis algorithm to play back the dialogue initiated by the large language model to the driver using a driver-customized prompt voice. The speech synthesis module includes the following three processing parts: Text analysis: Performing linguistic analysis on the input dialogue text to determine the sentence structure and the phoneme composition of each word; Speech synthesis: Extracting processed text from a speech synthesis library trained by deep learning, and converting linguistic descriptions into speech waveforms; Prosody processing: By processing the synthesized speech waveform, the output speech is made more coherent and natural, reducing driver discomfort.
3. The driver fatigue alerting device of claim 2, wherein, The deep learning-based speech synthesis algorithm employs at least one of the WaveNet, MelGAN, Tacotron, FastSpeech, and Deep Voice models, and supports drivers in customizing the timbre of their prompts.
4. The driver fatigue alerting device of claim 1, wherein, The automatic speech recognition module includes a microphone, a feature extraction unit, and a text output unit; The microphone is used to collect the driver's response voice and convert the response voice into a digital signal; The feature extraction unit is used to extract features from the digital signal using the Mel frequency cepstral coefficient algorithm to obtain pronunciation feature data; The text output unit is used to compare the pronunciation feature data with the data in the sound model to determine the corresponding text content of the input and obtain the answer text; The sound model includes an acoustic model, a language model, and a pronunciation model; the acoustic model is obtained by training a deep neural network and is used to map MFCC features to phonemes or speech units; the language model is trained based on a large amount of text data and is used to predict text content based on the probability of word occurrence; the pronunciation model uses a finite state machine or a rule-based system to process the mapping relationship between phonemes and words.
5. The driver fatigue wake-up device according to claim 4, characterized in that, The process by which the feature extraction unit extracts features from the digital signal using the Mel frequency cepstral coefficient algorithm includes: Audio capture: Recording sound signals through a microphone and converting them into digital signals; Preprocessing: The digital signal is preprocessed to obtain a preprocessed digital signal; the preprocessing includes removing DC offset and applying a pre-emphasis filter; Framing and windowing: The preprocessed digital signal is divided into multiple frames, and a window function is applied to each frame to reduce boundary effects, resulting in a windowed frame. Fourier Transform: Perform a Fast Fourier Transform on each windowed frame to convert the time-domain signal into a frequency-domain signal; Mel filter bank: A frequency domain signal is filtered through a set of Mel filters to obtain the frequency band of the Mel filter bank; the center frequencies of these filters are evenly distributed on the Mel scale. Logarithmic Transform: Perform a logarithmic transformation on the energy of each frequency band after passing through the Mel filter bank to obtain the logarithmic energy; Discrete Cosine Transform: Perform a discrete cosine transform on the logarithmic energy to obtain the MFCC coefficients as the pronunciation feature data.
6. The driver fatigue wake-up device according to claim 4, characterized in that, When comparing the pronunciation feature data with the data in the sound model, the text output unit matches the pronunciation feature data with the data feature templates in the acoustic model. The pronunciation feature data includes tone and vowel features, and the matching is indirectly completed through deep learning models and statistical models. When comparing similarity, the distance or similarity between the pronunciation feature data and the data feature templates stored in the acoustic model is calculated. Based on the calculation results, the phoneme or word most likely corresponding to the input speech is determined. The Euclidean distance algorithm or cosine similarity algorithm is used when calculating the distance or similarity.
7. The driver fatigue wake-up device according to claim 1, characterized in that, The Web voice chat room module consists of five parts: Front-end user interface: Users join voice chat rooms through a web interface or application; Audio capture and transmission: The user's voice is captured through a microphone and converted into digital audio data, which is then sent to a server or directly transmitted to other users via the Internet using a real-time transmission protocol; Server processing: The server is responsible for coordinating voice data streams, managing user connections and permissions, and handling data synchronization needs; Audio reception and playback: The received audio data is decoded and played at the user end; Chat room management features: Users can create chat rooms, invite other users to join, or manage chat room settings; The voice chat room is built using WebSocket and uses WebRTC technology for audio data transmission.
8. The driver fatigue wake-up device according to claim 7, characterized in that, It also includes a user terminal; the user terminal is used to enter the voice chat room and set the parameters of the voice chat room, and the user terminal is a mobile phone or tablet computer.
9. The driver fatigue wake-up device according to claim 8, characterized in that, It also includes a communication module and a monitoring center; the communication module communicates with the server, monitoring center and mobile terminal through the operator's network; the server has API access function for network information services, and can provide information sources such as traffic information platform, meteorological information platform, AI language model provider or other network consulting platform through API access, so as to provide network information services for the vehicle terminal.
10. A method for waking up a driver from driver fatigue, characterized in that, The driver fatigue wake-up device according to any one of claims 1-9, wherein the driver fatigue wake-up method comprises: When the fatigue monitoring system detects that the driver is fatigued, the command issuing module issues a voice interaction command. After receiving a voice interaction command, the intelligent voice dialogue module outputs dialogue text using a large language model. The large language model records the driver's historical dialogues and selects questions or topics that are more suitable for the driver's cognitive level as the dialogue text based on the historical dialogues. The speech synthesis module converts the dialogue text into output speech and performs prosodic processing on the speech waveform; When the driver responds to the output voice, the automatic speech recognition module recognizes the response voice to obtain the response text, and feeds the response text back to the intelligent voice dialogue module. The intelligent voice dialogue module continues to output dialogue text based on the response text. If the intelligent voice dialogue module fails to influence the driver, it automatically switches to the Web voice chat room module. The driver can then make voice calls with friends on the list through the Web voice chat room module to wake up the driver who has fallen into driving fatigue and to keep the driver in a real-time environment of communication to help prevent them from falling into a state of fatigue driving.