A safety control method and device for a voice interaction system in an intelligent cockpit
By evaluating the dangerous state of the driver and the vehicle and automatically switching the working mode, the problem that the existing voice interaction system cannot be dynamically adjusted is solved, and the safety and reliability of the system are improved.
Patent Information
- Application Number
- CN202510027166.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-08
- Publication Date
- 2025-05-27
- Estimated Expiration
- 2045-01-08
AI Technical Summary
The existing voice interaction system cannot dynamically adjust according to the real-time changing driving environment and driver status, resulting in the complexity of voice commands that reduce the safety and reliability of the system in emergencies.
By obtaining dynamic information of the vehicle and driver, using deep learning technology to evaluate the dangerous state of the driver and vehicle, and automatically switch the working mode of the smart cockpit voice interaction system, including ordinary working mode, safe working mode and hazardous environment mode.
It realizes automatic adjustment of voice interaction mode according to the real-time driving environment and driver status, improves the safety and reliability of the system, and ensures that the driver can obtain the most suitable operating mode in all situations.
Smart Images

Figure CN119418706B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the technical field of intelligent cockpit voice security and intelligent driving assistance, and particularly relates to a safety control method and device for an intelligent cockpit voice interaction system. Background Art
[0002] With the development of vehicle intelligence and the proposal of the intelligent cockpit concept, the voice interaction system has become an important way for drivers to interact with vehicles. However, current traditional voice control systems usually only rely on simple voice recognition and a fixed command set. Such systems cannot be dynamically adjusted according to the real-time changing driving environment and driver state. For example, when the driver is resting in the vehicle, the traditional system cannot timely adjust the voice interaction method. And in the face of emergencies, complex voice commands may not be as direct and reliable as physical buttons, further reducing the safety and reliability of the system. Summary of the Invention
[0003] Aiming at the problems existing in the prior art, the purpose of the embodiments of the present application is to provide a safety control method and device for an intelligent cockpit voice interaction system, aiming to automatically switch the working modes of the intelligent cockpit and the voice system according to the vehicle and driver states, and enhance the safety of the intelligent cockpit voice interaction system.
[0004] According to the first aspect of the embodiments of the present application, a safety control method for an intelligent cockpit voice interaction system is provided, including:
[0005] Obtain vehicle dynamic information and driver dynamic information, where the vehicle dynamic information includes vehicle speed, GPS data, and IMU data, and the driver dynamic information is the driver's facial video;
[0006] Judge whether the driver is in a fatigued or sleeping state according to the driver dynamic information, so as to realize the assessment of the driver's dangerous state;
[0007] Detect the state of the vehicle and whether there are dangerous situations around according to the vehicle dynamic information, so as to realize the assessment of the vehicle's dangerous state;
[0008] Switch the intelligent cockpit voice interaction system to the corresponding working mode according to the results of the driver dangerous state assessment and the vehicle dangerous state assessment, where the working modes include a normal working mode, a safety working mode, and a dangerous environment mode. In the normal working mode, the vehicle accepts all voice commands and executes corresponding operations. In the safety working mode, the vehicle accepts voice commands from registered personnel and executes corresponding operations. In the dangerous environment mode, the vehicle accepts physical button triggers and executes corresponding operations in cooperation with the driver's voice commands.
[0009] Further, based on the driver's dynamic information, it is determined whether the driver is in a fatigued or sleeping state, so as to evaluate whether the driver is in a dangerous state, including:
[0010] Preprocess the images of each frame in the driver's facial video;
[0011] Based on the preprocessed images, use the rPPG algorithm to extract the driver's physiological information;
[0012] Input the preprocessed images into the fatigue state prediction model, and combine the driver's physiological information to determine whether the driver is in a fatigued or sleeping state. If in a fatigued or sleeping state, it is determined that the driver is in a dangerous state.
[0013] Further, based on the vehicle's dynamic information, detect the state of the vehicle and whether there are dangerous situations around, including:
[0014] Based on the vehicle speed, GPS data, and IMU data, determine whether the vehicle is in a driving state or a parked state;
[0015] Detect whether there is a non - administrator - privileged user using the advanced voice commands for controlling the vehicle. If so, it is determined that there are dangerous situations around the vehicle.
[0016] Further, the voice commands are classified into advanced voice commands and low - level voice commands according to the safety sensitivity level. The advanced voice commands include basic driving control, body control, and driving assistance function control; the low - level voice commands include comfort control, entertainment system control, and in - vehicle interaction control.
[0017] Further, the process of determining whether the user who issues the voice command is an administrator - privileged user includes:
[0018] Obtain the multi - channel audio stream collected by the vehicle's internal microphone array;
[0019] Preprocess the multi - channel audio stream, including noise filtering and signal enhancement;
[0020] Input the preprocessed multi - channel audio stream into the pre - trained ECAPA - TDNN model to extract the voiceprint features;
[0021] Match the extracted voiceprint features with the registered voiceprint features stored in the voiceprint database. If the similarity exceeds a predetermined similarity threshold, the match is successful, that is, the user who issues the voice command is registered. Otherwise, the user who issues the voice command is not registered;
[0022] If the user who issues the voice command is registered, determine whether it is an administrator - privileged user according to the corresponding information stored in the voiceprint database.
[0023] Furthermore, according to the results of driver risk state assessment and vehicle risk state assessment, switch the intelligent cockpit voice system to the corresponding working mode, specifically as follows:
[0024] When the vehicle is not in a driving state, the driver is not in a risk state, and there are no dangerous situations around the vehicle, the intelligent cockpit voice interaction system is in the normal working mode;
[0025] If, when the vehicle is in a driving state or a parked state, it is detected that the driver is in a risk state or the driver is not detected in the driver's seat, the intelligent cockpit voice interaction system switches to the safety working mode;
[0026] If there are dangerous situations around the vehicle, the intelligent cockpit voice interaction system switches to the dangerous environment mode.
[0027] According to the second aspect of the embodiments of the present application, a safety control device for an intelligent cockpit voice interaction system is provided, including:
[0028] An acquisition module, configured to acquire vehicle dynamic information and driver dynamic information, where the vehicle dynamic information includes vehicle speed, GPS data, and IMU data, and the driver dynamic information is the driver's facial video;
[0029] A driver risk state assessment module, configured to determine whether the driver is in a fatigued or sleeping state according to the driver dynamic information, so as to implement driver risk state assessment;
[0030] A vehicle risk state assessment module, configured to detect the state of the vehicle and whether there are dangerous situations around according to the vehicle dynamic information, so as to implement vehicle risk state assessment;
[0031] A working mode switching module, configured to switch the intelligent cockpit voice interaction system to the corresponding working mode according to the results of driver risk state assessment and vehicle risk state assessment, where the working mode includes a normal working mode, a safety working mode, and a dangerous environment mode. In the normal working mode, the vehicle accepts all voice commands and executes corresponding operations. In the safety working mode, the vehicle accepts voice commands of registered personnel and executes corresponding operations. In the dangerous environment mode, the vehicle is triggered by physical buttons and executes corresponding operations in cooperation with the driver's voice commands.
[0032] According to the third aspect of the embodiments of the present application, a computer program product is provided, including computer programs / instructions, and when the computer programs / instructions are executed by a processor, the method described in the first aspect is implemented.
[0033] According to the fourth aspect of the embodiments of the present application, an electronic device is provided, including:
[0034] One or more processors;
[0035] A memory for storing one or more programs;
[0036] When the one or more programs are executed by the one or more processors, the one or more processors implement the method as described in the first aspect.
[0037] According to a fifth aspect of the embodiments of the present application, there is provided a computer-readable storage medium having computer instructions stored thereon, and when the instructions are executed by a processor, the steps of the method as described in the first aspect are implemented.
[0038] The present invention includes the following key links:
[0039] 1. Driver and vehicle state assessment: Based on sensor and camera data, analyze and evaluate the current state of the driver and the vehicle, including detecting whether the driver is in a fatigued or sleeping state and whether there are potential risk factors around the vehicle.
[0040] 2. Deep learning-based perception system: This system utilizes the driver and vehicle state assessment results to realize real-time switching of the intelligent cockpit working mode according to the driver and vehicle states. Through this perception technology, while ensuring the convenience of voice control, it can timely respond to various emergencies to protect the safety of the driver.
[0041] 3. Voice control system: In the normal working mode, the people in the vehicle can directly control the vehicle through voice commands; in the safe working mode, voice recognition will be combined with voiceprint recognition and voice command grading; when in the dangerous environment mode, it supports controlling the vehicle using physical buttons in combination with voice commands to ensure the safety of the driver.
[0042] The technical solutions provided by the embodiments of the present application may include the following beneficial effects:
[0043] As can be seen from the above embodiments, the present application uses devices such as in-vehicle sensors, cameras, and microphone arrays to obtain dynamic information of the vehicle and the driver, combines deep learning-based perception technology to real-time evaluate the risk state of the driver, can timely detect that the driver is in a dangerous state or the vehicle is in a dangerous environment, thereby automatically switching the intelligent cockpit system to the most suitable working mode and adjusting the voice control method, can effectively avoid interference from others and potential risk factors, ensure the safety and efficiency of operations, and thus enhance the overall driving experience and driving safety.
[0044] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit the present application. Description of the Drawings
[0045] The accompanying drawings here are incorporated into the specification and constitute a part of this specification, showing embodiments consistent with this application, and are used together with the specification to explain the principles of this application.
[0046] Figure 1 It is a flowchart of a safety control method for an intelligent cockpit voice interaction system shown according to an exemplary embodiment.
[0047] Figure 2 It is a schematic diagram of voice command classification shown according to an exemplary embodiment.
[0048] Figure 3 It is a schematic diagram of a voiceprint recognition process shown according to an exemplary embodiment.
[0049] Figure 4 It is a block diagram of a safety control device for an intelligent cockpit voice interaction system shown according to an exemplary embodiment.
[0050] Figure 5 It is a schematic diagram of an electronic device shown according to an exemplary embodiment. Detailed implementation manners
[0051] Here, the exemplary embodiments will be described in detail, and the examples are shown in the accompanying drawings. When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The implementation manners described in the following exemplary embodiments do not represent all implementation manners consistent with this application.
[0052] The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit this application. The singular forms "a", "the", and "said" used in this application are also intended to include the plural forms unless the context clearly indicates otherwise. It should also be understood that the term "and / or" used herein refers to and includes any or all possible combinations of one or more of the associated listed items.
[0053] It should be understood that although terms such as first, second, and third may be used in this application to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of this application, the first information may also be referred to as the second information, and similarly, the second information may also be referred to as the first information. Depending on the context, the word "if" as used herein may be interpreted as "when" or "while" or "in response to determining".
[0054] As Figure 1 shown, this application provides a safety control method for an intelligent cockpit voice interaction system, including:
[0055] (1)Obtain vehicle dynamic information and driver dynamic information;
[0056] Specifically, the vehicle dynamic information obtained through vehicle sensors includes vehicle speed, GPS data, IMU data, etc., and the driver dynamic information is the driver's facial video obtained through an in-vehicle camera.
[0057] (2)Based on the driver dynamic information, determine whether the driver is in a fatigued or sleeping state, thereby realizing the assessment of the driver's dangerous state;
[0058] Use vision-based driver state detection for the driver's behavior, including whether there is fatigue driving and whether the driver is in a sleeping state. Real-time performance and effectiveness are important indicators for driver behavior detection. Real-time performance requires being able to monitor this dynamic information in real time so as to immediately issue a warning or take corresponding safety measures when the driver is in a dangerous state, avoiding possible traffic accidents. At the same time, effectiveness requires being able to accurately identify the driver's true state and effectively avoid false alarms and missed detections, being able to accurately judge the driver's dangerous state, thereby triggering corresponding warning or intervention measures when needed to ensure driving safety, specifically including: preprocessing the images of each frame in the driver's facial video; based on the preprocessed images, using the rPPG algorithm to extract the driver's physiological information; inputting the preprocessed images into a fatigue state prediction model, and combining the driver's physiological information to determine whether the driver is in a fatigued or sleeping state. If in a fatigued or sleeping state, it is determined that the driver is in a dangerous state.
[0059] In a specific implementation, algorithms such as YOLOv4, YOLOv3, EfficientDet, Faster-RCNN, RetinaNet, R-FCN, etc. based on the CNN-based object detection method can be used for driver fatigue or sleep detection. In this embodiment, the YOLOv5 algorithm, which is more efficient and suitable for single-GPU training, is selected. Before detecting the driver's facial image through the YOLOv5 algorithm, the driver's facial image needs to be normalized. The normalization uses the z-score standardization method, which is based on the mean and standard deviation of the original data. The processed data conforms to the standard normal distribution, and the conversion function is:
[0060]
[0061] In the formula, is the data value of the normalized driver's facial image, x is the original data value of the driver's facial image, μ is the mean of all driver's facial image data, is the standard deviation of all driver's facial image data.
[0062] rPPG is a non-contact physiological signal detection technology that uses an ordinary camera to capture skin color changes to extract physiological signals such as heart rate, respiratory rate, and heart rate variability. The rPPG technology detects minute color changes related to the cardiac cycle by analyzing the changes in the light intensity reflected by the skin. These changes can be extracted through image processing and signal processing techniques for monitoring cardiovascular activities. In this application, the facial video of the driver to be detected is sent into the Yolov5 model and the rPPG model simultaneously. The length of the video input to the model should be at least 2 seconds, and in specific implementations, a segment of 10 - 15 seconds can be taken. The physiological information (such as heart rate, respiratory rate, heart rate variability, etc.) extracted from the video using the rPPG model is used to assist in determining whether the driver is in a sleep or fatigued state.
[0063] In one embodiment, the algorithm fusion process is as follows: 1) Perform image preprocessing operations such as cropping, scaling, and median filtering on the acquired sequence of driver behavior image frames, and normalize the image pixels using the above-mentioned z-score algorithm to obtain an image of 128 pixels × 128 pixels; 2) Input the preprocessed video into the rPPG algorithm to obtain the driver's heart rate, respiratory rate, and heart rate variability; 3) Input the preprocessed image into the Yolov5 model to obtain whether the driver is in a state such as yawning or closing eyes, and combine the results of the rPPG algorithm to determine whether the driver is in a fatigued or sleep state; specifically, if the duration of detecting that the driver is in a yawning state exceeds the first predetermined threshold and the heart rate variability decreases, it is determined that the driver is in a fatigued state; if the duration of monitoring that the driver is in a closed-eye state exceeds the second predetermined threshold and the heart rate and respiratory rate decrease, it is determined that the driver is in a sleep stage; if in a fatigued or sleep stage, it is determined that the driver is in a dangerous state at this time. It should be noted that the first predetermined threshold and the second predetermined threshold are set according to the actual situation, and this application does not limit them.
[0064] In this embodiment, the Yolov5 model needs to be pre-trained in advance, and the CSPDarknet53 feature extractor pre-trained under the COCO dataset is used.
[0065] (3)Detect the state of the vehicle and whether there are dangerous situations around according to the vehicle dynamic information, so as to realize the assessment of the vehicle dangerous state;
[0066] Specifically, the vehicle state and surrounding dangerous behaviors include: whether the vehicle is in a driving state, dangerous behaviors such as non-administrator privilege users attempting to use advanced voice commands, etc.;
[0067] (3.1) Based on the vehicle speed, GPS data, and IMU data, monitor the vehicle's movement trajectory and acceleration changes. When the acceleration, position changes, or vehicle speed is greater than 0, indicating that the vehicle is moving, it is determined that the vehicle is in a driving state.
[0068] (3.2) Detect whether there is a non - administrator - privileged user using the advanced voice commands to control the vehicle. If so, it is determined that there is a dangerous situation around the vehicle;
[0069] As Figure 2 shown, classify the voice commands into advanced voice commands and low - level voice commands according to the safety - sensitivity level; advanced voice commands include basic driving controls (such as starting / stopping the engine, setting cruise control, etc.), body controls (such as controlling the opening and closing of doors and windows, etc.), and driving assistance function controls (such as parking assistance, lane keeping, etc.); low - level voice commands include comfort controls (such as air - conditioning adjustment, non - driver seat adjustment, etc.), entertainment system controls (such as multimedia device control, etc.), and in - vehicle interaction controls (such as controlling the in - vehicle atmosphere lights, using the voice assistant, etc.). Only the personnel who have been pre - registered by voiceprint and set as administrator - privileged users can use the advanced voice commands, while other personnel can only use the low - level voice commands; voiceprint recognition uses a method based on time - delay neural network to perform voiceprint recognition on the voice, and determines whether the user issuing the voice command is an administrator - privileged user. As Figure 3 shown, specifically as follows:
[0070] (3.2.1) Obtain the multi - channel audio stream collected by the microphone array inside the vehicle;
[0071] The microphone array is installed at appropriate positions inside the vehicle (such as near the roof, seat headrest, etc.) to ensure that the voices of the driver and passengers can be clearly captured. The microphone array real - time collects the sound data inside the vehicle and generates a multi - channel audio stream.
[0072] (3.2.2) Pre - process the multi - channel audio stream, including noise filtering and signal enhancement;
[0073] Pre - process the collected multi - channel audio data, including noise filtering and signal enhancement. Specifically, use a band - pass filter to remove the background noise in non - target frequency bands and retain the sound signals in the target frequency band. The parameters of the band - pass filter should be adjusted according to the noise characteristics inside the vehicle and the target sound frequency range; further apply noise suppression algorithms, such as spectral subtraction or adaptive noise suppression, to reduce the background noise interference and improve the signal clarity; adopt the AGC algorithm to dynamically adjust the gain according to the intensity of the real - time audio signal and enhance the signal intensity of the target sound source to ensure that the signal can be clearly recognized at different volume levels; further apply signal enhancement algorithms, such as beamforming technology, to improve the signal - to - noise ratio of the target sound source by focusing the signals collected by multiple microphones.
[0074] (3.2.3) Input the preprocessed multi-channel audio stream into the pre-trained ECAPA-TDNN model to extract voiceprint features;
[0075] Input the preprocessed audio data into the pre-trained ECAPA-TDNN model. The ECAPA-TDNN model extracts voiceprint features from the audio signal through a deep convolutional neural network (CNN) and a time-delay neural network (TDNN). Feature extraction includes capturing the time-frequency characteristics of the audio signal and generating high-dimensional feature vectors.
[0076] (3.2.4) Match the extracted voiceprint features with the registered voiceprint features stored in the voiceprint database. If the similarity exceeds a predetermined similarity threshold, the match is successful, that is, the user who issued the voice command is registered; otherwise, the user who issued the voice command is not registered;
[0077] Match the extracted voiceprint features with the pre-registered voiceprint database, where the voiceprint database can set the permission level of the voiceprint. The driver can be set as an administrator permission in advance to ensure that high-level voice commands can be executed. Calculate the similarity between voiceprint features using methods such as cosine similarity or Euclidean distance. Determine the matching degree of the voiceprint features according to the similarity score. If the similarity exceeds the predetermined similarity threshold, it is determined that the match is successful, that is, the user who issued the voice command is registered; otherwise, the user who issued the voice command is not registered;
[0078] (3.2.5) If the user who issued the voice command is registered, determine whether it is an administrator permission user according to the corresponding information stored in the voiceprint database.
[0079] (4) According to the results of the driver's dangerous state assessment and the vehicle's dangerous state assessment, switch the intelligent cockpit voice interaction system to the corresponding working mode, where the working mode includes a normal working mode, a safe working mode, and a dangerous environment mode. In the normal working mode, the vehicle accepts all voice commands and performs corresponding operations. In the safe working mode, the vehicle accepts voice commands from registered personnel and performs corresponding operations. In the dangerous environment mode, the vehicle is triggered by physical buttons and performs corresponding operations in cooperation with the driver's voice commands;
[0080] Specifically, when the vehicle is not in a driving state, the driver is not in a sleeping state, and the surrounding of the vehicle is safe, the vehicle should be in the normal working mode and be able to receive voice commands normally; once the vehicle is in a driving state or when it is detected that the driver is in a dangerous state during parking or the driver is not detected in the driver's seat during parking, the intelligent cockpit will automatically switch to the safe working mode, where whether the driver is in the driver's seat can be judged by the Yolov5 algorithm or the pressure sensor on the driver's seat; if it is detected that a person without registered administrator privileges attempts to use a voice command to control an instruction of a high security level, the vehicle will be forced into the dangerous environment mode and trigger an all-vehicle external alarm.
[0081] In different intelligent cockpit working modes, the intelligent cockpit voice interaction system switches the corresponding voice control method according to the current working mode:
[0082] Normal working mode, at this time the vehicle can normally receive voice commands to perform corresponding operations.
[0083] Safe working mode, use the method in steps (3.2.1)-(3.2.5) to judge whether the voiceprint of the voice command has been registered. If it has been registered, at the same time judge its privilege level (administrator privilege or normal privilege). After successful authentication, the system will grant the corresponding operation privilege and execute the voice command issued by the user; after failed authentication, the system will refuse to execute the voice command and may issue a warning prompt.
[0084] Dangerous environment mode, at this time the vehicle will block all voice commands to ensure that the vehicle cannot be controlled by other people, improving the safety of the driver under dangerous conditions; only when the driver presses the physical button and passes the voiceprint recognition of the administrator privilege, the vehicle will execute the corresponding operation according to the voice command.
[0085] The voiceprint recognition in the above steps can monitor the audio information collected by the microphone array in real time and perform recognition and analysis. Only the voice characteristics of the pre-registered driver can pass the recognition, and the rest of the unregistered voices cannot pass the voiceprint recognition system, and different personnel control privileges are granted by judging the privilege level. The benefits of this design include: improving safety, preventing external interference, using voiceprint recognition technology, only allowing the intelligent cockpit voice interaction system to execute high-level voice commands issued by registered and privileged administrators, thus avoiding the impact on vehicle safety caused by interference from other people. Enhancing the user experience, the system can dynamically adjust the parameters of the voiceprint recognition algorithm according to changes in environmental noise and the driver's position to ensure accurate recognition and response to voice commands in various driving environments.
[0086] Corresponding to the embodiment of the safety control method of an intelligent cockpit voice interaction system described above, the present application also provides an embodiment of a safety control device of an intelligent cockpit voice interaction system.
[0087] Figure 4 It is a block diagram of a safety control device for an intelligent cockpit voice interaction system shown according to an exemplary embodiment. Referring to Figure 4 , the device may include:
[0088] An acquisition module 21, configured to acquire vehicle dynamic information and driver dynamic information, where the vehicle dynamic information includes vehicle speed, GPS data, and IMU data, and the driver dynamic information is the driver's facial video;
[0089] A driver dangerous state assessment module 22, configured to determine whether the driver is in a fatigued or sleeping state according to the driver dynamic information, so as to implement driver dangerous state assessment;
[0090] A vehicle dangerous state assessment module 23, configured to detect the state of the vehicle and whether there are dangerous situations around according to the vehicle dynamic information, so as to implement vehicle dangerous state assessment;
[0091] A working mode switching module 24, configured to switch the intelligent cockpit voice interaction system to a corresponding working mode according to the results of driver dangerous state assessment and vehicle dangerous state assessment, where the working modes include a normal working mode, a safety working mode, and a dangerous environment mode. In the normal working mode, the vehicle accepts all voice commands and executes corresponding operations. In the safety working mode, the vehicle accepts voice commands of registered personnel and executes corresponding operations. In the dangerous environment mode, the vehicle is triggered by physical buttons and cooperates with the driver's voice commands to execute corresponding operations.
[0092] Regarding the device in the above embodiment, the specific manners in which each module performs operations have been described in detail in the embodiment related to the method, and will not be elaborated here.
[0093] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can be referred to the partial description of the method embodiment. The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separated, and the components shown as units may or may not be physical units, that is, they may be located in one place, or may be distributed to multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the solution of the present application. Those of ordinary skill in the art can understand and implement it without creative efforts.
[0094] Correspondingly, the present application further provides an electronic device, including: one or more processors; a memory for storing one or more programs; when the one or more programs are executed by the one or more processors, the one or more processors implement the safety control method of an intelligent cockpit voice interaction system as described above. As Figure 5 shown, it is a hardware structure diagram of an apparatus for safety control of an intelligent cockpit voice interaction system provided by an embodiment of the present invention in any device with data processing capabilities. In addition to Figure 5 the processors, memory, and network interfaces shown, any device with data processing capabilities where the apparatus is located in the embodiment may usually include other hardware according to the actual functions of the device with data processing capabilities, which will not be elaborated here.
[0095] Correspondingly, the present application further provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the safety control method of an intelligent cockpit voice interaction system as described above is implemented. The computer-readable storage medium may be an internal storage unit of any device with data processing capabilities described in any of the foregoing embodiments, such as a hard disk or memory. The computer-readable storage medium may also be an external storage device, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Further, the computer-readable storage medium may also include both an internal storage unit of any device with data processing capabilities and an external storage device. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capabilities, and may also be used to temporarily store data that has been output or will be output.
[0096] Those skilled in the art will readily think of other implementation manners of the present application after considering the specification and practicing the content disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application.
Claims
1. A safety control method for an intelligent cockpit voice interaction system, characterized in that: include: Acquiring vehicle dynamic information and driver dynamic information, wherein the vehicle dynamic information includes vehicle speed, GPS data, and IMU data, and the driver dynamic information is a driver's facial video; Based on the driver's dynamic information, the driver's physiological information is extracted using the rPPG algorithm to determine whether the driver is in a fatigue or sleep state, thereby achieving a driver's dangerous state assessment; According to the vehicle dynamic information, the vehicle status and whether there are any dangerous situations around it are detected, so as to realize the dangerous status assessment of the vehicle; According to the results of the driver's dangerous state assessment and the vehicle's dangerous state assessment, the intelligent cockpit voice interaction system is switched to the corresponding working mode, wherein the working modes include normal working mode, safe working mode and dangerous environment mode. In the normal working mode, the vehicle accepts all voice commands and performs corresponding operations. In the safe working mode, the vehicle accepts voice commands of registered personnel and performs corresponding operations. In the dangerous environment mode, the vehicle accepts physical button triggers and performs corresponding operations in conjunction with the driver's voice commands. Among them, according to the results of the driver's dangerous state assessment and the vehicle's dangerous state assessment, the intelligent cockpit voice system is switched to the corresponding working mode, specifically: When the vehicle is not in motion, the driver is not in danger, and there is no danger around the vehicle, the intelligent cockpit voice interaction system is in normal working mode; If the driver is detected to be in a dangerous state or not in the driving position when the vehicle is in driving or parking state, the intelligent cockpit voice interaction system switches to a safe working mode, where whether the driver is in the driving position is determined by the Yolov5 algorithm or the pressure sensor on the driving position; If there is a dangerous situation around the vehicle, that is, the ECAPA-TDNN model detects that a person without registered administrator privileges attempts to use voice commands to control high-level safety level instructions, the intelligent cockpit voice interaction system switches to the dangerous environment mode.
2. The method according to claim 1, characterized in that Judging whether the driver is in a fatigue or sleep state based on the driver dynamic information, thereby evaluating whether the driver is in a dangerous state, including: Preprocessing the images of each frame in the driver's facial video; Based on the preprocessed images, the rPPG algorithm is used to extract the driver's physiological information; The preprocessed image is input into the fatigue state prediction model, and combined with the driver's physiological information, it is determined whether the driver is in a fatigue or sleep state. If the driver is in a fatigue or sleep state, it is determined that the driver is in a dangerous state.
3. The method according to claim 1, characterized in that According to the vehicle dynamic information, the vehicle status and whether there is a dangerous situation around the vehicle are detected, including: Based on the vehicle speed, GPS data and IMU data, determining whether the vehicle is in a driving state or a parking state; Detects whether there is an advanced voice command that a non-administrator user is using to control the vehicle. If so, determines that there is a dangerous situation around the vehicle.
4. The method according to claim 3, characterized in that The voice commands are classified into high-level voice commands and low-level voice commands according to the degree of security sensitivity. The high-level voice commands include basic driving control, body control, and driving assistance function control; the low-level voice commands include comfort control, entertainment system control, and in-vehicle interaction control.
5. The method according to claim 3, characterized in that: The process of determining whether the user issuing the voice command is an administrator user includes: Get the multi-channel audio stream collected by the microphone array inside the vehicle; Preprocessing the multi-channel audio stream, including noise filtering and signal enhancement; The preprocessed multi-channel audio stream is input into the pre-trained ECAPA-TDNN model to extract voiceprint features; The extracted voiceprint features are matched with the registered voiceprint features stored in the voiceprint database. If the similarity exceeds a predetermined similarity threshold, the match is successful, that is, the user who issued the voice command is registered, otherwise, the user who issued the voice command is not registered; If the user who issues the voice command has been registered, it is determined whether he is an administrator user based on the corresponding information stored in the voiceprint database.
6. A safety control device for an intelligent cockpit voice interaction system, characterized in that: include: An acquisition module, used to acquire vehicle dynamic information and driver dynamic information, wherein the vehicle dynamic information includes vehicle speed, GPS data, and IMU data, and the driver dynamic information is a driver's facial video; A driver danger state assessment module is used to extract the driver's physiological information based on the driver's dynamic information using the rPPG algorithm to determine whether the driver is in a fatigue or sleep state, thereby achieving driver danger state assessment; A vehicle dangerous state assessment module is used to detect the state of the vehicle and whether there are dangerous situations around it according to the vehicle dynamic information, so as to achieve vehicle dangerous state assessment; A working mode switching module is used to switch the intelligent cockpit voice interaction system to a corresponding working mode according to the results of the driver's dangerous state assessment and the vehicle's dangerous state assessment, wherein the working modes include a normal working mode, a safe working mode and a dangerous environment mode. In the normal working mode, the vehicle accepts all voice commands and performs corresponding operations. In the safe working mode, the vehicle accepts voice commands from registered personnel and performs corresponding operations. In the dangerous environment mode, the vehicle accepts physical button triggers and performs corresponding operations in conjunction with the driver's voice commands. Among them, according to the results of the driver's dangerous state assessment and the vehicle's dangerous state assessment, the intelligent cockpit voice system is switched to the corresponding working mode, specifically: When the vehicle is not in motion, the driver is not in danger, and there is no danger around the vehicle, the intelligent cockpit voice interaction system is in normal working mode; If the driver is detected to be in a dangerous state or not in the driving position when the vehicle is in driving or parking state, the intelligent cockpit voice interaction system switches to a safe working mode, where whether the driver is in the driving position is determined by the Yolov5 algorithm or the pressure sensor on the driving position; If there is a dangerous situation around the vehicle, that is, the ECAPA-TDNN model detects that a person without registered administrator privileges attempts to use voice commands to control high-level safety level instructions, the intelligent cockpit voice interaction system switches to the dangerous environment mode.
7. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instructions are executed by a processor, the method according to any one of claims 1 to 5 is implemented.
8. An electronic device, characterized in that: include: one or more processors; A memory for storing one or more programs; When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1 to 5.
9. A computer-readable storage medium having computer instructions stored thereon, characterized in that: When the instruction is executed by a processor, the steps of the method according to any one of claims 1 to 5 are implemented.
Citation Information
Patent Citations
User monitoring device based on facial expression and voice recognition and monitoring method thereof
CN112738364A
Driver dangerous driving behavior detection method and device, vehicle and storage medium
CN119049017A