Intelligent cabin interaction control method and system and vehicle
Through multimodal fusion decision-making technology, accurate vehicle control instructions are generated using EEG, eye movement, voice and gesture data, which solves the problem of vehicle control instruction conflicts in single-modal interaction technology and improves the accuracy of driver intention judgment and the vehicle's execution capability.
Patent Information
- Application Number
- CN202510764464.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-10
- Publication Date
- 2025-09-19
AI Technical Summary
Existing smart cockpit interaction technologies are mostly single-modal, making it difficult to accurately judge and quickly respond to the driver's intentions in complex driving scenarios, resulting in conflicts in vehicle control commands and the inability to execute them.
By acquiring the driver's EEG signals, eye images, voice and gesture image data, pre-processing and multimodal fusion decision-making are performed to generate vehicle control instructions. By using deep learning methods and multimodal fusion technology, the driver's intentions are comprehensively judged to generate accurate vehicle control instructions.
It achieves accurate judgment of the driver's intentions in complex driving scenarios, prevents conflicts in vehicle control commands, ensures that the vehicle can accurately execute the driver's intentions, and improves driving experience and safety.
Smart Images

Figure CN120669856A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to an intelligent cockpit of an automobile, and in particular to an intelligent cockpit interactive control method, system and vehicle. Background Art
[0002] With the continuous advancement of automotive technology, people's expectations for the interactive experience in car cockpits are increasing. Smart cockpit interaction technologies, including eye tracking, voice recognition, and gesture recognition, can perceive the driver's intentions and status from different perspectives, thereby improving the accuracy and convenience of interaction. For example, eye tracking can determine the driver's gaze direction and focus. When the driver looks at the navigation screen, the system can automatically pop up relevant operation options. Or, when the driver frequently checks the rearview mirror, the system can detect that there may be special circumstances behind and provide corresponding prompts. Voice recognition technology has already found some application in car cockpits, converting the driver's voice commands into text commands. Gesture recognition technology allows drivers or passengers to control in-cabin devices through specific gestures, such as adjusting the air conditioning temperature and switching music. In situations where voice control is inconvenient for the driver, gesture recognition provides a natural and intuitive way to interact. However, existing technical solutions focus on a single modality and have yet to form a multimodal collaborative interaction system. Eye tracking, voice recognition, gesture recognition, and other systems operate independently, which may conflict with each other. This makes it difficult to accurately judge the driver's intentions and respond quickly in complex driving scenarios. Summary of the Invention
[0003] In view of this, the purpose of the present invention is to provide an intelligent cockpit interactive control method, system and vehicle, which can obtain more accurate vehicle control instructions that meet the driver's expectations and prevent the vehicle from being unable to execute due to conflicts between vehicle control instructions output by different modules.
[0004] A smart cockpit interactive control method in the present invention includes: Acquiring physiological and behavioral data of the driver, wherein the physiological and behavioral data include: electroencephalogram signal data, eye image data, sound data, and gesture image data; Preprocessing the physiological and behavioral data to obtain the driver's intention tendency; Based on the driver's intention tendency, a multimodal fusion decision is adopted to generate vehicle control instructions.
[0005] Furthermore, the pre-processing of the physiological and behavioral data to obtain the driver's intention tendency includes: The EEG signal data is preprocessed, the EEG signal is amplified and filtered, and EEG feature signals related to the driver's state and intention are extracted. The driver's intention tendency is identified through deep learning methods such as convolutional neural network, recurrent neural network and end-to-end model to obtain a first intention tendency.
[0006] Furthermore, the pre-processing of the physiological and behavioral data to obtain the driver's intention tendency also includes: Preprocessing the eye image data to identify the pupil position, sight direction, and gaze time of the eyeball to obtain a second intention tendency; Preprocessing the sound data, performing noise reduction, feature extraction, speech recognition, and text conversion on the sound data to obtain a third intention tendency; The gesture image data is preprocessed to identify the gesture type and gesture parameters to obtain a fourth intention tendency.
[0007] Furthermore, the generating of vehicle control instructions based on the driver's intention tendency by adopting multimodal fusion decision-making includes: Based on a preset fusion rule and a weight coefficient, the first intention tendency, the second intention tendency, the third intention tendency, and the fourth intention tendency are fused to obtain multimodal fusion signal information; Decisions are made based on the multimodal fusion signal information, vehicle control instructions are generated, and control operation results are fed back.
[0008] Furthermore, the method further includes: executing corresponding vehicle control operations based on the vehicle control instructions.
[0009] An intelligent cockpit interactive control system in the present invention includes: an intention tendency acquisition unit, configured to acquire physiological and behavioral data of the driver and pre-process the physiological and behavioral data to obtain the driver's intention tendency; wherein the physiological and behavioral data include: electroencephalogram signal data, eye image data, sound data, and gesture image data; and a central fusion and control unit, configured to generate vehicle control instructions based on the driver's intention using multimodal fusion decision making.
[0010] Furthermore, the intention tendency acquisition unit specifically includes: The brain-computer interface subsystem is used to collect and pre-process EEG signal data to obtain the first intention tendency; And a multimodal interaction subsystem is used to collect and preprocess eye image data, sound data and gesture image data to obtain a second intention tendency, a third intention tendency and a fourth intention tendency.
[0011] Furthermore, the brain-computer interface subsystem is specifically used to: pre-process the EEG signal data, amplify and filter the EEG signals, extract EEG feature signals related to the driver's status and intention, identify the driver's intention tendency through deep learning methods such as convolutional neural networks, recurrent neural networks and end-to-end models, and obtain a first intention tendency.
[0012] Furthermore, the multimodal interaction subsystem includes: An eye tracking module is used to pre-process the eye image data, identify the pupil position, sight direction, and gaze duration, and obtain a second intention tendency; a speech recognition module, configured to pre-process the sound data, perform noise reduction, feature extraction, speech recognition, and text conversion on the sound data to obtain a third intention tendency; and a gesture recognition module, which is used to pre-process the gesture image data, identify the gesture type and gesture parameters, and obtain a fourth intention tendency.
[0013] Furthermore, the central fusion and control unit is specifically used to: based on preset fusion rules and weight coefficients, fuse the first intention tendency, the second intention tendency, the third intention tendency and the fourth intention tendency to obtain multimodal fusion signal information; make decisions based on the multimodal fusion signal information, generate vehicle control instructions, and feedback the control operation results to the central fusion and control unit.
[0014] Furthermore, the intelligent cockpit interactive control system also includes a device execution subsystem, which is used to: execute corresponding vehicle control operations based on the vehicle control instructions.
[0015] A vehicle in the present invention applies the above-mentioned intelligent cockpit interaction control method, or includes the above-mentioned intelligent cockpit interaction control system.
[0016] The beneficial effect of the present invention is that the present invention can obtain more accurate vehicle control instructions that meet the driver's expectations by preprocessing and multimodal fusion processing of EEG signal data, eye image data, sound data and gesture image data, thereby preventing the vehicle from being unable to execute due to conflicts between vehicle control instructions output by different modules. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] In order to make the purpose, technical solutions and beneficial effects of the present invention more clear, the present invention provides the following drawings for illustration: Figure 1 Schematic diagram of the flow of the intelligent cockpit interactive control method of the present invention; Figure 2 This is an architectural diagram of the intelligent cockpit interactive control system of the present invention. DETAILED DESCRIPTION
[0018] The technical solution of the present invention is described in detail below with reference to the accompanying drawings and embodiments.
[0019] Example 1: like Figure 1 As shown, a smart cockpit interactive control method in this embodiment includes: S1. Acquire physiological and behavioral data of the driver, wherein the physiological and behavioral data include: electroencephalogram signal data, eye image data, sound data, and gesture image data.
[0020] S2. Preprocess the physiological and behavioral data to obtain the driver's intention tendency.
[0021] This step specifically includes: S201. Preprocess the EEG signal data, amplify and filter the EEG signal, extract EEG feature signals related to the driver's state and intention, identify the driver's intention tendency through deep learning methods such as convolutional neural network, recurrent neural network and end-to-end model, and obtain a first intention tendency.
[0022] Specifically, the brain-computer interface subsystem includes an EEG signal processing unit and multiple EEG signal sensors. Multiple EEG signal sensors are distributed in locations such as the driver's seat headrest and steering wheel. After the EEG signal sensor collects the EEG signal data, the EEG signal processing unit amplifies and filters the EEG signal data to extract EEG feature signals related to the driver's status and intention. The EEG feature signals can reflect the driver's status (such as fatigue, concentration, emotions, etc.), and then directly output the intention classification results (such as wanting to turn on or off a device, turn, or overtake) through deep learning methods such as convolutional neural networks (CNN), recurrent neural networks (RNN) and end-to-end models.
[0023] CNNs automatically learn local and global features from EEG signals through components such as convolutional layers, pooling layers, and fully connected layers. When processing EEG signals, convolutional layers perform convolution operations on signals from different electrode locations to extract spatial features. Pooling layers reduce feature dimensionality, minimizing computational effort while retaining key features. After multiple layers of convolution and pooling, fully connected layers integrate the extracted features for subsequent classification or regression tasks.
[0024] RNNs process EEG signals with time series characteristics. By using hidden layers to remember information from previous moments and combining it with the current input, RNNs model the signal's time series characteristics. In a motor imagery-based brain-computer interface, as a user imagines different movements, EEG signals exhibit specific temporal patterns. Deep learning is used to identify the user's intent and determine the primary intention.
[0025] During normal driving, the brain-computer interface subsystem continuously collects EEG signals, collecting multiple sample points every second. After amplification, filtering, and feature extraction by the signal processing unit, a digital signal reflecting the driver's state and intention is generated. For example, if the proportion of alpha waves in the EEG signal increases, it may indicate that the driver is relaxed; if the proportion of beta waves increases, it may indicate that the driver is nervous or focused.
[0026] S202: Pre-process the eye image data to identify the pupil position, sight direction, and gaze time to obtain a second intention tendency.
[0027] Specifically, the multimodal interaction subsystem's eye tracking module includes an eye-image capture camera and an image processing algorithm. The eye-image capture camera is mounted directly above the dashboard, and the camera's angle and focal length are adjusted to clearly capture the driver's eye images. Camera parameters, such as resolution (e.g., 1280×720) and frame rate (e.g., 30fps), are initialized. Simultaneously, the eye-tracking image processing algorithm is initialized, setting thresholds for pupil recognition and error ranges for gaze direction determination. The eye-image capture camera acquires the driver's eye image data in real time, and the image processing algorithm analyzes this data to determine pupil position, gaze direction, and gaze duration, thereby deriving the secondary intention.
[0028] For example, the eye tracking module captures the driver's eye images at a rate of 30 frames per second, and identifies the pupil position, gaze direction, and gaze duration through image processing algorithms. If the driver's gaze stays on the navigation screen for more than 3 seconds, the system can determine that the driver may be interested in the navigation information.
[0029] S203 : Preprocess the sound data by performing noise reduction, feature extraction, speech recognition, and text conversion on the sound data to obtain a third intention tendency.
[0030] Specifically, the multimodal interaction subsystem's speech recognition module includes microphones and a speech recognition algorithm. The microphone array is mounted in locations such as below the front windshield and above the driver's seat. The microphones are calibrated to ensure phase and gain matching between each microphone. The speech recognition algorithm is initialized, setting the language type (e.g., Chinese, English), vocabulary (including commonly used terms for in-cabin interaction), and noise reduction algorithm parameters. The speech recognition module collects real-time sound signal data from within the vehicle and, through noise reduction, feature extraction, speech recognition, and text conversion, converts the sound signal data into text commands.
[0031] For example, the voice recognition module collects and analyzes the driver's voice data in real time. If it recognizes that the driver's voice contains instructions such as "turn on navigation" or "play music", it generates corresponding text instructions and uses the text instructions as the third intention tendency.
[0032] S204: Preprocess the gesture image data, identify gesture type and gesture parameters, and obtain a fourth intention tendency.
[0033] Specifically, the gesture recognition module of the multimodal interaction subsystem includes a gesture acquisition camera and a gesture recognition algorithm. The gesture acquisition camera is a wide-angle camera installed in the center of the vehicle's ceiling, covering the entire cabin space. The camera's viewing angle and focal length are adjusted to fully capture the driver's gestures. The gesture recognition algorithm is initialized to define gesture types (such as finger pointing, fist clenching, waving, and other corresponding operations), gesture recognition sensitivity, etc., and the gesture recognition algorithm is pre-set with a gesture command library that associates different gesture types with different operation commands. The gesture recognition module continuously monitors gestures in the cabin. When it detects that the driver makes a gesture type corresponding to the gesture command library, it determines the operation command the driver desires to implement based on the association relationship in the gesture command library and uses this operation command as the fourth intention tendency.
[0034] For example, the gesture recognition module continuously monitors gesture movements in the cabin. When it detects that the driver makes a fist movement, according to the definition in the gesture instruction library of the gesture recognition algorithm, the fist movement is associated with closing the window, and the fourth intention tendency is determined to be closing the window.
[0035] S3. Based on the driver's intention, a vehicle control instruction is generated using multimodal fusion decision making.
[0036] This step includes: S301: Based on preset fusion rules and weight coefficients, the first intention tendency, the second intention tendency, the third intention tendency and the fourth intention tendency are fused to obtain multimodal fusion signal information.
[0037] S302: Make a decision based on the multimodal fusion signal information and generate a vehicle control instruction.
[0038] The central fusion control unit's fusion rules can be weighted differently based on the driver's state, which can range from normal driving to fatigue. These weights can be used to prioritize the importance of different driver intentions. When conflicting or inconsistent driver intentions identified by different modules occur, fusion processing ensures that more accurate vehicle control commands that meet the driver's expectations are obtained.
[0039] Under normal driving conditions, the signal weight of the brain-computer interface subsystem may be low, while the signal weights of the eye tracking module and the speech recognition module are high. At this time, the weight of the eye tracking module is 0.3, the weight of the speech recognition module is 0.4, the weight of the gesture recognition module is 0.2, and the weight of the brain-computer interface subsystem is 0.1.
[0040] For example, under normal driving conditions, if the eye tracking module detects that the driver's gaze remains on the navigation screen for more than three seconds, it outputs a second intent tendency: "The driver may be interested in navigation information." If, at the same time, the speech recognition module recognizes that the driver's voice contains the phrase "turn on navigation," it outputs a third intent tendency: "turn on navigation." The gesture recognition module fails to recognize the gesture command, and the brain-computer interface subsystem outputs a first intent tendency: "turn on a device." At this point, the central fusion and control unit receives the first, second, third, and fourth intent tendencies, fuses them, and generates multimodal fusion signal information. Based on this multimodal fusion signal information, the central fusion and control unit makes a decision and generates the vehicle control instruction: "turn on navigation." Furthermore, due to the varying weights of different modules, if the gesture recognition module fails to recognize the gesture command, but instead recognizes the gesture command associated with "turn off navigation," the gesture recognition module has a lower weight, and the driver's intent tendencies for the eye tracking module, speech recognition module, and brain-computer interface subsystem are all related to "turn on navigation," the central fusion and control unit ultimately makes a decision based on the multimodal fusion signal information and generates the vehicle control instruction: "turn on navigation."
[0041] However, when the driver is driving fatigued or in a special state, the signal weight of the brain-computer interface subsystem will increase accordingly. Specifically, when the brain-computer interface subsystem detects fatigue characteristics in the driver's EEG signal, and the eye tracking module finds that the driver's vision has become blurred (such as decreased pupil focusing ability and frequent deviation of the vision range), and the speech recognition module detects that the driver's speech has become unclear or the speech speed has slowed down, these signals will be sent to the central fusion and control unit at the same time. The central fusion and control unit fuses these signals according to the fusion rules and weight coefficients preset by the fusion algorithm. In this special case, the signal weight of the brain-computer interface subsystem is increased (for example, adjusted to 0.4), and the weights of other modal signals are adjusted accordingly (such as the eye tracking module weight is adjusted to 0.2, the speech recognition module weight is adjusted to 0.2, and the gesture recognition module weight is adjusted to 0.2). Comprehensively judging that the driver is in a state of fatigue, special countermeasures need to be taken. The decision-making and implementation of special countermeasures are as follows: If the driver is deemed fatigued, the Central Fusion and Control Unit (CFCU) will generate a warning to the driver to rest. This warning will be issued via the in-vehicle voice prompt system, stating "You are fatigued, please take a break." A fatigue warning icon will also be displayed on the instrument panel. If the vehicle has an automated driving assistance feature and meets the requirements for automated driving (e.g., road conditions permitting, vehicle hardware functioning properly), the CFCU will also generate a command to activate the automated driving assistance feature, handing over partial or full driving control to the automated driving assistance system to ensure safe driving.
[0042] S4. Based on the vehicle control instruction, execute corresponding vehicle control operations and feed back control operation results to the central fusion and control unit.
[0043] After the vehicle control command is sent to the device execution subsystem, the device execution subsystem executes the command and feeds back the operation result to the central fusion and control unit, which can make further confirmation or adjustments based on the feedback information.
[0044] Vehicle control commands can control the seats, air conditioning, windows, and other functions within the cabin, as well as the content displayed on the in-cabin screen and navigation information. This step is primarily implemented by the device execution subsystem, which includes the seat control module, air conditioning control module, body control module, navigation module, in-cabin screen control module, and intelligent driving module.
[0045] Example 2: like Figure 2 As shown, an intelligent cockpit interactive control system in this embodiment includes: an intention tendency acquisition unit, configured to acquire physiological and behavioral data of the driver and pre-process the physiological and behavioral data to obtain the driver's intention tendency; wherein the physiological and behavioral data include: electroencephalogram signal data, eye image data, sound data, and gesture image data; and a central fusion and control unit, configured to generate vehicle control instructions based on the driver's intention using multimodal fusion decision making.
[0046] In this embodiment, the intention tendency acquisition unit specifically includes: The brain-computer interface subsystem is used to collect and pre-process EEG signal data to obtain the first intention tendency; And a multimodal interaction subsystem is used to collect and preprocess eye image data, sound data and gesture image data to obtain a second intention tendency, a third intention tendency and a fourth intention tendency.
[0047] Correspondingly, the central fusion and control unit is configured to: generate a vehicle control instruction using a multimodal fusion decision based on the first intention tendency, the second intention tendency, the third control instruction, and the fourth control instruction; In this embodiment, the intelligent cockpit interaction control system also includes a device execution subsystem, which is configured to execute corresponding vehicle control operations based on the vehicle control instructions. After the vehicle control instructions are sent to the device execution subsystem, the device execution subsystem executes the instructions and feeds back the operation results to the central fusion and control unit, which can then make further confirmation or adjustments based on the feedback information.
[0048] Vehicle control commands can control the seats, air conditioning, windows, and other functions within the cabin, as well as the onboard screen display content and navigation information. The device execution subsystem includes the seat control module, air conditioning control module, body control module, navigation module, onboard screen control module, and intelligent driving module.
[0049] In this embodiment, the brain-computer interface subsystem is specifically used to: preprocess the EEG signal data, amplify and filter the EEG signals, extract EEG feature signals related to the driver's status and intentions, identify the driver's intention tendencies through deep learning methods such as convolutional neural networks, recurrent neural networks, and end-to-end models, and obtain a first intention tendency.
[0050] Specifically, the brain-computer interface subsystem includes an EEG signal processing unit and multiple EEG signal sensors. Multiple EEG signal sensors are distributed in locations such as the driver's seat headrest and steering wheel. After the EEG signal sensor collects the EEG signal data, the EEG signal processing unit amplifies and filters the EEG signal data to extract EEG feature signals related to the driver's status and intention. The EEG feature signals can reflect the driver's status (such as fatigue, concentration, emotions, etc.), and then directly output the intention classification results (such as wanting to turn on or off a device, turn, or overtake) through deep learning methods such as convolutional neural networks (CNN), recurrent neural networks (RNN) and end-to-end models.
[0051] CNNs automatically learn local and global features from EEG signals through components such as convolutional layers, pooling layers, and fully connected layers. When processing EEG signals, convolutional layers perform convolution operations on signals from different electrode locations to extract spatial features. Pooling layers reduce feature dimensionality, minimizing computational effort while retaining key features. After multiple layers of convolution and pooling, fully connected layers integrate the extracted features for subsequent classification or regression tasks.
[0052] RNNs process EEG signals with time series characteristics. By using hidden layers to remember information from previous moments and combining it with the current input, RNNs model the signal's time series characteristics. In a motor imagery-based brain-computer interface, as a user imagines different movements, EEG signals exhibit specific temporal patterns. Deep learning is used to identify the user's intent and determine the primary intention.
[0053] During normal driving, the brain-computer interface subsystem continuously collects EEG signals, collecting multiple sample points every second. After amplification, filtering, and feature extraction by the signal processing unit, a digital signal reflecting the driver's state and intention is generated. For example, if the proportion of alpha waves in the EEG signal increases, it may indicate that the driver is relaxed; if the proportion of beta waves increases, it may indicate that the driver is nervous or focused.
[0054] In this embodiment, the multimodal interaction subsystem includes: An eye tracking module is used to pre-process the eye image data, identify the pupil position, sight direction, and gaze duration, and obtain a second intention tendency; a speech recognition module, configured to pre-process the sound data, perform noise reduction, feature extraction, speech recognition, and text conversion on the sound data to obtain a third intention tendency; and a gesture recognition module, which is used to pre-process the gesture image data, identify the gesture type and gesture parameters, and obtain a fourth intention tendency.
[0055] Specifically, the multimodal interaction subsystem's eye tracking module includes an eye-image capture camera and an image processing algorithm. The eye-image capture camera is mounted directly above the dashboard, and the camera's angle and focal length are adjusted to clearly capture the driver's eye images. Camera parameters, such as resolution (e.g., 1280×720) and frame rate (e.g., 30fps), are initialized. Simultaneously, the eye-tracking image processing algorithm is initialized, setting thresholds for pupil recognition and error ranges for gaze direction determination. The eye-image capture camera acquires the driver's eye image data in real time, and the image processing algorithm analyzes this data to determine pupil position, gaze direction, and gaze duration, thereby deriving the secondary intention.
[0056] For example, the eye tracking module captures the driver's eye images at a rate of 30 frames per second, and identifies the pupil position, gaze direction, and gaze duration through image processing algorithms. If the driver's gaze stays on the navigation screen for more than 3 seconds, the system can determine that the driver may be interested in the navigation information.
[0057] The multimodal interaction subsystem's speech recognition module includes microphones and a speech recognition algorithm. The microphone array is mounted in locations such as below the front windshield and above the driver's seat. The microphones are calibrated to ensure phase and gain matching between each microphone. The speech recognition algorithm is initialized, with settings for the language type (e.g., Chinese, English), vocabulary (including commonly used terms for in-cockpit interaction), and noise reduction algorithm parameters. The speech recognition module collects real-time sound signal data from within the vehicle and converts it into text commands through noise reduction, feature extraction, speech recognition, and text conversion.
[0058] For example, the voice recognition module collects and analyzes the driver's voice data in real time. If it recognizes that the driver's voice contains instructions such as "turn on navigation" or "play music", it generates corresponding text instructions and uses the text instructions as the third intention tendency.
[0059] The gesture recognition module includes a gesture acquisition camera and a gesture recognition algorithm. The gesture acquisition camera is a wide-angle camera installed in the center of the vehicle's ceiling, used to cover the entire cabin space. The camera's viewing angle and focal length are adjusted to enable it to fully capture the driver's gestures. The gesture recognition algorithm is initialized to define gesture types (such as corresponding operations such as finger pointing, fist clenching, and waving), gesture recognition sensitivity, etc., and the gesture recognition algorithm is pre-set with a gesture command library that associates different gesture types with different operation commands. The gesture recognition module continuously monitors gestures in the cabin. When it detects that the driver makes a gesture type corresponding to the gesture command library, it determines the operation command the driver expects to implement based on the association relationship in the gesture command library and uses this operation command as the fourth intention tendency.
[0060] For example, the gesture recognition module continuously monitors gesture movements in the cabin. When it detects that the driver makes a fist movement, according to the definition in the gesture instruction library of the gesture recognition algorithm, the fist movement is associated with closing the window, and the fourth intention tendency is determined to be closing the window.
[0061] In this embodiment, the central fusion and control unit is specifically used to: based on preset fusion rules and weight coefficients, fuse the first intention tendency, the second intention tendency, the third intention tendency and the fourth intention tendency to obtain multimodal fusion signal information; make decisions based on the multimodal fusion signal information to generate vehicle control instructions.
[0062] The central fusion control unit's fusion rules can be weighted differently based on the driver's state, which can range from normal driving to fatigue. These weights can be used to prioritize the importance of different driver intentions. When conflicting or inconsistent driver intentions identified by different modules occur, fusion processing ensures that more accurate vehicle control commands that meet the driver's expectations are obtained.
[0063] Under normal driving conditions, the signal weight of the brain-computer interface subsystem may be low, while the signal weights of the eye tracking module and the speech recognition module are high. At this time, the weight of the eye tracking module is 0.3, the weight of the speech recognition module is 0.4, the weight of the gesture recognition module is 0.2, and the weight of the brain-computer interface subsystem is 0.1.
[0064] For example, under normal driving conditions, if the eye tracking module detects that the driver's gaze remains on the navigation screen for more than three seconds, it outputs a second intent tendency: "The driver may be interested in navigation information." If, at the same time, the speech recognition module recognizes that the driver's voice contains the phrase "turn on navigation," it outputs a third intent tendency: "turn on navigation." The gesture recognition module fails to recognize the gesture command, and the brain-computer interface subsystem outputs a first intent tendency: "turn on a device." At this point, the central fusion and control unit receives the first, second, third, and fourth intent tendencies, and after fusing them, generates multimodal fusion signal information. The central fusion and control unit makes a decision based on the multimodal fusion signal information and generates a vehicle control instruction: "turn on navigation." Furthermore, due to the different weights of different modules, if the gesture recognition module fails to recognize the gesture command, but instead recognizes the gesture command associated with "turn off navigation," the gesture recognition module has a lower weight than the speech recognition module, and the driver's intent tendencies output by the eye tracking module, speech recognition module, and brain-computer interface subsystem are all related to "turn on navigation," the central fusion and control unit ultimately makes a decision based on the multimodal fusion signal information and generates the same vehicle control instruction: "turn on navigation."
[0065] However, when the driver is driving fatigued or in a special state, the signal weight of the brain-computer interface subsystem will increase accordingly. Specifically, when the brain-computer interface subsystem detects fatigue characteristics in the driver's EEG signal, and the eye tracking module finds that the driver's vision has become blurred (such as decreased pupil focusing ability and frequent deviation of the vision range), and the speech recognition module detects that the driver's speech has become unclear or the speech speed has slowed down, these signals will be sent to the central fusion and control unit at the same time. The central fusion and control unit fuses these signals according to the fusion rules and weight coefficients preset by the fusion algorithm. In this special case, the signal weight of the brain-computer interface subsystem is increased (for example, adjusted to 0.4), and the weights of other modal signals are adjusted accordingly (such as the eye tracking module weight is adjusted to 0.2, the speech recognition module weight is adjusted to 0.2, and the gesture recognition module weight is adjusted to 0.2). Comprehensively judging that the driver is in a state of fatigue, special countermeasures need to be taken. The decision-making and implementation of special countermeasures are as follows: If the driver is deemed fatigued, the Central Fusion and Control Unit (CFCU) will generate a warning to the driver to rest. This warning will be issued via the in-vehicle voice prompt system, stating "You are fatigued, please take a break." A fatigue warning icon will also be displayed on the instrument panel. If the vehicle has an automated driving assistance feature and meets the requirements for automated driving (e.g., road conditions permitting, vehicle hardware functioning properly), the CFCU will also generate a command to activate the automated driving assistance feature, handing over partial or full driving control to the automated driving assistance system to ensure safe driving.
[0066] Example 3: In this embodiment, a vehicle employs the above-described intelligent cockpit interaction control method or includes the above-described intelligent cockpit interaction control system. The vehicle may be, but is not limited to, a pure electric vehicle (PEV / BEV), a hybrid electric vehicle (HEV), a range-extended electric vehicle (REEV), a plug-in hybrid electric vehicle (PHEV), a new energy vehicle (NEV), or a fuel-powered vehicle.
[0067] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the purpose and scope of the technical solutions of the present invention, which should all be included in the scope of the claims of the present invention.
Claims
1. A smart cockpit interactive control method, characterized in that: include: Acquiring physiological and behavioral data of the driver, wherein the physiological and behavioral data include: electroencephalogram signal data, eye image data, sound data, and gesture image data; Preprocessing the physiological and behavioral data to obtain the driver's intention tendency; Based on the driver's intention tendency, a multimodal fusion decision is adopted to generate vehicle control instructions.
2. The intelligent cockpit interactive control method according to claim 1, characterized in that: The pre-processing of the physiological and behavioral data to obtain the driver's intention tendency includes: The EEG signal data is preprocessed, the EEG signal is amplified and filtered, and EEG feature signals related to the driver's state and intention are extracted. The driver's intention tendency is identified through deep learning methods such as convolutional neural network, recurrent neural network and end-to-end model to obtain a first intention tendency.
3. The intelligent cockpit interactive control method according to claim 2, characterized in that: The pre-processing of the physiological and behavioral data to obtain the driver's intention tendency further includes: Preprocessing the eye image data to identify the pupil position, sight direction, and gaze time of the eyeball to obtain a second intention tendency; Preprocessing the sound data, performing noise reduction, feature extraction, speech recognition, and text conversion on the sound data to obtain a third intention tendency; The gesture image data is preprocessed to identify the gesture type and gesture parameters to obtain a fourth intention tendency.
4. The intelligent cockpit interactive control method according to claim 3, characterized in that: The method of generating a vehicle control instruction based on the driver's intention tendency by adopting a multimodal fusion decision includes: Based on a preset fusion rule and a weight coefficient, the first intention tendency, the second intention tendency, the third intention tendency, and the fourth intention tendency are fused to obtain multimodal fusion signal information; A decision is made based on the multimodal fusion signal information to generate a vehicle control instruction.
5. An intelligent cockpit interactive control system, characterized in that: include: an intention tendency acquisition unit, configured to acquire physiological and behavioral data of the driver and pre-process the physiological and behavioral data to obtain the driver's intention tendency; wherein the physiological and behavioral data include: electroencephalogram signal data, eye image data, sound data, and gesture image data; and a central fusion and control unit, configured to generate vehicle control instructions based on the driver's intention using multimodal fusion decision making.
6. The intelligent cockpit interactive control system according to claim 5, characterized in that: The intention tendency acquisition unit specifically includes: The brain-computer interface subsystem is used to collect and pre-process EEG signal data to obtain the first intention tendency; And a multimodal interaction subsystem is used to collect and preprocess eye image data, sound data and gesture image data to obtain a second intention tendency, a third intention tendency and a fourth intention tendency.
7. The intelligent cockpit interactive control system according to claim 6, characterized in that: The brain-computer interface subsystem is specifically used to: pre-process the EEG signal data, amplify and filter the EEG signals, extract EEG feature signals related to the driver's status and intention, identify the driver's intention tendency through deep learning methods such as convolutional neural networks, recurrent neural networks and end-to-end models, and obtain a first intention tendency.
8. The intelligent cockpit interactive control system according to claim 6, characterized in that: The multimodal interaction subsystem includes: An eye tracking module is used to pre-process the eye image data, identify the pupil position, sight direction, and gaze duration, and obtain a second intention tendency; a speech recognition module, configured to pre-process the sound data, perform noise reduction, feature extraction, speech recognition, and text conversion on the sound data to obtain a third intention tendency; and a gesture recognition module, which is used to pre-process the gesture image data, identify the gesture type and gesture parameters, and obtain a fourth intention tendency.
9. The intelligent cockpit interactive control system according to claim 6, characterized in that: The central fusion and control unit is specifically configured to: fuse the first intention tendency, the second intention tendency, the third intention tendency, and the fourth intention tendency based on preset fusion rules and weight coefficients to obtain multimodal fusion signal information; A decision is made based on the multimodal fusion signal information to generate a vehicle control instruction.
10. A vehicle, characterized in that: The intelligent cockpit interaction control method according to any one of claims 1 to 4 is applied, or the intelligent cockpit interaction control system according to any one of claims 5 to 9 is included.
Citation Information
Cited By
Vehicle-mounted intelligent cabin dialogue method and system based on multi-modal interaction
CN122232648A
A Dialogue Method and System for In-Vehicle Intelligent Cockpit Based on Multimodal Interaction
CN122232648B
A brain-computer interaction-based vehicle-mounted non-inductive driving intention recognition and vehicle control method and system
CN122667064A