Fuel vehicle man-vehicle interaction system based on artificial intelligence
By employing a multimodal perception module, a dynamic adaptive engine, and cross-platform compatible middleware, the system addresses the challenges of speech recognition and emotion analysis in complex environments for human-vehicle interaction systems in gasoline-powered vehicles. This enables high-precision fatigue and emotion monitoring, ensures data privacy and security, adapts to different vehicle models, and improves the system's real-time performance and compatibility.
Patent Information
- Application Number
- CN202511181530.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-22
- Publication Date
- 2025-12-12
AI Technical Summary
Existing AI-based human-vehicle interaction systems for gasoline vehicles suffer from insufficient accuracy in speech recognition in complex environments, inaccurate emotion analysis and behavior prediction, data privacy and security issues, poor system adaptability, hardware and software compatibility problems, and insufficient real-time performance and computing power.
Employing a multimodal perception module, a dynamic adaptive engine, a privacy enhancement processing module, and cross-platform compatible middleware, the system collects data through devices such as directional microphone arrays, infrared cameras, millimeter-wave radar, and seat pressure sensors. It then combines this data with a federated learning framework to generate personalized interaction models, perform differential privacy protection and localization processing, and adapt to the hardware differences of different vehicle models to achieve cross-platform compatibility.
It improves the accuracy of voice recognition and the precision of fatigue/emotion monitoring in complex environments, ensures data privacy and security, adapts to different vehicle models, avoids over-reliance on technology, improves the real-time performance and compatibility of the system, and reduces the risk of misoperation and privacy leakage.
Smart Images

Figure CN121106064A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to an artificial intelligence-based human-vehicle interaction system for fuel-powered vehicles, belonging to the field of automotive intelligent interaction technology. Background Technology
[0002] Artificial intelligence-based human-vehicle interaction systems for gasoline-powered vehicles aim to enhance driving experience and safety through intelligent algorithms and perception technologies. Utilizing AI technologies such as voice recognition, natural language processing, and emotion analysis, the system can engage in natural conversation with the driver, enabling control of various in-vehicle functions, such as navigation, air conditioning, and entertainment. Furthermore, the system can intelligently provide feedback based on the driver's driving habits, emotional state, and physiological characteristics (such as fatigue levels), offering personalized driving suggestions and safety warnings. For example, when driver fatigue is detected, the system will remind the driver to rest or automatically adjust the in-vehicle environment to optimize the driving experience. Through continuous learning and adaptation, AI-based interaction systems can continuously improve the intelligence of human-vehicle interaction, providing users with a more convenient, safe, and comfortable driving experience.
[0003] Currently, while AI-based human-vehicle interaction systems for gasoline-powered vehicles have made some progress in improving the driving experience and safety, several shortcomings and challenges remain, including: **Accuracy of Voice Recognition:** Voice recognition is the core of in-vehicle AI interaction, but in complex environments (such as highway noise and background noise inside the car), the system's accuracy still suffers from significant deviations. Sometimes the system fails to accurately understand the driver's instructions or misunderstands them, leading to erroneous operations. **Limitations of Sentiment Analysis and Behavior Prediction:** AI systems provide personalized feedback by analyzing data such as the driver's emotions and fatigue, but current technology has not yet achieved perfect accuracy and real-time performance. Incorrect emotion recognition or fatigue prediction may affect the system's suggestions and interventions for the driver. For example, the system may underestimate the driver's fatigue level or misjudge the driver's emotional state, leading to inappropriate reactions. **Data Privacy and Security Issues:** AI-based interaction systems typically require the collection of large amounts of personal data, such as driving behavior, voice, facial expressions, and location. The storage and use of this data may lead to privacy breaches and data security issues. Furthermore, hacker attacks could compromise in-vehicle systems, creating security risks; poor system adaptability: although AI systems can learn driver habits and preferences, this learning process still has a certain lag. The system may need a considerable amount of time to accurately adapt to the needs of different drivers, and poor adaptation may occur when switching between different drivers; over-reliance on technology leading to driver negligence: as AI systems become more widespread in vehicles, some drivers may become overly reliant on technology, neglecting their driving responsibilities. For example, over-reliance on autonomous driving or voice control may prevent drivers from reacting quickly enough in critical moments, increasing the risk of accidents; hardware and software compatibility issues: currently, in-vehicle systems from different manufacturers vary significantly in hardware and software compatibility. Some AI functions may only be implemented on specific models and not work properly on others, limiting the system's widespread adoption and cross-platform application; challenges in real-time performance and computing power: AI systems require substantial data processing and computational support, placing high demands on in-vehicle computing platforms. Although more and more gasoline vehicles are equipped with more powerful hardware, computing power remains insufficient in some low-end or older models, limiting the reaction speed and real-time performance of AI systems.
[0004] To address this, an artificial intelligence-based human-vehicle interaction system for fuel-powered vehicles is proposed. Summary of the Invention
[0005] In view of this, the present invention provides an artificial intelligence-based human-vehicle interaction system for fuel vehicles to solve or alleviate the technical problems existing in the prior art, and at least provides a beneficial alternative.
[0006] The technical solution of the present invention is implemented as follows: an artificial intelligence-based human-vehicle interaction system for fuel vehicles, including a multimodal perception module, a dynamic adaptive engine, a privacy enhancement processing module, and cross-platform compatible middleware; The multimodal perception module is used to collect the driver's voice commands, facial expressions, physiological characteristics, and in-vehicle environment data; the dynamic adaptive engine generates personalized interaction models based on a federated learning framework and updates them in real time; the privacy enhancement processing module performs differential privacy protection and localization processing on sensitive data; and the cross-platform compatible middleware adapts to the hardware differences of different vehicle models and standardizes data interfaces.
[0007] More preferably, the multimodal perception module includes: a directional microphone array (supporting noise suppression and sound source localization), an infrared camera (monitoring the driver's eye opening and facial micro-expressions), a millimeter-wave radar (non-contact detection of heart rate variability and respiratory rate), a seat pressure sensor (analyzing sitting posture stability), and a CAN bus interface (acquiring vehicle status data such as vehicle speed and throttle depth).
[0008] More preferably, the dynamic adaptive engine includes: a local model training unit (which encrypts and trains personalized parameters based on the driver's historical interaction data), a global model update unit (which receives lightweight general model increments from the vehicle manufacturer's server), and a real-time decision-making unit (which integrates local and global model outputs to dynamically adjust the interaction strategy).
[0009] More preferably, the privacy enhancement processing module includes: a differential privacy encryption unit (adding controllable noise to voice command text and driving behavior logs), a local data storage unit (sensitive information is stored only in the vehicle encryption chip), and a secure communication unit (encrypting data transmission with the cloud via the SM algorithm).
[0010] More preferably, the cross-platform compatible middleware includes: a hardware abstraction layer (unifying the microphone / camera driver interface for different vehicle models), a function degradation strategy unit (automatically disabling non-core functions based on the vehicle's computing power, such as disabling millimeter-wave radar monitoring in low-end models), and a standardized API interface (allowing automakers to quickly access third-party sensors).
[0011] More preferably, the multimodal perception module fuses voice commands and facial expression data (timestamp synchronization error < 50ms) through a spatiotemporal alignment algorithm. When the voice recognition confidence is lower than the threshold, it automatically calls on facial micro-expressions and physiological features to assist in judging the driver's true intention.
[0012] More preferably, the dynamic adaptive engine uses multi-dimensional indicators for fatigue monitoring (including eye closure time > 3 seconds, heart rate variability reduced by 20%, and head droop angle > 15°). When a single indicator is abnormal, an early warning is triggered, and when multiple indicators are abnormal in tandem, lane keeping assist or automatic parking is forcibly activated.
[0013] More preferably, the privacy enhancement processing module performs local speech-to-text processing on the driver's voice commands (without uploading the original audio), and only encrypts the text commands and intent tags before transmitting them to the cloud for model optimization.
[0014] More preferably, the cross-platform compatible middleware supports OTA remote upgrades (by pushing hardware adaptation patches through the vehicle manufacturer's server) and is compatible with both the traditional CAN bus protocol and the emerging Ethernet communication protocol for gasoline vehicles.
[0015] More preferably, the system also includes a driver status visualization feedback unit (displaying the current fatigue level, emotional state, and system suggestions via the dashboard or HUD), allowing the driver to manually override AI decisions (such as rejecting rest reminders), and the system records human intervention data and optimizes subsequent strategies.
[0016] The embodiments of the present invention have the following advantages due to the adoption of the above technical solutions: I. This invention significantly improves environmental robustness and leads the industry in interactive accuracy. Traditional in-vehicle voice systems often suffer from a voice recognition error rate exceeding 15% in highway scenarios at speeds of 100 km / h due to wind noise (approximately 75-85 dB), tire noise (approximately 70-80 dB), and multiple conversations within the vehicle (more than 3 background noise sources). For example, "turn up the air conditioning temperature" might be misheard as "turn off navigation." This system utilizes a directional microphone array (beamforming technology to focus on the driver's voice source and suppress non-target direction noise by more than 30 dB) + The real-time noise suppression algorithm (based on deep learning-based spectral subtraction for dynamic background noise filtering) achieves a voice command recognition accuracy of ≥98% in high-speed scenarios (compared to approximately 85% in traditional systems), reducing the error rate to below 0.5% (compared to approximately 3% in traditional systems). When the confidence level of a voice command falls below a threshold (e.g., only 80% in noisy environments), the system automatically utilizes an infrared camera (to analyze the driver's lip movement trajectory and facial orientation) and physiological sensors (e.g., to determine whether the driver is turning their head based on seat pressure distribution). Through multi-source data fusion, the accuracy of parsing ambiguous commands is improved from 70% to 95% (e.g., reducing the misjudgment rate of distinguishing between "open the window" and "close the sunroof" by 80%).
[0017] II. This invention provides real-time and accurate fatigue / emotion monitoring with industry-leading safety warning response speed. Traditional systems rely on single indicators (such as steering wheel off-hand detection or simple eye closure monitoring), resulting in a misjudgment rate as high as 30% (e.g., a driver briefly closing their eyes and yawning is misjudged as severe fatigue). This system integrates millimeter-wave radar (non-contact monitoring of heart rate variability (HRV) decrease of more than 20%, reflecting autonomic nervous system fatigue), infrared cameras (eye closure duration > 3 seconds, blinking frequency < 3 times / minute, facial micro-expressions (drooping corners of the mouth / frowning eyebrows)), and seat pressure sensors (sitting posture offset angle > 15° or body slumping trend). Through a dynamic weighting algorithm (the contribution of different indicators adaptively adjusts with driving time), the fatigue warning accuracy rate is ≥ 92% (compared to about 70% for traditional systems). Furthermore, it can issue a "mild fatigue" warning 10-15 minutes in advance when fatigue is initially detected (e.g., when HRV slightly decreases), suggesting opening windows for ventilation or playing refreshing music. It also utilizes facial micro-expressions (contraction of the brow muscle when angry, and elevation of the orbicularis oculi muscle when happy) + By analyzing voice tone (increased pitch and faster speech when angry) and combining it with driving behavior data (frequency of rapid acceleration / braking), the system can distinguish between "brief agitation" (such as due to traffic jams) and "persistent anger" (which may lead to aggressive driving). When high-risk emotions are detected (such as anger + frequent overtaking intentions), the system automatically lowers the air conditioning fan speed (to reduce noise stimulation), switches to soothing music, and displays a non-intrusive reminder on the HUD that "please take a deep breath and stay calm," avoiding direct voice warnings that could escalate the conflict.
[0018] Third, this invention offers comprehensive and compliant privacy protection with zero risk of user data leakage. Traditional systems often upload sensitive data such as raw voice audio and driving trajectories to the cloud for analysis, posing a privacy risk (e.g., in 202X, a certain brand's in-vehicle system suffered a leak of millions of users' voice records due to an unencrypted cloud database). This system adopts a "local-first" strategy: voice commands are transcribed into text in real time via the in-vehicle ASR module (without storing the raw audio); differential privacy noise (ε=0.5, meeting the ε-differential privacy standard) is added to driving behavior logs (such as the number of emergency brakings and frequently used navigation destinations); only the desensitized intent labels (such as "need to rest" and "preferred music type") are encrypted before being transmitted to the vehicle manufacturer's server for model optimization; sensitive data (such as facial images and heart rate data) are stored only in the in-vehicle security chip (compliant with EAL5+ certification), so that even if the vehicle is physically disassembled, the original information cannot be extracted; data interaction between the system and the cloud / mobile APP is encrypted using the national cryptographic SM4 algorithm (with dynamically updated keys) to prevent man-in-the-middle attacks (such as hackers forging in-vehicle hotspots to steal data); users can choose the "local analysis only" mode (disabling all data uploads) through the vehicle settings.
[0019] IV. This invention offers flexible cross-vehicle adaptation, covering over 90% of the gasoline vehicle market. Traditional high-end models rely on sensors from specific brands (e.g., a certain brand's infrared camera is only compatible with its own models). However, this system abstracts hardware interfaces through cross-platform compatible middleware (e.g., a unified microphone driver protocol and camera resolution adaptation algorithm), supporting sensors from mainstream models on the market (including low-cost infrared modules and basic millimeter-wave radar). It only requires automakers to provide a basic CAN bus protocol for access. For low-end models (without millimeter-wave radar or with computing power <4 TOPS), the system automatically downgrades to basic functions (e.g., monitoring eye status only through infrared cameras and determining posture through seat sensors), ensuring that core interactions (navigation control, air conditioning adjustment) are not affected. It covers the entire price range from 50,000 RMB economy cars to 500,000 RMB luxury SUVs. Automakers can push hardware adaptation patches (e.g., adding driver support for a certain microphone model) and model optimization increments (e.g., special training for poor voice recognition when drivers wear gloves in winter) through the server. Users can obtain the latest functions without going to the store for upgrades (the actual function iteration cycle is shortened from 6 months in traditional systems to 2 weeks).
[0020] V. This invention features human-machine collaborative safety redundancy, avoiding over-reliance on technology. All system suggestions (such as "You have been driving continuously for 2 hours, it is recommended to rest") are non-mandatory prompts. Drivers can manually override AI decisions via voice ("Not needed for now"), physical buttons ("Ignore reminder" button on the center console), or gestures (tap the steering wheel twice). The system records data from each manual intervention (such as the frequency of refusing rest reminders and the time for manually turning off fatigue warnings) and dynamically adjusts subsequent strategies (such as increasing the intensity of visual warnings for drivers who frequently ignore reminders, such as red flashing on the dashboard). When the system detects that the driver is severely fatigued (such as a sudden drop in heart rate + unconscious nodding) and does not respond to reminders, the system automatically triggers LKA (Lane Keeping Assist) to limit vehicle deviation and gradually reduces the vehicle speed to a safe range (such as 60km / h). At the same time, it sends location information to preset emergency contacts through the vehicle network to prevent accidents.
[0021] The above overview is for illustrative purposes only and is not intended to be limiting in any way. In addition to the illustrative aspects, embodiments, and features described above, further aspects, embodiments, and features of the invention will become readily apparent from the accompanying drawings and the following detailed description. Attached Figure Description
[0022] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0023] Figure 1 This is a system schematic diagram of the present invention. Detailed Implementation
[0024] In the following description, only certain exemplary embodiments are briefly described. As those skilled in the art will recognize, the described embodiments can be modified in various ways without departing from the spirit or scope of the invention. Therefore, the drawings and description are considered to be exemplary in nature and not restrictive.
[0025] The embodiments of the present invention will now be described in detail with reference to the accompanying drawings.
[0026] Example like Figure 1 As shown, this embodiment of the invention provides an artificial intelligence-based human-vehicle interaction system for fuel vehicles, including a multimodal perception module, a dynamic adaptive engine, a privacy enhancement processing module, and cross-platform compatible middleware; The multimodal perception module is used to collect the driver's voice commands, facial expressions, physiological characteristics, and in-vehicle environment data; the dynamic adaptive engine generates personalized interaction models based on the federated learning framework and updates them in real time; the privacy enhancement processing module performs differential privacy protection and localization processing on sensitive data; and the cross-platform compatible middleware adapts to the hardware differences of different vehicle models and standardizes data interfaces.
[0027] Furthermore, the multimodal perception module includes: a directional microphone array (supporting noise suppression and sound source localization), an infrared camera (monitoring the driver's eye opening and facial micro-expressions), a millimeter-wave radar (non-contact detection of heart rate variability and respiratory rate), a seat pressure sensor (analyzing seat posture stability), and a CAN bus interface (acquiring vehicle status data such as vehicle speed and throttle depth).
[0028] Furthermore, the dynamic adaptive engine includes: a local model training unit (which encrypts and trains personalized parameters based on the driver's historical interaction data), a global model update unit (which receives lightweight general model increments from the car manufacturer's server), and a real-time decision-making unit (which integrates local and global model outputs to dynamically adjust the interaction strategy).
[0029] Furthermore, the privacy enhancement processing module includes: a differential privacy encryption unit (adding controllable noise to voice command text and driving behavior logs), a local data storage unit (sensitive information is stored only in the vehicle encryption chip), and a secure communication unit (encrypting data transmission with the cloud via the SM algorithm).
[0030] Furthermore, the cross-platform compatibility middleware includes: a hardware abstraction layer (unifying the microphone / camera driver interface for different vehicle models), a function degradation strategy unit (automatically disabling non-core functions based on the vehicle's computing power, such as disabling millimeter-wave radar monitoring in low-end models), and standardized API interfaces (allowing automakers to quickly access third-party sensors).
[0031] Furthermore, the multimodal perception module fuses voice commands and facial expression data (timestamp synchronization error < 50ms) through a spatiotemporal alignment algorithm. When the confidence level of voice recognition is lower than the threshold, it automatically calls upon facial micro-expressions and physiological features to assist in judging the driver's true intentions.
[0032] Furthermore, the dynamic adaptive engine uses multi-dimensional indicators for fatigue monitoring (including eye closure time > 3 seconds, heart rate variability reduced by 20%, and head droop angle > 15°). When a single indicator is abnormal, a warning is triggered, and when multiple indicators are abnormal in tandem, lane keeping assist or automatic parking is forcibly activated.
[0033] Furthermore, the privacy enhancement processing module performs local speech-to-text processing on the driver's voice commands (without uploading the original audio), and only encrypts the text commands and intent tags before transmitting them to the cloud for model optimization.
[0034] Furthermore, the cross-platform compatible middleware supports OTA remote upgrades (by pushing hardware adaptation patches through the vehicle manufacturer's server) and is compatible with both the traditional CAN bus protocol for gasoline vehicles and the emerging Ethernet communication protocol.
[0035] Furthermore, the system also includes a driver status visualization feedback unit (displaying current fatigue level, emotional state, and system suggestions via the dashboard or HUD), allowing drivers to manually override AI decisions (such as rejecting rest reminders), and the system records human intervention data and optimizes subsequent strategies.
[0036] This invention significantly improves environmental robustness and leads the industry in interactive accuracy. Traditional in-vehicle voice systems often suffer from a voice recognition error rate exceeding 15% in highway scenarios at speeds of 100 km / h due to wind noise (approximately 75-85 dB), tire noise (approximately 70-80 dB), and multiple conversations within the vehicle (more than 3 background noise sources). For example, "turn up the air conditioning temperature" may be misheard as "turn off navigation." This system, through a directional microphone array (beamforming technology focuses on the driver's voice source, suppressing non-target direction noise by more than 30 dB) + a real-time noise suppression algorithm (based on deep learning spectral subtraction to dynamically filter background noise), achieves a measured voice command recognition accuracy of ≥98% in high-speed scenarios (compared to approximately 85% for traditional systems), with the error rate reduced to 0.Below 5% (compared to approximately 3% for traditional systems), when the confidence level of a voice command falls below a threshold (e.g., only 80% in noisy environments), the system automatically utilizes an infrared camera (analyzing the driver's lip movements and facial orientation) and physiological sensors (e.g., seat pressure distribution to determine if the driver is turning their head). Through multi-source data fusion, the accuracy of resolving ambiguous commands is increased from 70% to 95% (e.g., reducing the misjudgment rate of distinguishing between "open the window" and "close the sunroof" by 80%). This invention provides real-time and accurate fatigue / emotion monitoring, with industry-leading safety warning response speed. Traditional systems rely on only a single indicator (e.g., steering wheel off-hand detection or simple eye closure monitoring), resulting in a misjudgment rate as high as 30% (e.g., a driver briefly closing their eyes and yawning is misjudged as severely fatigued). This system... Integrating millimeter-wave radar (non-contact monitoring of heart rate variability (HRV) decrease of more than 20%, reflecting autonomic nervous system fatigue), an infrared camera (eye closure duration >3 seconds, blinking frequency <3 times / minute, facial micro-expressions (drooping corners of the mouth / frowning eyebrows)), and a seat pressure sensor (sitting posture offset angle >15° or body slumping trend), and through a dynamic weighting algorithm (the contribution of different indicators adaptively adjusts with driving time), the fatigue warning accuracy rate is ≥92% (compared to about 70% for traditional systems). It can also issue a "mild fatigue" warning 10-15 minutes in advance when fatigue is initially detected (e.g., a slight decrease in HRV), suggesting opening windows for ventilation or playing refreshing music. It also detects facial micro-expressions (contraction of the brow muscle when angry, and elevation of the orbicularis oculi muscle when happy) + By analyzing voice tone (increased pitch and faster speech when angry) and combining it with driving behavior data (frequency of sudden acceleration / braking), the system can distinguish between "brief agitation" (such as due to traffic jams) and "persistent anger" (which may lead to aggressive driving). When high-risk emotions are detected (such as anger + frequent overtaking intentions), the system automatically lowers the air conditioning fan speed (to reduce noise stimulation), switches to soothing music, and displays a non-intrusive reminder on the HUD: "Please take a deep breath and stay calm." This avoids escalating conflicts with direct voice warnings. This invention fully complies with privacy protection regulations, with zero risk of user data leakage. Traditional systems often upload sensitive data such as raw voice audio and driving trajectory to the cloud for analysis, posing a privacy leakage risk (such as the leak of millions of users' voice records due to an unencrypted cloud database in a certain brand's in-vehicle system in 202X). This system adopts a "local priority" strategy: voice commands are converted to text in real time through the in-vehicle ASR module (without storing the raw audio), and differential privacy noise (ε=0) is added to the driving behavior log (such as the number of sudden brakings and frequently used navigation destinations).5. Meeting the ε-differential privacy standard, only anonymized intent tags (such as "need rest" or "preferred music genre") are encrypted before being transmitted to the vehicle manufacturer's server for model optimization; sensitive data (such as facial images and heart rate data) are stored only in the vehicle's security chip (compliant with EAL5+ certification), ensuring that the original information cannot be extracted even if the vehicle is physically disassembled. Data interaction between the system and the cloud / mobile app is encrypted using the national cryptographic SM4 algorithm (with dynamically updated keys) to prevent man-in-the-middle attacks (such as hackers spoofing in-vehicle hotspots to steal data); users can choose the "local analysis only" mode (disabling all data uploads) through the vehicle's settings. This invention offers flexible cross-vehicle adaptation and covers a wide range of applications. Covering over 90% of the gasoline vehicle market, the AI functions of traditional high-end models rely on sensors from specific brands (e.g., a certain brand's infrared camera is only compatible with its own models). This system, however, uses cross-platform compatible middleware to abstract hardware interfaces (such as a unified microphone driver protocol and camera resolution adaptation algorithm), supporting sensors from mainstream models (including low-cost infrared modules and basic millimeter-wave radar). It only requires automakers to provide a basic CAN bus protocol for access. For low-end models (without millimeter-wave radar or with computing power <4 TOPS), the system automatically downgrades to basic functions (e.g., only monitoring eye status via infrared camera + determining posture via seat sensors), ensuring core interactions (navigation) are maintained. Controls and air conditioning adjustments remain unaffected, covering the entire price range from 50,000 RMB economy cars to 500,000 RMB luxury SUVs. Automakers can push hardware adaptation patches (such as adding driver support for a specific microphone model) and model optimization increments (such as specialized training for poor voice recognition when drivers wear gloves in winter) via the server. Users can obtain the latest features without visiting a store for upgrades (actual testing shows the feature iteration cycle has been shortened from 6 months in traditional systems to 2 weeks). This invention features human-machine collaborative safety redundancy, avoiding over-reliance on technology. All system suggestions (such as "You have been driving continuously for 2 hours, we suggest you take a break") are non-mandatory prompts, which drivers can respond to via voice ("Not needed for now") or physical buttons. The AI decision can be manually overridden via either the "Ignore Reminder" button on the center console or a gesture (double-tapping the steering wheel). The system records data from each manual intervention (such as the frequency of ignoring rest reminders and the time spent manually turning off fatigue warnings) and dynamically adjusts subsequent strategies (e.g., increasing visual warning intensity for drivers who frequently ignore reminders, such as a red flashing on the instrument panel). When the system detects severe driver fatigue (such as a sudden drop in heart rate and unconscious head nodding) and failure to respond to reminders, it automatically triggers Lane Keeping Assist (LKA) to limit vehicle drift and gradually reduces the speed to a safe range (e.g., 60 km / h). Simultaneously, it sends location information to preset emergency contacts via the vehicle network to prevent accidents.
[0037] Hardware configuration: Multimodal perception module: 6-microphone directional array (above the windshield, supporting ±30° sound source localization), infrared camera (integrated in the rearview mirror, resolution 640×480, supporting near-infrared spectral analysis), 77GHz millimeter-wave radar (embedded in the steering wheel pillar, detection distance 0.5-3m, accuracy ±2mm), seat pressure sensor (16 pressure sampling points built into the leather seat, sampling frequency 10Hz), CAN bus interface (acquiring real-time data such as vehicle speed, accelerator depth, braking force, and steering angle).
[0038] Computing units: vehicle-mounted GPU (8 TOPS computing power, supports TensorRT acceleration), edge computing chip (runs federated learning models locally, 2 TOPS computing power).
[0039] Voice interaction scenarios: The driver, traveling at 120 km / h on the highway, says, "Turn the air conditioning to 22 degrees" (background noise approximately 80 dB, including children's conversations in the back). A directional microphone array uses beamforming technology to focus on the driver's direction (suppressing rear noise by 25 dB), achieving an initial voice recognition confidence level of 92% → direct execution. If a sudden tire noise causes the confidence level to drop to 80%, the system calls upon an infrared camera to analyze the driver's lip movement trajectory (the frequency and direction of lip movements corresponding to "air conditioning"), and combines this with a seat pressure sensor to determine that the driver has not turned their head (eliminating interference from rear passenger commands), ultimately confirming the target as "air conditioning adjustment."
[0040] When the driver gives the vague command "It's too hot here," the system uses multimodal fusion to infer that the infrared camera detects the driver frequently wiping away sweat (increased facial humidity), the seat pressure sensor shows the body leaning forward (close to the air conditioning vent), and combined with the current vehicle speed and ambient temperature (35℃ in summer), it automatically adjusts the air conditioning temperature from 25℃ to 22℃ and increases the airflow, while displaying "The air conditioning temperature has been lowered for you" on the instrument panel.
[0041] Fatigue monitoring scenarios: After driving for 1.5 hours, the millimeter-wave radar continuously monitored a 22% decrease in the driver's heart rate variability (HRV) (reflecting sympathetic nerve fatigue), the infrared camera detected a cumulative right eye closure duration of 4 seconds (maximum of 3.5 seconds in a single instance), and the seat pressure sensor showed a 18° shift in posture (body sliding to the right). The dynamic adaptive engine, through comprehensive weight calculation (fatigue index contribution: HRV 40%, eye closure 30%, posture 30%), determined it to be "moderate fatigue" → the instrument panel displayed a yellow warning icon "You may be fatigued, please pay attention to safety," while automatically dimming the interior lights (reducing visual stimulation) and playing soft music (frequency 80-100Hz, promoting relaxation); if driving continues for another 20 minutes and the driver's head drooping angle is >15° without manual intervention, the system forcibly activates LKA (limiting lane departure <30cm) and displays a red warning on the HUD "Please pull over immediately to rest."
[0042] Privacy protection scenarios: The text data of the driver's daily commute route (such as "home → company") and frequently used voice commands (such as "play Jay Chou's songs") are converted locally by ASR and differential privacy noise (random perturbation command frequency < 5%) is added. Only the desensitized "commuting time preferred music type = pop" label is encrypted and uploaded to the car manufacturer's server. Facial image data is only processed in the edge chip inside the vehicle (after extracting fatigue features, the original image is deleted). Even if the vehicle data is leaked, attackers will not be able to reconstruct the driver's identity or behavior trajectory.
[0043] The above description is merely a specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various variations or substitutions within the technical scope disclosed in the present invention, and these should all be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0044] Although embodiments of the invention have been shown and described, various modifications and variations can be made to the invention by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principle of the invention should be included within the scope of the claims of the invention.
Claims
1. A human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence, characterized in that: It includes a multimodal perception module, a dynamic adaptive engine, a privacy enhancement processing module, and cross-platform compatible middleware; The multimodal perception module is used to collect the driver's voice commands, facial expressions, physiological characteristics, and in-vehicle environment data; the dynamic adaptive engine generates personalized interaction models based on a federated learning framework and updates them in real time; the privacy enhancement processing module performs differential privacy protection and localization processing on sensitive data; and the cross-platform compatible middleware adapts to the hardware differences of different vehicle models and standardizes data interfaces.
2. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The multimodal perception module includes: a directional microphone array (supporting noise suppression and sound source localization), an infrared camera (monitoring the driver's eye opening and facial micro-expressions), a millimeter-wave radar (non-contact detection of heart rate variability and respiratory rate), a seat pressure sensor (analyzing sitting posture stability), and a CAN bus interface (acquiring vehicle status data such as vehicle speed and throttle depth).
3. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The dynamic adaptive engine includes: a local model training unit (which encrypts and trains personalized parameters based on the driver's historical interaction data), a global model update unit (which receives lightweight general model increments from the vehicle manufacturer's server), and a real-time decision-making unit (which integrates local and global model outputs to dynamically adjust the interaction strategy).
4. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The privacy enhancement processing module includes: a differential privacy encryption unit (adding controllable noise to voice command text and driving behavior logs), a local data storage unit (sensitive information is stored only in the vehicle encryption chip), and a secure communication unit (encrypting data transmission with the cloud via the SM algorithm).
5. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The cross-platform compatible middleware includes: a hardware abstraction layer (unifying the microphone / camera driver interface for different vehicle models), a function degradation strategy unit (automatically disabling non-core functions based on the vehicle's computing power, such as disabling millimeter-wave radar monitoring in low-end models), and a standardized API interface (allowing automakers to quickly access third-party sensors).
6. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The multimodal perception module fuses voice commands and facial expression data (timestamp synchronization error < 50ms) through a spatiotemporal alignment algorithm. When the voice recognition confidence is lower than the threshold, it automatically calls on facial micro-expressions and physiological features to assist in judging the driver's true intentions.
7. The artificial intelligence-based human-vehicle interaction system for fuel-powered vehicles according to claim 1, characterized in that: The dynamic adaptive engine uses multi-dimensional indicators for fatigue monitoring (including eye closure time > 3 seconds, heart rate variability reduced by 20%, and head droop angle > 15°). When a single indicator is abnormal, an early warning is triggered. When multiple indicators are abnormal in tandem, lane keeping assist or automatic parking is forcibly activated.
8. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The privacy enhancement processing module performs local speech-to-text processing on the driver's voice commands (without uploading the original audio), and only encrypts the text commands and intent tags before transmitting them to the cloud for model optimization.
9. The human-vehicle interaction system for fuel-powered vehicles based on artificial intelligence according to claim 1, characterized in that: The cross-platform compatible middleware supports OTA remote upgrades (by pushing hardware adaptation patches through the vehicle manufacturer's server) and is compatible with both the traditional CAN bus protocol for gasoline vehicles and the emerging Ethernet communication protocol.
10. The artificial intelligence-based human-vehicle interaction system for fuel-powered vehicles according to claim 1, characterized in that: The system also includes a driver status visualization feedback unit (displaying current fatigue level, emotional state, and system suggestions via dashboard or HUD), allowing drivers to manually override AI decisions (such as rejecting rest reminders), and the system records human intervention data and optimizes subsequent strategies.