Intelligent cockpit voice interaction test platform and method based on hardware simulation
By constructing a hardware simulation testing platform that integrates a driving simulator, a voice interaction box, and simulation software, the problems of environmental simulation and automated evaluation in intelligent cockpit voice interaction testing were solved, enabling efficient and objective testing and optimization suggestions.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHANGHAI VANABILI INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-02-06
- Publication Date
- 2026-05-01
AI Technical Summary
Existing technologies lack simulation of real driving environments in smart cockpit voice interaction testing. The test scenarios are limited, the evaluation dimensions are one-sided, the degree of automation is low, and it is difficult to assess the impact on driving safety, resulting in low testing efficiency and insufficiently objective conclusions.
This paper presents a hardware simulation-based intelligent cockpit voice interaction testing platform, which integrates driving simulator hardware, voice interaction testing box, simulation software module, data acquisition and processing module, and evaluation and report generation module to build a high-fidelity, repeatable testing environment that supports multi-dimensional quantitative evaluation and automated testing.
It enables comprehensive and quantitative evaluation in complex driving scenarios, improves testing efficiency and consistency, objectively assesses the impact of voice interaction on driving safety, and provides reliable data support and optimization directions.
Smart Images

Figure CN121963702A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of automotive electronics testing technology and artificial intelligence, specifically to a hardware simulation-based intelligent cockpit voice interaction testing platform and method, applicable in a laboratory environment for comprehensive, quantitative, and repeatable testing and evaluation of intelligent cockpit voice wake-up, recognition, semantic understanding, and interaction fluency. Background Technology
[0002] With the popularization of intelligent connected vehicles, voice interaction has become one of the core human-machine interfaces of intelligent cockpits, and its performance is directly related to user experience and driving safety. However, the development and testing of voice interaction face severe challenges: real-vehicle road testing is costly, time-consuming, and has poor repeatability, and it is difficult to cover dangerous or extreme scenarios; traditional laboratory pure software testing cannot reproduce real vehicle vibration, background noise (such as wind noise and road noise), driver status, and complex driving environments with multiple concurrent tasks.
[0003] Hardware simulation technology provides an efficient means for control testing, but its application in intelligent cockpit interaction, especially voice interaction testing, is still immature. Existing solutions often lack specificity: the test scenarios are limited and fail to integrate complex traffic events and noise models; the evaluation dimensions are one-sided, focusing mainly on recognition accuracy while neglecting the assessment of the impact on driving safety (such as distraction and reaction delay); the testing process relies on manual labor, has a low degree of automation, and data collection is not synchronized, resulting in low analysis efficiency and insufficient objectivity in the conclusions.
[0004] Therefore, there is an urgent need in this field for a dedicated testing platform and method that can simulate real driving environments, support multi-dimensional quantitative evaluation, and realize automated testing and data analysis, so as to accelerate the iterative optimization of intelligent cockpit voice interaction and ensure that it has both high intelligence and high safety. Summary of the Invention
[0005] This invention is made to solve the above-mentioned problems, and aims to provide a smart cockpit voice interaction testing platform and method based on hardware simulation.
[0006] This invention provides a hardware simulation-based intelligent cockpit voice interaction testing platform, characterized by comprising: a driving simulator hardware, a voice interaction testing box, a simulation software module, a data acquisition and processing module, and an evaluation and report generation module. Driving simulator hardware is used to provide a driving control environment and human-machine interface that closely resembles a real vehicle; The voice interaction test box is connected to the driving simulator hardware to collect voice signals in the cockpit and run the voice interaction to be tested, realizing voice wake-up, recognition, understanding and synthesis. The simulation software module is used to construct virtual driving scenarios that include roads, traffic events, and noise environments, and to drive the driving simulator hardware while simultaneously collecting vehicle dynamics and state data. The data acquisition and processing module is used to synchronously acquire and record voice interaction process data from the voice interaction test box, as well as vehicle and scene data from the simulation software module. The evaluation and report generation module is connected to the data acquisition and processing module. It is used to automatically analyze and calculate the collected data according to the preset quantitative evaluation index system and generate test reports.
[0007] The intelligent cockpit voice interaction test platform based on hardware simulation provided by the present invention also has the following features: the driving simulator hardware includes: a single-seat driving cockpit, a force feedback steering wheel and pedals, at least one central control display screen, multiple forward-facing displays for providing surround view scenarios, an eye tracker, and an in-vehicle camera.
[0008] The intelligent cockpit voice interaction test platform based on hardware simulation provided by this invention also has the following features: the voice interaction test box integrates an edge computing unit, a multi-microphone array, and a speaker; the edge computing unit is used to deploy an automatic speech recognition model, a large natural language processing model, and a text-to-speech model; the multi-microphone array supports beamforming and noise suppression.
[0009] The intelligent cockpit voice interaction test platform based on hardware simulation provided by this invention also has the following features: the edge computing unit is the NVIDIA Jetson AGX Xavier platform; the automatic speech recognition model is the SenseVoiceSmall model; the large natural language processing model is the Qwen2.5 model; and the text-to-speech model is the edgeTTS or CosyVoice model.
[0010] The intelligent cockpit voice interaction test platform based on hardware simulation provided by this invention also has the following features: the simulation software module is SILAB driving simulation software, which is used to draw a virtual test track and set multiple event trigger points on the track, including the appearance of obstacle vehicles, pedestrians crossing, sharp bends and autonomous driving takeover requests, as well as simulate different background noise environments such as engine noise, wind noise, and vibration noise.
[0011] The intelligent cockpit voice interaction test platform based on hardware simulation provided by this invention also has the following features: the quantitative evaluation index system includes four major categories of indicators: voice wake-up performance, voice recognition performance, intelligent interaction performance, and driving safety impact. Voice wake-up performance metrics include wake-up success rate and wake-up time under noise-free / noisy conditions, as well as false wake-up frequency; Speech recognition performance metrics include interaction success rate under noise-free / noisy conditions, and multilingual support. Intelligent interaction performance metrics include the success rate of multi-command processing, contextual understanding, and fuzzy command processing; Driving safety impact indicators include driver response time, vehicle speed standard deviation, lateral acceleration, and time to collision during voice interaction tasks.
[0012] The intelligent cockpit voice interaction testing platform based on hardware simulation provided by this invention also has the following features: the evaluation and report generation module assigns weights to indicators at all levels and performs weighted calculations to obtain a comprehensive performance score for voice interaction, and based on the indicator analysis results, identifies weak links and proposes optimization suggestions for voice activity detection sensitivity, historical memory function, model algorithm or noise suppression strategy.
[0013] The intelligent cockpit voice interaction test platform based on hardware simulation provided by this invention also has the following features: the data acquisition and processing module can achieve high-precision time synchronization of multimodal data, including voice audio streams, recognized text, interaction logs, vehicle CAN signals, eye-tracking trajectories, and video streams.
[0014] The intelligent cockpit voice interaction test platform based on hardware simulation provided by this invention also includes: an automated test control module, used to load a predefined test case sequence, control the simulation software module to automatically run different virtual driving scenarios, and guide or simulate the driver to execute preset voice interaction scripts, thereby achieving fully automated testing.
[0015] This invention also provides a hardware simulation-based intelligent cockpit voice interaction testing method, applicable to the testing platform of any one of claims 1 to 9, characterized by the following steps: S1: Set up the test platform, connect and initialize the driving simulator hardware, voice interaction test box and simulation software module; S2: Deploy the voice interaction to be tested in the voice interaction test box, build a virtual driving scenario and set test events in the simulation software module; S3: Start the test. During the simulated driving process, the driver engages in multiple rounds of dialogue with the voice interaction, and the platform simultaneously collects voice interaction data and vehicle dynamic data. S4: Based on the collected data, calculate and analyze according to the quantitative evaluation index system to evaluate the performance of voice interaction in various dimensions; S5: Generates test reports, identifies performance bottlenecks, and provides optimization suggestions.
[0016] The role and effect of invention The intelligent cockpit voice interaction testing platform and method based on hardware simulation disclosed in this invention constructs a closed, controllable, and highly realistic testing environment by integrating high-fidelity driving simulation hardware, programmable virtual driving scenarios, professional edge computing voice processing units, and multimodal synchronous data acquisition. This platform can comprehensively and quantitatively evaluate the performance of voice interaction under complex noise, dynamic traffic events, and multi-tasking driver workloads, and in particular, objectively assess its impact on driving safety. The fully automated and repeatable testing capabilities significantly improve testing efficiency and consistency, providing reliable data support and clear improvement directions for the rapid iterative optimization of voice interaction. Attached Figure Description
[0017] Figure 1 This is a schematic diagram of the overall architecture of the testing platform of this invention; Figure 2 This is a schematic diagram illustrating the analysis curves of vehicle speed versus time to collision (TTC) over time in a typical test. Detailed Implementation
[0018] To make the technical means, creative features, objectives and effects of this invention easier to understand, the following embodiments are described in detail with reference to the accompanying drawings. Example
[0019] The present invention aims to build a high-fidelity, repeatable, and automated virtual testing environment that can comprehensively evaluate the wake-up, recognition, and understanding performance of voice interaction in complex driving scenarios and its impact on driving safety, thereby providing objective and quantitative data support for design, optimization, and acceptance.
[0020] This embodiment provides a hardware simulation-based intelligent cockpit voice interaction test platform, including: driving simulator hardware, voice interaction test box, simulation software module, data acquisition and processing module, and evaluation and report generation module.
[0021] In this embodiment, the driving simulator hardware serves as the platform's physical human-machine interface, providing an immersive driving experience. It includes a single-seat driver's cockpit equipped with a high-precision force feedback steering wheel and pedals, a central control screen displaying vehicle instrument and infotainment information, a multi-display system providing views of the road ahead and the surrounding environment, an eye tracker for tracking driver attention, and an internal camera monitoring the driver's state. This creates a near-realistic physical interaction environment for testing.
[0022] In this embodiment, the voice interaction test box is the core computing and interaction node of the platform. It integrates a high-performance edge computing unit (such as the NVIDIA Jetson AGX Xavier), a multi-microphone array (supporting beamforming for directional sound pickup and background noise suppression), and a speaker. On the edge computing unit of this box, a complete voice interaction software stack under test is deployed, including: Automatic speech recognition module: Employs lightweight models such as SenseVoiceSmall to convert collected speech into text in real time.
[0023] Natural Language Processing Module: Integrates large language models such as Qwen2.5, and is responsible for semantic understanding, dialogue management, and response generation.
[0024] Text-to-speech module: Using models such as edgeTTS or CosyVoice, the response is synthesized into natural speech and played through a speaker.
[0025] The box connects to the driving simulator host via a high-speed data interface (such as USB or Ethernet) to ensure real-time linkage between voice interaction and driving scenarios.
[0026] In this embodiment, the simulation software module uses professional driving simulation software (such as SILAB) to construct a virtual test world. This module is used to: draw a virtual test track containing various lanes, slopes, and curves; precisely set dynamic traffic event trigger points on the track, such as a vehicle suddenly cutting in front, a pedestrian crossing the road, a sharp curve warning, and an autonomous driving takeover request; simulate different acoustic environments, such as injecting engine noise, wind noise, and road vibration noise with different signal-to-noise ratios, to test noise robustness; and calculate the vehicle dynamics model in real time, providing data such as vehicle speed, acceleration, yaw rate, and collision time with virtual obstacles, and publish this data via a network for other modules to subscribe to.
[0027] In this embodiment, the data acquisition and processing module is responsible for the synchronous acquisition and preprocessing of multimodal data throughout the entire testing process. This module records data in parallel through a time synchronization mechanism.
[0028] Voice interaction data: raw audio, ASR-recognized text, NLP understanding results, TTS-synthesized audio, wake-up timestamps for each interaction, and response latency.
[0029] Vehicle and scene data: Vehicle CAN bus data (speed, acceleration, steering wheel angle, etc.) from simulation software, event trigger flags, and virtual scene video streams.
[0030] Driver behavior data includes gaze trajectory data collected by an eye tracker and facial video captured by a camera. All data is timestamped and stored in a structured database for subsequent analysis.
[0031] In this embodiment, the evaluation and report generation module is the core analysis unit of this platform. It incorporates a meticulously designed quantitative evaluation index system, which consists of four primary indicators: Voice wake-up performance: Evaluate the sensitivity and reliability of wake-up, including wake-up success rate in no / noise conditions, average wake-up time, and false wake-up frequency.
[0032] Speech recognition performance: Evaluate the ability to understand instructions in different noise environments, including the success rate of interaction in no noise and with noise, and the support for multiple languages.
[0033] Intelligent interaction performance: The assessed "intelligence" includes the ability to process multiple consecutive instructions, understand contextual relationships, and cope with ambiguous or non-standard expressions.
[0034] Impact on driving safety: Assess the degree of interference of voice interaction with driving behavior. Key indicators include driver response delay when performing voice tasks, vehicle lateral control stability (such as lateral acceleration standard deviation), and potential risk indicators (such as minimum collision time TTC).
[0035] Specifically, the primary indicator is voice interaction; the secondary indicators and their weights are: voice wake-up 30%, voice recognition 40%, and intelligent interaction 30%; the tertiary indicators and their weights are: noiseless wake-up 30%, noisy wake-up 40%, false wake-up 30%, noiseless interaction 35%, noisy interaction 45%, multilingual recognition 20%, multi-command interaction 30%, contextual understanding 30%, and fuzzy command processing 40%; and the quaternary indicators and their weights are: wake-up time 50%, wake-up success rate 50%, wake-up time 50%, wake-up success rate 50%, false wake-up frequency 100%, interaction success rate 100%, interaction success rate 100%, variety richness 30%, interaction success rate 70%, interaction success rate 100%, interaction success rate 100%.
[0036] This module automatically extracts data from the database, calculates scores for various indicators, and performs weighted summation based on preset weights to arrive at an overall score. Finally, it automatically generates a visually appealing test report that not only presents the results but also, based on data comparison and analysis, identifies weaknesses in specific scenarios or indicators and proposes concrete optimization suggestions such as "adjusting the VAD threshold," "optimizing dialogue history management," and "enhancing the noise suppression model."
[0037] In this embodiment, the invention also includes a data preprocessing and fusion unit. This unit is located after the distributed sensing module and is responsible for timestamping, filtering and denoising (such as bandpass filtering of ECG signals and denoising of audio signals) and removing outliers from the raw multimodal data collected by each sensor. It also fuses the relevant data at the feature level (such as mapping the eye-tracking gaze coordinates to the visual image coordinate system) to provide a unified, high-quality, and semantically rich input data stream for subsequent modules.
[0038] This embodiment also provides a hardware simulation-based intelligent cockpit voice interaction testing method, applied to the aforementioned testing platform, including the following steps: S1: Set up the test platform, connect and initialize the driving simulator hardware, voice interaction test box and simulation software module; S2: Deploy the voice interaction to be tested in the voice interaction test box, build a virtual driving scenario and set test events in the simulation software module; S3: Start the test. During the simulated driving process, the driver engages in multiple rounds of dialogue with the voice interaction, and the platform simultaneously collects voice interaction data and vehicle dynamic data. S4: Based on the collected data, calculate and analyze according to the quantitative evaluation index system to evaluate the performance of voice interaction in various dimensions; S5: Generates test reports, identifies performance bottlenecks, and provides optimization suggestions.
[0039] The role and effect of the embodiments The intelligent cockpit voice interaction testing platform and method based on hardware simulation disclosed in this invention constructs a closed, controllable, and highly realistic testing environment by integrating high-fidelity driving simulation hardware, programmable virtual driving scenarios, professional edge computing voice processing units, and multimodal synchronous data acquisition. This platform can comprehensively and quantitatively evaluate the performance of voice interaction under complex noise, dynamic traffic events, and multi-tasking driver workloads, and in particular, objectively assess its impact on driving safety. The fully automated and repeatable testing capabilities significantly improve testing efficiency and consistency, providing reliable data support and clear improvement directions for the rapid iterative optimization of voice interaction.
[0040] The above embodiments are preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention.
Claims
1. A hardware simulation-based intelligent cockpit voice interaction test platform, characterized in that, include: The system includes driving simulator hardware, a voice interaction test box, simulation software modules, data acquisition and processing modules, and evaluation and report generation modules. The driving simulator hardware is used to provide a driving control environment and human-machine interface that closely resembles a real vehicle. The voice interaction test box is communicatively connected to the driving simulator hardware, used to collect voice signals in the cockpit and run the voice interaction to be tested, realizing voice wake-up, recognition, understanding and synthesis; The simulation software module is used to construct a virtual driving scenario that includes roads, traffic events, and noise environments, and to drive the driving simulator hardware while simultaneously collecting vehicle dynamics and state data. The data acquisition and processing module is used to synchronously acquire and record voice interaction process data from the voice interaction test box, as well as vehicle and scene data from the simulation software module. The evaluation and report generation module is connected to the data acquisition and processing module and is used to automatically analyze and calculate the acquired data according to a preset quantitative evaluation index system to generate a test report.
2. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 1, characterized in that: in, The driving simulator hardware includes: a single-seat driver's cockpit, a force feedback steering wheel and pedals, at least one central control display screen, multiple forward-facing displays for providing surround-view scenarios, an eye tracker, and an in-vehicle camera.
3. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 1, characterized in that: in, The voice interaction test box integrates an edge computing unit, a multi-microphone array, and a speaker; the edge computing unit is used to deploy automatic speech recognition models, large natural language processing models, and text-to-speech models; the multi-microphone array supports beamforming and noise suppression.
4. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 3, characterized in that: in, The edge computing unit is an NVIDIA Jetson AGX Xavier platform; the automatic speech recognition model is the SenseVoiceSmall model; the large natural language processing model is the Qwen2.5 model; and the text-to-speech model is either edgeTTS or CosyVoice.
5. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 1, characterized in that: in, The simulation software module is SILAB driving simulation software, which is used to draw a virtual test track and set multiple event trigger points on the track, including the appearance of obstacle vehicles, pedestrians crossing, sharp bends and autonomous driving takeover requests, as well as simulate different background noise environments such as engine noise, wind noise, and vibration noise.
6. The intelligent cockpit voice interaction test platform based on hardware simulation as described in claim 1, Its features are: The quantitative evaluation index system includes four categories of indicators: voice wake-up performance, voice recognition performance, intelligent interaction performance, and driving safety impact. The voice wake-up performance indicators include the wake-up success rate and wake-up time under noise-free / noisy conditions, as well as the frequency of false wake-ups; The speech recognition performance metrics include the interaction success rate under noise-free / noisy conditions, and multilingual support. The intelligent interaction performance indicators include the success rates of multi-command processing, contextual understanding, and fuzzy command processing. The driving safety impact indicators include driver response time, vehicle speed standard deviation, lateral acceleration, and collision time during voice interaction tasks.
7. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 6, characterized in that: in, The evaluation and report generation module assigns weights to indicators at each level and performs weighted calculations to obtain a comprehensive performance score for voice interaction. Based on the indicator analysis results, it identifies weak links and proposes optimization suggestions for voice activity detection sensitivity, historical memory function, model algorithm, or noise suppression strategy.
8. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 1, characterized in that: in, The data acquisition and processing module can achieve high-precision time synchronization of multimodal data, including voice audio streams, recognized text, interactive logs, vehicle CAN signals, eye-tracking trajectories, and video streams.
9. The intelligent cockpit voice interaction test platform based on hardware simulation according to claim 1, characterized in that, Also includes: The automated test control module is used to load predefined test case sequences, control the simulation software module to automatically run different virtual driving scenarios, and guide or simulate the driver to execute preset voice interaction scripts, thereby achieving fully automated testing.
10. A hardware simulation-based intelligent cockpit voice interaction testing method, applicable to the testing platform described in any one of claims 1 to 9, characterized in that, Includes the following steps: S1: Set up the test platform, connect and initialize the driving simulator hardware, voice interaction test box and simulation software module; S2: Deploy the voice interaction to be tested in the voice interaction test box, construct a virtual driving scenario and set test events in the simulation software module; S3: Start the test. During the simulated driving process, the driver engages in multiple rounds of dialogue with the voice interaction, and the platform simultaneously collects voice interaction data and vehicle dynamic data. S4: Based on the collected data, calculate and analyze according to the quantitative evaluation index system to evaluate the performance of voice interaction in various dimensions; S5: Generates test reports, identifies performance bottlenecks, and provides optimization suggestions.