Intelligent cab multi-modal data acquisition and analysis system and analysis method

By synchronously collecting multimodal data and quantitatively analyzing multidimensional performance indicators, the problems of inaccurate evaluation and low efficiency of intelligent cockpit fusion systems have been solved, achieving efficient and comprehensive system performance evaluation and problem localization.

CN121877407APending Publication Date: 2026-04-17INTELLIGENT CONNECTED TECH OF CAERI CO LTD +2
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
INTELLIGENT CONNECTED TECH OF CAERI CO LTD
Filing Date
2025-12-12
Publication Date
2026-04-17

AI Technical Summary

Technical Problem

Existing technologies for testing intelligent cockpit fusion systems suffer from problems such as data silos, spatiotemporal asynchrony, and limited analytical dimensions, leading to inaccurate and inefficient assessments that fail to fully reflect the overall system performance and safety.

Method used

A multimodal data synchronous acquisition method is adopted, and microsecond-level time synchronization is achieved by using a central synchronization control unit. Combined with multi-source signal cross-validation and event triggering mechanism, structured event data slice packages are generated, and multi-dimensional performance index quantitative analysis and weighted fusion evaluation are performed.

Benefits of technology

It enables high-precision and comprehensive performance evaluation of the cockpit fusion system, improves the accuracy and efficiency of testing, and supports full lifecycle management from R&D to operation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121877407A_ABST
    Figure CN121877407A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of intelligent networked automobile testing, in particular to an intelligent cab multi-modal data acquisition and analysis system and analysis method. The method comprises the following steps: by taking a PTP or PPS hardware clock of a central synchronous control unit as a reference, endowing a unified global timestamp to acquired multi-modal data; pre-defining a cab interaction event set; detecting an event in real time by adopting a multi-source signal cross validation mechanism, and judging that the event is valid when and only when at least two heterogeneous data sources report the same event feature in a set time window; taking a global timestamp T0 of event occurrence as an original point, automatically intercepting all modal data in a [T0-delta t1, T0 + delta t2] time window, generating structured event data slice packets, and calculating performance indexes in parallel for each event data slice packet; and dynamically adjusting the weight of the performance index of each dimension according to the evaluation scene, and outputting the comprehensive efficiency of the cockpit fusion system by adopting a weighted fusion algorithm. According to the technical scheme, the overall efficiency of the cab fusion system can be comprehensively evaluated.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent connected vehicle testing technology, specifically to an intelligent cockpit multimodal data acquisition and analysis system and analysis method. Background Technology

[0002] With the advancement of intelligent and connected technologies in the automotive industry, the cockpit has gradually evolved from a simple mechanical instrument panel into a highly integrated intelligent system that combines infotainment, navigation, vehicle control, and combined driver assistance functions—a cockpit fusion system. This transformation not only greatly enhances driving convenience and comfort but also places higher demands on vehicle safety and overall performance.

[0003] Against this backdrop, how to objectively, fairly, and comprehensively evaluate the performance, safety, and user experience of cockpit fusion systems has become a crucial issue for the industry. The increasing complexity and integration of cockpit fusion systems, involving multiple key areas, directly impacts driving safety, user satisfaction, and the market competitiveness of automotive products. Therefore, accurate and scientific evaluation plays a vital role in ensuring vehicle safety and reliability, and is a key link in promoting the continued healthy development of the automotive industry.

[0004] Currently, some preliminary automated data acquisition solutions have been introduced to facilitate the testing of smart cockpits, as detailed below: Distributed acquisition and simple overlay scheme: This scheme can simultaneously acquire vehicle CAN bus data, record driver video, and in-vehicle audio. However, each acquisition module uses its own independent clock for data acquisition, and after the acquisition is completed, time alignment is only attempted through software synchronization. This software synchronization method has limited accuracy, typically only reaching the order of 10-100 milliseconds, which is insufficient to meet the requirements of high-precision synchronization.

[0005] Post-processing and analysis mode: In this mode, after the test is completed, engineers need to spend a lot of time and effort manually reviewing hours of data and video recordings. They need to locate interaction points based on keyframe data, such as the moment a voice command was issued, and then search for the corresponding vehicle system response in the time-aligned data. This manual approach is not only inefficient but also prone to inaccurate results due to human factors, affecting the accuracy and reliability of the test.

[0006] A simplistic approach to metric calculation: This method only calculates a single, localized performance metric. For example, after identifying a voice command and system response, it only calculates the average response latency for a single interaction. This simplistic approach fails to comprehensively reflect the complex performance of a cockpit fusion system and cannot provide an accurate assessment of the overall system effectiveness.

[0007] In existing technologies, intelligent connected vehicles involve numerous sensors, control units (ECUs), communication modules, etc., resulting in a wide variety of data from complex sources, posing significant challenges to unified management and analysis. Furthermore, the lack of intuitive visualization of test data makes data analysis difficult for engineers, severely impacting testing efficiency and decision support capabilities. In addition, inconsistent data formats between different testing platforms and suppliers prevent data sharing between different systems, making it difficult to mutually recognize test results and limiting industry communication and collaboration. Specifically, the main technical problems are as follows: Data silo problem: Most existing testing solutions collect various types of data independently, such as conducting separate speech recognition rate tests, recording CAN bus data separately, and carrying out separate subjective evaluations of HMI (Human-Machine Interface). These independently collected data are difficult to align precisely in time and space, failing to effectively reflect the collaborative working status between different modules of the system, resulting in a lack of comprehensiveness and accuracy in the overall system performance evaluation.

[0008] Spatiotemporal asynchrony issue: Due to the lack of hardware-level synchronization mechanisms, there are significant errors between video, audio, and bus data in existing testing solutions, with error ranges reaching milliseconds or even seconds. This spatiotemporal asynchrony makes it impossible to accurately analyze the causal relationship between "driver issuing commands" and "vehicle executing actions," thus making it difficult to accurately assess the system's response performance and safety.

[0009] The problem of limited analytical dimensions: Most existing analytical methods focus on whether functions are implemented or only conduct simple subjective ratings, lacking quantitative analysis models that deeply integrate user interaction behavior, system response performance, vehicle dynamic changes, and driver status. This limited analytical dimension cannot comprehensively answer key questions such as "in what driving scenarios, which interaction method is safer and more efficient," and it is difficult to meet the needs of scientifically evaluating the overall performance of the cockpit fusion system. Summary of the Invention

[0010] The purpose of this invention is to propose a multimodal data acquisition and analysis system and method for intelligent cockpits, which can comprehensively evaluate the overall performance of cockpit fusion systems.

[0011] To achieve the above objectives, in a first aspect, the present invention proposes a method for multimodal data acquisition and analysis in an intelligent cockpit, comprising: Multimodal data is acquired synchronously, using the PTP or PPS hardware clock of the central synchronization control unit as a reference, performing microsecond-level time synchronization on all acquisition modules, and assigning a unified global timestamp to the acquired multimodal data; Based on event-triggered data association and slicing, a predefined set of cockpit interaction events is defined; a multi-source signal cross-validation mechanism is used to detect events in real time, and an event is deemed valid only if at least two heterogeneous data sources report the same event characteristics within a set time window; taking the global timestamp T0 of the event as the origin, all modal data within the time window [T0-Δt1,T0+Δt2] are automatically extracted to generate a structured event data slice package; Multidimensional performance index quantitative analysis: For each event data slice package, performance indicators are calculated in parallel. Performance indicators include task completion efficiency, interaction smoothness, system stability, driver attention distraction, vehicle control safety, and target recognition accuracy. The overall performance evaluation of the fusion system dynamically adjusts the weights of performance indicators across various dimensions based on the evaluation scenario, and outputs the overall performance of the cockpit fusion system using a weighted fusion algorithm, as shown in the following formula: Overall score = Σ (value of a certain dimension indicator × weight of that dimension).

[0012] Beneficial effects of the basic solution: This technical solution employs multi-source signal cross-validation to determine the validity of an event. At least two heterogeneous data sources must match the same event characteristics within a set time window to confirm the event's validity. This mechanism effectively filters out false triggers or data noise from a single sensor. For example, if only the voice recognition module recognizes the "turn on the air conditioner" command, but there is no corresponding operation record in the touch interaction data, it is determined to be an invalid event. This avoids analytical conclusions based on erroneous data, significantly improving the reliability of subsequent quantitative analysis and providing a precise data foundation for locating problems in the cockpit system.

[0013] The integrated performance evaluation process supports dynamic adjustment of indicator weights based on the evaluation scenario, adapting to the core requirements of different business scenarios. For example, during the development and testing phase, the weights of system stability and target recognition accuracy can be increased to focus on troubleshooting technical defects; during the road testing and verification phase, the weights of vehicle control safety and driver distraction can be increased to ensure human-machine collaborative safety on actual roads; and in fleet management scenarios, the focus is on task completion efficiency and interaction smoothness to optimize the operational service experience. This dynamic weighting mechanism frees the evaluation system from the limitations of fixed standards, enabling precise matching of evaluation objectives at different stages.

[0014] The parallel computing performance metrics designed for event data slices can simultaneously advance the quantitative analysis of multiple dimensions, significantly shortening the data processing cycle compared to traditional serial computing modes. In scenarios such as intelligent connected vehicle road tests that generate massive amounts of multimodal data, this feature enables rapid parsing of large-scale data, avoiding extended evaluation cycles due to data backlog, and providing efficient technical support for automakers to rapidly iterate and optimize cockpit systems.

[0015] Structured event data slices generated based on event timestamps integrate full-modal data before and after an event into independent data packets. This not only facilitates in-depth review of individual events but also enables end-to-end data traceability. For example, when a cockpit experiences interaction lag, the corresponding event slice can be directly retrieved to simultaneously view system operation data, driver operation data, and environmental perception data at that time. This allows for quick identification of whether the fault is caused by hardware response delays, algorithmic logic vulnerabilities, or external environmental interference, improving the efficiency of problem tracing and rectification.

[0016] This technical solution covers the entire chain from data collection, event correlation, indicator quantification to comprehensive evaluation, and is compatible with multiple scenarios such as development and testing, road testing and verification, and fleet management. It can run through the entire life cycle of intelligent cockpit from R&D to operation.

[0017] As a feasible preferred solution, Δt1 and Δt2 are dynamically adjusted according to the event type, where Δt1∈[0.5 s,2 s] and Δt2∈[1 s,5 s].

[0018] As a feasible preferred solution, the task completion efficiency is obtained by measuring the actual time taken from the issuance of the instruction to the completion of the task; The smoothness of the interaction is obtained by extracting the response delay time series of each sub-operation in the same interaction task and calculating its coefficient of variation as the volatility. The system stability is determined based on CPU / memory usage, application crashes, and communication packet loss information to identify lag or abnormalities. The driver's attention distraction level is calculated as the ratio of the cumulative time the driver's gaze deviates from the road ahead, as identified by the driver's facial camera, to the total duration of the event. The vehicle control safety is determined by analyzing whether the longitudinal / lateral acceleration changes in the vehicle bus data show sudden acceleration or braking that is inconsistent with the driver's operation. The target recognition accuracy is obtained by comparing the environmental perception recognition results with the real scene and calculating the missed recognition rate and false recognition rate.

[0019] As a feasible and preferred approach, the interaction smoothness metric is based on the coefficient of variation of the response delay sequence {ti} of the continuous sub-operation response delay, which measures the response delay volatility. The calculation method is as follows: Extract the sequence by extracting the response delay times of all sub-operations from its event data slices to form the sequence {t1, t2, ..., tn}; calculate the volatility by calculating the coefficient of variation of the sequence as the volatility.

[0020] Volatility = (Standard deviation {ti} / Mean {ti}) × 100% Where {ti} is the system response delay sequence corresponding to each sub-operation within the same interactive task.

[0021] As a feasible and preferred solution, the weighted fusion algorithm adaptively adjusts the process by increasing the weights of vehicle control safety and driver attention distraction to no less than 30% when the evaluation scenario is urban roads and safety is emphasized; and by increasing the weights of task completion efficiency and interaction fluency to no less than 30% when the evaluation scenario is highways and efficiency is emphasized. As a feasible and preferred solution, the comprehensive evaluation of integration effectiveness will also automatically generate a visual analysis report that includes data slice playback, indicator comparison curves, and problem location prompts.

[0022] Secondly, this invention also proposes an intelligent cockpit multimodal data acquisition and analysis system, which utilizes the aforementioned intelligent cockpit multimodal data acquisition and analysis method, wherein the data acquisition subsystem includes: The visual acquisition module is used to simultaneously acquire images of the driver's face, panoramic images of the vehicle interior, and images of the accelerator pedal area; The audio acquisition module is used to acquire in-vehicle voice commands and system feedback sounds in the form of a microphone array; The vehicle data acquisition module is used to collect vehicle speed, steering, braking and ADAS status parameters in real time via the vehicle CAN / CAN-FD bus; The interaction log collection module is used to acquire touch events, application switching, and navigation information via the vehicle's debugging interface or bypass monitoring. The environmental positioning module is used to output vehicle position, speed, acceleration, and attitude information through a high-precision GPS / IMU integrated navigation system. The environmental perception and acquisition module is used to acquire information about targets outside the vehicle through a forward-looking camera, a 128-line main LiDAR, a 32-line blind spot LiDAR, and a millimeter-wave radar.

[0023] As a feasible preferred embodiment, the visual acquisition module includes a driver's face camera, an in-vehicle panoramic camera, and an accelerator pedal camera. The driver's face camera is configured to be installed directly in front of the driver to acquire images for gaze direction and facial expression recognition; the in-vehicle panoramic camera is configured to be installed in the center of the vehicle's ceiling to acquire panoramic images for interactive action recognition; and the accelerator pedal camera is configured to be installed near the accelerator pedal to acquire foot images for driving operation recognition. The microphone array of the audio acquisition module is evenly distributed throughout the vehicle and has sound source orientation and gain adaptation functions. It is used to convert the acquired voice commands into text in real time and record the sound quality and volume characteristics of the system feedback sound.

[0024] As a feasible preferred solution, the vehicle data acquisition module achieves millisecond-level sampling by losslessly accessing the vehicle's CAN / CAN-FD bus, and supports synchronous recording of ADAS status words, brake master cylinder pressure, steering wheel angle, and wheel speed signals; The interaction log collection module adopts a bypass listening method, which obtains touch coordinates, application package name, interface control ID and response timestamp through USB / Ethernet debugging interface without modifying the vehicle system source code.

[0025] As a feasible preferred solution, the environmental positioning module uses a high-precision GPS / IMU integrated navigation system, which is installed at a suitable position on the top of the vehicle to provide the vehicle's position, speed, acceleration, and attitude information; The environmental perception and acquisition module includes a main lidar, a blind spot lidar, and a millimeter-wave radar. The main lidar is configured to be horizontally mounted on the roof of the vehicle to provide point cloud data with an angular resolution of 0.1° within a 360°×30° field of view. The blind spot lidar is configured to be obliquely mounted on the side of the vehicle to supplement the point cloud in the blind spot near the vehicle body. The millimeter-wave radar is configured as a forward-facing mid-range radar to provide target speed and distance information in rain, snow, and fog conditions. Attached Figure Description

[0026] Figure 1 This is a logical diagram of a multimodal data acquisition and analysis method for an intelligent cockpit.

[0027] Figure 2 This is a schematic diagram of the architecture of a multimodal data acquisition and analysis system for an intelligent cockpit. Detailed Implementation

[0028] To make the technical solution and advantages of this application clearer, the technical solution of the present invention will be further described in detail below with reference to the accompanying drawings. It is understood that the specific embodiments described herein are only some embodiments of the present invention, and are only used to explain this application, not to limit it. It should be noted that the technical features or combinations of technical features described in the following embodiments should not be considered isolated; they can be combined with each other to achieve better technical effects. The same reference numerals appearing in the accompanying drawings of the following embodiments represent the same features or components, and can be applied to different embodiments.

[0029] Furthermore, unless otherwise defined, the technical or scientific terms used in this invention description shall have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains.

[0030] The present invention will now be described in further detail with reference to the accompanying drawings.

[0031] Reference Figure 1This disclosure provides a method for acquiring and analyzing multimodal data in an intelligent cockpit, comprising the following steps.

[0032] Step S1, Construction of the integrated data acquisition subsystem, refer to Figure 2 The system includes: a visual acquisition module, an audio acquisition module, a vehicle data acquisition module, an interactive log acquisition module, an environmental positioning module, and an environmental perception acquisition module.

[0033] The visual acquisition module includes a driver's face camera, an in-vehicle panoramic camera, and a camera at the accelerator pedal.

[0034] A driver face camera, preferably installed directly in front of the driver, is used to monitor the driver's gaze direction and facial expression changes, and to analyze the driver's attention concentration during driving, such as determining whether the driver's gaze has been off the road for an extended period of time, thereby calculating the driver's attention distraction index.

[0035] The in-vehicle panoramic camera is preferably placed in the center of the top of the vehicle. It can capture the in-vehicle scene from all angles, including the driver's operation actions and the status of in-vehicle equipment. It is used to observe the interaction between the driver and the in-vehicle interactive devices and to help analyze indicators such as the smoothness of interaction.

[0036] Accelerator pedal camera: Installed near the accelerator pedal, it records the driver's foot operation on the accelerator pedal and analyzes the relationship between the driver's driving operation and the vehicle's power response by combining the vehicle's power data.

[0037] The audio acquisition module employs a high-fidelity microphone array, evenly distributed in suitable locations within the vehicle, to capture the driver's voice commands and system feedback sounds. The microphone array accurately captures the direction and intensity of sound, converts voice commands into text information, and records the sound quality and volume characteristics of system feedback sounds, providing data support for subsequent analysis of the efficiency and accuracy of voice interaction.

[0038] The vehicle data acquisition module connects to the vehicle's CAN / CAN bus and records hundreds of vehicle parameters in real time, including vehicle speed, steering, braking, and ADAS (Advanced Driver Assistance Systems) status. This reflects the dynamic changes of the vehicle during driving and assesses vehicle control safety and system stability. For example, by analyzing changes in braking parameters, the vehicle's braking performance in emergency situations can be determined.

[0039] The interaction log collection module acquires interaction logs such as touch events, application switching, and navigation information through the vehicle's infotainment system's debugging interface or bypass monitoring technology. The interaction logs record the driver's interaction with the vehicle's infotainment system, including information such as the type, time, and object of the operation. For example, it records the time and location when the driver clicks an application icon on the screen, as well as the application's response, providing data for analyzing interaction smoothness and task completion efficiency.

[0040] The environmental positioning module, using a high-precision GPS / IMU integrated navigation system, is installed at a suitable location on the vehicle's roof to provide the vehicle's position, speed, acceleration, and attitude information. GPS can accurately obtain the vehicle's geographical location, while the IMU (Inertial Measurement Unit) can measure the vehicle's acceleration and attitude changes. The combination of the two can accurately describe the vehicle's motion state in three-dimensional space, providing important data for analyzing vehicle control safety and driving scenarios.

[0041] The environmental perception and acquisition module is equipped with a forward-facing camera, a 128-line LiDAR, a 32-line blind-spot LiDAR, and millimeter-wave radar deployed outside the vehicle cabin. The forward-facing camera acquires visual information from the front of the vehicle, identifying road signs, traffic lights, and other targets. The 128-line LiDAR boasts high resolution, accurately measuring the distance and shape of surrounding objects. The 32-line blind-spot LiDAR supplements the perception information from the sides of the vehicle, expanding the sensing range. The millimeter-wave radar can operate in adverse weather conditions, detecting the speed and distance of surrounding objects. These devices work together to provide the vehicle with comprehensive environmental perception data, used to analyze indicators such as target recognition accuracy.

[0042] Step S2: Multimodal data synchronous acquisition. Before the test begins, all acquisition modules synchronize with the time of the central synchronization control unit. The central synchronization control unit uses PTP (Precise Time Protocol) or PPS (Pulse Per Second) signals to achieve hardware synchronization with microsecond-level precision.

[0043] During the data acquisition process, each video frame, each audio clip, each CAN bus message, and each interaction log is tagged with a unified, microsecond-level precision global timestamp. For example, when the driver's facial camera captures an image frame, it simultaneously records the precise time corresponding to that frame; when the high-fidelity microphone array captures an audio clip, it assigns that audio clip a timestamp with the same precision; when the vehicle CAN bus data acquisition unit receives a message, it also records the message's time information; and when the interaction log acquisition module obtains an interaction log, it also includes a precise timestamp.

[0044] Step S3, data association and slicing based on event triggers, includes: Step S31: Define the cockpit interaction events in advance. Define a series of preset cockpit interaction events as trigger conditions. When any event is detected, extract all related data within a specific time window before and after the event from the synchronized data stream to form a complete data slice package.

[0045] Cockpit interaction events include, but are not limited to: Voice command initiation: The driver issues voice commands, such as "navigate to the airport" or "turn on the air conditioning".

[0046] Touchscreen tap: The driver taps an icon or button on the vehicle's infotainment screen to switch applications, select functions, or perform other operations.

[0047] Physical button operation: The driver presses physical buttons on the vehicle, such as volume control buttons, air conditioning control buttons, etc.

[0048] Advanced driver assistance system warning trigger: When the vehicle's advanced driver assistance system detects a potential hazard, it issues a warning message, such as a forward collision warning or lane departure warning.

[0049] Navigation route replanning: Due to road changes or other reasons, the navigation system replans the driving route.

[0050] Step S32, Event Detection Mechanism: Event detection is achieved through multi-source signal cross-validation. Taking the "voice command initiation" event as an example, when the voice activation signal (VAD) energy threshold of the microphone array exceeds the threshold, and within 200 milliseconds, a specific message indicating "voice recognition module activation" appears on the CAN bus, while the vehicle system's interaction log records relevant information about the voice command, this is determined to be a valid "voice command initiation" event. Similar cross-validation methods are used for other events, combining information from multiple data sources for accurate judgment, avoiding false positives and false negatives.

[0051] Step S33 involves data slicing. In the background analysis system, data flow is monitored in real time using rules or algorithms. Once an event is detected, it is used as the time origin T0, and all modal data (video, audio, CAN data, etc.) within the time window from T0 - Δt1 to T0 + Δt2 are automatically extracted to form a structured event data slice package. For example, for the "voice command initiation" event, Δt1 is set to 1 second and Δt2 to 2 seconds, meaning all relevant data from 1 second before the voice command is initiated to 2 seconds after it is initiated is extracted, including driver facial expression video, voice command audio, vehicle status data on the CAN bus, and interaction logs. This data is integrated into a data slice package for subsequent centralized analysis.

[0052] Step S4: Quantitative analysis of multidimensional performance indicators. For each data slice package, calculate quantitative indicators from six dimensions: task completion efficiency, interaction fluency, system stability, driver attention distraction, and vehicle control safety.

[0053] Task completion efficiency: Calculate the total time from issuing the command to completing the task. For example, for the voice command "navigate to the airport", record the time interval from when the driver issues the command to when the navigation system begins planning the route and displaying relevant information. The shorter the time interval, the higher the task completion efficiency.

[0054] Interaction smoothness: This includes the number of errors during the operation, the system's secondary menu hierarchy, and the stability of response latency. The stability of response latency is measured by calculating the response latency volatility. The specific calculation method is as follows: Extracting sequences: For an interactive task containing multiple consecutive sub-operations, extract the response delay times of all sub-operations from its event data slices to form a sequence {t1, t2, ..., tn}. For example, in a touchscreen operation, which includes multiple clicks and swipes, record the time from triggering each operation to the system response to form a response delay sequence.

[0055] Calculate the volatility, and use the coefficient of variation of the sequence as the volatility.

[0056] Volatility = (Standard deviation {ti} / Mean {ti}) × 100% Where {ti} is the system response delay sequence corresponding to each sub-operation within the same interactive task.

[0057] The higher the volatility value, the more unstable the system response and the worse the smoothness of the interaction. For example, if the response delay times of multiple operations vary greatly, the calculated volatility will be higher, indicating that the system's response in that interactive task is unstable, affecting the smoothness of the interaction.

[0058] System stability is assessed by analyzing whether the vehicle's infotainment system experiences lag, application crashes, or packet loss during the incident. This is done by monitoring the system's operational status, such as CPU usage, memory consumption, and network communication status, combined with relevant information from video footage and interaction logs, to determine if any system anomalies have occurred. For example, if video stutters, system logs record application unresponsiveness, and CPU usage remains excessively high, it indicates a stability issue within the system during that time period.

[0059] Driver distraction level. Using images captured by a driver's facial camera, the system identifies the driver's gaze direction and calculates the total time and frequency during a specific event when the driver's gaze leaves the road ahead, representing a percentage of the total event time. For example, if, during a navigation command operation, the driver's gaze leaves the road ahead for a cumulative total of 5 seconds, and the total event duration is 20 seconds, then the driver's distraction level is 25%.

[0060] Vehicle control safety is assessed by analyzing the smoothness of lateral and longitudinal control during interaction, and identifying any abnormal driving behaviors such as sudden acceleration or braking. Combining parameters such as vehicle speed, steering, and braking recorded by the vehicle data acquisition module, with vehicle position and attitude information provided by the environmental positioning module, algorithms analyze the vehicle's trajectory and acceleration changes. For example, if the vehicle experiences a significant change in acceleration within a short period, and this change does not match the driver's commands, abnormal driving behavior can be identified, affecting vehicle control safety.

[0061] Target recognition accuracy is assessed by analyzing the accuracy of target perception acquired by the environmental perception module to identify any missed or misidentified targets. The target information identified by the environmental perception module is compared with the actual scene to statistically analyze the number of correctly identified targets, the number of missed targets, and the number of misidentified targets. For example, in a test route where there are 10 other vehicles, the environmental perception module identifies 8, and one of these is misidentified as a pedestrian. In this case, the target recognition accuracy is 80% (8 / 10), the misidentification rate is 10% (1 / 10), and the missed identification rate is 20% (2 / 10).

[0062] Step S5: Comprehensive evaluation of fusion performance. Based on the multi-dimensional indicators obtained in step S4, a weighted fusion algorithm is used to generate a comprehensive performance score of the cockpit fusion system for a specific interactive task or the overall test scenario, and a visual analysis report is generated.

[0063] The weighting coefficients in the weighted fusion algorithm described in step S5 can be dynamically adjusted according to different evaluation objectives. For example, in urban road scenarios, when focusing on safety evaluation, the weights of "vehicle control safety" and "driver attention distraction" are increased; in highway scenarios, when focusing on efficiency evaluation, the weights of "task completion efficiency" and "interaction fluency" are increased. A comprehensive score is calculated using the weighted fusion algorithm, as shown in the following formula: Overall score = Σ (value of a certain dimension indicator × weight of that dimension); Simultaneously, the system automatically generates a detailed evaluation report including data slice playback, indicator comparison curves, and problem location prompts. The data slice playback function allows evaluators to intuitively view various data at the time of the event, including video, audio, and vehicle status; the indicator comparison curves can compare and display the same dimension indicators under different events or different test scenarios, facilitating the analysis of indicator change trends; the problem location prompt function can point out potential problems and improvement directions of the cockpit fusion system based on the results of quantitative analysis.

[0064] The above content is merely an embodiment of the present invention. Commonly known structures and characteristics of the solutions are not described in detail here. Those skilled in the art are aware of all common technical knowledge in the field prior to the application date or priority date, are aware of all prior art in that field, and have the ability to apply conventional experimental methods prior to that date. Those skilled in the art can improve and implement this solution based on the guidance provided in this application and their own capabilities. Some typical well-known structures or systems should not be obstacles for those skilled in the art to implement this application. It should be noted that those skilled in the art can make several modifications and improvements without departing from the structure of the present invention. These should also be considered within the scope of protection of the present invention, and will not affect the effectiveness of the implementation of the present invention or the practicality of the patent. The scope of protection claimed in this application should be determined by the content of its claims, and the specific embodiments described in the specification can be used to interpret the content of the claims.

Claims

1. A smart cockpit multi-modal data acquisition and analysis method, characterized in that, include: Multimodal data is acquired synchronously, using the PTP or PPS hardware clock of the central synchronization control unit as a reference, performing microsecond-level time synchronization on all acquisition modules, and assigning a unified global timestamp to the acquired multimodal data; Based on event-triggered data association and slicing, a predefined set of cockpit interaction events is defined; a multi-source signal cross-validation mechanism is used to detect events in real time, and an event is deemed valid only if at least two heterogeneous data sources report the same event characteristics within a set time window; taking the global timestamp T0 of the event as the origin, all modal data within the time window [T0-Δt1,T0+Δt2] are automatically extracted to generate a structured event data slice package; Multidimensional performance index quantitative analysis: For each event data slice package, performance indicators are calculated in parallel. Performance indicators include task completion efficiency, interaction smoothness, system stability, driver attention distraction, vehicle control safety, and target recognition accuracy. The overall performance evaluation of the fusion system dynamically adjusts the weights of performance indicators across various dimensions based on the evaluation scenario, and outputs the overall performance of the cockpit fusion system using a weighted fusion algorithm, as shown in the following formula: Overall score = Σ (value of a certain dimension indicator × weight of that dimension).

2. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 1, characterized in that, Δt1 and Δt2 are dynamically adjusted according to the event type, where Δt1∈[0.5 s, 2 s] and Δt2∈[1 s, 5 s].

3. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 1, characterized in that, The task completion efficiency is derived from the actual time taken from the issuance of the instruction to the completion of the task. The smoothness of the interaction is obtained by extracting the response delay time series of each sub-operation in the same interaction task and calculating its coefficient of variation as the volatility. The system stability is determined based on CPU / memory usage, application crashes, and communication packet loss information to identify lag or abnormalities. The driver's attention distraction level is calculated as the ratio of the cumulative time the driver's gaze deviates from the road ahead, as identified by the driver's facial camera, to the total duration of the event. The vehicle control safety is determined by analyzing whether the longitudinal / lateral acceleration changes in the vehicle bus data show sudden acceleration or braking that is inconsistent with the driver's operation. The target recognition accuracy is obtained by comparing the environmental perception recognition results with the real scene and calculating the missed recognition rate and false recognition rate.

4. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 3, characterized in that, The interaction smoothness metric is based on the coefficient of variation of the response delay sequence {ti} of continuous sub-operations, which measures the volatility of the response delay. The calculation method is as follows: Extract the sequence and extract the response delay times of all sub-operations from its event data slice packet to form the sequence {t1,t2, ...,tn}; Calculate the volatility, and use the coefficient of variation of the sequence as the volatility. Volatility = (Standard deviation {ti} / Mean {ti}) × 100% Where {ti} is the system response delay sequence corresponding to each sub-operation within the same interactive task.

5. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 1, characterized in that, The weighted fusion algorithm adaptively adjusts the weights of vehicle control safety and driver attention distraction to no less than 30% when the evaluation scenario is urban road and safety is the primary focus; and when the evaluation scenario is highway and efficiency is the primary focus, the weights of task completion efficiency and interaction fluency to no less than 30%.

6. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 5, characterized in that, The comprehensive evaluation of integration effectiveness will also automatically generate a visual analysis report that includes data slice playback, indicator comparison curves, and problem location prompts.

7. A multimodal data acquisition and analysis system for an intelligent cockpit, characterized in that, The method employs a multimodal data acquisition and analysis method for an intelligent cockpit as described in any one of claims 1-6, wherein the data acquisition subsystem includes: The visual acquisition module is used to simultaneously acquire images of the driver's face, panoramic images of the vehicle interior, and images of the accelerator pedal area; The audio acquisition module is used to acquire in-vehicle voice commands and system feedback sounds in the form of a microphone array; The vehicle data acquisition module is used to collect vehicle speed, steering, braking and ADAS status parameters in real time via the vehicle CAN / CAN-FD bus; The interaction log collection module is used to acquire touch events, application switching, and navigation information via the vehicle's debugging interface or bypass monitoring. The environmental positioning module is used to output vehicle position, speed, acceleration, and attitude information through a high-precision GPS / IMU integrated navigation system. The environmental perception and acquisition module is used to acquire information about targets outside the vehicle through a forward-looking camera, a 128-line main LiDAR, a 32-line blind spot LiDAR, and a millimeter-wave radar.

8. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 5, characterized in that, The visual acquisition module includes a driver's face camera, an in-vehicle panoramic camera, and a camera at the accelerator pedal. The driver's face camera is configured to be installed directly in front of the driver to acquire images for eye direction and facial expression recognition. Inside the car A panoramic camera is configured to be mounted in the center of the vehicle's ceiling to capture panoramic images for interactive action recognition. A camera is configured to be mounted near the accelerator pedal to capture images of the feet for driving operation recognition. The microphone array of the audio acquisition module is evenly distributed throughout the vehicle and has sound source orientation and gain adaptation functions. It is used to convert the acquired voice commands into text in real time and record the sound quality and volume characteristics of the system feedback sound.

9. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 5, characterized in that, The vehicle data acquisition module achieves millisecond-level sampling by losslessly accessing the vehicle's CAN / CAN-FD bus, and supports synchronous recording of ADAS status words, brake master cylinder pressure, steering wheel angle, and wheel speed signals; The interaction log collection module adopts a bypass listening method, which obtains touch coordinates, application package name, interface control ID and response timestamp through USB / Ethernet debugging interface without modifying the vehicle system source code.

10. The method for multimodal data acquisition and analysis in an intelligent cockpit according to claim 5, characterized in that, The environmental positioning module uses a high-precision GPS / IMU integrated navigation system, which is installed at a suitable position on the top of the vehicle to provide the vehicle's position, speed, acceleration, and attitude information; The environmental perception and acquisition module includes a main lidar, a blind spot lidar, and a millimeter-wave radar. The main lidar is configured to be horizontally mounted on the roof of the vehicle to provide point cloud data with an angular resolution of 0.1° within a 360°×30° field of view. The blind spot lidar is configured to be obliquely mounted on the side of the vehicle to supplement the point cloud in the blind spot near the vehicle body. The millimeter-wave radar is configured as a forward-facing mid-range radar to provide target speed and distance information in rain, snow, and fog conditions.