Intelligent cabin multi-mode user interaction behavior fusion analysis and evaluation method

By combining timestamp synchronization and the fusion of infrared sensing and millimeter-wave radar, along with deep learning and reinforcement learning, the problem of asynchronous data in multimodal interaction analysis was solved. This enabled accurate fusion analysis and real-time evaluation of multimodal user behavior in the intelligent cockpit, improving the accuracy of user status recognition and the efficiency of interaction optimization.

CN121765627APending Publication Date: 2026-03-31TSINGHUA UNIVERSITY
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-31

AI Technical Summary

Technical Problem

In existing multimodal interaction analysis methods, the independent data collection from heterogeneous devices lacks a unified synchronization mechanism, resulting in time stamp offsets between eye movement trajectories and gestures, failure of feature correlation between facial expressions and voice signals, affecting the accuracy of user state recognition, and lacking a systematic evaluation framework, making it difficult to quantify the effect of interface layout optimization and failing to guarantee the robustness of data collection.

Method used

The eye-tracking data and facial expression features captured by the 3D camera are synchronized across devices using a timestamp synchronization mechanism. The finger movement path is tracked with sub-millimeter precision using a fusion scheme of infrared sensing and millimeter-wave radar. Cross-modal feature fusion processing is performed based on deep learning algorithms to generate a user state model. The cockpit interface layout and interaction path are adjusted in real time through a reinforcement learning mechanism. The optimization effect is verified using a closed-loop testing system.

Benefits of technology

It enables precise fusion analysis and real-time effectiveness evaluation of multimodal user interaction behaviors in intelligent cockpits, improves the accuracy of user status recognition and the closed-loop efficiency of interaction optimization, and enhances the system's adaptability and intelligence level.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121765627A_ABST
    Figure CN121765627A_ABST
Patent Text Reader

Abstract

The invention provides an intelligent cabin multi-mode user interaction behavior fusion analysis and evaluation method. According to the multi-modal user interaction behavior fusion analysis and evaluation method provided by the embodiment of the invention, accurate fusion analysis and real-time effectiveness evaluation of user multi-modal interaction behaviors (visual sense, auditory sense and tactile sense) in an intelligent cabin can be realized, and the accuracy of user state recognition and the reliability of system response are remarkably improved; and interactive design optimization and intelligent driving decision improvement are supported.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing technology, and in particular to a method for the fusion analysis and evaluation of multimodal user interaction behavior in intelligent cockpits. Background Technology

[0002] As a crucial carrier of automotive intelligence, the intelligent cockpit is widely used in human-machine interaction optimization and driving safety enhancement. With the development of in-vehicle perception technology, multimodal interaction systems are gradually evolving from traditional single-modality (such as voice or touch) to a fusion of vision, hearing, and touch. Related technologies utilize eye tracking, 3D cameras, voice acquisition, and motion capture to construct a basic framework for user state perception. Specifically, this technology system covers the entire process from data acquisition to behavior modeling, including key aspects such as micro-expression recognition, gesture trajectory analysis, physiological signal monitoring, and voice interaction evaluation. Among these, deep learning-based facial expression recognition algorithms have achieved an accuracy rate of over 85% on publicly available datasets, while 3D camera technology significantly improves the accuracy of gesture recognition, providing crucial support for intelligent cockpit interaction design.

[0003] However, existing multimodal interaction analysis methods directly collect data independently from heterogeneous devices without establishing a unified synchronization mechanism. This may lead to timestamp offsets between eye-tracking trajectories and gestures, or the failure of feature correlation between facial expressions and speech signals, thus affecting the accuracy of user state recognition. Specifically, traditional systems typically rely on a single modality (such as facial expressions alone) for analysis, but this has limitations such as missing multi-dimensional state features and incomplete behavioral modeling, making it difficult to fully reflect the real interaction needs of users in complex driving scenarios. Therefore, existing technologies lack a systematic evaluation framework, failing to quantify the effects of interface layout optimization and struggling to ensure the robustness of data collection in in-vehicle environments with noise interference and lighting changes. Ultimately, this results in delayed adjustments to interaction strategies, hindering the collaborative optimization capabilities of intelligent cockpits and autonomous driving systems. Summary of the Invention

[0004] The present invention aims to at least partially solve one of the technical problems in the related art.

[0005] Therefore, the first objective of this invention is to propose a method for the fusion analysis and evaluation of multimodal user interaction behavior in intelligent cockpits.

[0006] The second objective of this invention is to propose a device for the fusion analysis and evaluation of multimodal user interaction behavior in intelligent cockpits.

[0007] The third objective of this invention is to provide an electronic device.

[0008] The fourth objective of this invention is to provide a computer-readable storage medium.

[0009] The fifth objective of this invention is to provide a computer program product.

[0010] To achieve the above objectives, a first aspect of the present invention proposes a method for the fusion analysis and evaluation of multimodal user interaction behavior in an intelligent cockpit, comprising: S1, synchronously collecting multimodal interaction data of intelligent cockpit users, the data including eye movement trajectories, finger movement paths, physiological and motion signals, facial expression features, and voice interaction signals; S2, performing cross-modal feature fusion processing on the multimodal data based on a deep learning algorithm to generate a user state model, the model including quantitative assessment results of fatigue level, emotional fluctuations, and cognitive load; S3, constructing a dynamic evaluation strategy based on the user state model, and adjusting the cockpit interface layout, interaction path, and voice response parameters in real time through a reinforcement learning mechanism; S4, verifying the optimization effect using a closed-loop testing system, the closed-loop testing system integrating a simulated mouth, a noise interference player, and an online semantic recognition module, and iteratively updating the user state model based on feedback from real vehicle test data.

[0011] In one embodiment of the present invention, the synchronous acquisition of multimodal interaction data of smart cockpit users further includes: aligning eye-tracking trajectory data and facial expression features acquired by 3D cameras across devices using a timestamp synchronization mechanism; and using a fusion scheme of infrared sensing and millimeter-wave radar to track finger movement paths with sub-millimeter precision.

[0012] In one embodiment of the present invention, the cross-modal feature fusion processing of the multimodal data based on the deep learning algorithm further includes: constructing a multi-task learning framework to simultaneously process feature vectors of micro-expression recognition, gesture analysis, and voice interaction evaluation; and using an attention mechanism to perform weighted fusion of eye-tracking features and facial expression features to improve the accuracy of fatigue state recognition.

[0013] In one embodiment of the present invention, the verification of the optimization effect using a closed-loop testing system further includes: simulating voice input at different sound intensity levels using a simulated mouth to verify the wake-up rate and semantic recognition robustness of the voice interaction system; and testing the response stability of the voice interaction system under complex acoustic conditions by injecting in-vehicle environmental noise using an interference sound player.

[0014] In one embodiment of the present invention, the method further includes: constructing a personalized user profile based on the user's historical interaction data; adjusting the weight parameters in the dynamic evaluation strategy according to the profile; and transmitting the evaluation results of the user state model to the intelligent driving decision module in real time to optimize the path planning and interaction prompt logic of the autonomous driving system.

[0015] To achieve the above objectives, a second aspect of the present invention provides a device for the fusion analysis and evaluation of multimodal user interaction behavior in an intelligent cockpit, comprising: The multimodal interaction data synchronous acquisition module is used to synchronously acquire multimodal interaction data of the smart cockpit user. The data includes eye movement trajectory, finger movement path, physiological and motion signals, facial expression features and voice interaction signals. The cross-modal feature fusion processing module is used to perform cross-modal feature fusion processing on the multimodal data based on deep learning algorithms to generate a user state model, which includes quantitative assessment results of fatigue level, emotional fluctuations and cognitive load status. The dynamic evaluation strategy construction module is used to construct a dynamic evaluation strategy based on the user state model and adjust the cockpit interface layout, interaction path and voice response parameters in real time through a reinforcement learning mechanism. The closed-loop test verification module is used to verify the optimization effect using a closed-loop test system. The closed-loop test system integrates a simulated mouth, a noise interference player, and an online semantic recognition module, and iteratively updates the user state model based on feedback from real vehicle test data.

[0016] In one embodiment of the present invention, the multimodal interactive data synchronization acquisition module is further configured to: perform cross-device temporal alignment of eye-tracking trajectory data and facial expression features acquired by a 3D camera through a timestamp synchronization mechanism; and perform sub-millimeter-level precision tracking of finger movement paths using a fusion scheme of infrared sensing and millimeter-wave radar.

[0017] To achieve the above objectives, a third aspect of the present invention provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of the first aspects.

[0018] To achieve the above objectives, a fourth aspect of the present invention provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of the first aspects.

[0019] To achieve the above objectives, a fifth aspect of the present invention provides a computer program product that, when executed by a processor, implements the method described in any one of the first aspects.

[0020] The technical solutions provided by the embodiments of the present invention bring at least the following beneficial effects: they enable accurate fusion analysis and real-time effectiveness evaluation of multimodal user interaction behaviors (visual, auditory, tactile) in intelligent cockpits, thereby improving the accuracy of user status recognition and the closed-loop efficiency of interaction optimization.

[0021] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0022] The above and / or additional aspects and advantages of the present invention will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart illustrating a method for fusion analysis and evaluation of multimodal user interaction behavior in an intelligent cockpit, as provided in an embodiment of the present invention. Figure 2 This is a schematic diagram of the structure of a smart cockpit multimodal user interaction behavior fusion analysis and evaluation device provided in an embodiment of the present invention. Detailed Implementation

[0023] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0024] To address this issue, embodiments of the present invention provide a method for the fusion analysis and evaluation of multimodal user interaction behavior in intelligent cockpits. Figure 1 This is a flowchart illustrating a method for fusion analysis and evaluation of multimodal user interaction behavior in an intelligent cockpit, as provided in an embodiment of the present invention. Figure 1 As shown, the method includes the following steps: S1, synchronously collects multimodal interaction data of the smart cockpit user, including eye movement trajectory, finger movement path, physiological and motion signals, facial expression features and voice interaction signals.

[0025] Specifically, this step involves the simultaneous collection of multimodal interaction data from smart cockpit users, including eye-tracking trajectories, finger movement paths, physiological and motion signals, facial expression features, and voice interaction signals. This step is the core data input stage for building a prototype system for user behavior analysis and evaluation, and its technical implementation relies on multi-sensor fusion and a high-precision synchronous acquisition mechanism.

[0026] In some implementations, eye-tracking data is acquired using embedded eye-tracking devices (such as Tobii Pro or PupilLabs). These devices typically employ infrared light reflection, using binocular cameras to capture reflection points on the user's iris and cornea. Combined with head posture estimation algorithms, this enables three-dimensional localization of the gaze direction. The sampling frequency is generally set between 60Hz and 240Hz to ensure real-time capture of changes in user attention during dynamic driving scenarios. Eye-tracking data output includes gaze coordinates, gaze duration, and saccade paths, used to assess the distribution of user attention to cockpit interface elements. This helps understand the user's attention level to different interface elements and information prompts when operating the intelligent cockpit system, thereby optimizing interface design and placing important information where the user's gaze is easily focused, improving ease of operation.

[0027] Finger movement path acquisition relies on high-frame-rate RGB-D cameras (such as the Intel RealSense D455) or finger-tracking gloves based on inertial measurement units (IMUs). The camera fuses depth information with 2D images to identify finger joint angles, movement trajectories, and contact areas, with sampling frequencies typically ranging from 90Hz to 300Hz. Parameters such as dwell time and path meandering are used to analyze the smoothness of user operations and interface accessibility.

[0028] Various physiological and kinematic signals of the driver can be collected in real time by integrating high-precision sensors, triaxial accelerometers and other detection modules, which can be used to quantitatively assess the driver's fatigue level, emotional fluctuations and cognitive load.

[0029] Facial expression feature acquisition utilizes a high-resolution RGB camera and a 3D structured light sensor to achieve facial key point detection and micro-expression recognition. The advanced 3D camera technology further enhances the ability to acquire user limb position and movement depth information, significantly improving the accuracy of recognizing complex gestures, achieving an accuracy rate of over 90%. Simultaneously, the deep learning-based facial expression recognition algorithm is continuously optimized, achieving an average recognition accuracy rate exceeding 85% on public datasets. It can accurately identify various expressions such as joy, anger, and fatigue, providing strong support for analyzing user emotional states.

[0030] Voice interaction signals are acquired via a high-sensitivity microphone array, supporting multi-channel sound source localization and noise reduction. The system can simultaneously record user voice commands, system responses, and environmental noise, with a sampling rate typically ranging from 16kHz to 48kHz, meeting the accuracy requirements for speech recognition and semantic analysis. Key metrics such as voice wake-up rate and response latency are used to evaluate the performance of the voice interaction system.

[0031] This step, through the synchronous collection of multimodal data, provides a high-fidelity, multi-dimensional data foundation for subsequent user behavior modeling and personalized evaluation. Especially in intelligent driving assistance systems, it enables real-time perception and feedback optimization of user status, improving the safety and comfort of human-computer interaction.

[0032] Furthermore, S1 includes: S11 uses a timestamp synchronization mechanism to perform cross-device temporal alignment of eye-tracking data and facial expression features captured by a 3D camera.

[0033] Specifically, this step involves cross-device temporal alignment of eye-tracking data and facial expression features captured by 3D cameras using a timestamp synchronization mechanism. This is a key technical step in realizing multimodal user interaction behavior analysis in smart cockpits. In some implementations, this mechanism combines hardware trigger signals with software timestamp compensation to ensure high consistency of data from different acquisition devices in the temporal dimension.

[0034] At the technical implementation level, eye-tracking devices and 3D cameras operate at independent sampling frequencies. Typically, eye-tracking devices have sampling rates between 60Hz and 240Hz, while 3D cameras generally have frame rates between 30Hz and 60Hz. To achieve cross-device data alignment, the system introduces a unified hardware trigger signal (such as a GPIO signal or PTP protocol) at the acquisition end, enabling each device to begin data acquisition upon receiving a synchronization pulse. Simultaneously, at the software level, the system appends a high-precision timestamp (e.g., nanosecond level) to each frame of eye-tracking data and facial expression feature data, and employs interpolation algorithms (such as linear interpolation or spline interpolation) to perform time compensation for asynchronous data, thereby achieving millisecond-level timing alignment.

[0035] At the parameter level, the timestamp synchronization error should be controlled within ±5ms to meet the accuracy requirements for timing analysis of interactive behaviors in human factors engineering. Facial expression feature extraction is typically based on open-source tools such as OpenFace or MediaPipe, extracting parameters including facial action units (AUs), head pose angles (pitch / yaw / roll), and facial key point coordinates, while eye-tracking data includes gaze coordinates, pupil diameter, blink frequency, etc. The aligned data will be stored according to a unified timeline to facilitate subsequent multimodal fusion analysis.

[0036] At the application level, this step is widely used in the joint analysis of user attention and emotional state in smart cockpits. For example, in driver fatigue detection scenarios, the system needs to simultaneously analyze abnormal eye movements (such as prolonged fixation) and facial expression features (such as drooping corners of the mouth and furrowed brows) to improve the accuracy of the judgment. In addition, in the optimization of human-computer interaction experience, by aligning the timing of eye movements and facial expressions, the user's emotional response and attention allocation to specific interface elements can be assessed.

[0037] In terms of technical effectiveness, this step effectively solves the timing misalignment problem caused by the asynchronous operation of multimodal data acquisition devices, providing a reliable data foundation for subsequent fusion modeling and behavior analysis, and significantly improving the accuracy and robustness of the system in user state recognition and interactive behavior evaluation.

[0038] The S12 uses a fusion scheme of infrared sensing and millimeter-wave radar to track finger movement paths with sub-millimeter accuracy.

[0039] Specifically, in some implementations, this invention employs a fusion scheme of infrared sensing and millimeter-wave radar to achieve sub-millimeter-level accuracy in tracking the movement path of a user's finger. This scheme, through multi-sensor data fusion technology, combines infrared optical positioning with the non-contact motion sensing capabilities of millimeter-wave radar, effectively overcoming the limitations of single sensors in complex lighting, occlusion, or high-speed motion scenarios, thereby improving the robustness and accuracy of finger trajectory recognition.

[0040] From a technical implementation perspective, infrared sensing modules typically consist of an infrared emitter and a high-resolution infrared camera, used to capture the contours and movement trajectories of fingers under near-infrared light. Millimeter-wave radar, on the other hand, emits electromagnetic waves in the 24GHz or 60GHz band, receives reflected signals, and utilizes the Doppler effect and time-of-flight ranging technology to acquire the finger's position, velocity, and acceleration information in three-dimensional space. In the system, the infrared camera provides high-resolution two-dimensional image data, while the millimeter-wave radar supplements it with depth information and motion vectors. The two are fused through timestamp synchronization and spatial coordinate alignment. The fusion algorithm can employ methods such as Kalman filtering or particle filtering to achieve real-time, high-precision estimation of the finger's movement trajectory.

[0041] In terms of specifications, the infrared camera's frame rate is typically set to above 120Hz to ensure the ability to capture high-speed finger movements; the millimeter-wave radar achieves a ranging accuracy of 1mm, a velocity accuracy of 0.1m / s, and a spatial resolution of 5cm³. The fusion algorithm's output frequency can reach 100Hz, with trajectory tracking errors controlled within ±0.5mm, meeting sub-millimeter accuracy requirements. Furthermore, the system supports stable operation in environments with light intensity below 500 lux or with obstructions, complying with the functional safety standards for in-vehicle human-machine interaction systems in ISO 26262.

[0042] In terms of application scenarios, this technology is widely used in human factors engineering testing and interactive experience optimization in smart cockpits. For example, during operation of the in-vehicle central control screen, the system can accurately record the user's finger's touch path, dwell time, and operation smoothness in different functional areas, providing data support for interface layout optimization and interaction logic design. Simultaneously, this solution can also be used for performance evaluation of gesture control systems, especially in contactless interaction scenarios, such as operational efficiency analysis under a hybrid voice and gesture control mode.

[0043] In terms of technical effectiveness, this fusion tracking solution significantly improves the accuracy and stability of finger movement recognition, providing a reliable data foundation for multimodal user behavior analysis. Its high precision and low latency contribute to a more natural and intelligent cockpit interaction experience, enhancing system responsiveness and user satisfaction, and possessing significant engineering application value and innovative significance.

[0044] S2, based on deep learning algorithms, perform cross-modal feature fusion processing on the multimodal data to generate a user state model, which includes quantitative assessment results of fatigue level, emotional fluctuations, and cognitive load status.

[0045] Specifically, this step involves performing cross-modal feature fusion processing on multimodal data based on deep learning algorithms to generate a user state model that can quantitatively assess a user's fatigue level, mood swings, and cognitive load. In some implementations, this process employs multimodal neural network architectures, such as multimodal Transformers, cross-modal attention mechanisms, or fusion-type convolutional-recurrent neural networks (ConvLSTM), to achieve the alignment and fusion of visual, auditory, and tactile signals in the feature space.

[0046] At the technical implementation level, the data from modalities such as eye tracking, facial expression recognition, finger trajectory tracking, and dynamic signal capture are first preprocessed, including normalization, temporal alignment, and noise filtering. For example, eye tracking data is smoothed using Kalman filtering, facial expression data has facial key points and action units extracted using OpenFace or MediaPipe, and speech signals have acoustic features extracted using MFCC or Mel-spectrogram. Subsequently, the features from each modality are input into a shared feature fusion layer, which can employ multimodal feature concatenation, tensor fusion, or cross-modal attention mechanisms to achieve semantic association and complementary enhancement between features.

[0047] Regarding parameters, model training typically employs the Adam optimizer with a learning rate between 1e-4 and 5e-4, a batch size of 32 or 64, and 50-100 training epochs. The loss function can combine cross-entropy loss and mean squared error (MSE) to simultaneously optimize classification and regression tasks. Model evaluation metrics include accuracy, F1-score, AUC-ROC curve, and mean absolute error (MAE). Specifically, the MAE for fatigue assessment should be below 0.15, the accuracy for emotion recognition should reach over 85%, and the F1-score for cognitive load classification should be no less than 0.82.

[0048] In practical applications, this step can be deployed on an in-vehicle edge computing platform or cloud server to process multimodal perception data from the smart cockpit in real time, providing user status feedback to the driver assistance system. For example, during long-distance driving, the system can trigger a warning mechanism based on the fatigue index output by the fusion model, or adjust the voice interaction strategy according to emotional fluctuations, thereby improving the intelligence and safety of human-machine interaction.

[0049] The technical effect of this step is that it achieves efficient fusion and modeling of multimodal data through deep learning, which significantly improves the accuracy and timeliness of user status recognition, provides key support for personalized interaction and driving safety decision-making in smart cockpits, and has significant engineering practical value and innovation.

[0050] Furthermore, S2 includes: S21 constructs a multi-task learning framework that simultaneously processes feature vectors from micro-expression recognition, gesture analysis, and voice interaction evaluation.

[0051] Specifically, this step involves constructing a multi-task learning framework for simultaneously processing feature vectors from micro-expression recognition, gesture analysis, and voice interaction evaluation. This is a core technical component in realizing a prototype system for multimodal user interaction behavior analysis and evaluation in intelligent cockpits. In some implementations, this framework employs a hybrid architecture combining a shared underlying feature extraction module with a task-specific header structure to achieve efficient fusion of cross-modal features and multi-task collaborative optimization.

[0052] At the technical implementation level, the system first acquires raw user behavior data through a multimodal sensor acquisition module, including facial video captured by a high-resolution camera, finger and gesture motion trajectories acquired by a three-axis accelerometer and a 3D camera, and speech signals acquired by a microphone array. Subsequently, each modal data is processed through its corresponding feature extraction network: micro-expression recognition uses a 3D-CNN or LSTM structure to extract dynamic facial features; gesture action analysis uses OpenPose or MediaPipe to extract skeletal keypoints and combines them with a spatiotemporal graph convolutional network (ST-GCN) for action recognition; and voice interaction evaluation uses a speech recognition model (such as DeepSpeech or Wav2Vec2) to extract acoustic and semantic features. The extracted feature vectors are fused in a shared multi-task learning backbone network, which is typically based on a Transformer or a multimodal fusion module (such as Cross-Attention) to achieve cross-modal information interaction and alignment.

[0053] In terms of parameters, the dimension of the feature vector is usually controlled between 128 and 512 to balance the model's expressive power and computational efficiency. The frame rate for micro-expression recognition tasks needs to be no less than 30 FPS to ensure the continuity of dynamic expression capture; the key point detection accuracy for gesture action analysis should reach more than 95%, and the action recognition latency should be controlled within 200ms; in voice interaction evaluation, the speech recognition accuracy (WER) should be less than 10%, the wake-up rate should reach more than 98%, and the response latency should not exceed 500ms to meet the real-time requirements in the vehicle environment.

[0054] In application scenarios, this multi-task learning framework is deployed on an in-vehicle edge computing platform, supporting operation in real driving scenarios or simulated cockpit environments. It can simultaneously analyze the driver's facial emotions, gesture intentions, and voice commands, providing multi-dimensional interactive feedback and behavioral evaluation for the smart cockpit. For example, when the driver is in a low mood, the system can automatically adjust the voice interaction strategy or determine their attention state through gesture recognition, thereby optimizing the human-machine interface.

[0055] The technical benefits of this step are that by unifying the modeling of multimodal features, the system's perception accuracy and response efficiency for user behavior are significantly improved, while reducing redundant computation and information fragmentation caused by independent modeling of each modality. Simultaneously, the multi-task learning mechanism enhances the model's generalization ability and robustness, providing a solid data foundation and algorithmic support for subsequent personalized interaction optimization and intelligent driving assistance decision-making.

[0056] S22 improves the accuracy of fatigue state recognition by weighted fusion of eye movement trajectory features and facial expression features through an attention mechanism.

[0057] Specifically, in some implementations, this invention improves the accuracy of fatigue state recognition by introducing an attention mechanism to weightedly fuse eye-tracking features and facial expression features. The core technical principle of this step lies in using the attention mechanism to dynamically adjust the weights of different modal features, thereby highlighting the most discriminative feature information for fatigue state recognition during the multimodal fusion process.

[0058] In terms of specific implementation, firstly, eye-tracking features are collected through eye-tracking devices, including gaze coordinates, gaze duration, saccade path, and pupil diameter changes. These features are typically represented in time-series form and are normalized and temporally modeled by a feature extraction module, for example, using LSTM or Transformer structures for encoding. Facial expression features are collected by a high-resolution camera, and facial key points, action unit intensity, head pose, etc., are extracted through pre-trained deep learning models (such as FACET, OpenFace, or ResNet-based micro-expression recognition networks) to form a multi-dimensional feature vector.

[0059] In the feature fusion stage, a multi-head self-attention mechanism is employed to model the interaction between eye movement and facial expression features. The feature vector of each modality serves as the input to the attention module, and an attention score is calculated using a learnable weight matrix to generate the fused feature representation. Optionally, cross-modal attention can be introduced to establish a bidirectional dependency between eye movement features and facial expression features, enhancing semantic alignment between modalities.

[0060] Furthermore, the parameter settings for the attention mechanism in this step include the number of attention heads (usually 4-8), feature dimension (e.g., 128 or 256), activation function (e.g., ReLU or GELU), and normalization method (LayerNorm). During model training, a cross-entropy loss function is used, combined with an early stopping strategy and a learning rate decay mechanism to improve model convergence efficiency and generalization ability.

[0061] This technology has significant application value in intelligent cockpit scenarios, especially in driver fatigue monitoring systems. It can effectively integrate eye movement and facial information from the visual modality, improving the system's robustness and recognition accuracy in real-world driving environments with complex lighting and occlusion. Experiments show that this fusion strategy can improve fatigue recognition accuracy by more than 10%, reaching over 92% (based on transfer training on datasets such as KDEF and RAF-DB), significantly outperforming traditional linear weighted fusion methods.

[0062] S3. Based on the user state model, a dynamic evaluation strategy is constructed, and the cockpit interface layout, interaction path, and voice response parameters are adjusted in real time through a reinforcement learning mechanism.

[0063] Specifically, this step involves constructing a dynamic evaluation strategy based on the user state model and adjusting the interface layout, interaction path, and voice response parameters of the intelligent cockpit in real time through a reinforcement learning mechanism. This is a core component for achieving a personalized and adaptive user interaction experience. In some implementations, this process first relies on a user state model constructed by a multimodal perception module (such as eye tracking, finger trajectory, facial expression recognition, and physiological signal acquisition). This model is typically represented in the form of time-series data and includes key features such as user attention distribution, emotional state, operational intent, and cognitive load. By fusing data from visual, auditory, and tactile modalities, the system can generate a high-dimensional state vector, which serves as the input observation space for the reinforcement learning agent.

[0064] At the technical implementation level, the system employs a deep reinforcement learning framework, with user satisfaction, operational efficiency, and interaction fluency as the reward function design objectives. In each round of interaction, the agent receives the current user state information and outputs adjustment strategies for the cockpit interface layout (such as control positions and information hierarchy), interaction paths (such as gesture recognition priority and touch response logic), and voice response parameters (such as speech rate, tone, and wake-word sensitivity). Specific parameter settings include: an interface refresh rate controlled between 30-60Hz to ensure real-time performance, and a voice response latency of less than 500ms to meet the response time requirements for in-vehicle voice systems in ISO 21434.

[0065] In practical applications, this step can be deployed in the real-time interaction system of a smart cockpit, and is particularly suitable for adaptive human-machine interaction optimization in complex driving scenarios. For example, when user fatigue or inattention is detected, the system can automatically simplify the interface hierarchy, enhance the intensity of voice prompts, and adjust the interaction path to reduce the risk of misoperation. By continuously learning user behavior patterns, the system can achieve dynamic evolution of personalized interaction strategies, thereby improving user experience and driving safety.

[0066] The technical value of this step lies in the fact that by introducing a reinforcement learning mechanism, the system is equipped with online learning and real-time decision-making capabilities, effectively solving the problem that traditional static evaluation strategies cannot adapt to changes in user status, and significantly enhancing the adaptability and intelligence level of the intelligent cockpit system.

[0067] S4. The optimization effect is verified using a closed-loop testing system. The closed-loop testing system integrates a simulated mouth, a noise interference player, and an online semantic recognition module. The user state model is iteratively updated based on feedback from real vehicle test data.

[0068] Specifically, in some implementations, the "verification of optimization effect using a closed-loop testing system" step of this invention involves dynamically verifying and iteratively updating the optimization effect of the intelligent cockpit user state model through a closed-loop testing system integrating a simulated mouth, a noise interference player, and an online semantic recognition module. This closed-loop testing system simulates voice interaction scenarios in a real driving environment. It uses a simulated mouth to simulate user voice input and incorporates environmental noise (such as wind noise, engine noise, and road noise) through a noise interference player to test the robustness and accuracy of the voice recognition module under complex acoustic conditions. Simultaneously, the online semantic recognition module parses the semantic content of the voice commands in real time, verifying whether the system correctly understands the user's intent, and compares the recognition results with preset keywords, forming a closed-loop feedback.

[0069] In practical operation, the simulated mouth can be configured as a multi-angle sound source localization device, supporting voice input simulation at different distances (e.g., 0.3m to 1.5m), different azimuth angles (±60°), and different speech rates (120-180 words / minute) to cover diverse user scenarios. The interference sound player can load vehicle noise samples conforming to the ISO 362 standard, with noise intensity ranging from 40-85dB and frequency coverage from 20Hz to 20kHz, to simulate a real driving environment. The online semantic recognition module adopts a speech-semantic joint model based on the Transformer architecture, supporting multilingual (e.g., Chinese, English, Cantonese) and dialect recognition, with a semantic recognition accuracy of over 92% on the standard test set.

[0070] This step plays a crucial role in the verification and optimization of the entire system. Through continuous feedback from real-vehicle test data, parameters such as emotion recognition, attention distribution, and interaction intent in the user state model can be dynamically adjusted, thereby improving the model's generalization ability and real-time response performance. Furthermore, this closed-loop testing mechanism supports integration with the CI / CD process, ensuring that the effectiveness of each model update can be verified through automated testing, achieving continuous optimization and iterative upgrades of the intelligent cockpit interaction system.

[0071] The multimodal user interaction behavior fusion analysis and evaluation method of this invention realizes accurate perception and fusion analysis of multimodal interaction behaviors such as micro-expressions, voice, and actions of users in intelligent cockpits, improves the comprehensiveness and accuracy of user behavior evaluation, and optimizes human-computer interaction experience and the response capability of intelligent driving system.

[0072] Furthermore, S4 includes: S41 uses a simulated mouth to simulate speech input at different sound intensity levels, verifying the wake-up rate and semantic recognition robustness of the voice interaction system.

[0073] Specifically, this step simulates voice input at different sound intensity levels using a simulated mouth to verify the wake-up rate and semantic recognition robustness of the intelligent cockpit voice interaction system. In some implementations, the simulated mouth employs an acoustically driven mechanical sound-generating device combined with a programmable audio signal generator to achieve precise control and reproduction of the voice signal. Its core purpose is to simulate the pronunciation behavior of real users under different environmental noise and voice intensities, thereby evaluating the performance of the voice interaction system under complex acoustic conditions.

[0074] From a technical implementation perspective, a simulated mouth typically consists of a speaker unit, an acoustic cavity, and an adjustable mouth position mechanical structure, capable of simulating the position, direction, and distance of human voice pronunciation. The speech signal is generated through a pre-set speech library or a speech synthesis engine (such as a TTS system), covering various speech types including Mandarin, dialects, and male and female voices. During testing, the system can set the sound intensity level of the speech input, usually measured in dB SPL (sound pressure level), covering a range from 30dB to 90dB, to simulate various usage scenarios from whispering to shouting. Simultaneously, background noise (such as road noise, wind noise, music playback, etc.) can be introduced and superimposed onto the speech signal using an interference player to test the system's anti-interference capability in noisy environments.

[0075] In terms of parameter metrics, key performance indicators need to be set during testing, including wake-up rate (success rate of wake-up word recognition), false wake-up rate, semantic recognition accuracy, and response latency. The wake-up rate is usually expressed as a percentage and should not be less than 85% under different sound intensity levels; semantic recognition robustness is quantitatively evaluated through indicators such as confusion matrix and word error rate (WER), and the target recognition accuracy should be stable at over 90%.

[0076] In terms of application scenarios, this step is widely used in the development and verification phase of intelligent cockpit voice interaction systems, especially under extreme conditions that are difficult to reproduce in real vehicle environments, such as high-speed driving, strong winds, or multiple people speaking simultaneously. Standardized testing using simulated mouths ensures the stability and reliability of the voice system under different user voice characteristics and environmental noise levels.

[0077] In terms of technical effectiveness, this step effectively improves the environmental adaptability and user coverage of the voice interaction system, provides key data support for subsequent model optimization and system deployment, and is an important link in realizing closed-loop testing and continuous iteration of the intelligent cockpit multimodal interaction system.

[0078] S42, based on the injection of in-vehicle ambient noise by an interference sound player, tests the response stability of the voice interaction system under complex acoustic conditions.

[0079] Specifically, this step, "testing the response stability of the voice interaction system under complex acoustic conditions by injecting in-vehicle environmental noise using a noise interference player," is one of the key testing steps for multimodal user behavior analysis and evaluation of the voice interaction system in a smart cockpit. Its technical implementation is based on acoustic environment simulation and speech signal processing technology, aiming to evaluate the robustness and response stability of the voice interaction system in real or simulated complex noise environments.

[0080] In some implementations, this step involves deploying a high-fidelity noise interfering player in a test environment to inject preset environmental noise signals into the in-vehicle voice system, such as engine noise, wind noise, tire noise, broadcast sounds, and passenger conversations. These noise signals are typically played in multi-channel audio format to simulate the sound field distribution of sound sources in different locations (such as the front, rear, and outside the windows) to more realistically reproduce the in-vehicle acoustic environment. The noise interfering player supports multiple audio inputs and has programmable playback timing and volume control functions. Its output sound pressure level (SPL) can be adjusted to a range of 60-85 dB to cover typical noise levels in everyday driving scenarios.

[0081] Furthermore, during the test, an artificial mouth (articulatory head) was used to simulate user voice input. Its position and angle were adjustable to match the pronunciation position of a real driver. A microphone (such as an omnidirectional or directional microphone array) was used to collect the input and feedback audio signals of the voice interaction system, supporting sampling rates from 16kHz to 48kHz and a signal-to-noise ratio (SNR) of over 20dB. The test system verified the voice interaction results through an online semantic recognition module, determining whether the system correctly recognized and responded to preset keywords, thus forming a closed-loop test for voice interaction.

[0082] In practical applications, this step can be deployed in laboratory simulation environments or real-vehicle test platforms, and is particularly suitable for evaluating key performance indicators such as voice wake-up rate, voice recognition accuracy, and response latency. By injecting noise signals of different frequency bands, intensities, and timings, the stability of the voice system under different driving scenarios (such as high-speed driving, congested traffic, and air conditioning on) can be verified.

[0083] In terms of technical effectiveness, this step effectively improves the accuracy of the anti-interference capability assessment of the voice interaction system, provides reliable data support for subsequent model optimization and system deployment, and ensures the consistency and safety of the interactive experience of the smart cockpit in complex acoustic environments.

[0084] The multimodal user interaction behavior fusion analysis and evaluation method of this invention realizes accurate perception and fusion analysis of multimodal interaction behaviors such as micro-expressions, voice, and actions of users in intelligent cockpits, improves the comprehensiveness and accuracy of user behavior evaluation, and optimizes human-computer interaction experience and the response capability of intelligent driving system.

[0085] Furthermore, embodiments of the present invention also support the construction of personalized user profiles based on users' historical interaction data, and the adjustment of weight parameters in the dynamic evaluation strategy according to the profiles.

[0086] Specifically, this step involves constructing a personalized user profile based on historical user interaction data and dynamically adjusting the weight parameters in the evaluation strategy according to this profile. This is the core component for realizing personalized feedback and adaptive optimization in the intelligent cockpit user behavior analysis system. In some implementations, this process first acquires user behavior data in different interaction scenarios through a multimodal data acquisition module (including eye tracking, finger trajectory, facial expressions, voice interaction, etc.), covering key indicators such as attention distribution, operation path, emotional state, and voice response time. After preprocessing (such as noise reduction, time alignment, and feature extraction), this data is input into the user profile modeling module. Deep learning-based user feature fusion algorithms (such as multimodal attention mechanisms and graph neural networks) are used for feature weighting and clustering analysis to construct a high-dimensional user profile that includes dimensions such as user preferences, behavioral habits, and emotional tendencies.

[0087] At the parameter level, user profile models typically include multiple feature dimensions. For example, the threshold for dwell time in the area of ​​visual attention (AOI) is set to be more than 0.5 seconds to be considered valid attention. The meandering of the finger operation path is calculated using the ratio of Euclidean distance to the actual path length. Fatigue state recognition is based on facial action unit (AU) intensity and blinking frequency (more than 3 blinks per second per unit time is considered fatigue). The voice interaction response latency must be controlled within 300ms to comply with the safety standard for human-computer interaction response time in ISO 21434. In the dynamic evaluation strategy, the weight parameters of each modality are adjusted in real time according to the feature values ​​in the user profile. For example, for users with inattention, the system can increase the weight of voice feedback and reduce the frequency of visual cues to improve interaction efficiency.

[0088] This step is widely used in practical applications for personalized interaction optimization in smart cockpits. For example, during driving, the system can dynamically adjust the voice wake-up sensitivity or interface prompts based on the user's emotional state, thereby improving user experience and driving safety. Technically, this method significantly enhances the system's adaptability to individual user differences, improves the accuracy of behavioral analysis and the relevance of evaluation feedback, and provides a data foundation and strategic basis for subsequent reinforcement learning-based adaptive optimization, demonstrating high engineering practical value and innovation.

[0089] Furthermore, in this embodiment of the invention, the evaluation results of the user state model are transmitted to the intelligent driving decision module in real time to optimize the path planning and interactive prompt logic of the autonomous driving system.

[0090] Specifically, in some implementations, transmitting the evaluation results of the user state model to the intelligent driving decision-making module in real time is one of the key technical steps in realizing human-vehicle collaborative intelligent interaction in this invention. This step, by constructing an efficient data communication mechanism and model interface, transmits the output results of the user state model generated by fusing multimodal perception modules (such as eye tracking, finger trajectory, facial expressions, voice interaction, etc.) to the decision control unit of the intelligent driving system in a low-latency and highly reliable manner, thereby achieving dynamic optimization of path planning and interaction prompt logic.

[0091] At the technical implementation level, this step relies on a high-speed communication bus, such as CAN FD or Ethernet AVB, between the embedded system and the in-vehicle computing platform to ensure data transmission latency is controlled within 50ms. The output of the user state model typically includes key indicators such as user attention distribution, emotional state (e.g., pleasure, anxiety, fatigue), and cognitive load level. This information is encapsulated and transmitted through standardized middleware interfaces (e.g., ROS 2 or AUTOSAR). In the intelligent driving decision-making module, path planning algorithms (e.g., A*, Dijkstra, RRT*) and interactive prompting logic (e.g., HMI prompts, voice feedback, warning mechanisms) will be adjusted in real time based on the received user state information. For example, when a user is detected to be in a highly distracted state, the system can automatically reduce the frequency of non-critical information prompts or prioritize safer and more stable driving strategies in path planning.

[0092] In terms of parameter metrics, the output of the user state model must meet the requirements for response time and accuracy of human-computer interaction systems in ISO 21448 (expected functional safety) and ISO 26262 (functional safety). The model output frequency is recommended to be no less than 10Hz to ensure timely response to changes in user state. Emotion recognition accuracy should reach over 85%, and the confidence level of attention detection should be higher than 90% to ensure sufficient reliability of the input data for the decision-making module.

[0093] In terms of application scenarios, this step is widely applicable to autonomous driving systems in complex traffic environments, especially in high-risk scenarios such as driver emotional fluctuations, fatigue, or inattention, where it can significantly improve the system's active safety performance and user-friendliness. For example, on highways, if the system detects driver fatigue, it can automatically switch to a more conservative path planning strategy and enhance interactive guidance through voice or visual cues, thereby reducing the risk of accidents.

[0094] From a technical perspective, this step achieves a closed-loop integration of user state perception and intelligent driving decision-making, enhancing the environmental adaptability and personalized service capabilities of the autonomous driving system. By dynamically adjusting interactive prompts and path planning strategies, the system can better match the user's current psychological and physiological state, thereby enhancing driving safety and passenger comfort. This demonstrates the innovative value and practical significance of this invention in the collaborative optimization of intelligent cockpits and autonomous driving.

[0095] To achieve the above embodiments, the present invention also proposes a device for fusion analysis and evaluation of multimodal user interaction behavior in intelligent cockpits. Figure 2 This is a schematic diagram of the structure of a smart cockpit multimodal user interaction behavior fusion analysis and evaluation device provided in an embodiment of the present invention. Figure 2 As shown, the device includes: The multimodal interaction data synchronous acquisition module 100 is used to synchronously acquire multimodal interaction data of the smart cockpit user. The data includes eye movement trajectory, finger movement path, physiological and motion signals, facial expression features and voice interaction signals. The cross-modal feature fusion processing module 200 is used to perform cross-modal feature fusion processing on the multimodal data based on deep learning algorithms to generate a user state model, which includes quantitative assessment results of fatigue level, emotional fluctuation and cognitive load status. The dynamic evaluation strategy construction module 300 is used to construct a dynamic evaluation strategy based on the user state model and adjust the cockpit interface layout, interaction path and voice response parameters in real time through a reinforcement learning mechanism. The closed-loop test verification module 400 is used to verify the optimization effect using a closed-loop test system. The closed-loop test system integrates a simulated mouth, a noise interference player, and an online semantic recognition module, and iteratively updates the user state model based on feedback from real vehicle test data.

[0096] Regarding the apparatus in the above embodiments, the specific manner in which each module performs its operation has been described in detail in the embodiments related to the method, and will not be elaborated upon here.

[0097] To implement the above embodiments, the present invention also proposes an electronic device, comprising: a processor, and a memory communicatively connected to the processor; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to implement the method provided in the foregoing embodiments.

[0098] To implement the above embodiments, the present invention also proposes a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided in the foregoing embodiments.

[0099] To implement the above embodiments, the present invention also proposes a computer program product, including a computer program that, when executed by a processor, implements the methods provided in the foregoing embodiments.

[0100] The collection, storage, use, processing, transmission, provision, and disclosure of user personal information involved in this invention all comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0101] It should be noted that personal information collected from users should be used for legitimate and reasonable purposes and should not be shared or sold outside of these legitimate uses. Furthermore, such collection / sharing should only be conducted after receiving the user's informed consent, including but not limited to notifying the user to read the user agreement / user notice and sign an agreement / authorization that includes authorization of relevant user information before the user uses the function. In addition, any necessary steps must be taken to protect and safeguard access to such personal information data and ensure that others with access to personal information data comply with their privacy policies and procedures.

[0102] This invention is intended to provide implementation schemes for users to selectively prevent the use or access to personal information data. That is, this disclosure is intended to provide hardware and / or software to prevent or block access to such personal information data. Once personal information data is no longer needed, risks can be minimized by restricting data collection and deleting data. Furthermore, where applicable, such personal information can be de-identified to protect user privacy.

[0103] In the foregoing descriptions of the embodiments, the terms "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of the present invention. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0104] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this invention, "a plurality of" means at least two, such as two, three, etc., unless otherwise explicitly specified.

[0105] Any process or method description in the flowchart or otherwise herein can be understood as representing a module, segment, or portion of code comprising one or more executable instructions for implementing custom logic functions or processes, and the scope of preferred embodiments of the invention includes additional implementations in which functions may be performed not in the order shown or discussed, including substantially simultaneously or in reverse order depending on the functions involved, as will be understood by those skilled in the art to which embodiments of the invention pertain.

[0106] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus, or device (such as a computer-based system, a processor-included system, or other system that can fetch and execute instructions from, an instruction execution system, apparatus, or device). For the purposes of this specification, "computer-readable medium" can be any means that can contain, store, communicate, propagate, or transmit programs for use by, or in conjunction with, an instruction execution system, apparatus, or device. More specific examples (a non-exhaustive list) of computer-readable media include: an electrical connection having one or more wires (electronic device), a portable computer disk drive (magnetic device), random access memory (RAM), read-only memory (ROM), erasable and editable read-only memory (EPROM or flash memory), fiber optic devices, and portable optical disc read-only memory (CDROM). Alternatively, the computer-readable medium may be paper or other suitable media on which the program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other medium, followed by editing, interpreting, or otherwise processing as necessary, and then stored in a computer memory.

[0107] It should be understood that various parts of the present invention can be implemented in hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented in software or firmware stored in memory and executed by a suitable instruction execution system. For example, if implemented in hardware as in another embodiment, it can be implemented using any of the following techniques known in the art, or a combination thereof: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0108] Those skilled in the art will understand that all or part of the steps of the methods in the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, the program includes one or a combination of the steps of the method embodiments.

[0109] Furthermore, the functional units in the various embodiments of the present invention can be integrated into a processing module, or each unit can exist physically separately, or two or more units can be integrated into a module. The integrated module can be implemented in hardware or as a software functional module. If the integrated module is implemented as a software functional module and sold or used as an independent product, it can also be stored in a computer-readable storage medium.

[0110] The storage medium mentioned above can be a read-only memory, a disk, or an optical disk, etc. Although embodiments of the present invention have been shown and described above, it is to be understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions, and variations to the above embodiments within the scope of the present invention.

[0111] It should be understood that the various forms of processes shown above can be used, with steps reordered, added, or deleted. For example, the steps described in this invention can be executed in parallel, sequentially, or in different orders, as long as the desired result of the technical solution of this invention can be achieved, and this is not limited herein.

[0112] The specific embodiments described above do not constitute a limitation on the scope of protection of this invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this invention should be included within the scope of protection of this invention.

Claims

1. An intelligent cockpit multi-modal user interaction behavior fusion analysis and evaluation method, characterized in that, Comprise: S1, synchronously collect multi-modal interaction data of intelligent cockpit users, including eye movement trajectory, finger movement path, physiological and motion signal, facial expression feature and voice interaction signal; S2, cross-modal feature fusion processing of multi-modal data based on deep learning algorithm, generating user state model, the model contains fatigue degree, emotional fluctuation and cognitive load state quantitative evaluation results; S3, according to the user state model, a dynamic evaluation strategy is constructed, and cockpit interface layout, interaction path and voice response parameters are adjusted in real time through reinforcement learning mechanism; S4, the optimization effect is verified by using closed loop test system, the closed loop test system integrates simulation mouth, interference sound player and online semantic recognition module, and the user state model is updated based on real vehicle test data feedback iteration.

2. The method of claim 1, wherein, The synchronous collection of multi-modal interaction data of intelligent cockpit users also includes: Cross-device time sequence alignment of eye movement trajectory data and facial expression features collected by 3D camera through timestamp synchronization mechanism; Submillimeter level precision tracking of finger movement path by infrared induction and millimeter wave radar fusion scheme.

3. The method of claim 1, wherein, The cross-modal feature fusion processing of multi-modal data based on deep learning algorithm also includes: Constructing a multi-task learning framework to process the feature vectors of micro-expression recognition, gesture analysis and voice interaction evaluation at the same time; Weighted fusion of eye movement trajectory features and facial expression features through attention mechanism to improve fatigue state recognition accuracy.

4. The method of claim 1, wherein, The optimization effect is verified by using closed loop test system, the closed loop test system integrates simulation mouth, interference sound player and online semantic recognition module, and the user state model is updated based on real vehicle test data feedback iteration. Also includes: Based on user historical interaction data, construct personalized user portrait, adjust weight parameters in dynamic evaluation strategy according to the portrait; 5. The method of claim 1, wherein, The evaluation results of the user state model are transmitted to the intelligent driving decision module in real time, which is used to optimize the path planning and interaction prompt logic of the automatic driving system. Comprise: Multi-modal interaction data synchronous collection module, for synchronously collecting multi-modal interaction data of intelligent cockpit users, the data includes eye movement trajectory, finger movement path, physiological and motion signal, facial expression feature and voice interaction signal; 6. An intelligent cockpit multi-modal user interaction behavior fusion analysis and evaluation device, characterized in that, Cross-modal feature fusion processing module, for cross-modal feature fusion processing of multi-modal data based on deep learning algorithm, generating user state model, the model contains fatigue degree, emotional fluctuation and cognitive load state quantitative evaluation results; Dynamic evaluation strategy construction module, for constructing dynamic evaluation strategy according to the user state model, and adjusting cockpit interface layout, interaction path and voice response parameters in real time through reinforcement learning mechanism; Closed loop test verification module, for verifying the optimization effect by using closed loop test system, the closed loop test system integrates simulation mouth, interference sound player and online semantic recognition module, and the user state model is updated based on real vehicle test data feedback iteration. The multi-modal interaction data synchronous collection module is also used for: ​ 7. The apparatus of claim 6, wherein, ​ The eye movement track data and the facial expression features collected by the 3D camera are cross-device time sequence aligned through a timestamp synchronization mechanism; An infrared induction and millimeter wave radar fusion scheme is adopted to track the finger movement path with sub-millimeter level precision.

8. An electronic device, comprising: The method comprises the steps of: a processor, and a memory connected to the processor in communication; the memory stores computer execution instructions; the processor executes the computer execution instructions stored in the memory to realize the method of any one of claims 1-5.

9. A computer-readable storage medium, characterized in that, The computer readable storage medium stores computer execution instructions, and the computer execution instructions are executed by the processor to realize the method of any one of claims 1-5.

10. A computer program product, characterised in that, The computer program is executed by the processor to realize the method of any one of claims 1-5.