Cockpit interaction intention prediction method based on liquid network and large language cooperation

CN122863277APending Publication Date: 2026-10-02CHANGZHOU UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611004094.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-07-07
Publication Date
2026-10-02

AI Technical Summary

Technical Problem

驾驶员操作离散且时隙不均,LSTM、Transformer等离散时间模型依赖固定步长,处理多尺度时序依赖时精度不足,难以精确刻画行为随时间的连续演化

Benefits of technology

[0067](1)本发明通过引入液态神经网络(LNN)的闭式连续时间单元(CFC),能够对非均匀采样条件下的座舱交互序列进行连续时间建模,克服传统离散时序模型难以准确捕捉隐藏状态波动及动态演化规律的缺陷,从而提高意图预测的准确性和鲁棒性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122863277A_ABST
    Figure CN122863277A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of intelligent cockpit human-computer interaction and driver state perception, and particularly relates to a cockpit interaction intention prediction method based on liquid network and large language cooperation. The method comprises: training a prediction model using obtained vehicle dynamics and cockpit interaction data; inputting data at a to-be-predicted time into the trained model to obtain a cockpit interaction intention prediction result. The model comprises: a scene semantic encoding module, which converts vehicle dynamics data into structured scene semantic information through a large language model; a cockpit interaction data encoding module, which uses a closed continuous unit of a liquid neural network to model continuous time for a cockpit interaction data sequence and extract time sequence features; a double-level fusion module, which performs adaptive alignment and fusion on structured scene semantic information and time sequence features through semantic modulation scaling and a dynamic gate cross attention mechanism; and an intention prediction module. The present application can effectively improve the accuracy, stability and real-time performance of intelligent cockpit intention prediction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent cockpit human-computer interaction and driver state perception technology, specifically to a method for predicting cockpit interaction intent based on liquid network and large language collaboration. Background Technology

[0002] As intelligent cockpits evolve towards proactive intent understanding, driver-cockpit interaction intent prediction has become a key technology for realizing personalized proactive services. This type of intent specifically refers to cockpit control behaviors (such as air conditioning adjustment and media switching) triggered by the driver's physiological comfort or infotainment needs. Currently, driver monitoring research mainly focuses on surface-level action recognition such as fatigue or gestures, lacking the ability to deeply reason about interaction intents in driving scenarios.

[0003] Existing methods face two main challenges:

[0004] First, there is a semantic gap between physical signals and true intentions. Interaction intentions are essentially the product of "environmental states stimulating psychological needs, which in turn drive actions," and traditional numerical correlation models such as CNNs and RNNs struggle to model the causal logic involved. While large language models possess the potential for semantic reasoning, how to deeply integrate their qualitative knowledge with the high-frequency quantitative features of sensors remains an unsolved problem.

[0005] Secondly, cockpit operations exhibit continuous-time dynamic characteristics with non-uniform sampling. Pilot operations are discrete and time-slotted. Discrete-time models such as LSTM and Transformer rely on fixed step sizes, which are insufficient in accuracy when dealing with multi-scale temporal dependencies, making it difficult to accurately characterize the continuous evolution of behavior over time.

[0006] Furthermore, existing threshold- or rule-based methods are passive triggering mechanisms that lack the ability to predict in advance and cannot meet the real-world demand for both "environment-induced" and "advanced prediction" scenarios.

[0007] Therefore, there is an urgent need to build a unified prediction framework that combines semantic reasoning and continuous-time modeling capabilities to achieve a technological leap from passive identification to proactive prediction. Summary of the Invention

[0008] The technical problem to be solved by this invention is to overcome the defects of the prior art and provide a cockpit interaction intent prediction method based on liquid network and big language collaboration, thereby effectively improving the accuracy, stability and real-time performance of intelligent cockpit intent prediction, and providing reliable technical support for intelligent cockpit proactive services and real-time vehicle deployment.

[0009] To solve the above-mentioned technical problems, the technical solution of the present invention is: a cockpit interaction intent prediction method based on liquid network and large language collaboration, comprising:

[0010] Acquire vehicle dynamic data and cockpit interaction data sequences;

[0011] Constructing a predictive model includes:

[0012] The scene semantic encoding module is used to convert vehicle dynamic data into structured scene semantic information through a large language model. ;

[0013] The cockpit interaction data encoding module is used to perform continuous-time modeling of the cockpit interaction data sequence using closed continuous units of a liquid neural network, and to extract temporal features that characterize the dynamic evolution of the driver's interaction behavior. ;

[0014] The two-level fusion module is used to process the semantic information of structured scenes through a semantic modulation scaling mechanism and a dynamic gating cross-attention mechanism. With time series characteristics Adaptive alignment and deep fusion are performed to obtain multimodal fusion features. ;

[0015] The intent prediction module is used to fuse multimodal features. Map the interaction to the corresponding interaction category and output the driver's interaction intent prediction result;

[0016] A prediction model is trained using vehicle dynamic data and cockpit interaction data sequences to obtain a well-trained prediction model.

[0017] The vehicle dynamics data and cockpit interaction data sequence at the time to be predicted are input into the trained prediction model to obtain the cockpit interaction intention prediction result.

[0018] Furthermore, a large language model is used to convert vehicle dynamic data into structured scene semantic information. Specifically, this includes:

[0019] First, vehicle dynamics data is converted into scene semantic text S using a large language model fine-tuned by LoRA.

[0020] Then, the scene semantic text is mapped into high-level semantic text, i.e., structured scene semantic information, through a semantic projection encoder. The semantic projection encoder consists of two fully connected layers and a SiLU activation function.

[0021] Furthermore, closed-loop continuous units of a liquid neural network are used to perform continuous-time modeling of the cockpit interaction data sequence, extracting temporal features that characterize the dynamic evolution of the driver's interactive behavior. Specifically, this includes: First, projecting the cockpit interaction data sequence into the latent space through a linear mapping layer, and superimposing position encoding to form the initial temporal sequence Z;

[0022] Then, closed continuous units are used to extract features from the initial sequence Z to obtain the temporal hidden state sequence H.

[0023] Finally, attention pooling is performed on the temporal hidden state sequence H to obtain the global temporal feature H. pooled .

[0024] Furthermore, closed continuous unit traversal is used to extract features from the initial sequence Z, resulting in the temporal hidden state sequence H; specifically including:

[0025] First, for the initial time series Z, at any time step t, the backbone network receives the current input Z. t Compared with the previous hidden state H t-1 A fused input is formed through feature splicing. ;

[0026] Then, from the fused input via the backbone network Extracting fusion features ;

[0027] Based on fusion features Two candidate signals FF1 and FF2, and a time-gating coefficient τ are generated in parallel:

[0028] ;

[0029] ;

[0030] in, and These are the weights of candidate signals FF1 and FF2, respectively. The hyperbolic tangent activation function is used. Indicates the time span between adjacent time steps. and This is the learnable weight matrix in the time-gated branch; This represents the Sigmoid activation function;

[0031] Finally, based on the time gating coefficient τ, the two candidate signals FF1 and FF2 are dynamically weighted and fused to update the hidden state H at the current time. t :

[0032] ;

[0033] After traversing the entire initial time sequence Z, the temporal hidden state sequence H is obtained.

[0034] Furthermore, attention pooling is applied to the temporal hidden state sequence H to obtain the global temporal feature H. pooled Specifically, this includes:

[0035] First, calculate the global query vector Q;

[0036] Then, the temporal information is aggregated through a multi-head attention mechanism to obtain the attention output. ;

[0037] Finally, for Perform layer normalization to obtain global temporal features .

[0038] Furthermore, the two-level fusion module includes a semantic modulation scaling fusion (MSF) submodule and a dynamic gated cross-attention (DGCA) submodule. The DGCA submodule uses the temporal fusion features refined by the MSF submodule. As a query.

[0039] Furthermore, the working process of the Semantic Modulation Scaling Fusion (MSF) submodule includes:

[0040] First, global temporal features Perform semantically guided affine modulation to achieve dynamic recalibration at the feature level:

[0041] ;

[0042] Where γ1 and β1 are modulation parameters, and ⊙ denotes element-wise multiplication. Presentation layer normalization operation;

[0043] Subsequently, the modulated features Perform multi-head self-attention computation and preserve the original timing information through residual connections:

[0044] ;

[0045] in, This indicates multi-head self-attention computation;

[0046] Then, the temporal features are combined using a cross-attention mechanism. With structured scene semantic information To conduct

[0047] By aligning, we can obtain cross-modal interaction features. :

[0048] ;

[0049] in, Indicates alignment operation;

[0050] Finally, a second modulation and feedforward network transformation is performed to obtain the refined temporal fusion features. :

[0051] ;

[0052] Where γ2 and β2 are the second set of modulation parameters, FFN This represents a feedforward network.

[0053] Furthermore, the working process of the Dynamic Gated Cross-Attention (DGCA) submodule includes:

[0054] First, the temporal fusion features refined by the semantic modulation scaling fusion MSF submodule are used. As a query, it uses structured scene semantic information Using these as keys and values, semantically guided cross-attention outputs are computed. ;

[0055] Subsequently, learnable gating parameters were introduced. Dynamic gating weights are generated using the Sigmoid function. Adaptive adjustment of cross-attention output Features of temporal fusion Contribution percentage in the integration process:

[0056]

[0057] in, This represents element-wise multiplication;

[0058] Finally, the fusion features The input residual connection and layer normalization module are processed by a feedforward network for nonlinear transformation to obtain the final multimodal fusion features. .

[0059] Furthermore, multimodal fusion features Mapping to the corresponding interaction category, the driver's interaction intent prediction result is output; specifically including:

[0060] First, multimodal fusion features along the time dimension. Pooling is performed to obtain the global feature vector. ;

[0061] Then, the global feature vector The input is a fully connected layer, and after normalization by the Softmax activation function, the predicted probability distributions for each category are obtained;

[0062] Finally, the category corresponding to the highest predicted probability is selected as the final intention prediction result.

[0063] Furthermore, during the model training phase, the model parameters are optimized using a cross-entropy loss function that incorporates a label smoothing strategy. The loss function L is expressed as:

[0064]

[0065] Where K is the total number of categories, p k Let ỹ be the predicted probability of the model in the k-th class. k This represents the target label distribution after label smoothing.

[0066] By adopting the above technical solution, the present invention has the following beneficial effects:

[0067] (1) By introducing the closed continuous time unit (CFC) of liquid neural network (LNN), this invention can perform continuous time modeling of cabin interaction sequence under non-uniform sampling conditions, overcome the shortcomings of traditional discrete time series models that are difficult to accurately capture hidden state fluctuations and dynamic evolution laws, thereby improving the accuracy and robustness of intention prediction.

[0068] (2) This invention sets up a dual-level fusion module (DHFM), including two stages: Semantic Modulation Scaling (MSF) and Dynamic Gated Cross-Attention (DGCA). The former uses semantic priors to dynamically recalibrate temporal features, while the latter adaptively fuses semantic and temporal features through a gating mechanism. This avoids the insufficient information utilization caused by simply splicing semantic and physical temporal information or single attention fusion, enabling the model to automatically adjust the contributions of each modality according to different driving environments. It achieves adaptive alignment and deep integration of structured scene semantic information and temporal features, and can more effectively explore the correlation between multimodal information, thereby improving the stability and reliability of prediction results.

[0069] (3) The present invention has good application effect in real driving scenarios. The intention prediction accuracy can reach 78.5%, and the single reasoning time is only 6.2ms, which is significantly better than the existing benchmark model. It can meet the real-time application and deployment requirements of intelligent cockpit vehicles and provide reliable technical support for intelligent cockpit active services. Attached Figure Description

[0070] Figure 1 This is a framework diagram of the prediction model of the present invention;

[0071] Figure 2 This is a schematic diagram of the closed-loop continuous unit flow of the present invention;

[0072] Figure 3 This is a flowchart of the semantic modulation scaling fusion (MSF) submodule of the present invention;

[0073] Figure 4This is an example diagram illustrating the intent prediction of the present invention in a real driving scenario. Detailed Implementation

[0074] To make the content of this invention easier to understand, the invention will be further described in detail below with reference to specific embodiments and accompanying drawings.

[0075] like Figures 1 to 4 As shown, a cockpit interaction intent prediction method based on liquid network and large language collaboration includes:

[0076] Acquire vehicle dynamic data and cockpit interaction data sequences;

[0077] Constructing a predictive model includes:

[0078] The scene semantic encoding module is used to convert vehicle dynamic data into structured scene semantic information through a large language model. To bridge the semantic gap between physical signals and the psychological needs of drivers;

[0079] The cockpit interaction data encoding module is used to perform continuous-time modeling of the cockpit interaction data sequence using closed continuous units of a liquid neural network, and to extract temporal features that characterize the dynamic evolution of the driver's interaction behavior. To accurately depict the intention-triggered process and its changing patterns;

[0080] The two-level fusion module is used to process the semantic information of structured scenes through a semantic modulation scaling mechanism and a dynamic gating cross-attention mechanism. With time series characteristics Adaptive alignment and deep fusion are performed to obtain multimodal fusion features. ;

[0081] The intent prediction module is used to fuse multimodal features. Map the interaction to the corresponding interaction category and output the driver's interaction intent prediction result;

[0082] A prediction model is trained using vehicle dynamic data and cockpit interaction data sequences to obtain a well-trained prediction model.

[0083] The vehicle dynamics data and cockpit interaction data sequence at the time to be predicted are input into the trained prediction model to obtain the cockpit interaction intention prediction result.

[0084] Specifically, the prediction model input includes two types of data sequences constructed based on time windows: cockpit interaction data sequences and vehicle dynamic data. A 60-second sliding window is used to construct the input samples. Since the sampling frequency of vehicle dynamic data is 0.1Hz, the time step length within each time window is n=6. Cockpit interaction data includes the driver's interaction with the cockpit and the status information of cockpit equipment, including 24-dimensional features such as door position, gear position, and air conditioning. Vehicle dynamic data includes vehicle driving status and environmental status information, including 12-dimensional features such as vehicle speed, temperature, and mileage.

[0085] In this embodiment, as Figure 1 As shown, a large language model is used to convert vehicle dynamic data into structured scene semantic information. Specifically, this includes:

[0086] First, a large language model fine-tuned with LoRA is used to convert vehicle dynamic data into scene semantic text S. This involves semantic understanding and structured representation of time periods, vehicle states, road environment, temperature information, and driving-related states within the vehicle dynamic data, generating scene semantic prior information that can be used for subsequent intent prediction. The semantic text corresponding to a single moment can be represented as:

[0087]

[0088] Then, the scene semantic text is mapped into high-level semantic text, i.e., structured scene semantic information, through a semantic projection encoder. ;

[0089] The semantic projection encoder consists of two fully connected layers and a SiLU activation function, and its output can be expressed as:

[0090]

[0091] in, To unify the hidden dimension (set to 768), so as to provide contextual prior information for subsequent cross-modal fusion.

[0092] Specifically, by introducing a large language model fine-tuned with LoRA, vehicle dynamic data is transformed into structured scene semantic information. This effectively alleviates the semantic gap problem in the existing intelligent cockpit intent prediction process, improves the model's understanding of complex driving scenarios, and achieves an upgrade from traditional pattern matching to semantic understanding and cognitive reasoning. In this embodiment, the semantic information is not simply used as an auxiliary input for text embedding, but rather as a semantic prior for driving temporal feature reconstruction and cross-modal alignment. Existing technologies use Word2Vec to convert semantic text into vectors for attention fusion; this patent further uses structured scene semantics to modulate the interactive dynamic features extracted by CFC, allowing semantic factors such as "weather, temperature, road conditions, and driving status" to directly influence the generation process of interactive intent representation. Its beneficial effect is to better bridge the semantic gap between physical signals and the driver's psychological needs.

[0093] In this embodiment, the closed-loop continuous-time units of the liquid neural network are used to process cockpit interaction sequences under non-uniform sampling conditions to improve the ability to characterize the continuous changes in hidden states. The closed-loop continuous units of the liquid neural network are used to perform continuous-time modeling of the cockpit interaction data sequences, extracting temporal features that characterize the dynamic evolution of the driver's interactive behavior. Specifically, this includes: First, projecting the cockpit interaction data sequence into the latent space through a linear mapping layer, and superimposing position encoding to form the initial temporal sequence Z; the process can be represented as:

[0094]

[0095] In the formula, C represents cockpit interaction data. This represents a linear mapping, where P represents a positional encoding;

[0096] Then, closed continuous units are used to extract features from the initial sequence Z to obtain the temporal hidden state sequence H.

[0097] Finally, to further obtain temporal features containing global dynamic information, attention pooling is applied to the temporal hidden state sequence H to obtain the global temporal features Hi. pooled .

[0098] In this embodiment, as Figure 2 As shown, closed continuous units are used to extract features from the initial sequence Z, resulting in the temporal hidden state sequence H; specifically, this includes:

[0099] First, for the initial time series Z, at any time step t, the backbone network receives the current input Z. t Compared with the previous hidden state H t-1 A fused input is formed through feature splicing. :

[0100] ;

[0101] in, Indicates feature concatenation operation;

[0102] Then, the fusion features are extracted via the backbone network. :

[0103] ;

[0104] in, This indicates that feature extraction is performed using a backbone network;

[0105] Based on fusion features Two candidate signals FF1 and FF2, and a time-gating coefficient τ are generated in parallel:

[0106] ;

[0107] ;

[0108] in, Indicates time step The fusion features are extracted from the fusion input by the backbone network; and These represent two candidate state branches generated by different linear mappings, used to provide alternative information for the current hidden state update; and This is the learnable weight matrix corresponding to the candidate state branch; This is the hyperbolic tangent activation function, used to introduce a nonlinear mapping and restrict candidate signals to a stable range; It represents the time span between adjacent time steps and is used to reflect the time interval variation under non-uniform sampling conditions; and This is the learnable weight matrix in the time-gated branch; This represents the Sigmoid activation function, used to normalize the gated output to... Within the range; τ is a time gating coefficient with a value between 0 and 1, used to characterize the dynamic interpolation weight between information at different time scales at the current moment, that is, to control the contribution ratio of the two candidate state branches in the current hidden state update process;

[0109] Finally, based on the time gating coefficient τ, the two candidate signals FF1 and FF2 are dynamically weighted and fused to update the hidden state H at the current time. t

[0110] ;

[0111] After traversing the entire initial time sequence Z, the temporal hidden state sequence H is obtained.

[0112] In this embodiment, attention pooling is performed on the temporal hidden state sequence H to obtain the global temporal feature H. pooled Specifically, it includes:

[0113] First, calculate the global query vector Q:

[0114] ;

[0115] in, This indicates that average pooling is performed along the time dimension.

[0116] Then, the temporal information is aggregated through a multi-head attention mechanism to obtain the attention output. :

[0117] ;

[0118] in, This indicates the aggregation of multi-head attention mechanisms;

[0119] Finally, global temporal features are obtained through layer normalization. :

[0120] ;

[0121] in, Presentation layer normalization operation.

[0122] In this embodiment, as Figure 1 As shown, the two-level fusion module includes a semantic modulation scaling fusion (MSF) submodule and a dynamic gating cross-attention (DGCA) submodule. The DGCA submodule uses the temporal fusion features refined by the MSF submodule. As a query.

[0123] Figure 3 A flowchart of the semantic modulation scaling fusion (MSF) submodule is shown below. Figure 3 As shown, the working process of the Semantic Modulation Scaling Fusion (MSF) submodule includes:

[0124] First, global temporal features Perform semantically guided affine modulation to achieve dynamic recalibration at the feature level:

[0125] ;

[0126] Where γ1 and β1 are modulation parameters, and ⊙ denotes element-wise multiplication. Presentation layer normalization operation;

[0127] Subsequently, the modulated features Perform multi-head self-attention computation and preserve the original timing information through residual connections:

[0128] ;

[0129] in, This indicates multi-head self-attention computation;

[0130] Then, the temporal features are combined using a cross-attention mechanism. With structured scene semantic information To conduct

[0131] By aligning, we can obtain cross-modal interaction features. :

[0132] ;

[0133] in, Indicates alignment operation;

[0134] Finally, a second modulation and feedforward network transformation is performed to obtain the refined temporal fusion features. :

[0135] ;

[0136] Where γ2 and β2 are the second set of modulation parameters, FFN This represents a feedforward network.

[0137] In this embodiment, as Figure 1 As shown, the working process of the Dynamic Gated Cross-Attention (DGCA) submodule includes:

[0138] The Dynamic Gated Cross-Attention (DGCA) submodule receives the temporal fusion features output from the Semantic Modulation Scaling Fusion (MSF) submodule. And combined with scene semantic features To conduct cross-modal interaction. That is,

[0139] First, semantically guided cross-attention output is calculated using a cross-attention mechanism. The temporal fusion features refined by the semantic modulation scaling fusion MSF submodule are used. As a query, it uses structured scene semantic information As key and value:

[0140] =

[0141] W represents the cross-attention mechanism. Q WK and W V represents the projection matrices corresponding to the query, key, and value, respectively, and d represents the feature dimension.

[0142] Subsequently, adaptive adjustment of the contribution of different modes is further achieved by introducing learnable gating parameters. Dynamic gating weights are generated using the Sigmoid function. Adaptive adjustment of cross-attention output Features of temporal fusion Contribution percentage in the integration process:

[0143]

[0144] in, This represents element-wise multiplication; This represents the modal weights generated by learnable gating parameters. When scene semantic information is more instructive for judging the current interaction intent, the model can improve... The model is enhanced when the timing characteristics of cockpit interactions are more reliable. The function of. When When the size is large, the model pays more attention to the cross-attention features guided by scene semantics. ;when When the size is small, the model retains more of the temporal fusion features of the MSF output. Therefore, this structure can adaptively adjust the contribution ratio of semantic information and temporal information according to different driving scenarios and interaction states.

[0145] Finally, gated fusion features The output features of the dynamically gated cross-attention module are obtained by sequentially performing residual connections and layer normalization, followed by feedforward network nonlinear transformation. :

[0146]

[0147] in, Presentation layer normalization operation, Indicates a feedforward network. This represents the final multimodal fusion feature output by the two-level fusion module.

[0148] In this embodiment, multimodal fusion features are used. Mapping to the corresponding interaction category, the driver's interaction intent prediction result is output; specifically including:

[0149] First, multimodal fusion features along the time dimension. Pooling is performed to obtain the global feature vector. (A comprehensive feature used to characterize the pilot's cockpit interaction behavior and scene semantic information within the current time window):

[0150] Pool() represents pooling.

[0151] Then, the global feature vector The input is a fully connected layer, which is then normalized using the Softmax activation function to obtain the predicted probability distribution for each category. :

[0152]

[0153] in, This represents the learnable weight matrix of the fully connected layer of the classifier; This represents the classifier bias term; This represents the normalization exponential function, used to convert the classifier output into a predicted probability distribution for each intent category;

[0154] Finally, the category corresponding to the highest predicted probability is selected as the final intent prediction result:

[0155] .

[0156] In this embodiment, during the model training phase, the model parameters are optimized using a cross-entropy loss function that incorporates a label smoothing strategy. The loss function L is expressed as:

[0157]

[0158] Where K is the total number of categories, p k Let ỹ be the predicted probability of the model in the k-th class. k This represents the target label distribution after label smoothing.

[0159] The advantages of the solutions involved in the above embodiments will be introduced below with reference to specific experiments.

[0160] To verify the effectiveness of the method of this invention, the intelligent cockpit intent prediction method based on the collaboration of liquid neural network and large language model is compared with several existing typical prediction models, including TransDBC, Diff-Transformer, ETSformer, AFT-Full, and SGFormer. Its performance is evaluated using several metrics, including accuracy, F1 score, precision, and recall. The comparison results are shown in Table 1.

[0161] TransDBC 0.7290 0.6624 0.6825 0.6960 Diff-Transformer 0.7658 0.7251 0.7363 0.7244 ETSformer 0.6548 0.6129 0.6082 0.6335 AFT-Full 0.6093 0.6479 0.6090 0.6875 SGFormer 0.7174 0.7205 0.7125 0.7074 Our method 0.7850 0.8000 0.8492 0.7816

[0162] Accuracy measures the proportion of samples whose predicted class matches the true class. Higher accuracy indicates stronger overall predictive ability. The F1 score is the harmonic mean of precision and recall, comprehensively reflecting the model's overall performance in classification tasks. A higher F1 score indicates better performance in balancing precision and recall. Precision measures the proportion of samples correctly identified as belonging to a particular class. Higher precision indicates fewer false positives. Recall measures the proportion of samples correctly identified as belonging to a particular class. Higher recall indicates fewer false negatives.

[0163] As shown in Table 1, the experimental results of this invention outperform the comparative model in all evaluation metrics. Specifically, the accuracy of this invention reaches 0.7850, the F1 score reaches 0.8000, the precision reaches 0.8492, and the recall reaches 0.7816, all higher than the existing benchmark model SGFormer. This indicates that this invention, by introducing a large language model to generate structured scene semantic information, combining a liquid neural network to perform continuous-time modeling of non-uniform interaction sequences, and then achieving deep fusion of semantic priors and temporal features through a two-level fusion module, can effectively improve the accuracy and stability of predicting driver interaction intentions in intelligent cockpit scenarios.

[0164] To further illustrate the contribution of the Closed Continuous Time Unit (CFC) to the overall method in this invention, the CFC is replaced with various standard temporal modeling methods (including Transformer, LSTM, GRU, CNN, and MLP) while keeping the rest of the architecture unchanged.

[0165] Table 2 Comparison of CfC and Alternative Temporal Modeling Methods

[0166] Transformer 0.7699 0.7813 0.8282 0.7641 LSTM 0.7628 0.7705 0.8157 0.7558 GRU 0.7543 0.7682 0.8023 0.7543 CNN 0.7301 0.7356 0.7680 0.7240 MLP 0.7244 0.7349 0.7633 0.7261 CfC(Ours) 0.7850 0.8000 0.8492 0.7816

[0167] As shown in Table 2, the proposed CfC method achieved the best performance across all evaluation metrics, with an accuracy of 0.7850 and an F1 score of 0.8000. Specifically, CfC outperformed Transformer by 1.96% in accuracy and 2.39% in F1 score; compared to LSTM, CfC improved accuracy and F1 score by 2.91% and 3.83%, respectively; and compared to CNN and MLP, CfC's improvements were more significant, increasing accuracy by 7.52% and 8.36%, respectively, and F1 scores by 8.76% and 8.86%, respectively. These results demonstrate that the temporal convolution and attention mechanisms combined in CfC can effectively capture local patterns and global dependencies in driving scenarios, exhibiting superior temporal modeling capabilities.

[0168] Taking LSTM as an example, while existing technologies employ LSTM, it is essentially still a discrete-time series model, insufficient for representing the continuous evolution of non-uniform sampling, intermittent operations, and sudden cockpit interactions. This patent uses CFC to continuously update the hidden state over time, enabling a more accurate depiction of the dynamic changes in driver interaction intentions from environmental stimuli and psychological needs to operational behaviors, thereby improving prediction accuracy and robustness in complex driving scenarios.

[0169] To further illustrate the contribution of the two-layer fusion module (DHFM) to the overall method, DHFM is compared with several mainstream feature fusion methods to evaluate its effectiveness.

[0170] Table 3 Comparison of DHFM with different fusion methods

[0171] Concatenation + MLP 0.7699 0.7857 0.8229 0.7736 Fixed-weight Fusion 0.7727 0.7891 0.8329 0.7764 Single-layer Cross-Attn 0.6392 0.6317 0.6960 0.6194 FiLM Modulation 0.7741 0.7888 0.8450 0.7737 MHCA 0.7812 0.7934 0.8458 0.7724 DHFM (Ours) 0.7850 0.8000 0.8492 0.7816

[0172] As shown in Table 3, the proposed DHFM method achieves optimal performance across all evaluation metrics, with an accuracy of 0.7850 and an F1 score of 0.8000. Specifically, DHFM outperforms multi-layer attention mechanisms by 0.49% in accuracy and 0.83% in F1 score; compared to the FiLM modulation method, DHFM improves accuracy and F1 score by 1.41% and 1.42%, respectively; and compared to the basic feature concatenation method, DHFM's improvement is even more significant, increasing accuracy and F1 score by 1.96% and 1.82%, respectively. Notably, the single-layer cross-attention method performs significantly worse, validating the necessity of deep interaction structure design. These results demonstrate that DHFM, through its deep hybrid fusion mechanism, can more effectively integrate multi-source features, achieving optimal overall performance while maintaining high accuracy.

[0173] In summary, the experimental results demonstrate that this invention exhibits significant performance advantages in intelligent cockpit intent prediction tasks. It not only improves prediction accuracy but also balances model real-time performance with deployment efficiency. This showcases the strong generalization ability and application potential of this invention in proactive intelligent cockpit services. It can effectively enhance the accuracy and stability of driver interaction intent recognition in complex driving scenarios, providing strong support for the development of related intelligent cockpit technologies. These experimental results confirm the effectiveness of the method presented in this invention, indicating its good potential and adaptability in practical applications.

[0174] Based on the above-described preferred embodiments of the present invention, and through the foregoing description, those skilled in the art can make various changes and modifications without departing from the inventive concept. The technical scope of this invention is not limited to the contents of the specification, but must be determined according to the scope of the claims.

Claims

1. A method for predicting cockpit interaction intent based on liquid network and large language collaboration, characterized in that, include: Acquire vehicle dynamic data and cockpit interaction data sequences; Constructing a predictive model includes: The scene semantic encoding module is used to convert vehicle dynamic data into structured scene semantic information through a large language model. ; The cockpit interaction data encoding module is used to perform continuous-time modeling of the cockpit interaction data sequence using closed continuous units of a liquid neural network, and to extract temporal features that characterize the dynamic evolution of the driver's interaction behavior. ; The two-level fusion module is used to process the semantic information of structured scenes through a semantic modulation scaling mechanism and a dynamic gating cross-attention mechanism. With time series characteristics Adaptive alignment and deep fusion are performed to obtain multimodal fusion features. ; The intent prediction module is used to fuse multimodal features. Map the interaction to the corresponding interaction category and output the driver's interaction intent prediction result; A prediction model is trained using vehicle dynamic data and cockpit interaction data sequences to obtain a well-trained prediction model. The vehicle dynamics data and cockpit interaction data sequence at the time to be predicted are input into the trained prediction model to obtain the cockpit interaction intention prediction result.

2. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 1, characterized in that, Vehicle dynamic data is converted into structured scene semantic information using a large language model. ; Specifically, it includes: First, vehicle dynamics data is converted into scene semantic text S using a large language model fine-tuned by LoRA. Then, the scene semantic text is mapped into high-level semantic text, i.e., structured scene semantic information, through a semantic projection encoder. The semantic projection encoder consists of two fully connected layers and a SiLU activation function.

3. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 1, characterized in that, By utilizing closed-loop continuous units of a liquid neural network to perform continuous-time modeling of cockpit interaction data sequences, temporal features characterizing the dynamic evolution of driver interaction behavior are extracted. Specifically, this includes: First, projecting the cockpit interaction data sequence into the latent space through a linear mapping layer, and superimposing position encoding to form the initial temporal sequence Z; Then, closed continuous units are used to extract features from the initial sequence Z to obtain the temporal hidden state sequence H. Finally, attention pooling is performed on the temporal hidden state sequence H to obtain the global temporal feature H. pooled .

4. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 3, characterized in that, The closed continuous unit traversal extracts features from the initial sequence Z to obtain the temporal hidden state sequence H; specifically including: First, for the initial time series Z, at any time step t, the backbone network receives the current input Z. t Compared with the previous hidden state H t-1 A fused input is formed through feature splicing. ; Then, from the fused input via the backbone network Extracting fusion features ; Based on fusion features Two candidate signals FF1 and FF2, and a time-gating coefficient τ are generated in parallel: ; ; in, and These are the weights of candidate signals FF1 and FF2, respectively. The hyperbolic tangent activation function is used. Indicates the time span between adjacent time steps. and This is the learnable weight matrix in the time-gated branch; This represents the Sigmoid activation function; Finally, based on the time gating coefficient τ, the two candidate signals FF1 and FF2 are dynamically weighted and fused to update the hidden state H at the current time. t ; After traversing the entire initial time sequence Z, the temporal hidden state sequence H is obtained.

5. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 3, characterized in that, Attention pooling is applied to the temporal hidden state sequence H to obtain the global temporal feature H. pooled Specifically, this includes: First, calculate the global query vector Q; Then, the temporal information is aggregated through a multi-head attention mechanism to obtain the attention output. ; Finally, for Perform layer normalization to obtain global temporal features .

6. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 1, characterized in that, The two-level fusion module includes a Semantic Modulation Scaling Fusion (MSF) submodule and a Dynamic Gated Cross-Attention (DGCA) submodule. The DGCA submodule uses the temporal fusion features refined by the MSF submodule. As a query.

7. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 6, characterized in that, The working process of the Semantic Modulation Scaling Fusion (MSF) submodule includes: First, global temporal features Perform semantically guided affine modulation to achieve dynamic recalibration at the feature level: ; Where γ1 and β1 are modulation parameters, and ⊙ denotes element-wise multiplication. Presentation layer normalization operation; Subsequently, the modulated features Perform multi-head self-attention computation and preserve the original timing information through residual connections: ; in, This indicates multi-head self-attention computation; Then, the temporal features are combined using a cross-attention mechanism. With structured scene semantic information To conduct By aligning, we can obtain cross-modal interaction features. : ; in, Indicates alignment operation; Finally, a second modulation and feedforward network transformation is performed to obtain the refined temporal fusion features. : ; Where γ2 and β2 are the second set of modulation parameters, FFN This represents a feedforward network.

8. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 6, characterized in that, The working process of the Dynamic Gated Cross-Attention (DGCA) submodule includes: First, the temporal fusion features refined by the semantic modulation scaling fusion MSF submodule are used. As a query, it uses structured scene semantic information Using these as keys and values, semantically guided cross-attention outputs are computed. ; Subsequently, learnable gating parameters were introduced. Dynamic gating weights are generated using the Sigmoid function. Adaptive adjustment of cross-attention output Features of temporal fusion Contribution percentage in the integration process: ; in, This represents element-wise multiplication; Finally, the fusion features The input residual connection and layer normalization module are processed by a feedforward network for nonlinear transformation to obtain the final multimodal fusion features. .

9. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 1, characterized in that, Multimodal fusion features Mapping to the corresponding interaction category, the driver's interaction intent prediction result is output; specifically including: First, multimodal fusion features along the time dimension. Pooling is performed to obtain the global feature vector. ; Then, the global feature vector The input is a fully connected layer, and after normalization by the Softmax activation function, the predicted probability distributions for each category are obtained; Finally, the category corresponding to the highest predicted probability is selected as the final intention prediction result.

10. The cockpit interaction intent prediction method based on liquid network and large language collaboration according to claim 1, characterized in that, During the model training phase, the model parameters are optimized using a cross-entropy loss function that incorporates a label smoothing strategy. The loss function L is expressed as: Where K is the total number of categories, p k Let ỹ be the predicted probability of the model in the k-th class. k This represents the target label distribution after label smoothing.