Multi-modal sentiment analysis method for vehicle driver state detection

By integrating multimodal data with graph neural networks and network communication dynamics theory, the problem of incomplete emotional feature extraction in multimodal sentiment analysis is solved, accurate detection of the driver's emotional state and personalized interaction are achieved, and driving safety and user experience are improved.

CN120611344APending Publication Date: 2025-09-09CHONGQING UNIV OF POSTS & TELECOMM
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510737117.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-04
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

Existing technologies find it difficult to effectively integrate emotional information in multimodal data, resulting in incomplete emotional feature extraction and inaccurate capture of emotional associations in driver state detection.

Method used

Adopting the graph neural network and network communication dynamics theory in deep learning, through the dynamic routing node module and feature state transfer module, we construct high-level semantic features and deeply interact information between different modalities, and design a fusion classification module and loss function to stabilize emotional associations.

Benefits of technology

It achieves more accurate detection of the driver's emotional state, improves driving safety and in-car human-computer interaction experience, and enhances the system's personalized response capabilities.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120611344A_ABST
    Figure CN120611344A_ABST
Patent Text Reader

Abstract

The invention relates to a multi-mode sentiment analysis method for vehicle driver state detection, and belongs to the field of natural language processing. The method comprises the following steps: 1) processing an original data sample, and separating a text mode, a voice mode and a video mode from the original data sample; 2) constructing a dynamic routing node transfer module so as to convert the low-level semantic features into high-level semantic features; 3) designing a feature state transition module, and deepening deep interaction between different modes to obtain emotion association; 4) designing a fusion classification module and a loss function, wherein the loss function is used for updating model training parameters; and 5) verifying the effectiveness of the proposed method by comparing with various existing methods. According to the method, the problems that advanced semantic features cannot be effectively captured and utilized and deep interaction of different modal data cannot be deepened in an existing method can be solved, and the emotion recognition accuracy of the multi-modal emotion analysis model is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to multimodal sentiment analysis in the field of natural language processing, and relates to a multimodal sentiment analysis method for vehicle driver status detection. Background Art

[0002] With the rapid development of technologies such as artificial intelligence and intelligent driving, driving safety and driver status monitoring have become key research directions in intelligent transportation systems. Especially with the increasing popularity of intelligent assisted driving and autonomous driving, drivers remain an important and indispensable participant in the current transportation system, and their emotions and fatigue have a direct impact on driving safety. In recent years, traffic accidents caused by subjective factors such as driver fatigue, distraction, and emotional distress have become commonplace. Therefore, how to accurately perceive the driver's emotions and status in real time has become one of the core issues that urgently need to be addressed in intelligent driving technology. To achieve this goal, in-depth research on key emotion perception technologies for driver status detection is particularly important, especially the application of multimodal sentiment analysis, which is the key to this field.

[0003] Multimodal emotion analysis methods for vehicle driver status detection aim to integrate multi-source data, including speech, facial expressions, body movements, and physiological signals, to achieve comprehensive perception and intelligent analysis of the driver's mood, fatigue, and stress. Unlike traditional monitoring methods that rely on a single signal source, multimodal emotion analysis combines heterogeneous information from different sensors, processing and analyzing it through deep learning, machine learning, and multimodal fusion algorithms to extract useful emotional features. For example, by analyzing the driver's facial expressions, it is possible to identify the driver's happiness, anger, sadness, and joy; by analyzing the driver's speech, it is possible to identify the driver's emotional state, such as excitement or frustration.

[0004] Multimodal sentiment analysis is a technology that combines multiple modal data, such as text, voice, and video, to detect and analyze emotions. Multimodal sentiment analysis plays a crucial role in detecting the driver's emotional state. Because human emotional expression is multimodal, multimodal sentiment analysis can more accurately capture and understand the driver's emotional state by integrating information from multiple modalities. For example, if a driver says "I'm happy" accompanied by a smile and a cheerful tone of voice, multimodal sentiment analysis can more accurately determine the driver's true emotional state. In specific scenarios, in-vehicle systems can use cameras, microphones, and wearable devices to collect multimodal data in real time, including the driver's facial expressions, voice intonation, and textual information separated from the speech.

[0005] In the practical application of vehicle driver status detection, multimodal sentiment analysis can be used in multiple areas. First, in the precise identification and proactive intervention of emotional states, by integrating multimodal data such as text information, facial expressions, and voice features, it helps effectively distinguish different driver emotions such as anger, anxiety, and frustration, thereby supporting risk prediction and driving safety management. Second, in the in-vehicle human-computer interaction experience, by identifying the driver's emotional fluctuations, the in-vehicle system can dynamically implement humanized and personalized interaction methods. When the driver is depressed or stressed, the system can play soothing music at the right time to enhance user stickiness and driving pleasure.

[0006] However, practical applications present a complex and crucial challenge: the close and complex correlations between emotional information across modalities. This correlation is reflected in the fact that descriptions in the text modality may echo the intonation and rhythm of the speech modality, while facial expressions and body language in the video modality may complement the former two. This multimodal synergy makes it difficult to fully capture high-level semantic features and deep emotional connections using feature extraction techniques alone. To address this issue, we can try to use graph neural networks and network communication dynamics theory from deep learning to organically integrate information from different modalities, thereby extracting more comprehensive and accurate emotional features and identifying emotional connections between different modalities. Summary of the Invention

[0007] In view of this, the object of the present invention is to provide a multimodal emotion analysis method for vehicle driver status detection.

[0008] In order to achieve the above object, the present invention provides the following technical solutions:

[0009] A multimodal sentiment analysis method for vehicle driver state detection includes the following steps:

[0010] Step 1: Process the original data sample to separate the text modality, voice modality, and video modality;

[0011] Step 2: Build a dynamic routing node module to convert low-level semantic features into high-level semantic features;

[0012] Step 3: Design a feature state transfer module to deepen the interaction between different modalities to obtain emotional associations;

[0013] Step 4: Design the fusion classification module and loss function;

[0014] Step 5: Verify the effectiveness of the proposed method by comparing it with multiple existing methods.

[0015] Optionally, the specific process of step 1 includes: Since the original data is mixed with some redundant information irrelevant to the required modal information, we first perform some preliminary processing on the original data to accurately separate and extract the information of the three modalities of text, voice and video. The three input modalities are defined as follows: M = {T (text), A (audio), V (visual)}, then {X T ,X A ,X V}, m∈(A,V), In the i-th step, the emotional features extracted from modality m are expressed as follows: m is the dimension of the modal feature, and L m is the sequence length of the modality m; is the original sequence of words.

[0016] Optionally, the specific process of step 2 includes: dynamic routing node module: using a specific modal encoder to further extract relevant emotional features:

[0017] E m =f m (X m )

[0018] where m∈M, f m (·) represents the encoder of mode m, is the sequential representation of the modality encoding. For the speech and video modalities, two long short-term memory networks are selected as the encoders for the speech and video modalities. For the text modality, BERT is used as the text modality encoder due to its excellent language representation capabilities and wide utilization.

[0019] The previous method was to obtain E m Then directly construct E m The unimodal graph of nodes ignores the correlation between high-level semantic features. Constructing high-level nodes can capture high-level semantic features. Dynamic routing in capsule networks allows for efficient output of high-level capsules through coordination and communication between low-level capsules, enabling them to represent deeper, high-level semantic features. Therefore, dynamic routing is used to construct nodes from the output sequence. The encoded sequence can be represented as a capsule, and then the prediction vector can be constructed:

[0020]

[0021] where m∈M, Represents the prediction vector of the i-th capsule, used to construct the j-th node, are trainable parameters. The nodes are then defined by the weighted sum of the corresponding prediction vectors:

[0022]

[0023] in represents the embedding of the j-th node, Represents the vector assigned to the prediction During a total of k iterations, all routing coefficients are normalized using the softmax function and iteratively updated based on the inner product between the prediction vector and the node in each iteration step:

[0024]

[0025] in Represents the routing coefficient before normalization, which is initialized to zero before the iteration starts. After k iterations, the output capsule is obtained. As nodes in a unimodal graph, the output capsule can be viewed as the cluster center of the associated input capsules, possessing strong semantic representation capabilities. By capturing the correlations between these nodes rich in strong semantic features, highly abstract sentiment structures can be more accurately represented. Compared to clustering techniques that rely solely on data similarity and struggle to capture complex hierarchical structures, capsule networks can hierarchically represent information at different semantic levels, thereby obtaining high-level semantic features. The high-level semantic nodes derived from these high-level semantic features in the graph can more effectively capture the overall structure and relationships.

[0026] For the N nodes obtained by dynamic routing, this paper designs a learnable adjacency matrix O m , in order to adaptively learn the relationships between nodes in a manner similar to the self-attention mechanism.

[0027]

[0028] where m∈M, is the designed adjacency matrix, is a learnable parameter of the adjacency matrix; ReLU(x) = max(0,x) represents the activation function. The graph attention network then aggregates information from other nodes to update the nodes in the graph, ultimately generating a unimodal graph representation.

[0029] Optionally, the specific process of step three includes: feature state transfer module: Since the video feature V and the voice feature A are less robust than the text feature T, they are more susceptible to noise interference. In order to promote deep interaction and efficient fusion between subsequent features, the same structure is used to process the video and voice features input to the feature state transfer module. In the actual design, in order to avoid differences in the acquired features due to different structures, which is not conducive to the interaction between modalities, two cross-modal multi-head attention mechanisms are used to interact with text, video, and voice modalities. One of the cross-modal multi-head attention mechanisms is used to obtain the interactive feature V T and A T , and another one is used to obtain pseudo video features and pseudo-speech features

[0030] F T =MultiAtt(T,F,F;w F ),F∈(A,V)

[0031]

[0032] where w F and is a trainable parameter. In addition, to ensure that the initially established sentiment association information between different modalities is stable and reliable and to prevent instability due to the lack of effective constraints, this paper introduces the bulldozer distance as a loss function to constrain the aforementioned conversion process.

[0033] In the specific implementation of feature state transfer, virus features and several other state features need to be input to achieve the state transfer process through the state record matrix. Referring to the SIR model, and in order to reduce the negative impact that multimodal data misalignment may have on the state transfer process and subsequent analysis, we have made clear provisions for virus features and several other state features as follows:

[0034] (1) Viral features: The average value of the text modal feature T is obtained

[0035] (2) Initial state features: The video feature V and the speech feature A are averaged to obtain and

[0036] (3) Infection status characteristics: interaction characteristics V T and A T After taking the mean, we get and

[0037] (4) Restoring state features: for pseudo video features and pseudo-speech features After taking the mean, we get and

[0038] In the infection state, we first determine the position in the initial state that is susceptible to infection by the virus feature T based on the contact rate α. Then, we use the random matrix H rand The infection rate β is used to judge these susceptible locations, determine the truly infected parts, and then generate the infection status record matrix. The specific formula is as follows:

[0039]

[0040] in The infection state record matrix used to record the part of the initial state characteristics that are transformed into infection state characteristics after infection; H contact It is used to locate those initial state features that are susceptible to virus features The affected sensitive locations. abs(·) is used to take the modulus of each element in the tensor. To address the problem of subsequent training caused by the presence of 0 in the feature tensor, a matrix J with all 1s is introduced to obtain the initial state record matrix to record the locations in the initial state features that still retain valid information after infection, namely:

[0041]

[0042] in Represents the initial state characteristics after infection or The initial state record matrix of valid information positions.

[0043] In the recovery state transfer process, the contact rate α is used to determine the position where the infection state characteristics are easily transformed into the recovery state characteristics information. Then the random matrix is ​​introduced and recovery rate ρ, from the infection state record matrix The part that obtains the characteristic information of the restored state is specifically shown in the formula:

[0044]

[0045] in Is the record recovery status feature and The valid information position in It is responsible for recording the location information of those transitions from infection state characteristics to recovery state characteristics.

[0046] In addition, given that Some of the infection status characteristic information recorded in will be converted into The recovery status characteristic information recorded in Update, see the formula for details:

[0047]

[0048] in express Updated infection status record matrix.

[0049] After completing the above two state transition processes, the public sentiment features are obtained by integrating the effective information of different state features recorded by the state record matrix. The specific formula is as follows:

[0050]

[0051] Among them F share represents the public sentiment feature, and ⊙ represents the Hadamard product.

[0052] Optionally, in step 4, the classification module and loss function are integrated: the features obtained after processing by the dynamic routing node module are integrated with the features after state transfer, and then sent to the fully connected layer to obtain the final output. In order to obtain stable emotional association during cross-modal translation, this process is constrained. The specific formula is as follows:

[0053] L F =wass(F T ,T),F∈(V,A)

[0054] where wass(·) is the bulldozer distance calculation function, and function L F It is used to constrain this process, including the process of obtaining interactive features and obtaining pseudo video features and pseudo voice features. Finally, the loss value of the constraint process and the predicted loss value L are calculated. pred , the overall loss is defined as:

[0055] Loss = L V +L A +L pred

[0056] Among them L pred Is the mean absolute error function, and finally the overall learning of the model is performed by minimizing Loss.

[0057] The beneficial effects of the present invention are: proposing a multimodal sentiment analysis method for vehicle driver state detection, including: 1) processing the original data samples to separate the text modality, voice modality and video modality; 2) constructing a dynamic routing node module and a unimodal graph to convert low-level semantic features into high-level semantic features; 3) designing a feature state transfer module to deepen the deep interaction between different modalities to obtain sentiment association; 4) designing a fusion classification module and a loss function.

[0058] Other advantages, objects, and features of the present invention will be described in part in the following description and, in part, will be apparent to those skilled in the art upon examination of the following description or may be learned from practice of the present invention. The objects and other advantages of the present invention may be realized and obtained through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] In order to make the purpose, technical solutions and advantages of the present invention more clear, the present invention will be described in detail below with reference to the accompanying drawings, in which:

[0060] Figure 1 Overall process;

[0061] Figure 2 An overall framework diagram of a multimodal sentiment analysis method for vehicle driver status detection;

[0062] Figure 3 Dynamic routing node module structure diagram;

[0063] Figure 4 Feature state transfer module structure diagram. DETAILED DESCRIPTION

[0064] The following describes the embodiments of the present invention by means of specific examples, and those skilled in the art can easily understand other advantages and effects of the present invention from the contents disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present invention. It should be noted that the illustrations provided in the following embodiments are only schematic illustrations of the basic concept of the present invention, and the following embodiments and features in the embodiments can be combined with each other without conflict.

[0065] Among them, the accompanying drawings are only for illustrative purposes and represent only schematic diagrams rather than actual pictures, and should not be understood as limiting the present invention. In order to better illustrate the embodiments of the present invention, some parts of the accompanying drawings may be omitted, enlarged or reduced, and do not represent the dimensions of actual products. For those skilled in the art, it is understandable that some well-known structures and their descriptions may be omitted in the accompanying drawings.

[0066] The same or similar numbers in the drawings of the embodiments of the present invention correspond to the same or similar parts; in the description of the present invention, it should be understood that if there are terms such as "upper", "lower", "left", "right", "front", "back", etc. indicating directions or positional relationships, they are based on the directions or positional relationships shown in the drawings. They are only for the convenience of describing the present invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific direction, be constructed and operate in a specific direction. Therefore, the terms describing the positional relationship in the drawings are only used for illustrative purposes and cannot be understood as limiting the present invention. For ordinary technicians in this field, the specific meanings of the above terms can be understood according to specific circumstances.

[0067] See Figures 1 to 4 The present invention provides a multimodal emotion analysis method for vehicle driver status detection. Figure 1 The following is a flowchart for the specific implementation. Figure 2 This is the overall structure diagram of the present invention. Figure 3 、 Figure 4 They are Figure 2 The detailed structure of the module is described below with reference to the accompanying drawings, including the following steps:

[0068] Optionally, step 1 specifically includes: Since the original data contains some redundant information irrelevant to the required modal information, we first perform some preliminary processing on the original data to accurately separate and extract the information of the three modalities of text, voice, and video. The three input modalities are defined as follows: M = {T (text), A (audio), V (visual)}, then {X T ,X A ,X V}, m∈(A,V), In the i-th step, the emotional features extracted from modality m are expressed as follows: m is the dimension of the modal feature, and L m is the sequence length of the modality m; is the original sequence of words.

[0069] Optionally, step 2 specifically includes: Figure 2 This is the overall structure diagram of the present invention, which mainly includes the dynamic routing node module Figure 3 , Feature State Transfer Module Figure 4 The network structure of the proposed method is as follows Figure 2 The main structure is the dynamic routing node module and the feature state transfer module.

[0070] Dynamic routing node module: Use specific modality encoders to further extract relevant sentiment features:

[0071] E m =f m (X m )

[0072] where m∈M, f m (·) represents the encoder of mode m, is the sequential representation of the modality encoding. For the speech and video modalities, two long short-term memory networks are selected as the encoders for the speech and video modalities. For the text modality, BERT is used as the text modality encoder due to its excellent language representation capabilities and wide utilization.

[0073] The previous method was to obtain E m Then directly construct E m The unimodal graph of nodes ignores the correlation between high-level semantic features. Constructing high-level nodes can capture high-level semantic features. Dynamic routing in capsule networks allows for efficient output of high-level capsules through coordination and communication between low-level capsules, enabling them to represent deeper, high-level semantic features. Therefore, dynamic routing is used to construct nodes from the output sequence. The encoded sequence can be represented as a capsule, and then the prediction vector can be constructed:

[0074]

[0075] where m∈M, Represents the prediction vector of the i-th capsule, used to construct the j-th node, are trainable parameters. The nodes are then defined by the weighted sum of the corresponding prediction vectors:

[0076]

[0077] in represents the embedding of the j-th node, Represents the vector assigned to the prediction During a total of k iterations, all routing coefficients are normalized using the softmax function and iteratively updated based on the inner product between the prediction vector and the node in each iteration step:

[0078]

[0079] in Represents the routing coefficient before normalization, which is initialized to zero before the iteration starts. After k iterations, the output capsule is obtained. As nodes in a unimodal graph, the output capsule can be viewed as the cluster center of the associated input capsules, possessing strong semantic representation capabilities. By capturing the correlations between these nodes rich in strong semantic features, highly abstract sentiment structures can be more accurately represented. Compared to clustering techniques that rely solely on data similarity and struggle to capture complex hierarchical structures, capsule networks can hierarchically represent information at different semantic levels, thereby obtaining high-level semantic features. The high-level semantic nodes derived from these high-level semantic features in the graph can more effectively capture the overall structure and relationships.

[0080] For the N nodes obtained by dynamic routing, this paper designs a learnable adjacency matrix O m , in order to adaptively learn the relationships between nodes in a manner similar to the self-attention mechanism.

[0081]

[0082] where m∈M, is the designed adjacency matrix, is a learnable parameter of the adjacency matrix; ReLU(x) = max(0,x) represents the activation function. The graph attention network then aggregates information from other nodes to update the nodes in the graph, ultimately generating a unimodal graph representation.

[0083] Feature state module: Since the video feature V and the speech feature A are less robust than the text feature T, they are more susceptible to noise interference. In order to promote deep interaction and efficient fusion between subsequent features, the same structure is used to process the video and speech features input to the feature state transfer module. In the actual design, in order to avoid differences in the acquired features due to different structures, which is not conducive to the interaction between modalities, two cross-modal multi-head attention mechanisms are used to interact with text, video, and speech modalities. One of the cross-modal multi-head attention mechanisms is used to obtain the interactive feature V T and A T , and another one is used to obtain pseudo video features and pseudo-speech features

[0084] F T =MultiAtt(T,F,F;w F ),F∈(A,V)

[0085]

[0086] where w F and is a trainable parameter. In addition, to ensure that the initially established sentiment association information between different modalities is stable and reliable and to prevent instability due to the lack of effective constraints, this paper introduces the bulldozer distance as a loss function to constrain the aforementioned conversion process.

[0087] In the specific implementation of feature state transfer, virus features and several other state features need to be input to achieve the state transfer process through the state record matrix. Referring to the SIR model, and in order to reduce the negative impact that multimodal data misalignment may have on the state transfer process and subsequent analysis, we have made clear provisions for virus features and several other state features as follows:

[0088] (1) Viral features: The average value of the text modal feature T is obtained

[0089] (2) Initial state features: The video feature V and the speech feature A are averaged to obtain and

[0090] (3) Infection status characteristics: interaction characteristics V T and A T After taking the mean, we get and

[0091] (4) Restoring state features: for pseudo video features and pseudo-speech features After taking the mean, we get and

[0092] In the infection state, we first determine the position in the initial state that is susceptible to infection by the virus feature T based on the contact rate α. Then, we use the random matrix H rand The infection rate β is used to judge these susceptible locations, determine the truly infected parts, and then generate the infection status record matrix. The specific formula is as follows:

[0093]

[0094] in The infection state record matrix used to record the part of the initial state characteristics that are transformed into infection state characteristics after infection; H contact It is used to locate those initial state features that are susceptible to virus features The affected sensitive locations. abs(·) is used to take the modulus of each element in the tensor. To address the problem of subsequent training caused by the presence of 0 in the feature tensor, a matrix J with all 1s is introduced to obtain the initial state record matrix to record the locations in the initial state features that still retain valid information after infection, namely:

[0095]

[0096] in Represents the initial state characteristics after infection or The initial state record matrix of valid information positions.

[0097] In the recovery state transfer process, the contact rate α is used to determine the position where the infection state characteristics are easily transformed into the recovery state characteristics information. Then the random matrix is ​​introduced and recovery rate ρ, from the infection state record matrix The part that obtains the characteristic information of the restored state is specifically shown in the formula:

[0098]

[0099] in Is the record recovery status feature and The valid information position in It is responsible for recording the location information of those transitions from infection state characteristics to recovery state characteristics.

[0100] In addition, given that Some of the infection status characteristic information recorded in will be converted into The recovery status characteristic information recorded in Update, see the formula for details:

[0101]

[0102] in express Updated infection status record matrix.

[0103] After completing the above two state transition processes, the public sentiment features are obtained by integrating the effective information of different state features recorded by the state record matrix. The specific formula is as follows:

[0104]

[0105] Among them F share represents the public sentiment feature, and ⊙ represents the Hadamard product.

[0106] Fusion Classification Module and Loss Function: The features obtained after processing by the dynamic routing node module are fused with the features after state transfer and sent to the fully connected layer to obtain the final output. In order to obtain stable emotional association during cross-modal translation, this process is constrained. The specific formula is as follows:

[0107] L F =wass(F T ,T),F∈(V,A)

[0108] where wass(·) is the bulldozer distance calculation function, and function L F It is used to constrain this process, including the process of obtaining interactive features and obtaining pseudo video features and pseudo voice features. Finally, the loss value of the constraint process and the predicted loss value L are calculated. pred , the overall loss is defined as:

[0109] Loss = L V +L A +L pred

[0110] Among them L pred Is the mean absolute error function, and finally the overall learning of the model is performed by minimizing Loss.

Claims

1. A multimodal sentiment analysis method for vehicle driver status detection, characterized by: The method comprises the following steps: Step 1: Process the original data sample to separate the text modality, voice modality, and video modality; Step 2: Build a dynamic routing node module to convert low-level semantic features into high-level semantic features; Step 3: Design a feature state transfer module to deepen the interaction between different modalities to obtain emotional associations; Step 4: Design the fusion classification module and loss function; Step 5: Verify the effectiveness of the proposed method by comparing it with multiple existing methods.

2. The multimodal emotion analysis method for vehicle driver status detection according to claim 1, characterized in that: In step 1, since the original data contains some redundant information irrelevant to the required modal information, we first perform some preliminary processing on the original data to accurately separate and extract the information of the three modalities: text, voice, and video. The three input modalities are defined as follows: M = {T (text), A (audio), V (visual)}, then {X T ,X A ,X V }, m∈(A,V), In the i-th step, the emotional features extracted from modality m are expressed as follows: m is the dimension of the modal feature, and L m is the sequence length of the modality m; is the original sequence of words.

3. The multimodal emotion analysis method for vehicle driver status detection according to claim 2, characterized in that: In step 2, the dynamic routing node module uses a specific modality encoder to further extract relevant emotional features: E m =f m (X m ) where m∈M, f m (·) represents the encoder of mode m, is the sequential representation of the modality encoding. For the speech and video modalities, two long short-term memory networks are selected as the encoders for the speech and video modalities. For the text modality, BERT is used as the text modality encoder due to its excellent language representation capabilities and wide utilization. The previous method was to obtain E m Then directly construct E m The unimodal graph of nodes ignores the correlation between high-level semantic features. Constructing high-level nodes can capture high-level semantic features. Dynamic routing in capsule networks allows for efficient output of high-level capsules through coordination and communication between low-level capsules, enabling them to represent deeper, high-level semantic features. Therefore, dynamic routing is used to construct nodes from the output sequence. The encoded sequence can be represented as a capsule, and then the prediction vector can be constructed: where m∈ M, Represents the prediction vector of the i-th capsule, used to construct the j-th node, are trainable parameters. The nodes are then defined by the weighted sum of the corresponding prediction vectors: in represents the embedding of the j-th node, Represents the vector assigned to the prediction During a total of k iterations, all routing coefficients are normalized using the softmax function and iteratively updated based on the inner product between the prediction vector and the node in each iteration step: in Represents the routing coefficient before normalization, which is initialized to zero before the iteration starts. After k iterations, the output capsule is obtained. As nodes in a unimodal graph, the output capsule can be viewed as the cluster center of the associated input capsules, possessing strong semantic representation capabilities. By capturing the correlations between these nodes rich in strong semantic features, highly abstract sentiment structures can be more accurately represented. Compared to clustering techniques that rely solely on data similarity and struggle to capture complex hierarchical structures, capsule networks can hierarchically represent information at different semantic levels, thereby obtaining high-level semantic features. The high-level semantic nodes derived from these high-level semantic features in the graph can more effectively capture the overall structure and relationships. For the N nodes obtained by dynamic routing, this paper designs a learnable adjacency matrix O m , in order to adaptively learn the relationships between nodes in a manner similar to the self-attention mechanism. where m∈ M, is the designed adjacency matrix, is a learnable parameter of the adjacency matrix; ReLU(x) = max(0,x) represents the activation function. The graph attention network then aggregates information from other nodes to update the nodes in the graph, ultimately generating a unimodal graph representation.

4. The multimodal emotion analysis method for vehicle driver status detection according to claim 3, characterized in that: In the step three, the feature state transfer module: Since the video feature V and the speech feature A are less robust than the text feature T, they are more susceptible to noise interference. In order to promote deep interaction and efficient fusion between subsequent features, the same structure is used to process the video and speech features input to the feature state transfer module. In the actual design, in order to avoid differences in the acquired features due to different structures, which is not conducive to the interaction between modalities, two cross-modal multi-head attention mechanisms are used to interact with text, video, and speech modalities. One of the cross-modal multi-head attention mechanisms is used to obtain the interactive feature V T and A T , and another one is used to obtain pseudo video features and pseudo-speech features F T =MultiAtt(T,F,F;w F ),F∈(A,V) where w F and is a trainable parameter. In addition, to ensure that the initially established sentiment association information between different modalities is stable and reliable and to prevent instability due to the lack of effective constraints, this paper introduces the bulldozer distance as a loss function to constrain the aforementioned conversion process. In the specific implementation of feature state transfer, virus features and several other state features need to be input to achieve the state transfer process through the state record matrix. Referring to the SIR model, and in order to reduce the negative impact that multimodal data misalignment may have on the state transfer process and subsequent analysis, we have made clear provisions for virus features and several other state features as follows: (1) Viral features: The average value of the text modal feature T is obtained (2) Initial state features: The video feature V and the speech feature A are averaged to obtain and (3) Infection status characteristics: interaction characteristics V T and A T After taking the mean, we get and (4) Restoring state features: for pseudo video features and pseudo-speech features After taking the mean, we get and In the infection state, we first determine the position in the initial state that is susceptible to infection by the virus feature T based on the contact rate α. Then we use the random matrix H rand The infection rate β is used to judge these susceptible locations, determine the truly infected parts, and then generate the infection status record matrix. The specific formula is as follows: st in The infection state record matrix used to record the part of the initial state characteristics that are transformed into infection state characteristics after infection; H contact It is used to locate those initial state features that are susceptible to virus features The affected sensitive locations. abs(·) is used to take the modulus of each element in the tensor. To address the problem of subsequent training caused by the presence of 0 in the feature tensor, a matrix J with all 1s is introduced to obtain the initial state record matrix to record the locations in the initial state features that still retain valid information after infection, namely: in Represents the initial state characteristics after infection or The initial state record matrix of valid information positions. In the recovery state transfer process, the contact rate α is used to determine the position where the infection state characteristics are easily transformed into the recovery state characteristics information. Then the random matrix is ​​introduced and recovery rate ρ, from the infection state record matrix The part that obtains the characteristic information of the restored state is specifically shown in the formula: st in Is the record recovery status feature and The valid information position in It is responsible for recording the location information of those transitions from infection state characteristics to recovery state characteristics. In addition, given that Some of the infection status characteristic information recorded in will be converted into The recovery status characteristic information recorded in Update, see the formula for details: in express Updated infection status record matrix. After completing the above two state transition processes, the public sentiment features are obtained by integrating the effective information of different state features recorded by the state record matrix. The specific formula is as follows: Among them F share represents the public sentiment feature, and ⊙ represents the Hadamard product.

5. The multimodal emotion analysis method for vehicle driver status detection according to claim 4, characterized in that: In step 4, the classification module and loss function are integrated: the features obtained after processing by the dynamic routing node module are integrated with the features after state transfer, and then sent to the fully connected layer to obtain the final output. In order to obtain stable emotional association during cross-modal translation, this process is constrained. The specific formula is as follows: L F =wass(F T ,T),F∈(V,A) where wass(·) is the bulldozer distance calculation function, and function L F It is used to constrain this process, including the process of obtaining interactive features and obtaining pseudo video features and pseudo voice features. Finally, the loss value of the constraint process and the predicted loss value L are calculated. pred , the overall loss is defined as: Loss=L V +L A +L pred Among them L pred Is the mean absolute error function, and finally the overall learning of the model is performed by minimizing Loss.

6. The multimodal emotion analysis method for vehicle driver status detection according to claim 5, characterized in that: In step 5, the effectiveness of the proposed method is verified by comparing it with multiple existing methods. We conduct experiments on the CMU-MOSI and CMU-MOSEI datasets and verify its effectiveness by comparing it with the current state-of-the-art methods.