Multi-modal spatio-temporal alignment safe driving emotion recognition method based on sensitive word guidance

By using a mechanism that proactively guides responses to sensitive words, the analysis of driver voice and video data solves the problem of spatiotemporal misalignment of multimodal data in driving scenarios, enabling accurate identification of driver emotions and improving the real-time performance and accuracy of safe driving.

CN121188535BActive Publication Date: 2026-05-22HEFEI UNIV OF TECH
View PDF 5 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
HEFEI UNIV OF TECH
Filing Date
2025-08-22
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies struggle to effectively address the spatiotemporal misalignment of multimodal data in driving scenarios, resulting in insufficient accuracy and real-time performance in emotion recognition. In particular, the asynchronous nature of voice-sensitive words and visual facial emotion expression in time and space has not been fully utilized.

Method used

By using a mechanism that actively guides sensitive words, this method extracts sensitive word features from driving monitoring video and voice data, analyzes the temporal attention between sensitive words and voice text, determines multiple candidate time intervals, and analyzes the alignment relationship between multi-spatial region features of facial components and voice sensitive words within key time frames, thereby achieving spatiotemporally aligned multimodal emotion recognition.

Benefits of technology

It improves the accuracy and reliability of driver emotion recognition, meets the real-time and accurate requirements of safe driving for emotion recognition, and reduces misjudgments caused by spatiotemporal misalignment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121188535B_ABST
    Figure CN121188535B_ABST
Patent Text Reader

Abstract

The application discloses a kind of multi-modal space-time alignment safe driving emotion recognition methods based on sensitive word guide.The application extracts sensitive word features by driving monitoring video and voice data.With the analysis sensitive word-voice text time attention, determine the multiple candidate time intervals that possibly exist emotion in video.With the sensitive word-visual time cross attention, adaptively determine the key time frame that exists driving emotion.In key time frame, extract component features in the multiple space regions of human face component, analyze its alignment relationship with voice sensitive word, with the sensitive word-visual space cross attention, adaptively determine the key space region that exists driving emotion.The application fully considers the space-time asynchronization of sensitive word in multi-modal data, sequentially analyzes candidate time interval, key time frame, key space region, gradually finds the key clue of driving emotion, and can effectively improve the accuracy of safe driving emotion recognition.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of safe driving technology, and in particular to a multimodal spatiotemporal alignment method for safe driving emotion recognition based on sensitive words. Background Technology

[0002] With the rapid development of intelligent driving systems, driver emotion recognition has become an important research direction for ensuring driving safety. Driver emotional fluctuations during driving, such as anger, anxiety, and fatigue, can significantly affect driving behavior and decision-making abilities, thereby increasing the risk of traffic accidents. Existing technologies mainly focus on single-modal emotion recognition, such as facial expression recognition or voice emotion analysis, but these methods often struggle to accurately capture the complex and ever-changing emotional expressions in driving scenarios. Furthermore, although some systems attempt to combine multimodal data for emotion recognition, significant challenges remain in the accuracy and real-time performance of data fusion and spatiotemporal alignment.

[0003] Patent CN 114399818 A, entitled "Multimodal Facial Emotion Recognition Method and Device," proposes a facial emotion recognition method combining visual and speech modalities. This method improves emotion recognition efficiency by collecting data from real-world scenarios and incorporating an attention mechanism for temporal learning. However, this technique primarily focuses on recognizing and fusing data from single frames, resulting in low efficiency and frequent inconsistencies in the results. More importantly, this method fails to consider the spatiotemporal asynchrony between sensitive words and visual emotional expressions in driving scenarios, making it difficult to effectively address the unique emotion recognition needs of driving environments.

[0004] The patent CN 116189669 A, entitled "A Multimodal Emotion Recognition Method and System," integrates multiple modal information such as speech, text, and visual expressions, improving the accuracy of emotion recognition through multimodal information collection, encoding, and relevance analysis. However, this method does not specifically optimize for driving scenarios, lacks attention to sensitive words in driving situations, and does not establish a clear spatiotemporal alignment mechanism, thus limiting its application in real-world driving environments.

[0005] The patent CN 116052291 A, entitled "A Multimodal Emotion Recognition Method Based on Unaligned Sequences," addresses the problem of unaligned multimodal sequences by employing a cross-modal Transformer module to fuse features from different modalities. However, while this method considers the unalignment issue between modalities, it is not specifically designed for driving scenarios and does not fully utilize sensitive word features as key guiding information for emotion recognition, making it difficult to accurately identify specific emotional changes during driving.

[0006] Emotion recognition in driving scenarios faces unique challenges: First, drivers' emotional expression is often limited by the driving task, and facial expressions may not be obvious; second, there is often a spatiotemporal asynchrony between verbal and visual emotional cues in the driving environment. For example, a driver may express emotion verbally (such as "That's terrible") before a clear facial emotional change appears. Existing technologies have failed to effectively address this spatiotemporal misalignment problem, resulting in insufficient accuracy and real-time performance in emotion recognition. Therefore, there is an urgent need for a method that can accurately capture multimodal emotional information in driving situations and solve the spatiotemporal alignment problem to improve driving safety. Summary of the Invention

[0007] The purpose of this invention is to overcome the shortcomings of existing technologies and provide a multimodal spatiotemporal alignment-based method for safe driving emotion recognition based on sensitive words. This invention specifically addresses the spatiotemporal misalignment problem of multimodal data in driving scenarios, namely the asynchrony between voice-based sensitive words and visual facial emotion expressions in the temporal and spatial dimensions. This invention utilizes a sensitive word-guided mechanism, leveraging driving monitoring video and voice data to extract sensitive word features, analyze the temporal attention of sensitive words and voice text, and determine multiple candidate time intervals where emotions may exist in the video. Next, the alignment relationship between facial features and voice-based sensitive words in multiple time intervals is analyzed to identify key time frames where emotions exist. Within these key time frames, the alignment relationship between multi-spatial region features of facial components and voice-based sensitive words is further analyzed to identify key spatial regions where emotions exist, ultimately achieving accurate recognition of the driver's emotional state and providing timely alerts when abnormal emotions are detected to ensure driving safety.

[0008] This invention is achieved through the following technical solution:

[0009] A method for safe driving emotion recognition based on sensitive words and multimodal spatiotemporal alignment includes the following steps: S1: Acquire driver facial expression videos and audio, extract audio text sequences, define a set of sensitive words and driving emotion categories, and construct a multimodal emotion dataset; S2: Calculate sensitive word-guided speech attention based on the position of sensitive words in the audio text sequence, and select a set of candidate key time frames guided by sensitive words; S3: Construct a sensitive word temporal alignment attention model based on the set of candidate key time frames guided by sensitive words, and select a set of key time frames guided by sensitive words; S4: Extract multi-spatial region features based on the set of key time frames guided by sensitive words and the acquired driver facial expression videos, calculate sensitive word spatial alignment attention, and select a set of key spatial regions guided by sensitive words; S5: Input the multimodal emotion dataset, extract key multi-spatial region features and predict emotion categories based on the set of candidate key time frames guided by sensitive words, the set of key time frames guided by sensitive words, and the set of key spatial regions guided by sensitive words obtained in S2-S4, calculate the sensitive word-guided emotion recognition loss function and train it to obtain an emotion recognition model; S6: Based on the emotion recognition model obtained from training, predict the driver's emotion category according to the input driver's facial expression video, audio and the results of S2-S5.

[0010] The specific details of step S1 are as follows:

[0011] S1-1: Collect a multimodal sentiment dataset;

[0012] S1-1-1: Use cameras and microphones to collect video and audio of the driver's facial expressions;

[0013] S1-1-2: Decode the video into a sequence of video frames. , It is the t-th frame of the video image. =1, 2, ..., ;

[0014] S1-1-3: Extract the audio sequence synchronized with the video sequence , It is the audio sample amplitude of the t-th frame. =1, 2, ..., ;

[0015] S1-1-4: Output video frame sequence and audio sequences

[0016] S1-2: Extract the audio text sequence;

[0017] S1-2-1: Input audio sequence

[0018] S1-2-2: Using the Conformer audio recognition model, the audio sequence... Recognized as an audio text sequence , It is the text recognized from the audio of frame t. =1, 2, ..., ;

[0019] S1-2-3: Output audio text sequence ;

[0020] S1-3: Define the sensitive word set: Collect sensitive words that indicate abnormal driver emotions while driving to obtain the sensitive word set. , =1, 2, ..., K, For the first One sensitive word;

[0021] S1-4: Define driving emotion categories : , ;

[0022] S1-5: Annotated multimodal sentiment dataset;

[0023] S1-5-1: Input video frame sequence audio sequence Audio text sequence ;

[0024] S1-5-2: Manually label each video with corresponding emotion tags. , ;

[0025] S1-5-3: Output a labeled video. ;

[0026] S1-5-4: Repeat steps S1-5-1 to S1-5-3 to complete the annotation of all videos and obtain the multimodal emotion dataset.

[0027] The specific details of step S2 are as follows:

[0028] S2-1: Determine the location of the sensitive word frame;

[0029] S2-1-1: Input audio text sequence and audio sequences ;

[0030] S2-1-2: Input the set of sensitive words defined in S1-3 ;

[0031] S2-1-3: Traversing the set of sensitive words All sensitive words in ;

[0032] S2-1-4: Use the kth sensitive word Traverse the audio text sequence All of them ;

[0033] S2-1-5: Obtain the content containing the kth sensitive word of The set of subscripts t, that is, all frame positions t in which the k-th sensitive word appears;

[0034] S2-1-6: Obtain the set of all frame positions where the k-th sensitive word appears. ;

[0035] S2-1-7: In the sensitive word set In the middle, remove the set where the frame position set is empty;

[0036] S2-1-8: Merge all sensitive words to obtain the set of frame positions where all sensitive words appear. ;

[0037] S2-2: Calculate speech attention guided by sensitive words:

[0038] S2-2-1: Set of frame locations where sensitive words appear in the input Audio sequences Audio text sequence ;

[0039] S2-2-2: Based on the set of frame positions where sensitive words appear Given all indices k, obtain the sensitive word matching set. , It is an audio text sequence The sensitive words appearing in the sensitive word set Subscript in;

[0040] S2-2-3: Traverse the sensitive word matching set All the sensitive words inside ;

[0041] S2-2-4: Input sensitive words We use a BERT pre-trained language model to obtain the textual features (FT) of sensitive words;

[0042] S2-2-5: Set of frame locations where all sensitive words appear ;

[0043] S2-2-6: Based on the frame position set In Obtain frame position The corresponding set of audio sample amplitude values ​​is denoted as , , This refers to the audio sampling amplitude.

[0044] S2-2-7: Input audio sample amplitude set ,use The speech recognition model obtains the set of audio sampling amplitudes. In the middle, audio sampling amplitude Corresponding audio features , , This represents the audio feature corresponding to frame position t;

[0045] S2-2-8: Text features extracted from input S2-2-4 The text query features are calculated using a fully connected layer, denoted as... ;

[0046] S2-2-9: Input the audio features FA extracted in S2-2-7, and use a fully connected layer to calculate the speech key features, denoted as... , It is a matrix composed of the audio features of all sensitive words;

[0047] S2-2-10: Calculate speech attention guided by sensitive words.

[0048] ,

[0049] Where d represents the dimension of the text query features, and the sensitive word-guided speech attention reflects the relationship between driving emotion categories and all sensitive words, including attention at multiple sensitive word frame positions. ;

[0050] S2-3: Determine multiple time intervals guided by sensitive words;

[0051] S2-3-1: Set the speech attention threshold ;

[0052] S2-3-2: Input-sensitive word-guided speech attention ;

[0053] S2-3-3: Set of frame locations where sensitive words appear in the input Filter out all speech attention from them corresponding ;

[0054] S2-3-4: Obtain a set of candidate key time frames guided by sensitive words , where t satisfies .

[0055] The specific details of step S3 are as follows:

[0056] S3-1: Extract multi-time interval features guided by sensitive words;

[0057] S3-1-1: Input video frame sequence and a set of candidate key time frames guided by sensitive words ;

[0058] S3-1-2: A set of candidate key time frames guided by sensitive words In Obtain the video frame corresponding to t. Obtain a set of video frames across multiple time intervals guided by sensitive words, denoted as , ,in These are the video frames corresponding to the candidate key time frames;

[0059] S3-1-3: Multi-time-interval video frame set guided by inputting sensitive words ,use Image classification model for video frame sequences Each Obtain the corresponding video features Obtain video feature set ={ , ;

[0060] S3-2: Constructing a sensitive word-time aligned attention model:

[0061] S3-2-1: Input video features FV, ​​and use a fully connected layer to compute video key features. ;

[0062] S3-2-2: Define text query features calculated using S2-2-8 Constructing time-aligned attention guided by sensitive words:

[0063] ,

[0064] Where d represents the dimension of the text query features, and this attention reflects the relationship between driving emotion categories and video frames, including attention to multiple sensitive word frame locations. , This indicates the time-aligned attention guided by the sensitive word at frame position t;

[0065] S3-3: Output key time frames guided by sensitive words:

[0066] S3-3-1: Setting the time-aligned attention threshold ;

[0067] S3-3-2: Timing Alignment Attention Guided by Inputting Each Sensitive Word ;

[0068] S3-3-3: Set of frame locations where sensitive words appear in the input Filter out all speech attention from them Corresponding key time frames ;

[0069] S3-3-4: Obtain the set of key time frames guided by sensitive words , where t satisfies .

[0070] The specific details of step S4 are as follows:

[0071] S4-1: Input video frame sequence and key time frame set ,get middle Corresponding key video frames , gather , ,in These are the video frames corresponding to key time frames.

[0072] S4-2: Extracting features from multiple spatial regions:

[0073] S4-2-1: For video frames The MTCNN face detection model was used to extract the coordinates of five key points in the face spatial region, including the positions of the left eye, right eye, nose, left corner of mouth, and right corner of mouth.

[0074] S4-2-2: For each location, set the image patch size, perform image patch cropping, and obtain the face spatial region. , These are regional features, including the areas of the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth;

[0075] S4-2-3: Using Image Classification Extracting facial spatial regions In Regional characteristics, obtaining multi-spatial regional characteristics ;

[0076] S4-2-4: Merge the multi-spatial region features of all video frames ;

[0077] S4-3: Calculating Sensitive Word-Spatial Alignment Attention:

[0078] S4-3-1: Input multi-spatial region features of all video frames Spatial region key features are computed using fully connected layers. ;

[0079] S4-3-2: Using the Chinese text query features in S2-2-8 ;

[0080] S4-3-3: Calculate sensitive word-spatial alignment attention.

[0081] ,

[0082] Where d represents the dimension of the text query features, and attention reflects the relationship between driving emotion categories and video regions, including all keyframe locations t and their corresponding key spatial regions. Guided spatial alignment attention , This indicates that the sensitive word is at keyframe position t. Key spatial areas , Guided spatial alignment of attention;

[0083] S4-4: Output the key spatial region guided by sensitive words:

[0084] S4-4-1: Set the attention threshold for sensitive word spatial alignment ;

[0085] S4-4-2: Enter each sensitive word Guided spatial alignment attention ;

[0086] S4-4-3: From key time frames In the middle, from the facial spatial region Filter out all Corresponding key spatial regions

[0087] S4-4-4: Obtain the set of key spatial regions guided by sensitive words ,in satisfy .

[0088] The specific details of step S5 are as follows:

[0089] S5-1: Input a multimodal emotion dataset, where each video sample is... ;

[0090] S5-2: Set of key spatial regions guided by sensitive words obtained in step S4: Extract key multi-spatial region features:

[0091] S5-2-1: Using Image Classification Extract the set of key spatial regions guided by sensitive words In Regional characteristics, obtaining multi-spatial regional characteristics ;

[0092] S5-2-2: Output the set of key spatial region features = , ;

[0093] S5-3: Input a set of key spatial region features, concatenate the features, and use a fully connected layer to predict the probability of each emotion category. The emotion category with the highest probability is taken as the prediction result. ;

[0094] S5-4: Calculate the loss function for emotion recognition guided by sensitive words;

[0095] S5-4-1: Input Text Query Features Input visual key features Calculate the time alignment loss.

[0096] ,

[0097] in This represents the visual key feature corresponding to frame position t;

[0098] S5-4-2: Input Text Query Features Input region key features Calculate the spatial alignment loss.

[0099] ,

[0100] in Indicates frame position t, key spatial region Corresponding region key features;

[0101] S5-4-3: Input the predicted sentiment category And the manually labeled sentiment tag Y, using cross-entropy loss. Calculate the emotion classification loss.

[0102] ;

[0103] S5-4-4: Introducing Time Alignment Loss Weights Spatial alignment loss weights And emotion classification loss weights Calculate the total loss

[0104] ;

[0105] S5-5: Total Loss of Use Train the model using all results from steps S2 to S5-3 to obtain the optimal emotion recognition model. .

[0106] The specific details of step S6 are as follows:

[0107] S6-1: Input driver video frame sequence and audio sequences ;

[0108] S6-2: Use step S1-2 to process the audio sequence Recognized as an audio text sequence ;

[0109] S6-3: Input Emotion Recognition Model Step S2 is used to extract a set of candidate key time frames guided by sensitive words. ;

[0110] S6-4: Input Emotion Recognition Model Step S3 is used to extract the set of key time frames guided by sensitive words. ;

[0111] S6-5: Input Emotion Recognition Model According to step S4, the set of key spatial regions guided by sensitive words is obtained. ;

[0112] S6-6: Input Emotion Recognition Model Based on step S5-2, a set of key spatial region features is selected. = , ;

[0113] S6-7: Input Emotion Recognition Model The feature set of key spatial regions is concatenated, and a fully connected layer is used to predict the sentiment category. , , .

[0114] A safe driving emotion recognition device based on sensitive word-guided multimodal spatiotemporal alignment includes:

[0115] Data Acquisition Module: Acquires driver facial expression videos and audio, extracts audio text sequences, defines a set of sensitive words and driving emotion categories, and constructs a multimodal emotion dataset; Candidate Key Time Frame Set Filtering Module: Calculates sensitive word-guided speech attention based on the position of sensitive words in the audio text sequence and filters out a set of candidate key time frames guided by sensitive words; Key Time Frame Set Filtering Module: Constructs a sensitive word time alignment attention model based on the candidate key time frame set guided by sensitive words and filters out a set of key time frames guided by sensitive words; Key Spatial Region Set Filtering Module: Extracts multi-spatial region features based on the key time frame set guided by sensitive words and the acquired driver facial expression videos, calculates sensitive word spatial alignment attention, and filters out a set of key spatial regions guided by sensitive words; Emotion Recognition Model Training Module: Inputs the multimodal emotion dataset, extracts key multi-spatial region features and predicts emotion categories based on the obtained set of candidate key time frames guided by sensitive words, the set of key time frames guided by sensitive words, and the set of key spatial regions guided by sensitive words, calculates the sensitive word-guided emotion recognition loss function and trains the model to obtain the emotion recognition model; Driver Emotion Category Prediction Module: Based on the trained emotion recognition model, the module predicts the driver's emotion category according to the input driver's facial expression video, audio, and the obtained candidate key time frame set guided by sensitive words, key time frame set guided by sensitive words, and key spatial region set guided by sensitive words.

[0116] A computer device includes a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the steps of the aforementioned safe driving emotion recognition method based on sensitive word-guided multimodal spatiotemporal alignment.

[0117] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps of the aforementioned safe driving emotion recognition method based on sensitive word-guided multimodal spatiotemporal alignment.

[0118] The advantages of this invention are: This invention specifically solves the problem of spatiotemporal misalignment of multimodal data (sensitive words are not synchronized in the spatiotemporal data) in driving scenarios. Through the mechanism of active guidance of sensitive words, it first filters the candidate time intervals in which sensitive words appear, then accurately locates the key time frames, and then focuses on the core spatial regions of facial emotion expression (such as eyes, corners of mouth, and nose). Through a clear spatiotemporal alignment mechanism and targeted loss function training, it effectively improves the accuracy and reliability of driver emotion recognition, and can better meet the needs of safe driving for real-time and accurate emotion recognition. Attached Figure Description

[0119] Figure 1 This is a flowchart illustrating the implementation of an embodiment of the present invention;

[0120] Figure 2 A flowchart illustrating the construction of the multimodal emotion dataset mentioned in this invention;

[0121] Figure 3 This is a flowchart of the time-region filtering guided by sensitive words mentioned in this invention;

[0122] Figure 4 This is a flowchart illustrating the time alignment guided by sensitive words mentioned in this invention.

[0123] Figure 5 This is a flowchart illustrating the spatial alignment guided by sensitive words mentioned in this invention.

[0124] Figure 6 This is a flowchart of the training process for the sensitive word-guided emotion recognition model mentioned in this invention;

[0125] Figure 7 This is a flowchart of the driving emotion recognition mentioned in this invention. Detailed Implementation

[0126] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will now be described in detail with reference to the accompanying drawings and specific embodiments.

[0127] like Figure 1 As shown, a multimodal spatiotemporal alignment-based method for safe driving emotion recognition guided by sensitive words specifically includes the following steps:

[0128] like Figure 2 As shown, S1: Construction of the multimodal sentiment dataset, the specific steps are as follows:

[0129] S1-1: Collect a multimodal sentiment dataset;

[0130] S1-1-1: Use cameras and microphones to collect video and audio of the driver's facial expressions;

[0131] S1-1-2: Decode the video into a sequence of video frames. , It is the t-th frame of the video image. =1, 2, ..., ;

[0132] S1-1-3: Extract the audio sequence synchronized with the video sequence , It is the audio sample amplitude of the t-th frame. =1, 2, ..., ;

[0133] S1-1-4: Output video frame sequence and audio sequences

[0134] S1-2: Extract the audio text sequence;

[0135] S1-2-1: Input audio sequence

[0136] S1-2-2: Using the Conformer audio recognition model, the audio sequence... Recognized as an audio text sequence , It is the text recognized from the audio of frame t. =1, 2, ..., ;

[0137] S1-2-3: Output audio text sequence ;

[0138] S1-3: Define the sensitive word set: Collect sensitive words that indicate abnormal driver emotions while driving to obtain the sensitive word set. , =1, 2, ..., K, For the first One sensitive word;

[0139] S1-4: Define driving emotion categories : , ;

[0140] S1-5: Annotated multimodal sentiment dataset;

[0141] S1-5-1: Input video frame sequence audio sequence Audio text sequence ;

[0142] S1-5-2: Manually label each video with corresponding emotion tags. , ;

[0143] S1-5-3: Output a labeled video. ;

[0144] S1-5-4: Repeat steps S1-5-1 to S1-5-3 to complete the annotation of all videos and obtain the multimodal emotion dataset.

[0145] like Figure 3 As shown, S2: Time range filtering guided by sensitive words, the specific steps are as follows:

[0146] S2-1: Determine the location of the sensitive word frame;

[0147] S2-1-1: Input audio text sequence and audio sequences ;

[0148] S2-1-2: Input the set of sensitive words defined in S1-3 ;

[0149] S2-1-3: Traversing the set of sensitive words All sensitive words in ;

[0150] S2-1-4: Use the kth sensitive word Traverse the audio text sequence All of them ;

[0151] S2-1-5: Obtain the content containing the kth sensitive word of The set of subscripts t, that is, all frame positions t in which the k-th sensitive word appears;

[0152] S2-1-6: Obtain the set of all frame positions where the k-th sensitive word appears. ;

[0153] S2-1-7: In the sensitive word set In the middle, remove the set where the frame position set is empty;

[0154] S2-1-8: Merge all sensitive words to obtain the set of frame positions where all sensitive words appear. ;

[0155] S2-2: Calculate sensitive word-guided speech attention. The speech attention mechanism quantifies the correlation between sensitive words and audio features (e.g., the correlation between audio features and anger when the word "courting death" appears), filtering out candidate time frames that are more likely to contain emotions, thus narrowing down the scope for subsequent precise analysis.

[0156] S2-2-1: Set of frame locations where sensitive words appear in the input Audio sequences Audio text sequence ;

[0157] S2-2-2: Based on the set of frame positions where sensitive words appear Given all indices k, obtain the sensitive word matching set. , It is an audio text sequence The sensitive words appearing in the sensitive word set Subscript in;

[0158] S2-2-3: Traverse the sensitive word matching set All the sensitive words inside ;

[0159] S2-2-4: Input sensitive words We use a BERT pre-trained language model to obtain the textual features (FT) of sensitive words;

[0160] S2-2-5: Set of frame locations where all sensitive words appear ;

[0161] S2-2-6: Based on the frame position set In Obtain frame position The corresponding set of audio sample amplitude values ​​is denoted as , , This refers to the audio sampling amplitude.

[0162] S2-2-7: Input audio sample amplitude set ,use The speech recognition model obtains the set of audio sampling amplitudes. In the middle, audio sampling amplitude Corresponding audio features , , This represents the audio feature corresponding to frame position t;

[0163] S2-2-8: Text features extracted from input S2-2-4 The text query features are calculated using a fully connected layer, denoted as... ;

[0164] S2-2-9: Input the audio features FA extracted in S2-2-7, and use a fully connected layer to calculate the speech key features, denoted as... , It is a matrix composed of the audio features of all sensitive words;

[0165] S2-2-10: Calculate speech attention guided by sensitive words.

[0166] ,

[0167] Where d represents the dimension of the text query features, and the sensitive word-guided speech attention reflects the relationship between driving emotion categories and all sensitive words, including attention at multiple sensitive word frame positions. ;

[0168] S2-3: Determine multiple time intervals guided by sensitive words;

[0169] S2-3-1: Set the speech attention threshold ;

[0170] S2-3-2: Input-sensitive word-guided speech attention ;

[0171] S2-3-3: Set of frame locations where sensitive words appear in the input Filter out all speech attention from them corresponding ;

[0172] S2-3-4: Obtain a set of candidate key time frames guided by sensitive words , where t satisfies .

[0173] like Figure 4 As shown, S3: Time alignment guided by sensitive words, the specific steps are as follows:

[0174] S3-1: Extract multi-time interval features guided by sensitive words;

[0175] S3-1-1: Input video frame sequence and a set of candidate key time frames guided by sensitive words ;

[0176] S3-1-2: A set of candidate key time frames guided by sensitive words In Obtain the video frame corresponding to t. Obtain a set of video frames across multiple time intervals guided by sensitive words, denoted as , ,in These are the video frames corresponding to the candidate key time frames;

[0177] S3-1-3: Multi-time-interval video frame set guided by inputting sensitive words ,use Image classification model for video frame sequences Each Obtain the corresponding video features Obtain video feature set ={ , ;

[0178] S3-2: Constructing a sensitive word-temporally aligned attention mechanism to address the issue of "asynchronous speech and visual emotional signals": For example, a driver might first say "damn it" (sensitive word, speech frame t1), and then show an angry expression 1-2 seconds later (video frame t2). By using temporally aligned attention, the text features of the sensitive word are dynamically matched with the features of the video frame to find the keyframes that are truly relevant to the emotion (t2), avoiding misjudgments caused by temporal misalignment.

[0179] S3-2-1: Input video features FV, ​​and use a fully connected layer to compute video key features. ;

[0180] S3-2-2: Define text query features calculated using S2-2-8 Construct a time-aligned attention model guided by sensitive words:

[0181] ,

[0182] Where d represents the dimension of the text query features, and this attention reflects the relationship between driving emotion categories and video frames, including attention to multiple sensitive word frame locations. , This indicates the time-aligned attention guided by the sensitive word at frame position t;

[0183] S3-3: Output key time frames guided by sensitive words:

[0184] S3-3-1: Setting the time-aligned attention threshold ;

[0185] S3-3-2: Timing Alignment Attention Guided by Inputting Each Sensitive Word ;

[0186] S3-3-3: Set of frame locations where sensitive words appear in the input Filter out all speech attention from them Corresponding key time frames ;

[0187] S3-3-4: Obtain the set of key time frames guided by sensitive words , where t satisfies .

[0188] like Figure 5 As shown, S4: Spatial alignment guided by sensitive words, the specific steps are as follows:

[0189] S4-1: Input video frame sequence and key time frame set ,get middle Corresponding key video frames , gather , ,in These are the video frames corresponding to key time frames.

[0190] S4-2: Extracting features from multiple spatial regions:

[0191] S4-2-1: For video frames The MTCNN face detection model was used to extract the coordinates of five key points in the face spatial region, including the positions of the left eye, right eye, nose, left corner of mouth, and right corner of mouth.

[0192] S4-2-2: For each location, set the image patch size, perform image patch cropping, and obtain the face spatial region. , These are regional features, including the areas of the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth;

[0193] S4-2-3: Using Image Classification Extracting facial spatial regions In Regional characteristics, obtaining multi-spatial regional characteristics ;

[0194] S4-2-4: Merge the multi-spatial region features of all video frames ;

[0195] S4-3: Calculate sensitive word-spatial alignment attention. Different facial regions have different weights for expressing emotions (e.g., "downturned corners of the mouth" and "frowning" are more prominent when angry, while "wandering eyes" are more prominent when anxious). Spatial alignment attention automatically focuses on the regions where emotional expression is most pronounced, ignoring irrelevant areas (such as the forehead and cheeks), reducing noise interference. Combined with sensitive word guidance, this allows for a more accurate match between spatial features and textual emotion tags, improving the targeting of emotional cues in the spatial dimension.

[0196] S4-3-1: Input multi-spatial region features of all video frames Spatial region key features are computed using fully connected layers. ;

[0197] S4-3-2: Using the Chinese text query features in S2-2-8 ;

[0198] S4-3-3: Calculate sensitive word-spatial alignment attention.

[0199] ,

[0200] Where d represents the dimension of the text query features, and attention reflects the relationship between driving emotion categories and video regions, including all keyframe locations t and their corresponding key spatial regions. Guided spatial alignment attention , This indicates that the sensitive word is at keyframe position t. Key spatial areas , Guided spatial alignment of attention;

[0201] S4-4: Output the key spatial region guided by sensitive words:

[0202] S4-4-1: Set the attention threshold for sensitive word spatial alignment ;

[0203] S4-4-2: Enter each sensitive word Guided spatial alignment attention ;

[0204] S4-4-3: From key time frames In the middle, from the facial spatial region Filter out all Corresponding key spatial regions

[0205] S4-4-4: Obtain the set of key spatial regions guided by sensitive words ,in satisfy .

[0206] like Figure 6 As shown, S5: Training of the emotion recognition model guided by sensitive words, the specific steps are as follows;

[0207] S5-1: Input a multimodal emotion dataset, where each video sample is... ;

[0208] S5-2: Set of key spatial regions guided by sensitive words obtained in step S4: Extract key multi-spatial region features:

[0209] S5-2-1: Using Image Classification Extract the set of key spatial regions guided by sensitive words In Regional characteristics, obtaining multi-spatial regional characteristics ;

[0210] S5-2-2: Output the set of key spatial region features = , ;

[0211] S5-3: Input a set of key spatial region features, concatenate the features, and use a fully connected layer to predict the probability of each emotion category. The emotion category with the highest probability is taken as the prediction result. ;

[0212] S5-4: Calculate the loss function for emotion recognition guided by sensitive words;

[0213] S5-4-1: Input Text Query Features Input visual key features Calculate the time alignment loss.

[0214] ,

[0215] in This represents the visual key feature corresponding to frame position t;

[0216] S5-4-2: Input Text Query Features Input region key features Calculate the spatial alignment loss.

[0217] ,

[0218] in Indicates frame position t, key spatial region Corresponding region key features;

[0219] S5-4-3: Input the predicted sentiment category And the manually labeled sentiment tag Y, using cross-entropy loss. Calculate the emotion classification loss.

[0220] ;

[0221] S5-4-4: Introducing Time Alignment Loss Weights Spatial alignment loss weights And emotion classification loss weights Calculate the total loss

[0222] ;

[0223] S5-5: Total Loss of Use Train the model using all results from steps S2 to S5-3 to obtain the optimal emotion recognition model. .

[0224] like Figure 7 As shown, S6: Driver Emotion Recognition, the specific steps are as follows:

[0225] S6-1: Input driver video frame sequence and audio sequences ;

[0226] S6-2: Use step S1-2 to process the audio sequence Recognized as an audio text sequence ;

[0227] S6-3: Input Emotion Recognition Model Step S2 is used to extract a set of candidate key time frames guided by sensitive words. ;

[0228] S6-4: Input Emotion Recognition Model Step S3 is used to extract the set of key time frames guided by sensitive words. ;

[0229] S6-5: Input Emotion Recognition Model According to step S4, the set of key spatial regions guided by sensitive words is obtained. ;

[0230] S6-6: Input Emotion Recognition Model Based on step S5-2, a set of key spatial region features is selected. = , ;

[0231] S6-7: Input Emotion Recognition Model The feature set of key spatial regions is concatenated, and a fully connected layer is used to predict the sentiment category. , , .

[0232] This invention extracts sensitive word features from driving monitoring video and audio data. Due to the temporal asynchrony between spoken sensitive words and visual facial emotions, it utilizes temporal attention analysis of sensitive words and spoken text to identify multiple candidate time intervals where emotions may exist in the video. By analyzing the alignment relationship between facial features and spoken sensitive words across multiple time intervals, and employing sensitive word-visual temporal cross-attention, it adaptively determines key time frames where driving emotions exist. Since facial emotion features are hidden in numerous micro-expression details, in key time frames, component features are extracted from multiple spatial regions of facial components. Their alignment relationship with spoken sensitive words is analyzed, and using sensitive word-visual spatial cross-attention, it adaptively determines key spatial regions where driving emotions exist. This invention fully considers the spatiotemporal asynchrony of sensitive words in multimodal data, sequentially analyzing candidate time intervals, key time frames, and key spatial regions to progressively find key clues about driving emotions, effectively improving the accuracy of safe driving emotion recognition.

Claims

1. A method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment, characterized in that, Includes the following steps: S1: Collect driver facial expression videos and audio, extract audio text sequences, define a set of sensitive words and driving emotion categories, and construct a multimodal emotion dataset; S2: Calculate sensitive word-guided speech attention based on the position of sensitive words in the audio text sequence, and select a set of candidate key time frames guided by sensitive words; S3: Based on the set of candidate key time frames guided by sensitive words, construct a sensitive word time alignment attention model, and select a set of key time frames guided by sensitive words. S4: Based on the key time frame set guided by sensitive words and the collected driver expression video, extract multi-spatial region features, calculate the spatial alignment attention of sensitive words, and filter out the key spatial region set guided by sensitive words; S5: Input the multimodal emotion dataset, and based on the candidate key time frame set, key time frame set, and key spatial region set guided by sensitive words obtained in S2-S4, extract key multi-spatial region features and predict emotion categories, calculate the emotion recognition loss function guided by sensitive words and train it to obtain the emotion recognition model; S6: Based on the emotion recognition model obtained from training, predict the driver's emotion category according to the input driver's facial expression video, audio and the results of S2-S5.

2. The method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment according to claim 1, characterized in that, The specific details of step S1 are as follows: S1-1: Collect a multimodal sentiment dataset; S1-1-1: Use cameras and microphones to collect video and audio of the driver's facial expressions; S1-1-2: Decode the video into a sequence of video frames. , It is the t-th frame of the video image. =1, 2, ..., ; S1-1-3: Extract the audio sequence synchronized with the video sequence , It is the audio sample amplitude of the t-th frame. =1, 2, ..., ; S1-1-4: Output video frame sequence and audio sequence A; S1-2: Extract the audio text sequence; S1-2-1: Input audio sequence A; S1-2-2: Using the Conformer audio recognition model, the audio sequence... Recognized as an audio text sequence , It is the text recognized from the audio of frame t. =1, 2, ..., ; S1-2-3: Output audio text sequence ; S1-3: Define the sensitive word set: Collect sensitive words that indicate abnormal driver emotions while driving to obtain the sensitive word set. , =1, 2, ..., K, For the first One sensitive word; S1-4: Define driving emotion categories : , ; S1-5: Annotated multimodal sentiment dataset; S1-5-1: Input video frame sequence audio sequence Audio text sequence ; S1-5-2: Manually label each video with corresponding emotion tags. , ; S1-5-3: Output a labeled video. ; S1-5-4: Repeat steps S1-5-1 to S1-5-3 to complete the annotation of all videos and obtain the multimodal emotion dataset.

3. The method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment according to claim 2, characterized in that, The specific details of step S2 are as follows: S2-1: Determine the location of the sensitive word frame; S2-1-1: Input audio text sequence and audio sequences ; S2-1-2: Input the set of sensitive words defined in S1-3 ; S2-1-3: Traversing the set of sensitive words All sensitive words in ; S2-1-4: Use the kth sensitive word Traverse the audio text sequence All of them ; S2-1-5: Obtain the content containing the kth sensitive word of The set of subscripts t, that is, all frame positions t in which the k-th sensitive word appears; S2-1-6: Obtain the set of all frame positions where the k-th sensitive word appears. ; S2-1-7: In the sensitive word set In the middle, remove the set where the frame position set is empty; S2-1-8: Merge all sensitive words to obtain the set of frame positions where all sensitive words appear. ; S2-2: Calculate speech attention guided by sensitive words: S2-2-1: Set of frame locations where sensitive words appear in the input Audio sequences Audio text sequence ; S2-2-2: Based on the set of frame positions where sensitive words appear Given all indices k, obtain the sensitive word matching set. , It is an audio text sequence The sensitive words appearing in the sensitive word set Subscript in; S2-2-3: Traverse the sensitive word matching set All the sensitive words inside ; S2-2-4: Input sensitive words We use a BERT pre-trained language model to obtain the textual features (FT) of sensitive words; S2-2-5: Set of frame locations where all sensitive words appear ; S2-2-6: Based on the frame position set In Obtain frame position The corresponding set of audio sample amplitude values ​​is denoted as , , This refers to the audio sampling amplitude. S2-2-7: Input audio sample amplitude set ,use The speech recognition model obtains the set of audio sampling amplitudes. In the middle, audio sampling amplitude Corresponding audio features , , This represents the audio feature corresponding to frame position t; S2-2-8: Text features extracted from input S2-2-4 The text query features are calculated using a fully connected layer, denoted as... ; S2-2-9: Input the audio features FA extracted in S2-2-7, and use a fully connected layer to calculate the speech key features, denoted as... , It is a matrix composed of the audio features of all sensitive words; S2-2-10: Calculate speech attention guided by sensitive words. , Where d represents the dimension of the text query features, and the sensitive word-guided speech attention reflects the relationship between driving emotion categories and all sensitive words, including attention at multiple sensitive word frame positions. ; S2-3: Determine multiple time intervals guided by sensitive words; S2-3-1: Set the speech attention threshold ; S2-3-2: Input-sensitive word-guided speech attention ; S2-3-3: Set of frame locations where sensitive words appear in the input Filter out all speech attention from them corresponding ; S2-3-4: Obtain a set of candidate key time frames guided by sensitive words , where t satisfies .

4. The method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment according to claim 3, characterized in that, The specific details of step S3 are as follows: S3-1: Extract multi-time interval features guided by sensitive words; S3-1-1: Input video frame sequence and a set of candidate key time frames guided by sensitive words ; S3-1-2: A set of candidate key time frames guided by sensitive words In Obtain the video frame corresponding to t. Obtain a set of video frames across multiple time intervals guided by sensitive words, denoted as , ,in These are the video frames corresponding to the candidate key time frames; S3-1-3: Multi-time-interval video frame set guided by inputting sensitive words ,use Image classification model for video frame sequences Each Obtain the corresponding video features Obtain video feature set ={ , ; S3-2: Constructing a sensitive word-time aligned attention model: S3-2-1: Input video features FV, ​​and use a fully connected layer to compute video key features. ; S3-2-2: Define text query features calculated using S2-2-8 Constructing time-aligned attention guided by sensitive words: , Where d represents the dimension of the text query features, and this attention reflects the relationship between driving emotion categories and video frames, including attention to multiple sensitive word frame locations. , This indicates the time-aligned attention guided by the sensitive word at frame position t; S3-3: Output key time frames guided by sensitive words: S3-3-1: Setting the time-aligned attention threshold ; S3-3-2: Timing Alignment Attention Guided by Inputting Each Sensitive Word ; S3-3-3: Set of frame locations where sensitive words appear in the input Filter out all speech attention from them Corresponding key time frames ; S3-3-4: Obtain the set of key time frames guided by sensitive words , where t satisfies .

5. The method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment according to claim 4, characterized in that, The specific details of step S4 are as follows: S4-1: Input video frame sequence and key time frame set ,get middle Corresponding key video frames , gather , ,in These are the video frames corresponding to key time frames; S4-2: Extracting features from multiple spatial regions: S4-2-1: For video frames The MTCNN face detection model was used to extract the coordinates of five key points in the face spatial region, including the positions of the left eye, right eye, nose, left corner of mouth, and right corner of mouth. S4-2-2: For each location, set the image patch size, perform image patch cropping, and obtain the face spatial region. , These are regional features, including the areas of the left eye, right eye, nose, left corner of the mouth, and right corner of the mouth; S4-2-3: Using Image Classification Extracting facial spatial regions In Regional characteristics, obtaining multi-spatial regional characteristics ; S4-2-4: Merge the multi-spatial region features of all video frames ; S4-3: Calculating Sensitive Word-Spatial Alignment Attention: S4-3-1: Input multi-spatial region features of all video frames Spatial region key features are computed using fully connected layers. ; S4-3-2: Input text query features from S2-2-8 ; S4-3-3: Calculate sensitive word-spatial alignment attention. , Where d represents the dimension of the text query features, and attention reflects the relationship between driving emotion categories and video regions, including all key time frames t and their corresponding key spatial regions. Guided spatial alignment attention , This indicates that the sensitive word is set to t in the key time frame. Key spatial areas , Guided spatial alignment of attention; S4-4: Output the key spatial region guided by sensitive words: S4-4-1: Set the attention threshold for sensitive word spatial alignment ; S4-4-2: Enter each sensitive word Guided spatial alignment attention ; S4-4-3: From key time frames In the middle, from the facial spatial region Filter out all Corresponding key spatial regions ; S4-4-4: Obtain the set of key spatial regions guided by sensitive words ,in satisfy .

6. The method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment according to claim 5, characterized in that, The specific details of step S5 are as follows: S5-1: Input a multimodal emotion dataset, where each video sample is... ; S5-2: Set of key spatial regions guided by sensitive words obtained in step S4: Extract key multi-spatial region features: S5-2-1: Using Image Classification Extract the set of key spatial regions guided by sensitive words In Regional characteristics, obtaining multi-spatial regional characteristics ; S5-2-2: Output the set of key spatial region features = , ; S5-3: Input a set of key spatial region features, concatenate the features, and use a fully connected layer to predict the probability of each emotion category. The emotion category with the highest probability is taken as the prediction result. ; S5-4: Calculate the loss function for emotion recognition guided by sensitive words; S5-4-1: Input Text Query Features Input visual key features Calculate the time alignment loss. , in This represents the visual key feature corresponding to frame position t; S5-4-2: Input Text Query Features Input region key features Calculate the spatial alignment loss. , in Indicates frame position t, key spatial region Corresponding region key features; S5-4-3: Input the predicted sentiment category And the manually labeled sentiment tag Y, using cross-entropy loss. Calculate the emotion classification loss. ; S5-4-4: Introducing Time Alignment Loss Weights Spatial alignment loss weights And emotion classification loss weights Calculate the total loss ; S5-5: Total Loss of Use Train the model using all results from steps S2 to S5-3 to obtain the optimal emotion recognition model. .

7. A method for safe driving emotion recognition based on sensitive word-guided multimodal spatiotemporal alignment according to claim 6, characterized in that, The specific details of step S6 are as follows: S6-1: Input driver video frame sequence and audio sequences ; S6-2: Use step S1-2 to process the audio sequence Recognized as an audio text sequence ; S6-3: Input Emotion Recognition Model Step S2 is used to extract a set of candidate key time frames guided by sensitive words. ; S6-4: Input Emotion Recognition Model Step S3 is used to extract the set of key time frames guided by sensitive words. ; S6-5: Input Emotion Recognition Model According to step S4, the set of key spatial regions guided by sensitive words is obtained. ; S6-6: Input Emotion Recognition Model Based on step S5-2, a set of key spatial region features is selected. = , ; S6-7: Input Emotion Recognition Model The feature set of key spatial regions is concatenated, and a fully connected layer is used to predict the sentiment category. , , .

8. A safe driving emotion recognition device based on sensitive word-guided multimodal spatiotemporal alignment, characterized in that: Including: Data acquisition module: Collects driver facial expression videos and audio, extracts audio text sequences, defines a set of sensitive words and driving emotion categories, and constructs a multimodal emotion dataset; Candidate key time frame set filtering module: Based on the position of the sensitive word in the audio text sequence, calculate the speech attention guided by the sensitive word, and filter out the candidate key time frame set guided by the sensitive word; Key time frame set filtering module: Based on the candidate key time frame set guided by sensitive words, a sensitive word time alignment attention model is constructed to filter out the key time frame set guided by sensitive words; Key Spatial Region Set Filtering Module: Based on the set of key time frames guided by sensitive words and the collected driver facial expression videos, multi-spatial region features are extracted, sensitive word spatial alignment attention is calculated, and the set of key spatial regions guided by sensitive words is filtered out. The emotion recognition model training module: Inputting the multimodal emotion dataset, based on the obtained set of candidate key time frames guided by sensitive words, the set of key time frames guided by sensitive words, and the set of key spatial regions guided by sensitive words, it extracts key multi-spatial region features and predicts the emotion category. It calculates the emotion recognition loss function guided by sensitive words and trains the model to obtain the emotion recognition model. The driver emotion category prediction module: Based on the trained emotion recognition model, according to the input driver expression video and audio, and the obtained set of candidate key time frames guided by sensitive words, the set of key time frames guided by sensitive words, and the set of key spatial regions guided by sensitive words, it predicts the driver's emotion category.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that, When the processor executes the computer program, it implements the steps of the safe driving emotion recognition method based on sensitive word-guided multimodal spatiotemporal alignment as described in any one of claims 1-7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the steps of the safe driving emotion recognition method based on sensitive word-guided multimodal spatiotemporal alignment as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Multi-modal face emotion recognition method and device

    CN114399818A

  • Multi-modal emotion recognition method based on non-aligned sequence

    CN116052291A

  • Multi-modal emotion recognition method and system

    CN116189669A

  • Multi-modal emotion recognition method and system for guiding attention fusion based on text modal, and storage medium

    CN117786596A

  • Driver emotion recognition method based on multi-modal data fusion

    CN119502923A