Multi-modal dialogue emotion recognition method based on face recognition and emotion reasoning chain
Through the method based on face recognition and emotional reasoning chain, the facial features of the target speaker are extracted and the emotion recognition model is enhanced, and the problem of decreasing recognition accuracy in complex environments and multi-person scenarios in the prior art is solved, achieving a more efficient and robust emotion recognition effect.
Patent Information
- Application Number
- CN202411982100.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
When dealing with complex environments and multi-person scenarios, existing multimodal dialogue emotions recognition methods face problems such as lighting changes, background noise and low data quality, resulting in a decrease in recognition accuracy.
Using a method based on face recognition and emotional reasoning chain, the facial features of the target speaker are extracted through face detection and k-means clustering, reducing the impact of background noise, and using emotional reasoning chains to enhance the model's emotional recognition ability.
It improves the accuracy and interpretability of multimodal dialogue emotion recognition, reduces the calculation amount and processing time, and enhances the robustness of the model in complex scenarios.
Smart Images

Figure CN119939414A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of multimodal emotion recognition, and in particular to a multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain. Background Art
[0002] With the gradual deepening of research in the field of affective computing and human-computer interaction, emotion recognition based solely on single-modal information such as text or speech can no longer meet the current complex and diverse practical needs. Therefore, emotion recognition research based on multimodal information fusion has emerged and has become a hot topic, among which the multimodal dialogue emotion recognition (MERC) task is the key.
[0003] The multimodal conversation emotion recognition (MERC) task in deep learning refers to the use of deep learning technology to identify the emotional state (happy, angry, sad, etc.) of the participants in the conversation by analyzing multiple modal information in the conversation (such as text, voice, facial expressions, body movements, etc.), and fuse the emotional features from different modalities. The fusion method can be early fusion (directly fuse the features of different modalities after feature extraction), late fusion (recognize emotions on each modality separately, and then fuse the recognition results) or hybrid fusion (combining the characteristics of early and late fusion). The ultimate goal is to accurately judge the emotional state in the conversation based on the fused features through classifiers (such as support vector machines, deep neural networks, etc.), and be able to handle the dynamic changes of emotions during the conversation, providing a basis for subsequent applications such as dialogue strategies and emotional support. With the widespread application of multimodal large models, the MERC task can be better completed, enabling computer systems to understand the emotional content in the conversation like humans, so as to play a role in many fields such as intelligent customer service, emotional companion robots, and conference analysis.
[0004] At present, the conventional method of implementing MERC tasks mainly includes the following key steps. In the data preprocessing stage, comprehensive processing work is carried out on the collected multimodal data such as voice, text, and images. The first is data cleaning, which aims to remove all kinds of errors, redundancy, and interference information in the data to ensure the purity and availability of the data; the second is data labeling, according to specific emotion classification standards, to give each data segment a corresponding emotion label; finally, data alignment is performed to make different modal data accurately correspond in the time dimension or content level, laying a good foundation for subsequent processing.
[0005] In the feature extraction phase, for text feature extraction, the text is first segmented, and then the bag-of-words model and word embedding technology are used to convert the text into a vector representation suitable for computer processing. Subsequently, the neural network model is used to deeply explore the semantic structure, syntactic features and other deep-level information of the text to accurately interpret the emotional connotation contained in the text. In terms of image feature extraction, it mainly relies on convolutional neural networks to extract features from facial images or expression images in videos, and effectively captures feature information related to emotions by autonomously learning subtle changes in facial expressions.
[0006] In the modal fusion stage, a more direct approach is usually adopted, that is, the feature vectors extracted from different modalities are simply concatenated to generate a new joint feature vector as the input data of the subsequent emotion classifier.
[0007] In the emotion classification stage, the fused features are input into classifiers such as support vector machines, multi-layer perceptrons, and Softmax regression to perform detailed classification of the emotions in the conversation. Common emotion categories include happiness, sadness, anger, surprise, fear, disgust, etc.
[0008] Although this traditional method is feasible and simple to a certain extent, there are still many problems:
[0009] On the one hand, in actual application scenarios, environmental factors pose many obstacles to feature extraction. For example, dynamic changes in lighting conditions and cluttered background environments pose great challenges to feature extraction. Changes in lighting may cause confusion between target color and background color, leading to false detection and erroneous tracking, which seriously affects the accuracy and stability of feature extraction. In addition, the target motion in the video is complex and diverse, such as rapid movement, irregular motion, and sudden changes in motion direction and speed, which makes it extremely difficult to accurately capture the features of the target. At the same time, video data is often mixed with a large amount of irrelevant information, which constitutes noise interference, which affects the quality of feature extraction to a certain extent, and ultimately leads to a decrease in the accuracy of emotion recognition, making it difficult to meet the actual application requirements of high precision and high reliability.
[0010] On the other hand: the collection of multimodal data requires a lot of manpower, material resources and time, and the amount of high-quality multimodal dialogue data is small. At the same time, it is also challenging to accurately label multimodal data with emotions, and the subjectivity and inconsistency of the labeling results may affect the training effect of the model.
[0011] Therefore, exploring more advanced, efficient and robust multimodal dialogue emotion recognition methods and technical systems has become an important issue that needs to be urgently addressed in the current research field. Summary of the invention
[0012] In view of the shortcomings of the prior art, the present invention proposes a multimodal conversation emotion recognition method based on face recognition and emotion reasoning chain, aiming to improve the performance of the model in multimodal conversation emotion recognition and obtain more accurate emotion recognition results by providing a method that can accurately capture visual information, remove noise, and use emotion reasoning chain to solve the problem of low data quality.
[0013] The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain proposed in the present invention includes the following specific steps:
[0014] Step 1: Obtain a multi-round dialogue dataset; the multi-round dialogue dataset includes several multi-round dialogues, each of which includes several sentences, the speaker of each sentence, the emotion label corresponding to each sentence, and the dialogue video corresponding to each sentence;
[0015] Step 2: Determine the speaker in the conversation video corresponding to each sentence in the multi-round conversation dataset, and then obtain the speaker's face image sequence;
[0016] Step 2.1: Sampling the conversation video corresponding to each sentence to obtain a sampled image sequence; the image sequence includes the first video frame, the last video frame and a number of randomly sampled video frames of the conversation video;
[0017] The random sampling method is specifically as follows: reading the conversation video, constructing a frame list including all video frames of the conversation video, randomly generating an index list, sampling video frames from the frame list through the index list, and obtaining an image sequence F of length n;
[0018] Step 2.2: Use the face detection neural network to calculate the position information of each face in each video frame of the image sequence, and determine the number of people in the video conversation by the number of faces in the first video frame;
[0019] Step 2.2.1: For each video frame f in the image sequence F i , use the face detection neural network to calculate the face position and get the position information c of each face i,j , where i is the number of the video frame and j is the number of the face;
[0020] The position information includes the coordinates of the upper left corner of the face and the height and width of the face frame;
[0021] Step 2.2.2: According to the number of face position information output by the face detection network in the first video frame, determine the number of faces k in the first video frame and use it as the number of people in the video conversation;
[0022] Step 2.3: Based on the obtained face position information, all face images of each video frame of the image sequence are cropped, and the face images of different people are classified by the k-means algorithm according to the determined number of people in the video conversation, and the face images of the same person are sorted in chronological order to obtain a sequence of face images of different people;
[0023] Step 2.3.1: For each video frame f in the image sequence F i , according to the position information of the face in the video frame, the face image is cut out, and all face images constitute the face image set S face ;
[0024] Step 2.3.2: For the face image set S face For each face image in , use a deep learning neural network to calculate the feature vector v of the face image;
[0025] Step 2.3.3: Initialize the feature vectors of the k face images of the first video frame in the image sequence as k cluster centers μ;
[0026] Step 2.3.4: For the remaining face images except the k face images of the first video frame, calculate the Euclidean distance between the feature vector v of each face image and each cluster center μ, assign the face image to the cluster with the cluster center closest to its Euclidean distance, and obtain k clusters;
[0027] Step 2.3.5: For each cluster, update the cluster center according to the feature vector of the face image in the cluster;
[0028] Step 2.3.6: Determine whether the cluster center remains unchanged within the set round or whether the number of iterations reaches the set number of iterations. If so, execute step 2.3.8, otherwise execute step 2.3.7;
[0029] Step 2.3.7: According to the updated cluster centers, redistribute all face images to obtain k clusters and return to step 2.3.5;
[0030] Step 2.3.8: Sort the face images in each cluster according to the time sequence in the conversation video, and redistribute the clusters with multiple face images or missing face images at a certain time point, and finally obtain a face image sequence S of k people;
[0031] Specifically, for a cluster with redundant face images at a certain point in time, the Euclidean distance between all face images contained in the cluster at that point in time and other face images in the cluster is calculated, and the face image with the smallest average Euclidean distance is retained. The other face images will be reallocated, and the Euclidean distance between the face image to be reallocated and the cluster center of all clusters lacking face images at this point in time is calculated, and the face image is reallocated to the cluster with the smallest Euclidean distance;
[0032] Step 2.4: Using a facial key point detection tool to determine the speaker, and then determining the speaker's facial image sequence from facial image sequences of different persons;
[0033] Step 2.4.1: For each person's facial image sequence, use a facial key point detection tool to detect the coordinates of the facial key points of each facial image, and calculate the coordinate distance d between the upper lip and the lower lip; the facial key points include the upper lip and the lower lip;
[0034] Step 2.4.2: According to the coordinate distance d between the upper lip and the lower lip in each face image, calculate the variance σ of the coordinate distance between the upper lip and the lower lip in each person's face image sequence;
[0035] Step 2.4.3: Take the face image sequence with the largest variance σ as the speaker’s face image sequence S t ;
[0036] Step 3: preprocessing the speaker's facial image sequence to obtain a preprocessed facial image sequence;
[0037] Step 3.1: screening the face images in the face image sequence of the speaker to obtain a screened face image sequence of the speaker;
[0038] Specifically: remove face images with resolutions below a set threshold and remove face images with missing facial key points;
[0039] Step 3.2: Set the target size (w t ,h t ), scale the filtered face image to the target size (w t ,h t ), and then align the facial key points of the face image, where w t is the set width, h t is the set height;
[0040] Step 4: Based on the multi-round dialogue dataset obtained in step 1 and the pre-processed speaker face image sequence, construct several training templates;
[0041] Step 4.1: Build a sentence vector library D;
[0042] Step 4.1.1: Based on the multi-round dialogue data set, obtain training samples and then construct a training set including several training samples;
[0043] The training sample includes a conversation history, a video corresponding to a target sentence, and an emotion label of the target sentence. The conversation history includes a plurality of sentences and a speaker corresponding to each sentence. The target sentence is the last sentence in the conversation history.
[0044] The method for acquiring the historical dialogue is as follows: for each multi-round dialogue in the multi-round dialogue data set, the first w consecutive sentences are sequentially taken as a dialogue history using a window of size w, where w=1, 2, ..., W, and W is the number of sentences in a multi-round dialogue, that is, each sentence in the multi-round dialogue is taken as the last sentence in a dialogue history;
[0045] Step 4.1.2: Preprocess the sentences in the training set to obtain preprocessed sentences;
[0046] The preprocessing includes: removing the speaker of the sentence, replacing the character name in the sentence, and then ensuring that the difference in the number of sentences of various emotion tags is within a set range;
[0047] Step 4.1.3: Encode the preprocessed sentences to obtain the feature vector of each sentence;
[0048] Step 4.1.4: The preprocessed sentence u and its corresponding feature vector v u and the emotion label e as a piece of data in the sentence vector library, and construct a sentence vector library D including several pieces of data;
[0049] Step 4.2: For each target sentence in the training set, use the sentence similarity matching algorithm to retrieve the z sentences and their emotion labels that are most similar to the target sentence in the sentence vector library as candidate reference sentences;
[0050] Specifically, the cosine similarity between the feature vector of the target sentence and the feature vector of each other sentence in the sentence vector library except the sentence itself is calculated, and the z sentences with the highest cosine similarity and their emotion labels are taken as candidate reference sentences d h , h is the number of the candidate reference sentence;
[0051] Step 4.3: Set the instruction description and construct several input templates based on the training samples, candidate reference sentences and the speaker's face image sequence obtained in step 3;
[0052] The input template includes instruction description ins, dialogue history h, target sentence u target , target face image R and candidate reference sentence d h ;
[0053] The instruction manual provides background information of the emotion recognition task, clarifies the composition of the input data, and describes the tasks to be completed when processing the input data;
[0054] The target face image is a face image at the time corresponding to the target sentence obtained from the speaker's face image sequence;
[0055] Step 4.4: Prompt the general large language model to perform the dialogue emotion recognition task on the dialogue history in the input template in the form of thought chain, and obtain several emotion reasoning chains;
[0056] Step 4.4.1: Write several thought chain examples for conversation emotion recognition, where each thought chain example includes: conversation history, emotion classification thinking process, and emotion classification results;
[0057] Step 4.4.2: Adjust the several thought chain examples obtained in step 4.4.1 according to the emotion classification results to ensure that the difference in the number of thought chain examples corresponding to various emotion classification results is within the set range, and encode the dialogue history in each thought chain example through a sentence-level encoder to obtain a feature vector;
[0058] Step 4.4.3: Take the thought chain example and its corresponding feature vector as a data unit to form an example corpus including several data units;
[0059] Step 4.4.4: Use the sentence-level encoder to obtain a feature vector for each dialogue history in the input template, and calculate the cosine similarity with the feature vectors of the dialogue history in all thought chain examples in the example corpus except the thought chain example corresponding to the dialogue history, to obtain the thought chain example with the highest score;
[0060] Step 4.4.5: Take the historical dialogue in the input template as a question, and prompt the large model to complete the dialogue emotion recognition task based on the highest-scoring thought chain example, and obtain the emotion classification thinking process and emotion classification results. The emotion classification thinking process is called emotion reasoning chain c;
[0061] Step 4.4.6: Repeat steps 4.4.4 to 4.4.5 to obtain several emotion reasoning chains c and emotion classification results;
[0062] Step 4.5: Take an emotion reasoning chain c and the corresponding emotion label of the historical dialogue as an output template, and then obtain several output templates;
[0063] Step 4.6: Take a set of input templates and their corresponding output templates as a training template, and then obtain several training templates;
[0064] Step 5: Use several training templates to fine-tune the multimodal instructions of the multimodal large model to obtain a fine-tuned multimodal large model;
[0065] Step 6: Use the fine-tuned multimodal large model to perform emotion recognition on the conversation video to be recognized;
[0066] Step 6.1: Extract the conversation video slice corresponding to the last sentence from the conversation video, and obtain the face image sequence of the speaker in the conversation video slice corresponding to the last sentence according to the method of steps 2-3;
[0067] Step 6.2: Use the sentences in the conversation video to form a conversation history, and use the last sentence as the target sentence;
[0068] Step 6.3: Obtain candidate reference sentences from the sentence vector library according to step 4.2;
[0069] Step 6.4: Take the face image of the speaker at the time corresponding to the target sentence in the face image sequence as the target face image;
[0070] Step 6.5: Input the instruction description, dialogue history, target sentence, target face image and candidate reference sentence into the fine-tuned multimodal large model to obtain the emotion reasoning chain c and emotion classification results.
[0071] Compared with the prior art, the beneficial effects of the present invention are:
[0072] 1. Traditional methods may need to extract global features of each frame from the video (including irrelevant information such as background and environment), while face positioning and screenshots focus on the facial area of the person, reducing the amount of calculation and improving processing efficiency.
[0073] 2. By locating the face and focusing on the facial features of the target speaker, background noise and environmental interference can be reduced, and recognition accuracy can be improved. Especially in multi-person or complex scenes, face positioning can help accurately capture the emotional changes of the target person without being affected by other people or background.
[0074] 3. Only the facial features of the target person are extracted, which can better capture the person's expressions, movements and emotional changes. Compared with the global feature extraction of traditional methods, this method helps to accurately track the changes in the speaker's facial expressions, especially in multi-person scenes, avoiding confusion.
[0075] 4. The model training method based on the emotion reasoning chain designed by the present invention uses a closed-source large model that is more powerful than the fine-tuned model to generate the emotion reasoning chain. The emotion reasoning chain is an emotion reasoning process based on the conversation history. It uses a powerful large language model to generate the thinking process of conversation emotion analysis. By simulating the human emotion reasoning process, it improves the accuracy and interpretability of emotion recognition, makes the emotion classification results clearer, and optimizes the reasoning process. BRIEF DESCRIPTION OF THE DRAWINGS
[0076] Figure 1 is a flow chart of a multimodal conversation emotion recognition method based on face recognition and emotion reasoning chain in this embodiment;
[0077] Figure 2 This is a framework diagram of the method based on face recognition and emotion reasoning chain in this implementation. DETAILED DESCRIPTION
[0078] In order to facilitate the understanding of the present application, the specific embodiments of the present invention are further described in detail below in conjunction with the accompanying drawings and embodiments. The following embodiments are used to illustrate the present invention, but are not intended to limit the scope of the present invention. On the contrary, the purpose of providing these embodiments is to make the disclosure of the present application more thoroughly understood.
[0079] The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain of this embodiment is as follows: Figure 1 and Figure 2 As shown, the method comprises the following steps:
[0080] Step 1: Obtain a multi-round dialogue dataset; the multi-round dialogue dataset includes several multi-round dialogues, each of which includes several sentences, the speaker of each sentence, the emotion label corresponding to each sentence, and the dialogue video corresponding to each sentence;
[0081] Step 2: Determine the speaker in the conversation video corresponding to each sentence in the multi-round conversation dataset, and then obtain the speaker's face image sequence;
[0082] Step 2.1: Sampling the conversation video corresponding to each sentence to obtain a sampled image sequence; the image sequence includes the first video frame, the last video frame and a number of randomly sampled video frames of the conversation video;
[0083] The random sampling method is specifically as follows: reading the conversation video, constructing a frame list including all video frames of the conversation video, randomly generating an index list, sampling video frames from the frame list through the index list, and obtaining an image sequence F of length n;
[0084] When processing conversation videos, without scene changes, the content between adjacent video frames is usually very similar, which means that conversation videos often contain a lot of redundant information. If this redundant information is not filtered out, it will not only waste computing resources, but also may cause a decrease in recognition accuracy. To solve this problem, we first need to read the conversation video and perform effective frame sampling.
[0085] In this implementation, the input conversation video is read, the sampling number n is defined, and two lists are constructed. The first list is a frame list containing all video frames of the conversation video, and the second list is an index list containing the index of the sampled video frame. When reading the conversation video, each video frame of the conversation video is saved to the frame list. The number of video frames of the conversation video is N. The specific steps of constructing the index list are as follows:
[0086] 1. First, keep the first video frame and add 0 to the index list;
[0087] 2. Evenly divide 1 to N-2 into n-2 intervals, randomly generate a random integer from each interval, representing the index number of the sampled video frame, and add it to the index list;
[0088] 3. Keep the last video frame and add N-1 to the index list;
[0089] After the two lists are constructed, the sampled video frames are selected from the frame list using the index list to obtain the image sequence F.
[0090] Step 2.2: Use the face detection neural network (MTCNN) to calculate the position information of each face in each video frame of the image sequence, and determine the number of people in the video conversation by the number of faces in the first video frame;
[0091] In order to accurately detect the faces contained in each video frame in the conversation video while taking into account the processing speed, step 2.2 uses the face detection neural network MTCNN, which has a relatively good face recognition accuracy and processing speed. MTCNN accepts an image as input and outputs whether the image contains a face, the coordinates of the upper left corner of the face frame, and the width and height of the face frame. MTCNN corrects the prediction results multiple times through internal networks and operations such as non-maximum suppression (NMS), and finally generates a prediction result with a high accuracy.
[0092] Step 2.2.1: For each video frame f in the image sequence F i , use the face detection neural network to calculate the face position and get the position information c of each face i,j , where i is the number of the video frame and j is the number of the face;
[0093] In this embodiment, each video frame f in the image sequence Fi , use the face detection network MTCNN to calculate the location information of each face, and output whether there is a face in the video frame and the location information of each face. The location information includes
[0094] The coordinates of the upper left corner of the face (x, y) and the height and width of the face frame (h, w). Each video frame may contain multiple faces, so the face detection network will output multiple of the above location information.
[0095] Step 2.2.2: In order to separate different face images, it is necessary to determine how many people are involved in the conversation video. In a conversation scene, the first video frame often contains a group of people participating in the conversation, whether it is a frontal conversation or a side conversation. Therefore, according to the number of face position information output by the face detection network in the first video frame, the number of faces k in the first video frame is determined and used as the number of people in the video conversation. This will be used as the initialization number of cluster centers in the subsequent clustering steps;
[0096] Step 2.3: Based on the obtained face position information, all face images of each video frame of the image sequence are cropped, and the face images of different people are classified by the k-means algorithm according to the determined number of people in the video conversation, and the face images of the same person are sorted in chronological order to obtain a sequence of face images of different people;
[0097] The steps of the k-means algorithm include: 1. Randomly select k data points as the initial cluster centers, each cluster center represents a cluster category and serves as a reference point for subsequent calculations. 2. Assign each data point to the cluster center closest to it, that is, according to the Euclidean distance metric, assign each image to the cluster closest to its feature vector. 3. Recalculate the center of each cluster, that is, the mean of the feature vectors of all images in the cluster. The new cluster center will become the basis for the next round of allocation. 4. Repeat 2 and 3 until the cluster center no longer changes or the preset number of iterations is reached.
[0098] In this embodiment, the number of clusters is the number of people k in the video conversation. The face images are continuously reallocated to the clusters and the cluster centers are updated until the algorithm converges. Then the face images in each cluster are sorted in the order of their original time dimension in the conversation video. For each cluster, the redundant images in each time dimension are reallocated to ensure that each cluster contains the correct face image sequence, and finally a time-sorted k-person face image sequence S is obtained;
[0099] Step 2.3.1: For each video frame f in the image sequence F i , according to the position information of the face in the video frame, the face image is cut out, and all face images constitute the face image set S face ;
[0100] Step 2.3.2: For the face image set S face For each face image in , use a deep learning neural network to calculate the feature vector v of the face image;
[0101] Step 2.3.3: Initialize the feature vectors of the k face images of the first video frame in the image sequence as k cluster centers μ;
[0102] Step 2.3.4: For the remaining face images except the k face images of the first video frame, calculate the Euclidean distance between the feature vector v of each face image and each cluster center μ, assign the face image to the cluster with the cluster center closest to its Euclidean distance, and obtain k clusters;
[0103] The Euclidean distance calculation formula between the feature vector v and each cluster center μ is:
[0104]
[0105] Where D(v,μ) is the Euclidean distance between the feature vector v and the cluster center μ;
[0106] Step 2.3.5: For each cluster, update the cluster center according to the feature vector of the face image in the cluster;
[0107] Cluster center calculation formula:
[0108]
[0109] Among them, v a is the feature vector of the a-th face image in the current cluster, a is the number of face images in the current cluster, and b is the number of face images in the current cluster;
[0110] Step 2.3.6: Determine whether the cluster center remains unchanged within the set round or whether the number of iterations reaches the set number of iterations. If so, execute step 2.3.8, otherwise execute step 2.3.7;
[0111] Step 2.3.7: According to the updated cluster centers, redistribute all face images to obtain k clusters and return to step 2.3.5;
[0112] Step 2.3.8: Sort the face images in each cluster according to the time sequence in the conversation video, and redistribute the clusters with multiple face images or missing face images at a certain time point, and finally obtain a face image sequence S of k people;
[0113] Specifically, there may be misassigned face images in each cluster, that is, there may be redundant face images at a certain point in time, and there must be other clusters that lack face images at this point in time. For a cluster with redundant face images at a certain point in time, the Euclidean distance between all face images contained in it at that point in time and other face images in the cluster is calculated, and the face image with the smallest average Euclidean distance is retained. The other face images will be reallocated, and the Euclidean distance between the face image to be reallocated and the cluster center of all clusters that lack face images at this point in time is calculated, and the face image is reallocated to the cluster with the smallest Euclidean distance;
[0114] Step 2.4: Using a facial key point detection tool to determine the speaker, and then determining the speaker's facial image sequence from facial image sequences of different persons;
[0115] In this embodiment, Dlib is used as a facial key point detection tool. After the facial key points are detected by Dlib, the speaker in the conversation video is determined by the change in lip coordinates.
[0116] Step 2.4.1: For each person's facial image sequence, use a facial key point detection tool to detect the coordinates of the facial key points of each facial image, and calculate the coordinate distance d between the upper lip and the lower lip; the facial key points include the upper lip and the lower lip;
[0117] In this embodiment, the facial key points include left and right eyes, nose, upper and lower lips, etc., and only the coordinates of the upper and lower lips are used. In this embodiment, the Euclidean distance formula is used to calculate the coordinate distance d between the upper lip and the lower lip;
[0118] Step 2.4.2: According to the coordinate distance d between the upper lip and the lower lip in each face image, calculate the variance σ of the coordinate distance between the upper lip and the lower lip in each person's face image sequence;
[0119] Step 2.4.3: Take the face image sequence with the largest variance σ as the speaker’s face image sequence S t ;
[0120] Step 3: preprocessing the speaker's facial image sequence to obtain a preprocessed facial image sequence;
[0121] Step 3.1: screening the face images in the face image sequence of the speaker to obtain a screened face image sequence of the speaker;
[0122] Screen high-quality facial images of speakers by face image clarity and face stability, specifically: remove face images with resolutions below a set threshold and remove face images with missing facial key points;
[0123] In this embodiment, the definition of the face image is determined by the resolution of the face image, and the face images with low definition are filtered out. The face stability is mainly determined by determining whether the face key points in the face image are missing to determine whether the face image is usable. A qualified face image must have left and right eyes, nose, eyebrows and upper and lower lips. In order to ensure that each face image can better reflect the speaker's facial expression during dialogue emotion recognition, only face images with more complete facial key points are used.
[0124] Step 3.2: Set the target size (w t ,h t ), scale the filtered face image to the target size (w t ,h t ), and then align the facial key points of the face image, where w t is the set width, h t is the set height;
[0125] In this embodiment, in order to make the faces in different face images have consistent postures and scales, the face image sequence S of the speaker after the screening obtained in the previous step is t Each face image in is scaled and facial key points are aligned. The image scaling is as follows: 1. Select the face image sequence S of the speaker t The resolution of the first face image in is the target resolution, and the width and height are (w t ,h t );2. The speaker's face image sequence S t Each face image in the image is scaled by the ResizeCrop method. The original width and height of the face image are (w o ,h o ).if Then the image is scaled to Then select all pixels for width and all pixels for height. arrive pixels, round() is the rounding function. Then the image is scaled to Then select the width arrive The facial key point alignment specifically aligns the positions of the left and right eyes to a standard position (such as the left eye is aligned to the middle left part of the image, and the right eye is aligned to the middle right part). This makes each facial image have a similar format when processed, and makes the faces in different images have similar display effects.
[0126] Step 4: Based on the multi-round dialogue dataset obtained in step 1 and the pre-processed speaker face image sequence, construct several training templates;
[0127] Step 4.1: Build a sentence vector library D;
[0128] Step 4.1.1: Based on the multi-round dialogue data set, obtain training samples and then construct a training set including several training samples;
[0129] The training sample includes a conversation history, a video corresponding to a target sentence, and an emotion label of the target sentence. The conversation history includes a plurality of sentences and a speaker corresponding to each sentence. The target sentence is the last sentence in the conversation history.
[0130] The method for acquiring the historical dialogue is as follows: for each multi-round dialogue in the multi-round dialogue data set, the first w consecutive sentences are sequentially taken as a dialogue history using a window of size w, where w=1, 2, ..., W, and W is the number of sentences in a multi-round dialogue, that is, each sentence in the multi-round dialogue is taken as the last sentence in a dialogue history;
[0131] For example, for a multi-round dialogue, first select the first sentence as a dialogue history, use the emotion label of the first sentence as the emotion label of the target sentence, then take the first two consecutive sentences as a dialogue history, and use the emotion label of the last sentence (the second sentence) as the emotion label of the target sentence, then take the first three consecutive sentences as a dialogue history, and use the emotion label of the last sentence (the third sentence) as the emotion label of the target sentence, and so on, all sentences are the last sentence of a dialogue history and serve as the target sentence;
[0132] Step 4.1.2: Preprocess the sentences in the training set to obtain preprocessed sentences;
[0133] The preprocessing includes: removing the speaker of the sentence, replacing the character name in the sentence, and then balancing the number of sentences with various emotion tags;
[0134] In this embodiment, the preprocessing includes: speaker information removal and emotional label balancing. First, extract base_size sentences of each emotion from the sentences in the training set (base_size depends on the specific situation of different data sets), and then anonymize the sentences to ensure that the emotional analysis will not be affected by the identity of the speaker. First, remove the speaker of the sentence, the identity of the speaker has no effect on the emotional analysis, and then use the synonym replacement method for the characters in the sentence, and replace the character names with labels such as "role A" and "role B" that are irrelevant to the speaker's identity.
[0135] Step 4.1.3: Encode the preprocessed sentences to obtain the feature vector of each sentence;
[0136] In this implementation, the preprocessed statements are processed using Sen t ence BERT encodes the sentence to obtain the feature vector of each sentence. The formula is as follows:
[0137] v u =SentenceBERT(u)
[0138] Among them, u is the preprocessed statement, v u is the feature vector obtained by encoding the sentence through Sentence BERT, with a dimension of 768;
[0139] Step 4.1.4: The preprocessed sentence u and its corresponding feature vector v u and the emotion label e as a piece of data in the sentence vector library, and construct a sentence vector library D including several pieces of data;
[0140] For preprocessed statements with small data volume, you can directly store them in JSON format. If the data volume is too large, you can use vector database storage to speed up data retrieval.
[0141] Step 4.2: For each target sentence in the training set, use the sentence similarity matching algorithm to retrieve the z sentences and their emotion labels that are most similar to the target sentence in the sentence vector library as candidate reference sentences;
[0142] Specifically, the cosine similarity between the feature vector of the target sentence and the feature vector of each other sentence in the sentence vector library except the sentence itself is calculated, and the z sentences with the highest cosine similarity and their emotion labels are taken as candidate reference sentences d h , h is the number of the candidate reference sentence;
[0143] In this embodiment, the target sentence u target Use Sentence BERT encoding to get the feature vector v target ;
[0144] v target =SentenceBERT(u target )
[0145] The calculation formula of cosine similarity is as follows:
[0146]
[0147] in, is the feature vector of the pth sentence in the sentence vector library, v targetis the feature vector of the target sentence, · is the dot product operation of the vector, |||| is the Euclidean norm of the vector, indicating the length of the vector, and the value of cosine similarity is between -1 and 1. The closer the value is to 1, the more similar the two sentences are, and the closer the value is to -1, the less similar the two sentences are.
[0148] In this embodiment, after calculating the cosine similarities between all sentences and the target sentence, these cosine similarity values are sorted from high to low, and the z sentences with the highest cosine similarities are selected as candidate reference sentences. i , in the specific implementation, z is 1. The selected candidate sentences are selected together with their emotion labels as a reference for subsequent emotion classification;
[0149] Step 4.3: Set the instruction description and construct several input templates based on the training samples, candidate reference sentences and the speaker's face image sequence obtained in step 3;
[0150] The input template includes instruction description ins, dialogue history h, target sentence u target , target face image R and candidate reference sentence d h ;
[0151] The instruction description is a short explanatory text that aims to guide the model to understand the goals and requirements of the task and help the model understand the conversation emotion recognition task. It provides the model with background information about the emotion recognition task, clarifies the composition of the input data, and describes the tasks that the model should complete when processing these inputs.
[0152] The conversation history is w consecutive sentences obtained by segmenting multiple rounds of conversations using a window; the size of the window is W;
[0153] Since the input sequence length supported by the large model is limited, the input multi-round dialogue is split, and the dialogue history of the target sentence and its previous w-1 sentences is retained to form a dialogue history with a window size of w. Different datasets use different window sizes. For the MELD and EmoryNLP datasets, w is 8, and for the IEMOCAP dataset, w is 12.
[0154] The target sentence is the sentence to be sentimentally classified, and is the last sentence in the conversation history; in the specific implementation, '<' and '>' are used as highlight marks.
[0155] The target face image is a face image at the time corresponding to the target sentence obtained from the face image sequence of the speaker, which removes useless information in the video;
[0156] The candidate reference sentences are composed of sentences most similar to the target sentence and their emotion labels;
[0157] In this implementation, a manual instruction description is written, and then GPT-3.5 is used to adjust and polish it, and several instruction descriptions are generated. Finally, GPT-3.5 is used to score the instruction description with the highest score for subsequent experiments. The final instruction description is as follows:
[0158] You are an emotional master.You are about to complete the task ofanalyzing emotions in the conversation.You will receive a conversationhistory,and your task is to classify emotions for the last sentence in theconversation.You will also receive a candidate instance for reference.
[0159] Step 4.4: Prompt the general large language model to perform the dialogue emotion recognition task on the dialogue history in the input template in the form of thought chain, and obtain several emotion reasoning chains;
[0160] Step 4.4.1: Different participants manually and carefully write several thought chain examples of conversation emotion recognition, where each thought chain example includes: conversation history, emotion classification thinking process and emotion classification results;
[0161] In this implementation, participants are required to complete the emotion classification thinking process from three aspects: dialogue topic summary, target speaker emotion fluctuation analysis, and clue sentence discovery in the dialogue that may inspire the target speaker to produce certain emotions, and finally obtain the target speaker's emotion classification result;
[0162] In this embodiment, given a number of selected conversation histories, the emotional intensity and text quality in the conversation are taken into consideration during selection, with the goal of improving the quality of data generated by GPT-3.5 based on contextual learning through high-quality examples.
[0163] Each conversation history was given to three experimenters to perform an emotion recognition task. They were required to record their thinking process, especially the summary of conversation topics, analysis of the target speaker's emotional fluctuations, and the reason sentences in the conversation that may stimulate the target speaker to produce certain emotions. This is called the emotion reasoning chain. Then the emotion classification result was obtained, and finally it was scored by scorers, and the result with the highest score was selected.
[0164] Step 4.4.2: Adjust the several thought chain examples obtained in step 4.4.1 according to the emotion classification results to ensure that the number of thought chain examples corresponding to various emotion classification results is balanced, and encode the dialogue history in each thought chain example through a sentence-level encoder to obtain a feature vector;
[0165] Step 4.4.3: Take the thought chain example and its corresponding feature vector as a data unit to form an example corpus including several data units;
[0166] In this implementation, Sentence BERT is used to encode the conversation history of the thought chain example to obtain a feature vector, and the example corpus is stored in JSON format.
[0167] In this implementation, Sentence BERT is used to encode the conversation history to obtain a feature vector:
[0168] v h =SentenceBERT(h)
[0169] Step 4.4.4: Use the sentence-level encoder to obtain a feature vector for each dialogue history in the input template, and calculate the cosine similarity with the feature vectors of the dialogue history in all thought chain examples in the example corpus except the thought chain example corresponding to the dialogue history, to obtain the thought chain example with the highest score;
[0170] The calculation formula is as follows:
[0171]
[0172] in is the feature vector of the conversation history of the qth example in the example corpus, v h is the feature vector of a given conversation history, · is the dot product operation of the vector, ‖‖ is the Euclidean norm of the vector, indicating the length of the vector;
[0173] Step 4.4.5: Take the historical dialogue in the input template as the question, and complete the dialogue emotion recognition task based on the large model (GPT-3.5) with the highest-scoring thought chain example prompt capability, and obtain the thought process and emotion classification results of emotion classification. The thought process of emotion classification is called emotion reasoning chain c;
[0174] In the specific implementation, the HTML-like tags are used to standardize the output of the model for each component of the example. Specifically, query:> and answer:> are used to represent questions and answers respectively. Questions refer to the conversation history, and answers refer to the emotion reasoning chain and emotion classification results. <reason>Tags and <result>The labels encapsulate the emotion reasoning chain and emotion classification results respectively, which makes the reasoning process of the model clearer, easier to understand, and convenient for parsing.
[0175] Step 4.4.6: Repeat steps 4.4.4 to 4.4.5 to obtain several emotion reasoning chains c and emotion classification results;
[0176] Step 4.5: Take an emotion reasoning chain c and the corresponding emotion label of the historical dialogue as an output template, and then obtain several output templates;
[0177] Step 4.6: Take a set of input templates and their corresponding output templates as a training template, and then obtain several training templates;
[0178] Step 5: Use several training templates to fine-tune the multimodal instructions of the multimodal large model (LMM) to obtain a fine-tuned multimodal large model;
[0179] The training target is the emotion reasoning chain c and the true emotion label e. After integrating the above inputs and feeding them into the model, the model calculation output is represented by the following equation:
[0180] y m =LMM(x m ,θ)
[0181] Among them, y m represents the text generated by the large multimodal model (LMM) (emotional reasoning chain and true emotion label), x m Represents the formatted input of a large multimodal model (LMM), θ represents the parameters of the LMM, and the LMM prediction ends at the output <eos>The conditional probability p(γ|x) of generating y before each token γ m ,θ).
[0182] The next token prediction loss is used to measure the output error of the multimodal large model (LMM), and the loss function is defined as follows:
[0183]
[0184] Where γ represents the next token generated.
[0185] In this embodiment, the emotion reasoning chain and real emotion labels generated by GPT-3.5 are used as training targets during the training phase. During reasoning, the model can automatically generate emotion reasoning chains and emotion classification results. That is, only the emotion reasoning chain generated by GPT-3.5 is used as external knowledge to enhance the data set and then train the emotion reasoning ability of the model, without the need for external knowledge in the reasoning phase.
[0186] Step 6: Use the fine-tuned multimodal large model to perform emotion recognition on the conversation video to be recognized;
[0187] Step 6.1: Extract the conversation video slice corresponding to the last sentence from the conversation video, and obtain the face image sequence of the speaker in the conversation video slice corresponding to the last sentence according to the method of steps 2-3;
[0188] Step 6.2: Use the sentences in the conversation video to form a conversation history, and use the last sentence as the target sentence;
[0189] Step 6.3: Obtain candidate reference sentences from the sentence vector library according to step 4.2;
[0190] Step 6.4: Take the face image of the speaker at the time corresponding to the target sentence in the face image sequence as the target face image;
[0191] Step 6.5: Input the instruction description, dialogue history, target sentence, target face image and candidate reference sentence into the fine-tuned multimodal large model to obtain the emotion reasoning chain c and emotion classification results.< / eos> < / result> < / reason>
Claims
1. A multimodal conversation emotion recognition method based on face recognition and emotion reasoning chain, characterized in that: The specific steps include: Step 1: Obtain a multi-round dialogue dataset; the multi-round dialogue dataset includes several multi-round dialogues, each of which includes several sentences, the speaker of each sentence, the emotion label corresponding to each sentence, and the dialogue video corresponding to each sentence; Step 2: Determine the speaker in the conversation video corresponding to each sentence in the multi-round conversation dataset, and then obtain the speaker's face image sequence; Step 3: preprocessing the speaker's facial image sequence to obtain a preprocessed facial image sequence; Step 4: Based on the multi-round dialogue dataset obtained in step 1 and the pre-processed speaker face image sequence, construct several training templates; Step 5: Use several training templates to fine-tune the multimodal instructions of the multimodal large model to obtain a fine-tuned multimodal large model; Step 6: Use the fine-tuned multimodal large model to perform emotion recognition on the conversation video to be identified.
2. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 1 is characterized in that: Step 2 specifically includes: Step 2.1: Sampling the conversation video corresponding to each sentence to obtain a sampled image sequence; the image sequence includes the first video frame, the last video frame and a number of randomly sampled video frames of the conversation video; The random sampling method is specifically as follows: reading the conversation video, constructing a frame list including all video frames of the conversation video, randomly generating an index list, sampling video frames from the frame list through the index list, and obtaining an image sequence F of length n; Step 2.2: Use the face detection neural network to calculate the position information of each face in each video frame of the image sequence, and determine the number of people in the video conversation by the number of faces in the first video frame; Step 2.3: Based on the obtained face position information, all face images of each video frame of the image sequence are cropped, and the face images of different people are classified by the k-means algorithm according to the determined number of people in the video conversation, and the face images of the same person are sorted in chronological order to obtain a sequence of face images of different people; Step 2.4: Use a facial key point detection tool to determine the speaker, and then determine the speaker's facial image sequence from facial image sequences of different persons.
3. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 2 is characterized in that: Step 2.2 specifically includes: Step 2.2.1: For each video frame f in the image sequence F i , use the face detection neural network to calculate the face position and get the position information c of each face i,j , where i is the number of the video frame and j is the number of the face; The position information includes the coordinates of the upper left corner of the face and the height and width of the face frame; Step 2.2.2: According to the number of face position information output by the face detection network in the first video frame, determine the number of faces k in the first video frame and use it as the number of people in the video conversation.
4. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 2 is characterized in that: Step 2.3 specifically includes: Step 2.3.1: For each video frame f in the image sequence F i , according to the position information of the face in the video frame, the face image is cut out, and all face images constitute the face image set S face ; Step 2.3.2: For the face image set S face For each face image in , use a deep learning neural network to calculate the feature vector v of the face image; Step 2.3.3: Initialize the feature vectors of the k face images of the first video frame in the image sequence as k cluster centers μ; Step 2.3.4: For the remaining face images except the k face images of the first video frame, calculate the Euclidean distance between the feature vector v of each face image and each cluster center μ, assign the face image to the cluster with the cluster center closest to its Euclidean distance, and obtain k clusters; Step 2.3.5: For each cluster, update the cluster center according to the feature vector of the face image in the cluster; Step 2.3.6: Determine whether the cluster center remains unchanged within the set round or whether the number of iterations reaches the set number of iterations. If so, execute step 2.3.8, otherwise execute step 2.3.7; Step 2.3.7: According to the updated cluster centers, redistribute all face images to obtain k clusters and return to step 2.3.5; Step 2.3.8: Sort the face images in each cluster according to the time sequence in the conversation video, and redistribute the clusters with multiple face images or missing face images at a certain time point, and finally obtain a face image sequence S of k people; Specifically: for a cluster with redundant face images at a certain point in time, the Euclidean distance between all face images contained in the cluster at that point in time and other face images in the cluster is calculated, and the face image with the smallest average Euclidean distance is retained. The other face images will be reallocated, and the Euclidean distance between the face image to be reallocated and the cluster center of all clusters that lack face images at this point in time is calculated, and the face image is reallocated to the cluster with the smallest Euclidean distance.
5. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 2 is characterized in that: Step 2.4 specifically includes: Step 2.4.1: For each person's facial image sequence, use a facial key point detection tool to detect the coordinates of the facial key points of each facial image, and calculate the coordinate distance d between the upper lip and the lower lip; the facial key points include the upper lip and the lower lip; Step 2.4.2: According to the coordinate distance d between the upper lip and the lower lip in each face image, calculate the variance σ of the coordinate distance between the upper lip and the lower lip in each person's face image sequence; Step 2.4.3: Take the face image sequence with the largest variance σ as the speaker’s face image sequence S t .
6. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 1 is characterized in that: Step 3 specifically includes: Step 3.1: screening the face images in the face image sequence of the speaker to obtain a screened face image sequence of the speaker; Specifically: remove face images with resolutions below a set threshold and remove face images with missing facial key points; Step 3.2: Set the target size (w t ,h t ), scale the filtered face image to the target size (w t ,h t ), and then align the facial key points of the face image, where w t is the set width, h t is the set height.
7. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 1 is characterized in that: Step 4 specifically includes: Step 4.1: Build a sentence vector library D; Step 4.2: For each target sentence in the training set, use the sentence similarity matching algorithm to retrieve the z sentences and their emotion labels that are most similar to the target sentence in the sentence vector library as candidate reference sentences; Specifically, the cosine similarity between the feature vector of the target sentence and the feature vector of each other sentence in the sentence vector library except the sentence itself is calculated, and the z sentences with the highest cosine similarity and their emotion labels are taken as candidate reference sentences d h , h is the number of the candidate reference sentence; Step 4.3: Set the instruction description and construct several input templates based on the training samples, candidate reference sentences and the speaker's face image sequence obtained in step 3; The input template includes instruction description ins, dialogue history h, target sentence u target , target face image R and candidate reference sentence d h ; The instruction manual provides background information of the emotion recognition task, clarifies the composition of the input data, and describes the tasks to be completed when processing the input data; The target face image is a face image at the time corresponding to the target sentence obtained from the speaker's face image sequence; Step 4.4: Prompt the general large language model to perform the dialogue emotion recognition task on the dialogue history in the input template in the form of thought chain, and obtain several emotion reasoning chains; Step 4.5: Take an emotion reasoning chain c and the corresponding emotion label of the historical dialogue as an output template, and then obtain several output templates; Step 4.6: Take a set of input templates and their corresponding output templates as a training template, and then obtain several training templates.
8. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 7 is characterized in that: Step 4.1 specifically includes: Step 4.1.1: Based on the multi-round dialogue data set, obtain training samples and then construct a training set including several training samples; The training sample includes a conversation history, a video corresponding to a target sentence, and an emotion label of the target sentence. The conversation history includes a plurality of sentences and a speaker corresponding to each sentence. The target sentence is the last sentence in the conversation history. The method for acquiring the historical dialogue is as follows: for each multi-round dialogue in the multi-round dialogue data set, the first w consecutive sentences are sequentially taken as a dialogue history using a window of size w, where w=1, 2, ..., W, and W is the number of sentences in a multi-round dialogue, that is, each sentence in the multi-round dialogue is taken as the last sentence in a dialogue history; Step 4.1.2: Preprocess the sentences in the training set to obtain preprocessed sentences; The preprocessing includes: removing the speaker of the sentence, replacing the character name in the sentence, and then ensuring that the difference in the number of sentences of various emotion tags is within a set range; Step 4.1.3: Encode the preprocessed sentences to obtain the feature vector of each sentence; Step 4.1.4: The preprocessed sentence u and its corresponding feature vector v u and the emotion label e as a piece of data in the sentence vector library, and construct a sentence vector library D including several pieces of data.
9. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 7 is characterized in that: Step 4.4 specifically includes: Step 4.4.1: Write several thought chain examples for conversation emotion recognition, where each thought chain example includes: conversation history, emotion classification thinking process, and emotion classification results; Adjust the several thought chain examples obtained in step 4.4.1 according to the emotion classification results to ensure that the difference in the number of thought chain examples corresponding to various emotion classification results is within the set range, and encode the dialogue history in each thought chain example through a sentence-level encoder to obtain a feature vector; Step 4.4.3: Take the thought chain example and its corresponding feature vector as a data unit to form an example corpus including several data units; Step 4.4.4: Use the sentence-level encoder to obtain a feature vector for each dialogue history in the input template, and calculate the cosine similarity with the feature vectors of the dialogue history in all thought chain examples in the example corpus except the thought chain example corresponding to the dialogue history, to obtain the thought chain example with the highest score; Step 4.4.5: Take the historical dialogue in the input template as a question, and prompt the large model to complete the dialogue emotion recognition task based on the highest-scoring thought chain example, and obtain the emotion classification thinking process and emotion classification results. The emotion classification thinking process is called emotion reasoning chain c; Step 4.4.6: Repeat steps 4.4.4 to 4.4.5 to obtain several emotion reasoning chains c and emotion classification results.
10. The multimodal dialogue emotion recognition method based on face recognition and emotion reasoning chain according to claim 1, characterized in that: Step 6 specifically includes: Step 6.1: Extract the conversation video slice corresponding to the last sentence from the conversation video, and obtain the face image sequence of the speaker in the conversation video slice corresponding to the last sentence according to the method of steps 2-3; Step 6.2: Use the sentences in the conversation video to form a conversation history, and use the last sentence as the target sentence; Step 6.3: Obtain candidate reference sentences from the sentence vector library according to step 4.2; Step 6.4: Take the face image of the speaker at the time corresponding to the target sentence in the face image sequence as the target face image; Step 6.5: Input the instruction description, dialogue history, target sentence, target face image and candidate reference sentence into the fine-tuned multimodal large model to obtain the emotion reasoning chain c and emotion classification results.