A Dialogue Emotion Recognition Method and System Based on Dynamic Modeling of Personality

CN117195914BActive Publication Date: 2026-09-01QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202311247166.9
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-09-25
Publication Date
2026-09-01
Estimated Expiration
2043-09-25

AI Technical Summary

Technical Problem

然而这种上下文背景下我们仍很难判断男女双方具体的情绪是怎样的,即我们在无法获得双方性格的前提下,难以依据话语的文本内容判断说话人对某件事情的具体看法与情绪表达

Benefits of technology

[0040] (1) This disclosure provides a dialogue emotion recognition method and system based on dynamic modeling of personality. The scheme establishes the group personality and individual personality of the participants in the dialogue by using the proposed personality extraction module through historical conversations, and continuously updates the characteristics of the group personality and individual personality according to the progress of the dialogue. At the same time, in the emotional expression of each sentence, personality factors are used through the proposed personality participation module to influence the emotional expression tendency of the speech. In the whole process, the personality extraction module and the personality participation module are stacked to update the personality characteristics and conversation expression in a layered manner, thereby achieving more interpretable dialogue emotion recognition and more accurate emotion classification effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117195914B_ABST
    Figure CN117195914B_ABST
Patent Text Reader

Abstract

This disclosure provides a dialogue sentiment recognition method and system based on dynamic character modeling, comprising: acquiring historical conversations to be sentiment-recognized and extracting feature vectors for each utterance; based on the extracted feature vectors, acquiring the global context representation and individual context representation corresponding to each utterance; based on the global context representation of each utterance and the attention weight of each sentence in the group personality, obtaining the group personality characteristics of the speakers in the historical conversation; based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality, obtaining the individual personality characteristics of the speakers in the historical conversation; integrating the obtained group personality characteristics and individual personality characteristics into the utterance representation to obtain the utterance representation with group personality participation and the utterance representation with individual personality participation; and based on the fusion features of the obtained utterance representation and the initial context representation, using a multilayer perceptron to obtain the sentiment label of each utterance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure belongs to the field of emotion recognition technology, and in particular relates to a dialogue emotion recognition method and system based on dynamic modeling of personality. Background Technology

[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.

[0003] Emotion is an important biological attribute of humans, and dialogue is a common vehicle for expressing this attribute. Emotion recognition in dialogue aims to identify the personal emotions hidden behind each sentence in a conversation. This is a highly challenging task, and with the development of smart devices and the increase in dialogue scenarios, it has attracted more and more attention from researchers. However, emotion recognition in speech differs from conventional text emotion recognition. In text emotion recognition, only the text content needs to be considered; while dialogue, as a form of social interaction, naturally results in differences in how different individuals express emotions. Similar text content will produce different emotional expressions in different individuals, making it difficult to correctly distinguish the emotions in speech. This difference in the expression of emotions in speech is largely influenced by a person's fixed behavioral patterns and personal traits, which are attributed to personality factors in psychology. However, in recent work, many methods have ignored the role of personality in the expression of emotions in speech. Most methods focus on obtaining contextual background information. Ghosa et al. first proposed the DialogueGCN model, which uses graph convolutional networks (GCN) to capture remote contextual information in dialogue. DialogueGCN treats each utterance as a node and connects any node in the same window of the conversation. Hu et al. built a fully connected graph across multimodal discourses, using deep graph convolution pairs for representation updates. Shen et al. built graph structures based on different speakers and achieved excellent results by using recurrent neural networks in the graph fusion process. Li et al. used dynamic fusion to avoid information redundancy and achieved multimodal information complementarity in dialogue. Sun et al. built a more interpretable graph network by utilizing the discourse structure between discourses to directly obtain remote context. Besides the text itself, the contextual background information of discourse also includes a lot of emotional information from external knowledge. For example, Ghosa et al. extracted common sense information from discourse and integrated it into the dialogue to improve performance, while Li et al. improved the edge information in the graph structure during knowledge integration. Xie et al. used more fine-grained external knowledge to supplement the discourse content.

[0004] The inventors discovered that while contextual information is crucial, in real-world dialogue scenarios, even with similar backgrounds and textual information, different speakers can exhibit different emotional expressions. Neither recurrent neural network-based methods nor graph-based modeling can extract the speaker's personality traits from the discourse, thus limiting the accuracy and interpretability of discourse emotion recognition. Figure 1 As shown, based on the context, we know that the female character was unfairly treated and dismissed, while the male character expressed dissatisfaction with this outcome. However, even in this context, it is still difficult to determine the specific emotions of both parties. That is, without knowing their personalities, it is difficult to judge the speaker's specific views and emotional expressions based solely on the textual content of the discourse. Summary of the Invention

[0005] To address the aforementioned problems, this disclosure provides a dialogue sentiment recognition method and system based on dynamic character modeling. The scheme utilizes a proposed personality extraction module to establish the group and individual personalities of participants in a given dialogue session based on historical conversations, and continuously updates the characteristics of both group and individual personalities as the dialogue progresses. Simultaneously, in the sentiment representation of each sentence, personality factors are incorporated through a proposed personality participation module to influence the emotional expression tendency of the discourse. Furthermore, throughout the process, the personality extraction module and the personality participation module are stacked, allowing for hierarchical updates of personality features and dialogue representations, achieving more interpretable dialogue sentiment recognition and more accurate sentiment classification.

[0006] According to a first aspect of the embodiments of this disclosure, a dialogue emotion recognition method based on dynamic modeling of personality is provided, comprising:

[0007] Obtain the historical conversations to be sentiment identified, and extract the feature vector of each utterance in the historical conversations;

[0008] Based on the feature vector of each utterance extracted from the historical conversation, the global context representation and individual context representation corresponding to each utterance in the current historical conversation are obtained respectively;

[0009] Based on the global context representation of each utterance and the attention weight of each sentence in the group personality, the group personality characteristics of the speakers in the historical conversation are obtained; based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality, the individual personality characteristics of the speakers in the historical conversation are obtained.

[0010] The acquired group personality traits and individual personality traits are integrated into the discourse representation to obtain discourse representations involving group personality traits and discourse representations involving individual personality traits.

[0011] Based on the fusion features of discourse representations involving group personality, initial global context representation, discourse representations involving individual personality, and initial individual context representation, a multilayer perceptron is used to obtain the sentiment label of each discourse.

[0012] Furthermore, the step of obtaining the global context representation and individual context representation corresponding to each utterance in the current historical session is specifically as follows: two pre-trained bidirectional long short-term memory networks are used to model the global utterance order dependency and the individual utterance order dependency, respectively. Specifically, for modeling the global utterance order dependency, the input of the bidirectional long short-term memory network is the text features of each utterance in the historical session; for modeling the individual utterance order dependency, the input of the bidirectional long short-term memory network is the text features of adjacent utterances from the same speaker.

[0013] Furthermore, the acquisition of the attention weight for each sentence is specifically as follows:

[0014] The attention weight of each sentence in the group personality is calculated using the following formula:

[0015]

[0016] in, This provides a global context representation for each utterance in the previous layer. and These are training parameters, i,j∈[0,t-1], representing all utterances before time t; D is the hidden representation of the group personality at time t-1 in the l-th iteration layer. h With D g Both represent vector dimensions;

[0017] The attention weight of each sentence in an individual's personality is calculated using the following formula:

[0018]

[0019] in, This provides a speaker-level context representation for each utterance in the higher-level context. This represents the hidden representation of an individual's personality at time t-1 in the l-th iteration layer. and These are the training parameters, D h With D g Both represent vector dimensions.

[0020] Furthermore, the feature vector of each utterance in the historical conversation is extracted using a pre-trained language model, RoBERTa.

[0021] Furthermore, the process of integrating the obtained group personality characteristics and individual personality characteristics into the discourse representation to obtain discourse representations involving group personality and individual personality specifically involves: integrating group personality characteristics and individual personality characteristics into the discourse representation based on the update and reset mechanism of the gated loop unit.

[0022] Furthermore, the integration of the obtained group personality characteristics into the discourse representation specifically involves:

[0023] Calculate the update gate and reset gate based on the current input utterance and group personality characteristics;

[0024] By resetting the gate, the newly entered utterances are combined with the previous group personality traits;

[0025] The group's personality traits are passed to the next layer by using an update gate, and the dialogue is updated hierarchically by using a stacking method. The stacking method is as follows: the utterance representation generated by layer l is used as the initial input utterance of layer l+1.

[0026] Alternatively, the integration of the obtained individual personality traits into the discourse representation specifically includes:

[0027] Calculate update and reset gates based on individual historical discourse and personality traits;

[0028] By resetting the gate, the newly entered words are combined with the previous individual personality traits;

[0029] The update gate is used to pass individual personality traits to the next layer, and the dialogue is updated hierarchically using a stacking method. The stacking method is as follows: the utterance representation generated by layer l is used as the initial input utterance of layer l+1.

[0030] Furthermore, the fusion feature based on the discourse representation involving group personality, the initial global context representation of the discourse, the discourse representation involving individual personality, and the initial individual context representation of the discourse specifically involves: adding the discourse representation involving group personality to the initial global context representation of the discourse; simultaneously, adding the discourse representation involving individual personality to the initial individual context representation of the discourse, and concatenating the discourse representation obtained by adding the two parts to obtain the fusion feature of each discourse.

[0031] According to a second aspect of the present disclosure, a dialogue emotion recognition system based on dynamic modeling of personality traits is provided, comprising:

[0032] The discourse feature extraction unit is used to acquire the historical conversations to be sentiment identified and extract the feature vector of each discourse in the historical conversations.

[0033] The global and individual feature extraction units are used to obtain the global context representation and individual context representation corresponding to each utterance in the current historical session based on the feature vector of each utterance extracted in the historical session.

[0034] The personality extraction unit is used to obtain the group personality characteristics of the speakers in the historical conversation based on the global context representation of each utterance and the attention weight of each sentence in the group personality; and to obtain the individual personality characteristics of the speakers in the historical conversation based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality.

[0035] The personality participation unit is used to integrate the acquired group personality characteristics and individual personality characteristics into the discourse representation, thereby obtaining discourse representations with group personality participation and discourse representations with individual personality participation.

[0036] The emotion recognition unit is used to obtain the emotion label of each utterance based on the fusion features of utterance representations with group personality participation, initial global context representation, utterance representations with individual personality participation, and initial individual context representation, using a multilayer perceptron.

[0037] According to a third aspect of the present disclosure, an electronic device is provided, including a memory, a processor, and a computer program stored in the memory and running on the memory, wherein the processor executes the program to implement the aforementioned dialogue emotion recognition method based on dynamic modeling of personality.

[0038] According to a fourth aspect of the present disclosure, a non-transitory computer-readable storage medium is provided, on which a computer program is stored, which, when executed by a processor, implements the aforementioned dialogue emotion recognition method based on dynamic modeling of personality.

[0039] Compared with the prior art, the beneficial effects of this disclosure are:

[0040] (1) This disclosure provides a dialogue emotion recognition method and system based on dynamic modeling of personality. The scheme establishes the group personality and individual personality of the participants in the dialogue by using the proposed personality extraction module through historical conversations, and continuously updates the characteristics of the group personality and individual personality according to the progress of the dialogue. At the same time, in the emotional expression of each sentence, personality factors are used through the proposed personality participation module to influence the emotional expression tendency of the speech. In the whole process, the personality extraction module and the personality participation module are stacked to update the personality characteristics and conversation expression in a layered manner, thereby achieving more interpretable dialogue emotion recognition and more accurate emotion classification effect.

[0041] (2) The scheme described in this disclosure extracts the group and individual personalities of the participants in the dialogue, enabling them to participate in the emotional expression of the discourse. Extensive experimental results show that the scheme proposed in this disclosure achieves performance exceeding the baseline. Furthermore, through comprehensive evaluation and ablation studies, the advantages of the scheme described in this disclosure and the influence of different modules have been confirmed, improving the performance of discourse emotion recognition.

[0042] Advantages of this disclosure in some respects will be set forth in the description which follows, and in part will be obvious from the description, or may be learned by practice of this disclosure. Attached Figure Description

[0043] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.

[0044] Figure 1 This is an example of a segment of a dialogue from the I EMOCAP dataset described in the background section of this disclosure;

[0045] Figure 2 This is an overall flowchart of a dialogue emotion recognition method based on dynamic modeling of personality as described in the embodiments of this disclosure;

[0046] Figure 3 This is a schematic diagram of the personality discourse fusion module structure described in the embodiments of this disclosure. Detailed Implementation

[0047] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0048] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.

[0049] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0050] Where there is no conflict, the embodiments and features described herein can be combined with each other.

[0051] Example 1:

[0052] The purpose of this embodiment is to provide a dialogue emotion recognition method based on dynamic modeling of personality.

[0053] A dialogue sentiment recognition method based on dynamic character modeling includes:

[0054] Obtain the historical conversations to be sentiment identified, and extract the feature vector of each utterance in the historical conversations;

[0055] Based on the feature vector of each utterance extracted from the historical conversation, the global context representation and individual context representation corresponding to each utterance in the current historical conversation are obtained respectively;

[0056] Based on the global context representation of each utterance and the attention weight of each sentence in the group personality, the group personality characteristics of the speakers in the historical conversation are obtained; based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality, the individual personality characteristics of the speakers in the historical conversation are obtained.

[0057] The acquired group personality traits and individual personality traits are integrated into the discourse representation to obtain discourse representations involving group personality traits and discourse representations involving individual personality traits.

[0058] Based on the fusion features of discourse representations involving group personality, initial global context representation, discourse representations involving individual personality, and initial individual context representation, a multilayer perceptron is used to obtain the sentiment label of each discourse.

[0059] In specific implementation, the step of obtaining the global context representation and individual context representation corresponding to each utterance in the current historical session is as follows: Two pre-trained bidirectional long short-term memory networks are used to model the global utterance order dependency and the individual utterance order dependency, respectively. Specifically, for modeling the global utterance order dependency, the input of the bidirectional long short-term memory network is the text features of each utterance in the historical session; for modeling the individual utterance order dependency, the input of the bidirectional long short-term memory network is the text features of adjacent utterances from the same speaker.

[0060] In practice, the feature vector of each utterance in the historical conversation is extracted using a pre-trained language model, RoBERTa.

[0061] In specific implementation, the process of integrating the obtained group personality characteristics and individual personality characteristics into the discourse representation to obtain discourse representations with the participation of group personality and discourse representations with the participation of individual personality specifically involves: integrating group personality characteristics and individual personality characteristics into the discourse representation based on the update and reset mechanism of the gated loop unit.

[0062] In specific implementation, the integration of the obtained group personality characteristics into the discourse representation specifically includes:

[0063] Calculate the update gate and reset gate based on the current input utterance and group personality characteristics;

[0064] By resetting the gate, the newly entered utterances are combined with the previous group personality traits;

[0065] The group's personality traits are passed to the next layer by using an update gate, and the dialogue is updated hierarchically by using a stacking method. The stacking method is as follows: the utterance representation generated by layer l is used as the initial input utterance of layer l+1.

[0066] Alternatively, the integration of the obtained individual personality traits into the discourse representation specifically includes:

[0067] Calculate update and reset gates based on individual historical discourse and personality traits;

[0068] By resetting the gate, the newly entered words are combined with the previous individual personality traits;

[0069] The update gate is used to pass individual personality traits to the next layer, and the dialogue is updated hierarchically using a stacking method. The stacking method is as follows: the utterance representation generated by layer l is used as the initial input utterance of layer l+1.

[0070] In specific implementation, the fusion feature based on the discourse representation with group personality participation, the initial global context representation of the discourse, the discourse representation with individual personality participation, and the initial individual context representation of the discourse specifically involves: adding the discourse representation with group personality participation to the initial global context representation of the discourse; simultaneously, adding the discourse representation with individual personality participation to the initial individual context representation of the discourse, and concatenating the discourse representation obtained by adding the two parts to obtain the fusion feature of each discourse.

[0071] Specifically, for ease of understanding, the following detailed description of the solution in this embodiment is provided in conjunction with the accompanying drawings:

[0072] This embodiment provides a dialogue sentiment recognition method based on dynamic character modeling. In this method, we establish the group and individual personalities of participants in a dialogue segment using a proposed personality extraction module, based on historical conversations, and continuously update the characteristics of both group and individual personalities as the dialogue progresses. Simultaneously, in the sentiment representation of each sentence, we influence the emotional expression tendency of the utterances through a proposed personality participation module. Furthermore, throughout the process, we stack the personality extraction module and the personality participation module, enabling hierarchical updates of personality features and conversation representations, achieving more interpretable dialogue sentiment recognition and more accurate sentiment classification.

[0073] Specifically, such as Figure 2 As shown, the overall framework of the method described in this embodiment mainly includes the following three modules: a discourse feature extractor, a hierarchical fusion module for personality conversations, and a sentiment prediction module. The following is a detailed description of each of the three modules:

[0074] (I) Problem Definition

[0075] The task data consists of a transcript of a dialogue and speaker information for each constituent utterance. The aim is to identify the emotion of each utterance from a predefined set of emotions. In dialogue emotion recognition, the data comprises multiple dialogue segments {c1, c2, ..., c...}. n Each dialogue consists of several utterances. i =[u1,u2,...,u m ] and emotional tags Composition, where S represents the category of emotion. For a discourse, it consists of a series of words u i ={w1,w2,...,w L} Composed of. Session c i Each utterance in the string is made by a specific speaker and can be represented as p(c i )=[p(u1),p(u2),...,p(u r )] and p(u i )∈P, where P represents the speaker's category or name. Therefore, the entire problem can be represented as obtaining the sentiment label for each utterance based on the context and speaker information within a dialogue:

[0076] (II) Discourse Feature Extractor Module

[0077] We employ the widely used pre-trained language model RoBERTa to perform utterance-level feature extraction. RoBERTa Large follows the original BERT Large architecture, with 24 layers, 16 self-attention heads in each block, and a hidden layer dimension of 1024. Specifically, for each utterance u... i ={w1,w2,...,w L We use a special marker [CLS] to connect to the beginning of the utterance. Then we form the sequence {[CLS], w1, w2, ..., w L The input is fed into the RoBERTa model, and the [CLSW] labels from the last hidden layer are fed into a pooling layer to obtain the sentiment classification results. This completes the fine-tuning of the pre-trained RoBERTa model for the discourse-level sentiment classification task. After the fine-tuning process, in order to obtain the discourse-level feature vector u′ represented by the [CLSW] labels... iWe use {[CLS],w1,w2,···,w L Enter each utterance in the same input format:

[0078] u′ i =RoBERTa([CLS],w1,w2,...,w L (1)

[0079] in And d m This involves calculating the dimensions of the hidden states in RoBERTa and extracting the [CLS] token representation from the last layer to obtain the utterance-level feature vector for each utterance. Therefore, each session c is obtained. i The vectorized representation is {u′1, u′2, ..., u′} m}

[0080] In dialogue sentiment recognition, sequence relations provide context and development of the dialogue, helping the model eliminate ambiguity in certain statements and better understand the emotional changes of participants. We use a separate encoder to encode features for the perceived context of the dialogue text. To facilitate the extraction of personality profiles at both the group and individual levels, we utilize two bidirectional LSTMs in this section to model global discourse sequence dependencies and individual discourse sequence dependencies, respectively.

[0081] Specifically, in modeling global discourse order dependencies, the input is the text features of each discourse. Its global-level context representation is calculated as follows:

[0082]

[0083] To learn individual utterance order dependencies, we also used another bidirectional LSTM network to capture the mutual dependencies between adjacent utterances of the same speaker. Given the text features u′ for each utterance... i The individual-level context representation is computed as:

[0084]

[0085] Where p i ∈p(u i ), which represents the speaker p in the dialogue.

[0086] (III) Hierarchical Integration Module of Personality Conversation

[0087] (a) Integration of global personality discourse

[0088] In this step, to simulate the influence of group personality on the emotional expression of discourse in real-world dialogues and to enable the group personality of the speakers to participate in the updating of discourse representation, we set up a global personality-discourse fusion module to realize the interaction between discourse and group personality. This module consists of two main parts: personality extraction and personality participation, such as... Figure 3 As shown, we adopt the user interest modeling method used by NPA et al. in news recommendation to extract speaker personalities from the discourse, and use a gating mechanism to integrate speaker personalities into the discourse representation. Finally, we achieve hierarchical updates by stacking global personality-discourse fusion modules. In the first layer, we initialize the input discourse as follows:

[0089]

[0090] Next, we will describe in detail the specific implementation details of these two parts using the calculation steps of layer l.

[0091] 1) Group Personality Extraction Module

[0092] At this stage, we aim to learn from historical discourse the representation of the speaker group's character, such as... Figure 2 As shown on the left. By constructing personalized attention modules from historical dialogues, we obtain the collective personality characteristics of the speakers in the historical discourse. That is, in a dialogue, we perceive the speaking personalities of those around us, which in turn influences our own emotional expression. If the participants in the overall dialogue are open and extroverted, we may be more inclined to express emotions actively and directly. Conversely, if the participants are introverted and reserved, we may be more inclined to remain calm or subtle. To simulate the collective personality of the speakers in the dialogue, we use attention to fuse historical discourse into personality representations. We first initialize the collective personality vector of all speakers in the dialogue, represented as:

[0093]

[0094] in It represents the final expression of the personality of the group at the next higher level.

[0095] Note that the initialization of the group personality embedding in the first layer is as follows:

[0096]

[0097] Where k p Let r represent the initial embedding representation of the p-th speaker in the dialogue, and r represent the number of speakers in the dialogue.

[0098] Next, we will continuously update the group personality representation based on the progress of the conversation, specifically as follows:

[0099]

[0100] in, and These are the training parameters, D g This is the dimension for querying group personality, where l represents the number of layers in the iteration layer.

[0101] We represent the attention weight of the i-th sentence as α. i It is calculated by assessing the importance of the interaction between group personality queries and discourse representations, as shown below:

[0102]

[0103]

[0104] in and These are the training parameters, i,j∈[0,t-1], representing all utterances before time t. The final group personality representation of this layer at time t. It is the sum of historical dialogue contexts weighted by their level of attention:

[0105]

[0106] 2) Group personality participation module

[0107] In everyday conversations, the collective personality of speakers changes as the conversation progresses. Furthermore, because dialogue is a form of social communication, the verbal interactions between participants can influence each other's emotions and personality expressions. When people engage in conversation and communicate with others, this interaction leads to subtle changes in the speaker's understanding of collective personality. This dynamically changing collective personality then influences the speaker's subsequent emotional expression. For example... Figure 2 As shown on the right, to reconstruct this form of influence in the model, we borrow the update and reset mechanism of GRU units in this step to realize the influence of group personality on the emotional expression of discourse. First, we calculate the update gate and reset gate based on the current input discourse and the representation of group personality features, specifically:

[0108]

[0109]

[0110] Among them, W z and W r Both are training weight matrices, z t and r t These are the update door and the reset door, respectively.

[0111] Next, the reset gate is used to combine the newly input utterance with the previous group personality, which helps to capture important information in the current utterance that is influenced by the group personality. The specific calculation is as follows:

[0112]

[0113] Where W is the training weight matrix.

[0114] The next step uses an update gate to control the extent to which group personality information is incorporated into the current discourse state. This helps the model determine how much group personality information to pass to the next layer and be used to update the current discourse representation, calculated as follows:

[0115]

[0116] in This represents the discourse at time t in layer l.

[0117] Meanwhile, to achieve parameter sharing and improve generalization ability in the model, and to enable it to learn deep abstract representations of utterances and global features, we set up a stacked global personality utterance fusion module to update the dialogue hierarchy. That is, the utterance representation generated by layer l serves as the initial input utterance for layer l+1.

[0118] (b) Integration of individual personality discourse

[0119] In group personality discourse fusion, we update discourse representation using the group personality of the speakers. However, individual personality traits also influence the intensity and manner of emotional expression. For example, some people may be more inclined to express emotions strongly, while others may be more inclined to remain calm or control their emotions. An individual's personality type, to a certain extent, determines the directness, intensity, or restraint of their emotional expression. To enable individual personality to participate in the emotional expression of discourse in the model, we set up an individual personality discourse fusion module. Its steps are similar to those of group personality discourse fusion. Similarly, this module consists of two main parts: personality extraction and personality participation. In the first layer, we initialize the input discourse as follows:

[0120]

[0121] Next, we will describe in detail the implementation details of these two parts using the calculation steps of the lth layer in the individual personality discourse fusion module.

[0122] 1) Individual Personality Extraction Stage

[0123] In dialogues, individuals' personalities lead to significant differences in how they express the same emotion. For example, some individuals may tend to use strongly emotional language, while others prefer a softer approach. In this stage, to obtain the individual personality of each participant in the dialogue, we follow the implementation details of group personality extraction in group personality discourse fusion. We construct the individual personality of each speaker in the dialogue using their personal historical discourse. We first initialize the individual personality vector of each speaker in the dialogue, represented as:

[0124]

[0125] in This represents the final expression of the speaker's personality at the next higher level.

[0126] Note that the individual personality embedding of speaker p in the first layer is initialized as follows:

[0127]

[0128] Next, we will continuously update the individual personality representations based on the progress of the conversation. Unlike the group personality update, in this part, we will only utilize individual historical discourse from the dialogue. Since the calculation method is similar to that of group personality extraction, it will not be elaborated further. Specifically:

[0129]

[0130]

[0131]

[0132] Finally, the speaker at layer p is p t Individual personality at any given moment It is the sum of the individual's historical conversation context weighted according to their attention level:

[0133]

[0134] (2) Individual personality participation stage

[0135] Although individual personality is a relatively stable psychological trait, participants may exhibit different characteristics and behaviors as the dialogue situation progresses. For example, people may adjust their personality expression based on changes in the dialogue partner and purpose, leading to changes in the participant's emotional expression tendencies. Therefore, we also need to encourage dynamic changes in individual personality to participate in the expression of emotional emotions in discourse. In this step, we used the same method as in the group personality participation stage, so we will not repeat the explanation. The specific formula is:

[0136]

[0137]

[0138]

[0139]

[0140] Finally obtained This represents the speech of participant p at time t in layer l.

[0141] Meanwhile, we still use the stacked individual personality discourse fusion module to update the dialogue hierarchy. That is, the discourse representation generated at layer l serves as the initial input discourse for layer l+1.

[0142] (iv) Sentiment Prediction Module

[0143] In this step, we borrow the residual connection approach to represent discourse that has involved group personality traits. and the initial global discourse features h i Add them together, while simultaneously expressing the discourse that incorporates individual personality traits. and the initial individual discourse order dependency representation s i Add the two parts of the utterance representation together and concatenate them to generate the final feature for each utterance:

[0144]

[0145] Then e i The input is fed into an MLP with fully connected layers to predict the sentiment label y of the discourse. i :

[0146] l i =ReLU(W l e i +b l (27)

[0147] P i =Softmax(W smax l i +b smax (28)

[0148]

[0149] We use classification cross-entropy and L2 regularization as the loss function during training:

[0150]

[0151] Where N is the number of dialogues, c(i) is the number of utterances in dialogue i, and Pi,j y is the probability distribution of the predicted sentiment label for utterance j in dialogue i. i,j Let λ be the expected class label of utterance j in utterance i, λ be the L2 regularization weights, and θ be the set of all trainable parameters. We train our network using stochastic gradient descent based on the Adam optimizer and optimize the hyperparameters using grid search.

[0152] The scheme described in this embodiment can uncover the speaker's personality traits from limited dialogue information and incorporate them into the emotional expression of the utterances, which is particularly important for accurately identifying the correct emotion of each sentence. In daily conversations, a speaker's emotional expression is closely related to their personal characteristics and personality, and each individual expresses the same emotion in significantly different ways. However, traditional dialogue emotion recognition methods often ignore the differences and dynamic changes in speaker personality, leading to unreliable recognition results. Therefore, to solve this problem, this embodiment provides a dialogue emotion recognition scheme based on dynamic modeling of personality. In the scheme described in this embodiment, a personality capture module is designed to model the personality of characters in historical discourse in order to extract the speaker's personality profile. At the same time, there is a mutual influence between the speaker's personality tendencies and their language behavior, so the speaker's personality profile is continuously updated according to the progress of the discourse. Finally, since the personalities of other speakers in the conversation also affect the way and intensity of the speaker's emotional expression, we also use group personality to construct the conversation emotion. Next, at both the group and individual levels, the speaker's personality is fully integrated into the representation of each sentence through a personality participation stage, and the discourse representation and speaker personality are updated layer by layer in this process, thereby achieving more accurate dialogue emotion recognition. We conducted extensive experiments on four commonly used public benchmark datasets for dialogue emotion recognition to evaluate our proposed model. The results demonstrate the effectiveness of the scheme described in this embodiment and show that it achieves highly competitive results on popular metrics.

[0153] Example 2:

[0154] The purpose of this embodiment is to provide a dialogue emotion recognition system based on dynamic modeling of human personality.

[0155] A dialogue emotion recognition system based on dynamic modeling of personality traits, comprising:

[0156] The discourse feature extraction unit is used to acquire the historical conversations to be sentiment identified and extract the feature vector of each discourse in the historical conversations.

[0157] The global and individual feature extraction units are used to obtain the global context representation and individual context representation corresponding to each utterance in the current historical session based on the feature vector of each utterance extracted in the historical session.

[0158] The personality extraction unit is used to obtain the group personality characteristics of the speakers in the historical conversation based on the global context representation of each utterance and the attention weight of each sentence in the group personality; and to obtain the individual personality characteristics of the speakers in the historical conversation based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality.

[0159] The personality participation unit is used to integrate the acquired group personality characteristics and individual personality characteristics into the discourse representation, thereby obtaining discourse representations with group personality participation and discourse representations with individual personality participation.

[0160] The emotion recognition unit is used to obtain the emotion label of each utterance based on the fusion features of utterance representations with group personality participation, initial global context representation, utterance representations with individual personality participation, and initial individual context representation, using a multilayer perceptron.

[0161] Furthermore, the system described in this embodiment corresponds to the method described in Embodiment 1, and its technical details have been described in detail in Embodiment 1, so they will not be repeated here.

[0162] In further embodiments, the following is also provided:

[0163] An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.

[0164] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0165] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.

[0166] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.

[0167] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory; the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.

[0168] Those skilled in the art will recognize that the units, i.e., algorithm steps, of the various examples described in connection with this embodiment can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this disclosure.

[0169] The above embodiments provide a dialogue emotion recognition method and system based on dynamic modeling of personality, which has broad application prospects.

[0170] The above description is merely a preferred embodiment of this disclosure and is not intended to limit this disclosure. Various modifications and variations can be made to this disclosure by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this disclosure should be included within the scope of protection of this disclosure.

Claims

1. A dialogue emotion recognition method based on dynamic modeling of personality, characterized in that, include: Obtain the historical conversations to be sentiment identified, and extract the feature vector of each utterance in the historical conversations; Based on the feature vector of each utterance extracted from the historical conversation, the global context representation and individual context representation corresponding to each utterance in the current historical conversation are obtained respectively; Based on the global context representation of each utterance and the attention weight of each sentence in the group personality, the group personality characteristics of the speakers in the historical conversation are obtained. Based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality, the individual personality characteristics of the speakers in the historical conversation are obtained; the acquisition of the attention weight of each sentence is specifically as follows: The attention weight of each sentence in the group personality is calculated using the following formula: in, This provides a global context representation for each utterance in the previous layer. and These are training parameters. ,express All words spoken before that moment; For the first In the iterative layer Hidden expressions of group personality at all times Both represent vector dimensions; The attention weight of each sentence in an individual's personality is calculated using the following formula: in, This provides a speaker-level context representation for each utterance in the higher-level context. For the first l In the iterative layer t- The hidden manifestations of an individual's personality at a given moment. and These are training parameters. Both represent vector dimensions. That is to say, in the dialogue Speaker Expressing words; The acquired group personality traits and individual personality traits are integrated into the discourse representation to obtain discourse representations involving group personality traits and discourse representations involving individual personality traits. The process of integrating the obtained group personality traits into discourse representation specifically involves: Calculate the update gate and reset gate based on the current input utterance and group personality characteristics; By resetting the gate, the newly entered utterances are combined with the previous group personality traits; The update gate is used to pass group personality traits to the next layer, while a stacking method is used to update the dialogue hierarchy. The stacking method is as follows: l The discourse generated by the layer represents, as l Initialization input utterance for +1 layer; Alternatively, the integration of the obtained individual personality traits into the discourse representation specifically includes: Calculate update and reset gates based on individual historical discourse and personality traits; By resetting the gate, the newly entered words are combined with the previous individual personality traits; An update gate is used to pass individual personality traits to the next layer, while a stacking method is used to update the dialogue hierarchy. The stacking method is as follows: l The discourse generated by the layer represents, as l Initialization input utterance for +1 layer; Based on the fusion features of discourse representations involving group personality, initial global context representation, discourse representations involving individual personality, and initial individual context representation, a multilayer perceptron is used to obtain the sentiment label of each discourse.

2. The dialogue emotion recognition method based on dynamic character modeling as described in claim 1, characterized in that, The step of obtaining the global context representation and individual context representation corresponding to each utterance in the current historical session is specifically as follows: Two pre-trained bidirectional long short-term memory networks are used to model the global utterance order dependency and the individual utterance order dependency, respectively. Specifically, for modeling the global utterance order dependency, the input of the bidirectional long short-term memory network is the text features of each utterance in the historical session; for modeling the individual utterance order dependency, the input of the bidirectional long short-term memory network is the text features of adjacent utterances from the same speaker.

3. The dialogue emotion recognition method based on dynamic character modeling as described in claim 1, characterized in that, The feature vector of each utterance in the historical conversation is extracted, specifically using the pre-trained language model RoBERTa.

4. The dialogue emotion recognition method based on dynamic character modeling as described in claim 1, characterized in that, The process of integrating the obtained group personality characteristics and individual personality characteristics into the discourse representation to obtain discourse representations involving group personality and individual personality involves the following: integrating group personality characteristics and individual personality characteristics into the discourse representation based on the update and reset mechanism of the gated loop unit.

5. The dialogue emotion recognition method based on dynamic character modeling as described in claim 1, characterized in that, The method, based on the fusion features of discourse representation with group personality participation, initial global context representation, discourse representation with individual personality participation, and initial individual context representation, uses a multilayer perceptron to obtain the sentiment label of each discourse. Specifically, the discourse representation with group personality participation is added to the initial global context representation of the discourse; at the same time, the discourse representation with individual personality participation and the initial individual context representation of the discourse are added together, and the discourse representation obtained by adding the two parts is concatenated to obtain the fusion features of each discourse.

6. A dialogue emotion recognition system based on dynamic modeling of personality, characterized in that, include: The discourse feature extraction unit is used to acquire the historical conversations to be sentiment identified and extract the feature vector of each discourse in the historical conversations. The global and individual feature extraction units are used to obtain the global context representation and individual context representation corresponding to each utterance in the current historical session based on the feature vector of each utterance extracted in the historical session. The personality extraction unit is used to obtain the group personality characteristics of the speakers in the historical conversation based on the global context representation of each utterance and the attention weight of each sentence in the group personality. Based on the individual context representation of each utterance and the attention weight of each sentence in the individual personality, the individual personality characteristics of the speakers in the historical conversation are obtained. The acquisition of attention weight for each sentence is as follows: The attention weight of each sentence in the group personality is calculated using the following formula: in, This provides a global context representation for each utterance in the previous layer. and These are training parameters. ,express All words spoken before that moment; For the first In the iterative layer Hidden expressions of group personality at all times Both represent vector dimensions; The attention weight of each sentence in an individual's personality is calculated using the following formula: in, This provides a speaker-level context representation for each utterance in the higher-level context. For the first l The hidden representation of an individual's personality at time t-1 in the iterative layer. and These are training parameters. Both represent vector dimensions. That is to say, in the dialogue Speaker Expressing words; A personality participation unit is used to integrate the obtained group personality characteristics and individual personality characteristics into the discourse representation, thereby obtaining discourse representations with group personality participation and discourse representations with individual personality participation; the integration of the obtained group personality characteristics into the discourse representation specifically includes: Calculate the update gate and reset gate based on the current input utterance and group personality characteristics; By resetting the gate, the newly entered utterances are combined with the previous group personality traits; The update gate is used to pass group personality traits to the next layer, while a stacking method is used to update the dialogue hierarchy. The stacking method is as follows: l The discourse generated by the layer represents, as l Initialization input utterance for +1 layer; Alternatively, the integration of the obtained individual personality traits into the discourse representation specifically includes: Calculate update and reset gates based on individual historical discourse and personality traits; By resetting the gate, the newly entered words are combined with the previous individual personality traits; An update gate is used to pass individual personality traits to the next layer, while a stacking method is used to update the dialogue hierarchy. The stacking method is as follows: l The discourse generated by the layer represents, as l Initialization input utterance for +1 layer; The emotion recognition unit is used to obtain the emotion label of each utterance based on the fusion features of utterance representations with group personality participation, initial global context representation, utterance representations with individual personality participation, and initial individual context representation, using a multilayer perceptron.

7. An electronic device, comprising a memory, a processor, and a computer program stored in the memory and running thereon, characterized in that, When the processor executes the program, it implements a dialogue emotion recognition method based on dynamic modeling of personality as described in any one of claims 1-5.

8. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by the processor, the program implements a dialogue emotion recognition method based on dynamic modeling of personality as described in any one of claims 1-5.

Citation Information

Patent Citations

  • Artificial intelligence emotion accompanying word memorizing learning system in dialogue chat mode

    CN112632243A

  • Personalized meta-cognition enhanced heuristic emotion support dialogue method

    CN116702794A