Emotion classification method and device based on artificial intelligence, computer device and medium
By optimizing the feature embedding and emotion probability distribution of each sentence in the dialogue text, the problem of semantic information being difficult to represent due to the emotional inertia of the same speaker is solved, thus improving the accuracy of emotion classification.
Patent Information
- Application Number
- CN202211618712.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-12-15
- Publication Date
- 2025-12-30
- Estimated Expiration
- 2042-12-15
AI Technical Summary
Existing emotion classification methods suffer from low accuracy because the semantic information is difficult to accurately and completely represent due to the emotional inertia of the same speaker in dialogue texts.
Each sentence in the dialogue text is input into a trained encoder for feature embedding to obtain a sentence embedding vector. The emotion probability distribution is determined by a trained classifier. The sentence set is divided according to the speaker, the transition probability between sentences is calculated, and an objective function is constructed for optimization to determine the emotion classification result.
By jointly modeling the emotional changes and semantic information of different speakers in the dialogue text, the accuracy of emotion classification is improved.
Smart Images

Figure CN115983283B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to an emotion classification method, apparatus, computer device, and medium based on artificial intelligence. Background Technology
[0002] Currently, with the development of artificial intelligence technology, sentiment classification technology based on dialogue text has been gradually applied to various application scenarios such as product recommendation and public opinion analysis. For example, based on the dialogue text between users of the same product, we can understand the user experience of the product and thus help provide suitable product recommendations for new users.
[0003] Existing emotion classification methods typically involve directly inputting each sentence from a dialogue text into an emotion classification model. The model then uses the contextual semantic information of each sentence to obtain the corresponding emotion classification result. However, because the emotions of the same speaker in a dialogue text can exhibit emotional inertia, the semantic information cannot accurately and completely represent all the effective information in the dialogue text, resulting in low emotion classification accuracy. Therefore, improving the accuracy of emotion classification has become an urgent problem to be solved. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide an artificial intelligence-based emotion classification method, apparatus, computer device, and medium to solve the problem of low accuracy in emotion classification.
[0005] In a first aspect, embodiments of the present invention provide an artificial intelligence-based emotion classification method, the emotion classification method comprising:
[0006] Each sentence in the acquired dialogue text is input into the trained encoder for feature embedding, resulting in a sentence embedding vector for each sentence.
[0007] For any sentence, the sentence embedding vector of the sentence is input into a trained classifier to obtain the probability that the sentence belongs to N preset emotion categories. The N probabilities are then used to determine the emotion probability distribution of the sentence, and the emotion probability distribution of each sentence is obtained, where N is a positive integer.
[0008] All sentences in the dialogue text are divided into at least two sentence sets according to the speaker. For any sentence set, the transition probability between the sentiment probability distributions of any two sequentially adjacent sentences in the sentence set is calculated according to the text order of the dialogue text to obtain the speaker transition probability vector of the sentence set.
[0009] Based on the speaker transition probability vector of all sentence sets and the emotion probability distribution of each sentence in the dialogue text, an objective function is constructed, the objective function is optimized and solved, and the result corresponding to each sentence is determined from the solution results according to the text order as the emotion classification result of the corresponding sentence.
[0010] Secondly, embodiments of the present invention provide an artificial intelligence-based emotion classification device, the emotion classification device comprising:
[0011] The feature embedding module is used to input each sentence in the acquired dialogue text into the trained encoder for feature embedding, and obtain the sentence embedding vector of each sentence.
[0012] The sentence classification module is used to input the sentence embedding vector of any sentence into a trained classifier to obtain the probability that the sentence belongs to N preset emotion categories, determine the emotion probability distribution of the sentence composed of the N probabilities, and obtain the emotion probability distribution of each sentence, where N is an integer greater than zero.
[0013] The probability calculation module is used to divide all sentences in the dialogue text into at least two sentence sets according to the speaker. For any sentence set, according to the text order of the dialogue text, the module calculates the transition probability between the emotion probability distributions of any two sequentially adjacent sentences in the sentence set to obtain the speaker transition probability vector of the sentence set.
[0014] The emotion classification module is used to construct an objective function based on the speaker transition probability vector of all sentence sets and the emotion probability distribution of each sentence in the dialogue text, optimize the objective function, and determine the result corresponding to each sentence from the solution results according to the text order as the emotion classification result of the corresponding sentence.
[0015] Thirdly, embodiments of the present invention provide a computer device, the computer device including a processor, a memory, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the emotion classification method as described in the first aspect.
[0016] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the emotion classification method as described in the first aspect.
[0017] The beneficial effects of the embodiments of the present invention compared with the prior art are as follows:
[0018] Each sentence in the acquired dialogue text is input into a trained encoder for feature embedding, resulting in a sentence embedding vector for each sentence. For any given sentence, the sentence embedding vector is input into a trained classifier to obtain the probability that the sentence belongs to N preset emotion categories. The N probabilities form the emotion probability distribution of the sentence, resulting in the emotion probability distribution for each sentence. All sentences in the dialogue text are divided into at least two sentence sets based on the speaker. For any sentence set, the transition probability between the emotion probability distributions of any two sequentially adjacent sentences in the sentence set is calculated according to the text order of the dialogue text, thus obtaining the speaker transition probability of the sentence set. The transition probability vector is used to construct an objective function based on the speaker transition probability vectors of all sentence sets and the emotion probability distribution of each sentence in the dialogue text. The objective function is optimized and solved. According to the text order, the result corresponding to each sentence is determined as the emotion classification result of the corresponding sentence from the solution results. The dialogue text is divided into multiple sentence sets according to the speaker, and transition probability modeling is performed on each sentence set separately. That is, transition probability modeling is performed between non-continuous texts, which effectively extracts the emotion change information of different speakers in the dialogue text. Thus, the emotion classification result is solved jointly based on the emotion change information and semantic information, thereby improving the accuracy of emotion classification. Attached Figure Description
[0019] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0020] Figure 1 This is a schematic diagram of an application environment for an artificial intelligence-based emotion classification method provided in Embodiment 1 of the present invention;
[0021] Figure 2 This is a flowchart illustrating an artificial intelligence-based emotion classification method provided in Embodiment 1 of the present invention.
[0022] Figure 3 This is a flowchart illustrating an artificial intelligence-based emotion classification method provided in Embodiment 2 of the present invention.
[0023] Figure 4 This is a schematic diagram of the structure of an artificial intelligence-based emotion classification device provided in Embodiment 3 of the present invention;
[0024] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Detailed Implementation
[0025] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0026] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.
[0027] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.
[0028] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," "in response to determination," or "in response to detection." Similarly, the phrase "if determined" or "if [described condition or event] is detected" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once [described condition or event] is detected," or "in response to detection of [described condition or event]."
[0029] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.
[0030] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.
[0031] The embodiments of this invention can acquire and process relevant data based on artificial intelligence technology. Artificial intelligence (AI) refers to the theories, methods, technologies, and application systems that utilize digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to obtain optimal results.
[0032] Foundational technologies for artificial intelligence generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interactive systems, and mechatronics. AI software technologies mainly encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0033] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0034] To illustrate the technical solution of the present invention, specific embodiments are described below.
[0035] The first embodiment of this invention provides an artificial intelligence-based emotion classification method, which can be applied to, for example... Figure 1 In this application environment, the client communicates with the server. Clients include, but are not limited to, handheld computers, desktop computers, laptops, ultra-mobile personal computers (UMPCs), netbooks, cloud terminal devices, and personal digital assistants (PDAs). The server can be a standalone server or a cloud server providing basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0036] See Figure 2 This is a flowchart illustrating an artificial intelligence-based emotion classification method provided in Embodiment 1 of the present invention. The above-described emotion classification method can be applied to... Figure 1The client, a computer device connected to the server, obtains the dialogue text to be classified for emotion. The client's computer device contains a trained encoder and a trained classifier. The encoder extracts sentence embeddings from the dialogue text, and the classifier classifies these embeddings to obtain the emotion probability distribution corresponding to each sentence. Figure 2 As shown, this emotion classification method may include the following steps:
[0037] Step S201: Input each sentence in the acquired dialogue text into the trained encoder for feature embedding to obtain the sentence embedding vector of each sentence.
[0038] The dialogue text can refer to the text obtained by converting the speech of at least two speakers. A complete sentence in the dialogue text is a sentence in the dialogue text, and a sentence in the dialogue text belongs to one of the speakers.
[0039] The trained encoder can be used to embed features into sentences to obtain sentence embedding vectors. Sentence embedding vectors can refer to the vector representation of sentences in the dialogue text, so that the sentence embedding vectors can be calculated by the model later.
[0040] Specifically, suppose the dialogue text contains C speakers, which can be represented as pe0, pe1, ..., pe C-1 Let the speaker be pe c The total number of words spoken in the dialogue text was len. C A dialogue text contains Q sentences, which can be represented as x0, x1, ..., x2. Q-1 Then the speaker pe c The spoken sentence can be represented as Among them, pe c [len c [This can refer to the speaker, pe.] c The first len C One sentence.
[0041] In this embodiment, the trained encoder can be the encoder part of a language model such as a trained BERT model, a trained Word2Vec model, or a trained Transformer model. The resulting sentence embedding vectors can be used to represent the semantic features of sentences in the dialogue text. It should be noted that the sentence embedding vectors have the same size so that a fixed classifier architecture can be used to process multiple sentence embedding vectors in the future.
[0042] Optionally, each sentence in the acquired dialogue text is input into the trained encoder for feature embedding, resulting in a sentence embedding vector for each sentence, including:
[0043] For any sentence, the sentence is segmented into characters to obtain at least one character, and all characters are concatenated into a character vector according to the order of the sentence;
[0044] The character vectors are input into the trained encoder for feature embedding. The feature embedding result is determined as the sentence embedding vector of the sentence. All sentences are traversed to obtain the sentence embedding vector of each sentence.
[0045] Here, a character can refer to the basic building block of a sentence, and character segmentation can be performed using a pre-defined dictionary model. The sentence order can refer to the word order of a sentence, and the order of all characters is determined by the position of each character in the sentence.
[0046] A character vector can be a vector representation of all the characters in a sentence. A character vector can be formed by concatenating all the characters in the order they appear.
[0047] Specifically, the character vector corresponding to each sentence is input into the trained encoder for feature embedding to obtain the sentence embedding vector of the corresponding sentence.
[0048] In this embodiment, the trained encoder uses the encoder part of the trained Transformer model. Each character in the sentence is first vector-encoded and position-encoded to obtain character embedding features and position embedding features. The character embedding features and position embedding features are added together to obtain the character sub-vector corresponding to the character. The character sub-vectors of all characters are concatenated to obtain the character vector corresponding to the sentence. The character vector is then input into the trained encoder to obtain the sentence embedding vector. The size of the sentence embedding vector is the same as the size of the character vector.
[0049] In this embodiment, the sentence is vector-encoded and position-encoded on a character-by-character basis to obtain character sub-vectors. The concatenation result of the character sub-vectors is used as the character vector input to the trained encoder for feature embedding. The character vector can effectively represent the positional information and contextual semantic information of each character, so that the obtained sentence embedding vector can fully extract the semantic features of the sentence, thereby facilitating the subsequent improvement of the accuracy of the emotion classification process based on the sentence embedding vector.
[0050] Optionally, the sentence is segmented into characters to obtain at least one character, and all characters are concatenated into a character vector according to the sentence order, including:
[0051] The sentence is segmented into characters to obtain at least one character. All characters are then concatenated in the order of the sentence to obtain the concatenated result.
[0052] The preset start delimiter is appended to the beginning of the concatenation result, and the preset end delimiter is appended to the end of the concatenation result to obtain the character vector.
[0053] Here, character-based segmentation can refer to segmenting according to characters, and the concatenation result can refer to the concatenation vector obtained by concatenating all the characters in the sentence.
[0054] The start separator can be used to indicate the beginning of the corresponding sentence, and the end separator can be used to indicate the end of the corresponding sentence.
[0055] Specifically, let the sentence be x. q That is, any one of the Q sentences above, sentence x q After segmenting the characters and then concatenating them, the resulting concatenation can be represented as [x q,0 ,x q,1 ,…,x q,L-1 ], where L is sentence x q The number of characters contained.
[0056] In this embodiment, the start delimiter can be represented as [cls], the end delimiter can be represented as [sep], and the character vector can be represented as [[cls], x q,0 ,x q,1 ,…,x q,L-1 ,[sep]。
[0057] In this embodiment, character vectors are constructed by adding preset delimiters, which can provide effective sentence semantic information, avoid incorrect semantic representation caused by sentence segmentation errors, and thus reduce the accuracy of subsequent emotion classification, thereby effectively improving the accuracy of subsequent emotion classification.
[0058] The above steps involve inputting each sentence in the acquired dialogue text into a trained encoder for feature embedding to obtain a sentence embedding vector for each sentence. This process of embedding sentences into features improves the effectiveness of sentence semantic representation, thereby effectively improving the accuracy of subsequent sentence sentiment classification.
[0059] Step S202: For any sentence, input the sentence embedding vector into the trained classifier to obtain the probability that the sentence belongs to N preset emotion categories, determine the emotion probability distribution of the sentence composed of N probabilities, and obtain the emotion probability distribution of each sentence.
[0060] Here, N is a positive integer, the preset emotion category can refer to a preset speaker emotion category, and the trained classifier can be used to classify the input sentence embedding vector. The emotion probability distribution can contain the probability that the sentence belongs to each preset emotion category.
[0061] Specifically, the preset emotion categories can include happy, angry, neutral, sad, excited, and furious categories, etc. A trained classifier can be implemented using trained fully connected layers, trained decision trees, etc.
[0062] In this embodiment, the trained classifier can be a trained fully connected layer. The trained fully connected layer can be used for multi-classification tasks, that is, to determine the probability that the input sentence embedding vector belongs to multiple categories. After the input sentence embedding vector is input, the trained fully connected layer first outputs the prediction value for N preset emotion categories, and obtains N prediction values. The N prediction values are then normalized using a normalized exponential function to obtain N prediction probabilities, which are the probabilities that the sentence corresponding to the sentence embedding vector belongs to N preset emotion categories.
[0063] In one implementation, a trained classifier can be constructed using multiple trained fully connected layers. In this case, each trained fully connected layer performs a binary classification task, classifying each preset emotion category separately. Finally, the probabilities of all trained fully connected layers belonging to the corresponding preset emotion categories are normalized again to obtain the probabilities of sentences belonging to N preset emotion categories. By using multiple fully connected layers to construct the classifier, the mutual influence between the classification tasks of each preset emotion category can be avoided, further improving the accuracy of emotion classification.
[0064] The above steps involve inputting the sentence embedding vector into a trained classifier for any given sentence to obtain the probability that the sentence belongs to N preset emotion categories, determining the emotion probability distribution of the sentence composed of N probabilities, and obtaining the emotion probability distribution of each sentence. Using the emotion probability distribution of the sentence obtained through the trained classifier as a basis, it is convenient to combine emotion inertia information to jointly optimize and solve the emotion classification result, thereby improving the emotion classification accuracy of the sentence.
[0065] Step S203: Divide all sentences in the dialogue text into at least two sentence sets according to the speaker. For any sentence set, calculate the transition probability between the emotion probability distributions of any two sequentially adjacent sentences in the sentence set according to the text order of the dialogue text, and obtain the speaker transition probability vector of the sentence set.
[0066] Here, the sentence set can refer to a set containing all sentences of the same speaker, and sequentially adjacent sentences can refer to two adjacent sentences in the arrangement result after arranging all sentences of the same speaker in the sentence set according to the text order.
[0067] Transition probability can refer to the probability of shifting the emotion from the earlier sentence to the later sentence in sequentially adjacent sentences. A speaker transition probability vector can contain the transition probability between the emotion probability distributions of any two sequentially adjacent sentences.
[0068] Specifically, in this embodiment, the transition probability is calculated for any two sequentially adjacent sentences in any set of sentences in the dialogue text. That is, the transition probability of two sentences from the same speaker is calculated. The speaker transition probability vector can be represented as a two-dimensional matrix of size N*N, where N is a preset emotion category. The rows of the two-dimensional matrix can represent the preset emotion category corresponding to the emotion probability distribution of the preceding sentence, and the columns of the two-dimensional matrix can represent the preset emotion category corresponding to the emotion probability distribution of the following sentence. The matrix element values determined by the rows and columns of the two-dimensional matrix are the transition probabilities between the emotion probability distributions of two sequentially adjacent sentences. It can be known that the sum of the matrix element values in the same column of the two-dimensional matrix is 1.
[0069] In this embodiment, the speaker transition probability vector can be predicted by a trained prediction model. The input of the trained prediction model can be an emotion sequence of the same speaker that conforms to the text order, and the output can be a two-dimensional matrix of size N*N mentioned above.
[0070] For example, let the first row of the two-dimensional matrix be the happiness category, indicating that the preset emotion category corresponding to the emotion probability distribution of the preceding sentence is the happiness category. The first column of the two-dimensional matrix is the happiness category, and the second column is the neutral category. Then, the probability that the speaker's emotion changes from the happiness category to the happiness category can be determined by the element value of the first row and first column of the two-dimensional matrix, that is, the speaker's emotion remains unchanged. The probability that the speaker's emotion changes from the happiness category to the neutral category can be determined by the element value of the first row and second column of the two-dimensional matrix, that is, the speaker's emotion changes from happiness to neutral.
[0071] Optionally, after obtaining the speaker transition probability vector for each sentence set, the following is also included:
[0072] The global transition probability vector is obtained by calculating the transition probability between the sentiment probability distributions of any two sequentially adjacent sentences in the dialogue text, according to the text order of the dialogue text.
[0073] Accordingly, based on the speaker transition probability vectors of all sentence sets and the sentiment probability distribution of each sentence in the dialogue text, the objective function is constructed as follows:
[0074] The objective function is constructed based on the speaker transition probability vector of all sentence sets, the global transition probability vector, and the emotion probability distribution of each sentence in the dialogue text.
[0075] The global transition probability vector can include the transition probability between any two sequentially adjacent sentences in the dialogue text.
[0076] Specifically, in this embodiment, the transition probability is calculated for any two sequentially adjacent sentences in the dialogue text, that is, without distinguishing the speakers, only considering the order of the sentences in the dialogue text.
[0077] In this embodiment, the transition probability between sequentially adjacent sentences in the dialogue text is calculated to obtain a global transition probability vector, which represents the relationship of mutual influence between different speakers' emotions, thereby further improving the accuracy of emotion change information, and thus improving the accuracy of emotion classification based on emotion change information.
[0078] Optionally, based on the speaker transition probability vector of the entire sentence set, the global transition probability vector, and the sentiment probability distribution of each sentence in the dialogue text, the objective function can be constructed as follows:
[0079] For any given sentence, determine the preset emotion category corresponding to the highest probability in the emotion probability distribution of the sentence as the reference category, and iterate through all sentences to obtain the reference category corresponding to each sentence;
[0080] For any two consecutive sentences in a sentence set, the first transition probability distribution between the reference category of the sentence that comes first and the reference category of the sentence that comes second is determined based on the speaker transition probability vector of the sentence set.
[0081] Iterate through all consecutive pairs of sentences in the entire sentence set to obtain M first transition probability distributions, where M is a positive integer;
[0082] Based on the global transition probability vector, determine K second transition probability distributions between the reference category of the preceding sentence and the reference category of the following sentence in K groups of sequentially adjacent sentences in the dialogue text.
[0083] Add the M first transition probability distributions, K second transition probability distributions, and the sentiment probability distributions of all sentences in the dialogue text, and determine the sum as the objective function.
[0084] The reference category can refer to the preset emotion category that the input sentence is most likely to belong to, as determined by the trained classifier. The first transition probability distribution can refer to the transition probability distribution between the emotions of two adjacent sentences from the same speaker. The second probability distribution can refer to the transition probability distribution between the emotions of two adjacent sentences in the dialogue text, where the adjacent sentences are usually from different speakers.
[0085] This embodiment models the emotional inertia of the same speaker and the emotional influence relationship between different speakers by using speaker transition probability vectors and global transition probability vectors, thereby improving the ability to represent emotional change information and thus improving the accuracy of emotion classification.
[0086] The above steps divide all sentences in the dialogue text into at least two sentence sets according to the speaker. For any sentence set, the transition probability between the emotional probability distributions of any two sequentially adjacent sentences in the sentence set is calculated according to the text order of the dialogue text to obtain the speaker transition probability vector of the sentence set. The speaker transition probability vector can effectively represent the emotional inertia of the same speaker, improve the ability to represent emotional change information, and thus improve the accuracy of emotion classification.
[0087] Step S204: Based on the speaker transition probability vector of all sentence sets and the emotion probability distribution of each sentence in the dialogue text, construct an objective function, optimize and solve the objective function, and determine the result corresponding to each sentence from the solution results according to the text order as the emotion classification result of the corresponding sentence.
[0088] Here, the objective function can be a function containing parameters to be optimized, the solution result can be the optimized parameters, the parameters can be the emotion category corresponding to each sentence, and the emotion classification result can be the emotion category corresponding to the corresponding sentence in the solution result.
[0089] Optionally, optimizing the objective function includes:
[0090] The maximum likelihood optimization method is used to construct the likelihood function of the objective function, and the negative number of the likelihood function is used as the optimization loss for the optimization solution.
[0091] The objective function is optimized using a dynamic programming algorithm until the optimization loss converges, and the solution is obtained.
[0092] Here, the likelihood function can refer to the function to be solved, the optimization loss can be used to provide optimization direction for the likelihood function solution process, and the dynamic programming algorithm can adopt Viterbi algorithm, etc.
[0093] This embodiment uses a dynamic programming algorithm to solve the objective function, which has fewer parameters, higher computational efficiency, and improves the efficiency of emotion classification.
[0094] The above steps involve constructing an objective function based on the speaker transition probability vector of all sentences and the emotion probability distribution of each sentence in the dialogue text, optimizing and solving the objective function, and determining the result corresponding to each sentence as the emotion classification result of the corresponding sentence from the solution results according to the text order. Compared with conventional dialogue text modeling models, this method has fewer parameters and higher computational efficiency by constructing and optimizing the objective function.
[0095] In this embodiment, the dialogue text is divided into multiple sentence sets according to the speaker, and transition probability modeling is performed on each sentence set. That is, transition probability modeling is performed between non-continuous texts, which effectively extracts the emotional change information of different speakers in the dialogue text. Thus, the emotion classification result is solved by jointly solving the emotion change information and semantic information, thereby improving the accuracy of emotion classification.
[0096] See Figure 3 This is a flowchart illustrating an artificial intelligence-based emotion classification method provided in Embodiment 2 of the present invention. The construction of the objective function in this emotion classification method includes the following steps:
[0097] Step S301: For any sentence, determine the preset emotion category corresponding to the highest probability in the emotion probability distribution of the sentence as the reference category, and iterate through all sentences to obtain the reference category corresponding to each sentence.
[0098] Step S302: For any two sentences that are sequentially adjacent in any sentence set, determine the transition probability distribution between the reference category of the sentence that comes first and the reference category of the sentence that comes second, based on the speaker transition probability vector of the sentence set.
[0099] Step S303: Traverse all consecutive pairs of sentences in the entire sentence set to obtain M transition probability distributions;
[0100] Step S304: Add the M transition probability distributions and the sentiment probability distributions of all sentences in the dialogue text, and determine the sum as the objective function.
[0101] Where M is a positive integer, and the reference category can refer to the preset sentiment category to which the sentence output by the classifier is most likely to belong.
[0102] Specifically, the objective function is obtained by summing the emotion probability distribution of each sentence in the dialogue text with the speaker transition probability vector of the entire sentence set. The objective function can be expressed as max(P(y|X)), where X can refer to the dialogue text, represented as [p0, p1, ..., p...]. Q-1 ], p q This can represent the probability distribution of emotions corresponding to the q-th sentence, where y can refer to a preset sequence of emotion categories, represented as [y0, y1, ..., y]. Q-1 ], where y q It can represent the true emotional category of the q-th sentence.
[0103] The objective function s can then be further expressed as:
[0104]
[0105] Where W can be the number of sentences in the sentence set, Q can be the number of sentences in the dialogue text, and len i Let MP be the number of sentences contained in the i-th sentence set. pei[j-1],i[] It can refer to the transition probability from the emotion category of the (j-1)th sentence in the i-th sentence set to the emotion category of the j-th sentence. The total number of transition probabilities can be represented as M.
[0106] In this embodiment, the objective function is constructed and optimized using a conditional random field model. Compared with conventional dialogue text modeling models, it has fewer parameters and higher efficiency in model training and inference, thereby improving the accuracy of emotion classification.
[0107] Corresponding to the AI-based emotion classification method in the above embodiments, Figure 4 A structural block diagram of an AI-based emotion classification device according to Embodiment 3 of the present invention is shown. This emotion classification device is applied to a client-side computer device connected to a server to obtain the dialogue text to be classified. The client-side computer device is equipped with a trained encoder and a trained classifier. The trained encoder can be used to extract sentence features from the dialogue text to obtain sentence embedding features, and the trained classifier can be used to classify the sentence embedding features to obtain the emotion probability distribution corresponding to the sentence. For ease of explanation, only the parts relevant to the embodiments of the present invention are shown.
[0108] See Figure 4 The emotion classification device includes:
[0109] The feature embedding module 41 is used to input each sentence in the acquired dialogue text into the trained encoder for feature embedding, and obtain the sentence embedding vector of each sentence.
[0110] The sentence classification module 42 is used to input the sentence embedding vector into the trained classifier for any sentence, obtain the probability that the sentence belongs to N preset emotion categories, determine the emotion probability distribution of the sentence composed of N probabilities, and obtain the emotion probability distribution of each sentence, where N is an integer greater than zero.
[0111] The probability calculation module 43 is used to divide all sentences in the dialogue text into at least two sentence sets according to the speaker. For any sentence set, according to the text order of the dialogue text, the transition probability between the emotional probability distributions of any two sequentially adjacent sentences in the sentence set is calculated to obtain the speaker transition probability vector of the sentence set.
[0112] The emotion classification module 44 is used to construct an objective function based on the speaker transition probability vector of all sentence sets and the emotion probability distribution of each sentence in the dialogue text, optimize the objective function, and determine the result corresponding to each sentence from the solution results in the order of the text as the emotion classification result of the corresponding sentence.
[0113] Optionally, the feature embedding module 41 mentioned above includes:
[0114] The character segmentation unit is used to segment any sentence into at least one character by word, and then concatenate all the characters into a character vector according to the order of the sentence.
[0115] The character encoding unit is used to input character vectors into the trained encoder for feature embedding, determine the feature embedding result as the sentence embedding vector, traverse all sentences, and obtain the sentence embedding vector of each sentence.
[0116] Optionally, the above character segmentation unit includes:
[0117] The character concatenation subunit is used to segment a sentence into at least one character by word, and then concatenate all the characters in the order of the sentence to obtain the concatenation result.
[0118] The delimiter concatenation subunit is used to concatenate a preset start delimiter before the concatenation result and a preset end delimiter after the concatenation result to obtain a character vector.
[0119] Optionally, the aforementioned emotion classification module 44 includes:
[0120] The category determination unit is used to determine the preset emotion category corresponding to the highest probability in the emotion probability distribution of any sentence as the reference category, and to traverse all sentences to obtain the reference category corresponding to each sentence.
[0121] The distribution determination unit is used to determine the transition probability distribution between the reference category of the sentence that comes first and the reference category of the sentence that comes second in any set of sentences, based on the speaker transition probability vector of the sentence set.
[0122] The sentence traversal unit is used to traverse all consecutive pairs of sentences in the entire sentence set, and obtain M transition probability distributions, where M is an integer greater than zero;
[0123] The first function construction unit is used to add the M transition probability distributions and the sentiment probability distributions of all sentences in the dialogue text, and determine the sum as the objective function.
[0124] Optionally, the aforementioned emotion classification module 44 includes:
[0125] The loss building unit is used to construct the likelihood function of the objective function using the maximum likelihood optimization method, and the negative number of the likelihood function is used as the optimization loss for the optimization solution.
[0126] The optimization unit is used to optimize the objective function using a dynamic programming algorithm until the optimization loss converges and the solution is obtained.
[0127] Optionally, the aforementioned emotion classification device also includes:
[0128] The global probability calculation module is used to calculate the transition probability between the sentiment probability distributions of any two sequentially adjacent sentences in the dialogue text, according to the text order of the dialogue text, and obtain the global transition probability vector.
[0129] Accordingly, the aforementioned emotion classification module 44 includes:
[0130] The second function construction unit is used to construct the objective function based on the speaker transition probability vector, the global transition probability vector, and the emotion probability distribution of each sentence in the dialogue text for all sentence sets.
[0131] Optionally, the second function building unit mentioned above includes:
[0132] The reference category subunit is used to determine the preset emotion category corresponding to the highest probability in the emotion probability distribution of any sentence as the reference category. It iterates through all sentences to obtain the reference category corresponding to each sentence.
[0133] The first distribution subunit is used to determine the first transition probability distribution between the reference category of the sentence that comes first and the reference category of the sentence that comes second in any two sentences that are sequentially adjacent in any sentence set, based on the speaker transition probability vector of the sentence set.
[0134] The distribution traversal subunit is used to traverse all consecutive pairs of sentences in the entire sentence set, and obtain M first transition probability distributions, where M is an integer greater than zero;
[0135] The second distribution subunit determines, based on the global transition probability vector, K second transition probability distributions between the reference categories of the preceding and following sentences in the K groups of sequentially adjacent sentences in the dialogue text.
[0136] The distribution superposition subunit is used to add M first transition probability distributions, K second transition probability distributions, and the sentiment probability distributions of all sentences in the dialogue text, and determine the sum as the objective function.
[0137] It should be noted that the information interaction and execution process between the above-mentioned module units and sub-units are based on the same concept as the method embodiments of the present invention. For details on their specific functions and technical effects, please refer to the method embodiments section, which will not be repeated here.
[0138] Figure 5 This is a schematic diagram of the structure of a computer device provided in Embodiment 4 of the present invention. Figure 5 As shown, the computer device of this embodiment includes: at least one processor ( Figure 5 Only one is shown in the diagram), a memory, and a computer program stored in the memory and capable of running on at least one processor, which, when executing the computer program, implements the steps in any of the above-described emotion classification method embodiments.
[0139] This computer device may include, but is not limited to, a processor and memory. Those skilled in the art will understand that... Figure 5 The examples of computer devices are merely examples and do not constitute a limitation on computer devices. Computer devices may include more or fewer components than shown in the illustration, or combinations of certain components, or different components, such as network interfaces, displays, and input devices.
[0140] The processor referred to can be a CPU, but it can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.
[0141] Memory includes readable storage media, internal memory, etc., wherein internal memory can be the RAM of a computer device, providing an environment for the operation of the operating system and computer-readable instructions stored in the readable storage media. The readable storage media can be the hard drive of a computer device, or in other embodiments, it can be an external storage device of the computer device, such as a plug-in hard drive, Smart Media Card (SMC), Secure Digital (SD) card, or Flash Card. Furthermore, memory can include both internal storage units and external storage devices of a computer device. Memory is used to store the operating system, applications, bootloader, data, and other programs, such as program code for computer programs. Memory can also be used to temporarily store data that has been output or will be output.
[0142] Those skilled in the art will understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the functions described above can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above device can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here. If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the present invention can implement all or part of the processes in the methods of the above embodiments by instructing related hardware through a computer program. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the above method embodiments. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. A computer-readable medium can include at least: any entity or device capable of carrying computer program code, a recording medium, a computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media. Examples include USB flash drives, portable hard drives, magnetic disks, or optical disks. In some jurisdictions, according to legislation and patent practice, computer-readable media cannot be electrical carrier signals or telecommunication signals.
[0143] The present invention can implement all or part of the processes in the methods of the above embodiments, or it can be accomplished by a computer program product. When the computer program product is run on a computer device, the computer device executes the steps in the above method embodiments.
[0144] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0145] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0146] In the embodiments provided by this invention, it should be understood that the disclosed apparatus / computer devices and methods can be implemented in other ways. For example, the apparatus / computer device embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0147] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0148] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. An artificial intelligence-based emotion classification method, characterized by, The emotion classification method comprises: inputting each sentence in the obtained dialogue text into a trained encoder for feature embedding to obtain a sentence embedding vector of each sentence; for any sentence, inputting the sentence embedding vector of the sentence into a trained classifier to obtain probabilities that the sentence respectively belongs to N preset emotion categories, determining that the N probabilities form an emotion probability distribution of the sentence, obtaining the emotion probability distribution of each sentence, and N is an integer greater than zero; dividing all sentences in the dialogue text into at least two sentence sets according to speakers, for any sentence set, calculating a transition probability between emotion probability distributions of any two sequentially adjacent sentences in the sentence set according to a text order of the dialogue text to obtain a speaker transition probability vector of the sentence set, wherein the sentence set refers to a set containing all sentences of the same speaker, and any two sequentially adjacent sentences in the sentence set refer to two adjacent sentences in the arrangement result after arranging all sentences of the same speaker in the sentence set according to the text order; constructing an objective function according to the speaker transition probability vectors of all sentence sets and the emotion probability distribution of each sentence in the dialogue text, and optimizing and solving the objective function to determine, according to the text order, that a result corresponding to each sentence is an emotion classification result of the corresponding sentence; after obtaining the speaker transition probability vector of each sentence set, the method further comprises: calculating a transition probability between emotion probability distributions of any two sequentially adjacent sentences in the dialogue text according to the text order of the dialogue text to obtain a global transition probability vector; correspondingly, the constructing an objective function according to the speaker transition probability vectors of all sentence sets and the emotion probability distribution of each sentence in the dialogue text comprises: constructing an objective function according to the speaker transition probability vectors of all sentence sets, the global transition probability vector and the emotion probability distribution of each sentence in the dialogue text.
2. The emotion classification method of claim 1, wherein, the inputting each sentence in the obtained dialogue text into a trained encoder for feature embedding to obtain a sentence embedding vector of each sentence comprises: for any sentence, cutting the sentence into at least one character according to a word, and concatenating all characters into a character vector according to the order of the sentence; inputting the character vector into a trained encoder for feature embedding to determine that a feature embedding result is a sentence embedding vector of the sentence, and traversing all sentences to obtain a sentence embedding vector of each sentence.
3. The emotion classification method of claim 2, wherein, the cutting the sentence into at least one character according to a word, and concatenating all characters into a character vector according to the order of the sentence comprises: cutting the sentence into at least one character according to a word, and concatenating all characters according to the order of the sentence to obtain a concatenation result; concatenating a preset start delimiter to the front of the concatenation result and concatenating a preset end delimiter to the back of the concatenation result to obtain the character vector.
4. The emotion classification method of claim 1, wherein, The constructing a target function according to the speaker transition probability vector of all sentence sets and the emotion probability distribution of each sentence in the dialogue text comprises: For any sentence, determining a preset emotion category corresponding to a maximum probability in the emotion probability distribution of the sentence as a reference category, and traversing all sentences to obtain a reference category corresponding to each sentence; For any two sequentially adjacent sentences in a sentence set, determining a transition probability distribution between the reference category of the sentence in front and the reference category of the sentence behind according to the speaker transition probability vector of the sentence set; Traversing all two sequentially adjacent sentences in all sentence sets to obtain M transition probability distributions, M being an integer greater than zero; Adding the M transition probability distributions and the emotion probability distribution of all sentences in the dialogue text to determine a sum result as the target function.
5. The emotion classification method of claim 1, wherein, The optimizing and solving the target function comprises: adopting a maximum likelihood optimization method to construct a likelihood function of the target function, and taking a negative number of the likelihood function as an optimization loss of the optimizing and solving; adopting a dynamic programming algorithm to optimize and solve the target function until the optimization loss converges, to obtain the solving result.
6. The emotion classification method of claim 1, wherein, The constructing a target function according to the speaker transition probability vector of all sentence sets, the global transition probability vector and the emotion probability distribution of each sentence in the dialogue text comprises: For any sentence, determining a preset emotion category corresponding to a maximum probability in the emotion probability distribution of the sentence as a reference category, and traversing all sentences to obtain a reference category corresponding to each sentence; For any two sequentially adjacent sentences in a sentence set, determining a first transition probability distribution between the reference category of the sentence in front and the reference category of the sentence behind according to the speaker transition probability vector of the sentence set; Traversing all two sequentially adjacent sentences in all sentence sets to obtain M first transition probability distributions, M being an integer greater than zero; Determining K second transition probability distributions between the reference category of the sentence in front and the reference category of the sentence behind among K groups of two sequentially adjacent sentences in the dialogue text according to the global transition probability vector; Adding the M first transition probability distributions, the K second transition probability distributions and the emotion probability distribution of all sentences in the dialogue text to determine a sum result as the target function.
7. An emotion classification apparatus based on artificial intelligence, characterized by, The emotion classification device comprises: a feature embedding module configured to input each sentence in the acquired dialogue text into a trained encoder for feature embedding, to obtain a sentence embedding vector of each sentence; a sentence classification module configured to input the sentence embedding vector of any sentence into a trained classifier, to obtain probabilities of the sentence belonging to N preset emotion categories, determine N probabilities as an emotion probability distribution of the sentence, and obtain an emotion probability distribution of each sentence, N being an integer greater than zero. The probability calculation module is configured to divide all sentences in the dialogue text into at least two sentence sets according to speakers, calculate a transition probability between emotion probability distributions of any two sequentially adjacent sentences in any sentence set according to a text order of the dialogue text, and obtain a speaker transition probability vector of the sentence set, wherein the sentence set refers to a set of all sentences of a same speaker, and the any two sequentially adjacent sentences in the sentence set refer to two adjacent sentences in a result of arranging all sentences of the same speaker in the sentence set according to the text order. The emotion classification module is configured to construct a target function according to the speaker transition probability vector of all sentence sets and the emotion probability distribution of each sentence in the dialogue text, perform optimization solving on the target function, and determine, according to the text order, a result corresponding to each sentence from a result of the solving as an emotion classification result of the corresponding sentence. Further comprising: The global probability calculation module is configured to, after obtaining the speaker transition probability vector of each sentence set, calculate a transition probability between emotion probability distributions of any two sequentially adjacent sentences in the dialogue text according to the text order of the dialogue text, and obtain a global transition probability vector. Correspondingly, the emotion classification module comprises: The target function is constructed according to the speaker transition probability vector of all sentence sets, the global transition probability vector, and the emotion probability distribution of each sentence in the dialogue text.
8. A computer device, comprising: The computer device comprises a processor, a memory, and a computer program stored in the memory and executable on the processor, and the processor implements the emotion classification method of any one of claims 1 to 6 when executing the computer program.
9. A computer-readable storage medium storing a computer program, the computer program comprising instructions that, when executed by a computer, cause the computer to perform the method of any one of claims 1 to 8. The computer program is executed by the processor to implement the emotion classification method of any one of claims 1 to 6.
Citation Information
Patent Citations
Emotion analysis method and device, electronic equipment and storage medium
CN115292495A