CASA: Method, apparatus, and program for sentiment analysis of conversation modes for dialogue understanding
By adapting aspect-based sentiment analysis to extract internal knowledge from conversations, the method addresses the challenges of passive behavior and inconsistency in multi-turn scenarios, enhancing the interpretability and accuracy of sentiment analysis in conversation contexts.
Patent Information
- Application Number
- JP2023547681
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Priority Date
- 2021-10-18
- Filing Date
- 2022-08-25
- Publication Date
- 2025-06-18
- Estimated Expiration
- 2042-08-25
AI Technical Summary
Existing sentiment analysis models struggle in multi-turn conversation scenarios, often exhibiting passive behavior and inconsistent responses due to the lack of explicit representation of knowledge graphs and common sense knowledge.
The proposed method adapts aspect-based sentiment analysis to conversation scenarios, enabling the extraction of internal knowledge from conversations to understand fine-grained sentiment information. This involves extracting sentiment expressions, polarity values, and corresponding mentions from conversations, which helps chatbots plan subsequent topics and improve proactive engagement.
The solution enhances the interpretability of sentiment analysis models by alleviating data sparsity and allowing for the combination of extracted knowledge with other knowledge sources, leading to more accurate and consistent responses in multi-turn conversations.
Smart Images

Figure 0007694967000028 
Figure 0007694967000029 
Figure 0007694967000030
Abstract
Description
Technical Field
[0001] Cross - reference to Related Applications This application claims the benefit of priority to U.S. Application No. 17 / 503,584, filed Oct. 18, 2021, the entire disclosure of which is incorporated herein by reference.
[0002] Embodiments of the present disclosure relate to the field of sentiment analysis. More specifically, the present disclosure relates to dialogue understanding such as dialogue response generation and conversational question - answering.
Background Art
[0003] Modeling chat conversations is an important area due to its potential to facilitate human - computer communication. Most previous research has focused on the design of end - to - end neural networks that consume only surface features. However, these models are not satisfactory in multi - turn conversation scenarios. Specifically, these models have problems such as passive behavior during conversations and often inconsistent multi - turn responses.
[0004] To generate meaningful responses, the influence of knowledge graphs (KGs), common sense knowledge, personality, and sentiment has been investigated. However, such knowledge, e.g., related KGs, is not usually explicitly represented in conversations and thus, for it to be meaningful, human annotation is required along with benchmark datasets. Furthermore, KGs are difficult to obtain in real - world scenarios and often require entity linking as a necessary step, so using related KGs may introduce additional errors.
Summary of the Invention
Means for Solving the Problems
[0005] The present disclosure addresses one or more technical problems. The present disclosure proposes a method and / or apparatus for extracting internal knowledge from conversations that can be used to understand fine-grained sentiment information and assist in understanding conversations. The present disclosure adapts aspect-based sentiment analysis to sentiment analysis in conversation scenarios. As an example, according to embodiments of the present disclosure, conversation-mode sentiment analysis can extract a user's opinions, polarities, and corresponding mentions from a conversation. Based on the understanding that humans often express their emotions in relation to the entities they are talking about, extracting emotions, polarities, and mentions can lead to useful features and general domain understanding. More specifically, accurately extracting people's emotions and corresponding entities from conversations can help chatbots plan subsequent topics and be more proactive in multi-turn conversations. Another advantage of explicitly extracting emotions and mentions is that the same pair of emotions and mentions may appear in various texts, which includes alleviating the sparsity of data in order to enhance the interpretability of the model and make it easier to combine this knowledge with other knowledge (e.g., KG).
[0006] The present disclosure includes a method and an apparatus for sentiment analysis for multi-turn conversations, comprising a memory configured to store computer program code, and one or more processors configured to access the computer program code and operate when instructed by the computer program code. The computer program code includes a first acquisition code configured to cause at least one processor to acquire an input dialogue, a first extraction code configured to cause at least one processor to extract a sentiment expression based on an embedding of a sentence corresponding to the input dialogue, a first generation code configured to cause at least one processor to generate a polarity value based on an embedding of a sentence corresponding to the input dialogue, and a first determination code configured to cause at least one processor to determine a target reference associated with at least one of the sentiment expressions based on the sentiment expression and the sentence embedding. The first determination code includes a second generation code configured to cause at least one processor to generate a rich context expression based on the sentence embedding and the sentiment expression, and a second determination code configured to cause at least one processor to determine a target reference based on a calculated boundary, wherein the calculated boundary is generated using the rich context expression.
[0007] According to an embodiment, the second generation code includes a third generation code configured to cause at least one processor to generate a turn-wise distance based on the sentence embedding, a fourth generation code configured to cause at least one processor to generate speaker information based on the sentence embedding, wherein the speaker information indicates whether the input dialogue is from the same speaker, and a first concatenation code configured to cause at least one processor to concatenate the turn-wise distance, the speaker information, and the sentiment expression.
[0008] According to an embodiment, the second determination code includes a fifth generation code configured to cause at least one processor to generate a distribution based on rich context expressions and sentiment expressions using one or more attention layers, and a third determination code configured to cause at least one processor to determine a target reference based on the boundary of the distribution.
[0009] According to an embodiment, the step of generating a distribution includes the step of determining the product of the distributions of each of the one or more attention layers.
[0010] According to an embodiment, the step of determining a target reference based on the boundary of the distribution includes the step of selecting the boundary of the distribution based on the highest score from a plurality of scores, and the plurality of scores are generated by determining the product of the distributions of each of the one or more attention layers.
[0011] According to an embodiment, the first extraction code includes a sixth generation code configured to cause at least one processor to generate a plurality of tags using a pre-trained machine learning model, and a first inference code configured to cause at least one processor to infer a sentiment expression based on the plurality of tags.
[0012] According to an embodiment, the first generation code includes a sixth generation code configured to cause at least one processor to generate a plurality of tags using a pre-trained machine learning model, and a first inference code configured to cause at least one processor to infer a sentiment expression based on the plurality of tags.
[0013] According to an embodiment, the polarity value is one of positive, negative, or neutral.
[0014] According to an embodiment, the embedding of the sentence is generated based on the input dialogue.
[0015] [1] Further features, properties, and various advantages of the subject matter of the present disclosure will become more apparent from the following detailed description and the accompanying drawings.
Brief Description of the Drawings
[0016]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Best Mode for Carrying Out the Invention
[0017] The present disclosure relates to the extraction of internal knowledge from conversations that can be used to understand fine-grained sentiment information and assist in understanding conversations. The present disclosure adapts aspect-based sentiment analysis to sentiment analysis of conversation scenarios. As an example, according to embodiments of the present disclosure, conversation-mode sentiment analysis can extract a user's opinions, polarities, and corresponding mentions from a conversation. Based on the understanding that humans often express their emotions in relation to the entities they are talking about, extracting emotions, polarities, and mentions can provide useful features and general domain understanding. More specifically, accurately extracting people's emotions and corresponding entities from a conversation can help a chatbot plan subsequent topics and be more proactive in multi-turn conversations. Another advantage of explicitly extracting emotions and mentions is that the same pair of emotions and mentions may appear in various texts, which includes alleviating the sparsity of data to enhance the interpretability of the model and make it easier to combine this knowledge with other knowledge (e.g., KG).
[0018] Consider the example of the multi-turn conversation in Table 1.
[0019]
Table 1
[0020] Accurately extracting people's emotions and corresponding entities from conversations can help chatbots plan subsequent topics and be more proactive in multi-turn conversations. As an example, if a user mentions that they are a big fan of the soccer player "Lionel Messi", the chatbot can mention the latest news about Messi. Additionally, explicit emotion, polarity, and / or mention extraction can include understanding the entire conversation history, making it easier to combine the extraction with other knowledge (e.g., external KG) and making the model more interpretable. Continuing with the "Lionel Messi" example, by combining the analysis results of emotion and model extraction with an external KG, the chatbot can even recommend the latest matches of Messi's soccer club "Football Club Barcelona".
[0021] In available datasets, sentiment analysis includes a very limited number of instances, which cover only a few domains (such as hotel and restaurant reviews), while daily conversations are open-domain. Additionally, in these datasets, sentiment expressions are usually close to their corresponding aspects or mentioned in short sentences. However, in reality, sentiment expressions and their mentions or aspects are some descriptions that are disjoint, and ellipses and references may introduce more complex inferences. As an example, consider the sentence from Table 1: The mention of "Messi" appears in the 3rd utterance, while the corresponding sentiment word "amazing" is in the 5th utterance. Additionally, "Neymar" introduces further challenges as a very confusing candidate mention. This is, of course, just a 3-turn example with the complexity of internal folding at more turns.
[0022] According to an embodiment, emotion extraction can find all emotional expressions from the last user utterance and determine the polarity of each extracted emotional expression. According to an embodiment, reference extraction can extract corresponding references from the dialogue history for each emotional expression. Reference extraction can include understanding the entire dialogue history using rich features such as information about the speaker and speaker ID for each sentence to assist in long-distance dependency modeling.
[0023] In some embodiments, the exemplary or training dataset can be manually annotated. As an example, the dataset can include many dialogues from multiple datasets, and each dialogue can include multiple sentences. As a first pass, human and / or expert annotators may be asked to annotate and / or label each dialogue. In some embodiments, they may be asked to annotate based on state-of-the-art guidelines. The annotation can include not only the emotional expressions in the sentence but also the polarity values of each reference. The annotation can follow other guidelines. As an example, the annotated reference must be specific. For multiple references corresponding to the same entity, only the most specific one must be annotated; in order to train the model against explicit user opinions, only the references related to the corresponding emotional expression can be annotated.
[0024] The proposed functions described below can be used separately or combined in any order. Further, embodiments may be implemented by a processing circuit (e.g., one or more processors or one or more integrated circuits). In one example, one or more processors execute a program stored in a non-transitory computer-readable medium.
[0025] FIG. 1 is a diagram of an environment 100 in which the methods, apparatuses, and systems described herein according to an embodiment can be implemented.
[0026] As shown in FIG. 1, the environment 100 may include a user device 110, a platform 120, and a network 130. The devices in the environment 100 can be interconnected by a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0027] The user device 110 includes one or more devices that can receive, generate, store, process, and / or provide information related to the platform 120. For example, the user device 110 may include a computing device (e.g., a desktop computer, a laptop computer, a tablet computer, a handheld computer, a smart speaker, a server, etc.), a mobile phone (e.g., a smartphone, a wireless phone, etc.), a wearable device (e.g., smart glasses or a smartwatch), or a similar device. In some implementations, the user device 110 can receive information from the platform 120 and / or send information to the platform.
[0028] The platform 120 includes one or more devices described elsewhere in this document. In some implementations, the platform 120 may include a cloud server or a group of cloud servers. In some implementations, the platform 120 may be modularly designed such that software components can be swapped in or out. Therefore, the platform 120 may be easily and / or quickly restored for different applications.
[0029] In some implementations, as shown, the platform 120 may be hosted in a cloud computing environment 122. In particular, the implementations described in this document are described as having the platform 120 hosted in a cloud computing environment 122, but in some implementations, the platform 120 may not be cloud-based (i.e., implemented outside of a cloud computing environment) or may be partially cloud-based.
[0030] The cloud computing environment 122 includes an environment that hosts the platform 120. The cloud computing environment 122 may provide computing services, software services, data access services, storage services, etc. that do not require an end user (e.g., the user device 110) to recognize the physical location and configuration of one or more systems and / or one or more devices that provide the platform 120 by hosting. As shown, the cloud computing environment 122 may include a group of computing resources 124 (collectively referred to as "computing resources 124" and individually referred to as "computing resource 124").
[0031] The computing resources 124 include one or more personal computers, workstation computers, server devices, or other types of computing and / or communication devices. In some implementations, the computing resources 124 may host the platform 120. The cloud resources may include computing instances running on the computing resources 124, storage devices provided within the computing resources 124, data transfer devices provided by the computing resources 124, etc. In some implementations, the computing resources 124 can communicate with other computing resources 124 through a wired connection, a wireless connection, or a combination of a wired connection and a wireless connection.
[0032] As further shown in FIG. 1, the computing resources 124 include a group of cloud resources such as one or more applications ("APP") 124-1, one or more virtual machines ("VM") 124-2, virtualized storage ("VS") 124-3, one or more hypervisors ("HYP") 124-4, etc.
[0033] Application 124-1 includes one or more software applications that can be provided to, or accessed by, user device 110 and / or platform 120. Application 124-1 may eliminate the need to install and execute software applications on user device 110. For example, Application 124-1 may include software associated with platform 120 and / or any other software that can be provided via cloud computing environment 122. In some implementations, one instance of Application 124-1 can send and receive information with one or more other instances of Application 124-1 through virtual machine 124-2.
[0034] Virtual machine 124-2 includes a software implementation of a machine (e.g., a computer) that executes programs like a physical machine. Virtual machine 124-2 can be either a system virtual machine or a process virtual machine, depending on the use by virtual machine 124-2 and the degree of correspondence with any physical machines. A system virtual machine can provide a complete system platform that supports the execution of a complete operating system ("OS"). A process virtual machine can execute a single program and support a single process. In some implementations, virtual machine 124-2 can operate on behalf of a user (e.g., user device 110) and manage the infrastructure of cloud computing environment 122, such as data management, synchronization, or long-term data transfer.
[0035] The virtualized storage 124-3 includes one or more storage systems and / or one or more devices that use virtualization techniques within the storage system or device of the computing resource 124. In some implementations, within the context of the storage system, the types of virtualization may include block virtualization and file virtualization. Block virtualization can refer to the abstraction (or separation) of logical storage from physical storage so that the storage system can be accessed regardless of the physical storage or heterogeneous structure. The separation can enable flexibility in the way the storage system administrator manages storage for end users. File virtualization can eliminate the dependency between the data accessed at the file level and the location where the files are physically stored. This can enable optimization of storage usage, server consolidation, and / or the performance of non-disruptive file migration.
[0036] The hypervisor 124-4 can provide a hardware virtualization technique that enables multiple operating systems (e.g., "guest operating systems") to be run simultaneously on a host computer such as the computing resource 124. The hypervisor 124-4 can present a virtual operating platform to the guest operating system and manage the execution of the guest operating system. Multiple instances of various operating systems can share the virtualized hardware resources.
[0037] Network 130 includes one or more wired and / or wireless networks. For example, Network 130 may include a cellular network (e.g., a fifth-generation (5G) network, a long-term evolution (LTE) network, a third-generation (3G) network, a code division multiple access (CDMA) network, etc.), a public land mobile network (PLMN), a local area network (LAN), a wide area network (WAN), a metropolitan area network (MAN), a telephone network (e.g., a public switched telephone network (PSTN)), a private network, an ad hoc network, an intranet, the Internet, an optical fiber-based network, etc., and / or combinations thereof or other types of networks.
[0038] The number and arrangement of the devices and networks shown in FIG. 1 are provided as an example. In practice, there may be additional devices and / or networks, fewer devices and / or networks, different devices and / or networks, or devices and / or networks with different arrangements compared to those shown in FIG. 1. Further, two or more devices shown in FIG. 1 may be implemented within a single device, or a single device shown in FIG. 1 may be implemented as a plurality of distributed devices. Additionally or alternatively, a set of devices (e.g., one or more devices) of Environment 100 may perform one or more functions described as being performed by another set of devices of Environment 100.
[0039] FIG. 2 is a block diagram of exemplary components of one or more devices of FIG. 1.
[0040] Device 200 may correspond to User Device 110 and / or Platform 120. As shown in FIG. 2, Device 200 may include a bus 210, a processor 220, a memory 230, a storage component 240, an input component 250, an output component 260, and a communication interface 270.
[0041] Bus 210 includes components that enable communication among the components of device 200. Processor 220 is implemented in hardware, firmware, or a combination of hardware and software. Processor 220 is a central processing unit (CPU), a graphics processing unit (GPU), an accelerated processing unit (APU), a microprocessor, a microcontroller, a digital signal processor (DSP), a field programmable gate array (FPGA), an application specific integrated circuit (ASIC), or another type of processing component. In some implementations, processor 220 includes one or more processors that can be programmed to perform functions. Memory 230 includes random access memory (RAM), read only memory (ROM), and / or another type of dynamic or static storage device (e.g., flash memory, magnetic memory, and / or optical memory) for storing information and / or instructions for use by processor 220.
[0042] Storage component 240 stores information and / or software related to the operation and use of device 200. For example, storage component 240 may include, along with a corresponding drive, a hard disk (e.g., a magnetic disk, an optical disk, a magneto-optical disk, and / or a solid state disk), a compact disc (CD), a digital versatile disc (DVD), a floppy disk, a cartridge, a magnetic tape, and / or another type of non-transitory computer-readable medium.
[0043] The input component 250 includes components that enable the device 200 to receive information via user input (such as a touch screen display, keyboard, keypad, mouse, button, switch, and / or microphone). Additionally or alternatively, the input component 250 may include sensors for detecting information (such as a global positioning system (GPS) component, accelerometer, gyroscope, and / or actuator). The output component 260 includes components that provide output information from the device 200 (such as a display, speaker, and / or one or more light emitting diodes (LEDs)).
[0044] The communication interface 270 includes transceiver-like components (such as a transceiver and / or separate receiver and transmitter) that enable the device 200 to communicate with other devices via a wired connection, wireless connection, or a combination of wired and wireless connections. The communication interface 270 may enable the device 200 to receive information from another device and / or provide information to another device. For example, the communication interface 270 may include an Ethernet interface, optical interface, coaxial interface, infrared interface, radio frequency (RF) interface, universal serial bus (USB) interface, Wi-Fi interface, cellular network interface, and the like.
[0045] The device 200 can execute one or more of the processes described herein. The device 200 may execute these processes in response to the processor 220 executing software instructions stored by a non-transitory computer-readable medium such as the memory 230 and / or the storage component 240. The computer-readable medium is defined herein as a non-transitory memory device. The memory device includes a memory space within a single physical storage device or a memory space that spans multiple physical storage devices.
[0046] Software instructions may be read into memory 230 and / or storage component 240 from another computer-readable medium or from another device via communication interface 270. The software instructions stored in memory 230 and / or storage component 240, when executed, may cause processor 220 to execute one or more processes described herein. Additionally or alternatively, instead of or in combination with software instructions, hardwired circuitry may be used to execute one or more processes described herein. Thus, the implementations described herein are not limited to any particular combination of hardware circuitry and software.
[0047] The number and arrangement of components shown in FIG. 2 are provided as an example. In practice, device 200 may include additional components, fewer components, different components, or components in a different arrangement than those shown in FIG. 2. Additionally or alternatively, a set of components (e.g., one or more components) of device 200 may perform one or more functions described as being performed by another set of components of device 200.
[0048] FIG. 3 is a schematic diagram showing an exemplary model 300 for emotion extraction according to an embodiment of the present disclosure.
[0049] According to an embodiment, the input for the conversation mode sentiment analysis for understanding a multi-turn conversation may be one or more input dialogues. The multi-turn conversation may be an utterance of a dialogue including one or more sentences from one or more speakers. As an example, the multi-turn conversation may be a conversation before and after where the context of the previous question and / or sentence affects the response or the next question and / or sentence. The input dialogue may include one or more sentences. In some embodiments, the input for the conversation mode sentiment analysis for understanding a multi-turn conversation may be one or more input dialogues and / or sentences broken down into words. As an example, a list of utterances of a dialogue is X1, X2, ..., X i , where X i is a sentence of the utterance of the dialogue, and
Number
Number
[0050] Sentiment extraction may include extracting all sentiment expressions from the input dialogue. Polarity extraction may include extracting a polarity value corresponding to each sentiment. As an example, sentiment and / or polarity extraction (360) may include extracting all sentiment expressions {s1, ..., s M} and their polarity values {p1, ..., p i} from X M (sentiment extraction, SE). In some embodiments, each sentiment expression may be a word and / or phrase of the input dialogue. As an example, the sentiment expression s j may be a word or phrase in the order X i , and its polarity value p j is selected from three possible values: -1 (negative), 0 (neutral), and +1 (positive).
[0051] In some embodiments, a sentence encoder (320) can be used to identify emotional expressions and polarity values from an input dialogue. As an example, the sentence encoder (320) can be used, and the sentence encoder (320) can be modeled to handle the extraction of emotional expressions and the detection of their polarities as a sequence labeling task. In some embodiments, the sentence encoder (320) employs a pre-trained model such as a pre-trained BERT model to input words (310)
Number
Number
Number
[0052] In some embodiments, the context-dependent sentence embedding (330) may be input to a neutral network and / or a machine-learned model (340) to generate a plurality of tags for each input word, sentence, and / or dialogue. As an example, the context-dependent sentence embedding (330)
Number
Number
[0053] Figure 4 is a schematic diagram showing an exemplary model 400 for reference extraction according to an embodiment of the present disclosure.
[0054] In some embodiments, emotional expressions and their polarities can be input to a reference extractor model to extract corresponding references for at least one emotional expression. In some embodiments, for each emotional expression s j a corresponding reference m j can be extracted by employing a reference encoder (420). In some embodiments, reference extraction may be based on emotional expressions and context-dependent sentence embeddings. In some embodiments, reference extraction may be based on input concatenation (410) based on emotional expressions and context embeddings. As an example, all turns of the dialogue
Number
Number
Number
[0055] Referential extraction may require longer-distance inferences throughout the conversation. In some embodiments, rich features including turn-wise distance and speaker information for modeling cross-sentence correlations can be used. In some embodiments, a feature extractor (430) can be used to generate rich features including turn-wise distance and speaker information to model cross-sentence correlations. In some embodiments, the turn-wise distance may be the relative distance to the current turn bucketed into [0, 1, 2, 3, 4, 5+, 8+, 10+]. The speaker information may be a binary feature indicating whether the tokens in the conversation history are from the same speaker as the current turn. Both types of information may be represented by embeddings. As an example,
Number
Number
Number
Number
Number
Number
Number
[0056] In some embodiments, the average vector representation (450) representing the overall sentiment expression s j can be generated by averaging the context representations of all tokens therein. The average vector representation (450) can be represented using Equation (5).
Number
Number
Number
Number
[0057] According to an embodiment, a target mention (st, ed) can be generated by selecting both of the boundaries st and ed that result in the highest score from φ[st, ed], where st ≦ ed and st and ed may be in the same utterance.
[0058] FIG. 5 is a simplified flowchart showing an exemplary process 500 for sentiment analysis of a conversation mode according to an embodiment of the present disclosure.
[0059] In operation 510, sentiment expressions can be extracted from embeddings of sentences corresponding to the input dialogue, sentences, and / or words. As an example, sentiment expressions can be extracted using the input words (310). In some examples, sentiment expressions may be extracted from the input dialogue, sentences, and / or words using an encoder. As an example, a sentence encoder (320) can be used to extract sentiment expressions. In some embodiments, a specific sentence encoder can be used. In some embodiments, any method and / or model can be used as the encoder.
[0060] In some embodiments, there may be a preliminary operation performed before extracting sentiment expressions including obtaining the input dialogue, sentences, and / or words. In some embodiments, extracting sentiment expressions can include generating a plurality of tags using a pre-trained machine learning model and inferring a sentiment expression based on the plurality of tags. As an example, extracting sentiment expressions can include generating one or more tags (350) using a pre-trained machine learning model (340) and inferring a sentiment expression based on the plurality of tags. As an example, a pre-trained BERT model and / or an attention layer can be used to generate a plurality of tags and infer a sentiment expression from the tags.
[0061] In operation 520, the polarity value can be extracted from the embedding of the sentence corresponding to the input dialogue, sentence, and / or word. As an example, the input word (310) can be used to extract the polarity value. The polarity value can be associated with one or more sentiment expressions. In some embodiments, each polarity value can be associated with a sentiment expression. In some examples, the polarity value may be extracted from the input dialogue, sentence, and / or word using an encoder. As an example, the sentence encoder (320) can be used to extract the polarity value. In some embodiments, a specific sentence encoder can be used. In some embodiments, any method and / or model can be used as the encoder.
[0062] In operation 530, a target mention can be determined based on the sentiment expression, the polarity value, and / or the sentence embedding. In some embodiments, the target mention can be associated with at least one sentiment expression. Determining the target mention of the sentiment expression can include, at 540, generating a rich context expression based on the sentence embedding and the sentiment expression. Determining the target mention of the sentiment expression can also include, at 550, determining the target mention based on the calculated boundary, where the calculated boundary is generated using the rich context expression. As an example, the rich context expression (440) can be generated using the distance embedding, speaker embedding, sentence embedding, and / or context embedding generated by the mention encoder (420) and / or the feature extractor (430). In some embodiments, the rich context expression (440) and the average vector expression (450) may be used as inputs to one or more attention models (460) to calculate the boundary.
[0063] FIG. 6 is a simplified flowchart showing an exemplary process 600 for sentiment analysis of a conversation mode according to an embodiment of the present disclosure.
[0064] In operation 610, the input dialogue can be obtained. The input dialogue can include one or more sentences and / or words. In some embodiments, the input dialogue can include a multi-turn conversation with one or more speakers.
[0065] In operation 620, sentence embeddings can be generated using a sentence encoder. As an example, sentence embeddings can be generated using a sentence encoder (320). As an example, the sentence encoder (320) can be used, and the sentence encoder (320) can be modeled to handle the extraction of emotional expressions and the detection of their polarities as a sequence labeling task. In some embodiments, the sentence encoder (320) employs a pre-trained model such as a pre-trained BERT model to generate context-dependent embeddings for the input words (310)
Number
Number
[0066] In some embodiments, in operation 630, one or more tags can be generated based on the sentence embeddings using a pre-trained model. As an example, the context-dependent sentence embeddings (330) can be input into a neutral network and / or a machine-learned model (340) to generate a plurality of tags for each input word, sentence, and / or dialogue. In some embodiments, the context-dependent sentence embeddings (330) are the input words (310) (e.g.,
Number
[0067] FIG. 7 is a simplified flowchart showing an exemplary process 700 for sentiment analysis of a conversation mode according to an embodiment of the present disclosure.
[0068] In operation 710, sentiment expressions and sentence embeddings can be input to one or more models. By way of example, the sentiment expressions and sentence embeddings may be input to a reference encoder (420) and / or a feature extractor (430).
[0069] In operation 720, rich context representations can be generated using the sentiment expressions and sentence embeddings. In some embodiments, one or more models can be used to generate rich context representations based on the sentiment expressions and sentence embeddings. By way of example, a reference encoder (420) and / or a feature extractor (430) can be used to generate rich context representations based on the sentiment expressions and sentence embeddings.
[0070] In some embodiments, generating a rich context representation based on a sentence embedding and a sentiment expression can include generating a turn-wise distance based on the sentence embedding, generating speaker information based on the sentence embedding, and concatenating the turn-wise distance, speaker information, and sentiment expression to generate the rich context representation. In some embodiments, the speaker information can indicate whether the input dialogue is from the same speaker. In some embodiments, generating rich context information can also include generating an average vector representation representing the entire sentiment expression by averaging the context representations of all tokens therein.
[0071] In some embodiments, the referring encoder (420) can be implemented using one or more encoders based on self-attention and / or pre-trained BERT to obtain context embeddings. In some embodiments, the feature extractor (430) can be implemented using one or more encoders based on self-attention and / or pre-trained BERT to obtain context embeddings.
[0072] In operation 730, based on the rich context information, a distribution can be generated using at least two attention layers and / or attention models. As an example, the rich context representation (440) and the average vector representation (450) may be input into one or more attention models (460) to obtain one or more distributions.
[0073] In operation 740, the product of the distributions generated from each of the one or more attention layers can be determined. In some embodiments, determining the product of the generated distributions can include generating a plurality of scores.
[0074] In operation 750, a target associated with at least one sentiment expression reference can be determined based on the boundaries of the distribution. In some embodiments, determining the target reference can include selecting a boundary of the distribution based on the highest score from a plurality of scores. In some embodiments, determining the target reference can include selecting a boundary of the distribution based on the highest score from a plurality of scores, and the plurality of scores are generated by determining the product of the distributions of each of the one or more attention layers. As an example, the target reference may be generated by selecting a boundary from each attention model that yields the highest score from the product of the distributions. In some embodiments, the boundary selected from one attention model may be smaller than the boundary selected from the other attention model. In some embodiments, both boundaries can belong to the same utterance.
[0075] Exemplary advantages of the present disclosure can be described as follows.
[0076] Table 2 shows the performance of embodiments of the present disclosure. As can be seen in Table 2, the present disclosure using the BERT model presents the best scores in the identification of emotions and references in multi-turn conversations.
[0077]
Table 2
[0078] Table 3 shows the performance of embodiments of the present disclosure. As can be seen in Table 3, the present disclosure using one or more transformers yields the best scores in the identification of emotions and references in multi-turn conversations.
[0079]
Table 3
[0080] The average length of the knowledge utilized, as evidenced in the column "Avg.KN Len." as seen in Table 3. Using the complete news document, the BLEU score increases slightly, but the diversity of the output decreases as shown by the Distinct score. Taking only the selected segments according to embodiments of the present disclosure improves the diversity regarding the Distinct score and shows an equivalent BLEU score. More importantly, in embodiments of the present disclosure, only an average of 29 Chinese characters are selected, while the baseline for the entire document uses 765 characters. This indicates that embodiments of the present disclosure can save 96% of the memory usage to represent the relevant knowledge.
[0081] Figures 5 through 7 illustrate exemplary blocks of processes 500, 600, and 700. However, in implementations, processes 500, 600, and 700 may include additional blocks, fewer blocks, different blocks, or blocks in a different arrangement than those shown in Figures 5 through 7. In embodiments, any blocks of processes 500, 600, and 700 may be combined or arranged in any amount or order as needed. In embodiments, two or more of the blocks of processes 500, 600, and 700 may be executed in parallel.
[0082] The foregoing techniques may be implemented using computer-readable instructions, as computer software physically stored on one or more computer-readable media, or by one or more specifically configured hardware processors. For example, FIG. 1 illustrates an environment 100 suitable for implementation of various embodiments.
[0083] Computer software may be encoded using any suitable machine code or computer language that may be subject to mechanisms such as assembly, compilation, linking, etc. to create code including instructions that can be executed directly, or via interpretation, execution of microcode, etc., by a computer central processing unit (CPU), a graphics processing unit (GPU), etc.
[0084] The instructions may be executed on various types of computers or computer components, including, for example, personal computers, tablet computers, servers, smartphones, gaming consoles, Internet of Things devices, etc.
[0085] Although several exemplary embodiments of the present disclosure have been described, there are modifications, substitutions, and various alternative equivalents within the scope of the present disclosure. Accordingly, it will be understood that those skilled in the art can devise numerous systems and methods that embody the principles of the present disclosure and thus fall within the spirit and scope of the present disclosure, even though not explicitly shown or described herein.
Description of Symbols
[0086] 100 Environment 110 User Device 120 Platform 122 Cloud Computing Environment 124 Computing Resources 124-1 Application 124-2 Virtual Machine 124-3 Virtualized Storage 124-4 Hypervisor 130 Network 200 Device 210 Bus 220 Processor 230 Memory 240 Storage Component 250 Input Component 260 Output Component 270 Communication Interface 300 Model 310 Input Word 320 Sentence Encoder 330 Context-Dependent Embedding 330 Context-Dependent Sentence Embedding 340 Pre-Trained Machine Learning Model 350 Tag 360 Sentiment and / or Polarity Extraction 400 Model 410 Input Concatenation Based on Context Embedding 420 Referral Encoder 430 Feature Extractor 440 Rich Context Representation 450 Average Vector Representation 460 Attention Model 470 Distribution
Claims
1. A method for sentiment analysis for multi-turn conversations, the method comprising: obtaining the input conversation; extracting sentiment expressions based on the embedding of the sentences corresponding to the input conversation; generating a polarity value based on the embedding of the sentences corresponding to the input conversation; determining a target mention associated with at least one of the sentiment expressions based on the sentiment expression and the embedding of the sentences, wherein the step of determining the target mention generates a rich context expression based on the embedding of the sentences and the sentiment expression, generates a turn-wise distance which is a distance embedding representing the relative distance to the current turn, based on the embedding of the sentences; generates speaker information based on the embedding of the sentences, the speaker information indicating whether the input conversation is from the same speaker; generates the rich context expression by concatenating the turn-wise distance, the speaker information, and the sentiment expression; and determines the target mention based on the calculated boundary, generates a dialogue history expression by concatenating the rich context expressions; calculates the distribution of the start and end boundaries of the target mention using one or more attention layers; generates the target mention by selecting the start and end boundaries that result in the highest score from a plurality of scores, the plurality of scores being generated by determining the product of the distributions of each of the one or more attention layers; and A method comprising.
2. The step of determining the target reference based on the calculated boundary comprises: generating a distribution based on the rich context expression and the sentiment expression using the one or more attention layers; determining the target reference based on the boundary of the distribution. The method according to claim 1.
3. The embedding of the text is generated based on the input dialogue, the method according to claim 1.
4. The step of extracting the sentiment expression from the text embedding comprises: generating a plurality of tags using a pre-trained machine learning model; inferring the sentiment expression based on the plurality of tags. The method according to claim 1.
5. The step of generating the polarity value from the text embedding comprises: generating a plurality of tags using a pre-trained machine learning model; inferring the polarity value based on the plurality of tags. The method according to claim 1.
6. The polarity value is one of positive, negative, or neutral, the method according to claim 1.
7. An apparatus for sentiment analysis for multi-turn conversations, configured to perform the method according to any one of claims 1 to 6.
8. A program which, when executed by a computer, causes the computer to perform the method according to any one of claims 1 to 6.
Citation Information
Patent Citations
Tag estimation device, tag estimation method and program
JP2020052611A
Generating responses in automated chatting
US20200159997A1