Program, information processing device and information processing method

A program using AI to analyze conversation data for both emotional and stylistic insights enhances user interaction by providing real-time emotional and conversational style feedback, improving satisfaction and response efficiency.

JP7810962B2Active Publication Date: 2026-02-04PKUTECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022040588
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-15
Publication Date
2026-02-04
Estimated Expiration
2042-03-15

AI Technical Summary

Technical Problem

Existing technologies are unable to obtain style information relating to a speaker's conversation style, limiting the comprehensive understanding of user emotions.

Method used

A program that utilizes two learning models to output emotion and style information by analyzing conversation data, employing AI to classify emotions into categories and identify conversation styles, displaying the results chronologically.

Benefits of technology

Enables the simultaneous output of affective and stylistic information, allowing operators to adapt their interaction style to user preferences, improving user satisfaction and reducing response times.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007810962000001
    Figure 0007810962000001
  • Figure 0007810962000002
    Figure 0007810962000002
  • Figure 0007810962000003
    Figure 0007810962000003
Patent Text Reader

Abstract

To provide a program and the like that enable output of both emotion information concerning user's emotion and style information concerning a conversation style.SOLUTION: A program according to an aspect causes a computer to execute processing of: acquiring conversation data on a user who talks with an operator; inputting the acquired conversation data to a first learning model learned to output emotion information concerning user's emotion in a case where the conversation data is input, and outputting the emotion information on the user; inputting the acquired conversation data to a second learning model learned to output style information concerning a conversation style of the user in the case where the conversation data is input, and outputting the style information on the user; and outputting the emotion information and the style information to an operation terminal of the operator.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a program, an information processing device, and an information processing method. [Background technology]

[0002] In recent years, there has been active development of technology for estimating a speaker's emotions based on dialogue data (voice data) obtained by voice recognition technology. For example, Patent Document 1 discloses an information processing device that obtains the content of an utterance and the speaker's emotions when recognizing the voice uttered by the speaker. [Prior art documents] [Patent documents]

[0003] [Patent Document 1] Patent Publication No. 2021-124530 Summary of the Invention [Problem to be solved by the invention]

[0004] However, the invention of Patent Document 1 has a problem in that it is not possible to obtain style information relating to the conversation style of the speaker (user).

[0005] One aspect of the present invention is to provide a program or the like that is capable of outputting both emotion information relating to a user's emotion and style information relating to a conversation style. [Means for solving the problem]

[0006] A program according to one aspect acquires conversation data of a user who is conversing with an operator, inputs the acquired conversation data into a first learning model that has been trained to output emotion information relating to the user's emotions when the conversation data is input, and outputs emotion information of the user that is classified into a plurality of levels for each of categories of calm, pleasure, discomfort, and depression, and when the conversation data is input: Slow speaking type and quick speaking typeThe acquired conversation data is input into a second learning model that has been trained to output style information regarding the user's conversation style, including the acquired conversation data, and the style information of the user is output. The second learning model is then trained to output style information regarding the user's conversation style, including the acquired conversation data, and the emotion information, including the emotion classification and level, and the style information are output to the operator's operation terminal. The computer is then caused to execute a process of displaying icons indicating the emotion classification and level in a chronological order in a predetermined time sequence. [Effects of the Invention]

[0007] In one aspect, it is possible to output both affective information relating to the user's emotions and style information relating to the conversation style. [Brief explanation of the drawings]

[0008] [Figure 1] FIG. 1 is an explanatory diagram illustrating an overview of an emotion analysis system. [Figure 2] FIG. 2 is a block diagram illustrating an example of the configuration of a server. [Figure 3] FIG. 2 is an explanatory diagram showing an example of the record layout of a training data DB and a user DB. [Figure 4] FIG. 10 is an explanatory diagram showing an example of a record layout of a conversation data DB. [Figure 5] FIG. 2 is a block diagram illustrating a configuration example of an operator terminal. [Figure 6] FIG. 1 is an explanatory diagram of an emotion identification model. [Figure 7] FIG. 1 is an explanatory diagram of a style identification model. [Figure 8] 10 is a flowchart showing a processing procedure when emotion information and style information are output. [Figure 9] FIG. 10 is an explanatory diagram showing an example of a display screen for emotion information and style information. [Figure 10] FIG. 10 is a block diagram showing an example of the configuration of a server in Modification 1. [Figure 11] FIG. 10 is an explanatory diagram illustrating an example of a record layout of an advice DB. [Figure 12]10 is a flowchart showing a processing procedure for identifying advice. [Figure 13] FIG. 10 is an explanatory diagram showing an example of a record layout of a training data DB in the second embodiment. [Figure 14] FIG. 10 is an explanatory diagram of an emotion identification model in the second embodiment. [Figure 15] FIG. 11 is an explanatory diagram showing an example of a display screen for emotion information and style information in the second embodiment. [Figure 16] FIG. 11 is an explanatory diagram showing an example of a record layout of a conversation data DB in the third embodiment. [Figure 17] FIG. 11 is an explanatory diagram showing an example of a display screen for emotion information and style information in the third embodiment. [Figure 18] 10 is a flowchart showing a processing procedure for outputting both user emotion information and operator emotion information. DETAILED DESCRIPTION OF THE INVENTION

[0009] The present invention will be described in detail below with reference to the drawings showing embodiments thereof.

[0010] (Embodiment 1) The first embodiment relates to a form in which, based on conversation data of a user who is conversing with an operator, emotion information relating to the emotion of the user and style information relating to the conversation style are output using artificial intelligence (AI).

[0011] 1 is an explanatory diagram showing an overview of an emotion analysis system. The system of this embodiment includes an information processing device 1, an information processing terminal 2, and a receiving terminal 3, and each device transmits and receives information via a network N such as the Internet.

[0012] The information processing device 1 is an information processing device that processes, stores, and transmits / receives various types of information. The information processing device 1 is, for example, a server device, a personal computer, or a general-purpose tablet PC (personal computer). In this embodiment, the information processing device 1 is a server device, and for simplicity, will be referred to as server 1 below.

[0013] The information processing terminal 2 is an operator's operating terminal device that receives and displays the user's emotion information and style information. The information processing terminal 2 is an information processing device such as a smartphone, a mobile phone, a wearable device such as Apple Watch (registered trademark), a tablet, or a personal computer terminal. For simplicity, the information processing terminal 2 will be referred to as the operator terminal 2 below.

[0014] The receiving terminal 3 is a terminal device such as a telephone, a mobile phone, a smartphone, or a tablet capable of making calls.

[0015] The server 1 according to this embodiment acquires conversation data of a user who is conversing with an operator. The server 1 inputs the acquired conversation data into an emotion identification model (first learning model) that has been trained to output emotion information related to the user's emotion when the conversation data is input, and outputs the emotion information of the user.

[0016] The server 1 inputs the acquired conversation data into a style identification model (second learning model) that has been trained to output style information related to the user's conversation style when conversation data is input, and outputs the style information of the user. The server 1 outputs the user's emotion information output from the emotion identification model and the style information output from the style identification model to the operator terminal 2 of the operator who is conversing with the user.

[0017] 2 is a block diagram showing an example of the configuration of the server 1. The server 1 includes a control unit 11, a storage unit 12, a communication unit 13, an input unit 14, a display unit 15, a reading unit 16, and a large-capacity storage unit 17. Each component is connected by a bus B.

[0018] The control unit 11 includes an arithmetic processing device such as a CPU (Central Processing Unit), an MPU (Micro-Processing Unit), a GPU (Graphics Processing Unit), an FPGA (Field Programmable Gate Array), a DSP (Digital Signal Processor), or a quantum processor. The control unit 11 reads and executes a control program 1P (program product) stored in the storage unit 12, thereby performing various information processing, control processing, and the like related to the server 1.

[0019] The control program 1P can be deployed to run on a single computer, or on multiple computers located at one site, or distributed across multiple sites and interconnected by a communications network. While the control unit 11 is illustrated in FIG. 2 as a single processor, it may also be a multiprocessor.

[0020] The storage unit 12 includes memory elements such as RAM (Random Access Memory) and ROM (Read Only Memory), and stores the control program 1P or data required for the control unit 11 to execute processing. The storage unit 12 also temporarily stores data required for the control unit 11 to execute arithmetic processing. The communication unit 13 is a communication module for performing processing related to communication.

[0021] The input unit 14 is an input device such as a mouse, keyboard, touch panel, or button, and outputs received operation information to the control unit 11. The display unit 15 is a liquid crystal display, an organic EL (electroluminescence) display, or the like, and displays various information according to instructions from the control unit 11.

[0022] The reading unit 16 reads a portable storage medium 1a including a CD (Compact Disc)-ROM or a DVD (Digital Versatile Disc)-ROM. The control unit 11 may read the control program 1P from the portable storage medium 1a via the reading unit 16 and store it in the mass storage unit 17. Alternatively, the control unit 11 may download the control program 1P from another computer via a network N or the like and store it in the mass storage unit 17. Furthermore, the control unit 11 may read the control program 1P from the semiconductor memory 1b.

[0023] The mass storage unit 17 includes a recording medium such as an HDD (Hard disk drive), an SSD (Solid State Drive), etc. The mass storage unit 17 includes an emotion identification model (first learning model) 171, a style identification model (second learning model) 172, a training data DB (database) 173, a user DB 174, and a conversation data DB 175.

[0024] The emotion identification model 171 is a classifier that identifies emotion information related to the emotion of a user who is conversing with an agent based on conversation data of the user, and is a trained model generated by machine learning. The style identification model 172 is a classifier that identifies style information related to the conversation style of a user who is conversing with an agent based on conversation data of the user, and is a trained model generated by machine learning.

[0025] The training data DB 173 stores training data for constructing (generating) the emotion identification model 171 and the style identification model 172. The user DB 174 stores information about users. The conversation data DB 175 stores conversation data of conversations between users and operators, as well as emotion information and style information identified based on the conversation data.

[0026] In this embodiment, the storage unit 12 and the large-capacity storage unit 17 may be configured as an integrated storage device. Furthermore, the large-capacity storage unit 17 may be configured by a plurality of storage devices. Furthermore, the large-capacity storage unit 17 may be an external storage device connected to the server 1.

[0027] The server 1 may execute various information processing and control processing on a single computer, or may execute the processing in a distributed manner on multiple computers. The server 1 may also be realized by multiple virtual machines provided in a single server, or may be realized by using a cloud server.

[0028] FIG. 3 is an explanatory diagram showing an example of the record layout of the training data DB 173 and the user DB 174. The training data DB 173 includes a data type column, an input data column, and an output data column. The data type column stores the type of training data. In this embodiment, the types of training data include "emotion" and "style." "Emotion" is the type of training data for constructing the emotion identification model 171. "Style" is the type of training data for constructing the style identification model 172.

[0029] The input data string stores conversation data of the user and the operator. The output data string stores user emotion information when the type of training data is "emotion," and stores user style information when the type of training data is "style." Emotion information and style information will be described later.

[0030] The user DB 174 includes a user ID column, a name column, and a gender column. The user ID column stores a unique user ID to identify each user. The name column stores the user's name. The gender column stores the user's gender. Note that the information stored in the user DB 174 is not limited to the user ID, name, and gender. For example, information such as the user's age or address may also be stored in the user DB 174.

[0031] 4 is an explanatory diagram showing an example of a record layout of the conversation data DB 175. The conversation data DB 175 includes a conversation data ID column, a user ID column, a conversation data column, an emotion column, and a conversation style column. The conversation data ID column stores a unique ID of each conversation data item to identify the conversation data item.

[0032] The user ID column stores a user ID that identifies a user. The conversation data column stores conversation data of a user who conversed with an operator. The emotion column stores emotion information of the user that is identified by the emotion identification model 171 based on the conversation data of the user. The conversation style column stores style information of the user that is identified by the style identification model 172 based on the conversation data of the user.

[0033] The storage format of each DB described above is an example, and other storage formats may be used as long as the relationships between the data are maintained.

[0034] 5 is a block diagram showing an example of the configuration of the operator terminal 2. The operator terminal 2 includes a control unit 21, a storage unit 22, a communication unit 23, an input unit 24, and a display unit 25. Each component is connected by a bus B.

[0035] The control unit 21 includes an arithmetic processing unit such as a CPU, an MPU, etc., and performs various information processing, control processing, etc. related to the operator terminal 2 by reading and executing a control program 2P (program product) stored in the storage unit 22. Note that although the control unit 21 is described in Fig. 5 as being a single processor, it may be a multiprocessor.

[0036] The storage unit 22 includes memory elements such as RAM and ROM, and stores the control program 2P or data required for the control unit 21 to execute processing. The storage unit 22 also temporarily stores data required for the control unit 21 to execute arithmetic processing.

[0037] The communication unit 23 is a communication module for performing communication-related processing, and transmits and receives information to and from the server 1, etc. via the network N. The input unit 24 may be a keyboard, a mouse, or a touch panel integrated with the display unit 25. The display unit 25 is a liquid crystal display, an organic EL display, or the like, and displays various information according to instructions from the control unit 21.

[0038] The operator terminal (edge ​​terminal) 2 may have all or part of the functions of the server 1.

[0039] 6 is an explanatory diagram of the emotion identification model 171. The emotion identification model 171 is used as a program module that is part of artificial intelligence software. The emotion identification model 171 is a learning model that outputs emotion information related to the emotion of a user when conversation data in which the user converses with an operator is input.

[0040] In this embodiment, the emotional information includes at least two of calm, pleasure, discomfort, and depression. Note that the emotional information is not limited to the above-mentioned categories of emotions, and may include, for example, anger, fear, sadness, surprise, etc.

[0041] The emotion identification model 171 according to this embodiment performs emotion information identification processing using, for example, a BERT (Bidirectional Encoder Representations from Transformers) model. The emotion identification model 171 has a neural network structure in which a plurality of neurons are interconnected. The emotion identification model 171 includes an input layer that receives one or more pieces of data as input, an intermediate layer that performs arithmetic processing on the data received in the input layer, and an output layer that aggregates the arithmetic results of the intermediate layer and outputs one or more values.

[0042] The emotion identification model 171 is a trained model that has undergone a learning process in advance. The learning process is a process of setting appropriate values ​​for the coefficients and thresholds of each neuron that constitutes the neural network using a large amount of training data that has been given in advance. The emotion identification model 171 according to this embodiment undergoes learning process using training data stored in a training data DB 173. Input data for the training data is, for example, conversation data for the first 30 seconds of a conversation, and output data is emotion information related to emotions.

[0043] The above-described learning process may be performed by another computer (not shown) and the emotion identification model 171 may be deployed. Instead of constructing the emotion identification model 171, emotion information may be identified by using a WEB API (Application Programming Interface) that uses a machine learning model.

[0044] When the server 1 acquires conversation data of a user, it inputs the acquired conversation data into the emotion identification model 171 and converts the conversation data into text. The server 1 then outputs emotion information identified from the converted text data. The emotion information is neutral, pleasant, unpleasant, or depressed. Note that the BERT model is an existing technology, so a detailed description will be omitted. As shown in the figure, the server 1 inputs the conversation data into the emotion identification model 171 and outputs emotion information that has been identified as "pleasant."

[0045] The emotion identification model 171 is not limited to BERT, and may be realized by other models such as a deep neural network (DNN), a universal sentence encoder, a recurrent neural network (RNN), a long short-term memory (LSTM), a logistic regression, an SVM, a k-NN, a decision tree, a naive Bayes classifier, or a random forest. Alternatively, the emotion identification model 171 may be realized by a speech recognition framework using a Transformer. The speech recognition framework may be, for example, wav2vec 2.0, which employs a self-supervised learning method using a small amount of labeled speech data and a large amount of unlabeled speech data.

[0046] Note that the process of identifying emotional information is not limited to using the emotion identification model 171. For example, emotional information may be identified based on the feature quantities of conversation data. Specifically, the server 1 extracts feature quantities based on the pitch (pitch indicating the high and low of the voice), speech rate (speech rate or tempo) or intonation of the user's speech from the user's conversation data. The server 1 identifies emotional information based on the feature quantities of the extracted conversation data or the amount of change in the feature quantities.

[0047] Furthermore, emotional information can be identified based on the conversation data that has been converted into text. For example, the server 1 converts the conversation data into text and extracts words that particularly express emotions from the converted text data. The server 1 then performs an emotional information identification process based on the extracted words.

[0048] 7 is an explanatory diagram of the style identification model 172. The style identification model 172 is used as a program module that is part of artificial intelligence software. The style identification model 172 is a learning model that outputs style information related to the conversation style of a user when conversation data in which the user converses with an operator is input.

[0049] The style identification model 172 according to this embodiment performs a process of identifying style information using, for example, a BERT model. The style identification model 172 has a neural network structure in which a plurality of neurons are interconnected.

[0050] The style identification model 172 includes an input layer that accepts one or more pieces of data as input, an intermediate layer that performs calculations on the data accepted by the input layer, and an output layer that aggregates the calculation results of the intermediate layer and outputs one or more values. The output layer includes, for example, a sigmoid function or a softmax function, and outputs probability values ​​for various estimated conversation styles based on the calculation results of the intermediate layer. The probability values ​​are, for example, values ​​greater than 0 and less than 1.

[0051] The style identification model 172 is a trained model that has undergone a learning process in advance. The style identification model 172 undergoes the learning process using training data stored in the training data DB 173. The input data of the training data is, for example, conversation data for the first 30 seconds of a conversation, and the output data is style information related to the conversation style.

[0052] The above-described learning process may be performed by another computer (not shown) and the style identification model 172 may be deployed. Instead of constructing the style identification model 172, style information related to conversational styles may be identified by using a WEB API that uses a machine learning model.

[0053] When the server 1 acquires conversation data of a user, it inputs the acquired conversation data into the trained style identification model 172 and converts the conversation data into text. Then, the server 1 outputs style information identified from the converted text data. The style information in this embodiment includes a slow speaking type and a brisk speaking type. Note that the conversation style classification is not limited to the above-mentioned type, and conversation styles may be classified into, for example, a slow speaking type, a normal speaking type, and a brisk speaking type. Alternatively, conversation styles may be classified into a poor speaking type, a normal speaking type, and a good speaking type.

[0054] As shown in the figure, the classification results output for the conversation data are "0.16" and "0.84" as probability values ​​for "slow speaking type" and "quick speaking type," respectively.

[0055] Furthermore, a predetermined threshold value may be used to output the classification result. For example, if the server 1 determines that the probability value (0.84) of the "quick-speaking type" is equal to or greater than a predetermined threshold value (e.g., 0.80), the server 1 outputs the "quick-speaking type" as the classification result. Note that, without using the above-mentioned threshold value, the conversation style corresponding to the highest probability value from the probability values ​​of the various conversation styles classified by the style classification model 172 may be output as the classification result.

[0056] The style identification model 172 is not limited to BERT, and may be realized by other models such as DNN, logistic regression, SVM (Support Vector Machine), k-NN (k-Nearest Neighbor algorithm), decision tree, naive Bayes classifier, or random forest. Alternatively, the style identification model 172 may be realized by a speech recognition framework using Transformer (e.g., wav2vec 2.0).

[0057] 8 is a flowchart showing the processing steps when emotion information and style information are output. The control unit 11 of the server 1 acquires conversation data of a user who is conversing with an operator in real time from the receiving terminal 3 used by the user via the communication unit 13 (step S101). Note that the control unit 11 may acquire the user's conversation data from the receiving terminal 3 at predetermined time intervals (for example, every 10 seconds).

[0058] The control unit 11 determines whether the style information of the user has been identified (step S102). For example, if a flag (variable) for determining the identification status of the style information is provided, the control unit 11 sets the initial value of the flag to "FALSE." Once the control unit 11 has identified the style information based on the conversation data, the control unit 11 sets the flag to "TRUE." Then, the control unit 11 may determine whether the style information has been identified based on the value of the flag.

[0059] If the control unit 11 determines that the style information has been identified (YES in step S102), the process proceeds to step S105, which will be described later. If the control unit 11 determines that the style information has not yet been identified (NO in step S102), the control unit 11 identifies the user's style information using the acquired conversation data (step S103). Specifically, the control unit 11 inputs the user's conversation data to the style identification model 172, and outputs the identification result of the user's conversation style as style information. The style information includes the classification of the identified conversation style (slow speaking type or quick speaking type), a probability value, etc.

[0060] The order in which the emotion information identification process and the style information identification process are performed does not matter. For example, the emotion information identification process may be performed in parallel with the style information identification process.

[0061] The control unit 11 transmits the identified user's style information to the operator terminal 2 via the communication unit 13 (step S104). The control unit 21 of the operator terminal 2 receives the user's style information transmitted from the server 1 via the communication unit 23 (step S201). The control unit 21 displays the received user's style information on the display unit 25 (step S202).

[0062] The control unit 11 of the server 1 identifies the user's emotional information using the acquired conversation data (step S105). Specifically, the control unit 11 inputs the acquired user's conversation data into the emotion identification model 171, and outputs the identification result of the user's emotion as emotional information. The emotional information includes the classification of the identified emotion (calm, pleasant, unpleasant, or depressed), etc.

[0063] The control unit 11 transmits the identified user emotion information to the operator terminal 2 via the communication unit 13 (step S106). The control unit 21 of the operator terminal 2 receives the user emotion information transmitted from the server 1 via the communication unit 23 (step S203). The control unit 21 displays the received user emotion information on the display unit 25 (step S204).

[0064] The control unit 11 of the server 1 determines whether or not the dialogue between the user and the operator has ended through the receiving terminal 3 (step S107). When the control unit 11 determines that the dialogue has not ended (NO in step S107), the process returns to step S101.

[0065] If control unit 11 determines that the dialogue has ended (YES in step S107), control unit 11 stores the identified emotion information and style information in conversation data DB 175 of mass storage unit 17 in association with the conversation data (step S108). Specifically, control unit 11 assigns a conversation data ID to the conversation data. Control unit 11 stores the user ID, conversation data, emotion information, and style information as one record in conversation data DB 175 in association with the assigned conversation data ID. Control unit 11 then terminates the process.

[0066] In this embodiment, once a user's style information has been identified, the identification process is not performed again on the user's style information. Of course, similar to the identification process for emotion information, the identification process for style information based on conversation data may be performed in real time.

[0067] 9 is an explanatory diagram showing an example of a display screen for emotion information and style information. The screen includes a user information display field 11a, a style information display field 11b, and a user emotion information display field 11c. The user information display field 11a is a display field for displaying information about the user. The style information display field 11b is a display field for displaying style information related to the user's conversation style. The user emotion information display field 11c is a display field for displaying emotion information related to the user's emotions.

[0068] The server 1 acquires user information including the user's name, gender, etc. from the user DB 174 based on the user ID. The server 1 transmits the acquired user information to the operator terminal 2. The operator terminal 2 displays the user information transmitted from the server 1 in the user information display field 11a. As shown in the figure, the user information including the user ID, name, and gender is displayed in the user information display field 11a.

[0069] When a user is conversing with an operator using the receiving terminal 3, the server 1 acquires the user's conversation data in real time from the receiving terminal 3. The server 1 may also acquire the user's conversation data at predetermined time intervals (for example, every 10 seconds).

[0070] The server 1 identifies the style information of the user using the style identification model 172 so as to output style information of the user when conversation data of the user is input. The style information includes the classification and probability value of the identified conversation style. Note that the style information identification process may be executed only once. The server 1 transmits the identified style information of the user to the operator terminal 2. The operator terminal 2 displays the style information transmitted from the server 1 in the style information display field 11b. As shown in the figure, the style information that has been determined to be "84% efficient speaking type" is displayed in the style information display field 11b.

[0071] When the server 1 receives user conversation data, it uses the emotion identification model 171 to identify the emotional information of the user during conversation. The emotional information includes the classification of the identified emotion. The server 1 transmits the identified emotional information of the user to the operator terminal 2. The operator terminal 2 receives the emotional information transmitted from the server 1 and displays the received emotional information in chronological order in the user emotion information display field 11c.

[0072] As shown in the figure, icons indicating emotion classifications (calm, pleasant, unpleasant, or depressed) at each time point (0:00, 0:10, 0:20, etc.) of the conversation data are displayed in the user emotion information display field 11c. The operator terminal 2, for example, stores time-series conversation data and emotion classifications corresponding to the conversation data at each time point as spool data in the storage unit 22. When the operator terminal 2 acquires an emotion classification corresponding to the conversation data at the current time point (e.g., 0:50), it acquires emotion classifications at each past time point (e.g., 0:00, 0:10, 0:20, 0:30, and 0:40) that are accumulated in the storage unit 22. The operator terminal 2 displays icons indicating emotion classifications in chronological order from left to right based on the acquired emotion classifications at each past time point and the current emotion classification.

[0073] According to this embodiment, it is possible to identify the emotion information of a user using the emotion identification model 171 based on the conversation data of the user.

[0074] According to this embodiment, it is possible to identify style information of a user using the style identification model 172 based on the conversation data of the user.

[0075] According to this embodiment, the operator appropriately changes the speaking style to suit the user's preference in accordance with the user's emotion information and style information, thereby making it possible to improve the user's satisfaction.

[0076] According to this embodiment, even operators who are unfamiliar with handling complaints can accurately grasp the user's emotions, making it possible to reduce the operator's response time or reduce the turnover rate.

[0077] <Variation 1> In this modification, a process of specifying advice according to emotion information output from emotion identification model 171 and style information output from style identification model 172 will be described.

[0078] Fig. 10 is a block diagram showing an example of the configuration of the server 1 in Modification 1. Note that the same reference numerals are used to denote the same parts as in Fig. 2, and the description thereof will be omitted. The mass storage unit 17 stores an advice DB 176. The advice DB 176 stores advice corresponding to emotion information and style information.

[0079] The advice is determined in advance based on the emotion information and style information, and stored in the advice DB 176. For example, the advice corresponding to "emotion information: unpleasant, style information: quick speaking type" may be "Respond immediately!". Alternatively, the advice corresponding to "emotion information: calm, style information: slow speaking type" may be "Keep it up!".

[0080] 11 is an explanatory diagram showing an example of a record layout of the advice DB 176. The advice DB 176 includes an advice ID column, an emotion column, a conversation style column, and an advice column. The advice ID column stores a uniquely specified advice ID to identify each piece of advice. The emotion column stores emotion information related to emotions. The conversation style column stores style information related to conversation styles. The advice column stores the content of advice.

[0081] Fig. 12 is a flowchart showing the processing steps for identifying advice. Note that the same reference numerals are used to designate the same parts as in Fig. 8, and the description thereof will be omitted. After executing the processing of step S105, the control unit 11 of the server 1 identifies advice according to the identified emotion information and style information (step S111).

[0082] Specifically, the control unit 11 acquires advice corresponding to the identified emotion information and style information from the advice DB 176 of the mass storage unit 17. When emotion information and style information are input, the control unit 11 may specify advice using an advice specification model or the like that has been trained to output advice corresponding to the emotion information and style information. The advice specification model may be constructed by fine tuning a language model such as GPT (Generative Pre-trained Transformer), GPT-2, or GPT-3 that uses a deep learning technique called Transformer.

[0083] The control unit 11 transmits the identified emotion information and the specified advice to the operator terminal 2 via the communication unit 13 (step S112). The control unit 21 of the operator terminal 2 receives the emotion information and advice transmitted from the server 1 via the communication unit 23 (step S211). The control unit 21 displays the received emotion information and advice on the display unit 25 (step S212).

[0084] According to this modification, it is possible to instantly present to the operator advice that is specified according to the user's emotion information and style information.

[0085] (Embodiment 2) The second embodiment relates to a form in which each of the categories of calm, pleasure, discomfort, and depression is classified into a plurality of levels using the emotion identification model 171. Note that a description of the contents overlapping with the first embodiment will be omitted.

[0086] In the first embodiment, an example was described in which the emotion classification included at least two of calm, pleasure, discomfort, and depression. In the present embodiment, a process for further classifying emotions within each classification will be described. For example, emotions within each classification may be further classified into two levels. Specifically, the "calm" emotion is classified into "calm level 1" and "calm level 2." "Calm level 1" indicates a calm mood, and "calm level 2" indicates a very calm mood.

[0087] "Pleasure" emotions are classified into "Pleasure Level 1" and "Pleasure Level 2." "Pleasure Level 1" indicates a feeling of slight happiness, and "Pleasure Level 2" indicates a feeling of good mood. "Unpleasant" emotions are classified into "Unpleasant Level 1" and "Unpleasant Level 2." "Unpleasant Level 1" indicates feeling irritated or anxious, and "Unpleasant Level 2" indicates feeling angry.

[0088] The "depressed" emotion is classified into "depression level 1" and "depression level 2." "Depression level 1" indicates feeling depressed, and "depression level 2" indicates feeling very depressed. Note that, although the present embodiment has described an example in which emotions in each category are classified into two levels, this is not limiting, and emotions may be classified into multiple levels (for example, three levels) according to actual needs.

[0089] FIG. 13 is an explanatory diagram showing an example of a record layout of the training data DB 173 in the second embodiment. Note that description of content overlapping with FIG. 3 will be omitted. When the type of training data is "emotion," the output data column stores emotion information including the emotion classification and the emotion level. For example, when each classification of calm, pleasure, discomfort, and depression is further classified into two levels, the output data column stores "calm level 1," "calm level 2," "pleasure level 1," "pleasure level 2," "discomfort level 1," "discomfort level 2," "depression level 1," or "depression level 2."

[0090] Fig. 14 is an explanatory diagram of the emotion identification model 171 in embodiment 2. Note that explanations of the contents overlapping with Fig. 6 will be omitted. The emotion identification model 171 is a learning model that outputs emotion information including the classification (type; category) and level of the emotion of the user when conversation data in which the user talks with an operator is input.

[0091] The emotion identification model 171 according to this embodiment performs an emotion information identification process using, for example, a BERT model. The emotion identification model 171 performs the process using training data stored in a training data DB 173. Input data for the training data is, for example, conversation data for the first 30 seconds of a conversation, and output data is emotion classifications and levels. The server 1 trains the emotion identification model 171 using the training data.

[0092] When the server 1 acquires conversation data of a user, it inputs the acquired conversation data into the trained emotion identification model 171 and converts the conversation data into text. Then, the server 1 outputs emotion information identified from the converted text data. The emotion information includes the classification of the emotion (calm, pleasant, unpleasant, depressed, etc.) and the level of the emotion (level 1, level 2, etc.).

[0093] As shown in the figure, the classification results output for the conversation data are "Calm Level 1," "Calm Level 2," "Pleasure Level 1," "Pleasure Level 2," "Discomfort Level 1," "Discomfort Level 2," "Depression Level 1," and "Depression Level 2," with respective probability values ​​of "0.76," "0.10," "0.03," "0.02," "0.03," "0.03," "0.01," and "0.02."

[0094] Furthermore, a predetermined threshold value may be used to output the classification result. For example, if the server 1 determines that the probability value (0.76) of "calm level 1" is equal to or greater than a predetermined threshold value (e.g., 0.70), it outputs "calm level 1" as the classification result. Note that, without using the above-mentioned threshold value, the emotion corresponding to the highest probability value from the probability values ​​of various emotions classified by the emotion classification model 171 may be output as the classification result.

[0095] Fig. 15 is an explanatory diagram showing an example of a display screen for emotion information and style information in embodiment 2. Note that the same reference numerals are used to designate the same contents as in Fig. 9, and the description thereof will be omitted.

[0096] When the server 1 receives conversation data of a user, it uses the emotion identification model 171 to identify the category and level of the emotion of the user during the conversation so as to output emotion information of the user. The server 1 transmits the identified category and level of the emotion of the user to the operator terminal 2. The operator terminal 2 receives the category and level of the emotion transmitted from the server 1, and displays the received category and level of the emotion in chronological order in the user emotion information display field 11c.

[0097] As shown in the figure, icons indicating the emotional classification (calm, pleasant, unpleasant, or depressed) and level at each time point (0:00, 0:10, 0:20, etc.) of the conversation data are displayed in the user emotional information display field 11c. The levels include level 1, which is indicated by diagonal hatching sloping downward to the right, and level 2, which is indicated by diamond-shaped hatching enclosed in a frame. Note that, although the emotional levels are indicated by hatching patterns in FIG. 15, this is not limiting. For example, the emotional levels may be displayed in different colors. Alternatively, the emotional classification may be accompanied by a number indicating the quantitative level.

[0098] Furthermore, advice prepared according to emotion information including the category and level of emotion and style information can be stored in advance in the advice DB 176. For example, the advice DB 176 may store advice such as "Be careful not to make the user feel uncomfortable!" in response to "emotion information: unpleasant level 1, style information: quick talker," and advice such as "Talk while soothing the user!" in response to "emotion information: unpleasant level 2, style information: quick talker."

[0099] When the server 1 acquires emotion information including the classification and level of the user's emotion and style information using the emotion identification model 171, the server 1 acquires corresponding advice from the advice DB 176 according to the acquired emotion information and style information. The server 1 transmits the acquired advice to the operator terminal 2.

[0100] The process of specifying advice is not limited to the above. For example, advice may be provided based only on emotion information including the emotion classification and level. In this case, the server 1 acquires the appropriate advice from the advice DB 176 based on the emotion classification and level. The server 1 transmits the acquired advice to the operator terminal 2.

[0101] According to this embodiment, by further classifying emotions into a plurality of levels for each category, it is possible to improve the accuracy of identifying emotion information.

[0102] (Embodiment 3) The third embodiment relates to a form in which both the user's emotion information and the operator's emotion information are output. Note that a description of the content that overlaps with the first and second embodiments will be omitted.

[0103] 16 is an explanatory diagram showing an example of a record layout of the conversation data DB 175 in embodiment 3. The conversation data DB 175 includes a conversation data ID column, a user column, and an operator column. The conversation data ID column stores a conversation data ID that identifies conversation data. The user column includes a user ID column, a conversation data column, an emotion column, and a conversation style column. Note that the user ID column, conversation data column, emotion column, and conversation style column are the same as those in FIG. 4, so their description will be omitted.

[0104] The operator sequence includes an operator ID sequence, a conversation data sequence, an emotion sequence, and a conversation style sequence. The operator ID sequence stores an operator ID that identifies an operator. The conversation data sequence stores conversation data of the operator. The emotion sequence stores emotion information of the agent identified by the emotion identification model 171 based on the conversation data of the agent. The conversation style sequence stores style information of the agent identified by the style identification model 172 based on the conversation data of the agent.

[0105] Fig. 17 is an explanatory diagram showing an example of a display screen for emotion information and style information in embodiment 3. Description of content that overlaps with Fig. 9 will be omitted. This screen includes an operator emotion information display field 11d. The operator emotion information display field 11d is a display field that displays emotion information related to the operator's emotions.

[0106] The server 1 acquires both conversation data of a user who is conversing with an operator and conversation data of the operator. The server 1 inputs the acquired conversation data of the user to the emotion identification model 171 and outputs emotion information of the user. The server 1 inputs the acquired conversation data of the operator to the emotion identification model 171 and outputs emotion information of the operator. The server 1 transmits the emotion information of the user and the emotion information of the operator output from the emotion identification model 171 to the operator terminal 2.

[0107] The operator terminal 2 receives the user emotion information and the operator emotion information transmitted from the server 1. The operator terminal 2 displays the received user emotion information in chronological order in the user emotion information display field 11c, and displays the received operator emotion information in chronological order in the operator emotion information display field 11d. As shown in the figure, icons indicating the emotion classification of the user's conversation data at each point in time are displayed in the user emotion information display field 11c, and icons indicating the emotion classification of the operator's conversation data at each point in time are displayed in the operator emotion information display field 11d.

[0108] In this embodiment, emotion information including emotion classification is illustrated, but this is not limiting. For example, emotion information including emotion classification and level may be illustrated, similar to FIG.

[0109] 18 is a flowchart showing the processing steps when outputting both the user's emotion information and the operator's emotion information. The control unit 11 of the server 1 acquires conversation data of a user who is conversing with an operator in real time from the receiving terminal 3 used by the user via the communication unit 13 (step S121). The control unit 11 identifies the user's emotion information using the acquired user's conversation data (step S122). Specifically, the control unit 11 inputs the acquired user's conversation data into the emotion identification model 171, and outputs the identification result of the user's emotion as emotion information.

[0110] The control unit 11 of the server 1 acquires conversation data of the operator who is conversing with the user in real time from the receiving terminal 3 used by the operator via the communication unit 13 (step S123). The control unit 11 identifies emotion information of the operator using the acquired conversation data of the operator (step S124). Specifically, the control unit 11 inputs the acquired conversation data of the operator to the emotion identification model 171, and outputs the identification result of the emotion of the operator as emotion information.

[0111] The control unit 11 transmits the identified user emotion information and operator emotion information to the operator terminal 2 via the communication unit 13 (step S125). The control unit 21 of the operator terminal 2 receives the user emotion information and operator emotion information transmitted from the server 1 via the communication unit 23 (step S221). The control unit 21 displays the received user emotion information and operator emotion information on the display unit 25 (step S222).

[0112] The control unit 11 determines whether the dialogue between the user and the operator has ended through the receiving terminal 3 used by the user or the receiving terminal 3 used by the operator (step S126). If the control unit 11 determines that the dialogue has not ended (NO in step S126), the process returns to step S121.

[0113] If the control unit 11 determines that the dialogue has ended (YES in step S126), the control unit 11 stores the identified user emotion information and operator emotion information in the conversation data DB 175 of the mass storage unit 17 (step S127). Specifically, the control unit 11 assigns a conversation data ID, and stores the user ID, user conversation data, user emotion information, operator ID, operator conversation data, and operator emotion information as one record in association with the assigned conversation data ID in the conversation data DB 175. The control unit 11 then ends the process.

[0114] Furthermore, it is possible to output both the user's style information (for example, a slow speaking type or a brisk speaking type) and the operator's style information. Specifically, the server 1 acquires both conversation data of a user who is conversing with an operator and conversation data of the operator. The server 1 inputs the acquired user's conversation data into the style identification model 172 and outputs the user's style information. The server 1 inputs the acquired operator's conversation data into the style identification model 172 and outputs the operator's style information. The server 1 transmits the output user's style information and operator's style information to the operator terminal 2.

[0115] Furthermore, the server 1 acquires emotion information and style information of the user and emotion information and style information of the operator by using the emotion identification model 171 and the style identification model 172. The server 1 may transmit the acquired emotion information and style information of the user and emotion information and style information of the operator to the operator terminal 2.

[0116] According to this embodiment, it is possible to output both the user's emotion information and the operator's emotion information.

[0117] The embodiments disclosed herein are to be considered in all respects as illustrative and not restrictive. The scope of the present invention is defined by the claims, not by the above meaning, and is intended to include all modifications within the meaning and scope of the claims. [Explanation of symbols]

[0118] 1. Information processing device (server) 11 Control section 12 Storage section 13 Communications Department 14 Input section 15 Display 16 Reading unit 17 Mass storage 171 Emotion Identification Model (First Learning Model) 172 Style Discrimination Model (Second Learning Model) 173 Training Data DB 174 User DB 175 Conversation Data DB 176 Advice DB 1a Portable storage media 1b semiconductor memory 1P control program 2. Information processing terminal (operator terminal) 21 Control section 22 Memory section 23 Communications Department 24 Input section 25 Display section 2P control program 3. Receiving terminal

Claims

1. Acquire conversation data of a user who is interacting with an operator; inputting the acquired conversation data into a first learning model that has been trained to output emotion information relating to a user's emotions when conversation data is input, and outputting emotion information of the user that is classified into a plurality of levels for each of categories of calm, pleasure, discomfort, and depression; inputting the acquired conversation data into a second learning model that has been trained to output style information regarding a user's conversation style, including a slow speaking type and a brisk speaking type, when conversation data is input, and outputting the user's style information; outputting the emotion information including the emotion classification and level and the style information to an operation terminal of the operator; The icons indicating the emotion classifications and levels are displayed in a chronological order. A program that causes a computer to perform a process.

2. The first learning model classifies emotional information into at least two of neutral, pleasant, unpleasant, and depressed. The program according to claim 1.

3. The second learning model classifies the speech into style information including a slow speaking type and a quick speaking type. The program according to claim 1 or 2.

4. Identifying advice according to emotion information output from the first learning model and style information output from the second learning model; The specified advice is output to the operation terminal of the operator.

4. The program according to claim 1.

5. Acquire conversation data of the operator; The acquired conversation data is input into the first learning model, and emotion information of the operator is output.

5. The program according to claim 1.

6. The emotional information of the user, the conversation style of the user, and the emotional information of the operator are displayed in association with each other in a time series.

6. The program according to claim 1.

7. An information processing device including a control unit, The control unit Acquire conversation data of a user who is interacting with an operator; inputting the acquired conversation data into a first learning model that has been trained to output emotion information relating to a user's emotions when conversation data is input, and outputting emotion information of the user that is classified into a plurality of levels for each of categories of calm, pleasure, discomfort, and depression; inputting the acquired conversation data into a second learning model that has been trained to output style information regarding a user's conversation style, including a slow speaking type and a brisk speaking type, when conversation data is input, and outputting the user's style information; outputting the emotion information including the emotion classification and level and the style information to an operation terminal of the operator; The icons indicating the emotion classifications and levels are displayed in a chronological order. Information processing device.

8. Acquire conversation data of a user who is interacting with an operator; inputting the acquired conversation data into a first learning model that has been trained to output emotion information relating to a user's emotions when conversation data is input, and outputting emotion information of the user that is classified into a plurality of levels for each of categories of calm, pleasure, discomfort, and depression; inputting the acquired conversation data into a second learning model that has been trained to output style information regarding a user's conversation style, including a slow speaking type and a brisk speaking type, when conversation data is input, and outputting the user's style information; outputting the emotion information including the emotion classification and level and the style information to an operation terminal of the operator; The icons indicating the emotion classifications and levels are displayed in a chronological order. An information processing method that causes a computer to execute a process.

Citation Information

Patent Citations

  • Information processing apparatus, information processing method, video data, program, and information processing system

    JP2019029984A

  • Information processing device, information processing method, and information processing program

    JP2021012303A

  • Information processor, information processing method and program

    JP2021124530A

  • Information processor, information processing method, information processing program, and storage medium

    JP2021162627A

  • Information processor, information processing method, information processing program, and storage medium

    JP2021162628A