Conversation analysis device

The conversation analysis apparatus uses emotion estimation from tone and text to accurately assess conversation sincerity and empathy, addressing misjudgment in user-robot interactions.

JP7713415B2Active Publication Date: 2025-07-25NTT DOCOMO INC
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
JP2022042451
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-03-17
Publication Date
2025-07-25
Estimated Expiration
2042-03-17

AI Technical Summary

Technical Problem

Existing conversation analysis techniques fail to accurately determine the establishment of a conversation between a user and a robot, particularly when the user is not sincerely responding, leading to misjudgment.

Method used

A conversation analysis apparatus that includes a voice control unit, a first estimation unit, and a first determination unit to estimate emotions from tone color and text of user responses, determining conversation establishment based on the agreement between these emotions.

Benefits of technology

Reduces misjudgment of conversation establishment by assessing the sincerity of user responses, allowing for accurate evaluation of empathy and growth factors in user interactions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007713415000001
    Figure 0007713415000001
  • Figure 0007713415000002
    Figure 0007713415000002
  • Figure 0007713415000003
    Figure 0007713415000003
Patent Text Reader

Abstract

To prevent erroneous determination when a user does not give a response in earnest in a conversation between the user and a robot.SOLUTION: A conversation analysis device 40 comprises a voice control unit 412, a first estimation unit 414, a second estimation unit 415, and a first determination unit 416. The voice control unit 412 acquires voice data Dv1 indicating a voice of a user's response in a conversation between a robot and a user. The first estimation unit 414 generates a first feeling vector V1 representing a feeling expressed in a speech tone of the user's response on the basis of the voice data Dv1. The second estimation unit 415 generates a second feeling vector V2 representing a feeling expressed in words of the user's response on the basis of the voice data Dv1. The first determination unit 416 determines whether or not a conversation is established between the robot and the user on the basis of the degree of matching indicating the degree of matching between the feeling represented by the first feeling vector V1 and the feeling represented by the second feeling vector V2.SELECTED DRAWING: Figure 8
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a conversation analysis device.

Background Art

[0002] Patent Document 1 discloses a technique for determining whether a conversation is established based on the semantic content of an utterance. Patent Document 2 discloses a technique for calculating a conversation establishment degree indicating the degree to which a conversation is established based on the ratio of a section in which the conversation voice of one speaker is voiced and the conversation voice of the other speaker is unvoiced, and the ratio of a section in which both are voiced or unvoiced.

Prior Art Documents

Non-Patent Documents

[0003]

Patent Document 1

Patent Document 2

Summary of the Invention

Problems to be Solved by the Invention

[0004] In recent years, various robots that converse with users by voice have been proposed. When a user replies playfully or gives an appropriate response to a query from a robot, for example, when the user is not sincerely responding, it cannot be said that the conversation between the robot and the user is established. However, in a technique for determining the success or failure of a conversation based on the semantic content of a user's utterance or the ratio of sections that are voiced or unvoiced in each of the robot's conversation voice and the user's conversation voice, if the user is not sincerely responding, there may be a case where it is erroneously determined that the conversation is established.

Means for Solving the Problems

[0005] In order to solve the above problems, a conversation analysis apparatus according to a preferred embodiment of the present disclosure includes a voice control unit, a first estimation unit, a second estimation unit, and a first determination unit. The voice control unit acquires voice data representing the voice of the user's response in the conversation between the robot and the user. The first estimation unit estimates the emotion contained in the response based on the tone color of the voice represented by the voice data. The second estimation unit estimates the emotion contained in the response based on the text represented by the voice data. The first determination unit determines whether the conversation is established based on a degree of agreement indicating the degree of agreement between the emotion estimated by the first estimation unit and the emotion estimated by the second estimation unit.

Effect of the Invention

[0006] According to the conversation analysis apparatus of the present disclosure, even when the user is not sincerely responding in the conversation with the robot, the misjudgment that the conversation is established is reduced.

Brief Description of the Drawings

[0007]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Embodiments for Carrying Out the Invention

[0008] (A. Embodiment) (A-1: Overall Configuration) FIG. 1 is a block diagram showing a configuration example of the conversation system 1 according to an embodiment of the present disclosure. As shown in FIG. 1, the conversation system 1 includes a conversation analysis device 40, a user device 10, and a robot 30. The conversation partner U1 is a user who uses the robot 30. The robot 30 conducts a conversation with the conversation partner U1 under the control of the conversation analysis device 40. A typical example of the conversation partner U1 is a child such as an elementary school student. The robot 30 is given to the conversation partner U1 by the guardian U2 of the conversation partner U1. The robot 30 is arranged, for example, in the study room of the conversation partner U1. The user device 10 communicates with the conversation analysis device 40 via the communication network NW. The user device 10 is used by the guardian U2.

[0009] FIG. 2 is a front view showing an example of the appearance of the robot 30. The robot 30 has a head 31 and a body 32. The robot 30 operates on the power of a battery. A speaker 370 and a human sensor 350 are arranged on the body 32. Inside the head 31, a microphone 360 (not shown) is arranged. The robot 30 uses the human sensor 350 to detect that the conversation partner U1 has approached. The robot 30 switches the operation mode based on the detection result by the human sensor 350. The operation modes of the robot 30 include a conversation mode for conversing with the conversation partner U1 and a sleep mode for not conversing with the conversation partner U1. The power consumption of the robot 30 in the sleep mode is smaller than the power consumption of the robot 30 in the conversation mode. In the sleep mode, by supplying power to the human sensor 350, it is detected that the conversation partner U1 has approached the robot 30, while the supply of power to other components is restricted. By having the conversation mode and the sleep mode, the robot 30 can converse when the conversation partner U1 is nearby, and can save power consumption when the conversation partner U1 is not nearby.

[0010] In the present embodiment, the robot 30 is used for the sentiment education of the conversation partner U1 through conversation. In the present embodiment, by having the conversation analysis device 40 analyze the voice of the response of the conversation partner U1 in the conversation with the robot 30, the growth factors of the conversation partner U1 are grasped. The growth factors of the conversation partner U1 refer to things with low empathy of the conversation partner U1 and with prospects for future growth.

[0011] (A-2: User device) FIG. 3 is a block diagram showing a configuration example of the user device 10. Specific examples of the user device 10 include a smartphone or a tablet terminal. The user device 10 includes a processing device 110, a storage device 120, a display panel 130, a communication device 140, and an input device 150. Each element of the user device 10 is interconnected by one or more buses for communicating information. Note that the term "device" in this specification may be read as other terms such as a circuit, a device, or a unit. Each element of the user device 10 may be composed of one or more devices. Some elements of the user device 10 may be omitted.

[0012] The processing device 110 is a processor that controls the entire user device 10. The processing device 110 is composed of, for example, one or more chips. The processing device 110 is composed of, for example, a central processing unit (CPU) including an interface with peripheral devices, an arithmetic unit, and registers. Note that some or all of the functions of the processing device 110 may be realized by hardware such as a DSP (Digital Signal Processor), an ASIC (Application Specific Integrated Circuit), a PLD (Programmable Logic Device), or an FPGA (Field Programmable Gate Array). The processing device 110 executes various processes in parallel or sequentially. The processing device 110 functions as a control center of the user device 10 by operating according to the program PR1 stored in the storage device 120.

[0013] The memory device 120 is a recording medium readable by the processing device 110. The memory device 120 stores a plurality of programs including the program PR1, and various information used by the processing device 110. The memory device 120 may be constituted by at least one of, for example, ROM (Read Only Memory), EPROM (Erasable Programmable ROM), EEPROM (Electrically Erasable Programmable ROM), RAM (Random Access Memory), etc. The program PR1 may be transmitted from another device such as the conversation analysis device 40 via the communication network NW. Specific examples of the information stored in the memory device 120 and used by the processing device 110 include first identification information uniquely indicating the robot 30 in the communication network NW, and second identification information uniquely indicating the conversation analysis device 40 in the communication network NW. Specific examples of the first identification information include the communication address assigned to the robot 30. Similarly, specific examples of the second identification information include the communication address assigned to the conversation analysis device 40.

[0014] The display panel 130 is a device that displays images. Specific examples of the display panel 130 include a liquid crystal display panel and an organic EL (Electro Luminescence) display panel, etc.

[0015] The communication device 140 is hardware (a transmission / reception device) for communicating with other devices. The communication device 140 is also called, for example, a network device, a network controller, a network card, a communication module, etc.

[0016] The input device 150 is a device that receives external input. As a specific example of the input device 150, a touch panel provided integrally with the display panel 130 can be mentioned. The input device 150 may include a plurality of operators operable by the user. The input device 150 outputs input information corresponding to the operation of the user, that is, the caregiver U2, to the processing device 110. In the present embodiment, the caregiver U2 can specify a thing to be determined as a growth factor for the conversation partner U1 by operating the input device 150. The processing device 110 transmits information indicating the thing specified by the caregiver U2 and the first identification information to the device indicated by the second identification information, that is, the conversation analysis device 40, using the communication device 140. Hereinafter, the information indicating the thing specified by the caregiver U2 is referred to as determination target information.

[0017] (A-3: Robot) FIG. 4 is a block diagram showing a configuration example of the robot 30. The robot 30 includes a processing device 310, a storage device 320, a motor 330, a communication device 340, a human sensor 350, a microphone 360, and a speaker 370. Each element of the robot 30 is interconnected by a single or a plurality of buses for communicating information.

[0018] The processing device 310 is a processor that controls the entire robot 30. The processing device 310 is composed of one or more chips. The processing device 310 functions as a control center of the robot 30 by operating according to the program PR3 stored in the storage device 320.

[0019] The storage device 320 is a recording medium readable by the processing device 310. The storage device 320 stores a plurality of programs including the program PR3 and various information used by the processing device 310. The storage device 320 may be composed of at least one of, for example, ROM, EPROM, EEPROM, and RAM.

[0020] The motor 330 operates under the control of the processing device 310. By driving the motor 330, the head 31 or the arm of the robot 30 shown in FIG. 2 moves. The communication device 340 is hardware (a transmission / reception device) for communicating with other devices. The communication device 340 communicates with the conversation analysis device 40. The human sensor 350 detects that a person has approached and outputs a detection signal indicating the detection result to the processing device 310. When a detection signal is input from the human sensor 350, the processing device 310 switches the operation mode to the conversation mode and transmits a start signal indicating the start of operation in the conversation mode to the conversation analysis device 40. After switching the operation mode to the conversation mode, if no detection signal is input from the human sensor 350 for a predetermined time, the processing device 310 switches the operation mode from the conversation mode to the sleep mode. Also, when the processing device 310 switches the operation mode from the conversation mode to the sleep mode, it transmits an end signal indicating the end of operation in the conversation mode to the conversation analysis device 40.

[0021] The microphone 360 generates an audio signal indicating the voice of the conversation partner U1 by converting sound into an electrical signal. The audio signal generated by the microphone 360 is converted into audio data by an AD converter (not shown). This audio data is output to the processing device 310. The processing device 310 transmits the audio data to the conversation analysis device 40 via the communication device 340.

[0022] The speaker 370 outputs the voice of the robot 30 to the conversation partner U1 under the control of the processing device 310. The speaker 370 includes a DA converter. Although details will be described later, in this embodiment, the audio data representing the voice of the robot in the conversation with the conversation partner U1 is generated by the conversation analysis device 40. The audio data generated by the conversation analysis device 40 is transmitted to the robot 30 via the communication network NW. The processing device 310 outputs the audio data received via the communication network NW to the speaker 370. The speaker 370 uses the DA converter to convert the audio data output from the processing device 310 into an audio signal. The speaker 370 is driven by this audio signal.

[0023] (A-4: Conversation Analysis Device) FIG. 5 is a block diagram showing a configuration example of the conversation analysis device 40. The conversation analysis device 40 is, for example, a server. The conversation analysis device 40 includes a processing device 410, a storage device 420, a display panel 430, a communication device 440, and an input device 450. Each element of the conversation analysis device 40 is interconnected by one or more buses for communicating information. The hardware configuration of the conversation analysis device 40 is the same as that of the user device 10. However, the processing speed of the processing device 410 is preferably higher than the processing speed of the processing device 110. Also, the storage capacity of the storage device 420 is preferably larger than the storage capacity of the storage device 120.

[0024] The processing device 410 functions as a control center of the conversation analysis device 40 by operating according to the program PR4. The storage device 420 stores a plurality of programs including the program PR4, a topic table TBL1, a word table TBL2, and various information used by the processing device 410. Specific examples of the various information used by the processing device 410 include a pair of the first identification information received from the user device 10 and the determination target information.

[0025] FIG. 6 is a diagram showing an example of the topic table TBL1. As shown in FIG. 6, a plurality of combinations of theme data and story data are stored in the topic table. The theme data is text data representing things to be judged as growth factors. As a specific example of the theme data, as shown in FIG. 6, text data representing character strings such as "kindness" and "etiquette" can be cited. The story data is text data representing a story related to the thing represented by the theme data associated with the story data. As shown in FIG. 6, in the topic table TBL1 in the present embodiment, story data A is associated with the theme data indicating "kindness". Also, in the present embodiment, story data B is associated with the theme data indicating "etiquette". The topic table TBL1 shown in FIG. 6 stores two types of theme data, "kindness" and "etiquette", but instead of or in addition to "kindness" and "etiquette", theme data indicating "friendship", "courage", etc. and story data corresponding to the theme data may be stored in the topic table TBL1.

[0026] The story data A in the present embodiment represents, for example, a story about kindness such as "A person walking in front dropped something. When I picked it up...".

[0027] Also, the story data B in the present embodiment represents, for example, a story about etiquette such as "Since the person greeted me loudly with 'Hello', when I replied 'Hello'...".

[0028] FIG. 7 is a diagram showing an example of the word table TBL2. As shown in FIG. 7, a plurality of pairs of word data and emotion data are stored in the word table TBL2. The word data is text data representing a word that is assumed to be used in the response of the conversation partner U1 in the conversation with the robot 30. For example, the word table TBL2 shown in FIG. 7 stores word data representing the words "fun", "kind", "boring", and "sad". The emotion data is text data representing an emotion that is assumed to be included in the response including the word represented by the word data associated with the emotion data. In the present embodiment, emotion data indicating "joy" is associated with the word data "fun". Also, in the present embodiment, emotion data indicating "slight joy" is associated with the word data "kind". In the present embodiment, emotion data indicating "anger" is associated with the word data "boring". In the present embodiment, emotion data indicating "sadness" is associated with the word data "sad". Although details will be described later, the word table TBL2 is used when estimating the emotion included in the words of the response of the conversation partner U1 in the conversation with the robot 30.

[0029] FIG. 8 is a block diagram showing the functions realized by the processing device 410 according to the program PR4. The processing device 410 functions as a presentation unit 411, an audio control unit 412, a second determination unit 413, a first estimation unit 414, a second estimation unit 415, a first determination unit 416, and a third estimation unit 417 by reading and executing the program PR4 from the storage device 420.

[0030] The prompting unit 411 prompts the conversation partner U1 with a story related to a thing to be determined whether it is a growth factor prior to the start of the conversation between the robot 30 and the conversation partner U1. When the prompting unit 411 receives the start signal SS from the robot 30 via the communication network NW, it reads out the determination target information stored in the storage device 420 in association with the first identification information corresponding to the transmission source of the start signal SS. Next, the prompting unit 411 reads out the story data stored in the topic table TBL1 in association with the theme data indicating the same thing as the read determination target information. For example, when the thing to be determined is "kindness", the prompting unit 411 reads out the story data A from the topic table TBL1.

[0031] Next, the prompting unit 411 synthesizes the voice data DS representing the voice of "Can you listen to me again today?" and the voice of reading out the story represented by the read story data. The prompting unit 411 transmits the synthesized voice data DS to the robot 30. FIG. 9 is a diagram showing an example of the conversation between the robot 30 and the conversation partner U1. The robot 30 that has received the voice data DS outputs the voice Q11 of "Can you listen to me again today?" as shown in FIG. 9, and then outputs the voice Q12 of "The person walking in front...". By outputting the voice Q12 of reading out the story represented by the story data A from the robot 30, the story represented by the story data A is presented to the conversation partner U1.

[0032] After the presentation unit 411 presents the story, the voice control unit 412 causes the robot 30 to carry out a conversation with the conversation partner U1 until it receives an end signal SE from the robot 30. The voice control unit 412 generates voice data Dv2 representing the voice of the query from the robot 30 in the conversation with the conversation partner U1. The voice control unit 412 transmits the generated voice data Dv2 to the robot 30 via the communication network NW. The robot 30 that has received the voice data Dv2 outputs the voice represented by the voice data Dv2. For example, when the voice data Dv2 representing the voice Q13 of "What do you think of this story?" is transmitted from the conversation analysis device 40 to the robot 30, the robot 30 outputs the voice Q13 as shown in FIG. 9. Note that the voice Q13 may be the voice of "Did you feel warm?" Also, the voice control unit 412 acquires from the robot 30 the voice data Dv1 representing the voice of the response of the conversation partner U1 to the query from the robot 30. For example, when a response A11 of "I thought it was kind." is made as shown in FIG. 9 in response to the query by the voice Q13, the voice control unit 412 acquires from the robot 30 the voice data Dv1 representing the response A11.

[0033] When the voice data Dv1 representing the response is not acquired by the voice control unit 412 until a predetermined time such as 30 seconds has elapsed since the voice data Dv2 was transmitted to the robot 30, the second determination unit 413 determines that the conversation has not been established. This is because not responding to the query indicates that the conversation partner U1 does not have the will to talk to the robot 30.

[0034] Even if the voice control unit 412 acquires the voice data Dv1 representing a response before a predetermined time elapses since the voice data Dv2 was transmitted to the robot 30, the second determination unit 413 may determine that the conversation has not been established if the response represented by the voice data Dv1 is a response that diverts the topic. Specific examples of responses that divert the topic include responses that explicitly prompt a change of topic, such as "Let's talk about something else rather than that," or responses regarding things unrelated to the story presented by the presentation unit 411, such as "Isn't ○○ interesting?" (○○ is something unrelated to the story presented by the presentation unit 411). Diverting the topic is an indication that the conversation partner U1 does not have the will to talk to the robot 30 regarding the thing that is the subject of the determination as to whether it is a growth factor or not.

[0035] The first estimation unit 414 estimates the emotion incorporated in the response represented by the voice data Dv1 acquired by the voice control unit 412 based on the tone color represented by the voice data Dv1. Tone color refers to the time change in pitch in the voice, the time change in the intensity of the sound in the voice, and the time change in the length of the sound in the voice. The first estimation unit 414 generates tone color data representing the tone color of the voice represented by the voice data Dv1 by analyzing the voice data Dv1 acquired by the voice control unit 412. Specific examples of tone color data include data representing some or all of 12-dimensional MFCC (Mel-Frequency Cepstrum Coefficients), loudness, fundamental frequency (F0), voice probability, zero-crossing rate, HNR (Harmonics-to-Noise-Ratio), and the first derivatives of these, and the second derivatives of MFCC and loudness.

[0036] Next, the first estimation unit 414 generates a first emotion vector V1 having as components the evaluation values of each of the emotions of "joy", "anger", "sorrow", and "happiness" regarding the emotion incorporated in the response voice represented by the voice data Dv1 based on the timbre data generated from the voice data Dv1. It is common for the emotion of the person who uttered the voice to be incorporated in the voice. The first emotion vector V1 is a four-dimensional vector representing the intensity of each of the emotions of "joy", "anger", "sorrow", and "happiness" incorporated in the response voice represented by the voice data Dv1. As for the algorithm for estimating the emotion incorporated in the voice based on the timbre, an existing timbre analysis algorithm may be appropriately used.

[0037] The second estimation unit 415 estimates the emotion incorporated in the response based on the text of the response represented by the voice data Dv1. The second estimation unit 415 performs voice recognition on the voice data Dv1 to generate text data representing the sentence of the response represented by the voice data Dv1. Next, the second estimation unit 415 performs morphological analysis or the like on the sentence represented by the text data generated from the voice data Dv1 to extract the words constituting the document. Then, the second estimation unit 415 specifies the corresponding emotion for each word extracted from the sentence represented by the text data by referring to the word table TBL2. Then, the second estimation unit 415 generates a second emotion vector V2 representing the emotion incorporated in the response represented by the voice data Dv1 based on the emotion specified for each word extracted from the sentence represented by the text data.

[0038] It is common for the person who utters the voice to select the words constituting the voice according to the emotion at that time. The second emotion vector V2 is a four-dimensional vector representing the intensity of each of the emotions of "joy", "anger", "sorrow", and "happiness" incorporated in the text of the response represented by the voice data Dv1. As for the algorithm for estimating the emotion from the text of the voice, an existing emotion analysis algorithm may be appropriately used.

[0039] The first determination unit 416 determines whether the conversation between the robot 30 and the conversation partner U1 is established based on the degree of coincidence indicating the degree of coincidence between the emotion estimated from the tone color of the response indicated by the voice data Dv1 and the emotion estimated from the text of the response indicated by the voice data Dv1. The reason why it is possible to determine whether the conversation between the robot 30 and the conversation partner U1 is established based on the emotion estimated from the tone color of the response of the conversation partner U1 and the emotion estimated from the text of the response of the conversation partner U1 is as follows.

[0040] If the conversation partner U1 answers sincerely to the query from the robot 30, that is, if the conversation between the robot 30 and the conversation partner U1 is established, there should be no discrepancy between the emotion estimated from the tone color of the response and the emotion estimated from the text of the response, that is, the degree of coincidence between the two should be high. For example, in response to the query by the voice Q13, as shown in FIG. 9, if a response A11 of "I thought it was kind." is returned and the emotion estimated from the tone color of the response A11 is "joy". In this case, it is considered that the conversation partner U1 is truly answering the query by the voice Q13 with the emotion of "joy" and the conversation between the robot 30 and the conversation partner U1 is established.

[0041] On the other hand, when the degree of agreement between the emotion estimated from the tone of the response and the emotion estimated from the text of the response is not high, it is considered that the conversation partner U1 is answering the question from the robot 30 in a joking manner or giving an appropriate comeback. For example, in response to the question by voice Q13, as shown in FIG. 9, a response A11 of "I thought it was kind" is given, and it is assumed that the emotion estimated from the tone of the response A11 is "indifferent" or "emotionless". In this case, it is considered that the conversation partner U1 is giving an appropriate comeback to the question from the robot 30. Also, when a sad story is presented by the presentation unit 411 and, despite the response being "sad", the emotion of "happiness" is incorporated in the tone of the response, it is considered that the conversation partner U1 is answering in a joking manner. When the conversation partner U1 is answering the question from the robot 30 in a joking manner or giving an appropriate comeback, it cannot be said that the conversation between the robot 30 and the conversation partner U1 is established. The above is the reason why it is possible to determine whether the conversation between the robot 30 and the conversation partner U1 is established based on the degree of agreement between the emotion estimated from the tone of the response of the conversation partner U1 and the emotion estimated from the text of the response of the conversation partner U1.

[0042] In this embodiment, the first determination unit 416 calculates an angle θ (0 ≤ θ ≤ 180°) formed by the first emotion vector V1 and the second emotion vector V2 as a value indicating the degree of coincidence between the emotion estimated from the tone of voice of the response and the emotion estimated from the text of the response. The angle θ is obtained as the inverse cosine of the value obtained by dividing the inner product of the first emotion vector V1 and the second emotion vector V2 by the product of the norm of the first emotion vector V1 and the norm of the second emotion vector V2. Then, if the angle θ is less than a predetermined threshold (for example, 30°), the first determination unit 416 determines that the conversation between the robot 30 and the conversation partner U1 is established. Conversely, if the angle θ is greater than or equal to the above threshold, it is determined that the conversation between the robot 30 and the conversation partner U1 is not established. In this embodiment, the angle formed by the first emotion vector V1 and the second emotion vector V2 is used as a value indicating the degree of coincidence between the emotion estimated from the tone of voice of the response and the emotion estimated from the text of the response, but the cosine value of the angle may also be used. The cosine value of the angle formed by the first emotion vector V1 and the second emotion vector V2 is the value obtained by dividing the inner product of the first emotion vector V1 and the second emotion vector V2 by the product of the norm of the first emotion vector V1 and the norm of the second emotion vector V2. In the aspect of using the cosine value of the angle formed by the first emotion vector V1 and the second emotion vector V2 as a value indicating the degree of coincidence, the first determination unit 416 may determine that the conversation between the robot 30 and the conversation partner U1 is established when the value is greater than or equal to a predetermined threshold (for example, 0.6). Also, the distance between the end point of the first emotion vector V1 and the end point of the second emotion vector V2 may be used as a value indicating the degree of coincidence between the emotion estimated from the tone of voice of the response and the emotion estimated from the text of the response.

[0043] When the first determination unit 416 determines that a conversation is established, the third estimation unit 417 estimates a degree of empathy of the conversation partner U1 with respect to the thing indicated by the determination target information based on the first emotion vector V1 and the second emotion vector V2. In the present embodiment, the third estimation unit 417 calculates the degree of empathy by weighted addition of the norm of the first emotion vector V1 and the norm of the second emotion vector V2, but the inner product of the first emotion vector V1 and the second emotion vector V2 may be used as the degree of empathy. Note that the third estimation unit 417 determines whether the emotion of the conversation partner U1 based on the first emotion vector V1 and the second emotion vector V2 is a positive emotion or a negative emotion with respect to the thing indicated by the determination target thing, and calculates the degree of empathy according to the degree of affirmation.

[0044] Then, when the calculated degree of empathy is less than a predetermined threshold, the third estimation unit 417 stores the determination target information in the storage device 420 as growth factor information indicating a growth factor in which further growth of the conversation partner U1 can be expected. Note that when the calculated degree of empathy is equal to or greater than a predetermined threshold, the third estimation unit 417 may store the determination target information in the storage device 420 as strength information indicating the strengths of the conversation partner U1. Further, the third estimation unit 417 may store the degree of empathy in the storage device 420 in association with the growth factor information or the strength information.

[0045] Also, the processing device 410 executes the conversation analysis method of the present disclosure by executing the program PR4. FIG. 10 is a flowchart showing the flow of the conversation analysis method executed by the processing device 410 according to the program PR4. As shown in FIG. 10, this conversation analysis method includes each process from step SA110 to step SA150.

[0046] In step SA110, the processing device 410 functions as the presentation unit 411. In step SA110, the processing device 410 causes the robot 30 to output the voice of a story about a thing to be determined whether it is a growth factor, triggered by the reception of the start signal SS.

[0047] In step SA120 that follows step SA110, the processing device 410 functions as the voice control unit 412. In step SA120, the processing device 410 determines whether it has received the end signal SE from the robot 30. If the end signal SE is received from the robot 30, the determination result of step SA120 is "Yes". If the determination result of step SA120 is "Yes", the processing device 410 ends the execution of this conversation analysis method. Conversely, if the determination result of step SA120 is "No", the processing device 410 executes the processing after step SA130.

[0048] In step SA130, the processing device 410 functions as the voice control unit 412. In step SA130, the processing device 410 causes the robot 30 to execute a conversation with the conversation partner U1. That is, in step SA130, the processing device 410 causes the robot 30 to output the voice of the query to the conversation partner U1, while acquiring the voice data representing the response voice of the conversation partner U1 from the robot 30.

[0049] In step SA140 that follows step SA130, the processing device 410 functions as the second determination unit 413, the first estimation unit 414, the second estimation unit 415, and the first determination unit 416. In step SA140, the processing device 410 first functions as the second determination unit 413. If the voice data Dv1 representing the response is not acquired until the elapse of a predetermined time from the time when the voice data Dv2 representing the query voice is transmitted to the robot 30, the processing device 410 determines that the conversation between the robot 30 and the conversation partner U1 has not been established. If the voice data Dv1 representing the response is not acquired until the elapse of a predetermined time from the time when the voice data Dv2 representing the query voice is transmitted to the robot 30, the determination result of step SA130 is "No".

[0050] When the voice data Dv1 representing the response is acquired before the elapse of a predetermined time from the time when the voice data Dv2 representing the inquiry voice is transmitted to the robot 30, the processing device 410 functions as the first estimation unit 414, the second estimation unit 415, and the first determination unit 416. Based on the acquired voice data Dv1, the processing device 410 calculates the above-described first emotion vector V1 and second emotion vector V2. Then, the processing device 410 determines whether or not a conversation between the robot 30 and the conversation partner U1 is established based on the degree of coincidence between the emotion represented by the first emotion vector V1 and the emotion represented by the second emotion vector V2. When it is determined that the conversation between the robot 30 and the conversation partner U1 is established, the determination result of step SA140 becomes "Yes". Conversely, when it is determined that the conversation between the robot 30 and the conversation partner U1 is not established, the determination result of step SA140 becomes "No".

[0051] When the determination result of step SA140 is "Yes", the processing device 410 executes the process of step SA150 and then executes the process of step SA110 again. On the other hand, when the determination result of step SA140 is "No", the processing device 410 executes the process of step SA110 again without executing the process of step SA150.

[0052] In step SA150, the processing device 410 functions as the third estimation unit 417. In step SA150, the processing device 410 calculates a degree of empathy indicating the strength of the conversation partner U1's empathy for the thing indicated by the determination target information based on the first emotion vector V1 and the second emotion vector V2. Then, when the calculated degree of empathy is less than a predetermined threshold value, the processing device 410 stores the determination target information in the storage device 420 as growth element information.

[0053] As described above, since the conversation analysis device 40 determines whether a conversation with the robot 30 is established based on the degree of agreement between the emotion estimated from the tone of the response of the conversation partner U1 in the conversation with the robot 30 and the emotion estimated from the text of the response, it is possible to reduce the misjudgment that the conversation is established when the conversation partner U1 is not sincerely responding. In addition, since the conversation analysis device 40 can estimate the strength of the empathy of the conversation partner U1 for the thing indicated by the determination target information based on the response when it is determined that the conversation with the robot 30 is established, misjudgment of growth factors can also be reduced.

[0054] (B: Modification example) The present disclosure is not limited to the embodiments illustrated above. Specific modification modes are as follows. Two or more modes arbitrarily selected from the following examples may be combined. (B-1: Modification example 1) When the empathy of the conversation partner U1 is less than the threshold for any of a plurality of types of things, the processing device 410 may determine that the conversation partner U1 is poor at emotional expression. Further, the processing device 410 may generate a vocabulary database by associating the vocabulary used by the conversation partner U1 in the response with the emotion estimated from the tone of the response and writing it into the storage device 420. For example, assume that the vocabulary "terrible" is associated with both "happiness" and "fear" in the vocabulary database. In this way, when the same vocabulary is used for a plurality of types of emotions, it is estimated that the conversation partner U1 has a shortage of vocabulary (or lack of expressiveness). In the mode of generating the vocabulary database, the processing device 410 may estimate the presence or absence of a shortage of vocabulary (or lack of expressiveness) for the conversation partner U1 from the stored content of the vocabulary database. Then, when it is estimated that the conversation partner U1 has a shortage of vocabulary (or lack of expressiveness), the processing device 410 may refer to the stored content of the word table TBL2 and propose a paraphrase using other words to the conversation partner U1. For example, when expressing "happiness", it may be proposed to use "fun" or "excited" instead of "terrible".

[0055] Also, in the aspect of generating the vocabulary database, the processing device 410 may store at least one of information indicating the age of the conversation partner U1, information indicating the region where the conversation partner U1 resides, and information indicating the growth status of the conversation partner U1 in the vocabulary database in association with the vocabulary used by the conversation partner U1 in the response and the emotion estimated from the response. In the vocabulary database generated by this aspect, the usage (or distribution) of vocabulary specific to the age, residential area, or growth status of the conversation partner U1 is stored in the database. By referring to this vocabulary database, when creating commercials or the like targeting the conversation partner U1 of a specific age or the conversation partner U1 living in a specific region, it becomes possible to select vocabulary that is likely to evoke empathy and create commercials or the like. Also, by using this vocabulary database, based on the vocabulary used by the conversation partner U1 in a conversation regarding a thing to be determined whether it is a growth factor, it is possible to estimate the age, residential area, or growth status of the conversation partner U1, that is, to profile the conversation partner U1.

[0056] (B-2: Variant Example 2) The presentation unit 411, the second determination unit 413, and the third estimation unit 417 are not essential and may be omitted. The main point is that the conversation analysis device of the present disclosure only needs to include the voice control unit 412, the first estimation unit 414, the second estimation unit 415, and the first determination unit 416. The voice control unit 412 acquires voice data Dv1 representing the voice of the response of the conversation partner U1 in the conversation between the robot 30 and the conversation partner U1. The first estimation unit 414 estimates the emotion incorporated in the response indicated by the voice data Dv1 based on the tone color of the voice represented by the voice data Dv1. The second estimation unit 415 estimates the emotion incorporated in the response indicated by the voice data Dv1 based on the text of the response represented by the voice data Dv1. The first determination unit 416 determines whether the conversation between the robot 30 and the conversation partner U1 is established based on the degree of agreement indicating the degree of agreement between the emotion estimated by the first estimation unit 414 and the emotion estimated by the second estimation unit 415. The conversation analysis device of this aspect can reduce the misjudgment of the establishment of the conversation when the conversation partner U1 is not sincerely responding.

[0057] (B-3: Variant Example 3) In the above-described embodiment, the conversation analysis device 40 and the robot 30 were separate devices, but the conversation analysis device 40 may be included in the robot 30. Further, the program PR4 in the above-described embodiment may be manufactured or sold alone.

[0058] (C: Others) (1) In the above-described embodiments, the storage devices 120, 320, and 420 were exemplified by ROM and RAM, etc., but flexible disks, magneto-optical disks (e.g., compact disks, digital versatile disks, Blu-ray (registered trademark) disks), smart cards, flash memory devices (e.g., cards, sticks, key drives), CD-ROM (Compact Disc-ROM), registers, removable disks, hard disks, floppy (registered trademark) disks, magnetic strips, databases, servers, and other appropriate storage media. Further, the program may be transmitted from a network via a telecommunication line. Also, the program may be transmitted from a communication network via a telecommunication line.

[0059] (2) In the above-described embodiments, the information, signals, etc. described may be represented using any of various different technologies. For example, the data, instructions, commands, information, signals, bits, symbols, chips, etc. that may be referred to throughout the above description may be represented by voltage, current, electromagnetic waves, magnetic fields or magnetic particles, optical fields or photons, or any combination thereof.

[0060] (3) In the above-described embodiments, the input / output information, etc. may be stored in a specific location (e.g., memory) or may be managed using a management table. The input / output information, etc. may be overwritten, updated, or appended. The output information, etc. may be deleted. The input information, etc. may be transmitted to other devices.

[0061] (4) In the above-described embodiments, the determination may be made based on a value represented by 1 bit (either 0 or 1), a boolean value (Boolean: true or false), or a numerical comparison (e.g., comparison with a predetermined value).

[0062] (5) The processing procedures, sequences, flowcharts, etc. exemplified in the above-described embodiments may be reordered as long as there is no contradiction. For example, regarding the methods described in the present disclosure, the elements of various steps are presented using an exemplary order and are not limited to the presented specific order.

[0063] (6) Each function exemplified in FIG. 8 is realized by any combination of at least one of hardware and software. Also, the method of realizing each functional block is not particularly limited. That is, each functional block may be realized using one physically or logically combined device, or two or more physically or logically separated devices may be directly or indirectly (e.g., using wired, wireless, etc.) connected and realized using these multiple devices. The functional block may be realized by combining software with the above one device or the above multiple devices.

[0064] (7) The programs exemplified in the above-described embodiments should be broadly interpreted to mean instructions, instruction sets, codes, code segments, program codes, programs, subprograms, software modules, applications, software applications, software packages, routines, subroutines, objects, executable files, execution threads, procedures, functions, etc., regardless of whether the software is called a hardware description language, firmware, middleware, microcode, or by another name.

[0065] Also, software, instructions, information, etc. may be transmitted and received via a transmission medium. For example, when software is transmitted from a website, server, or other remote source using at least one of wired technologies (such as coaxial cable, fiber optic cable, twisted pair, digital subscriber line (DSL), etc.) and wireless technologies (such as infrared, microwave, etc.), at least one of these wired and wireless technologies is included within the definition of the transmission medium.

[0066] (8) In each of the above-described embodiments, the terms "system" and "network" are used interchangeably.

[0067] (9) The information, parameters, etc. described in this disclosure may be represented using absolute values, relative values from a predetermined value, or corresponding other information.

[0068] (10) In the above-described embodiments, the user device 10 may include a case where it is a mobile station (MS). A mobile station may also be called by those skilled in the art as a subscriber station, mobile unit, subscriber unit, wireless unit, remote unit, mobile device, wireless device, wireless communication device, remote device, mobile subscriber station, access terminal, mobile terminal, wireless terminal, remote terminal, handset, user agent, mobile client, client, or some other appropriate term. Also, in this disclosure, terms such as "mobile station", "user terminal", "user equipment (UE)", "terminal", etc. may be used interchangeably.

[0069] (11) In the above-described embodiments, the terms "connected" and "coupled", or any variations thereof, mean any direct or indirect connection or coupling between two or more elements, and can include the presence of one or more intermediate elements between two elements "connected" or "coupled" to each other. The coupling or connection between elements can be physical, logical, or a combination thereof. For example, "connected" may be read as "accessed". As used in this disclosure, two elements can be considered to be "connected" or "coupled" to each other using at least one of one or more wires, cables, and printed electrical connections, as well as, by way of some non-limiting and non-exhaustive examples, electromagnetic energy having wavelengths in the radio frequency region, microwave region, and optical (both visible and invisible) region, etc.

[0070] (12) In the above-described embodiments, the description "based on" does not mean "based only on" unless otherwise specified. In other words, the description "based on" means both "based only on" and "based at least on".

[0071] (13) As used in this disclosure, the terms "determining" may encompass a wide variety of operations. "Determining" may include, for example, judging, calculating, computing, processing, deriving, investigating, looking up (e.g., searching in a table, database, or other data structure), ascertaining, and considering something as having been "determined". Also, "determining" may include receiving (e.g., receiving information), transmitting (e.g., transmitting information), inputting, outputting, accessing (e.g., accessing data in memory), and considering something as having been "determined". Further, "determining" may include resolving, selecting, choosing, establishing, comparing, etc., and considering something as having been "determined". That is, "determining" may include considering that some operation has been "determined". Also, "determining" may be replaced by "assuming", "expecting", "considering", etc.

[0072] (14) In the embodiments described above, when the terms "include", "including" and their variants are used, these terms are intended to be inclusive, similar to the term "comprising". Further, the term "or" as used in this disclosure is not intended to be an exclusive disjunction.

[0073] (15) In this disclosure, for example, when articles are added by translation, such as a, an, and the in English, this disclosure may include that the nouns following these articles are in the plural form.

[0074] (16) In the present disclosure, the term "A and B are different" may mean "A and B are different from each other". Note that this term may also mean "A and B are each different from C". Terms such as "separate" and "coupled" may also be interpreted in the same way as "different".

[0075] (17) Each aspect / embodiment described in the present disclosure may be used alone, in combination, or switched and used during execution. Also, the notification of predetermined information (for example, the notification of "being X") is not limited to being explicitly performed, and may be performed implicitly (for example, by not performing the notification of the predetermined information).

[0076] (D: Aspects grasped from the above-described forms or modified examples) As described above in detail, it is clear to those skilled in the art that the present disclosure is not limited to the embodiments described in the present disclosure. The present disclosure can be implemented as modified and changed aspects without departing from the spirit and scope of the present disclosure determined by the claims. Therefore, the description of the present disclosure is for illustrative purposes and does not have any limiting meaning for the present disclosure. The following aspects can be grasped from at least one of the above-described embodiments or modified examples.

[0077] The conversation analysis apparatus according to the first aspect includes an audio control unit, a first estimation unit, a second estimation unit, and a first determination unit. The audio control unit acquires audio data representing the audio of the user's response in the conversation between the robot and the user. The first estimation unit estimates the emotion contained in the user's response based on the tone of voice of the response, that is, the tone of voice represented by the audio data acquired by the audio control unit. The second estimation unit estimates the emotion contained in the user's response based on the text of the response, that is, the text represented by the audio data acquired by the audio control unit. The first determination unit determines whether the conversation between the user and the robot is established based on the degree of agreement indicating the degree of agreement between the emotion estimated by the first estimation unit and the emotion estimated by the second estimation unit. Since the conversation analysis apparatus according to the first aspect determines whether the conversation between the user and the robot is established based on the degree of agreement between the emotion estimated from the tone of voice of the user's response to the robot and the emotion estimated from the text of the response, it is possible to reduce misjudgment when the user is not sincerely responding.

[0078] The conversation analysis apparatus in the example of the first aspect (second aspect) may further include a second determination unit that determines that the conversation is not established when there is no response to the query to the user or when the topic is diverted by the response. The conversation analysis apparatus according to the second aspect can determine that the conversation between the user and the robot is not established when there is no response to the query to the user or when the topic is diverted by the response.

[0079] In the example of the first aspect or the second aspect (third aspect), the topic in the conversation between the robot and the user may be a thing that is the object of determination of the strength of the user's empathy. The conversation analysis apparatus according to the third aspect may further include the following third estimation unit. When it is determined by the first determination unit that the conversation is established, the third estimation unit estimates the degree of empathy indicating the strength of the user's empathy for the thing that is the object of determination of the strength of the user's empathy based on the emotion estimated by the first estimation unit and the emotion estimated by the second estimation unit. The conversation analysis apparatus according to the third aspect can estimate the strength of the user's empathy for the thing that is the topic in the conversation between the robot and the user.

[0080] The conversation analysis device in the example of the third aspect (the fourth aspect) may further include a presentation unit that presents to the user a story related to the thing to be the topic in the conversation between the robot and the user, that is, the thing to be the determination target of the strength of the user's empathy, prior to the start of the conversation between the robot and the user. The conversation analysis device of the fourth aspect can present to the user a story related to the thing to be the topic in the conversation between the robot and the user, that is, the thing to be the determination target of the strength of the user's empathy, prior to the conversation.

Explanation of Signs

[0081] 1... Conversation system, 10... User device, 30... Robot, 40... Conversation analysis device, 411... Presentation unit, 412... Voice control unit, 413... Second determination unit, 414... First estimation unit, 415... Second estimation unit, 416... First determination unit, 417... Third estimation unit.

Claims

1. An audio control unit that acquires audio data representing the voice of the user's response in the conversation between the robot and the user, A first estimation unit that estimates the emotion contained in the response based on the tone color of the voice represented by the audio data, A second estimation unit that estimates the emotion contained in the response based on the words represented by the audio data, A first determination unit that determines whether the conversation is established based on a degree of agreement indicating the degree of agreement between the emotion estimated by the first estimation unit and the emotion estimated by the second estimation unit, A conversation analysis device comprising the above.

2. The conversation analysis device according to claim 1, further comprising a second determination unit that determines that the conversation is not established when there is no response to the query to the user or when the topic is diverted by the response.

3. The topic in the conversation is a thing to be judged for the strength of the user's empathy, When it is determined by the first determination unit that the conversation is established, a third estimation unit that estimates an empathy degree indicating the strength of the user's empathy for the thing based on the emotion estimated by the first estimation unit and the emotion estimated by the second estimation unit, The conversation analysis device according to claim 1 or claim 2, further comprising:

4. The conversation analysis device according to claim 3, further comprising a presentation unit that presents a story related to the thing to the user prior to the start of the conversation.

Citation Information

Patent Citations

  • Interaction controller with interaction correcting function by feeling utterance detection

    JP2005258235A

  • Information providing device and program

    JP2014178621A

  • Psychological analyzer, psychological analysis method and program

    JP2017211586A

  • Voice interactive device and voice interactive method

    JP2017215468A

  • Speech processing device and speech processing method

    WO2012042768A1