Virtual conversation system, virtual conversation method, and program
The virtual conversation system addresses the challenge of customer harassment by integrating non-verbal expressions into conversation practice, enhancing interaction skills and improving prevention measures.
Patent Information
- Application Number
- JP2025066897
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-11-06
- Estimated Expiration
- 2045-04-15
AI Technical Summary
Existing systems fail to effectively address customer harassment by not considering non-verbal expressions in conversation practice, limiting companies' ability to implement comprehensive measures.
A virtual conversation system that allows users to interact with a virtual face-to-face person, incorporating non-verbal expressions through voice recognition, candidate display, and response determination, enabling users to practice conversations that consider both verbal and non-verbal cues.
Enables users to learn conversations that take non-verbal expressions into account, improving the effectiveness of customer harassment prevention by enhancing interaction skills.
Smart Images

Figure 0007765138000001_ABST
Abstract
Description
[Technical Field]
[0001] The present invention relates to a virtual conversation system, a method, and a program for allowing a user to have a conversation with a virtual person displayed on a screen. [Background technology]
[0002] According to a corporate survey conducted by the Ministry of Health, Labor and Welfare (October 2020), customer harassment has been increasing in recent years, becoming the third most common social problem after power harassment and sexual harassment, with 19.5% of companies experiencing it. Workers who have experienced it repeatedly reported serious problems, such as "losing sleep" (21.2%) and "having to go to the hospital or take medication" (8.8%). In February 2022, the Ministry published a "Corporate Manual for Countermeasures against Customer Harassment," and measures against customer harassment are being implemented in all industries.
[0003] As part of their duty of care, companies have a legal obligation to take measures to prevent customer harassment. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Japanese Patent Application Publication No. 10-326074 Summary of the Invention [Problem to be solved by the invention]
[0005] However, companies and other organizations do not know how to implement measures to combat customer harassment, and many are limited to symptomatic treatment or superficial responses, so a fundamental solution is needed.
[0006] Patent Document 1 discloses a control method for a language training device that displays a speech guide that indirectly suggests one of a single or multiple utterance sentences to a learner practicing conversation, and guides the learner to utter one of the utterance sentences prepared on the device side. However, since customer harassment complainers do not simply state the content of their complaints in a matter-of-fact manner, but rather use non-verbal expressions to express their complaints, it is difficult to practice countermeasures against customer harassment using this control method that uses only language.
[0007] In addition, a system is desired that can provide not only conversation practice for customer harassment prevention but also other conversation practice that takes non-verbal expressions into consideration.
[0008] The present invention has been made in consideration of the above points, and aims to provide a virtual conversation system, a virtual conversation method, and a program that allow users to learn conversations that take non-verbal expressions into consideration. [Means for solving the problem]
[0009] According to the present invention, A virtual conversation system for allowing a user to have a conversation with a virtual face-to-face person displayed on a screen, a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising: And, further comprising means for recognizing non-verbal expressions of the user; the candidate display control means displays at least one candidate by a combination of a character string indicating the content of the user's utterance corresponding to the content of the utterance of the virtual face-to-face partner and a character string indicating a non-verbal expression of the user; the candidate determination means determines, for the at least one candidate, whether the user has selected the candidate based on a pair of a character string indicating the recognized content of the user's utterance and a character string indicating the recognized non-verbal expression of the user, and the pair of character strings related to the candidate; Virtual Conversation System is provided. [Effects of the Invention]
[0010] According to the present invention, a user can learn conversations that take non-verbal expressions into consideration. [Brief explanation of the drawings]
[0011] [Figure 1] 1 is a diagram showing an example of a form in which a user uses a virtual conversation system in an embodiment of the present invention. [Figure 2] FIG. 10 is a diagram showing an example of an image displayed on a screen in the embodiment of the present invention. [Figure 3] 1 is a functional block diagram showing a partial configuration of a virtual conversation system according to an embodiment of the present invention. [Figure 4] FIG. 10 is a functional block diagram showing another partial configuration of the virtual conversation system according to the embodiment of the present invention. [Figure 5] 1 is a flowchart (1 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. [Figure 6] 1 is a flowchart (2 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. [Figure 7] 1 is a flowchart (3 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. [Figure 8] 1 is a flowchart (4 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. [Figure 9] 1 is a flowchart (5 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. [Figure 10]6 is a flowchart (6 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. [Figure 11] 7 is a flowchart (7 / 7) illustrating an example of a virtual conversation method according to an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION
[0012] DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.
[0013] [First embodiment]
[0014] 1, the virtual conversation system according to this embodiment is used, for example, in a seminar room 101. In the seminar room 101, there are a user 102 wearing VR goggles 104 and a plurality of participants 103 sitting in chairs.
[0015] The seminar room 101 is also equipped with a large monitor 105 and speakers 106 .
[0016] The user 102 can view images displayed on a screen 201 (see FIGS. 2 and 3) of the VR goggles 104 and hear sounds output from a speaker (first sound output unit 316 (see FIG. 3)) of the VR goggles 104. The VR goggles 104 also include a microphone (sound input unit 313 (see FIG. 3)) for inputting sounds uttered by the user 102. The VR goggles 104 also include a position detection unit 314 for detecting the position of the user 102. The position detection unit 314 can detect the position where the user 102 is located at each time.
[0017] 2, a virtual complainant 202 is displayed on a screen 201 of the VR goggles 104 as an example of a virtual interpersonal party. A character string 204 representing the content of the utterance of the virtual complainant 202 is also displayed on the screen 201. Furthermore, a plurality of (for example, four) candidate character strings 205#1 to 205#4 for the content of the utterance of the user corresponding to the content of the utterance of the virtual complainant 202 are displayed on the screen 201. Furthermore, a character string 203 representing the content of the utterance of the user 102 is displayed on the screen 201.
[0018] The speaker of the VR goggles 104 outputs the content of the utterance of the virtual complainant 202 as voice.
[0019] The same image as that displayed on the screen 201 of the VR goggles 104 is displayed on the screen 301 of a large monitor installed in the seminar room 101 .
[0020] A speaker 106 installed in the seminar room 101 outputs the speech content of the virtual complainant 202 as audio. The speaker 106 also outputs the amplified voice of the user.
[0021] Next, the basic usage of the virtual conversation system will be explained.
[0022] The virtual claimant 202 makes a speech containing some kind of content. The user 102 can hear this speech and can also read and recognize the character string 204.
[0023] The user 102 needs to respond in some way depending on the content of the utterance of the virtual complainant 202. When using the virtual conversation system according to this embodiment, the user selects one of a plurality of candidates displayed on the screen 201 by character strings 205#1 to 205#4 and reads the character string of the selected candidate as if speaking. The example in FIG. 2 shows the screen 201 when the user 102 selects the fourth candidate #4. Character string 203 obtained as a result of recognizing the voice of the user 102 reading character string 205#4 is displayed. That is, candidate #4 is "Is there anything I can help you with?", but the user 102 utters "Is there anything I can help you with?", which is reflected in character string 203.
[0024] Furthermore, since the user 102 selected the fourth candidate #4, only the frame surrounding the character string 205#4 becomes thicker. Even if the user does not read out the candidates accurately, it is possible to determine which candidate the user selected as long as the user does so to a certain extent. This is because, for example, the determination is made based on the cosine similarity between the vector formed by the character strings of the candidates and the vector formed by the character string obtained from the user's utterance. If the cosine similarity is greater than a predetermined threshold, it can be determined that the candidate was selected.
[0025] Cosine similarity is calculated by converting the candidate string (Q) and the string obtained from the utterance (A) into vectors, and then calculating the similarity based on the angle between those vectors.
[0026]
number
[0027] Regarding the vectorization method, each string is vectorized in the following way: Bag-of-Words (BoW): A simple representation using the number of occurrences of each word. ·TF-IDF: A representation that takes into account the frequency of occurrence of words and their importance across the entire document. ·Word Embeddings: Embedding vectors that take into account the meaning of words, such as Word2Vec, GloVe, and FastText. Sentence Vectors: Vectorize entire sentences using models such as BERT, RoBERTa, and GPT.
[0028] The participant 103 can know the actual exchange between the virtual claimant 202 and the user 102 through the image displayed on the screen 301 and the sound output from the speaker 106 .
[0029] Next, the configuration of the virtual conversation system according to this embodiment will be described with reference to FIG.
[0030] The first audio output control unit 301 generates audio from the speech of the virtual contact person (for example, a virtual complainant) and outputs the audio to the first audio output unit 316 (speaker of the VR goggles 104). To this end, the first audio output control unit 301 inputs a character string of the speech content of the virtual contact person from the speech content control unit in the form of, for example, character code, and performs speech synthesis based on the character string.
[0031] The candidate display control unit 302 displays, on the screen 201, a plurality of candidates for the content of the utterance of the user 102, which correspond to the content of the utterance of the virtual face-to-face person 202, as character strings. To this end, the candidate display control unit 302 receives as input identification information of the content of the utterance of the virtual face-to-face person 202 from the utterance content control unit 305, acquires identification information of the plurality of candidates by referring to the correspondence table 306, and acquires character strings (character code strings) corresponding to each candidate from the second character string storage unit 312. The candidate display control unit 302 also outputs character code strings indicating the character strings of each candidate to the screen 201. A character code to character image conversion unit (not shown) provided in the screen 201 displays the character strings based on the character code strings from the candidate display control unit 302.
[0032] Here, first identification information is assigned to the content of each utterance of the virtual face-to-face meeting 202. Second identification information is assigned to each candidate corresponding to the content of each utterance of the virtual face-to-face meeting 202. The correspondence table 306 holds a correspondence relationship between the first identification information and the second identification information. That is, the correspondence table 306 holds a correspondence relationship between the first identification information indicating the content of each utterance of the virtual face-to-face meeting 202 and the second identification information indicating the candidate to be displayed accordingly. The correspondence table 306 also holds a correspondence relationship between the second identification information indicating the candidate selected by the user 102 and the first identification information indicating the content to be uttered by the virtual face-to-face meeting 202 in accordance with the candidate selected by the user 102. That is, the correspondence table 306 holds a one-to-many correspondence relationship between the utterance content indicated by the first identification information and the candidate indicated by the second identification information, and a one-to-one correspondence relationship between each candidate indicated by the second identification information and the utterance content indicated by the first identification information.
[0033] The voice recognition unit 303 recognizes the voice of the user input by the voice input unit 313 (a microphone provided in the VR goggles 104) and converts it into a character string indicating the content of the utterance of the user 102. The character string is usually expressed by a character code string.
[0034] The candidate determination unit 304 determines which candidate the user 102 selected by comparing the character string of each candidate input from the candidate display control unit 302 with the character string input from the voice recognition unit 303. The comparison between character strings is performed, for example, by comparing character code strings with character code strings. The determination result is indicated by second identification information.
[0035] The speech content control unit 305 controls the content of the next utterance of the virtual contact person 202 based on the candidate selected by the user 102. That is, the speech content control unit 305 acquires first identification information corresponding to the second identification information input from the candidate determination unit 304 by referring to the correspondence table 306. Then, the speech content control unit 305 acquires a character string (character code string) of the utterance content of the virtual contact person to which the acquired first identification information is assigned from the first character string storage unit 311. Then, the acquired character string (character code string) is output to the first audio output control unit 301.
[0036] The virtual conversation system according to this embodiment allows the user to not only speak but also to respond to the virtual person through movements. In other words, in practice, for example, it may be effective for the person responding to the virtual person to move closer to the complainant, and this can be practiced using the virtual conversation system according to this embodiment.
[0037] Therefore, among the multiple candidates, in addition to those that indicate the content of the utterance of the user 102 in response to the content of the utterance of the virtual face-to-face person 202, those that indicate the behavior of the user 102 in response to the content of the utterance of the virtual face-to-face person 202 may be included.
[0038] When the user 102 selects a candidate for an action, the user 102 may perform the action according to the candidate. In this case, the action of the user 102 is identified by the position detection unit 314 and the action recognition unit 315. The action recognition unit 315 outputs the identified action to the candidate determination unit 304 in the form of a character string (character code string). Therefore, the candidate determination unit 304 can treat the action in the same way as an utterance.
[0039] When the user 102 selects a candidate for an action, the user 102 may read out the term indicating the action displayed as a character string of the candidate. In this case, there is no need to use the position detection unit 314 and the action recognition unit 315.
[0040] The virtual face-to-face person display control means 307 displays on the screen 201 the virtual face-to-face person 202 that has a physical attitude corresponding to the candidate selected by the user 102. To this end, for example, image data of the virtual face-to-face person 202 that corresponds to the first identification information is stored in a storage unit (not shown), and such image data is used. The physical attitude may include one or more elements selected from gestures, hand gestures, and facial expressions.
[0041] The tone control unit 308 causes the voice uttered by the virtual face-to-face person 202 to have a tone corresponding to the candidate selected by the user 102 via the first voice output control unit 301. To achieve this, for example, tone data indicating a tone corresponding to the first identification information is stored in a storage unit (not shown) and such tone data is used. When synthesizing voice, the first voice output control unit 301 synthesizes voice in accordance with the tone data.
[0042] The first character string display control unit 309 displays a character string 203 indicating the content of the user's utterance on the screen 201. That is, the first character string display control unit 309 acquires the result of the voice recognition from the voice recognition unit 303 as a character string (character code string), and outputs this to the screen 201. A character code to character image conversion unit (not shown) provided in the screen 201 displays the character string based on the character code string from the first character string display control unit 309.
[0043] The second character string display control unit 310 displays a character string indicating the content of the utterance of the virtual interpersonal partner 202 on the screen 201. That is, the second character string display control unit 310 obtains the character string (character code string) of the utterance content of the virtual interpersonal partner from the utterance content control unit 305, and outputs it to the screen 201. A character code to character image conversion unit (not shown) provided in the screen 201 displays the character string based on the character code string from the second character string display control unit 310.
[0044] 4, the external display control unit 402 causes the second screen 301 to display the same image as the image displayed on the first screen 201. Note that the large monitor 105 may receive an input of a video signal for the first screen 201 from the VR goggles 104, or the VR goggles 104 and the large monitor 105 may receive an input of a common video signal.
[0045] 3, the screen 201 receives data separately from the candidate display control unit 302, the virtual face-to-face person display control unit 307, the first character string display control unit 309, and the second character string display control unit 310. However, an image processing unit (not shown) for synthesizing images to be displayed on the screen 201 may receive data from these units, consolidate the data, and then output the data to the screen 201. The image processing unit may then output similar data to the screen 301.
[0046] The second audio output control unit 401 combines the audio generated by the first audio output control unit 301 and the audio input by the audio input unit 313 provided in the VR goggles 104 and outputs the combined audio to the second audio output unit 106 .
[0047] An example of a virtual conversation method performed by the virtual conversation system will be described with reference to Figures 5 to 11. In this example, the virtual face-to-face person 202 is a virtual complainer.
[0048] In S501 and S502, the virtual complainer 202 complains to the hotel receptionist played by the user 102.
[0049] The user 102 selects one of the four candidates in S503, S504, S505, and S506, and reads out the selected candidate as if speaking it.
[0050] If the user 102 selects the candidate in S503, the process proceeds to S507 (FIG. 6). If the user 102 selects the candidate in S504, the process proceeds to S521 (FIG. 7). If the user 102 selects the candidate in S505, the process proceeds to S529 (FIG. 8). If the user 102 selects the candidate in S506, the process proceeds to S537 (FIG. 9).
[0051] Referring to FIG. 6, in S507, the virtual complainant 202 makes an utterance corresponding to the candidate selected by the user in S503.
[0052] The user 102 selects one of the four candidates S508, S511, S512, and S518, and reads out the selected candidate as if speaking it.
[0053] If the user 102 selects the candidate in S508, the process proceeds to S509. If the user 102 selects the candidate in S511, the process proceeds to S539 (FIG. 9). If the user 102 selects the candidate in S505, the process proceeds to S513. If the user 102 selects the candidate in S518, the process proceeds to S519.
[0054] In S509, the virtual complainant 202 shows an attitude of "extremely angry", and the process ends with the evaluation result of "inappropriate" (S510).
[0055] The processing in S539 will be described later.
[0056] If the process proceeds to S513, the process goes through S514, S515, and S516, and ends with a processing result of "inappropriate" (S517). In S516, the virtual complainant 202 shows an "extremely angry" attitude.
[0057] In step 519, the virtual complainant 202 shows a "slightly angry" attitude, and the process ends with an evaluation result of "ifficult" (S520).
[0058] Referring to FIG. 7, in S521, the virtual complainant 202 makes an utterance corresponding to the candidate selected by the user in S504.
[0059] The user 102 selects one of the four candidates S522, S526, S527, and S528, and reads out the selected candidate as if speaking it.
[0060] If the user 102 selects the candidate in S522, the process proceeds to S523. If the user 102 selects the candidate in S526, the process also proceeds to S523. If the user 102 selects the candidate in S527, the process also proceeds to S523. If the user 102 selects the candidate in S528, the process proceeds to S537 (FIG. 9).
[0061] If the process proceeds to S523, the process goes through S524 and ends with a processing result of "inappropriate" (S525). In S524, the virtual complainant 202 shows an "extremely angry" attitude.
[0062] Referring to FIG. 8, in S529, the virtual complainant utters a message corresponding to the candidate selected by the user in S505.
[0063] The user 102 selects one of the four candidates in S530, S533, S534, and S535, and reads out the selected candidate as if speaking it.
[0064] If the user 102 selects the candidate in S530, the process proceeds to S531. If the user 102 selects the candidate in S533, the process also proceeds to S531. If the user 102 selects the candidate in S534, the process also proceeds to S531. If the user 102 selects the candidate in S535, the process proceeds to S531 via S536.
[0065] In S531, the virtual complainant 202 shows an "extremely angry" attitude, and the process ends with an evaluation result of "inappropriate" (S532).
[0066] Referring to FIG. 9, in S537, the virtual complainant 202 utters a message corresponding to the option selected by the user in S506 or S528.
[0067] If the user 102 selects the candidate in S538, the process proceeds to S539. If the user 102 selects the candidate in S540, the process also proceeds to S539. If the user 102 selects the candidate in S541, the process also proceeds to S539. If the user 102 selects the candidate in S542, the process proceeds to S543.
[0068] In S543, the virtual complainant 202 shows an attitude of "extremely angry," and the process ends with an evaluation result of "inappropriate" (S544).
[0069] In S539, the virtual complainant 202 speaks as if he or she is speaking his or her true feelings.
[0070] In response to S539, the user selects one of the four candidates in S535 (see FIG. 8), S545 (see FIG. 10), S546 (see FIG. 10), and S550 (see FIG. 10), and the selected candidate is read aloud.
[0071] If the user 102 selects the candidate in S535 (see FIG. 8), as already explained, the process proceeds to S531 via S536. In S531, the virtual complainant 202 shows an "extremely angry" attitude, and the process ends with an evaluation result of "inappropriate" (S532).
[0072] If the user 102 selects the candidate in S545, the process proceeds to S507 (see FIG. 6). The process from S507 onwards has already been explained, so a duplicate explanation will be omitted.
[0073] If the user 102 selects the candidate in S546, the virtual complainant 202 becomes "extremely angry" (S547 and S548), and the process ends with an evaluation result of "inappropriate" (S549).
[0074] If the user 102 selects the candidate in S550, the process proceeds to S551 (see FIG. 11).
[0075] In S551, the virtual complainant 202 takes a restrained attitude toward anger.
[0076] In response to S551, the user selects one of the four candidates in S552, S555, S559, and S562, and the selected candidate is read aloud.
[0077] If the user 102 selects the candidate in S552, in S553 the virtual complainant 202 shows an attitude that is neither angry nor convinced, and the process ends with an evaluation result of "ambiguous" (S554).
[0078] If the user 102 selects the candidate in S555, the virtual complainant 202 becomes "extremely angry" (S556 and S557), and the process ends with an evaluation result of "inappropriate" (S558).
[0079] If the user 102 selects the candidate in S559, the attitude of the virtual complainant 202 changes to one of understanding and gratitude in S560, and the process ends with an evaluation result of "perfect" (S561).
[0080] If the user 102 selects the candidate in S562, the attitude of the virtual complainant 202 changes to normal in S563, and the process ends with an evaluation result of "average" (S564).
[0081] The candidate for S550 (Fig. 10) is "It's an important day with your girlfriend, isn't it? Thank you for your understanding." However, this can be changed to "(Move to the side of the complaint.)" and "It's an important day with your girlfriend, isn't it? Thank you for your understanding." can be inserted between the revised S550 and S551 (Fig. 11).
[0082] In this case, the user 102 is given the option of moving to the side of the complainant after S539. When the user 102 actually moves to the side of the virtual complainant 202, the movement can be detected by the position detection unit 314 and the movement recognition unit 315. If the movement is detected, only one candidate 205#1, "It's an important day with your girlfriend, isn't it? ...I understand.", is displayed on the screen 201 instead of the four candidates 205#1 to 205#4. When the user 201 reads this out loud, the process proceeds to S551. Note that the user 102 may read out "(Move to the side of the complainant.)" instead of actually moving to the side of the virtual complainant 202. In this case, the process proceeds in the same way. To accommodate this, when the action recognition unit 315 recognizes the action of the virtual complainant 202, the candidate discrimination unit 304 outputs the identification information of the selected candidate, "It's an important day with your girlfriend, isn't it?... I understand." to the candidate display control unit 302 via a signal line (not shown), instead of outputting the identification information of the selected candidate to the speech content control unit 305.
[0083] The above change can be considered as changing the candidates in S550 from only linguistic expressions to a combination of linguistic and non-linguistic expressions. With this configuration, the user 102 can learn that a "perfect" evaluation result cannot be obtained unless non-linguistic expressions are used in addition to linguistic expressions.
[0084] The candidate in S550 may be "(Action: move to the side of the complainant) It's an important day with your girlfriend, isn't it? ... I understand." In order to accommodate this, when the voice recognition unit 303 recognizes "It's an important day with your girlfriend, isn't it? ... I understand." and the action recognition unit 315 recognizes the action of the virtual complainant 202, the candidate determination unit 304 may output identification information of the selected candidate to the utterance content control unit 305.
[0085] Instead of or in addition to the action of moving to the side of the complainant, a non-verbal expression related to a facial expression such as "smile" may be used. The smile of the user 102 can be detected even when the user 102 is wearing the VR goggles 104 by providing the VR 104 with a camera for photographing the eyes of the user 102. A camera for photographing the face of the user 102 may be used to detect a smile based on the corners of the user 102's mouth.
[0086] Non-verbal expressions may also be added to other candidates such as S552, S559, and S562.
[0087] The explanations are written near the blocks representing each step in Figures 6 to 11, so they will not be repeated, but in steps S509, S513, S515, S516, S519, etc., the virtual complainant 202 has the physical attitude and tone of voice as shown in the figures.
[0088] To deal with customer harassment, it is effective to take the following three steps: <Step 1> Mental preparation <Step 2> Hearing and approval <Step 3> Negotiation
[0089] Regarding the "mental preparation" in Step 1, the person acting as a customer harasser must understand that deep down, customers have a "fear" and must be prepared to resolve that "fear." This mental preparation will be reflected in the person's nonverbal expressions. Nonverbal expressions here include facial expressions, tone of voice, gestures, and hand movements. These nonverbal expressions convey to the customer that the person acting as a customer harasser is an ally, not an enemy. This prepares the customer to listen. Although not shown in the figure, the user acting as a customer harasser begins <Step 1> simultaneously with S501.
[0090] Regarding step 2, "listening and acknowledging," the person handling the customer needs to have the perspective of wondering why the customer is saying what they are. The person handling the customer needs to clarify the customer's concerns by asking questions such as, "What happened?", "What do you mean specifically?", and "Why do you think that?" By asking questions, the person handling the customer acknowledges the feelings expressed in the process of asking questions, which builds trust between the customer and the person handling the customer. The person handling the customer can deepen their understanding of the problem by imagining themselves in the customer's position and looking at it from that perspective. It is effective for the person handling the customer to gradually ask questions related to the customer's position.
[0091] Regarding step 3, "Negotiation," you basically clearly communicate that you cannot resolve the problem as the customer wishes. Step 3 consists of the following three substeps. <Step 3-1> Preframe (Conclusion) Explanation <Step 3-2> Explain the reasons and grounds <Step 3-3> Propose a solution to the problem within the scope of the rules
[0092] Regarding "Explaining the Preframe (Conclusion)" in Step 3-1, clearly communicate that you cannot solve the problem as the customer wishes. By clearly communicating the preframe (conclusion) to the customer, you will be able to lay the groundwork for the conversation, and you will be able to proceed along that groundwork.
[0093] Regarding step 3-2, "Explaining the reasons and evidence," by explaining "official" reasons and evidence, you can convince the customer of the preframe (conclusion) and gain their trust. Here, it would be a mistake to explain your own reasons instead of the "official" reasons.
[0094] Regarding step 3-3, "Providing a solution that fits within the rules," solving the problem within the rules will make the customer feel grateful that the person handling the problem has made concessions and thought about them. This will help the customer accept the solution.
[0095] If the user plays the role of the person in charge and makes utterances corresponding to the responses in <Step 1>, <Step 2>, <Step 3-1>, <Step 3-2>, and <Step 3-3> in order, the process will end with an evaluation result of "perfect."
[0096] Let us review the virtual conversation method shown in Figures 5 to 11.
[0097] <Step 2> corresponds to S506 and the following S538, S540, and S541. The questions are asked gradually in two stages.
[0098] <Step 2> also corresponds to S528 and the subsequent S538, S540, and S541. The questions are asked gradually in two stages.
[0099] S539, which follows S538, S540, and S541, reveals the customer's position and the details of their concerns.
[0100] <Step 3-1> corresponds to S550's "Sorry, but as a general rule, we cannot change rooms."
[0101] <Step 3-2> corresponds to S550's response: "We have reserved the period and room you requested for you, so please understand that it has been reserved for other guests."
[0102] <Step 3-3> is handled by S559. The problem-solving presented by S559 is within the scope of the rules and is based on an understanding of the customer's position and the content of their concerns.
[0103] Reaching S559 means that the user has made utterances corresponding to the responses in <Step 1>, <Step 2>, <Step 3-1>, <Step 3-2>, and <Step 3-3> in order, so in this case the processing ends with an evaluation result of "perfect" (S561).
[0104] S555 is placed in a position where S552, S559, and S562 can be selected as alternatives. Even though we have come this far, S555 states that "rules are rules" and states our own reasons. Therefore, the evaluation result is "inappropriate" and the process ends.
[0105] Since S562 is a certain degree of concession, the process ends with an evaluation result of "normal."
[0106] Since S545 is an explanation of the reason on our side, the evaluation result is "inappropriate" and the process ends.
[0107] S508, S522, 526, and S527 are also explanations of our reasons, so the evaluation result is "inappropriate" and the process ends.
[0108] S530, S533, S534, and S535 are also explanations of our reasons, so the evaluation result is "inappropriate" and the process ends.
[0109] Since S546 is also an explanation of our reasons, the evaluation result is "inappropriate" and the process ends.
[0110] Since S555 is also an explanation of our reasons, the evaluation result is "inappropriate" and the process ends.
[0111] The time from when the character string 204, 205#1 to 205#4 is switched to when the user 102 starts speaking may be measured, and statistical information (for example, average value and variance) about this may be obtained. The duration of the speech may also be measured, and statistical information about this may be obtained.
[0112] The VR goggles 104 may be equipped with a built-in camera for detecting the line of sight of the user 102. Then, a line of sight analysis of the user 102 may be performed. Such line of sight analysis may be performed for multiple users 102 belonging to an organization.
[0113] Finally, to finish the practice, the user 102 may be asked to recite a conversation that gets an evaluation result of "perfect" by rote as the virtual complainant 202. In other words, the user 102 may check whether he / she can recite a conversation that gets an evaluation result of "perfect" as the virtual complainant 202 without reading the character string 204 indicating the content of the utterance of the virtual complainant 202 or the candidate character strings 205#1 to 205#4. In that case, the user 102 uses the image of the virtual complainant 202 displayed on the first screen and the voice of the virtual complainant 202 heard from the speaker of the VR goggles 104 as main input information, and speaks the content that he / she has recited while showing non-verbal expressions.
[0114] As an example of a non-verbal expression, the timing, length, and tone of the user's 102 speech may be measured. Another example of a non-verbal expression is measuring the user's 102 gaze. For example, it may be measured which part of the virtual complainant 202 the user 102 is looking at, or whether the user 102 is looking at the virtual complainant 202 at all. Another example of a non-verbal expression is measuring the user's 102 gestures, hand movements, and other movements. If the user 102 moves, their movements may be measured. Another example of a non-verbal expression is measuring the user's facial expression. For this purpose, the user 102 does not wear VR goggles. A microphone is provided to input the voice of the user's 102 speech. Furthermore, a video camera is provided to capture the user's face. The user 102 interacts with the virtual complainant 202, for example, displayed on a screen (second screen) 301 of the large monitor 105. The user 102 also listens to the voice of the virtual complainant 202 output from the speaker 106, or wears headphones (not shown) and listens to the voice of the virtual complainant 2020 through the headphones.
[0115] The evaluation of the user 102 may be based not only on the content of the user's 102 utterances but also on the user's non-verbal expressions. For example, even if the content of the user's 102 utterances is "perfect," if the user's gaze is not directed at the virtual complainant 202 or the tone of the utterance is inappropriate for the service, the user may be given an overall evaluation lower than "perfect." Also, the content of the utterances may be "perfect," but the non-verbal expressions may be given an evaluation of "average." An evaluation may then be given for each element of the non-verbal expressions.
[0116] In the above embodiment, the user plays the role of a customer service representative dealing with complaints. However, the virtual conversation system according to this embodiment can also be applied when the user plays a role in other roles. For example, the virtual conversation system according to this embodiment can be applied when the user plays a person who serves important customers, a parent raising a child, or a caregiver caring for a care recipient in a nursing home. Furthermore, the virtual conversation system according to this embodiment can be applied when the user plays a counselor, a coach, or a consultant. Furthermore, the virtual conversation system according to this embodiment can be used for language practice, such as English conversation. Furthermore, the virtual conversation system according to this embodiment can be applied when the user plays a salesperson or an immigration bureau official. Furthermore, the virtual conversation system according to this embodiment can be applied when the user plays a school teacher or a politician.
[0117] The virtual call system can be realized by hardware, software, or a combination of these. The virtual call method performed by the virtual call system can also be realized by hardware, software, or a combination of these. "Realized by software" here means that the method is realized by a computer reading and executing a program.
[0118] The program can be stored and supplied to a computer using various types of non-transitory computer-readable media. Non-transitory computer-readable media include various types of tangible storage media. Examples of non-transitory computer-readable media include magnetic recording media (e.g., flexible disks, magnetic tapes, hard disk drives), magneto-optical recording media (e.g., magneto-optical disks), CD-ROMs (Read Only Memory), CD-Rs, CD-R / Ws, and semiconductor memories (e.g., mask ROMs, PROMs (Programmable ROMs), EPROMs (Erasable PROMs), flash ROMs, and RAMs (Random Access Memory)). The program may also be supplied to a computer by various types of transitory computer-readable media. Examples of transitory computer-readable media include electrical signals, optical signals, and electromagnetic waves. The transitory computer-readable media can supply the program to a computer via a wired communication path such as an electric wire or optical fiber, or via a wireless communication path.
[0119] The present invention can be embodied in various other forms without departing from its spirit or main characteristics. Therefore, the above-described embodiments are merely examples and should not be interpreted as limiting. The scope of the present invention is defined by the claims and is not limited to the text of the specification. Furthermore, all modifications and variations within the equivalent range of the claims are within the scope of the present invention.
[0120] A part or all of the above-described embodiments can be described as, but not limited to, the following supplementary notes.
[0121] [Appendix 1] A virtual conversation system for allowing a user to have a conversation with a virtual face-to-face person displayed on a screen, a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising:
[0122] [Appendix 2] further comprising means for recognizing non-verbal expressions of the user; the candidate display control means displays at least one candidate by a combination of a character string indicating the content of the user's utterance corresponding to the content of the utterance of the virtual face-to-face partner and a character string indicating a non-verbal expression of the user; the candidate determination means determines, for the at least one candidate, whether the user has selected the candidate based on a pair of a character string indicating the recognized content of the user's utterance and a character string indicating the recognized non-verbal expression of the user, and the pair of character strings related to the candidate; 10. The virtual conversation system of claim 1.
[0123] [Appendix 3] further comprising means for recognizing non-verbal expressions of the user; At least one candidate includes a combination of the content of the user's utterance and a non-verbal expression of the user corresponding to the content of the utterance of the virtual face-to-face person; The candidate display control means and the candidate determination means perform the following operations with respect to the at least one candidate: (i) the candidate display control means displays a character string indicating the non-verbal expression of the user related to the candidate; (ii) the candidate determination means determines whether the user has selected the candidate based on a character string indicating the recognized non-verbal expression of the user and a character string indicating a non-verbal expression related to the candidate; (iii) the candidate display control means displays a character string indicating the content of the user's utterance related to the candidate; (iv) the candidate determination means determines whether the user has selected the candidate based on a character string indicating the content of the recognized utterance of the user and the character string indicating the content of the utterance related to the candidate; 10. The virtual conversation system of claim 1.
[0124] [Appendix 4] the candidate display control means displays at least one candidate by a combination of a character string indicating the content of the user's utterance corresponding to the content of the utterance of the virtual face-to-face partner and a character string indicating a non-verbal expression of the user; the candidate determination means determines whether the user has selected the at least one candidate based on a pair of a character string indicating the recognized content of the user's utterance and a character string indicating the user's linguistically recognized non-verbal expression, and the pair of character strings related to the candidate; 10. The virtual conversation system of claim 1.
[0125] [Appendix 5] At least one candidate includes a combination of the content of the user's utterance and a non-verbal expression of the user corresponding to the content of the utterance of the virtual face-to-face person; The candidate display control means and the candidate determination means perform the following operations with respect to the at least one candidate: (i) the candidate display control means displays a character string indicating the non-verbal expression of the user related to the candidate; (ii) the candidate determination means determines whether the user has selected the candidate based on a character string representing the linguistically recognized non-verbal expression of the user and the character string representing the non-verbal expression related to the candidate; (iii) the candidate display control means displays a character string indicating the content of the user's utterance related to the candidate; (iv) the candidate determination means determines whether the user has selected the candidate based on a character string indicating the content of the recognized utterance of the user and the character string indicating the content of the utterance related to the candidate; 10. The virtual conversation system of claim 1.
[0126] [Appendix 6] The system further includes a correspondence table that stores a correspondence relationship between the content of the utterance of the virtual face-to-face meeting and a plurality of corresponding candidates, and a correspondence relationship between the candidate selected by the user and the content of the next utterance of the virtual face-to-face meeting corresponding to the candidate, the candidate display control means and the speech content control means refer to the correspondence table. 6. A virtual conversation system according to any one of appendices 1 to 5.
[0127] [Appendix 7] The system further includes a tone control means for causing the voice uttered by the virtual face-to-face contact to have a tone corresponding to the candidate selected by the user via the first voice output control means, 7. A virtual conversation system according to any one of appendices 1 to 6.
[0128] [Appendix 8] The virtual face-to-face person is a virtual complainer who makes a complaint against the user, The virtual conversation system operates repeatedly until it receives a rating of "perfect," "average," "slight," or "inappropriate," When the user selects a candidate for final "perfection" in each iteration, the content of the user's utterance includes, in this order: (A) hearing and approval; (B) explanation of the pre-frame for the claim; (C) explanation of the official grounds and reasons for the pre-frame; and (D) presentation of a solution to the claim within the scope of the rules. 8. A virtual conversation system according to any one of appendices 1 to 7.
[0129] [Appendix 9] If the user selects a candidate that ultimately leads to a rating of "Poor" in any iteration, or if the user selects a candidate that ultimately leads to a rating of "Mellow" in any iteration, it is not possible to reach a rating of "Perfect" or "Average." 9. The virtual conversation system of claim 8.
[0130] [Appendix 10] the candidate display control means causes a display mode of the candidate determined by the candidate determination means to be selected by the user to differ from a display mode of other candidates; 10. A virtual conversation system according to any one of appendices 1 to 9.
[0131] [Appendix 11] The system further includes a first character string display control means for displaying a character string indicating the content of the user's utterance on the screen. 11. A virtual conversation system according to any one of appendices 1 to 10.
[0132] [Appendix 12] The system further includes a second character string display control means for displaying a character string indicating the content of the utterance of the virtual face-to-face contact on the screen. 12. A virtual conversation system according to any one of claims 1 to 11.
[0133] [Appendix 13] The system further includes a second audio output control means for combining the audio generated by the audio output means and the audio generated by the user, and outputting the combined audio to another audio output means. 13. A virtual conversation system according to any one of appendices 1 to 12.
[0134] [Appendix 14] further comprising an external display control means for displaying the same image as the image displayed on the screen on another screen; 14. A virtual conversation system according to any one of claims 1 to 13.
[0135] [Appendix 15] means for measuring non-verbal expressions of said user; means for evaluating the user based on the non-verbal expression; 15. The virtual conversation system of any one of claims 1 to 14, further comprising:
[0136] [Appendix 16] A virtual conversation method for allowing a user to have a conversation with a virtual face-to-face person displayed on a screen, comprising: a first voice output control step of generating voice uttered by the virtual face-to-face contact and outputting the voice to a voice output means; a candidate display control step of displaying, on the screen, a plurality of candidates for the content of the user's utterance according to the content of the utterance of the virtual face-to-face person as character strings; a speech recognition step of recognizing a speech uttered by the user and converting the speech uttered by the user into a character string indicating the content of the speech; a candidate determination step of determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and character strings of the plurality of candidates; a speech content control step of controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control step of displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation method having the steps:
[0137] [Appendix 17] A program for causing a computer to function as the virtual conversation system described in any one of appendices 1 to 15. [Explanation of symbols]
[0138] 102 User 103 participants 104 VR goggles 105 Large Monitor 106 Speaker
Claims
1. A virtual conversation system for allowing a user to have a conversation with a virtual face-to-face person displayed on a screen, a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising: further comprising means for recognizing non-verbal expressions of the user; the candidate display control means displays at least one candidate by a combination of a character string indicating the content of the user's utterance corresponding to the content of the utterance of the virtual face-to-face partner and a character string indicating a non-verbal expression of the user; the candidate determination means determines, for at least one candidate, whether the candidate has been selected by the user, based on a pair of a character string indicating the recognized content of the user's utterance and a character string indicating the recognized non-verbal expression of the user, and the pair of character strings related to the candidate; Virtual conversation system.
2. A virtual conversation system for allowing a user to have a conversation with a virtual person displayed on a screen, comprising: a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising: further comprising means for recognizing non-verbal expressions of the user; At least one candidate includes a combination of the content of the user's utterance and a non-verbal expression of the user corresponding to the content of the utterance of the virtual face-to-face person; The candidate display control means and the candidate determination means perform the following operations with respect to the at least one candidate: (i) the candidate display control means displays a character string indicating the non-verbal expression of the user related to the candidate; (ii) the candidate determination means determines whether the user has selected the candidate based on a character string indicating the recognized non-verbal expression of the user and the character string indicating the non-verbal expression related to the candidate; (iii) the candidate display control means displays a character string indicating the content of the user's utterance related to the candidate; (iv) the candidate determination means determines whether the user has selected the candidate based on a character string indicating the content of the recognized utterance of the user and the character string indicating the content of the utterance related to the candidate; 10. The virtual conversation system of claim 1.
3. A virtual conversation system for allowing a user to have a conversation with a virtual person displayed on a screen, comprising: a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising: the candidate display control means displays at least one candidate by a combination of a character string indicating the content of the user's utterance corresponding to the content of the utterance of the virtual face-to-face partner and a character string indicating a non-verbal expression of the user; the candidate determination means determines whether the user has selected at least one of the candidates based on a pair of a character string indicating the recognized content of the user's utterance and a character string indicating the user's linguistically recognized non-verbal expression, and the pair of character strings related to the candidate; 10. The virtual conversation system of claim 1.
4. A virtual conversation system for allowing a user to have a conversation with a virtual person displayed on a screen, comprising: a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising: At least one candidate includes a combination of the content of the user's utterance and a non-verbal expression of the user corresponding to the content of the utterance of the virtual face-to-face person; The candidate display control means and the candidate determination means perform the following operations with respect to the at least one candidate: (i) the candidate display control means displays a character string indicating the non-verbal expression of the user related to the candidate; (ii) the candidate determination means determines whether the user has selected the candidate based on a character string representing the linguistically recognized non-verbal expression of the user and the character string representing the non-verbal expression related to the candidate; (iii) the candidate display control means displays a character string indicating the content of the user's utterance related to the candidate; (iv) the candidate determination means determines whether the user has selected the candidate based on a character string indicating the content of the recognized utterance of the user and the character string indicating the content of the utterance related to the candidate; 10. The virtual conversation system of claim 1.
5. A virtual conversation system for allowing a user to have a conversation with a virtual person displayed on a screen, comprising: a first voice output control means for generating voice uttered by the virtual face-to-face person and causing a voice output means to output the voice; a candidate display control means for displaying, on the screen, a plurality of candidates for the content of the user's utterance in accordance with the content of the utterance of the virtual face-to-face person, as character strings; a speech recognition means for recognizing a speech uttered by the user and converting it into a character string representing the content of the user's speech; a candidate determination means for determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and the character strings of the plurality of candidates; a speech content control means for controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control means for displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation system comprising: The virtual face-to-face person is a virtual complainer who makes a complaint against the user, The virtual conversation system operates repeatedly until it receives a rating of "perfect," "average," "slight," or "inappropriate," When the user selects a candidate for final "perfection" in each iteration, the content of the user's utterance includes, in this order: (A) hearing and approval; (B) explanation of the pre-frame for the claim; (C) explanation of the official grounds and reasons for the pre-frame; and (D) presentation of a solution to the claim within the scope of the rules.
10. The virtual conversation system of claim 1.
6. If the user selects a candidate that ultimately leads to a rating of "Poor" in any iteration, or if the user selects a candidate that ultimately leads to a rating of "Mellow" in any iteration, the user is unable to reach a rating of "Perfect" or "Average." 6. The virtual conversation system of claim 5.
7. The system further includes a correspondence table that stores a correspondence relationship between the content of the utterance of the virtual face-to-face meeting and a plurality of corresponding candidates, and a correspondence relationship between the candidate selected by the user and the content of the next utterance of the virtual face-to-face meeting corresponding to the candidate, the candidate display control means and the speech content control means refer to the correspondence table. A virtual conversation system according to any one of claims 1 to 6.
8. The system further includes a tone control means for causing the speech of the virtual face-to-face contact to have a tone corresponding to the candidate selected by the user via the first voice output control means. A virtual conversation system according to any one of claims 1 to 6.
9. the candidate display control means causes a display mode of the candidate determined by the candidate determination means to be selected by the user to differ from a display mode of other candidates; A virtual conversation system according to any one of claims 1 to 6.
10. a first character string display control means for displaying a character string indicating the content of the user's utterance on the screen; A virtual conversation system according to any one of claims 1 to 6.
11. The system further includes a second character string display control means for displaying a character string indicating the content of the utterance of the virtual face-to-face contact on the screen. A virtual conversation system according to any one of claims 1 to 6.
12. The voice output control means may further include a second voice output control means for combining the voice generated by the voice output means and the voice generated by the user, and outputting the combined voice to another voice output means. A virtual conversation system according to any one of claims 1 to 6.
13. further comprising an external display control means for displaying the same image as the image displayed on the screen on another screen; A virtual conversation system according to any one of claims 1 to 6.
14. means for measuring non-verbal expressions of said user; means for evaluating the user based on the non-verbal expression; 7. A virtual conversation system according to any one of claims 1 to 6, further comprising:
15. A virtual conversation method for allowing a user to have a conversation with a virtual face-to-face person displayed on a screen, comprising: a first voice output control step of generating voice uttered by the virtual face-to-face contact and causing voice output means to output the voice; a candidate display control step of displaying, on the screen, a plurality of candidates for the content of the user's utterance according to the content of the utterance of the virtual face-to-face person as character strings; a speech recognition step of recognizing a speech uttered by the user and converting the speech uttered by the user into a character string indicating the content of the speech; a candidate determination step of determining which of the plurality of candidates has been selected by the user based on a character string indicating the content of the recognized utterance of the user and character strings of the plurality of candidates; a speech content control step of controlling the content of the next speech of the virtual face-to-face person based on the candidate selected by the user; a virtual face-to-face person display control step of displaying on the screen the virtual face-to-face person who takes a physical attitude corresponding to the candidate selected by the user; A virtual conversation method comprising: further comprising the step of recognizing non-verbal expressions of the user; In the candidate display control step, at least one candidate is displayed by a combination of a character string indicating the content of the user's utterance corresponding to the content of the utterance of the virtual face-to-face partner and a character string indicating a non-verbal expression of the user; In the candidate determination step, for at least one candidate, it is determined whether the candidate has been selected by the user based on a pair of a character string indicating the content of the recognized utterance of the user and a character string indicating the recognized non-verbal expression of the user, and the pair of character strings related to the candidate. Virtual conversation method.
16. A program for causing a computer to function as the virtual conversation system according to any one of claims 1 to 5.
Citation Information
Patent Citations
Apparatus and method for training with an interpersonal interaction simulator
JP2002530724A
Emotionally appealing counseling systems and methods
JP2010531478A
Control method for language training device
JP1998326074A