Information processing program and information processing apparatus
The information processing program addresses the challenge of flexible feedback and content creation in conversational training by recognizing conversation steps and utterance tags, enhancing the effectiveness of conversational training through tailored educational content.
Patent Information
- Application Number
- JP2024080263
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2024-05-16
- Publication Date
- 2025-11-28
AI Technical Summary
Existing conversational training methods lack flexibility in providing feedback and educational content tailored to the specific type of conversation, leading to personalized skill acquisition and a heavy burden on instructors.
An information processing program that recognizes conversation steps and utterance tags within conversations, allowing for the flexible creation of educational content and feedback based on the type of conversation.
Enables intuitive understanding of conversation progress and provides flexible feedback, improving the effectiveness of conversational training by tailoring educational content to the specific context.
Smart Images

Figure 2025174158000001_ABST
Abstract
Description
[Technical Field]
[0001] An embodiment of the present invention relates to an information processing program and an information processing device. [Background technology]
[0002] In conversational tasks such as answering inquiries as a call center operator or facilitating a workshop, individuals are required to gain experience in order to acquire the skills required for the job. Because skills are acquired through individual experience, the acquired skills tend to become personalized. In addition, in order to pass on these skills, it is important for instructors to provide feedback to inexperienced newcomers every time they perform a task, but providing feedback while the instructor is still performing the task places a heavy burden on the instructor.
[0003] A known technology for supporting feedback on conversational work involves visualizing the content and scores of individual utterances in a conversation related to work, thereby enabling real-time feedback on the conversation. Another known technology for supporting the creation of learning content for conversational work involves, for example, a technology in which an evaluator assigns an evaluation based on multiple predetermined items to responses made in a call center, and presents learning content that corresponds to the evaluation as the correct answer. [Prior art documents] [Patent documents]
[0004] [Patent Document 1] Patent No. 7231894 [Patent Document 2] Patent No. 4400052 Summary of the Invention [Problem to be solved by the invention]
[0005] In order to properly pass on the skills required for conversational tasks, it is desirable that the creation of educational content and the accompanying feedback be carried out flexibly for each type of conversation.
[0006] The embodiments provide an information processing program and an information processing device that can flexibly provide feedback according to the type of conversation and flexibly create educational content. [Means for solving the problem]
[0007] One embodiment of an information processing program is an information processing program that supports learning of tasks performed through conversation, and causes a processor to perform the following operations: based on conversation information including information about conversations held in conversations between speakers, recognize conversation steps that can occur from the start to the end of a conversation and are composed of one or more semantically consistent utterances; recognize utterance tags that are tags that indicate the intention of each utterance in the conversation related to the conversation information; and associate the utterances included in each conversation step with the utterance tags. [Brief explanation of the drawings]
[0008] [Figure 1] FIG. 1 is a block diagram showing an information processing apparatus according to the first embodiment. [Figure 2] FIG. 2 is a diagram showing an example of conversation information in the first embodiment. [Figure 3] FIG. 3 is a diagram showing an example of conversation step data. [Figure 4] FIG. 4 is a diagram showing an example of a step label. [Figure 5] FIG. 5 is a diagram showing an example of an utterance tag list. [Figure 6] FIG. 6 is a diagram illustrating an example of a hardware configuration of an information processing device. [Figure 7] FIG. 7 is a flowchart showing the operation of the information processing device according to the first embodiment. [Figure 8]FIG. 8 is a diagram showing a display screen of a display device as an example of information output in the first embodiment. [Figure 9] FIG. 9 is a block diagram showing an information processing apparatus according to the second embodiment. [Figure 10] FIG. 10 is a diagram showing an example of conversation information in the second embodiment. [Figure 11] FIG. 11 is a diagram showing an example of speaker state data. [Figure 12] FIG. 12 is a diagram showing an example of analysis result data. [Figure 13] FIG. 13 is a flowchart showing the operation of the information processing device according to the second embodiment. [Figure 14] FIG. 14 is a diagram showing a display screen of a display device as a first example of information output in the second embodiment. [Figure 15] FIG. 15 is a diagram showing a display screen of a display device as a second example of information output in the second embodiment. [Figure 16] FIG. 16 is a diagram showing a display screen of a display device as a third example of information output in the second embodiment. [Figure 17] FIG. 17 is a block diagram showing an information processing apparatus according to the third embodiment. [Figure 18] FIG. 18 is a diagram showing an example of teaching material content data. [Figure 19] FIG. 19 is a flowchart showing the operation of the information processing device according to the third embodiment. [Figure 20] FIG. 20 is a diagram showing a display screen of a question on a display device as an example of question output in the third embodiment. DETAILED DESCRIPTION OF THE INVENTION
[0009] Hereinafter, an embodiment will be described with reference to the drawings.
[0010] (First embodiment) An information processing device according to a first embodiment receives conversation information including text of a conversation held for business purposes, recognizes conversation steps and utterance tags from the conversation information, and visualizes the recognized information. A conversation in the embodiment typically includes utterances made by two or more speakers. Such business may include responding to inquiries at a call center, sales talks by a salesperson, facilitation of a workshop, etc. The information processing device is provided, for example, in an operator response evaluation system at a call center.
[0011] 1 is a block diagram showing an information processing device according to the first embodiment. The information processing device 100 according to the first embodiment includes a receiving unit 101, a conversation step recognition unit 102, an utterance tag recognition unit 103, a corresponding unit 104, and an output unit 105.
[0012] The receiving unit 101 receives the input conversation information. Fig. 2 is a diagram showing an example of conversation information in the first embodiment. Fig. 2 shows conversation information about how to handle an inquiry at a call center.
[0013] The conversation information includes text data of L (L is an integer equal to or greater than 2) utterance sentences 11, 12, ..., 1L associated with an utterance ID. The utterance ID is an ID assigned to each utterance. The text data of sentences 11-1L may be generated by transcribing the utterances made during the conversation. The conversation information in Figure 2 is composed of an utterance ID and text data of the utterance sentences. In contrast, the conversation information may include data other than text data of the utterance sentences, such as audio data of the utterances, data on the speaker's name for each utterance, and data describing the business content, such as "responding to inquiries at a call center."
[0014] The conversation step recognition unit 102 recognizes conversation steps from the conversation information and generates conversation step data based on the recognition results. The conversation step data is information related to conversation steps. A conversation step is a step that can occur from the start to the end of a conversation and is composed of one or more semantically consistent utterances. In an embodiment, one conversation may be composed of two or more, typically three or more, conversation steps. A business conversation usually begins with a conversation step "start of conversation" and ends with a conversation step "end of conversation." However, the conversation steps in between may differ depending on the type of conversation. The conversation step recognition unit 102 recognizes conversation steps from the conversation information and generates conversation step data by assigning step labels corresponding to each conversation step. The conversation step recognition unit 102 may assign the same step label to multiple conversation steps, or may not necessarily assign all predetermined types of step labels.
[0015] 3 is a diagram showing an example of conversation step data. The conversation step data includes a list of M utterance IDs (M is an integer equal to or greater than 2) each associated with a step ID, and step labels 21, 22, ..., 2L. A step ID is an ID assigned to each conversation step. An utterance ID list is a list of utterance IDs included in the corresponding conversation step. A step label is a label that indicates the content of the conversation step.
[0016] FIG. 4 is a diagram showing example step labels. FIG. 4 shows step labels for "Call Center Inquiry Response," "Sales Talk," "Brainstorming Workshop," "One-on-One Meeting from Supervisor to Subordinate for Performance Evaluation Feedback," and "Research Presentation at Academic Conference." Step labels for "Call Center Inquiry Response" include, for example, "Start of Conversation," "Understanding the Situation," "Explaining the Response," "Reaching Consensus," and "End of Conversation." Conversation steps for "Sales Talk" include, for example, "Start of Conversation," "Chat," "Understanding the Situation," "Product Description," "Heard Customer Feedback," "Confirm Purchase Intention," "Explaining the Response," "Reaching Consensus," and "End of Conversation." Step labels for "Brainstorming Workshop" include, for example, "Start of Conversation," "Icebreaker," "Explaining the Purpose," "Idea Creation," "Idea Grouping," "Wholesaling Workshop," and "End of Conversation." Step labels for "One-on-One Meeting from Supervisor to Subordinate for Performance Evaluation Feedback" include, for example, "Start of Conversation," "Chat," "Explaining the Purpose," "Explaining the Evaluation," "Evaluation Feedback," "Heard Questions," "Heard Requests," and "End of Conversation." Step labels for "research presentation at a conference" include, for example, "start of conversation," "background explanation," "purpose explanation," "related research explanation," "proposed method explanation," "experimental results explanation," "future prospects explanation," "question response," and "conversation end."
[0017] 4 shows step labels for "responding to inquiries at a call center," "sales talk by a salesperson," "brainstorming workshop," "one-on-one meeting from a supervisor to a subordinate to provide performance evaluation feedback," and "research presentation at an academic conference." However, the conversations that are the subject of processing in the embodiment are not limited to those shown in FIG. 4. In other words, the conversations that are the subject of processing in the embodiment can include any conversation that can be described as conversation steps, that is, that progresses according to a predetermined flow to some extent.
[0018] The step labels shown in Figure 4 are also examples. For example, the step labels for "responding to inquiries at a call center" are not limited to those that include only "start conversation," "understand the situation," "explain how to respond," "reach a consensus," and "end conversation." For example, step labels for responding to inquiries on the spot while operating the screen may include "explain how to operate the screen" and "check screen operation," etc.
[0019] The speech tag recognition unit 103 recognizes speech tags for each utterance in the conversation information, and generates an utterance tag list based on the recognition result.
[0020] FIG. 5 is a diagram showing an example of an utterance tag list. The utterance tag list includes L utterance tags 31, 32, ..., 3L, each associated with an utterance ID. The utterance tags are tags that represent the intention of each utterance included in the conversation information, such as "greeting," "apology," "question," "explanation," "confirmation," "sympathy," and "proposal." The utterance tags may differ depending on the type of conversation. For example, different utterance tags may be assigned to the same type of utterance if the type of conversation is different.
[0021] The correspondence unit 104 associates each utterance in the conversation step data with each utterance tag in the utterance tag list. The correspondence unit 104 associates each utterance in the conversation step data with each utterance tag by referring to the utterance ID list in the conversation step data and the utterance ID in the utterance tag list.
[0022] The output unit 105 receives the conversation step data and the utterance tag list associated by the association unit 104. Based on the conversation step data and the utterance tag list, the output unit 105 performs processing to output information representing conversation steps and the intention of the utterance in each conversation step. For example, the output unit 105 performs processing necessary to visualize and display the intention of each utterance for each conversation step on a display device. The processing necessary for the visualized display includes processing such as creating a display screen.
[0023] 6 is a diagram showing an example of the hardware configuration of the information processing device 100. The information processing device 100 is a computer, and has, as hardware, for example, a processor 201, a memory 202, an input device 203, a display device 204, a communication device 205, and a storage 206. The processor 201, the memory 202, the input device 203, the display device 204, the communication device 205, and the storage 206 are connected to a bus 207.
[0024] The processor 201 is a processor that controls the overall operation of the information processing device 100. The processor 201 operates as the receiving unit 101, the conversation step recognition unit 102, the utterance tag recognition unit 103, the corresponding unit 104, and the output unit 105, for example, by executing a program stored in the storage 206. The processor 201 is, for example, a CPU. The processor 201 may be an MPU, a GPU, an ASIC, an FPGA, or the like. The processor 201 may be a single CPU or the like, or may be multiple CPUs or the like.
[0025] The memory 202 includes a ROM and a RAM. The ROM is a non-volatile memory. The ROM stores a startup program for the information processing device 100 and the like. The RAM is a volatile memory. The RAM is used as a working memory for processing by the processor 201, for example.
[0026] The input device 203 is an input device such as a touch panel, a keyboard, a mouse, etc. When the input device 203 is operated, a signal corresponding to the operation content is input to the processor 201 via the bus 207. The processor 201 performs various processes in response to this signal.
[0027] The display device 204 is a display device such as a liquid crystal display or an organic EL display. Various information output devices such as a printer may be provided instead of or in addition to the display device 204. Furthermore, the display device 204 does not necessarily have to be provided in the information processing device 100, and may be an external display device capable of communicating with the information processing device 100.
[0028] The communication device 205 is a communication device that enables the information processing device 100 to communicate with external devices. The communication device 205 may be a communication device for wired communication or a communication device for wireless communication.
[0029] The storage 206 is, for example, a storage such as a hard disk drive or a solid state drive. The storage 206 stores various programs executed by the processor 201, such as an information processing program 2061. The storage 206 may also store natural language processing models for various processes such as recognizing conversation steps and recognizing utterance tags.
[0030] The bus 207 is a data transfer path for exchanging data between the processor 201 , memory 202 , input device 203 , display device 204 , communication device 205 , and storage 206 .
[0031] Next, the operation of the information processing device according to the first embodiment will be described. Fig. 7 is a flowchart showing the operation of the information processing device according to the first embodiment. In the following, the explanation will be continued assuming that conversation information in response to an inquiry at a call center shown in Fig. 2 is input to the information processing device.
[0032] In step S101, the conversation step recognition unit 102 recognizes conversation steps from the conversation information. Specifically, the conversation step recognition unit 102 recognizes semantically inconsistent sections between multiple utterances as step boundaries by analyzing the semantic structure of each utterance included in the conversation information. The conversation step recognition unit 102 then recognizes each conversation step by dividing the utterance sentence before and after the step boundary, and generates an utterance ID list based on the recognition results. Step boundaries can be recognized, for example, by inputting the text of each utterance sentence as conversation information into a natural language processing model. For example, the conversation step recognition unit 102 analyzes the semantic structure of the sentences by inputting the text of the sentences into a model capable of performing anaphora resolution, which analyzes the referential relationship between demonstratives or pronouns, and recognizes a section between sentences that has no referential relationship as a step boundary. Alternatively, the conversation step recognition unit 102 may input the text of the sentences into a model that outputs the probability that the section between two consecutive sentences is a step boundary, compare the output probability with a threshold, and recognize a section between sentences where the output probability is equal to or greater than the threshold as a step boundary. Furthermore, when the conversation step recognition unit 102 is configured to accept all sentences included in the conversation information as input at once, either a model capable of performing anaphora resolution or a model capable of outputting the probability of a step boundary, the conversation step recognition unit 102 may recognize the step boundary from an index representing the step boundary created based on the output of each model, or may recognize the step boundary from the start and end positions of the step boundary created based on the output of each model.
[0033] For example, in the case of the conversation information in Figure 2, "equipment" in sentence 12 is considered to refer to "equipment" in sentence 13, so the conversation step recognition unit 102 recognizes that sentences 12 and 13 are included in the same semantically consistent conversation step. On the other hand, since sentences 11 and 12 do not have a reference relationship, the conversation step recognition unit 102 recognizes the space between sentences 11 and 12 as a step boundary. In this case, the conversation step recognition unit 102 stores utterance ID_1 in the utterance ID list of step ID_1, and stores utterance ID_2 and ID_3 in the utterance ID list of the next step ID_2, as shown in Figure 3.
[0034] In step S102, the conversation step recognition unit 102 assigns a step label to each of the recognized conversation steps. For example, if step labels have been determined, the conversation step recognition unit 102 may classify the conversation steps using a natural language processing model and assign a step label corresponding to the class to each classified conversation step. Alternatively, the conversation step recognition unit 102 may select representative vocabulary from the vocabulary in the sentences of utterances included in each conversation step and assign the selected vocabulary as the step label of the corresponding conversation step. Alternatively, the conversation step recognition unit 102 may summarize the utterances included in each conversation step to the phrase level using a natural language processing model and assign the summary as the step label of the corresponding conversation step. Alternatively, the conversation step recognition unit 102 may vectorize a label list predefined based on the vocabulary and conversation type in the utterances included in each conversation step using a natural language processing model capable of vectorizing vocabulary, and assign the label with the highest similarity between the center of gravity vector of the vocabulary vector and the vector of the label in the label list as the step label.
[0035] For example, when an utterance ID list is created as shown in FIG. 3, sentence 11 of utterance ID_1 included in the utterance ID list of step ID_1 is the beginning of the conversation and is a greeting sentence when starting to serve. Therefore, the conversation step recognition unit 102 assigns "start of conversation" as the step label of step ID_1. Furthermore, sentence 12 of utterance ID_2 included in the utterance ID list of step ID_2 is a sentence of inquiry from the customer, and sentence 13 of utterance is a sentence of apology from the operator. In other words, since sentences 12 and 13 are considered to be the beginning of understanding the specific content of the inquiry, the conversation step recognition unit 102 assigns "understanding the situation" as the step label of step ID_2.
[0036] In step S103, the utterance tag recognition unit 103 assigns an utterance tag to each utterance included in the conversation information. The utterance tag recognition unit 103 may classify each utterance using a natural language processing model, for example, and assign an utterance tag corresponding to the class to each utterance. Alternatively, the utterance tag recognition unit 103 may assign an utterance tag to each utterance using a natural language processing model that outputs a likely corresponding utterance tag when an utterance sentence is input. Alternatively, the utterance tag recognition unit 103 may vectorize the vocabulary in the utterance sentence and a predefined tag list using a natural language processing model that can vectorize vocabulary, and assign the tag with the highest similarity between the center of gravity vector of the vocabulary vector and the vector of a tag in the tag list as the utterance tag.
[0037] For example, in the case of the conversation information of Fig. 2, the sentence of the utterance ID_1 is a greeting sentence, as shown in Fig. 5. Therefore, the sentence of the utterance ID_1 is associated with the utterance tag "greeting." Similarly, the utterance ID_2 is associated with the utterance tag "question," and the utterance ID_3 is associated with the utterance tag "apology."
[0038] In step S104, the correspondence unit 104 associates the conversation step data with the utterance tag list.
[0039] In step S105, the output unit 105 performs processing for outputting information representing the conversation steps and the intention of the utterances in each conversation step to, for example, a display device, based on the associated conversation step data and utterance tag list. After that, the processing in FIG. 7 ends.
[0040] 7, the processing by the conversation step recognition unit 102 in steps S101 and S102 and the processing by the utterance tag recognition unit 103 in step S103 are performed in this order. On the other hand, the processing by the conversation step recognition unit 102 and the processing by the utterance tag recognition unit 103 may be performed in the reverse order or in parallel.
[0041] 8 is a diagram showing a display screen of a display device as an example of information output in the first embodiment. One example of the display screen displays a state transition diagram showing the intention of each utterance from the start of the utterance to the end of the utterance. The state transition diagram displays each utterance and its corresponding utterance tag on a two-dimensional plane, with the horizontal axis representing the change in utterance over time and the vertical axis representing the change in conversation step accompanying the change in utterance.
[0042] The state transition diagram includes an utterance ID 301, a conversation step 302, a conversation state node 303, a state transition arrow 304, and an utterance tag display 305. The utterance ID 301 is displayed on the horizontal axis and indicates the utterance ID of each of the L utterances from the beginning to the end of the conversation. The conversation step 302 is displayed on the vertical axis and indicates each of the M conversation steps from the beginning to the end of the conversation. The conversation steps 302 are displayed in the order in which they appear, starting from the bottom along the vertical axis. If the order in which the conversation steps occur can be defined, the conversation steps 302 may be displayed in a predefined order. The conversation state node 303 is a node that indicates the state of the conversation and represents the change in conversation step each time an utterance is made. The conversation state node 303 includes N nodes from the start state of the conversation to the end state of the conversation. The start state of the conversation is the state before the first utterance is made. The end state of the conversation is the state after the last utterance is made. Therefore, the number N of conversation state nodes 303 is one more than the number L of utterances. The state transition arrows 304 are directed arrows that represent transitions of conversation state nodes 303 due to utterances. The utterance tag displays 305 are displayed near the state transition arrows 304 and indicate the utterance tags of utterances made between the nearby conversation state nodes 303, i.e., the intention of the utterance.
[0043] As described above, according to the first embodiment, conversation steps and utterance tags are recognized from business conversations. Based on the recognition results, as shown in FIG. 8, conversation state transitions are visualized in association with conversation steps and utterance tags, allowing intuitive understanding of the situation, the intention of each utterance, and how the conversation progressed as a result. For example, in FIG. 8, it can be seen that after utterances intended as "questions" or "apologies" are made in a scene where a conversation is taking place to "understand the situation," the conversation transitions to the next conversation step. Furthermore, in FIG. 8, the conversation monotonically transitions from the lower left to the upper right on a two-dimensional plane, with no backtracking of conversation steps. From a state transition diagram such as FIG. 8, a user can intuitively understand that a good conversation is taking place with little unnecessary conversation. By providing feedback, etc. while presenting a state transition diagram such as FIG. 8 to the user, easy-to-understand feedback, etc. can be provided to the user. Furthermore, by recognizing conversation steps and utterance tags, the state of the conversation according to the type of conversation, etc., and the intention of each utterance are recognized. This allows flexible feedback, etc. to be provided according to the type of conversation, etc.
[0044] (Second embodiment) The information processing device according to the second embodiment further has a function of analyzing a conversation and visualizing the analysis results, in addition to the functions described in the first embodiment.
[0045] 9 is a block diagram showing an information processing device according to the second embodiment. The information processing device 100 according to the second embodiment includes a receiving unit 101, a conversation step recognition unit 102, an utterance tag recognition unit 103, a corresponding unit 104, an output unit 105, a speaker state recognition unit 106, an analysis unit 107, and a storage unit 108. Hereinafter, descriptions of elements similar to those described in the first embodiment will be omitted or simplified as appropriate.
[0046] The receiving unit 101 receives conversation information in the same manner as in the first embodiment. In the second embodiment, the conversation information includes meta-information. FIG. 10 is a diagram showing an example of conversation information in the second embodiment. As in the first embodiment, the conversation information includes text data of L (L is an integer equal to or greater than 2) utterance sentences 11, 12, ..., 1L associated with utterance IDs. Furthermore, the conversation information includes, as meta-information, for example, a speaker, a start time, and an end time. Furthermore, the conversation information includes audio data. The speaker is information indicating the speaker who made the utterance of the corresponding utterance ID. The start time is information indicating the start time of the utterance of the corresponding utterance ID. The end time is information indicating the end time of the utterance of the corresponding utterance ID. The audio data is recorded data of the utterance of the corresponding utterance ID.
[0047] The conversation step recognition unit 102 recognizes conversation steps from the conversation information and generates conversation step data based on the recognition results. In the second embodiment, the conversation step recognition unit 102 can recognize conversation steps not only from spoken sentences but also from meta-information. For example, the conversation step recognition unit 102 recognizes conversation steps based on predetermined rules for the meta-information. Alternatively, the conversation step recognition unit 102 may recognize conversation steps in the conversation information using, for example, a natural language processing model that can recognize conversation steps including meta-information.
[0048] The utterance tag recognition unit 103 recognizes utterance tags for each utterance in the conversation information and generates an utterance tag list based on the recognition results. In the second embodiment, the utterance tag recognition unit 103 can recognize utterance tags not only from sentences of an utterance but also from meta-information. For example, the utterance tag recognition unit 103 recognizes conversation steps based on predetermined rules for the meta-information. Alternatively, the utterance tag recognition unit 103 may recognize utterance tags using a natural language processing model that can recognize utterance tags including meta-information, for example.
[0049] The speaker state recognition unit 106 recognizes the speaker state for each utterance in the conversation information, and generates speaker state data based on the recognition result.
[0050] 11 is a diagram showing an example of speaker state data. The speaker state data indicates the speaker's state for each utterance, and includes L speaker state scores 31-3L associated with utterance IDs. Each speaker state score 41-4L includes, for example, scores for "anger level," "relief level," "satisfaction level," "tension level," and "level of understanding of explanation" as elements. The scores may be quantitative scores such as real numbers in the range [0, 1], or may be qualitative scores such as "high" or "low."
[0051] The analysis unit 107 analyzes the conversation of the analysis target speaker by referring to the conversation information, the conversation step data, the speech tag list, and the speaker state data, and generates analysis result data. The analysis target speaker is a speaker who is the subject of analysis among the speakers included in the conversation information. The analysis target speaker is designated, for example, in a predetermined order or by a user who evaluates the analysis target speaker.
[0052] 12 is a diagram showing an example of analysis result data. The analysis result data includes O analysis items each associated with an analysis ID, a result, and whether feedback is required 51, 52, 53, ..., 50.
[0053] The analysis items are O items among the items predetermined for analysis. The analysis items are specified by the user who will be evaluating the speaker to be analyzed. The analysis items do not necessarily have to include all of the predetermined items. The results are the results of the analysis for each analysis item. The feedback requirement is a flag indicating whether feedback to the speaker to be analyzed for each analysis item is necessary.
[0054] For example, Figure 12 shows the analysis items "Backtracking," "Speech Tag Ratio," "Required Time," and "Difference in Required Time." "Backtracking" refers to whether a backtracking occurred in a conversation step, and if so, the number of times it occurred. For example, in a "call center inquiry response" scenario, if a conversation for "explaining the situation" follows a conversation for "explaining the response," it is determined that a backtracking occurred. Such backtracking indicates inefficient operations. Here, if the same step label is permitted to be assigned to multiple conversation steps, consecutive conversation steps with the same step label may also be counted as "backtracking." "Speech Tag Ratio" refers to the ratio of speech tags registered in the speech tag list. "Required Time" refers to the time required for each conversation step transition. "Difference in Required Time" refers to the difference between the required time of a specified reference person, such as a veteran, and the "required time," or the difference between the average "required time" based on past analysis results and the current "required time." The results for each analysis item are compared with a predetermined criterion for determining whether or not feedback is necessary, and whether or not feedback is necessary is determined based on the comparison result.
[0055] Returning now to the description of Fig. 10, the storage unit 108 stores the conversation information, conversation step data, utterance tag list, speaker state data, and information on the speaker to be analyzed used in the analysis by the analysis unit 107 in association with the analysis result data.
[0056] As in the first embodiment, the correspondence unit 104 associates each utterance in the conversation step data with each utterance tag in the utterance tag list. Furthermore, in the second embodiment, the correspondence unit 104 further associates each utterance and utterance tag in the conversation step with a speaker state. The correspondence unit 104 associates each utterance in the conversation step data with the utterance tag and the speaker state data by referring to the utterance ID list in the conversation step data, the utterance ID in the utterance tag list, and the utterance ID in the speaker state data.
[0057] Next, the operation of the information processing device according to the second embodiment will be described. Fig. 13 is a flowchart showing the operation of the information processing device according to the second embodiment. In the following, the explanation will be continued assuming that conversation information in response to an inquiry at a call center shown in Fig. 10 is input to the information processing device. In addition, in the following, explanations that overlap with Fig. 7 will be omitted or simplified as appropriate.
[0058] In step S201, the conversation step recognition unit 102 recognizes conversation steps from the conversation information. In the second embodiment, the conversation information includes meta information. Therefore, the conversation step recognition unit 102 may recognize conversation steps taking the meta information into consideration. For example, in the case of a "one-on-one meeting for performance evaluation feedback from a superior to a subordinate," when the subordinate speaks after the superior speaks, it is highly likely that the subordinate is speaking to confirm the subordinate's understanding or questions after the feedback from the superior. Therefore, for the conversation information of the "one-on-one meeting for performance evaluation feedback from a superior to a subordinate," which includes speaker information, a rule may be established to recognize the interval between the superior's speech and the subordinate's speech as a step boundary. Furthermore, in the "response to a call center inquiry," if there is a gap between speeches of a certain time or more, it is considered that the operator is currently checking the situation. In this case, it is highly likely that the "explanation of the response" will begin from the next speech. Therefore, for the conversation information of the "response to a call center inquiry," which includes information on the start and end times of the speeches, a rule may be established to recognize the interval between speeches of a certain time or more as a step boundary. In the following, it is assumed that the same conversation steps as those in FIG. 3 are recognized in response to the input of the conversation information shown in FIG.
[0059] In step S202, the conversation step recognition unit 102 assigns a step label to each of the recognized conversation steps. The conversation step recognition unit 102 may assign step labels taking meta-information into consideration. For example, in the case of a "brainstorming workshop," if only a specific speaker speaks for a certain period of time, it is highly likely that "purpose explanation" or "summary of the workshop" is being performed. Therefore, for the conversation information of a "brainstorming workshop" that includes speaker information, if only a specific speaker speaks for a certain period of time, a rule may be established to assign either the step label "purpose explanation" or "summary of the workshop" to that conversation step. Which step label to assign can be determined by a method using the natural language processing model described above, or the like. Furthermore, in the case of a "research presentation at an academic conference," if two or more speakers are speaking, it is natural to assume that the speech is for "answering questions." Therefore, for the conversation information of "Shop for research presentations at academic conferences" that includes speaker information, if there are utterances from two or more speakers, a rule can be established that assigns a step label of "question answering" to that conversation step. Note that in the following, it is assumed that the same step labels as in Figure 3 are assigned to the input of the conversation information shown in Figure 10.
[0060] In step S203, the utterance tag recognition unit 103 assigns an utterance tag to each utterance included in the conversation information. The utterance tag recognition unit 103 may recognize the utterance tag taking meta information into consideration. For example, in the case of "responding to inquiries at a call center," the first and last utterances of the operator are likely to be "greetings." Therefore, for the conversation information of "responding to inquiries at a call center," which includes speaker information, a rule may be established to assign the utterance tag "greetings" to the first and last utterances of the operator. Furthermore, in the case of "research presentations at academic conferences," utterances by speakers other than the presenter are likely to be "questions." Therefore, for the conversation information of "research presentations at academic conferences," which includes speaker information, a rule may be established to assign the utterance tag "question" to utterances by speakers other than the presenter. Note that, hereinafter, it is assumed that the same utterance tags as those in FIG. 5 are recognized for the input of the conversation information shown in FIG. 10.
[0061] In step S204, the speaker state recognition unit 106 recognizes the speaker state for each utterance in the conversation information and generates speaker state data. The speaker state recognition unit 106 recognizes a score representing the degree of each speaker state item, such as anger level or relief level. For example, the speaker state recognition unit 106 may recognize the score of each item for the voice data input as conversation information as a regression problem using an acoustic processing model, or may recognize the score based on the difference when comparing acoustic features such as sound pressure, amplitude, and frequency with average values. The speaker state recognition unit 106 may also recognize the speaker state based on the vocabulary and polarity contained in the sentences of the conversation information using a natural language processing model, or may recognize the speaker state as a regression problem using a natural language processing model. Furthermore, the speaker state recognition unit 106 may recognize the speaker state using both the speaker state acoustic processing model and the natural language processing model simultaneously. The speaker state recognition unit 106 may also recognize the speaker state independently for each speaker using speaker information as meta-information. In the following, it is assumed that the same speaker state as in FIG. 11 is recognized in response to the input of conversation information shown in FIG.
[0062] In step S205, the analysis unit 107 analyzes the analysis target speaker based on the conversation information, conversation step data, conversation tag list, and speaker state data. For example, the analysis unit 107 analyzes the analysis target speaker's utterances for each analysis item and determines whether feedback is necessary based on the analysis results. The analysis unit 107 may analyze, for example, whether a relapse in conversation steps has occurred. If a relapse has occurred, the analysis unit 107 may evaluate the severity of the number of conversation steps involved in the relapse, or may identify the cause of the relapse by analyzing the utterances after the relapse. For example, in the case of "responding to inquiries at a call center," it may be analyzed that forgetting to confirm information is the cause of the relapse. The analysis unit 107 may also analyze the proportion of utterance tags in the utterance tag list. The analysis unit 107 may also analyze the duration of each conversation step based on the start time of the first utterance and the end time of the last utterance in each conversation step. Furthermore, the analysis unit 107 can analyze differences and biases in the required time by calculating the average of the required time for conversation steps in other conversation information stored in the storage unit 108. Similarly, the analysis unit 107 can analyze differences and biases in the proportions of utterance tags by calculating the average of the proportions of utterance tags in other conversation information stored in the storage unit 108. The analysis unit 107 may perform these analyses using a natural language processing model that can output the analysis results of a specific analysis item, or may perform these analyses using a natural language processing model that can evaluate multiple analysis items at once. After analyzing each analysis item, the analysis unit 107 generates analysis result data by associating the analysis item, the result, and whether feedback is required using an analysis ID.
[0063] For example, suppose that an analysis of analysis ID_1 was performed on the conversation step data shown in Figure 3 to determine whether or not a conversation step backtracking occurred. Three different step labels, "Start of conversation," "Understanding the situation," and "Explanation of response," are recognized in step ID_1, step ID_2, and step ID_3 in Figure 3. If any of "Start of conversation," "Understanding the situation," or "Explanation of response" is recognized in the conversation step of the next step ID, it can be determined that a conversation step backtracking occurred. The number of backtrackings can be determined by reconfirming the recognition of the step label starting from the utterance where the backtracking occurred. The results of analysis ID_1 store the analysis results for "backtracking." Furthermore, if feedback is required when backtracking occurs two or more times, for example, three backtrackings occurred in the example of Figure 12, so "required" is stored in the feedback requirement field for analysis ID_1.
[0064] Similarly, analysis is performed for other analysis items. In Figure 12, the results of the analysis of the difference in required time for analysis ID_O show that the time required for the conversation step "start responding" is 30 seconds shorter than the standard, and the time required for the conversation step "understand the situation" is 3 minutes longer than the standard. For example, if it is set to provide feedback when the difference in required time exceeds 2 minutes, the difference in required time for the conversation step "understand the situation" is 3 minutes, so "required" is stored in the feedback requirement field for analysis ID_O.
[0065] In step S206, the analysis unit 107 stores the conversation information, conversation step data, speech tag list, speaker state data, and information on the speaker to be analyzed used in the analysis in the storage unit 108 in association with the analysis result data.
[0066] In step S207, the correlation unit 104 correlates the conversation step data, the conversation tag list, and the speaker state data for the designated speaker to be analyzed.
[0067] In step S208, the output unit 105 performs processing to output, for example, to a display device, information indicating the conversation steps and the intention of the utterances in each conversation step, based on the associated conversation step data, utterance tag list, and speaker state data. After that, the processing in FIG. 13 ends.
[0068] 13, the processing by the conversation step recognition unit 102 in steps S201 and S202, the processing by the utterance tag recognition unit 103 in step S203, and the processing by the speaker state recognition unit 106 in step S204 are performed in this order. On the other hand, the processing by the conversation step recognition unit 102, the processing by the utterance tag recognition unit 103, and the processing by the speaker state recognition unit 106 may be performed in the reverse order or in parallel.
[0069] 14 is a diagram showing a display screen of a display device as a first example of information output in the second embodiment. The display screen of the first example displays a state transition diagram showing the intention of each utterance from the start of the utterance to the end of the utterance. Furthermore, in the second embodiment, the state transition diagram shows the speaker state.
[0070] In Figure 14, the state transition diagram includes an utterance ID 301, a conversation step 302, a conversation state node 303, a state transition arrow 304, and an utterance tag display 305. The utterance ID 301, the conversation step 302, the state transition arrow 304, and the utterance tag display 305 are the same as those shown in Figure 8, except that only those corresponding to the speaker to be analyzed are extracted. Meanwhile, with regard to the conversation state nodes 303, in addition to only those corresponding to the speaker to be analyzed being extracted, the speaker states of speakers other than the speaker to be analyzed are indicated by the darkness of the display. For example, if the speaker state is "anger," conversation state nodes 303 with a high level of "anger" are displayed dark, and conversation state nodes 303 with a low level of "anger" are displayed light.
[0071] For example, Fig. 14 shows an example of a conversation about "responding to inquiries at a call center." In Fig. 14, the call center operator is the speaker to be analyzed, and the customer is a speaker other than the speaker to be analyzed. Therefore, Fig. 14 displays an utterance ID 301, a conversation step 302, a state transition arrow 304, and an utterance tag display 305 for the operator's utterance.
[0072] Assuming that the speaker state data is as shown in FIG. 11 , the customer's anger level at the time of the customer's first utterance, utterance ID_2, is "0.5." This anger level corresponds to the customer's anger level at the time of receiving the operator's first utterance, utterance ID_1. Therefore, the display of the first conversation state node 303 is slightly darkened, corresponding to the anger level of "0.5." That is, in FIG. 14 , the darkness of each conversation state node 303 reflects the speaker states of speakers other than the speaker being analyzed immediately after the analysis. Even when the display is limited to the speaker being analyzed, the speaker states of speakers other than the speaker being analyzed are visualized in the conversation state nodes, allowing the user to check the speaker states of speakers other than the speaker being analyzed on the same screen. Furthermore, by reflecting the speaker state of the speaker being analyzed in the darkness of each conversation state node 303, it is possible to check, for example, whether the operator himself is always responding calmly.
[0073] 14, only the utterance IDs 301, conversation steps 302, conversation state nodes 303, state transition arrows 304, and utterance tag displays 305 corresponding to the speaker being analyzed are extracted. In contrast, as shown in FIG. 8, the utterance IDs 301, conversation steps 302, conversation state nodes 303, state transition arrows 304, and utterance tag displays 305 corresponding to all speakers may be displayed. In this case, too, the display intensity of each conversation state node 303 may be adjusted to correspond to the speaker state of the speaker in the immediately following utterance.
[0074] FIG. 15 is a diagram showing a display screen of a display device as a second example of information output in the second embodiment. The display screen of the second example visualizes the occurrence of conversation step backtracking. In the example of FIG. 15, a rectangular frame 306 is displayed to highlight the conversation state node where the conversation step backtracking occurred, the state transition arrow, and the utterance tag display, surrounding the conversation state node where the conversation step backtracking occurred, the state transition arrow, and the utterance tag display. By displaying such a rectangular frame 306, the user can intuitively confirm which utterance tag in which scene caused the conversation step backtracking. Such confirmation is useful when the user considers how to improve the conversation. Instead of displaying the rectangular frame 306, other highlighting methods may be used, such as changing the color of the utterance tag display 305 of the utterance where the backtracking occurred.
[0075] Here, the display of Fig. 15 may be displayed when at least one conversation step back occurs, or may be displayed only when it is determined that feedback is necessary. Furthermore, in addition to the display of Fig. 15, if the cause of the backtracking has been identified, information about the cause may also be displayed. Furthermore, in addition to the display of Fig. 15, a display showing the speaker state shown in Fig. 14 may also be displayed.
[0076] FIG. 16 is a diagram showing a display screen of a display device as a third example of information output in the second embodiment. The display screen of the third example visualizes the difference in required time. In the example of FIG. 16, a rectangular frame 307 is displayed to surround conversation steps with a large difference in required time in order to highlight conversation steps with a large difference in required time. A conversation step with a large difference in required time is, for example, a conversation step determined to require feedback. In FIG. 16, a rectangular frame 307 is displayed around the conversation step "understanding the situation." By displaying such a rectangular frame 307, the user can intuitively confirm which conversation step's utterances require attention. This confirmation is also useful when the user considers ways to improve the conversation. Instead of displaying the rectangular frame 307, other highlighting methods may be used, such as changing the color of conversation steps 302 with a large difference in the predetermined time.
[0077] Here, in addition to the display of FIG. 16, the display showing the speaker state shown in FIG. 14 and / or the display showing the occurrence of a backtrack shown in FIG. 15 may also be displayed.
[0078] As described above, in the second embodiment, in addition to visualizing the correspondence between conversation steps and utterance tags, an analysis of one's own conversation and / or an analysis comparing it with past conversations are performed. By outputting the results of these analyses, the user can more intuitively grasp the utterances that need to be confirmed.
[0079] Here, when it is determined that feedback is necessary, the state transition diagram of the speaker being analyzed and the state transition diagram of a reference speaker, such as an experienced speaker, may be displayed superimposed on the same screen. This display allows the user to intuitively grasp the differences between the speaker and the reference speaker. This display is also useful when the user is considering ways to improve the conversation.
[0080] (Third embodiment) The information processing device according to the third embodiment has, in addition to the functions described in the second embodiment, a function of creating learning material content for conversation learning from past analysis results.
[0081] 17 is a block diagram showing an information processing device according to a third embodiment. The information processing device 100 according to the third embodiment includes a receiving unit 101, a conversation step recognition unit 102, an utterance tag recognition unit 103, a speaker state recognition unit 106, an analysis unit 107, a storage unit 108, a correspondence unit 104, and an output unit 105, as well as a recommendation unit 109 and a learning material creation unit 110. Hereinafter, descriptions of elements similar to those described in the second embodiment will be omitted or simplified as appropriate. Specifically, the receiving unit 101, the conversation step recognition unit 102, the utterance tag recognition unit 103, the speaker state recognition unit 106, the analysis unit 107, the storage unit 108, the correspondence unit 104, and the output unit 105 are similar to those in the second embodiment, and therefore descriptions thereof will be omitted.
[0082] The recommendation unit 109 searches the storage unit 108 for conversation information to recommend as learning material content, based on the conversation information received from the receiving unit 101. Then, the recommendation unit 109 inputs information indicating the searched conversation information and the conversation information received from the receiving unit 101 to the learning material creation unit 110. The information indicating the conversation information may be a conversation ID assigned to each piece of conversation information stored in the storage unit 108.
[0083] The learning material creation unit 110 acquires conversation information, conversation step data, an utterance tag list, and a speaker state from the storage unit 108 based on information indicating the conversation information searched for by the recommendation unit 109. The learning material creation unit 110 also acquires, from the storage unit 108, analysis result data for the conversation information received from the receiving unit 101. The learning material creation unit 110 then creates learning material content data based on the conversation information, conversation step data, utterance tag list, speaker state, and analysis result data acquired from the storage unit 108. FIG. 18 is a diagram showing an example of learning material content data. The learning material content data includes I conversation IDs associated with question IDs, question utterance IDs, and correct answers 61, 62, 63, ..., 6I. The conversation IDs are IDs of the conversation information used for the question of the corresponding question ID. The question utterance IDs are utterance IDs used for the question of the corresponding question ID. The correct answers are utterance tags that indicate the correct answer for the corresponding question ID.
[0084] In the third embodiment, the correspondence unit 104 associates conversation information with learning material content by referring to the utterance ID of the conversation information and the question utterance ID of the learning material content.
[0085] In the third embodiment, the output unit 105 uses the conversation information and the learning material content associated by the association unit 104 to perform processing for outputting questions for learning speaking to, for example, a display device.
[0086] Next, the operation of the information processing device according to the third embodiment will be described. Fig. 19 is a flowchart showing the operation of the information processing device according to the third embodiment. In the following, explanations that overlap with those in Fig. 13 will be omitted or simplified as appropriate. That is, the processing of steps S301-S306 in Fig. 19 is the same as the processing of steps S201-S206 in Fig. 13. However, the conversation information input to the information processing device 100 in steps S301-S306 in Fig. 19 is conversation information of the user as a participant in speech learning.
[0087] In step S307 after the conversation information received from the receiving unit 101, the conversation step data recognized from the conversation information, the utterance tag list, the speaker state data, and the analysis result data analyzed from these are stored, the recommendation unit 109 searches the storage unit 108 for conversation information to recommend as learning material content based on the conversation information received from the receiving unit 101. The recommendation unit 109 then inputs information indicating the searched conversation information to the learning material creation unit 110. The recommendation unit 109 searches the storage unit 108 for conversation information having text similar to the text of the conversation information received from the receiving unit 101, for example. The search for similar conversation information can be performed based on the similarity between the conversation information from the receiving unit 101, which has been converted into vectors using a natural language processing model that can convert text into vectors, and the conversation information stored in the storage unit 108. The search for similar conversation information may also be performed based on the matching rate between keywords extracted from the text of the conversation information in the receiving unit 101 and keywords extracted from the text of the conversation information in the storage unit 108, using parts of speech and dependency relationships using a morphological analyzer and a syntactic analyzer. The search for similar conversation information may also be performed based on the word matching rate in summarized text of the conversation information generated by a natural language processing model capable of generating a conversation summary. Here, in a search based on similarity, a threshold for determining how many conversation information items with the highest similarity will be adopted must be defined in advance.
[0088] Here, the search for conversation information to be recommended as learning material content does not necessarily have to be based on the similarity of the text of the conversation information. For example, the search may be based on the similarity of meta information of the conversation information, such as the speaker and the date and time of the conversation, or the similarity of utterance tags and / or the similarity of transitions of conversation steps.
[0089] In step S308, the learning material creation unit 110 acquires conversation information, conversation step data, an utterance tag list, and a speaker state from the storage unit 108 based on information indicating the conversation information searched for by the recommendation unit 109. The learning material creation unit 110 also acquires, from the storage unit 108, analysis result data for the conversation information received from the receiving unit 101. The learning material creation unit 110 then creates learning material content data based on the conversation information, conversation step data, utterance tag list, speaker state, and analysis result data acquired from the storage unit 108. The learning material creation unit 110 sets, as learning items, analysis items for which feedback has been determined to be "needed" with reference to the analysis result data. The learning material creation unit 110 then identifies areas for improvement in the conversation information received from the receiving unit 101 for each learning item. For example, if the feedback for "backtracking" is "needed," the learning material creation unit 110 identifies the utterance in which the backtracking occurred and the utterance that caused it as areas for improvement. For example, in a situation where a backtracking is determined to have occurred due to a failure to confirm information, the utterance of the target speaker where the backtracking occurred is often associated with a "question" tag. Therefore, the utterance tagged with "question" before the backtracking, i.e., the utterance with utterance ID_5, may be identified as an area for improvement. Furthermore, for example, if the feedback for "difference in required time" is "needed," unnecessary utterances in the corresponding conversation step are identified as areas for improvement. For example, since the utterance of the target speaker in the conversation step "understanding the situation" is usually associated with a "question" tag, utterances associated with tags other than the "question" tag may be identified as areas for improvement. After identifying areas for improvement, the learning material creation unit 110 identifies utterances similar to the identified areas for improvement from the conversation information recommended by the recommendation unit 109. Identification of similar utterances may be performed based on the similarity between the utterances in the areas for improvement, each converted into a vector using a natural language processing model capable of converting text into a vector, and the utterances in the conversation information recommended by the recommendation unit 109. In addition, similar utterances may be identified based on the matching rate between keywords extracted from the utterance to be improved and keywords extracted from the text of the conversation information recommended by the recommendation unit 109, using parts of speech and dependency relationships using a morphological analyzer and a syntactic analyzer.Alternatively, similar utterances may be identified based on the degree of agreement or similarity between the utterance tag assigned to the utterance in the improvement area and the transition between the utterance tags before and after it. After identifying utterances similar to the utterance in the improvement area, the learning material creation unit 110 creates learning material content data by associating the conversation ID of the conversation information recommended by the recommendation unit 109 with the utterance ID of the utterance similar to the utterance in the improvement area included in this conversation information and the utterance tag associated with this utterance, and then assigning a unique question ID to the utterance.
[0090] In step S309, the correspondence unit 104 associates the learning material content data created by the learning material creation unit 110 with the conversation information stored in the storage unit 108. The correspondence unit 104 acquires conversation information of the conversation IDs associated with each question ID from the storage unit 108. Then, the correspondence unit 104 associates the question ID with an utterance ID that occurs before the question utterance ID associated with each question ID in the conversation information acquired from the storage unit 108. For example, if the fifth utterance has the question utterance ID, the correspondence unit 104 associates the first to fifth utterances in the corresponding conversation information with the corresponding question ID.
[0091] In step S310, the output unit 105 performs processing to output questions for utterance learning to, for example, a display device, using the conversation information and the learning material content data associated by the association unit 104. Thereafter, the processing in FIG. 19 ends.
[0092] 20 is a diagram showing a display screen of a question on a display device as an example of question output in the third embodiment. In one example, the display screen of the question has a status display area 310, a question display area 320, and an answer history display area 330.
[0093] The situation display area 310 is a display area for allowing the user to recognize the situation leading up to the conversation in the question. The situation display area 310 is formed based on the conversation information associated with the learning material content data of the corresponding question ID. The situation display area 310 includes a state transition diagram 311 and a conversation text display 312. The state transition diagram 311 is a state transition diagram up to the utterance before the question utterance ID of the corresponding question ID. The conversation state nodes displayed in the state transition diagram 311 may be only the conversation state nodes corresponding to the speaker to be analyzed, or may be all the conversation state nodes. Furthermore, the speaker state may also be displayed in the conversation state node. The conversation text display 312 displays the text of the utterance up to the utterance before the question utterance ID of the corresponding question ID.
[0094] The question display area 320 is a display area for the question. The output unit 105 creates question sentences according to the type of improvement area. For example, if a backtrack occurred due to forgetting to confirm information and an utterance tagged with "Confirm" before the backtrack occurred is identified as an improvement area, the output unit 105 generates a question sentence asking the user as a student what should be confirmed. Furthermore, if the time required for a conversation step is long due to unnecessary utterances, the output unit 105 generates a question sentence asking what kind of utterance should be made next in the improvement area. In FIG. 20 , a question sentence asking what kind of utterance should be made next is displayed. This question sentence further displays options for utterance tags indicating the intention of appropriate utterances. These options include a correct utterance tag associated with the corresponding question ID and an utterance tag other than the correct answer. The utterance tags other than the correct answer may be displayed by those associated with the question ID, or may be displayed randomly each time the question sentence is displayed. When the user selects one of the options, the output unit 105 compares the speech tag of the option selected by the user with the speech tag of the correct answer to determine whether the question is correct or incorrect, and notifies the user of the result. Note that instead of options, the user may actually speak. In this case, the speech tag recognition unit 103 recognizes speech tags from the speech input by the user. The output unit 105 compares the speech tag recognized by the speech tag recognition unit 103 with the speech tag of the correct answer to determine whether the question is correct or incorrect, and notifies the user of the result. Here, in FIG. 20, one question is displayed on one screen, and once the answer is completed, the next question is displayed. However, multiple questions may be displayed simultaneously on one screen, or the questions may be displayed so that they can be switched using tabs or the like.
[0095] The answer history display area 330 is an area for displaying the user's answer history. The answer history display area 330 may display an answer history 331. The answer history 331 may display, for example, the user's answer for each question, the correct answer, and a check box for whether to request an explanation. The user's answer is an answer given by the user in the past. The correct answer is the correct answer for the question. Instead of the correct answer, a correct / incorrect result for the question may be displayed. The check box is checked when the user requests an explanation for the correct answer for the question. For a checked question, the output unit 105 displays a preset text for explaining the question, for example, after the user has answered all the questions. In addition, the output unit 105 may request feedback from a benchmark such as a veteran by email or the like.
[0096] As described above, in the third embodiment, learning material content data is created based on conversation information accumulated in the past and the analysis results of the user's conversation information. The past conversation information used to create the learning material content data is then recommended based on areas of improvement in the student's own conversation. Conversation-related questions are posed to the user based on this learning material content data. This allows the student to acquire skills efficiently. Furthermore, because questions are based on actual conversations, it is expected that the student will also learn tacit knowledge that is not included in the analysis items. Furthermore, by using speech tags for the correct answers to the questions, it is possible to perform accuracy determination that is less affected by individual differences in spoken sentences, such as vocabulary and phrasing.
[0097] The instructions shown in the processing procedures described in the above-described embodiments can be executed based on a software program. A general-purpose computer system can store this program in advance and, by loading this program, achieve effects similar to those of the information processing device described above. The instructions described in the above-described embodiments are recorded as a computer-executable program on a magnetic disk (flexible disk, hard disk, etc.), an optical disk (CD-ROM, CD-R, CD-RW, DVD-ROM, DVD±R, DVD±RW, Blu-ray (registered trademark) Disc, etc.), semiconductor memory, or similar recording medium. The recording medium may take any storage format as long as it is readable by a computer or embedded system. A computer can achieve operations similar to those of the information processing device described in the above-described embodiments by loading the program from the recording medium and having the CPU execute the instructions described in the program based on the program. Of course, the computer may acquire or load the program via a network. In addition, an OS (operating system), database management software, network middleware, etc. running on a computer may execute some of the processes required to realize this embodiment based on instructions from a program installed on the computer or embedded system from a recording medium. Furthermore, the recording medium in this embodiment is not limited to a medium independent of a computer or an embedded system, but also includes a recording medium that stores or temporarily stores a program downloaded via a LAN, the Internet, or the like. Furthermore, the number of recording media is not limited to one, and cases where the processing in this embodiment is executed from multiple media are also included in the recording media in this embodiment, and the media may have any configuration.
[0098] The computer or embedded system in this embodiment is for executing each process in this embodiment based on a program stored on a recording medium, and may be configured as either a device consisting of a single device such as a personal computer or a microcomputer, or a system in which multiple devices are connected to a network. Furthermore, the computer in this embodiment is not limited to a personal computer, but also includes an arithmetic processing unit, a microcomputer, etc. included in information processing equipment, and is a general term for equipment or devices that can realize the functions in this embodiment by a program.
[0099] Although several embodiments of the present invention have been described, these embodiments are presented as examples and are not intended to limit the scope of the invention. These novel embodiments can be embodied in various other forms, and various omissions, substitutions, and modifications can be made without departing from the spirit of the invention. These embodiments and their modifications are included within the scope and spirit of the invention, and are also included in the scope of the invention and its equivalents as defined in the claims. [Explanation of symbols]
[0100] 100 Information processing device, 101 Receiving unit, 102 Conversation step recognition unit, 103 Speech tag recognition unit, 104 Corresponding unit, 105 Output unit, 106 Speaker state recognition unit, 107 Analysis unit, 108 Storage unit, 109 Recommendation unit, 110 Teaching material creation unit, 201 Processor, 202 Memory, 203 Input device, 204 Display device, 205 Communication device, 206 Storage, 207 Bus, 2061 Information processing program.
Claims
1. An information processing program for supporting learning of tasks performed through conversation, Based on conversation information including information about utterances made in a conversation between speakers, recognizing conversation steps that may occur from the beginning to the end of the conversation, the conversation steps being composed of one or more semantically consistent utterances; Recognizing an utterance tag that is a tag indicating the intention of each utterance in the conversation related to the conversation information; Associating utterances included in each of the conversation steps with the utterance tags; An information processing program for causing a processor to execute the above.
2. displaying, on a display device, each utterance included in the conversation information, a conversation step corresponding to each utterance, and an utterance tag corresponding to each utterance; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
3. In displaying the utterances, the conversation steps, and the utterance tags on the display device, the utterances and the utterance tags are displayed on a two-dimensional plane in which the horizontal axis represents changes in the utterances over time and the vertical axis represents changes in the conversation steps associated with changes in the utterances; 3. The information processing program according to claim 2, further causing the processor to execute the following:
4. displaying the conversation step and the utterance tag on the display device, a node representing a conversation state and a directed arrow representing a transition of the conversation state accompanying the utterance; 4. The information processing program according to claim 3, further causing the processor to execute the following:
5. conducting an analysis to evaluate whether improvement is necessary for one or more analysis items from the conversation information, the conversation steps recognized from the conversation information, and the utterance tags; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
6. Displaying, on a display device, each utterance included in the conversation information, a conversation step corresponding to each utterance, and an utterance tag corresponding to each utterance; highlighting an area including the speech tag corresponding to an analysis item that needs improvement in the analysis; 6. The information processing program according to claim 5, further causing the processor to execute the following:
7. Displaying, on a display device, each utterance included in the conversation information, a conversation step corresponding to each utterance, and an utterance tag corresponding to each utterance; highlighting an area including the conversation step corresponding to an analysis item that needs improvement in the analysis; 6. The information processing program according to claim 5, further causing the processor to execute the following:
8. conducting an analysis to detect backtracking of said conversational steps; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
9. conducting an analysis to calculate the time taken for said conversation step transitions; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
10. performing an analysis to compare first conversation steps and first utterance tags recognized from first conversation information made by a first speaker with second conversation steps and second utterance tags recognized from one or more second conversation information of the same type as the first conversation information made by a second speaker different from the first speaker; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
11. recognizing a speaker state indicating a speaker state for each utterance related to the conversation information; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
12. Displaying, on a display device, each utterance included in the conversation information, a conversation step corresponding to each utterance, and an utterance tag corresponding to each utterance; displaying a speaker status indicating the status of the speaker of each of the utterances; The information processing program according to claim 11 , further causing the processor to execute the following:
13. recommending, in response to first conversation information, second conversation information that is different from the first conversation information and similar in content to the first conversation information; 2. The information processing program according to claim 1, further comprising causing the processor to execute the following:
14. extracting from the second conversation information a portion of the first conversation information that needs improvement; creating teaching materials for the speaker of the first conversation information from the extracted passages; The information processing program according to claim 13 , further causing the processor to execute the following:
15. An information processing device that supports learning of tasks performed through conversation, a conversation step recognition unit that recognizes conversation steps that may occur from the start to the end of a conversation between speakers, the conversation steps being composed of one or more semantically consistent utterances, based on conversation information including information on utterances made in the conversation between the speakers; an utterance tag recognition unit that recognizes an utterance tag that is a tag indicating the intention of each utterance in the conversation related to the conversation information; a correspondence unit that associates utterances included in each of the conversation steps with the utterance tags; An information processing device comprising:
Citation Information
Patent Citations
Learning support method and apparatus
JP4400052B2
Computer-assisted conversation support method
JP7231894B1