Visualization information generation device, visualization information generation method, and program
The visualization information generation device addresses the challenge of manual keyword setting by automatically analyzing and visualizing conformance of operator utterances to talk scripts, enhancing the efficiency of script adherence verification.
Patent Information
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- NIPPON TELEGRAPH & TELEPHONE CORP
- Filing Date
- 2021-12-22
- Publication Date
- 2026-04-14
AI Technical Summary
Existing methods for verifying operator utterances against a talk script require manual setting of comparison keywords, which is costly and difficult to implement, especially when the script is in sentence form.
A visualization information generation device that analyzes the conformance of operator utterances to a talk script by generating visualization information on the ranges of conformance and non-conformance, using a compliance estimation processing unit to split, match, and visualize the utterance text against the script.
Enables efficient estimation of whether operator utterances conform to the talk script, providing visualization of compliance and non-compliance ranges, and offering proposed revisions for improvement.
Smart Images

Figure 0007846134000001 
Figure 0007846134000002 
Figure 0007846134000003
Abstract
Description
Technical Field
[0001] The present invention relates to a visualization information generation device, a visualization information generation method, and a program.
Background Art
[0002] In a contact center (or also called a call center), generally, an operator determines a talk script when dealing with a customer so that there is no difference in customer service among operators. Here, the talk script refers to the content of speech, the speech procedure, etc. determined in the contact center. In the talk script, for example, it is necessary to speak sentences, keywords, phrases, etc. defined for each item or each scene such as the first greeting (opening), the inquiry content, customer identification (name, date of birth, etc.), response, and the last greeting (closing).
[0003] Also, in order for an administrator to check whether each operator has appropriately dealt with a customer, for example, the administrator checks the recording of the voice call between the operator and the customer, or takes a questionnaire from the customer and analyzes the results. For example, there is known a technique for estimating the appropriateness of an operator's response to a customer by comparing the text obtained by voice recognition of the voice call between the operator and the customer with a predetermined keyword (Patent Document 1).
Prior Art Documents
Patent Documents
[0004]
Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0005] However, when it is necessary to verify whether an operator's utterance conforms to a talk script, prior art such as Patent Document 1 requires manual setting of comparison keywords for each item in the talk script, incurring setting costs. Furthermore, when the talk script is expressed in sentences (for example, when the talk script is in a script format consisting of sentences representing the content of the operator's utterances), it can be difficult to set keywords to appropriately verify whether or not these sentences were uttered.
[0006] One embodiment of the present invention has been made in view of the above points and aims to estimate whether or not it conforms to a talk script. [Means for solving the problem]
[0007] To achieve the above objective, a visualization information generation device according to one embodiment includes a visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance. [Effects of the Invention]
[0008] It is possible to estimate whether or not the talk script is being followed. [Brief explanation of the drawing]
[0009] [Figure 1] This figure shows an example of the overall configuration of the contact center system according to this embodiment. [Figure 2] This figure shows an example of the hardware configuration of the estimation device according to this embodiment. [Figure 3] This figure shows an example of the functional configuration of the estimation device according to this embodiment. [Figure 4] This is a diagram (part 1) showing an example of a talk script. [Figure 5] This is a diagram (part 2) showing an example of a talk script. [Figure 6] This is a diagram (part 3) showing an example of a talk script. [Figure 7] This is a diagram (number 4) showing an example of a talk script. [Figure 8] This figure shows an example of the detailed functional configuration of the compliance estimation processing unit according to this embodiment. [Figure 9] This figure shows an example of a processing flow for saving compliance history and visualizing the scope of compliance and non-compliance. [Figure 10] This is a diagram (part 1) illustrating an example of how correspondence information is generated. [Figure 11] This is a diagram (part 2) illustrating an example of the generation of correspondence information. [Figure 12] This is a diagram (part 3) illustrating an example of the generation of correspondence information. [Figure 13] This is a diagram (part 4) illustrating an example of the generation of correspondence information. [Figure 14] This is a diagram showing an example of compliance history. [Figure 15] This figure shows an example of a conformance history when multiple utterances are integrated. [Figure 16] This is Figure (1) showing an example of the visualization results of the scope of compliance and non-compliance. [Figure 17] This is Figure (2) showing an example of the visualization results of the scope of compliance and non-compliance. [Figure 18] This diagram shows an example of a processing flow when visualizing compliance status. [Figure 19] This figure shows an example of the visualization results of compliance status. [Figure 20] This figure shows an example of a processing flow when visualizing proposed revisions, compliance rates, operator utterances, and related information. [Figure 21] This figure shows an example of a compliance history that combines call evaluations and related information. [Figure 22] This is a diagram (part 1) showing an example of the visualization results of the revised proposal. [Figure 23]This is a diagram (part 2) showing an example of the visualization results of the revised proposal. [Figure 24] This is Figure (1) showing an example of the visualization results of the compliance rate. [Figure 25] This is a diagram (part 1) showing an example of the visualization results of the operator utterance list. [Figure 26] This figure shows an example of the visualization results of related information. [Figure 27] This is Figure (2) showing an example of the visualization results of the compliance rate. [Figure 28] This is Figure (2) showing an example of the visualization results of the operator utterance list. [Figure 29] This is Figure (3) showing an example of the visualization results of the operator utterance list. [Figure 30] This is Figure (4) showing an example of the visualization results of the operator utterance list. [Modes for carrying out the invention]
[0010] One embodiment of the present invention will be described below. In this embodiment, a contact center system 1 will be described that includes an estimation device 10 capable of estimating whether the operator's utterances when responding to customer inquiries conform to a talk script, targeting contact center operators.
[0011] However, contact centers are just one example; the same method can be applied to other situations, such as sales representatives for products or services, or store counter staff, when trying to estimate whether their utterances conform to a talk script (or equivalent conversation manual or script). More generally, it can also be applied to individuals who engage in conversations with one or more people, when trying to estimate whether their utterances conform to a talk script (or equivalent conversation manual or script).
[0012] In the following explanation, contact center operators primarily perform tasks such as handling inquiries with customers via voice calls. However, this is not limited to this, and the same principles can be applied to cases where work is performed via text chat (including those that allow sending and receiving stamps, attachments, etc., in addition to text), video calls, etc.
[0013] <Overall configuration of Contact Center System 1> Figure 1 shows the overall configuration of the contact center system 1 according to this embodiment. As shown in Figure 1, the contact center system 1 according to this embodiment includes an estimation device 10, an operator terminal 20, a supervisor terminal 30, a PBX (Private Branch eXchange) 40, and a customer terminal 50. Here, the estimation device 10, the operator terminal 20, the supervisor terminal 30, and the PBX 40 are installed in the contact center environment E, which is the system environment of the contact center. Note that the contact center environment E is not limited to a system environment in the same building, but may be, for example, a system environment in multiple buildings geographically separated.
[0014] The estimation device 10 estimates whether the operator's utterances when responding to customer inquiries conform to the talk script. The estimation device 10 is a general-purpose server and other devices that visualize various information on the operator terminal 20 and supervisor terminal 30 based on the estimation results.
[0015] The operator terminal 20 is a type of terminal, such as a PC (personal computer), used by an operator to handle customer inquiries, and functions as an IP (Internet Protocol) telephone. The operator terminal 20 may also be, for example, a smartphone, tablet, or wearable device.
[0016] The supervisor terminal 30 is a PC or other type of terminal used by an administrator (such an administrator is also called a supervisor) who manages the operators. The supervisor terminal 30 may also be, for example, a smartphone, tablet, or wearable device.
[0017] PBX40 is a telephone exchange (IP-PBX) connected to a communication network 60, which includes a VoIP (Voice over Internet Protocol) network and a PSTN (Public Switched Telephone Network). Note that PBX40 may also be a cloud-based PBX (i.e., a general-purpose server providing call control services as a cloud service).
[0018] Customer terminals 50 are various devices used by customers, such as smartphones, mobile phones, and landline phones.
[0019] Note that the overall configuration of the contact center system 1 shown in Figure 1 is just one example, and other configurations are possible. For example, in the example shown in Figure 1, the estimation device 10 is included in the contact center environment E (i.e., the estimation device 10 is on-premise), but all or part of the functions of the estimation device 10 may be implemented by cloud services or the like. Also, although the operator terminal 20 is shown to function as an IP phone, for example, a telephone separate from the operator terminal 20 may be included in the contact center system 1.
[0020] <Hardware configuration of estimation device 10> Figure 2 shows the hardware configuration of the estimation device 10 according to this embodiment. As shown in Figure 2, the estimation device 10 according to this embodiment is implemented with the hardware configuration of a general computer or computer system and includes an input device 101, a display device 102, an external I / F 103, a communication I / F 104, a processor 105, and a memory device 106. Each of these hardware components is connected to each other via a bus 107 for communication.
[0021] The input device 101 is, for example, a keyboard, mouse, or touch panel. The display device 102 is, for example, a display. The estimation device 10 does not necessarily have to have at least one of the input device 101 and the display device 102.
[0022] External I / F 103 is an interface to external devices such as recording media 103a. The estimation device 10 can read from and write to the recording media 103a via the external I / F 103. Examples of recording media 103a include CD (Compact Disc), DVD (Digital Versatile Disc), SD memory card (Secure Digital memory card), and USB (Universal Serial Bus) memory card.
[0023] The communication interface 104 is an interface for the estimation device 10 to communicate with other devices and equipment. The processor 105 is, for example, a CPU (Central Processing Unit) or a GPU (Graphics Processing Unit), among other types of computing devices. The memory device 106 is, for example, a HDD (Hard Disk Drive), an SSD (Solid State Drive), RAM (Random Access Memory), ROM (Read Only Memory), or flash memory, among other types of storage devices.
[0024] The estimation device 10 according to this embodiment can perform various processes described later by having the hardware configuration shown in Figure 2. Note that the hardware configuration shown in Figure 2 is just one example, and the estimation device 10 may have other hardware configurations. For example, the estimation device 10 may have multiple processors 105 or multiple memory devices 106.
[0025] <Functional configuration of the estimation device 10> Figure 3 shows the functional configuration of the estimation device 10 according to this embodiment. As shown in Figure 3, the estimation device 10 according to this embodiment includes a speech recognition unit 201, a conformance estimation processing unit 202, and a storage unit 203. The speech recognition unit 201 and the conformance estimation processing unit 202 are implemented, for example, by processing that one or more programs installed in the estimation device 10 cause the processor 105 to execute. The storage unit 203 is implemented, for example, by a memory device 106. The storage unit 203 may also be implemented, for example, by a storage device connected to the estimation device 10 via a communication network.
[0026] The speech recognition unit 201 converts the voice call between the operator and the customer into text using speech recognition. The speech recognition unit 201 may also remove fillers (such as "um," "uh," "uh," etc.) included in the voice call. Hereafter, such text will also be referred to as "speech text." The speech text may be a transcription of both the operator's and the customer's voices, or a transcription of only the operator's voice. In the following, we will primarily assume that the speech text is a transcription of only the operator's voice, with fillers removed.
[0027] This embodiment assumes a voice call between a contact center operator and a customer, and therefore assumes two speakers, but it is not limited to this. For example, this embodiment can be applied to cases with three or more speakers. However, in this case, the talk script must be designed to accommodate speech between three or more people. Furthermore, the relationship between the speakers is not limited to an operator and a customer. Moreover, the speakers are not necessarily limited to humans; at least some of the speakers among multiple speakers may be robots or agents, etc.
[0028] The compliance estimation processing unit 202 estimates whether the operator's utterance conforms to the talk script based on the utterance text and the talk script. The compliance estimation processing unit 202 also visualizes various information on the operator terminal 20 and supervisor terminal 30 based on the estimation result. This information includes, for example, the extent to which the operator's utterance conforms to (or does not conform to) the talk script, the compliance status of each operator, proposed revisions to the talk script or utterances, each operator's compliance rate, each operator's utterances, and related information concerning inquiries in calls from which utterance text was obtained. The detailed functional configuration of the compliance estimation processing unit 202 will be described later.
[0029] The memory unit 203 stores information such as utterance text, talk scripts, and compliance history. The compliance history, as will be described later, refers to historical information indicating whether each operator utterance conforms to the talk script.
[0030] In the example shown in Figure 3, the estimation device 10 is assumed to have a speech recognition unit 201. However, if, for example, no voice calls are made between the operator terminal 20 and the customer terminal 50, and only text chats are conducted, the estimation device 10 does not need to have a speech recognition unit 201.
[0031] <Talk Script> As mentioned above, a talk script refers to the content and procedures of speech determined by the contact center. Several specific examples of talk scripts are described below. However, the talk scripts described below are all examples, and this embodiment can be applied to any talk script. Talk scripts often specify the sentences, content, keywords, or key phrases that the operator needs to speak, but in addition to these, they may also specify sentences, content, keywords, or key phrases that the customer is expected to speak, or they may further specify the necessary operating procedures for speaking (for example, operating procedures for FAQ searches).
[0032] ≪Example of a Talk Script 1≫ In the talk script shown in Figure 4, each item representing a scene or similar situation has a script that specifies the sentences the operator needs to speak in that particular situation.
[0033] For example, the item "Opening Greeting" has a script that reads "Thank you for calling...". This means that the operator must say a phrase like "Thank you for calling..." as their opening greeting. The same applies to the other items: "Confirming Inquiry Details," "Customer Identification (Name, Date of Birth, etc.)," "Handling the Call," and "Closing Greeting."
[0034] In Figure 4, the talk script shows that the inquiry process proceeds in the following order: "Initial greeting (opening)", "Confirmation of inquiry content", "Customer identity verification (name, date of birth, etc.)", "Response", and "Final greeting (closing)".
[0035] ≪Example of a Talk Script 2≫ In the talk script shown in Figure 5, similar to Figure 4, each item representing a scene or similar situation specifies the sentences, utterances, keywords, or phrases that the operator must speak in that item. The number of turns for each item is also defined. A turn represents the exchange of utterances between the customer and the operator. For example, "one turn" refers to the customer speaking in response to the operator's utterance, or the operator speaking in response to the customer's utterance.
[0036] For example, the item "Opening" has a script (Example 1) that includes phrases like "Thank you for calling." This indicates, as in Figure 4, that the operator is required to say a sentence such as "Thank you for calling" during the opening.
[0037] Furthermore, for example, the item "Opening" has a script (Example 2) that specifies "Express gratitude." This indicates that the operator needs to express gratitude during the opening (for example, "Thank you very much," or "Thank you very much," etc.).
[0038] Furthermore, for example, the "Opening" section includes "Telephone" and "Thank you" as script examples (Example 3). This indicates that operators are required to include keywords (or phrases) such as "Telephone" and "Thank you" in their opening remarks.
[0039] Furthermore, the item "Opening" is defined as "the first 3 turns," meaning that the first 3 turns of the inquiry process constitute the opening.
[0040] The same applies to the other items: "Customer Verification," "Identity Verification," "Verification of a Phone Number Where You Can Call Back," and "Closing."
[0041] In the examples shown in Figure 5, "Script (Example 1)" is a so-called "reading script type," "Script (Example 2)" is a so-called "item enumeration type" (or "utterance content enumeration type"), and "Script (Example 3)" is a so-called "keyword type." Generally, a script is defined using one of these types, but a script may be defined using two or more types. For example, for a certain item in a talk script, both utterance content and keywords may be defined.
[0042] ≪Example 3 of a talk script≫ Figure 6 shows an example of a talk script used, for example, to handle inquiries about malfunctions. Such a talk script is represented as a tree structure in which the utterances (scripts) that the operator needs to speak are nodes, and the transition relationships between utterances are directed edges (branches).
[0043] For example, the root node of the talk script shown in Figure 6 defines the utterance "Function A is not working" as the script, and it is shown that if the customer's response to this utterance is YES, the process proceeds to the left child node, and if it is NO, it proceeds to the right child node. Furthermore, the talk script shown in Figure 6 represents the progress of the inquiry process (i.e., the progress of the talk script) from the root node to the leaf nodes.
[0044] In the example shown in Figure 6, each node defines the utterances that the operator needs to speak as a script, but this is not limited to this. For example, each node may define sentences that the operator needs to speak as a script, or keywords or phrases that the operator needs to speak as a script. Furthermore, each node may also define the customer's utterances (or sentences, keywords, phrases, etc.). Moreover, instead of nodes, the edge may define the utterances (or sentences, keywords, phrases, etc.) as a script.
[0045] ≪Specific Example of a Talk Script 4≫ Figure 7 shows an example of a talk script used for handling inquiries that involve complex question-and-answer sessions (for example, inquiries regarding contracts for insurance or financial products). Such a talk script can be represented, for example, as a directed graph where the utterances (scripts) that the operator needs to speak are nodes, and the transition relationships between utterances are directed edges.
[0046] For example, in the talk script shown in Figure 7, node 0 defines the utterance "You should have a smartphone" as part of the script, and it is shown that if you want to provide a counter-argument to that utterance, you should proceed to node 1, and if you want to provide a reason, you should proceed to node 2. Also, in the talk script shown in Figure 7, the inquiry process progresses in the direction of the directed edge (i.e., the talk script progresses).
[0047] In the example shown in Figure 7, similar to the talk script shown in Figure 6, each node defines the utterances that the operator needs to speak as a script. However, this is not limited to this; for example, each node may define sentences that the operator needs to speak as a script, or keywords or phrases that the operator needs to speak as a script. Furthermore, each node may also define the customer's utterances (or sentences, keywords, phrases, etc.). Moreover, instead of nodes, the edge may define the utterances (or sentences, keywords, phrases, etc.) as a script.
[0048] The talk scripts in Examples 1-4 above are all illustrative, and this embodiment is applicable to any talk script. In addition to the talk scripts exemplified in Examples 1-4 above, there are also talk scripts that are expressed in a format in which labels representing items are attached to the utterances, and talk scripts that do not define items or scenes, but simply list sentences that the operator needs to utter, and this embodiment is similarly applicable to such talk scripts as well. Furthermore, as mentioned above, this embodiment is also applicable when the speaker is a robot or agent, and the talk script may be applied to a computer or program that implements such a robot or agent. Examples of talk scripts that can be applied to a computer or program include those described in International Publication No. 2019 / 172205.
[0049] <Detailed Functional Configuration of the Compliance Estimation Processing Unit 202> Figure 8 shows the detailed functional configuration of the compliance estimation processing unit 202 according to this embodiment. As shown in Figure 8, the compliance estimation processing unit 202 according to this embodiment includes a division unit 211, a matching unit 212, a corresponding information generation unit 213, a compliance estimation unit 214, a compliance range visualization unit 215, an aggregation unit 216, a compliance status visualization unit 217, an evaluation unit 218, a proposed revision identification unit 219, a proposed revision visualization unit 220, and a compliance rate visualization unit 221.
[0050] The splitting unit 211 divides the utterance text and the script included in the talk script into certain units. Hereinafter, the utterance text and script divided into certain units will also be referred to as "divided utterance text" and "divided script," respectively.
[0051] The matching unit 212 matches the divided utterance text and the divided script in that unit.
[0052] The correspondence information generation unit 213 generates correspondence information that represents the ranges that are matched between the divided utterance text and the divided script.
[0053] The conformance estimation unit 214 uses correspondence information to estimate whether the utterance text conforms to the talk script (or whether there is an utterance text that conforms to the talk script).
[0054] The compliance range visualization unit 215 visualizes on the operator terminal 20 or supervisor terminal 30 the ranges in the spoken text that comply with the talk script and the ranges that do not (or the ranges in the talk script where spoken text that complies with the script exists and where it does not).
[0055] The aggregation unit 216 creates a compliance history by aggregating the estimation results from the compliance estimation unit 214 and stores it in the storage unit 203.
[0056] The compliance status visualization unit 217 visualizes the compliance status of multiple operators' utterances in the same talk script on the operator terminal 20 or supervisor terminal 30.
[0057] The evaluation unit 218 evaluates the operator or talk script based on the call evaluation and related information. The evaluation unit 218 also calculates the compliance rate, as described later. Here, the call evaluation refers to information representing the results of a manual evaluation of a call between an operator and a customer. Related information refers to information related to the inquiry in the call, such as search keywords for FAQs or response manuals related to the inquiry (more specifically, the search keywords used by the operator to search the FAQ system or response manual when receiving the inquiry), browsing history of FAQs and response manuals, results of adding links to text representing inquiry handling records (links to FAQs), and escalation information to the supervisor. However, if other information can be obtained, such as information about the customer during the call (FAQ search history from past inquiries, past inquiry information, service contract information, etc.), this information may also be considered related information. In addition to FAQs and response manuals, if there is any support system that the operator can use while handling customer inquiries, information such as the usage history of that support system may also be considered related information.
[0058] Furthermore, call evaluations are not limited to those performed manually; they may also be performed automatically by a system. In this case, evaluations may be based on the number of turns, for example, with shorter call durations being considered better; automatic evaluations may be performed on a sentence or scene basis using a machine learning model; or evaluations may be based on customer reactions, such as the appropriateness of the operator's utterances and whether paraphrasing is possible. Additionally, call evaluations may be based on information evaluated for each call (i.e., calls with the same call ID), or on each utterance (for example, information evaluated for each segmented utterance text). Moreover, when obtaining a call evaluation for a single call from information evaluated for each utterance, for example, the information evaluated for each utterance may be scored and then the average may be calculated.
[0059] The proposed revision identification unit 219 identifies scripts that should be added to the talk script, unnecessary scripts, and unnecessary utterances in the utterance text as proposed revisions, based on the evaluation results from the evaluation unit 218. An unnecessary script is, for example, a script that, if followed, would lower (or potentially lower) the call evaluation.
[0060] The revised proposal visualization unit 220 visualizes the revised proposal on the operator terminal 20 or the supervisor terminal 30.
[0061] The compliance rate visualization unit 221 visualizes the compliance rate of the utterance text of operators belonging to a certain group that conforms to the talk script, and the compliance rate of the utterance text of a particular operator that conforms to the talk script, on the operator terminal 20 or supervisor terminal 30. In addition to the compliance rate, the compliance rate visualization unit 221 also visualizes the utterance text of each operator, related information, etc., on the operator terminal 20 or supervisor terminal 30.
[0062] The compliance range visualization unit 215, compliance status visualization unit 217, proposed revision visualization unit 220, and compliance rate visualization unit 221 may be collectively referred to as the "visualization information generation unit," etc. In addition, in the example shown in Figure 8, the utterance text and talk script are provided to the splitting unit 211, but other information such as call ID and operator ID may also be provided.
[0063] <Processing flow for saving compliance history and visualizing compliance and non-compliance scope> Figure 9 shows the processing flow for saving the compliance history and visualizing the scope of compliance and non-compliance. Here, the scope of compliance refers to the range in the utterance text that conforms to the talk script, or the range in the talk script where there is utterance text that conforms to the script. On the other hand, the scope of non-compliance refers to the range in the utterance text that does not conform to the talk script, or the range in the talk script where there is no utterance text that conforms to the script.
[0064] Steps S101 to S106 (or some of the steps therein) may be performed in real time while a call is being made between the operator and the customer, or they may be performed using pre-stored speech text or segmented speech text.
[0065] Step S101: First, the splitting unit 211 splits the utterance text and the script included in the talk script into predetermined units to create split utterance text and split script. The predetermined unit represents the unit for which it is desired to estimate whether the utterance text conforms to the talk script. Hereinafter, one split script will be assumed to represent one item or scene. In this case, since it is estimated whether the operator's utterance conforms to the item on an item-by-item basis, the item may be called a "conformity item," etc. However, one item or scene may be represented by multiple split scripts.
[0066] (How to split the script) In addition to dividing the script by the items or scenes mentioned above, it may also be divided by certain segments or sentences.
[0067] Furthermore, when splitting a script, it is split according to the order in which the talk script progresses. For example, in the case of a tree structure like that shown in Figure 6, a split script is created by sequentially arranging and expanding the scripts that exist along the path from the root node to the leaf nodes. Alternatively, in the case of a graph structure like that shown in Figure 7, a split script is created by sequentially arranging and expanding the scripts that exist along the path traversing directed edges from a predetermined initial node to the end node. However, the number of expansions may be limited using some kind of metric.
[0068] (Method of dividing spoken text) For example, the text may be divided into word units, phrase units, or certain segment units, or it may be divided into utterance units using existing text segmentation techniques. In this case, if the utterance text is text from a text chat, it can be divided as is, but if it is text converted by speech recognition, it may be divided after processing to improve readability, such as removing fillers.
[0069] However, the utterance text and script do not necessarily need to be split, and either the utterance text or the script, or both, may remain unsplit. Furthermore, since the utterance text can be considered as a split utterance text with a division count of 1, the term "split utterance text" below may include cases where it is not split. Similarly, since the split script can be considered as a split script with a division count of 1, the term "split script" below may include cases where it is not split.
[0070] Step S102: Next, the matching unit 212 matches the divided utterance text and the divided script in that unit and calculates a matching score that represents the degree of matching.
[0071] Step S103: Next, the correspondence information generation unit 213 uses the matching score calculated in step S102 to generate correspondence information representing the range that is matched between the divided utterance text and the divided script.
[0072] The following describes an example of matching in step S102 and generating correspondence information in step S103. However, in addition to the example described below, correspondence information may also be generated by determining the correspondence range between the segmented utterance text and the segmented script using the method described in Reference 1 (a method for determining sentence correspondence using a neural network).
[0073] (Example 1 of matching and correspondence information generation) This section explains how to generate correspondence information by solving a matching problem as a combinatorial problem.
[0074] Procedure 1-1: The matching unit 212 converts each segmented utterance text and each segmented script into features. Any method can be used for the conversion to features, but for example, one of the following methods 1 to 3 could be used. Alternatively, the above conversion to features may be performed by a device other than the estimation device 10, and then the matching unit 212 may input these features.
[0075] ·Method 1 Morphological analysis is performed on the segmented speech text to extract morphemes (keywords), and the word vectors representing the extracted morphemes are used as features. Similarly, morphological analysis is performed on the segmented script to extract morphemes (keywords), and the word vectors representing the extracted morphemes are used as features.
[0076] ·Method 2 Morphological analysis is performed on the segmented speech text to extract morphemes (keywords), and then the vectors obtained by transforming the extracted morphemes using Word2Vec are used as features. Similarly, morphological analysis is performed on the segmented script to extract morphemes (keywords), and then the vectors obtained by transforming the extracted morphemes using Word2Vec are used as features.
[0077] ·Method 3 The vectors obtained by transforming the segmented utterance text using text2vec are used as features. Similarly, the vectors obtained by transforming the segmented script using text2vec are used as features.
[0078] Step 1-2: The matching unit 212 uses the features calculated in Step 1-1 above to calculate a matching score between each segmented utterance text and each segmented script. Specifically, for example, if the i-th segmented utterance text is "Segmented Utterance Text i" and the j-th segmented script is "Segmented Script j", then for each i and j, the matching score s between segmented utterance text i and segmented script j is calculated. ij Calculate the matching score s. ijFor example, one can calculate the similarity (e.g., cosine similarity) between the features of the segmented utterance text i and the features of the segmented script j.
[0079] Step 1-3: The matching unit 212 uses the matching score calculated in Step 1-2 above to identify the correspondence between the segmented utterance text and the segmented script. For example, the correspondence is identified using dynamic programming as an elastic matching problem. In this embodiment, since similarity is used as the matching score, when identifying the correspondence using dynamic programming, the matching score value is converted from similarity to a cost representing distance before the calculation. However, the correspondence may also be identified using, for example, integer linear programming.
[0080] For example, suppose the matching scores shown in Figure 10 have been calculated. In Figure 10, the matching score is written in parentheses in each cell. For example, the matching score between split utterance text 1 and split script 1 is 0.8, the matching score between split utterance text 1 and split script 2 is 0.2, and the matching score between split utterance text 1 and split script 3 is 0.1.
[0081] In this case, the split utterance text 1 is identified as corresponding to split script 1, split utterance text 2 as corresponding to split script 2, split utterance text 4 as corresponding to split script 2, and split utterance text 5 as corresponding to split script 4. Therefore, in this case, split utterance text 1 corresponds to the range of items represented by split script 1, split utterance text 2 and split utterance text 4 correspond to the range of items represented by split script 2, and split utterance text 5 corresponds to the range of items represented by split script 4.
[0082] For example, if there is a split utterance text whose matching score with all split scripts is below a predetermined threshold, this split utterance text may be excluded in advance. Similarly, if there is a split script whose matching score with all split utterance texts is below a predetermined threshold, this split script may be excluded in advance. Figure 10 shows an example where split utterance text 3 and split script 3 may be excluded in advance.
[0083] Furthermore, when identifying correspondences, the matching score may be adjusted using auxiliary information such as turns. For example, the score may be adjusted by adding a certain score to the matching score of a split script belonging to a predetermined turn. To give a specific example, one could uniformly add 0.2 to the matching score of a split script belonging to the first three turns.
[0084] When identifying correspondences by solving an elastic matching problem, matching can be performed by considering the order in which the divided utterance texts and divided utterances progress. However, if the order of the divided scripts can be ignored, each divided utterance text may be associated with one divided script whose matching score is above a predetermined threshold (e.g., 0.5), or the correspondences may be identified by solving a maximum matching problem of a bipartite graph.
[0085] Step 1-4: The correspondence information generation unit 213 generates correspondence relationship information that represents the correspondence relationship identified in Step 1-3 above.
[0086] (Example of matching and correspondence information generation, part 2) This section explains how to generate correspondence information by solving a matching problem as an extraction problem.
[0087] Step 2-1: The matching unit 212 converts each segmented utterance text and each segmented script into features. Any method can be used for converting the features, but for example, a pre-trained language model fine-tuned for a machine reading comprehension task that extracts answers to question sentences from a text to be read can be used to convert each segmented utterance text and each segmented script into hidden layer vectors, and these vectors can be used as features. In this embodiment, the case where BERT (Bidirectional Encoder Representations from Transformers) is used as the pre-trained language model will be described, but any other pre-trained language model that can perform similar processing may be used. BERT is a pre-trained natural language model used in machine reading comprehension technology, etc. For example, see Reference 2. When the segmented utterance text and segmented script are input to BERT, they are divided into predetermined units called tokens (for example, words or subwords). Hereafter, the fine-tuned pre-trained language model described above will be referred to as the "correspondence model".
[0088] Step 2-2: The matching unit 212 calculates a matching score between each segmented utterance text and each segmented script using the features calculated in Step 2-1 within the mapping model. In a machine reading task where the answer to a question is extracted from the text to be read, the start and end points of the range that corresponds to the answer to the question are output within the text to be read. These start and end points are determined by calculating the start and end scores (hereinafter also referred to as the start score and end score) for each token in the text to be read, and then summing them (hereinafter also referred to as the overall score). Therefore, considering the segmented script as the question and the segmented utterance text as the text to be read, the mapping model (in this embodiment, the finely tuned BERT) calculates the start score and end score for each token included in the segmented utterance text, and these start and end scores are used as the matching score. When performing the fine-tuning described above, a training dataset consisting of multiple sets of three pieces of information—(split script, split utterance text, and reference range)—is used.
[0089] However, when calculating the starting and ending scores using the correspondence model, the segmented utterance text may be treated as the question text and the segmented script as the text to be read.
[0090] Step 2-3: The matching unit 212 uses the matching score calculated in Step 2-2 above to identify the correspondence between the segmented utterance texts and segmented scripts. That is, for example, for each segmented script, the range with the highest overall score is used to create correspondence information for that segmented script. However, if the segmented utterance text is considered a question and the segmented script is considered a text to be read, then the range with the highest overall score for each segmented utterance is used to create correspondence information for that segmented utterance text.
[0091] Hereinafter, specific examples of the above-mentioned steps 2-2 to 2-3 will be described. Note that the number of divisions in each of the following specific examples is just an example, and the number of divisions of the speech text, script, divided speech tokens, and divided script can be determined independently of each other.
[0092] · Specific Example 1 A specific example where the speech text is not divided in the above step S101 and only the script is divided will be described.
[0093] For example, as shown in FIG. 11, assume that the script is divided into divided scripts 1 to 4, and when the speech text is input into the association model, this speech text is divided into tokens x1, ···, x 20 Let it be divided into. Hereinafter, these tokens x1, ···, x 20 will also be referred to as "speech tokens". When the association model is BERT, special tokens representing the beginning of a sentence, the end of a sentence, etc. are also input, but for simplicity, the description thereof is omitted (the same omission is made in the following specific examples 2 and 3).
[0094] At this time, in this specific example, each speech token and each divided script are matched by the association model, and a start score where each speech token is the starting point and an end score where it is the ending point are calculated for each divided script. That is, if the k-th speech token is x k and the j-th divided script is "divided script j", then for the divided script j, the start score s k where the speech token x kj is the starting point and the end score e kj are calculated.
[0095] And for the divided script j, the start score s kj and the end score s k'jThe range in which the sum of k is maximized (where k ≤ k') becomes the corresponding range of the split script j, and correspondence information representing this corresponding range is created. For example, in the example shown in Figure 11, the corresponding range of split script 1 is utterance tokens x1 to x6, and the corresponding range of split script 2 is utterance tokens x7 to x 12 The range of the split script 3 corresponds to speech token x9~x 16 The range of the split script 4 corresponds to speech token x 17 ~x 20 This indicates that...
[0096] Furthermore, it is possible that multiple corresponding ranges can be obtained for a given split script j. For example, the corresponding range of split script 4 may be speech tokens x3~x5 and speech token x 17 ~x 20 This is the case. In such cases, for example, one can solve a combination problem as explained in Matching and Correspondence Information Generation Example 1 to identify one of the two, or one can select the correspondence range with the highest overall score. However, if the correspondence range with the highest overall score is selected, the order of the script's progress may be ignored, so auxiliary information such as turns may be used to ensure that the order of progress is taken into consideration. These points are also true for the following specific examples 2 and 3.
[0097] • Specific example 2 The following provides a specific example of what happens when both the spoken text and the script are split in step S101 above.
[0098] For example, as shown in Figure 12, suppose the utterance text is divided into divided utterance text 1 to divided utterance text 5, and the script is divided into divided script 1 to divided script 4, and when the divided utterance text i (i=1,···,5) is input into the mapping model, this divided utterance text i corresponds to utterance token x1 i ,···,x4 iIt is assumed that the text is divided into these segments. As mentioned above, these division numbers are just examples, and the number of divisions for the utterance text, script, and divided utterance text can be determined independently. For example, in the example shown in Figure 12, all divided utterance texts are divided into four utterance tokens, but the number of divisions into utterance tokens may differ for each divided utterance text.
[0099] In this specific example, for each segmented utterance text, the correspondence model matches each utterance token with each segmented script, and for each segmented script, a starting score is calculated where each utterance token is the starting point and an ending score where it is the ending point. That is, for segmented script j, utterance token x k i The starting point is the starting score s kj i , the final destination score e kj i This is calculated.
[0100] Then, for the split script j, the starting score s kj i and the final score s k'j i The range in which the sum of k is maximized (where k ≤ k') becomes the corresponding range for the split script j, and correspondence information representing this corresponding range is created. For example, in the example shown in Figure 12, the corresponding range for split script 1 is speech token x1 1 ~x3 1 The scope of the split script 2 is utterance token x1 2 ~x4 2 The scope of the split script 3 is 1 speech token 3 ~x4 3 and x1 4 ~x4 4 The scope of the split script 4 is utterance token x1 5 ~x4 5 This indicates that...
[0101] • Specific example 3 This section describes a specific example of matching each utterance token in a segmented utterance text with each token in a segmented script (hereinafter also referred to as "script token"). This example can be implemented, for example, using the method described in Reference 3 (a method for finding word correspondences between two texts). Therefore, in this example, the model described in Reference 3 will be used as the correspondence model.
[0102] For example, as shown in Figure 13, suppose the utterance text is divided into divided utterance text 1 to divided utterance text 5, and the script is divided into divided script 1 to divided script 4. Also, when inputting divided utterance text i (i=1,···,5) into the mapping model, utterance token x1 i ,···,x4 i It is divided into parts, and when inputting the divided script j(i=1,···,4) into the mapping model, the script token y1 j ,y2 j It is assumed that the text is divided into these. As mentioned above, these division numbers are just examples, and the number of divisions for the utterance text, script, divided utterance text, and divided script can each be determined independently. For example, in the example shown in Figure 13, all divided utterance texts are divided into four utterance tokens, and all divided scripts are divided into two script tokens, but the number of divisions into utterance tokens may differ for each divided utterance text, and similarly, the number of divisions into script tokens may differ for each divided script.
[0103] In this specific example, for each segmented utterance text, the correspondence model matches each utterance token with each script token in each segmented script, and for each script token in each segmented script, a starting score where the utterance token is the starting point and an ending score where it is the ending point are calculated. That is, the script token y of segmented script j m j For the utterance token x k i The starting point is the starting score s kmj i, the final destination score e knj i This is calculated.
[0104] And the script token y of the split script j m j For the starting score s kmj i and the final score s k'mj i The range (where k ≤ k') in which the sum is maximized is the script token y m j This represents the scope of correspondence, and correspondence information representing this scope is created. For example, in the example shown in Figure 13, the script token y1 of split script 1 1 The range of correspondence is 1 speech token 1 ~x3 1 , script token y2 of split script 1 1 The range of correspondence is 4 speech tokens 1 , script token y1 of split script 2 2 The range of correspondence is 1 speech token 2 ~x3 2 , script token y2 of split script 2 2 The range of correspondence is 4 speech tokens 2 This indicates that, among other things, that... Note that in the example shown in Figure 13, utterance token x1 4 ~x3 4 There is no corresponding script token.
[0105] Step S104: Next, the conformance estimation unit 214 uses the correspondence information generated in step S103 above to estimate, based on predetermined estimation conditions, whether the utterance text conforms to the talk script, or whether there is an utterance text that conforms to the talk script. Hereinafter, if the utterance text conforms to the talk script, it will be referred to as "utterance conformance," and if it does not, it will be referred to as "utterance non-conformance." On the other hand, if there is an utterance text that conforms to the talk script, it will be referred to as "script conformance," and if there is no such utterance text, it will be referred to as "script non-conformance."
[0106] The above-mentioned predetermined estimation conditions include, for example, whether or not a corresponding text exists as correspondence information for the text to be evaluated, where the text to be evaluated is referred to as the "text to be evaluated" and the text to be evaluated is referred to as the "text to be evaluated." When using these estimation conditions, if a corresponding split script (text to be evaluated) exists for a given split utterance text (text to be evaluated), then that split utterance text is estimated to be utterance-compliant. On the other hand, if a corresponding split script does not exist, then that split utterance text is estimated to be utterance-non-compliant.
[0107] Furthermore, if a split script (text to be judged) has a corresponding split utterance text (text to be judged), the split script is presumed to be script-compliant. On the other hand, if no corresponding split utterance text exists, the split script is presumed to be script-non-compliant.
[0108] However, even if a corresponding text to be judged exists as corresponding information, if the matching score is below a certain predetermined threshold, it may be presumed to be non-conforming to speech or script. This corresponds to using a presumption condition that is further limited by the matching score to the condition "whether or not a corresponding text to be judged exists as corresponding information for the text to be judged."
[0109] The compliance estimation unit 214 may also estimate whether a call (i.e., all utterances in a single call) conforms to the talk script. For example, the compliance estimation unit 214 may estimate that a call conforms to the talk script if the proportion of segmented utterance texts estimated to be "compliant" among the segmented utterance texts in a call satisfies a certain condition (e.g., 80% or more). Alternatively, the compliance estimation unit 214 may estimate that a call conforms to the talk script if it conforms to an item in the talk script that must be followed, or it may estimate whether a call conforms to the talk script using various other rule-based methods.
[0110] Step S105: Next, the aggregation unit 216 creates a compliance history from the estimation results of step S104 (whether the divided utterance text conforms to the utterance or not, whether each divided script conforms to the script or not), and stores the compliance history in the storage unit 203.
[0111] An example of a compliance history is shown in Figure 14. In the compliance history shown in Figure 14, the call ID, operator ID, item, script, utterance ID, utterance, matching score, script compliance / non-compliance, and utterance compliance / non-compliance are associated with each other. In addition to these, other items such as script ID and script item ID may also be associated.
[0112] Here, the call ID is an ID that identifies the call between the operator and the customer, the operator ID is an ID that identifies the operator, and the item is a compliant item of the talk script. The script is the script that belongs to that compliant item, and in the example shown in Figure 14, it is one split script. The utterance ID is an ID that identifies a certain utterance unit of the operator, and the utterance is the utterance text of that utterance unit, and in the example shown in Figure 14, it is one split utterance text. The matching score is the matching score between the split script and the split utterance text, and in the example shown in Figure 14, the matching score is calculated using the method described in the specific example shown in Figure 13, and the average of that matching score across the split utterance text (or split script) is used. Script compliance / non-compliance and utterance compliance / non-compliance are the estimation results of step S104 above.
[0113] Furthermore, in the example shown in Figure 14, the corresponding ranges between the script and the utterance are shown in bold. For example, in the third line of the script in the example shown in Figure 14, "Could you tell me your phone number and your name?", the part "Could you tell me your name?" is in bold, meaning that a corresponding utterance exists. Similarly, the utterance "Please tell me your name." is in bold, meaning that a corresponding script exists. On the other hand, in the fourth line of the script in the example shown in Figure 14, "Could you tell me your phone number and your name?", it means that there is no utterance corresponding to "Could you tell me your name?". Based on this correspondence information, it is determined whether or not there is a corresponding range between the script and the utterance.
[0114] In the example shown in Figure 14, lines 3 and 4 of the conformance history show that while there are utterances corresponding to the script and scripts corresponding to the utterances, the matching score is below a certain threshold (e.g., 0.5). Therefore, the estimated results for script conformance / non-conformance and utterance conformance / non-conformance are set to non-conformance, respectively.
[0115] Here, the aggregation unit 216 may merge utterances if multiple utterances are associated with the same compliance item. In this case, the values set for script compliance / non-compliance and utterance compliance / non-compliance may be changed by adding the matching score of the merged utterances.
[0116] For example, Figure 15 shows the compliance history obtained by integrating the third and fourth lines of the compliance history shown in Figure 14. In the example shown in Figure 15, as a result of integrating the third and fourth lines of the compliance history shown in Figure 14, the matching score in the third line of the compliance history shown in Figure 15 becomes 0.9, and as a result, both script compliance / non-compliance and utterance compliance / non-compliance are changed to "compliant".
[0117] Furthermore, as described above, if multiple split utterances are associated with a single split script, it would be possible to further highlight (for example, by highlighting in red) the range of the corresponding split script when the cursor or other cursor is placed over any of the multiple split utterances.
[0118] Step S106: The compliance range visualization unit 215 generates information (for example, screen information to be displayed on the user interface; hereinafter also referred to as visualization information) for visualizing the ranges in the spoken text that comply with the talk script and the ranges that do not comply (hereinafter also referred to as "spoken compliance range" and "spoken non-compliance range", respectively), or the ranges in the talk script where there are spoken texts that comply with the script and the ranges where there are no such ranges (hereinafter also referred to as "script compliance range" and "script non-compliance range", respectively), and transmits the generated visualization information to the operator terminal 20 or supervisor terminal 30. As a result, the spoken compliance range and spoken non-compliance range, script compliance range and script non-compliance range, etc. are visualized on the display of the operator terminal 20 or supervisor terminal 30. Note that this step does not necessarily have to be performed after step S105, and may be performed after step S103 described above. However, if this is performed after step S103 above, only the correspondence information will be visualized (for example, as shown in the example in Figure 15, the script or utterance will be visualized with the range where correspondence information exists displayed in bold).
[0119] Figure 16 shows an example of the visualization results for utterance-compliant and non-compliant ranges. In the example shown in Figure 16, the range of utterance text that conforms to each item (utterance-compliant range) is shown in bold. On the other hand, ranges that are not in bold represent the utterance-non-compliant range. This allows operators and supervisors to check which ranges of utterance text conform to which items in the talk script.
[0120] Furthermore, Figure 17 shows an example of the visualization results of script-compliant and script-non-compliant ranges. In the example shown in Figure 17, for each script belonging to an item, the range of scripts in which compliant utterance text exists (script-compliant range) is shown in bold. On the other hand, the range that is not bold represents the script-non-compliant range. This allows operators and supervisors to check which scripts belonging to each item have utterance text that complies with that script.
[0121] Here, the visualization information for the speech-compliant and non-compliant ranges, and the visualization information for the script-compliant and non-compliant ranges are created from the estimation results of step S104 above (or the compliance history, which is a history of these estimation results), but they may also be created from the correspondence information. For example, if step S106 is executed after step S103 above, the visualization information will be created from the correspondence information. Alternatively, the visualization information for the speech-compliant and non-compliant ranges, and the visualization information for the script-compliant and non-compliant ranges may be created from both the estimation results of step S104 above (or the compliance history, which is a history of these estimation results) and the correspondence information, respectively. In this case, it may be possible to switch which visualization information is based on, for example, according to the user's selection or settings.
[0122] In the examples shown in Figures 16 and 17, the speech-compliant range and script-compliant range are shown in bold, but bolding is just one example, and bolding is not required as long as it is different from the non-compliant range. For example, the speech-compliant and script-compliant ranges could be colored differently, highlighted, etc.
[0123] Furthermore, only one of the following—the utterance compliance range and the utterance non-compliance range, or the script compliance range and the script non-compliance range—may be visualized on the operator terminal 20 or supervisor terminal 30, or both may be visualized. In addition to the utterance compliance range and script compliance range, compliance ratios, compliance counts, matching scores, etc., may also be visualized. In this case, if the compliance ratio, compliance counts, matching scores, etc., are visualized along with the utterance compliance range and script compliance range, the visual effects may be changed, for example, by changing the bold size or color of the utterance compliance range and script compliance range according to the values of the compliance ratio, compliance counts, matching scores, etc. When calculating the compliance ratio and compliance counts, for example, compliance or non-compliance may be counted on an item-by-item basis in the talk script, or compliance or non-compliance may be calculated on a segmented script basis.
[0124] <Processing flow for visualizing compliance status> Figure 18 shows the processing flow for visualizing compliance status. Here, compliance status refers to the total number of compliances for each script in the talk script.
[0125] Step S201: First, the aggregation unit 216 aggregates the compliance history stored in the storage unit 203. For example, the aggregation unit 216 aggregates the number of script compliances for each script (i.e., the total number of script compliances / non-compliances set to "compliant"). This aggregated result represents the compliance status of utterances by multiple operators in the same talk script. When aggregating, for example, only the number of script compliances for utterances belonging to a specific group (e.g., a specific department, a group that handles specific inquiries, a specific incoming number, etc.) may be aggregated. Alternatively, for example, the compliance history may be aggregated when the same operator handles the same talk script multiple times (this will allow the operator to see which parts of the talk script are more compliant and which are not in the compliance status visualization results described later). Alternatively, for example, the compliance history may be aggregated on a daily basis so that the compliance status visualization results described later can be viewed on a daily basis (especially in chronological order) (this will allow for verification such as "will compliance improve as experience accumulates?").
[0126] Step S202: The compliance status visualization unit 217 then generates visualization information of the compliance status of multiple operators' utterances in the same talk script and transmits the generated visualization information to the operator terminal 20 or supervisor terminal 30. This visualizes the compliance status on the display of the operator terminal 20 or supervisor terminal 30. An example of the compliance status visualization result is shown in Figure 19. In the example shown in Figure 19, scripts such as "Thank you for calling," "May I have your phone number and name?", "Could you tell me your date of birth?", and "Could you tell me your contract number?" are visualized, and scripts with a higher number of script compliances are visualized in larger font (i.e., highlighted). Note that visualizing scripts with a higher number of script compliances in larger font is just one example; any method of highlighting scripts with a higher number of script compliances is acceptable. This allows the operator or supervisor to know which scripts are easy (or difficult) to comply with.
[0127] <Processing flow for visualizing proposed revisions, compliance rates, operator utterances, and related information> Figure 20 shows the processing flow for visualizing proposed revisions, compliance rates, operator utterances, and related information. Here, proposed revisions refer to utterances that are currently non-compliant with the script but are considered to be better incorporated into the script (proposed script additions), scripts that are considered to be better removed from the talk script (proposed script deletions), and unnecessary utterances that are non-compliant with the script (proposed utterance revisions).
[0128] Furthermore, related information that is highly relevant to the utterance text that has been proposed for script addition (for example, search keywords that are frequently used in FAQs when the utterance text is spoken, or links to FAQs) may be included as a proposed revision along with the proposed script addition.
[0129] Step S301: First, the aggregation unit 216 combines the call evaluation and related information with the compliance history stored in the storage unit 203. Figure 21 shows the result of combining the call evaluation and related information with the compliance history shown in Figure 15. In the example shown in Figure 21, the call evaluation is assumed to be a graded evaluation such as "A", "B", "C", etc., but it is not limited to this and may be a numerical value such as a score, for example.
[0130] Step S302: Next, the evaluation unit 218 calculates an evaluation score in a certain unit (for example, per operator or per talk script) using the compliance history stored in the memory unit 203. Examples of evaluation scores include compliance rate, precision, recall, and F-score. Note that compliance rate, precision, and recall do not necessarily have to be percentages or ratios, and may be called conformance, fit, recall, etc.
[0131] Operator-level compliance can be, for example, the percentage of segmented utterances in the operator's segmented utterance text that are estimated to be speech-compliant. Operator-level precision can be calculated as "(number of segmented utterances in the operator's segmented utterance text that conform to the talk script) / (total number of segmented utterances in the operator's segmented utterance text)". Operator-level recall can be calculated as "(number of conforming items in the talk script that are conforming in the operator's utterance text) / (total number of conforming items in the talk script)". Operator-level F-score can be calculated as the harmonic mean of operator-level precision and operator-level recall.
[0132] The conformance rate at the talk script level should be the percentage of split scripts within the talk script that are estimated to be script-compliant. The precision rate at the talk script level should be calculated as "(number of split utterances that conform to the talk script when the talk script is used) / (total number of split utterances when the talk script is used)". The recall rate at the talk script level should be calculated as "(number of conforming items in the talk script that conform in the utterances when the talk script is used) / (total number of conforming items in the talk script)". The F-score at the talk script level should be the harmonic mean of the precision rate and the recall rate at the talk script level.
[0133] In addition to the above, evaluation scores may also be calculated for each operator belonging to a specific group (e.g., a specific department, a group handling specific inquiries, a specific incoming number, etc.). Alternatively, evaluation scores may be calculated for each item in the talk script. Furthermore, evaluation scores may be calculated for each operator and each item in the talk script.
[0134] For example, the compliance rate for each operator and each item of the talk script can be calculated as the percentage of segmented utterances that are estimated to be utterance-compliant with the item, out of the segmented utterances of that operator for that item. Other evaluation scores can be calculated similarly using utterances filtered by item as appropriate.
[0135] Step S303: Next, the revision proposal identification unit 219 uses the evaluation score calculated in step S302 to identify either or both of the script revision proposals and the speech revision proposals.
[0136] Here, one possible approach to adding scripts is to identify the utterances of operators who have high call evaluations but low compliance rates. Another possible approach to removing scripts is to identify the utterances of operators who have low call evaluations but high compliance rates, or to identify scripts for compliance items that have both low call evaluations and low compliance rates. A possible approach to revising utterances is to identify utterances that have both low call evaluations and low compliance rates. These are just examples, and script additions, deletions, and revisions may also be identified using precision, recall, F-score, etc.
[0137] Step S304: Next, the revision proposal visualization unit 220 generates visualization information for the revision proposals (script addition proposals, script deletion proposals, and speech revision proposals) identified in step S303, and transmits the generated visualization information to the operator terminal 20 or supervisor terminal 30. As a result, the revision proposals (script addition proposals, script deletion proposals, and speech revision proposals) are visualized on the display of the operator terminal 20 or supervisor terminal 30. Preferably, for example, the script addition proposals and script deletion proposals are visualized on the supervisor terminal 30, and the speech revision proposals are visualized on the operator terminal 20.
[0138] Figure 22 shows an example of the visualization results of proposed script additions. In the example shown in Figure 22, the operator's utterance text is visualized under "Non-compliant utterances." This utterance text has a high call evaluation ("A" in the example shown in Figure 22), but it does not conform to the talk script. Therefore, supervisors can use this utterance text as a reference to consider what kind of script should be added to the talk script.
[0139] In the example shown in Figure 22, the items that the preceding and succeeding utterances conform to (the preceding and succeeding conforming items) are also visualized. This allows the supervisor to see in what context the non-conforming utterance was spoken. In this case, the preceding and succeeding utterances may also be visualized.
[0140] Figure 23 shows an example of the visualization results of proposed speech corrections. In the example shown in Figure 23, the operator's speech text is visualized under "Non-compliant speech." This speech text indicates that the call evaluation was low ("C" in the example shown in Figure 23) and that the speech did not conform to the talk script. Therefore, operators can use this speech text as a reference to consider whether their own speech was inappropriate (for example, whether they made unnecessary speech that was not in the talk script). Supervisors can also use this speech text to check whether something unexpected happened to the operator, and can provide education and guidance to the operator.
[0141] In the example shown in Figure 22, a call evaluation of "A" is considered high, but for example, a call evaluation of "A" and "B" could also be considered high. In other words, there may be multiple values or a range of values that are judged to be high in call evaluation. In this case, the visualization results of the script addition proposals may allow for sorting and filtering of the utterance text based on the call evaluation. Similarly, there may be multiple values or a range of values that are judged to be low in call evaluation, and in this case, the visualization results of the utterance revision proposals may also allow for sorting and filtering of the utterance text based on the call evaluation.
[0142] Step S305: The compliance rate visualization unit 221 generates visualization information of the compliance rate, which is one of the evaluation scores in step S302 above, and transmits the generated visualization information to the operator terminal 20 or supervisor terminal 30. As a result, the compliance rate is visualized on the display of the operator terminal 20 or supervisor terminal 30.
[0143] Figure 24 shows an example of the visualization results of the compliance rate of a particular operator (hereinafter referred to as "Operator A"). In the example shown in Figure 24, for each item (scene) of the talk script, the average compliance rate of the operators for that item and Operator A's compliance rate are visualized. In addition, in this case, areas where Operator A's compliance rate is particularly low (for example, areas below a certain predetermined threshold) are displayed in a different manner than others. In the example shown in Figure 24, Operator A's compliance rate of "20%" for the item "Confirm a phone number where you can call back" is visualized in a conspicuous manner. This allows Operator A or the supervisor to know which items (scenes) have particularly low compliance rates.
[0144] As shown in Figure 24, the example allows us to compare the compliance rate of a typical operator with that of a specific operator, enabling us to identify, for example, items that a particular operator struggles with. Furthermore, if a specific operator has a low compliance rate, and the average compliance rate for all operators is also low, it indicates that this is an item that is difficult for any operator to achieve compliance with.
[0145] In the example shown in Figure 24, the average operator compliance rate and the compliance rate of a particular operator are visualized for each item in the talk script. However, this is just one example, and compliance rates can be visualized using various other criteria.
[0146] For example, for each talk script, the compliance rate for calls with a call evaluation of "A" and the compliance rate for calls with a call evaluation of "C" may be visualized. In this case, items with low compliance rates in calls with a call evaluation of "A" and items with high compliance rates in calls with a call evaluation of "C" may be visualized in a more prominent manner. This is because items with high call evaluations but low compliance rates may contain unnecessary scripts, allowing for consideration of script revision. Similarly, items with low call evaluations but high compliance rates may also contain unnecessary scripts, allowing for consideration of script revision. Note that high or low compliance rates can be determined simply by comparing them to a threshold, but they can also be determined by performing statistical tests, for example, to see if there is a significant difference.
[0147] Step S306: The compliance rate visualization unit 221 generates visualization information of the operator's utterances and transmits the generated visualization information to the operator terminal 20 or supervisor terminal 30. This visualizes the operator's utterances on the display of the operator terminal 20 or supervisor terminal 30. For example, by selecting a desired item from the compliance rate visualization results, the operator or supervisor can visualize a list of utterance texts (operator utterances) that comply with that item.
[0148] Figure 25 shows an example of the operator utterances when the item "Confirm a phone number where you can call back" is selected in the compliance rate visualization results shown in Figure 24. Note that in Figure 25, the utterance text for the item "Confirm a phone number where you can call back" is visualized, but for example, a list of utterance texts for all items could be displayed, and then the utterance texts could be filtered and visualized in Figure 25 when the item "Confirm a phone number where you can call back" is selected in the compliance rate visualization results shown in Figure 24.
[0149] In the example shown in Figure 25, the utterance text for the item "Confirm a phone number where you can call back" from operators A, B, and C is visualized. The call ID in which the utterance text was spoken, and the call evaluation for that call, are also visualized. This allows operators or supervisors to see the utterances and call evaluations for the relevant item from various operators. Furthermore, this operator utterance list may allow for sorting and filtering of utterance texts, for example, by call evaluation. While the script is not visualized in the example shown in Figure 25, it may be visualized as well.
[0150] Step S307: The compliance rate visualization unit 221 generates visualization information of related information and transmits the generated visualization information to the operator terminal 20 or supervisor terminal 30. As a result, the related information is visualized on the display of the operator terminal 20 or supervisor terminal 30. This means that, for example, the operator or supervisor can visualize the related information by performing an operation to display the related information in the compliance rate visualization result. This allows, for example, the operator to find out what they were having trouble with if they were not compliant with the script, and can be used to correct the script or FAQ, etc.
[0151] Figure 26 shows an example of the visualization results of related information. In the example shown in Figure 26, "FAQ search keyword ranking," "FAQ browsing history," and "SV escalation information" are visualized as examples of related information for a particular operator. Note that this related information does not necessarily represent the related information of a single operator; for example, it may be an aggregate of related information from multiple operators.
[0152] In step S306 above, the operator's compliance rate was visualized, but the compliance rate of the talk script may also be visualized. For example, the compliance rate visualization results shown in Figure 27 may be visualized. In the example shown in Figure 27, for each item of the talk script, the compliance rate of calls with a high evaluation result for that item (e.g., calls with a call evaluation above a predetermined threshold) and the compliance rate of calls with a low evaluation result (e.g., calls with a call evaluation below a predetermined threshold) are visualized. Furthermore, in this case, by selecting a desired item in the compliance rate visualization results shown in Figure 27, a list of utterance texts (operator utterances) that comply with that item can be visualized. For example, the example shown in Figure 28 is the visualization result when the item "Identity Verification" is selected in the compliance rate visualization results shown in Figure 27 (i.e., when the cell in the 4th row and 1st column is selected in the visualization results shown in Figure 27). Since the visualization result shown in Figure 28 is the same as that in Figure 25, a detailed explanation is omitted.
[0153] In the above example, the operator or supervisor selected the cell in the first column representing the item in the visualization result shown in Figure 27. However, any other desired cell other than the first column may be selected in the visualization result shown in Figure 27. For example, the example shown in Figure 29 is the visualization result when the cell in the 5th row and 4th column is selected in the visualization result shown in Figure 27 (i.e., when the cell representing the compliance rate (high evaluation result call) for the item "Confirm a phone number that can be returned" is selected). The visualization result shown in Figure 29 is a filtered display of operator utterances for the item "Confirm a phone number that can be returned" that also have a high evaluation result.
[0154] As another example, Figure 30 shows the visualization result when the cell in the 5th row and 5th column is selected in the visualization result shown in Figure 24 (i.e., when the cell for the compliance rate (Operator A) is selected in the item "Check a phone number where you can call back"). The visualization result shown in Figure 30 is filtered to display only the operator utterances of Operator A in the item "Check a phone number where you can call back".
[0155] As shown in Figures 24 and 27, when a desired cell is selected in the visualization results, the operator utterance corresponding to that cell (and its corresponding items, operator ID, call ID, call evaluation, etc.) is displayed in a list.
[0156] The present invention is not limited to the embodiments specifically disclosed above, and various modifications, changes, and combinations with known technologies are possible without departing from the scope of the claims.
[0157] The following additional information is disclosed regarding the embodiments described above.
[0158] (Note 1) Memory and At least one processor connected to the memory, Includes, The aforementioned processor, Taking information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, visualization information is generated to visualize the range of the utterance content represented by either the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. Visualization information generation device.
[0159] (Note 2) The utterance text and the script are pre-divided into one or more segmented utterance texts and one or more segmented scripts, respectively. The visualization information generating device according to Appendix 1, wherein the information representing at least one of compliance and non-compliance is information representing at least one of compliance and non-compliance between the divided utterance text and the divided script.
[0160] (Note 3) The aforementioned processor, Based on the information representing at least one of the conformance and non-conformance, a score is calculated representing the degree of conformance between the utterance content represented by the divided utterance text and the utterance content represented by the divided script. A visualization information generation device as described in Appendix 2, which generates visualization information for visualizing the score for each divided script of the same script.
[0161] (Note 4) The aforementioned score includes, For all of the aforementioned segmented utterance texts, a degree of relevance is given to the extent to which the segmented utterance texts that are presumed to be relevant occupy the text. For all of the aforementioned split scripts, the degree of reproducibility represents the extent to which the split scripts that are presumed to conform to the scripts are present. A visualization information generation device as described in Appendix 3, which includes either one of the following.
[0162] (Note 5) The aforementioned processor, A visualization information generating device according to any one of the appendices 2 to 4, which generates visualization information for visualizing, for each item of the same script or each of the same split scripts, the item and the split utterance text of the utterance content that is presumed to be in accordance with the utterance content representing the split script corresponding to the item.
[0163] (Note 6) The aforementioned processor, A visualization information generating device according to any one of the appendices 2 to 4, which generates visualization information for visualizing a list of the speech content of divided speech texts that are presumed to be non-conforming to the speech content represented by the divided script.
[0164] (Note 7) The aforementioned processor, A visualization information generation device as described in Appendix 6, which generates visualization information for visualizing a list of utterances of segmented utterance texts that are presumed to be non-conforming to the utterance content represented by the segmented script, for each utterance subject unit of the utterance text.
[0165] (Note 8) The aforementioned processor, A visualization information generating device according to Appendix 6 or 7, which generates visualization information for visualizing a list of utterances of split utterance texts that are presumed to be non-conforming to the utterance content represented by the split script, and for visualizing the items before and after the item corresponding to the split script, or the split utterance texts before and after the split utterance text.
[0166] (Note 9) The aforementioned processor, A visualization information generating device according to any one of the appendices 2 to 8, which acquires evaluation information for the utterance text from an external source, and generates visualization information for visualizing in a list the acquired evaluation information, the divided utterance text, and information representing at least one of the conformance and non-conformance between the divided utterance text and the divided script.
[0167] (Note 10) The aforementioned processor, Related information related to the aforementioned utterance text is obtained from an external source, and visualization information is generated to visualize the obtained related information and the aforementioned segmented utterance text in a list. The visualization information generating device described in any one of the appendices 2 to 9, wherein the related information includes at least one of the following: a search keyword for the FAQ used by the speaker of the utterance text when he / she spoke; the browsing history of the FAQ used by the speaker of the utterance text when he / she spoke; and history information of other inquiries made by the speaker of the utterance text when he / she spoke.
[0168] (Note 11) The aforementioned processor, The visualization information generation device according to Appendix 1, which has a specification unit that acquires evaluation information for the utterance text from an external source and identifies proposed revisions to the utterance text or the script based on the acquired evaluation information and information representing at least one of compliance and non-compliance.
[0169] (Note 12) The aforementioned processor, A visualization information generating device according to Appendix 3, which identifies proposed modifications to the spoken text or the script based on information representing at least one of compliance and non-compliance, and the score.
[0170] (Note 13) The aforementioned processor, Among the speech content represented by speech texts whose evaluation information is rated above a predetermined level, the speech content that is presumed to be non-compliant with the speech content represented by the script is identified as a proposed revision representing the speech content to be added to the script. Among the speech content represented by speech texts whose evaluation information is below a predetermined level, the speech content that is presumed to be consistent with the speech content represented by the script is identified as a proposed revision representing the speech content that should be deleted from the script. A visualization information generating device as described in Appendix 11, which identifies, among the speech content represented by speech texts whose evaluation information is below a predetermined evaluation, speech content that is estimated to be non-conforming to the speech content represented by the script as a proposed revision representing unnecessary speech content in the speech text.
[0171] (Note 14) The aforementioned processor, A visualization information generating device according to any one of the appendices 11 to 13, which generates visualization information for visualizing the aforementioned revised proposal.
[0172] (Note 15) A non-temporary storage medium that stores a program executable by a computer to perform visualization information generation processing, The aforementioned visualization information generation process is: Taking information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, visualization information is generated to visualize the range of the utterance content represented by either the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. Non-transitory storage medium.
[0173] [References] Reference 1: Katsuki Chousa, Masaaki Nagata, Masaaki Nishino. Bilingual Text Extraction as Reading Comprehension, arXiv:2004.14517v1. Reference 2: Jacob Devlin, Ming-Wei Chang, Kenton Lee, Kristina Toutanova. BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, arXiv:1810.04805v2. Reference 3: Masaaki Nagata, Chousa Katsuki, Masaaki Nishino. A Supervised Word Alignment Method based on Cross-Language Span Prediction using Multilingual BERT, arXiv:2004.14516v1. [Explanation of Symbols]
[0174] 1. Contact Center System 10 Estimation device 20 Operator terminals 30 Supervisor terminals 40 PBX 50 Customer terminals 60 Communication Networks 101 Input Device 102 Display device 103 External I / F 103a Recording medium 104 Communication I / F 105 Processors 106 Memory device 107 Bus 201 Speech Recognition Unit 202 Compliant Estimation Processing Unit 203 Storage section 211 Split section 212 Matching Department 213 Corresponding Information Generation Unit 214 Reference Estimation Unit 215 Compliance Scope Visualization Section 216 Aggregation Department 217 Compliance Status Visualization Section 218 Evaluation Department 219 Amendment Specification Section 220 Amendment Visualization Department 221 Compliance Rate Visualization Section
Claims
1. A visualization information generation unit takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. It has, The utterance text and the script are pre-divided into one or more segmented utterance texts and one or more segmented scripts, respectively. The information representing at least one of compliance and non-compliance is information representing at least one of compliance and non-compliance between the divided utterance text and the divided script, The aforementioned visualization information generation unit, A visualization information generation device that generates visualization information for visualizing a list of utterances of segmented utterance texts that are presumed to be non-conforming to the utterance content represented by the segmented script.
2. A visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance. It has, The aforementioned visualization information generation unit, A visualization information generation device that generates visualization information for visualizing a list of speech content of speech texts that are presumed to be non-conforming to the speech content represented by the script.
3. A visualization information generation unit that takes an utterance text and information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance. It has, The aforementioned visualization information generation unit, A visualization information generating device that generates visualization information for visualizing, in association with the utterance text and information representing at least one of the conformance and non-conformance between the utterance text and the script, in a list.
4. A visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance. It has, The script is a first utterance that the user, as the speaker, needs to utter, or a second utterance that the user is expected to utter. A visualization information generating device, wherein the visualization information includes the utterance text or the first utterance or the second utterance, and when displaying the utterance text or the first utterance or the second utterance, the information is for visualizing the range that is presumed to conform in a manner different from the range that is presumed to be non-conforming.
5. A visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance, An identification unit takes evaluation information, which is information representing the evaluation result of a predetermined evaluation performed on the utterance text, as input, and identifies proposed revisions to the utterance text or the script based on the evaluation information and information representing at least one of compliance and non-compliance, It has, The specified part is, A visualization information generating device that identifies, among the utterances represented by utterance text that meet predetermined criteria, the utterances that are estimated to be non-conforming to the utterances represented by the script, as proposed revisions representing utterances to be added to the script.
6. A visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance, An identification unit takes evaluation information, which is information representing the evaluation result of a predetermined evaluation performed on the utterance text, as input, and identifies proposed revisions to the utterance text or the script based on the evaluation information and information representing at least one of compliance and non-compliance, It has, The specified part is, A visualization information generating device that identifies, among the speech content represented by speech texts whose evaluation information does not meet predetermined criteria, speech content within the range that is presumed to conform to the speech content represented by the script as a proposed revision representing speech content that should be deleted from the script.
7. A visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance, An identification unit takes evaluation information, which is information representing the evaluation result of a predetermined evaluation performed on the utterance text, as input, and identifies proposed revisions to the utterance text or the script based on the evaluation information and information representing at least one of compliance and non-compliance, It has, The specified part is, A visualization information generating device that identifies, among the speech content represented by speech texts in which the evaluation information does not meet predetermined criteria, speech content that is presumed to be non-conforming to the speech content represented by the script as a proposed revision representing unnecessary speech content in the speech text.
8. A visualization information generation unit that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is estimated to conform to the utterance content represented by the other, in a manner different from the range that is estimated to be non-conformance. It has, The aforementioned visualization information generation unit, Taking related information associated with the utterance text as input, visualization information is generated to visualize the related information and the utterance text in a list, associating them with each other. A visualization information generating device, wherein the related information includes at least one of the following: a search keyword for the FAQ used when the first speaker of the utterance text spoke; a browsing history of the FAQ used when the first speaker spoke; a history of inquiries made to other parties when the first speaker spoke; and information concerning the second speaker of the utterance text.
9. The visualization information generating device according to any one of claims 1 to 8, wherein the range that is presumed to conform is the range in the token sequence constituting the utterance text that conforms to the utterance content represented by the script, or the range in the token sequence constituting the script that conforms to the utterance content represented by the utterance text.
10. The visualization information generating device according to any one of claims 5 to 7, wherein the evaluation evaluates the quality of the utterance text or the content of a call corresponding to the utterance text, and includes an evaluation of the number of turns between the first user and the second user who uttered the utterance content represented by the divided utterance text that constitutes the utterance text, for each utterance text or each divided utterance text that constitutes the utterance text, an automatic evaluation by a machine learning model on a sentence or scene basis, or an evaluation of the validity or paraphrasability of the second user's utterance based on the first user's response.
11. The visualization information generation device according to any one of claims 1 to 10, wherein the visualization information is screen information for displaying the range estimated to be compliant and the range estimated to be non-compliant on a user interface.
12. The system includes a compliance rate calculation unit that calculates a score representing the degree of compliance between the utterance content represented by the divided utterance text and the utterance content represented by the divided script, based on information representing at least one of the compliance and non-compliance, for any given divided utterance text, a score that includes at least a degree of relevance representing the degree to which the divided utterance text is estimated to be compliant. The aforementioned visualization information generation unit, A visualization information generation device according to claim 1, which generates visualization information for visualizing the score for each divided script of the same script.
13. A compliance rate calculation unit that calculates a score representing the degree of compliance between the utterance content represented by the divided utterance text and the utterance content represented by the divided script, based on information representing at least one of compliance and non-compliance, The aforementioned visualization information generation unit, The visualization information generating device according to claim 1, which generates visualization information for visualizing the score for each segmented utterance text of the same utterance text.
14. The aforementioned visualization information generation unit, A visualization information generating device according to claim 1, which generates visualization information for visualizing a list of utterances of segmented utterance texts that are presumed to be non-conforming to the utterance content represented by the segmented script, for each utterance subject unit of the utterance text.
15. The aforementioned visualization information generation unit, The visualization information generating device according to claim 14, which generates visualization information for visualizing a list of utterances of divided utterance texts that are presumed to be non-conforming to the utterance content represented by the divided script, and which generates visualization information for visualizing the items before and after the item corresponding to the divided script, or the divided utterance texts before and after the divided utterance text.
16. The aforementioned visualization information generation unit, Taking related information associated with the aforementioned utterance text as input, visualization information is generated to visualize the related information and the segmented utterance text in a list, associating them with each other. The visualization information generating device according to claim 14 or 15, wherein the related information includes at least one of the following: a search keyword for the FAQ used by the speaker of the utterance text when he / she spoke; a browsing history of the FAQ used by the speaker of the utterance text when he / she spoke; and a history of other inquiries made by the speaker of the utterance text when he / she spoke.
17. The system includes an identification unit that takes evaluation information, which is information representing the evaluation result of a predetermined evaluation performed on the utterance text, as input, and identifies proposed revisions to the utterance text or the script based on the evaluation information and information representing at least one of compliance and non-compliance. The visualization information generating apparatus according to claim 1, wherein the proposed modification includes at least one of the proposed addition of speech content based on the speech text to be added to the script, the proposed deletion of speech content to be removed from the script, and the proposed modification of speech content represented by the speech text that is not compliant with the script.
18. The visualization information generation device according to claim 1, which takes evaluation information, which is information representing an evaluation result obtained by performing a predetermined evaluation on the utterance text, as input, and has a specification unit that identifies proposed modifications to parts of the utterance text that would lower the evaluation result or have the potential to lower the evaluation result, based on the evaluation information and information representing at least one of compliance and non-compliance.
19. A specification unit that takes evaluation information, which is information representing the evaluation result of performing a predetermined evaluation on the utterance text, as input, and identifies proposed revisions to the utterance text or the script based on the evaluation information and information representing at least one of compliance and non-compliance, The specified part is, Among the speech content represented by speech texts whose evaluation information is rated above a predetermined level, the speech content that is presumed to be non-compliant with the speech content represented by the script is identified as a proposed revision representing the speech content to be added to the script. Among the speech content represented by speech texts whose evaluation information is below a predetermined level, the speech content that is presumed to be consistent with the speech content represented by the script is identified as a proposed revision representing the speech content that should be deleted from the script. A visualization information generating device according to any one of claims 1 to 4, wherein, among the speech content represented by a speech text whose evaluation information is below a predetermined evaluation, the speech content that is estimated to be non-compliant with the speech content represented by the script is identified as a proposed revision that represents unnecessary speech content in the speech text.
20. A visualization information generation procedure that takes information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. The computer executes this, The utterance text and the script are pre-divided into one or more segmented utterance texts and one or more segmented scripts, respectively. The information representing at least one of compliance and non-compliance is information representing at least one of compliance and non-compliance between the divided utterance text and the divided script, The aforementioned visualization information generation procedure is as follows: A method for generating visualization information, which generates visualization information to visualize a list of utterances of segmented utterance texts that are presumed to be non-conforming to the utterance content represented by the segmented script.
21. A visualization information generation procedure that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. The computer executes this, The aforementioned visualization information generation procedure is as follows: A method for generating visualization information, which generates visualization information for visualizing a list of speech content of speech texts that are presumed to be non-conforming to the speech content represented by the script.
22. A visualization information generation procedure that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. The computer executes this, The aforementioned visualization information generation procedure is as follows: A method for generating visualization information, which generates visualization information for visualizing a list of the utterance text and information representing at least one of the conformance and non-conformance between the utterance text and the script, in association with each other.
23. A visualization information generation procedure that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. The computer executes this, The script is a first utterance that the user needs to speak, or a second utterance that the user is expected to speak. A method for generating visualization information, wherein the visualization information includes the utterance text or the first utterance or the second utterance, and when displaying the utterance text or the first utterance or the second utterance, the information is for visualizing the range that is presumed to conform in a manner different from the range that is presumed to not conform.
24. A visualization information generation procedure that takes information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance, A procedure for identifying proposed revisions to the utterance text or script, based on evaluation information, which is information representing the evaluation results of a predetermined evaluation performed on the utterance text, and based on the evaluation information and information representing at least one of compliance and non-compliance, The computer executes this, The aforementioned identification procedure is, A visualization information generation method that identifies, among the utterances represented by utterance text that meet predetermined criteria, the utterances that are estimated to be non-conforming to the utterances represented by the script, as proposed revisions representing utterances to be added to the script.
25. A visualization information generation procedure that takes information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance, A procedure for identifying proposed revisions to the utterance text or script, based on evaluation information, which is information representing the evaluation results of a predetermined evaluation performed on the utterance text, and based on the evaluation information and information representing at least one of compliance and non-compliance, The computer executes this, The aforementioned identification procedure is, A visualization information generation method that identifies, among the utterances represented by utterance texts that do not meet predetermined criteria, the utterances within the range that are presumed to conform to the utterances represented by the script, as proposed revisions representing utterances that should be deleted from the script.
26. A visualization information generation procedure that takes information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script as input, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance, A procedure for identifying proposed revisions to the utterance text or script, based on evaluation information, which is information representing the evaluation results of a predetermined evaluation performed on the utterance text, and based on the evaluation information and information representing at least one of compliance and non-compliance, The computer executes this, The aforementioned identification procedure is, A visualization information generation method that identifies, among the speech content represented by speech texts in which the evaluation information does not meet predetermined criteria, speech content that is presumed to be non-conforming to the speech content represented by the script as a proposed revision representing unnecessary speech content in the speech text.
27. A visualization information generation procedure that takes as input information representing at least one of the conformance and non-conformance between the utterance content represented by the utterance text and the utterance content represented by a predetermined script, and generates visualization information for visualizing the range of the utterance content represented by one of the utterance text or the script that is presumed to conform to the utterance content represented by the other, in a manner different from the range that is presumed to be non-conformance. The computer executes this, The aforementioned visualization information generation procedure is as follows: Taking related information associated with the utterance text as input, visualization information is generated to visualize the related information and the utterance text in a list, associating them with each other. A method for generating visualization information, wherein the related information includes at least one of the following: a search keyword for the FAQ used by the speaker of the utterance text when he / she spoke; a browsing history of the FAQ used by the speaker of the utterance text when he / she spoke; and a history of other inquiries made by the speaker of the utterance text when he / she spoke.
28. A program that causes a computer to function as a visualization information generation device according to any one of claims 1 to 19.
Citation Information
Patent Citations
Operator business support system
JP2008123447A
Talk script use state calculation system and talk script use state calculation program
JP2012003702A
Business support system
JP2015156062A
Telephone conversation content analysis display device, telephone conversation content analysis display method, and program
JP2016143909A