Display data generating device, display data generating method, and display data generating program

The display data generation device enhances visualization of annotation information by correlating text with background colors, improving user comprehension of customer interactions in contact centers.

JP7758965B2Active Publication Date: 2025-10-23NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023509990
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-03-30
Publication Date
2025-10-23
Estimated Expiration
2041-03-30

AI Technical Summary

Technical Problem

Existing systems fail to effectively visualize annotation information, such as scene estimation results, making it difficult for users to recognize and understand the context of customer interactions in contact centers.

Method used

A display data generation device and method that determines background colors and positions to display annotation information alongside text, allowing for a clear correspondence between the text and annotation information, thereby enhancing visualization.

Benefits of technology

Enables intuitive recognition of annotation information, facilitating better understanding of customer interactions by clearly associating text with its corresponding context.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007758965000001
    Figure 0007758965000001
  • Figure 0007758965000002
    Figure 0007758965000002
  • Figure 0007758965000003
    Figure 0007758965000003
Patent Text Reader

Abstract

A display data generation device (1) according to the present disclosure comprises: an input unit (11) that receives input of a text series and target data which contains pieces of annotation information corresponding to respective text included in the text series; and a display preparation unit (14) that determines a background color of a display screen of a display device (4) and annotation expression information indicating the position and the range in which the background color is displayed in order to express the correspondence between a text and annotation information when the display device displays the text on the basis of the annotation information, and that generates display data which is used to display the text series and the pieces of annotation information in accordance with the series in the text series, and which is used to display the background color indicated by the annotation expression information at the position and in the range indicated by the annotation expression information.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present disclosure relates to a display data generating device, a display data generating method, and a display data generating program. [Background technology]

[0002] Contact center operators are required to receive inquiries from customers about products, services, etc., and to provide support to solve customer problems. In order to analyze customer inquiries and improve the quality of responses, operators create a history of customer interactions and share this information within the contact center.

[0003] Non-Patent Document 1 discloses a system that supports operators by presenting appropriate information to the operator on duty based on the purpose of a customer who has called a contact center (call center). The system disclosed in Non-Patent Document 1 displays the text of utterances between the operator and the customer on the left side of the screen, and displays similar questions with high scores and their answers from FAQs searched from the utterance text indicating the customer's purpose or the utterance text confirming the operator's purpose on the right side of the screen. Furthermore, Non-Patent Document 1 infers the scene for each utterance, then extracts keywords by narrowing down to only the utterances from the specified scene, and searches the FAQs. (A scene is a classification of spoken text according to the type of situation in a conversation between an operator and a customer. For example, a conversation begins with the operator introducing themselves, followed by the customer explaining why they called, the operator confirming the reason, confirming the contract holder and contract details, and then responding to the customer's request, ending with a thank you. This can be classified into scenes such as "opening," "understanding the inquiry," "response," and "closing." The results of this scene estimation are assigned as labels to the spoken text.) [Prior art documents] [Non-patent literature]

[0004] [Non-Patent Document 1] Takaaki Hasegawa and three others, "Automatic Knowledge Support System for Supporting Operators' Responses," NTT Technical Journal, pp. 16-19, 2019, vol. 31, No. 7 Summary of the Invention [Problem to be solved by the invention]

[0005] In the technology described in Non-Patent Document 1, the user refers to the spoken text of the operator and the customer, as well as similar questions (with high scores in FAQs automatically searched from the spoken text conveying the customer's needs or the spoken text confirming the operator's needs) and their answers. However, labels (annotation information) such as scene estimation results are not presented, making it difficult to visualize the annotation information in a way that is easy for users to recognize.

[0006] The purpose of the present disclosure, made in consideration of the above-mentioned problems, is to provide a display data generation device, a display data generation method, and a display data generation program that can visualize annotation information. [Means for solving the problem]

[0007] In order to solve the above problem, the present disclosure provides a display system that includes an input unit that accepts input of target data including a text sequence according to the present disclosure and annotation information corresponding to each piece of text included in the text sequence; and a display preparation unit that determines, based on the annotation information, a background color of a display screen of the display device and annotation expression information that indicates a position and range in which to display the background color, in order to express a correspondence between the text and the annotation information when the display device displays the text, and generates display data for displaying the text sequence and the annotation information according to a sequence in the text sequence, the display data being for displaying the background color indicated by the annotation expression information at the position and range indicated by the annotation expression information.

[0008] Furthermore, in order to solve the above-mentioned problem, a display data generation method according to the present disclosure includes the steps of: accepting input of target data including a text sequence and annotation information corresponding to each piece of text included in the text sequence; determining, based on the annotation information, a background color of a display screen of the display device and annotation expression information indicating a position and range in which to display the background color, in order to express a correspondence between the text and the annotation information when the display device displays the text; and generating display data for displaying the text sequence and the annotation information according to a sequence in the text sequence, the display data being for displaying the background color indicated by the annotation expression information at the position and range indicated by the annotation expression information.

[0009] In order to solve the above problem, a display data generation program according to the present disclosure causes a computer to function as the display data generation device described above. [Effects of the Invention]

[0010] According to the display method, display data generating device, and display data generating program of the present disclosure, annotation information can be visualized. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is an overall schematic diagram of a display data generating device according to a first embodiment. [Figure 2] 2 is a diagram showing an example of target data input received by an input unit shown in FIG. 1; FIG. [Figure 3] 2 is a diagram showing an example of correspondence between annotation information and colors stored in a color storage unit shown in FIG. 1; FIG. [Figure 4] 2 is a diagram showing an example of display data generated by a display preparation unit shown in FIG. 1. FIG. [Figure 5]2 is an example of a screen displayed by the display data output unit shown in FIG. 1. [Figure 6] 2 is a flowchart showing an example of the operation of the display data generating device shown in FIG. [Figure 7] FIG. 10 is an overall schematic diagram of a display data generating device according to a second embodiment. [Figure 8] FIG. 8 is a diagram showing an example of a gradation rule stored in a gradation rule storage unit shown in FIG. 7. [Figure 9] 8 is a diagram showing an example of display data generated by a display preparation unit shown in FIG. 7. FIG. [Figure 10] 8 is an example of a screen displayed by the display data output unit shown in FIG. 7. [Figure 11] 8 is a flowchart showing an example of the operation of the display data generating device shown in FIG. 7. [Figure 12] FIG. 10 is an overall schematic diagram of a display data generating device according to a third embodiment. [Figure 13] 13 is a diagram showing an example of target data input received by the input unit shown in FIG. 12. FIG. [Figure 14] 13 is a diagram showing an example of a gradation rule stored in a gradation rule storage unit shown in FIG. 12. FIG. [Figure 15] 15 is a diagram for explaining in detail the annotation expression information determined by the gradation rule shown in FIG. 14. FIG. [Figure 16] 13 is a diagram showing an example of display data generated by a display preparation unit shown in FIG. 12. FIG. [Figure 17] 13 is an example of a screen displayed by the display data output unit shown in FIG. 12. [Figure 18] 13 is a flowchart showing an example of the operation of the display data generating device shown in FIG. [Figure 19] 8 is an example of a screen displayed by a first modified example of the display data output unit shown in FIG. 7. [Figure 20] 8 is an example of a screen displayed by a second modified example of the display data output unit shown in FIG. 7. [Figure 21] 8 is an example of a screen displayed by a third modified example of the display data output unit shown in FIG. 7. [Figure 22] 8 is an example of a screen displayed by a fourth modified example of the display data output unit shown in FIG. 7. [Figure 23] 10 is an example of a screen displayed by a fifth modified example of the display data output unit shown in FIG. 7. [Figure 24] FIG. 2 is a hardware block diagram of the display data generating device. DETAILED DESCRIPTION OF THE INVENTION

[0012] First, an embodiment of the present disclosure will be described with reference to the drawings.

[0013] First Embodiment The overall configuration of the first embodiment will be described with reference to Fig. 1. Fig. 1 is a schematic diagram of a display data generating device 1 according to this embodiment.

[0014] (Functional configuration of the display data generating device) As shown in FIG. 1, the display data generating device 1 according to the first embodiment includes an input unit 11, a target data storage unit 12, a display rule storage unit 13, a display preparation unit 14, a display data storage unit 15, and a display data output unit 16. The input unit 11 is configured with an input interface that accepts input of information. The input interface may be a keyboard, a mouse, a microphone, or the like, or may be an interface for accepting information received from another device via a communication network. The target data storage unit 12, the display rule storage unit 13, and the display data storage unit 15 are configured with, for example, a ROM or storage. The display preparation unit 14 constitutes a control unit (controller). The control unit may be configured with dedicated hardware such as an ASIC (Application Specific Integrated Circuit) or an FPGA (Field-Programmable Gate Array), a processor, or both. The display data output unit 16 is configured with an output interface that outputs information.

[0015] The input unit 11 accepts input of target data including a text sequence and annotation information corresponding to each piece of text included in the text sequence, as shown in FIG. 2 . The target data may further include a text ID (identifier) ​​for identifying the text. The target data may further include a sequence order in which each piece of spoken text is arranged. The sequence order is information indicating an order when there is an order between the texts included in the text sequence. In each embodiment, the text may be text obtained by speech recognition of speech data, text obtained by transcription of speech, text included in a chat, text of minutes of a meeting, text of a story, etc., but is not limited to this. The sequence order is information for arranging the utterances of multiple speakers in chronological order in a voice dialogue or chat between multiple speakers. Furthermore, the sequence order is the order of text in a sentence in text of minutes of a meeting or a story. The sequence order can be a meaningful order for arranging text from the beginning to the end in a text sequence. In this embodiment, the sequence order is indicated by a text ID, but is not limited to this. It is not essential that the target data includes a text ID, and in a configuration in which the target data does not include a text ID, the spoken text may include information indicating the sequence order.

[0016] A spoken text is a text indicating the content of an utterance made by each of multiple speakers in a dialogue conducted by the multiple speakers. A single spoken text is a text output in units of speech completion (units in which it is determined whether an operator or a customer has finished speaking or whether they have fully expressed what they wanted to say) based on the results of speech recognition. The spoken text may be text data. The multiple speakers may be, for example, an operator at a call center and a customer making an inquiry to the call center. The following describes an example of target data including annotation information related to a dialogue between an operator and a customer. However, in each embodiment described in this specification, the multiple speakers who utter uttered texts included in the target data are not limited to an operator and a customer. A single spoken text is a segment of a spoken text uttered by any one of the multiple speakers. A segment of a spoken text may be determined by an arbitrary rule, by the operation of the speaker who uttered the utterance text, or by a computer that performs speech recognition using an arbitrary algorithm. If the text is spoken text, it may further include speaker information indicating the speaker who uttered the spoken text. Also, if the text is spoken text, a text ID for identifying the spoken text is called an utterance ID. In the following, although the explanation will be given using spoken text as an example of text, the text included in the target data processed by the display data generation device of this embodiment is not limited to spoken text and can be any text.

[0017] The annotation information refers to information (metadata) associated with each utterance text. The annotation information may be the topic of the utterance text, the scene in which the utterance text was spoken, or some kind of classification label.

[0018] The target data storage unit 12 stores the target data input by the input unit 11.

[0019] The display rule storage unit 13 stores rules that are used by the display preparation unit 14 to determine annotation expression information for spoken text based on annotation information.

[0020] The annotation expression information is information indicating the background color of the display screen of the display device 4 and the position and range in which the background color is displayed, in order to express the correspondence between the uttered text and the annotation information when the display device 4 displays the uttered text. The position and range in which the background color is displayed may include the display position and display range of the annotation information, respectively. In the first embodiment, the annotation expression information is the background color of the annotation information.

[0021] The display rule storage unit 13 includes a color storage unit 131. The color storage unit 131 stores rules indicating the association between annotation information and annotation expression information. In the first embodiment, as shown in FIG. 3, the color storage unit 131 stores color arrangement rules indicating the association between annotation information and annotation expression information (background color of the display screen). The annotation expression information associated with annotation information in the color arrangement rules may be determined by a computer using an arbitrary algorithm, or may be determined by an administrator of the display data generating device 1.

[0022] Based on the annotation information, the display preparation unit 14 determines a background color of the display screen of the display device 4 and annotation expression information indicating the position and range in which to display the background color, in order to express the correspondence between the text and the annotation information when the display device 4 displays the spoken text. The display preparation unit 14 may divide the spoken text and determine the annotation expression information for the divided spoken text. Hereinafter, the divided spoken text will be referred to as a "divided spoken text." When divided spoken text and undivided spoken text are to be distinguished, the divided spoken text will be referred to as a "divided spoken text," and the undivided spoken text will be simply referred to as a "spoken text." However, when divided spoken text and undivided spoken text are not to be distinguished, both the divided spoken text and the undivided spoken text may be simply referred to as a "spoken text."

[0023] Specifically, first, the display preparation unit 14 divides the utterance text included in the target data input by the input unit 11. The display preparation unit 14 can divide the utterance text using any algorithm. At this time, the display preparation unit 14 uniquely identifies the divided utterance text and assigns a determination unit ID indicating the utterance text sequence of the divided utterance text. For example, the display preparation unit 14 may divide the utterance text into a portion before a period and a portion after the period. In the example shown in FIG. 2, the utterance text corresponding to utterance ID "1" is "I'm BB from AA Insurance. Is CC at home?" Therefore, the display preparation unit 14 divides this utterance text at the period into "I'm BB from AA Insurance." and "I'm CC at home?" as shown in FIG. 4, and associates the determination unit ID "1" and the determination unit ID "2" with each of them. Furthermore, the display preparation unit 14 determines that the annotation information of the divided utterance text is the annotation information of the original utterance text. In the example shown in Fig. 4, the display preparation unit 14 determines that the topic, which is the annotation information of the utterance text corresponding to the determination unit IDs "1" and "2", is "opening".

[0024] As described above, the display preparation unit 14 divides the spoken text into a portion before the punctuation mark and a portion after the punctuation mark, but this is not limited to this. For example, the display preparation unit 14 may divide the spoken text for each word, or may divide the spoken text into a portion before the punctuation mark and a portion after the punctuation mark. Note that the display preparation unit 14 does not have to divide the spoken text. In such a configuration, for example, the spoken text included in the target data may be undivided spoken text.

[0025] The display preparation unit 14 forms a group (hereinafter referred to as an "utterance text group") consisting of utterance texts that have the same annotation information and are consecutive when arranged in the above-described sequential order. The display preparation unit 14 determines annotation expression information that indicates a color corresponding to the utterance text group, using the coloring rule stored in the color storage unit 131. Specifically, the display preparation unit 14 determines that the annotation expression information of the utterance text group is a color that corresponds to the annotation information of the utterance text group in the coloring rule.

[0026] Furthermore, after determining the annotation expression information for the utterance text group, the display preparation unit 24 determines whether the annotation expression information for all utterance texts has been determined. If the display preparation unit 24 determines that the annotation expression information for some utterance texts has not been determined, the display preparation unit 24 forms an utterance text group for the utterance texts for which the annotation expression information has not been determined, and repeats the process of determining the annotation expression information for the utterance text group. If the display preparation unit 24 determines that the annotation expression information for all utterance texts has been determined, the display preparation unit 24 generates display data for displaying the text sequence and the annotation expression information in accordance with the sequential order in the text sequence, and for displaying the background color indicated by the annotation expression information at the position and range indicated by the annotation expression information. The display data may include, for example, a determination unit ID, speaker information, utterance text, annotation information, and annotation expression information, as shown in FIG. 4.

[0027] The display data storage unit 15 stores the display data generated by the display preparation unit 14.

[0028] The display data output unit 16 outputs display data. The display data output unit 16 may output the display data to a display device 4 such as a liquid crystal panel or an organic EL display, or may output the display data to another device via a communication network.

[0029] As a result, the display device 4 displays a display screen based on the display data. Specifically, as shown in FIG. 5, the display device 4 displays the utterance text included in the display data in the above-described utterance text sequence. The display device 4 then displays annotation information corresponding to the utterance text in association with the utterance text, and further displays the background of the annotation information in a color indicated by the annotation expression information included in the display data. The display device 4 may also display one or more of an utterance ID and speaker information in association with the utterance text and the annotation information. Note that the gray color displayed as the background of the "Opening" section, the green color displayed as the background of the "Accident Situation" section, the blue color displayed as the background of the "Injury Situation" section, and the orange color displayed as the background of the "Injury Situation" section are indicated by different black and white binary hatching in FIG. 5. Furthermore, as described above, since the annotation information includes scenes, the display device 4 can display the utterances collectively by scene, which allows the operator to grasp the overall flow of the dialogue in order to understand the dialogue. When the display data output unit 16 transmits the display data to another device via a communication network, the other device displays a display screen based on the display data, just like the display device 4.

[0030] (Operation of the display data generator) Here, the operation of the display data generating device 1 according to the first embodiment will be described with reference to Fig. 6. Fig. 6 is a flowchart showing an example of the operation of the display data generating device 1 according to the first embodiment. The operation of the display data generating device 1 described with reference to Fig. 6 corresponds to the display method of the display data generating device 1 according to the first embodiment.

[0031] In step S11, the input unit 11 receives input of target data including an utterance text sequence and annotation information corresponding to each of the texts included in the utterance text sequence. In this example, the target data further includes an utterance ID.

[0032] In step S12, the display preparation unit 14 divides the utterance text included in the target data input by the input unit 11.

[0033] In step S13, the display preparation unit 14 forms a spoken text group made up of consecutive spoken texts having the same annotation information.

[0034] In step S14, the display preparation unit 14 determines, based on the annotation information and the sequential order, annotation expression information indicating the background color of the display screen of the display device 4 and the position and range in which to display the background color, in order to express the correspondence between the utterance text and the annotation information when the utterance text is displayed by the display device 4. In this example, the display preparation unit 14 determines, based on the annotation information, annotation expression information indicating a color corresponding to the utterance text group.

[0035] In step S15, the display preparation unit 14 determines whether or not annotation expression information corresponding to all of the utterance text groups has been determined.

[0036] If it is determined in step S15 that annotation expression information corresponding to some of the utterance text groups has not been determined, the process returns to step S13, and the display preparation unit 14 repeats the process. If it is determined in step S15 that annotation expression information corresponding to all of the utterance text groups has been determined, the display preparation unit 14 generates display data in step S16 for displaying the utterance text sequence and the annotation information according to the sequence in the utterance text sequence, and for displaying the background color indicated by the annotation expression information at the position and in the range indicated by the annotation expression information.

[0037] In step S17, the display data storage unit 15 stores the display data.

[0038] Thereafter, the display data output unit 16 outputs the display data at an arbitrary timing. The display data output unit 16 may output the display data to a display device 4 such as a liquid crystal panel or an organic EL display, or may output the display data to another device via a communication network. The arbitrary timing may be, for example, the timing when a display command is input to the input unit 11 by a user's operation. As a result, the display device 4 displays a display screen based on the display data. Specifically, the display device 4 displays the spoken text and annotation information based on the display data, and displays the background color indicated by the annotation expression information at the position and in the range indicated by the annotation expression information.

[0039] In the above description, the display data generating device 1 executes the process of step S12, but this is not a limitation. For example, the display data generating device 1 does not have to execute the process of step S12.

[0040] As described above, according to the first embodiment, the display data generating device 1 determines, based on the annotation information, the background color of the display screen of the display device 4 and annotation expression information indicating the position and range in which to display the background color, in order to represent the correspondence between the utterance text and the utterance annotation information when the display device 4 displays the utterance text. The display data generating device 1 then generates display data for displaying the utterance text sequence and the annotation information according to the sequence in the utterance text sequence, and for displaying the background color indicated by the annotation expression information at the position and range indicated by the annotation expression information. This allows the user to intuitively grasp the annotation information from the background color of the display screen. Therefore, the content of the target data including the utterance text corresponding to the annotation information can be quickly recognized.

[0041] <Second embodiment> The overall configuration of the display data generating device 2 of the second embodiment will be described with reference to Fig. 7. Fig. 7 is a schematic diagram of the display data generating device 2 according to this embodiment.

[0042] (Functional configuration of the display data generating device) 7, the display data generating device 2 according to the second embodiment includes an input unit 21, a target data storage unit 22, a display rule storage unit 23, a display preparation unit 24, a display data storage unit 25, and a display data output unit 26. The input unit 21 is configured by an input interface that accepts input of information, similar to the input unit 11 of the first embodiment. The target data storage unit 22, the display rule storage unit 23, and the display data storage unit 25 are configured by memory, similar to the target data storage unit 12, the display rule storage unit 13, and the display data storage unit 15 of the first embodiment. Furthermore, the display preparation unit 24 and the display data output unit 26 constitute a control unit, similar to the display preparation unit 14 and the display data output unit 16 of the first embodiment.

[0043] The input unit 21 and the target data storage unit 22 are similar to the input unit 11 and the target data storage unit 12 of the display data generating device 2 according to the first embodiment. In the second embodiment, the input unit 21 accepts input, and the target data stored in the target data storage unit 22 further includes a sequence order in addition to the text sequences included in the target data of the first embodiment and annotation information corresponding to each of the texts included in the text sequences.

[0044] The display rule storage unit 23 includes a color storage unit 231 and a gradation rule storage unit 232. The color storage unit 231 stores color scheme rules, similar to the color storage unit 131 of the display data generating device 1 according to the first embodiment. In the color scheme rules of the second embodiment, the colors corresponding to the respective pieces of annotation information may be different or the same. In the following specific examples, the annotation information is a topic.

[0045] The gradation rule storage unit 232 stores gradation rules for determining annotation expression information. As shown in Fig. 8, the gradation rules in the second embodiment are rules that indicate gradations corresponding to annotation information and series. In the second embodiment, the annotation expression information is information that indicates colors and gradations.

[0046] In the example gradation rule shown in FIG. 8, when the spoken text included in the spoken text group includes the first spoken text in the target data but does not include the last spoken text, the annotation expression information corresponding to the utterance text group is a gradation that continuously changes from the color corresponding to the topic to white from the start point to the end point. Here, when the utterances included in the target data are displayed in an utterance text sequence in an arrangement direction (from top to bottom in the example shown in FIG. 10, which will be referred to later), the start point is the end of the column displaying the topic on the start side of the arrangement direction (the upper end in the example shown in FIG. 10). The end point is the end of the column displaying the topic on the end side of the arrangement direction (the lower end in the example shown in FIG. 10). The color corresponding to the topic is a color stored in association with the topic in the color scheme rule.

[0047] Furthermore, in the example gradation rule shown in Figure 8, if the utterance texts included in the utterance text group do not include the first utterance text in the target data and do not include the last utterance text, the annotation expression information corresponding to the utterance text group is a gradation that continuously changes from white to the color corresponding to the topic from the starting point toward the midpoint, and continuously changes from the color corresponding to the topic to white from the midpoint toward the end point.

[0048] Furthermore, in the example gradation rule shown in Figure 8, if the speech texts included in the speech text group do not include the first speech text in the target data but do include the last speech text, the annotation expression information corresponding to the speech text group is a gradation that continuously changes from white to the color corresponding to the topic as it moves from the starting point to the end point.

[0049] Furthermore, in the example gradation rule shown in Figure 8, if the spoken text included in the spoken text group includes the first spoken text in the target data and the last spoken text, the annotation expression information corresponding to the spoken text group has no gradation.

[0050] However, the gradation rule is not limited to the example shown in Fig. 8, and can be any rule that does not clearly change the color corresponding to the topic. For example, in another example gradation rule, if the spoken texts included in an utterance text group do not include the first utterance text in the target data and do not include the last utterance text, the annotation expression information corresponding to the utterance text group may be a gradation that continuously changes from the color corresponding to the topic to white from the start point toward the midpoint, and continuously changes from white to the color corresponding to the topic from the midpoint toward the end point.

[0051] The display preparation unit 24 determines annotation expression information of the utterance text corresponding to the utterance text sequence and the annotation information based on the annotation information and the utterance text sequence. At this time, the display preparation unit 24 may divide the utterance text and determine the annotation expression information based on the divided utterance text, the annotation information of the utterance text, and the utterance text sequence.

[0052] Specifically, first, the display preparation unit 24, like the display preparation unit 14 of the first embodiment, divides the utterance text included in the target data input by the input unit 11. Note that, like the display preparation unit 14 of the first embodiment, the display preparation unit 24 does not have to perform the process of dividing the utterance text. In such a configuration, for example, the utterance text included in the target data may be divided utterance text.

[0053] The display preparation unit 24 forms utterance text groups in the same manner as the display preparation unit 14 of the first embodiment. In the example shown in FIG. 9, the display preparation unit 24 forms a group constituted by utterance texts corresponding to determination unit IDs "1" to "6" having the same annotation information "opening". The display preparation unit 24 also forms a group constituted by utterance texts corresponding to determination unit IDs "7" and "8" having the same annotation information "accident situation". Similarly, the display preparation unit 24 forms a group constituted by utterance texts corresponding to determination unit IDs "9" to "14" having the same annotation information "injury situation". Similarly, the display preparation unit 24 forms a group constituted by utterance texts corresponding to determination unit ID "15" having the same annotation information "repair situation".

[0054] The display preparation unit 24 determines the annotation expression information so that the background color gradually changes toward the boundary between the utterance text group and the adjacent utterance text group, where the annotation information differs between the adjacent utterance text group and the adjacent adjacent utterance text group. In this embodiment, the display preparation unit 24 determines the annotation expression information corresponding to the utterance text group using a coloring rule and a gradation rule.

[0055] In the example using the gradation rule shown in Fig. 8, when the utterance texts included in the utterance text group include the first utterance text in the target data but not the last utterance text, the display preparation unit 24 determines that the annotation expression information has a gradation that continuously changes from a color corresponding to the topic to white from the start point to the end point (a gradation from gray to white). As a result, as shown in Fig. 9, the display preparation unit 24 determines that the annotation expression information of the group made up of utterance texts corresponding to the determination unit IDs "1" to "6" has a gradation that continuously changes from gray to white from the start point to the end point. Here, gray is the color that corresponds to "opening" in the color scheme rule.

[0056] In the example using the gradation rule shown in FIG. 8 , if the utterance texts included in the utterance text group do not include the first utterance text in the target data and do not include the last utterance text, the display preparation unit 24 determines that the annotation expression information is a gradation in which the color continuously changes from white to a color corresponding to the topic from the starting point toward the midpoint, and continuously changes from the color corresponding to the topic to white from the midpoint toward the end point (a gradation with white at both ends and green in the center). Here, the midpoint is the midpoint between the starting point and the end point in the arrangement direction. As a result, as shown in FIG. 9 , the display preparation unit 24 determines that the annotation expression information of the group consisting of utterance texts corresponding to the determination unit IDs "7" and "8" is a gradation in which the color continuously changes from white to green from the starting point toward the midpoint, and continuously changes from green to white from the midpoint toward the end point. Here, green is the color corresponding to "accident situation" in the color scheme rule. Similarly, the display preparation unit 24 determines that the annotation expression information of the group consisting of the spoken text corresponding to the judgment unit IDs "9" to "14" is a gradation that changes continuously from white to blue from the start point to the midpoint, and continuously changes continuously from blue to white from the midpoint to the end point (a gradation with white on both ends and blue in the center). Here, blue is the color that corresponds to "injury status" in the color scheme rules.

[0057] 8, when the utterance text in the utterance text group does not include the first utterance text in the target data but does include the last utterance text, the display preparation unit 24 determines that the annotation expression information corresponding to the utterance text group has a gradation that continuously changes from white to a color corresponding to the topic from the start point to the end point. As a result, as shown in FIG. 9, the display preparation unit 24 determines that the annotation expression information of the group formed by the utterance text corresponding to the determination unit ID "15" has a gradation that continuously changes from orange to white from the start point to the end point (a gradation from white to orange). Here, orange is the color corresponding to "repair status" in the color scheme rule.

[0058] 8, when the utterance texts included in an utterance text group include the first utterance text in the target data and the last utterance text, the display preparation unit 24 determines that the annotation expression information corresponding to the utterance text group does not have gradation. Note that in the example of FIG. 8, there is no utterance text group that includes the first utterance text and the last utterance text.

[0059] After determining the annotation expression information for the utterance text group, the display preparation unit 24 determines whether the annotation expression information for all utterance texts has been determined. If the display preparation unit 24 determines that the annotation expression information for some utterance texts has not been determined, the display preparation unit 24 forms an utterance text group for the utterance texts for which the annotation expression information has not been determined, and repeats the process of determining the annotation expression information for the utterance text group. Furthermore, if the display preparation unit 24 determines that the annotation expression information for all utterance texts has been determined, it generates display data in which the determination unit ID, speaker information, utterance text, topic of each utterance text group, and annotation expression information are associated with each other, as shown in FIG. 9.

[0060] The display data storage unit 25 stores the display data generated by the display preparation unit 24.

[0061] The display data output unit 26 outputs display data. The display data output unit 26 may output the display data to a display device 4 such as a liquid crystal panel or an organic EL display, or may output the display data to another device via a communication network.

[0062] As a result, the display device 4 displays a display screen based on the display data. Specifically, as shown in FIG. 10, the display device 4 displays the utterance text included in the display data in the above-described sequence. The display device 4 then displays annotation information corresponding to the utterance text in association with the utterance text, and further displays the background color of the annotation information with a color gradation indicated by the annotation expression information included in the display data. Note that the gray and white background color of the "Opening" gradation, the green and white background color of the "Accident Situation" gradation, the blue and white background color of the "Injury Situation" gradation, and the orange and white background color of the "Injury Situation" gradation are all shown as black and white gradations in FIG. 10. The same applies to FIGS. 17, 19 to 23 referenced below. The display device 4 may further display one or more of an utterance ID and speaker information in association with the utterance text and the annotation. Note that when the display data output unit 26 transmits the display data to another device via a communication network, the other device displays a display screen based on the display data, similar to the display device 4.

[0063] (Operation of the display data generator) Here, the operation of the display data generating device 2 according to the second embodiment will be described with reference to Fig. 11. Fig. 11 is a flowchart showing an example of the operation of the display data generating device 2 according to the second embodiment. The operation of the display data generating device 2 described with reference to Fig. 11 corresponds to the display method of the display data generating device 2 according to the second embodiment.

[0064] In step S21, the input unit 21 receives input of target data including an utterance text sequence and annotation information corresponding to each piece of text included in the utterance text sequence.

[0065] In step S22, the display preparation unit 24 divides the utterance text included in the target data input by the input unit 21.

[0066] In step S23, the display preparation unit 24 forms a spoken text group made up of consecutive spoken texts having the same annotation information.

[0067] In step S24, the display preparation unit 24 determines, based on the annotation information and the sequential order, annotation expression information indicating the background color of the display screen of the display device 4 and the position and range in which to display the background color, in order to express the correspondence between the utterance text and the annotation information when the utterance text is displayed on the display device 4. In this example, the display preparation unit 24 determines annotation expression information indicating the color and gradation corresponding to the utterance text group.

[0068] In step S25, the display preparation unit 24 determines whether or not annotation expression information corresponding to all of the utterance text groups has been determined.

[0069] If it is determined in step S25 that annotation expression information corresponding to some of the utterance text groups has not been determined, the process returns to step S23, and the display preparation unit 24 repeats the process. If it is determined in step S25 that annotation expression information corresponding to all of the utterance text groups has been determined, the display preparation unit 24 generates display data in step S26 for displaying the utterance text sequence and the annotation information according to the sequence in the utterance text sequence, and for displaying the background color indicated by the annotation expression information at the position and in the range indicated by the annotation expression information.

[0070] In step S27, the display data storage unit 25 stores the display data.

[0071] Thereafter, the display data output unit 26 outputs the display data at an arbitrary timing. The display data output unit 26 may output the display data to the display device 4, or may output the display data to another device via a communication network. The arbitrary timing may be, for example, the timing when a display command is input to the input unit 21. As a result, the display device 4 displays a display screen based on the display data. Specifically, the display device 4 displays the spoken text and annotation information based on the display data, and displays the background color indicated by the annotation expression information at the position and in the range indicated by the annotation expression information.

[0072] In the above description, the display data generating device 2 executes the process of step S22, but this is not a limitation. For example, the display data generating device 2 does not have to execute the process of step S22.

[0073] Here, the effects of the second embodiment compared with the first embodiment will be described.

[0074] In target data containing multiple speech texts uttered by multiple speakers, one speech text may have more than one topic. For example, multiple topics may be interpreted as corresponding to one speech text, and the topic may change midway through one speech text. In such cases, it is difficult to display the speech text and the topic in a way that allows users to accurately recognize the topic. For example, if one of multiple topics corresponding to the speech text is displayed in association with the speech text, the user may not be able to recognize the other topics corresponding to the speech text. Furthermore, if a speech text with a topic change is divided according to the change and the corresponding topic is displayed for each divided speech text, the user may have difficulty understanding the content of the speech text simply by referring to the divided speech text. In other words, if the speech text is displayed collectively by label (annotation information), such as a scene estimation result, the user can recognize the speech text for each label. However, speech text does not necessarily correspond to a single label. When multiple labels can be associated with a single speech text, it has been difficult to visualize the annotation information in a way that is easy for users to recognize. For example, there may be multiple possible interpretations of a label corresponding to one utterance text, or the utterance text may be long and the corresponding label may change midway through.

[0075] Taking the target data shown in FIG. 2 as an example, the utterance text "Now, let me confirm some details about this accident" at the beginning of the utterance text sequence is a standard opening phrase, and therefore the topic of the utterance text is interpreted as "opening." Furthermore, since the utterance text includes the phrase "some details about the accident," the topic of the utterance text is also interpreted as "accident situation." In such a case, if two topics, "opening" and "accident situation," are displayed in response to the above utterance text, the user may have difficulty understanding the topic of the utterance text. Furthermore, if one of the two topics, "opening" or "accident situation," is displayed in response to the above utterance text, the user will not be able to recognize the other topic.

[0076] In the example shown in FIG. 2, the customer uttered the utterance text "Your rear bumper hit the wall and missed, and you were shocked" (utterance ID "7"), and then the operator uttered the utterance text "That was terrible. I'm worried about you. Are you OK?" (utterance ID "8"). Here, the topic of the utterance text "That was terrible" is "the accident situation," and the topics of the utterance texts "I'm worried about you" and "Are you OK?" are "the injury situation." In this case, if the utterance text is divided by a period and the topics corresponding to the utterance texts "That was terrible," "I'm worried about you," and "Are you OK?" are displayed, it becomes difficult for the user to understand the target of the utterance text "Are you OK?", and as a result, it becomes difficult for the user to recognize the content of the target data.

[0077] In contrast, according to the second embodiment, the display data generating device 2 determines annotation expression information so that the background color gradually changes toward the boundary between the different annotation information before and after a text sequence. This allows the display data generating device 2 to visualize annotation information even when multiple pieces of annotation information correspond to one utterance text. This allows the user to recognize that the topic of the utterance text may be a topic indicated by color, and that the topic of the utterance text may also be a topic not indicated by color. In the example shown in FIG. 10 , the user can recognize that the topic of the utterance text corresponding to utterance ID “7” is “accident situation” and may also be “injury situation.” Therefore, the user can understand that the subject of “That must have been tough” included in the utterance text corresponding to utterance ID “8” following utterance ID “7” may be “injury situation.” Therefore, the user can intuitively grasp utterance text-related information based on the background color of the information and quickly and appropriately recognize the content of target data including the utterance text.

[0078] Similarly, the background of the topic "Opening" (utterance IDs "1" to "5") is displayed with a gradient that changes from gray to white from the start point to the end point. Furthermore, the background of the topic "Accident Situation" (utterance IDs "6" and "7") is displayed with a gradient that changes from white to green from the start point to the middle point. This allows the user to recognize that the topic of the utterance text corresponding to ID "5," which is the last of the utterance text group corresponding to the topic "Opening" (utterance IDs "1" to "5"), is not only "Opening," but may also be "Accident Situation." This also allows the user to intuitively grasp the information related to the utterance text from the background color of the information, and to quickly and appropriately recognize the content of the target data including the utterance text.

[0079] Furthermore, if the spoken text were not divided and a spoken text such as "That was terrible. I'm worried about your health. Are you okay?" were displayed using gradation without being divided, the range of gradation would be narrow, and it would be difficult for the user to tell where the "accident situation" ends and the "injury situation" begins. In contrast, in this embodiment, the display data generating device 2 displays the spoken text divided into three parts using periods, for example, using gradation, as shown in utterance ID8 in Fig. 10, so the range of gradation is wider, making it easier for the user to intuitively grasp the boundary between the "accident situation" and the "injury situation."

[0080] <Third embodiment> The overall configuration of the display data generating device 3 of the third embodiment will be described with reference to Fig. 12. Fig. 12 is a schematic diagram of the display data generating device 3 according to this embodiment.

[0081] (Functional configuration of the display data generating device) 12, the display data generating device 3 according to the third embodiment includes an input unit 31, a target data storage unit 32, a display rule storage unit 33, a display preparation unit 34, a display data storage unit 35, and a display data output unit 36. The input unit 31 is configured by an input interface that accepts input of information, similar to the input unit 21 of the second embodiment. The target data storage unit 32, the display rule storage unit 33, and the display data storage unit 35 are configured by memory, similar to the target data storage unit 22, the display rule storage unit 23, and the display data storage unit 25 of the second embodiment. Furthermore, the display preparation unit 34 and the display data output unit 36 ​​constitute a control unit, similar to the display preparation unit 24 of the second embodiment.

[0082] The input unit 31 accepts input of target data including an utterance text sequence and annotation information corresponding to each piece of text included in the utterance text sequence, as shown in FIG. 13, and further including accuracy indicating the accuracy of the annotation information. The target data may further include speaker information. The accuracy of the topic may be determined for the utterance text by an arbitrary algorithm, or may be input by a user's operation. In the third embodiment, the annotation information is the topic to which the content of the utterance text belongs, but is not limited to this.

[0083] The target data storage unit 32 stores the target data input by the input unit 31.

[0084] The display rule storage unit 33 stores rules that are used by the display preparation unit 34 to determine annotation expression information for utterance text based on annotation information. The display rule storage unit 33 includes a color storage unit 331 and a gradation rule storage unit 332. The color storage unit 331 is similar to the color storage unit 231 of the display data generating device 2 according to the second embodiment.

[0085] The gradation rule storage unit 332 stores gradation rules such as those shown in Fig. 14, which are used by the display data output unit 36 ​​to determine annotation expression information used when displaying spoken text related information and the background of the information. The gradation rules in the third embodiment are gradations determined based on the annotation information, the sequence of the spoken text, and the accuracy of the annotation information.

[0086] Fig. 15 is a diagram showing an example of applying the gradation rule of "the next topic follows the last utterance text of the topic" shown in Fig. 14 when the accuracy is 60%. "The next topic follows the last utterance text of the topic" indicates that the topic of the utterance text is different from the topic of the utterance text uttered next to the utterance text.

[0087] As shown in FIG. 15 , when the utterance text to be determined is "the last utterance text of a topic, and the next topic continues," and the accuracy of the topic is not 100%, the annotation expression information has a color corresponding to the topic from the start point to a position corresponding to the accuracy of the topic (the 60% position in the example of FIG. 15 ), where 100% is the range from the start point to the end point, and the color changes from the color corresponding to the topic to white as it approaches the end point. Here, the start point is the end (the upper end in the example of FIG. 17 ) of a column displaying topics (one utterance text) in the arrangement direction (from top to bottom in the example of FIG. 17 , which will be referred to later), as in the second embodiment. The end point is the end (the lower end in the example of FIG. 17 ) of a column displaying topics (one utterance text) in the arrangement direction. Furthermore, if the utterance text to be determined is "the last utterance text of a topic, followed by the next topic," and the accuracy of the topic is 100%, the annotation expression information is the color corresponding to the topic, without gradation.

[0088] In the example gradation rule shown in FIG. 14 , when the relationship between the utterance text to be determined and the topic associated with the utterance text is “the first utterance text of the topic, and continues from the previous topic,” and the accuracy of the topic is not 100%, the annotation expression information has a gradation that changes from white to the color corresponding to the topic as it moves from the start point toward a position corresponding to (100-topic accuracy)%, and the color corresponds to the topic from the position corresponding to (100-topic accuracy)% to the end point. Note that “the first utterance text of the topic, and continues from the previous topic” indicates that the topic of the utterance text is different from the topic of the utterance text spoken before the topic text. Furthermore, when the relationship between the utterance text to be determined and the topic associated with the utterance text is “the first utterance text of the topic, and continues from the previous topic,” and the accuracy of the topic is 100%, the annotation expression information has a color corresponding to the topic without gradation.

[0089] Furthermore, if the relationship between the speech text to be determined and the topic associated with the speech text is "a speech text in which the topic changes midway," the annotation expression information is a gradation from the color of the topic before the change to white from the start point to a position corresponding to the certainty of the topic, and a gradation from white to the color of the topic after the change from the position corresponding to the certainty of the topic to the end point.

[0090] If the utterance text to be determined does not satisfy any of the above conditions, the annotation expression information is in the color of the topic of the utterance text from the start point to the end point without gradation.

[0091] The display preparation unit 34 determines the annotation expression information so that the background color gradually changes toward a boundary where annotation information differs before and after a sequence in a spoken text sequence. In this embodiment, the display preparation unit 34 determines the annotation expression information further based on the accuracy. The display preparation unit 34 may determine annotation expression information that indicates the degree to which the background color changes further based on the accuracy. In a third embodiment, the annotation expression information is information that indicates a color and a gradation. In this case, the display preparation unit 34 may divide the spoken text and determine the annotation expression information based on the divided spoken text, the annotation information of the spoken text, and the sequence.

[0092] Specifically, first, the display preparation unit 34 divides the utterance text included in the target data input by the input unit 11, similar to the display preparation unit 24 of the second embodiment. Note that the display preparation unit 34 does not need to perform the process of dividing the utterance text, similar to the display preparation unit 24 of the second embodiment. In such a configuration, for example, the utterance text included in the target data may be divided utterance text. Note that in the example shown in FIG. 16, the display preparation unit 34 does not divide the utterance text, and therefore the utterance text corresponding to the determination unit ID in the display data is the same as the utterance text corresponding to the utterance ID in the target data shown in FIG. 13.

[0093] The display preparation unit 34 determines a color and gradation corresponding to the spoken text using the color scheme rule and the gradation rule. In the example using the gradation rule shown in Fig. 14, the display preparation unit 34 determines the annotation expression information based on the annotation information of the spoken text and the annotation information of the spoken text arranged before or after the spoken text in the spoken text sequence. Specifically, if the spoken text to be determined is "the last spoken text of a topic, and the next topic follows," and the accuracy of the topic is not 100%, the display preparation unit 34 determines that the annotation expression information will be a color corresponding to the topic up to the accuracy of the topic, and a gradation that changes from the color corresponding to the topic to white as the accuracy of the topic increases.

[0094] 14, when the utterance text is "the first utterance text of a topic, which continues from the previous topic" and the accuracy of the topic is not 100%, the display preparation unit 34 determines that the annotation expression information will be the color corresponding to the topic up to the accuracy of the topic, with a gradation that changes from the color corresponding to the topic to white from the accuracy of the topic. Furthermore, whether the utterance text to be determined is "the last utterance text of a topic, which continues from the next topic" or "the first utterance text of a topic, which continues from the previous topic," and the accuracy of the topic is 100%, the display preparation unit 34 determines that the annotation expression information will be the color corresponding to the topic without gradation.

[0095] Furthermore, in an example using the gradation rule shown in Figure 14, when the spoken text is "spoken text in which the topic switches midway," the display preparation unit 34 determines that the annotation expression information is a gradation that changes from the color of the topic before the switch to white up to the topic certainty, and from the topic certainty, a gradation that changes from white to the color of the topic after the switch.

[0096] In addition, in an example using the gradation rule shown in Figure 14, if the spoken text to be determined does not satisfy any of the above conditions, the display preparation unit 34 determines that the annotation expression information is the color of the topic of the spoken text without gradation.

[0097] Furthermore, after determining the annotation expression information for the spoken text, the display preparation unit 34 determines whether the annotation expression information for all of the spoken text has been determined. If the display preparation unit 34 determines that the annotation expression information for some of the spoken text has not been determined, the display preparation unit 34 repeats the process of determining the annotation expression information for the spoken text for which the annotation expression information has not been determined. If the display preparation unit 34 determines that the annotation expression information for all of the spoken text has been determined, the display preparation unit 34 generates display data in which the annotation expression information is associated with each of the spoken texts included in the target data.

[0098] The display data storage unit 35 stores the display data generated by the display preparation unit 34.

[0099] The display data output unit 36 ​​outputs display data. The display data output unit 36 ​​may output the display data to a display device 4 such as a liquid crystal panel or an organic EL display, or may output the display data to another device via a communication network.

[0100] As a result, the display device 4 displays a display screen based on the display data. Specifically, as shown in Fig. 17, the display device 4 displays the spoken text included in the display data in association with the annotation information corresponding to the spoken text, and further displays the background color of the annotation information in a color gradation as indicated by the annotation expression information included in the display data. The display device 4 may also display one or more of an ID and speaker information in association with the spoken text. Note that when the display data output unit 36 ​​transmits the display data to another device via a communication network, the other device displays a display screen based on the display data, similar to the display device 4.

[0101] (Operation of the display data generator) Here, the operation of the display data generating device 3 according to the third embodiment will be described with reference to Fig. 18. Fig. 18 is a flowchart showing an example of the operation of the display data generating device 3 according to the third embodiment. The operation of the display data generating device 3 described with reference to Fig. 18 corresponds to the display method of the display data generating device 3 according to the third embodiment.

[0102] In step S31, the input unit 31 receives input of target data including a spoken text sequence, annotation information corresponding to each spoken text included in the spoken text sequence, and accuracy of the annotation information.

[0103] In step S32, the display preparation unit 34 divides the utterance text included in the target data input by the input unit 31.

[0104] In step S33, the display preparation unit 34 determines, based on the accuracy of the annotation information in addition to the annotation information and the sequential order, annotation expression information indicating the background color of the display screen of the display device 4 and the position and range in which to display the background color, in order to express the correspondence between the uttered text and the annotation information when the display device 4 displays the uttered text. Specifically, the display preparation unit 24 determines annotation expression information indicating the color and gradation corresponding to the uttered text.

[0105] In step S34, the display preparation unit 34 determines whether or not the annotation expression information for all of the spoken texts has been determined.

[0106] If it is determined in step S34 that the annotation expression information for some of the utterance texts has not been determined, the process returns to step S33, and the display preparation unit 34 repeats the process. If it is determined in step S34 that the annotation expression information for all of the utterance texts has been determined, the display preparation unit 34 generates display data in step S35 for displaying the utterance text sequence and the annotation information according to the sequence in the utterance text sequence, and for displaying the background color indicated by the annotation expression information at the position and in the range indicated by the annotation expression information.

[0107] In step S36, the display data storage unit 35 stores the display data.

[0108] Thereafter, the display data output unit 36 ​​outputs the display data at an arbitrary timing. The display data output unit 36 ​​may output the display data to the display device 4, or may output the display data to another device via a communication network. The arbitrary timing may be, for example, the timing when a display command is input to the input unit 31. As a result, the display device 4 displays a display screen based on the display data. Specifically, the display device 4 displays the spoken text and annotation information based on the display data, and displays the background color indicated by the annotation expression information in the position and range indicated by the annotation expression information.

[0109] In the above description, the display data generating device 3 executes the process of step S32, but this is not a limitation. For example, the display data generating device 3 does not have to execute the process of step S32.

[0110] As described above, according to the third embodiment, the target data further includes a degree of certainty indicating the certainty of the annotation information, and the display preparation unit 34 determines the annotation expression information further based on the degree of certainty. This allows the user to recognize that the annotation information corresponding to the spoken text is annotation information corresponding to a color, and also recognize that the annotation information may not be annotation information corresponding to a color. Furthermore, the user can intuitively grasp the certainty that the annotation information corresponding to the spoken text is annotation information corresponding to a color. Therefore, the user can more quickly and appropriately understand the content of the target data including the spoken text.

[0111] In the second embodiment described above, the display data generating device 2 displays utterance texts from multiple speakers in the same column. However, this is not limited to this. For example, as shown in FIG. 19 , the display data generating device 3 displays utterance texts from one speaker and utterance texts from another speaker in different columns, displays annotation information in the row where the utterance texts are displayed, and displays a gradation behind the annotation information. In the example shown in FIG. 19 , the display data generating device 2 displays the utterance texts on the display device 4 so that they are arranged in a sequence of utterance texts from top to bottom of the screen. Regarding the target data of this example, in a dialogue, the operator utters the utterance text corresponding to utterance ID "8," "Are you okay?", and at almost the same time, the customer utters the utterance text corresponding to utterance ID "9," "Yes, everything's fine." In such a case, in the example shown in FIG. 10 , one of the utterance texts simultaneously uttered by multiple speakers is displayed first, and the other is displayed later. In contrast, since the target data includes the time at which the utterance text was uttered, in the example shown in FIG. 19, the display data generating device 2 can display utterance texts uttered by multiple speakers almost simultaneously on the same line based on the time included in the target data. This allows the user to clearly understand that multiple utterance texts were uttered simultaneously by multiple speakers. Therefore, a user who refers to the utterance text based on the target data displayed by the display data generating device 2 can easily understand the utterance text uttered by each speaker and can efficiently recognize the content of the target data. The same applies to the display data generating device 1 according to the first embodiment and the display data generating device 3 according to the third embodiment.

[0112] The display preparation unit 24 of the display data generating device 2 may further determine important utterance texts from among the multiple utterance texts. The display preparation unit 24 can determine important utterance texts using any algorithm. For example, the display data generating device 2 may determine important utterance texts using a model previously generated by learning based on a large amount of important utterance texts, or may store important words and phrases in a memory and determine utterance texts containing the words and phrases stored in the memory as important utterance texts. The display preparation unit 24 may also determine important utterance texts based on a user's operation. In such a configuration, as shown in FIG. 20 , the display data output unit 26 highlights utterance texts determined to be important utterance texts and displays them on the display device 4. For example, the display data generating device 2 may display characters indicating utterance texts determined not to be important utterance texts (other utterance texts) in black on the display device 4, and characters indicating utterance texts determined to be important utterance texts in a color (e.g., red) different from the other utterance texts on the display device 4. In the example shown in Fig. 20, important spoken text is shown in bold, but highlighting is not limited to this. This allows the user to easily grasp important spoken text and efficiently recognize the content of the target data. The same applies to the display data generating device 1 according to the first embodiment and the display data generating device 3 according to the third embodiment.

[0113] As shown in FIG. 21 , the display data output unit 26 of the display data generating device 2 may further display only the utterance text determined to be important, without displaying the utterance text not determined to be important. This allows the user to more easily grasp the important utterance text and more efficiently recognize the content of the target data. In this configuration, the display data output unit 26 may switch between a state in which other utterance text is displayed and a state in which other utterance text is not displayed, in response to a user operation. For example, if the user determines that the target data cannot be fully understood because the other utterance text is not displayed, the user can perform an operation to display the other utterance text and try to understand the target data fully by referring to the other utterance text. This also applies to the display data generating device 1 according to the first embodiment and the display data generating device 3 according to the third embodiment.

[0114] In the second embodiment described above, the annotation information is a topic, but this is not limiting. As shown in FIG. 22, the annotation information may be a "scene" indicating the situation in which the spoken text is uttered. In this example, a "scene" refers to a classification of the spoken text according to the type of situation in a conversation between an operator and a customer. For example, a conversation that begins with an operator introducing himself / herself as a greeting, explaining the reason for the customer's call, the operator confirming the reason, confirming the customer and the contract details, and then responding to the reason, and finally ending with a thank you, is classified into scenes such as "opening," "understanding the inquiry," "response," and "closing." The results of such scene estimation are assigned as labels to the spoken text.

[0115] For example, in an inbound call center where an operator receives calls from customers, the items may include "opening," "understanding the inquiry," "identity verification," "response," and "closing." Furthermore, the display data output unit 26 of the display data generating device 2 may display the spoken text included in the target data, and display the background of the spoken text, which is the information-related portion, on the display device 4 using a color gradation. In other words, in this example, the information-related portion is the background of the spoken text. Furthermore, the display data output unit 26 may display, on the display device 4, a "whole call" button and buttons indicating each item included in the scene, which is annotation information.

[0116] In such a configuration, when any button is operated by a user, the input unit 21 receives information indicating that the operation has been performed, and the display device 4 displays the utterance text based on the information.

[0117] For example, when the "Entire Call" button is pressed by a user's operation, the input unit 21 receives information indicating that the "Entire Call" button has been pressed. Then, based on this information, the display device 4 displays the entire utterance text included in the target data. Also, when the "Opening" button is pressed by a user's operation, the input unit 21 receives information indicating that the "Opening" button has been pressed. Then, based on this information, the display device 4 displays the utterance text included in the target data whose scene is "Opening."

[0118] Furthermore, when the "Understanding Inquiry" button is pressed by a user operation, the display device 4 may display detailed information regarding the "Understanding Inquiry." The detailed information regarding the "Understanding Inquiry" may include at least one of a "Topic," a "Subject," and a "Subject Confirmation" generated by an arbitrary algorithm based on the spoken text corresponding to the scene of the "Understanding Inquiry." The display device 4 may display, together with the "Subject" and "Subject Confirmation," operation objects for performing operations to change the "Topic," "Subject," and "Subject Confirmation." Note that the display device 4 may also display detailed information regarding the "Understanding Inquiry" when the "Entire Call" button is pressed by a user operation.

[0119] Furthermore, when the "identity verification" button is pressed by a user operation, the display device 4 may display detailed information regarding the "identity verification." The detailed information regarding the "identity verification" may include at least one of the customer's "name," "address," and "telephone number" generated by an arbitrary algorithm based on the spoken text corresponding to the "identity verification" scene. The display data output unit 26 may display, on the display device 4, the "name," "address," and "telephone number," as well as operation objects for performing operations to change the "name," "address," and "telephone number." Note that the display data output unit 26 may also display detailed information regarding the "identity verification" on the display device 4 when the "entire call" button is pressed by a user operation.

[0120] Additionally, the display device 4 may display the time period in which the utterance text included in the target data was uttered, along with displaying the utterance text included in the target data. Additionally, the display device 4 may display an audio playback button (a triangular arrow in FIG. 22) near the utterance text for playing back audio data corresponding to the utterance text. In such a configuration, the display data generating device 2 plays back audio data when the user presses the audio playback button.

[0121] The display data generating device 1 according to the first embodiment and the display data generating device 3 according to the third embodiment can also execute the aspect described with reference to FIG. 22 in a similar manner.

[0122] In the embodiment described with reference to FIG. 22 , the annotation information is a “scene.” However, as shown in FIG. 23 , the annotation information may be both a “scene” and a “dialogue act type” indicating the type of action when the utterance text was uttered. For example, in an outbound call center where an operator makes calls to customers, the “scene” may include “opening,” “injury,” “self-driving,” “grade,” “insurance response,” “repair status,” “accident status,” “contact information,” and “closing.” Furthermore, the display device 4 to which the display data is output from the display data generating device 2 may display the background color of the utterance text included in the target data using a gradation. Furthermore, in this example, the “dialogue act type” may include “interrogation,” “explanation,” “question,” and “answer.” The “interrogation” is an utterance text in which the operator is interviewing the customer, the “explanation” is an utterance text in which the operator is explaining to the customer, the “question” is an utterance text in which the customer is asking the operator a question, and the “answer” is an utterance text in which the customer responds to the operator’s inquiry.

[0123] The display device 4 may display a "whole call" button, buttons indicating each item included in the annotation information "scene," and buttons indicating each item included in the annotation information "dialogue act type." In this configuration, when a user operates any button, the input unit 21 receives information indicating that the operation has been performed, and the display device 4 displays the utterance text based on the information. In this example, the buttons indicating each item included in the "dialogue act type" are configured as check buttons so that one or more buttons can be selected, but this is not limited to this, and any button configuration can be used as appropriate. In the example shown in Figure 23, the "answer" button is checked, and only utterance texts for which the annotation information "dialogue act type" is associated with "answer" are displayed.

[0124] 22, the display device 4 may display speaker information and the time period in which the utterance text included in the target data was spoken, along with displaying the utterance text included in the target data. The display device 4 may also display an audio playback button (the triangular arrow in FIG. 23) for playing audio data corresponding to the utterance text, near the portion where the utterance text is displayed. In such a configuration, the display data generating device 2 plays the audio data when the user presses the audio playback button.

[0125] The display data generating device 1 according to the first embodiment and the display data generating device 3 according to the third embodiment can similarly execute the aspect described with reference to FIG.

[0126] Furthermore, in the second embodiment described above, the colors corresponding to the annotation information stored in the color storage unit 331 are different from one another. However, this is not limited to this; the colors corresponding to the annotation information may be the same. Even in this configuration, the display data output unit 36 ​​can display the background on the display device 4 with a color gradation based on the annotation expression information indicating the color and gradation generated by the display preparation unit 34 based on the gradation rule stored in the gradation rule storage unit 232. This allows the user to recognize that the topic corresponding to the utterance text group can be interpreted as multiple topics, not just one. Furthermore, in this configuration, the display data generating device 2 does not need to include the color storage unit 231, thereby reducing the memory capacity. The same applies to the third embodiment.

[0127] Furthermore, the display modes, gradation rules, etc. described in the first to third embodiments are merely examples, and the present invention is not limited to these. Furthermore, the display data generating devices 1 to 3 according to the first to third embodiments may further include various functions used by operators when creating call histories. For example, the display data generating devices 1 to 3 may further include a function for displaying utterance text for each topic, a function for editing utterance text and topics, a search function for searching utterance text, a comparison function for comparing target data, etc.

[0128] A computer 100 capable of executing program instructions can be used to function as the display data generation device 1 described above. FIG. 24 is a block diagram showing a schematic configuration of a computer 100 that functions as the display data generation device 1. Here, the computer 100 may be a general-purpose computer, a special-purpose computer, a workstation, a personal computer (PC), an electronic notepad, or the like. The program instructions may be program code, code segments, or the like for executing necessary tasks. Similarly, a computer 100 capable of executing program instructions can be used to function as the display data generation device 2, and a computer 100 capable of executing program instructions can be used to function as the display data generation device 3.

[0129] <Hardware configuration> 24, computer 100 includes processor 110, ROM (Read Only Memory) 120, RAM (Random Access Memory) 130, storage 140, input unit 150, output unit 160, and communication interface (I / F) 170. Each component is connected to each other via bus 180 so as to be able to communicate with each other. Processor 110 is specifically a CPU (Central Processing Unit), MPU (Micro Processing Unit), GPU (Graphics Processing Unit), DSP (Digital Signal Processor), SoC (System on a Chip), etc., and may be configured by multiple processors of the same or different types.

[0130] The processor 110 controls each component and executes various arithmetic processes. That is, the processor 110 reads a program from the ROM 120 or the storage 140 and executes the program using the RAM 130 as a work area. The processor 110 controls each component and executes various arithmetic processes in accordance with the program stored in the ROM 120 or the storage 140. In this embodiment, the program according to the present disclosure is stored in the ROM 120 or the storage 140.

[0131] The program may be recorded on a recording medium readable by computer 100. Using such a recording medium, the program can be installed on computer 100. Here, the recording medium on which the program is recorded may be a non-transitory recording medium. The non-transitory recording medium is not particularly limited, and may be, for example, a CD-ROM, a DVD-ROM, or a USB (Universal Serial Bus) memory. Furthermore, the program may be downloaded from an external device via a network.

[0132] The ROM 120 stores various programs and various data. The RAM 130 temporarily stores programs or data as a working area. The storage 140 is configured with an HDD (Hard Disk Drive) or SSD (Solid State Drive) and stores various programs including the operating system and various data.

[0133] The input unit 150 includes one or more input interfaces that receive input operations from a user and acquire information based on the user operations. For example, the input unit 150 is a pointing device, a keyboard, a mouse, etc., but is not limited to these.

[0134] The output unit 160 includes one or more output interfaces for outputting information, such as, but not limited to, a display for outputting information visually or a speaker for outputting information audibly.

[0135] The communication interface 170 is an interface for communicating with other devices such as external devices, and uses standards such as Ethernet (registered trademark), FDDI, and Wi-Fi (registered trademark).

[0136] The following additional notes are provided regarding the above-described embodiments.

[0137] (Additional note 1) A display data generating device including a control unit, The control unit Accepting input of target data including a text sequence and annotation information corresponding to each piece of text included in the text sequence; A display data generating device that determines, based on the annotation information, a background color of the display screen of the display device, as well as annotation expression information indicating the position and range in which to display the background color, in order to express the correspondence between the text and the annotation information when the display device displays the text, and generates display data for displaying the text sequence and the annotation information according to a sequence in the text sequence, the display data being for displaying the background color indicated by the annotation expression information at the position and range indicated by the annotation expression information. (Additional note 2) The display data generating device described in Appendix 1, wherein the control unit determines the annotation expression information so that the background color gradually changes toward a boundary where the annotation information before and after the series in the text series is different. (Additional note 3) the target data further includes a degree of certainty indicating the certainty of the annotation information; 3. The display data generating device according to claim 2, wherein the control unit determines the annotation expression information further based on the accuracy. (Additional note 4) The display data generating device according to claim 3, wherein the control unit determines the annotation expression information indicating a degree of change in the background color based further on the degree of accuracy. (Additional note 5) 5. The display data generating device according to any one of claims 1 to 4, wherein the control unit divides the spoken text and determines the annotation expression information of the divided spoken text. (Additional note 6) A display data generating device described in any one of appendix 1 to 5, wherein the display data includes the annotation information, and the position and range for displaying the background color respectively include the display position and display range of the annotation information. (Additional note 7) receiving input of target data including a text sequence and annotation information corresponding to each piece of text included in the text sequence; determining, based on the annotation information, a background color of a display screen of the display device and annotation expression information indicating a position and range in which the background color is to be displayed, in order to express a correspondence between the text and the annotation information when the display device displays the text, and generating display data for displaying the text sequence and the annotation information according to the sequence in the text sequence, the display data being for displaying the background color indicated by the annotation expression information at the position and range indicated by the annotation expression information; A display data generation method including: (Additional note 8) A non-transitory storage medium storing a program executable by a computer, the program causing the computer to function as the display data generating device described in any one of appendix 1 to 6.

[0138] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, and technical standard was specifically and individually indicated to be incorporated by reference.

[0139] Although the above-described embodiments have been described as typical examples, it will be apparent to those skilled in the art that many modifications and substitutions can be made within the spirit and scope of the present disclosure. Therefore, the present invention should not be interpreted as being limited by the above-described embodiments, and various modifications or alterations are possible without departing from the scope of the claims. For example, multiple building blocks shown in the block diagrams of the embodiments can be combined into one, or one building block can be divided. [Explanation of symbols]

[0140] 1, 2, 3 Display data generator 4 Display device 11, 21, 31 Input section 12, 22, 32 Target data storage unit 13, 23, 33 Display rule memory section 14, 24, 34 Display preparation section 15, 25, 35 Display data storage section 16, 26, 36 Display data output section 131, 231, 331 Color memory section 232, 332 Gradation rule memory 100 computers 110 processors 120 ROM 130 RAM 140 Storage 150 Input section 160 Output section 170 Communication Interface (I / F) 180 Bus

Claims

1. an input unit that receives input of target data including a text sequence including utterances that occur in chronological order or texts arranged in a sentence, a sequence order that is the chronological order or the arrangement order in the sentence, and metadata assigned to each of the texts included in the text sequence; a display preparation unit that determines, based on rules indicating correspondence between metadata and expressions, an expression corresponding to the metadata assigned to each of the texts as a background expression to be applied to a position and range corresponding to each of the texts on a display screen of the display device when the display device displays the text, and generates display data for displaying the texts and the metadata assigned to each of the texts in accordance with the sequential order, the display data being for applying the determined background expression to the position and range corresponding to each of the texts; Equipped with The metadata assigned to each of the texts includes a topic of the utterance, a scene in which the utterance was uttered, or a scene in the sentence; The display preparation unit determines the representation of the background so that the representation of the background gradually changes toward a boundary where metadata before and after the sequential order differ.

2. The target data further includes a degree of certainty indicating the certainty of the metadata assigned to each of the texts; The display data generating device according to claim 1 , wherein the display preparation unit determines a representation of the background based further on the likelihood.

3. The display data generating device according to claim 2 , wherein the display preparation unit determines a degree of change in the representation of the background based further on the probability.

4. an input unit that receives input of target data including a text sequence including utterances that occur in chronological order or texts arranged in a sentence, a sequence order that is the chronological order or the arrangement order in the sentence, and metadata assigned to each of the texts included in the text sequence; a display preparation unit that determines, based on rules indicating correspondence between metadata and expressions, an expression corresponding to the metadata assigned to each of the texts as a background expression to be applied to a position and range corresponding to each of the texts on a display screen of the display device when the display device displays the text, and generates display data for displaying the texts and the metadata assigned to each of the texts in accordance with the sequential order, the display data being for applying the determined background expression to the position and range corresponding to each of the texts; Equipped with The display preparation unit generates the display data as display data for displaying the text in a first column and the metadata assigned to each of the text in a second column different from the first column, in the sequential order, so that each of the texts and the metadata assigned to each of the texts are lined up in the same row, and the display data generation device generates the display data for applying the determined background expression to the display position and display range of the metadata assigned to each of the texts in the second column.

5. The display data generating device according to claim 4 , wherein the metadata assigned to each of the texts includes a topic of the utterance, a scene in which the utterance was uttered, a scene in the sentence, or a classification label of each of the texts.

6. the rules are coloration rules that indicate correspondence between metadata and colors, 6. A display data generating device as described in any one of claims 1 to 5, wherein the display preparation unit determines, based on the color scheme rules, a color corresponding to the metadata assigned to each of the texts as a background color to be displayed at a position and in an area corresponding to each of the texts on the display screen of the display device when the display device displays the text, and generates the display data for displaying the determined background color at the position and in the area corresponding to each of the texts.

7. 7. The display data generating device according to claim 1, wherein the display preparation unit groups, from among the text included in the text sequence, text that has the same assigned metadata and that is continuous when arranged in the sequential order, determines, based on the rule, an expression corresponding to the metadata assigned to each of the text groups as a background expression to be applied to a position and range corresponding to each of the text groups on a display screen of the display device when the display device displays the text, and generates, as the display data, display data for displaying the text and the metadata assigned to each of the text groups in the sequential order, the display data for applying the determined background expression to the position and range corresponding to each of the text groups.

8. receiving an input of target data including a text sequence including utterances occurring in chronological order or text arranged in a sentence, a sequence order that is the chronological order or the order in the sentence, and metadata assigned to each piece of text included in the text sequence; determining, based on a rule indicating the correspondence between metadata and expressions, an expression corresponding to the metadata assigned to each of the texts as a background expression to be applied to a position and range corresponding to each of the texts on a display screen of the display device when the display device displays the text, and generating display data for displaying the texts and the metadata assigned to each of the texts in accordance with the sequential order, the display data being for applying the determined background expression to the position and range corresponding to each of the texts; Including, The metadata assigned to each of the texts is a topic of the utterance, a scene in which the utterance was uttered, or a scene in the sentence; The display data generating method includes determining a representation of the background in such a way that the representation of the background gradually changes toward a boundary where the metadata before and after the sequential order differ.

9. receiving an input of target data including a text sequence including utterances occurring in chronological order or text arranged in a sentence, a sequence order that is the chronological order or the arrangement order in the sentence, and metadata assigned to each piece of text included in the text sequence; determining, based on a rule indicating the correspondence between metadata and expressions, an expression corresponding to the metadata assigned to each of the texts as a background expression to be applied to a position and range corresponding to each of the texts on a display screen of the display device when the display device displays the text, and generating display data for displaying the texts and the metadata assigned to each of the texts in accordance with the sequential order, the display data being for applying the determined background expression to the position and range corresponding to each of the texts; Including, The generating step is a display data generation method for generating display data for displaying the text in a first column and the metadata assigned to each of the text in a second column different from the first column in the sequential order so that each of the texts and the metadata assigned to each of the texts are lined up in the same row, and for applying the determined background expression to the display position and display range of the metadata assigned to each of the texts in the second column.

10. Computer, an input unit that receives input of target data including a text sequence including utterances that occur in chronological order or texts arranged in a sentence, a sequence order that is the chronological order or the arrangement order in the sentence, and metadata assigned to each of the texts included in the text sequence; a display preparation unit that determines, based on rules indicating correspondence between metadata and expressions, an expression corresponding to the metadata assigned to each of the texts as a background expression to be applied to a position and range corresponding to each of the texts on a display screen of the display device when the display device displays the text, and generates display data for displaying the texts and the metadata assigned to each of the texts in accordance with the sequential order, the display data being for applying the determined background expression to the position and range corresponding to each of the texts; Equipped with The metadata assigned to each of the texts is a topic of the utterance, a scene in which the utterance was uttered, or a scene in the sentence; A display data generation program for causing the display preparation unit to function as a display data generation device that determines the representation of the background so that the representation of the background gradually changes toward the boundary where the metadata before and after the series order is different.

11. Computer, an input unit that receives input of target data including a text sequence including utterances that occur in chronological order or texts arranged in a sentence, a sequence order that is the chronological order or the arrangement order in the sentence, and metadata assigned to each of the texts included in the text sequence; a display preparation unit that determines, based on rules indicating correspondence between metadata and expressions, an expression corresponding to the metadata assigned to each of the texts as a background expression to be applied to a position and range corresponding to each of the texts on a display screen of the display device when the display device displays the text, and generates display data for displaying the texts and the metadata assigned to each of the texts in accordance with the sequential order, the display data being for applying the determined background expression to the position and range corresponding to each of the texts; Equipped with The display preparation unit is a display data generation program that functions as a display data generation device that generates display data for displaying the text in a first column and the metadata assigned to each of the text in a second column different from the first column in the sequential order so that each of the texts and the metadata assigned to each of the texts are lined up in the same row, and that applies the determined background expression to the display position and display range of the metadata assigned to each of the texts in the second column.

Citation Information

Patent Citations

  • Document editing device

    JP1993143588A

  • Information processing device, information processing method, program, and recording medium

    JP2006216022A

  • FMEA sheet creation support system and creation support program

    JP2011008355A

  • Sentence display device, program and control method

    WO2016056402A1