Conversation information extraction device and conversation information extraction method

The dialogue information extraction device automates the analysis of dialogue structure and viewpoint identification, addressing the inefficiencies of manual content checking and enhancing compliance by providing real-time, automated dialogue status monitoring.

JP7801175B2Active Publication Date: 2026-01-16HITACHI LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
JP2022085015
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2022-05-25
Publication Date
2026-01-16
Estimated Expiration
2042-05-25

AI Technical Summary

Technical Problem

Existing dialogue analysis systems require manual checking of conversation content, which is costly and prone to missing important information, especially in contexts like call centers where compliance with regulations is critical.

Method used

A dialogue information extraction device that analyzes dialogue between multiple speakers using an estimation model to identify dialogue structure and extract utterances with specific viewpoints, such as agreement or disagreement, enabling efficient and automated confirmation of dialogue situations.

Benefits of technology

Enables efficient and automated extraction of dialogue information with high granularity, reducing the risk of missing critical points and improving compliance by allowing real-time monitoring of dialogue status.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007801175000001
    Figure 0007801175000001
  • Figure 0007801175000002
    Figure 0007801175000002
  • Figure 0007801175000003
    Figure 0007801175000003
Patent Text Reader

Abstract

To extract information on a dialogue having a viewpoint to which attention should be paid from a dialogue made by a plurality of speakers and enable a dialogue situation to be efficiently checked.SOLUTION: A dialogue information extraction device 100 includes: an utterance information input unit 111 that for each of utterances constituting a dialogue, inputs utterance information including a written text of the utterance; a dialogue structure estimation unit 112 that using an estimation model having learned a dialogue structure, estimates the dialogue structure of the dialogue from utterance information on a plurality of utterances in the dialogue; and a viewpoint-based dialogue extraction unit 113 that determines whether or not there is a predetermined viewpoint to which attention should be paid in an utterance group composed of utterances of a plurality of speakers generated based on the dialogue structure, thereby extracting an utterance group having the viewpoint.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The present invention relates to a dialogue information extraction device and a dialogue information extraction method, and is suitable for application to a dialogue information extraction device and a dialogue information extraction method for analyzing a dialogue between multiple speakers. [Background technology]

[0002] In recent years, various efforts have been made to transcribe spoken dialogues in order to expand the value of digital dialogue services.

[0003] For example, when a bank employee or staff member provides an explanation to a customer and proceeds with a procedure at a call center or counter, they must proceed with the conversation while checking the conversation status, with particular emphasis on the customer's consent. In this case, any interaction without clearly obtaining the customer's consent poses compliance risks, such as violating the Financial Instruments Sales Act. Therefore, it is extremely important to confirm the content of what was said and the consent status thereof during the conversation or afterwards.

[0004] As a conventional technique relating to the above-mentioned background, for example, Patent Document 1 discloses a call center conversation content display system that displays responses from each speaker in a chat format. [Prior art documents] [Patent documents]

[0005] [Patent Document 1] International Publication No. 2019 / 003395 Summary of the Invention [Problem to be solved by the invention]

[0006] However, the system disclosed in Patent Document 1 is relatively simple, displaying the responses of each speaker in a chat format, and when checking the content of the conversation from a specific viewpoint, it is necessary to manually check the text of the utterances. In this case, not only is the cost of manual work enormous, but there is also the risk of missing something.

[0007] The present invention has been made in consideration of the above points, and aims to propose a dialogue information extraction device and a dialogue information extraction method that extract dialogue information having noteworthy viewpoints from a dialogue between multiple speakers, thereby enabling efficient confirmation of the dialogue situation. [Means for solving the problem]

[0008] In order to solve such problems, the present invention provides a dialogue information extraction device that analyzes a dialogue between multiple speakers based on a viewpoint to be noted, comprising: an utterance information input unit that inputs utterance information including a text transcript of each utterance that constitutes the dialogue; a dialogue structure estimation unit that estimates the dialogue structure of the dialogue from the utterance information of the multiple utterances in the dialogue using an estimation model that has learned the dialogue structure; and a viewpoint-specific dialogue extraction unit that extracts a group of utterances that have the viewpoint by determining whether or not the viewpoint is present in a group of utterances made up of utterances from multiple speakers that are generated based on the dialogue structure.

[0009] In addition, in order to solve such problems, the present invention provides a dialogue information extraction method using a dialogue information extraction device that analyzes a dialogue between multiple speakers based on a viewpoint to be noted, the dialogue information extraction method comprising: an utterance information input step in which the dialogue information extraction device inputs utterance information including a text transcript of each utterance that constitutes the dialogue; a dialogue structure estimation step in which the dialogue information extraction device estimates the dialogue structure of the dialogue from the utterance information of the multiple utterances in the dialogue inputted in the utterance information input step, using an estimation model that has learned the dialogue structure; and a viewpoint-specific dialogue extraction step in which the dialogue information extraction device extracts a group of utterances that have the viewpoint by determining whether or not the viewpoint is present in a group of utterances made up of utterances from multiple speakers, which is generated based on the estimation result of the dialogue structure estimation step. [Effects of the Invention]

[0010] According to the present invention, it is possible to extract information on a dialogue having a point of view that should be noted from a dialogue between a plurality of speakers, and to efficiently check the dialogue situation. [Brief explanation of the drawings]

[0011] [Figure 1] 1 is a block diagram showing an example of the configuration of a conversation information extraction device 100 according to a first embodiment of the present invention. [Figure 2] 10 is a flowchart illustrating an example of a processing procedure for dialogue information extraction processing. [Figure 3] FIG. 10 is a diagram illustrating an example of utterance information. [Figure 4] FIG. 10 is a diagram showing an example of a past utterance information group. [Figure 5] FIG. 10 is a diagram showing an example of a dialogue structure estimation result. [Figure 6] FIG. 10 is a diagram illustrating an example of dialogue information regarding a set of utterances having a viewpoint. [Figure 7] FIG. 10 is a diagram showing an example of a display of an extraction result of dialogue information. [Figure 8] FIG. 10 is a block diagram showing an example of the configuration of a conversation information extraction device 101 according to a second embodiment of the present invention. [Figure 9] 10 is a flowchart illustrating an example of a processing procedure for dialogue information extraction processing according to the second embodiment. [Figure 10] FIG. 2 is a block diagram showing an example of the hardware configuration of conversation information extraction devices 100 and 101. DETAILED DESCRIPTION OF THE INVENTION

[0012] Hereinafter, embodiments of the present invention will be described in detail with reference to the drawings.

[0013] Note that the following description and drawings are examples for explaining the present invention, and have been omitted or simplified as appropriate for clarity of explanation. Furthermore, not all of the combinations of features described in the embodiments are necessarily essential to the solution of the invention. The present invention is not limited to the embodiments, and all application examples consistent with the concept of the present invention are included in the technical scope of the present invention. Those skilled in the art can make various additions and modifications to the present invention within the scope of the present invention. The present invention can also be implemented in various other forms. Unless otherwise specified, each component may be plural or singular.

[0014] In the following explanation, various types of information may be described using expressions such as "table," "list," "queue," etc., but the various types of information may also be expressed using data structures other than these. To indicate that it is not dependent on the data structure, "XX table," "XX list," etc. may be referred to as "XX information." When describing the content of each piece of information, expressions such as "identification information," "identifier," "name," "ID," "number," etc. are used, but these are interchangeable.

[0015] In addition, in the following description, when describing elements of the same type without distinguishing between them, reference signs or common numbers in reference signs will be used, and when describing elements of the same type with distinction between them, the reference signs of those elements will be used or an ID assigned to those elements will be used instead of the reference signs.

[0016] Furthermore, although the following description may describe processing performed by executing a program, the program is executed by at least one processor (e.g., a CPU) to perform a predetermined process using storage resources (e.g., memory) and / or interface devices (e.g., communication ports) as appropriate, and therefore the processor may be the subject of the processing. Similarly, the subject of the processing performed by executing a program may be a controller, device, system, computer, node, storage system, storage device, server, management computer, client, or host having a processor. The subject of the processing performed by executing a program (e.g., a processor) may include a hardware circuit that performs part or all of the processing. For example, the subject of the processing performed by executing a program may include a hardware circuit that performs encryption and decryption or compression and decompression. The processor operates as a functional unit that realizes a predetermined function by operating in accordance with the program. Apparatuses and systems including a processor are apparatuses and systems that include these functional units.

[0017] A program may be installed on a device such as a computer from a program source. The program source may be, for example, a program distribution server or a computer-readable storage medium. When the program source is a program distribution server, the program distribution server includes a processor (e.g., a CPU) and storage resources, and the storage resources may further store a distribution program and a program to be distributed. The processor of the program distribution server may then execute the distribution program, causing the processor of the program distribution server to distribute the program to be distributed to other computers. In the following description, two or more programs may be realized as one program, and one program may be realized as two or more programs.

[0018] In the following description, a "perspective" refers to a perspective (a viewpoint of interest) focused on an utterance, and specifically includes, for example, agreement or disagreement with an utterance. When the perspective carries some meaning through the dialogue (in other words, in a text transcribed from multiple utterances), i.e., in the above example, when it can be determined that there is agreement or disagreement with the utterance, it is expressed as "there is a perspective." The perspective can be set arbitrarily, and may be specified by a user or determined by a program, for example. In this description, a series of dialogues consisting of utterances from multiple speakers is treated as a "dialogue scenario," and the dialogue information extraction device according to the present invention analyzes the dialogue based on the perspective of interest for each dialogue scenario.

[0019] (1) First embodiment Fig. 1 is a block diagram showing an example of the configuration of a dialogue information extraction device 100 according to a first embodiment of the present invention. The dialogue information extraction device 100 shown in Fig. 1 includes, as processing units having predetermined processing functions, an utterance information input unit 111, a dialogue structure estimation unit 112, a viewpoint-specific dialogue extraction unit 113 having an utterance set generation unit 114 and a viewpoint determination unit 115, and an extraction result display unit 116, and includes, as storage units for storing data, a dialogue structure estimation model storage unit 121, a viewpoint determination model storage unit 122, a past utterance information storage unit 123, a dialogue structure storage unit 124, and an extraction result storage unit 125.

[0020] The utterance information input unit 111 has a function of inputting, for each utterance, utterance information including text transcribed from the utterance to the dialogue information extraction device 100. The utterance information input unit 111 may be configured to acquire text from outside the dialogue information extraction device 100 (such as a dialogue recording device that records dialogue content as text data), or may be configured to transcribe text from audio data of the utterance itself.

[0021] The dialogue structure estimation unit 112 has a function of estimating the dialogue structure in the current dialogue scenario using an estimation model learned by a statistical method.

[0022] The perspective-specific dialogue extraction unit 113 has a function of extracting information about dialogues having a predetermined perspective in a dialogue scenario based on the estimation result of the dialogue structure in the dialogue scenario. Of the perspective-specific dialogue extraction unit 113, the utterance set generation unit 114 has a function of generating an utterance set indicating a combination of utterances that have a strong dialogue relationship based on the estimation result of the dialogue structure. In addition, the perspective determination unit 115 has a function of determining whether or not the utterance set generated by the utterance set generation unit 114 has a perspective.

[0023] The extraction result display unit 116 has a function of outputting dialogue information relating to an utterance set that satisfies a user-specified condition, out of the utterance sets having a predetermined viewpoint extracted by the viewpoint-specific dialogue extraction unit 113 .

[0024] The dialogue structure estimation model storage unit 121 stores an estimation model used when the dialogue structure estimation unit 112 estimates a dialogue structure.

[0025] The viewpoint determination model storage unit 122 stores a determination model that is used when the viewpoint determination unit 115 determines whether or not a viewpoint exists in an utterance set.

[0026] The past utterance information storage unit 123 stores a past utterance information group that is a collection of utterance information (past utterance information) of a series of dialogues in each dialogue scenario. A specific example of the past utterance information group is shown in Fig. 4, which will be described later.

[0027] The dialogue structure storage unit 124 stores information indicating the result of estimating the dialogue structure in the dialogue scenario. The dialogue structure is estimated by the dialogue structure estimation unit 112, and a specific example of the result of estimating the dialogue structure is shown in Fig. 5, which will be described later.

[0028] The extraction result storage unit 125 stores the dialogue information to be extracted that has been determined by the viewpoint-specific dialogue extraction unit 113 (viewpoint determination unit 115).

[0029] FIG. 2 is a flowchart showing an example of the processing procedure of the dialogue information extraction processing. The dialogue information extraction processing shown in FIG. 2 is executed by each unit of the dialogue information extraction device 100. The dialogue information extraction processing may be executed online sequentially while a dialogue between two or more speakers is in progress, or may be executed after the dialogue has ended. Below, the processing procedure of the dialogue information extraction processing shown in FIG. 2 will be explained, taking the case where the processing is executed sequentially as an example, with reference to other drawings as appropriate. Note that the viewpoint of attention for the utterance is specified, for example, by the user before the processing in FIG. 2 is started.

[0030] According to FIG. 2, first, the utterance information input unit 111 inputs text data (utterance information) obtained by transcribing the latest utterance (step S101).

[0031] FIG. 3 is a diagram showing an example of utterance information. The utterance information 210 shown in FIG. 3 has the following fields: utterance ID 211, utterance time 212, utterance user ID 213, and text 214. The utterance ID 211 indicates an identifier (utterance ID) assigned to each utterance. The utterance ID may be an identifier assigned in accordance with the order of utterances regardless of the dialogue scenario, or may be an identifier assigned in accordance with the order of utterances for each dialogue scenario. The utterance time 212 indicates the time when the utterance occurred. The utterance user ID 213 indicates the identifier (user ID) of the speaker of the utterance. The text 214 indicates the text obtained by transcribing the content of the utterance.

[0032] Following step S101, the utterance information input unit 111 refers to the group of past utterance information stored in the past utterance information storage unit 123 and checks whether or not there is utterance information (past utterance information) relating to past utterances in the current dialogue scenario (step S102).

[0033] 4 is a diagram showing an example of a past utterance information group. The past utterance information group 220 shown in FIG. 4 has items of an utterance ID 221, an utterance time 222, a user ID 223, and a text 224. Basically, each item of the past utterance information group 220 corresponds to each item 211 to 214 of the utterance information 210 illustrated in FIG. 3. The utterance ID 221 indicates an identifier capable of identifying the order of an utterance in a dialogue scenario corresponding to the past utterance information group 220. Note that if the utterance ID 211 of the utterance information 210 is an identifier assigned in the order of utterances regardless of the dialogue scenario, when registering the utterance information 210 in the past utterance information group 220, the value of the utterance ID 211 may be converted to a value corresponding to the order of utterances in the dialogue scenario before being registered in the utterance ID 221.

[0034] If no past utterance information exists in step S102 (NO in step S102), the utterance information input in step S101 is utterance information indicating the first utterance in the current dialogue scenario. In this case, no dialogue has been formed in the dialogue scenario at this time, and it is not yet time to extract dialogue information. Therefore, the utterance information input unit 111 stores the utterance information input in step S101 in the past utterance information storage unit 123 as a group of past utterance information for the new dialogue scenario (step S110), and ends the dialogue information extraction process.

[0035] On the other hand, if past utterance information exists in step S102 (YES in step S102), the utterance information input unit 111 inputs the utterance information input in step S101 and a group of past utterance information of the current dialogue scenario stored in the past utterance information storage unit 123 to the dialogue structure estimation unit 112 (step S103).

[0036] Then, the dialogue structure estimation unit 112 estimates the dialogue structure of the current dialogue scenario based on the information input in step S103 (step S104), and stores the dialogue structure estimation result in the dialogue structure storage unit 124.

[0037] Here, the estimation of the dialogue structure by the dialogue structure estimation unit 112 in step S104 will be described in detail.

[0038] The dialogue structure estimation unit 112 uses an estimation model trained by a statistical method to calculate the probability of a relationship between two utterances by different speakers, and estimates a combination of utterances that have a dialogue relationship based on the calculation result, thereby estimating the dialogue structure. The estimation model is stored in the dialogue structure estimation model storage unit 121.

[0039] The dialogue information extraction device 100 (e.g., the dialogue structure estimation unit 112) can train an estimation model stored in the dialogue structure estimation model storage unit 121 using training data to which a dialogue structure has already been assigned. Specifically, for example, when training an estimation model, the following inputs are prepared: The spoken sentence (text 224) is vectorized using Word2Vec, Glove, a pre-trained language model, or the like. Furthermore, the utterance time (utterance time 222) is treated as a vector using a numerical value as is. Furthermore, the speaker (user ID 223) is vectorized as a one-hot vector.

[0040] By training an estimation model using the above-described training data, the input feature vector for training the estimation model or for estimation using the estimation model has the following characteristics: "difference in speech time between two utterances to be subject to relationship estimation," "number of utterances per speaker between two utterances to be subject to relationship estimation," and "similarity of utterance sentences between two utterances to be subject to relationship estimation." Note that the above-described "two utterances to be subject to relationship estimation" are a combination of one utterance and one utterance that precedes that utterance, and the number of previous utterances to be combined can be set arbitrarily. Furthermore, "similarity of utterance sentences" indicates, for example, whether semantically similar words or phrases appear in the utterance sentences.

[0041] Fig. 5 is a diagram showing an example of the dialogue structure estimation result. The dialogue structure data 230 shown in Fig. 5 is data showing an example of the dialogue structure estimated by the dialogue structure estimation unit 112, and a combination of utterances having a dialogue relationship is shown using the utterance ID of each utterance. The dialogue structure data 230 has a record for each dialogue structure, and each record is configured with items of utterance ID 231 of utterance A and utterance ID 232 of utterance B. The values ​​of the utterance IDs 231 and 232 correspond to the value of the utterance ID 221 in the past utterance information group 220 of Fig. 4.

[0042] In this description, when one utterance in a dialogue is related to another utterance, the former utterance will be referred to as "utterance A" and the latter as "utterance B." In other words, utterance B is an utterance that occurred earlier than utterance A, and the speaker of utterance A and the speaker of utterance B are different people.

[0043] Furthermore, in the dialogue structure estimation result, the combination of utterance A and utterance B does not necessarily have to be one-to-one. Specifically, for example, in the case of Fig. 5, utterance B with utterance ID "4" has a dialogue structure (has a dialogue relationship) with utterance A with utterance ID "5" and utterance A with utterance ID "6".

[0044] As described above, after the dialogue structure estimation unit 112 estimates the dialogue structure in step S104, the utterance set generation unit 114 of the perspective-specific dialogue extraction unit 113 generates an utterance set having a strong dialogue relationship based on the dialogue structure estimation result (step S105).

[0045] An example of a method for generating an utterance set will be described. The utterance set generation unit 114 extracts one utterance that satisfies a predetermined condition (for example, the utterance that was most recent) from among past utterances that are related to one utterance at a probability value equal to or greater than an arbitrary threshold, and creates an utterance set by combining these utterances. The predetermined condition is, empirically, the past utterance that was most recent, but is not limited to this. The threshold is set by the user in the range from 0 to 1, with 0.5 as the base, for example, depending on the purpose of extraction. To increase the recall rate (to prevent extraction omissions), the threshold can be set low. To increase the recall rate, the extraction target for the utterance set can be broadened.

[0046] As a modified example of the method for generating an utterance set, the utterance set generation unit 114 may extract, for one utterance, multiple utterances that are close in utterance time from among past utterances, and set the size of the utterance set to 2 or more. Alternatively, multiple utterances that are close in utterance time may be extracted from past utterances that have a relationship with each other with a probability value equal to or greater than an arbitrary threshold value.

[0047] Returning to the description of Fig. 2, after generating the utterance set in step S105, the utterance set generation unit 114 inputs the generated utterance set to the viewpoint determination unit 115 (step S106).

[0048] Next, the viewpoint determination unit 115 determines whether or not a viewpoint is present in each utterance set by using a determination model trained by a statistical method (step S107). The determination model is stored in the viewpoint determination model storage unit 122.

[0049] An example of a method for determining the presence or absence of a viewpoint by the viewpoint determination unit 115 in step S107 will be described in detail. First, a pre-trained language model is used as the determination model. The viewpoint determination unit 115 then uses the utterances in the utterance set arranged in chronological order and their labels (presence or absence of a viewpoint) as one sample, and fine-tunes the pre-trained language model using this sample set to train the determination model. For example, when the number of utterances constituting the utterance set is two, it can be treated as a sentence pair classification problem. Also, when the number of utterances constituting the utterance set is three or more, it can be treated as a classification problem in the same way.

[0050] Next, the viewpoint-specific dialogue extraction unit 113 (e.g., viewpoint determination unit 115) determines an utterance set having a viewpoint as an extraction target based on the determination result of step S107 (step S108). Then, the viewpoint-specific dialogue extraction unit 113 stores dialogue information related to the determined utterance set to be extracted in the extraction result storage unit 125.

[0051] As described above, in this embodiment, "having a viewpoint" means that some meaning is given from the viewpoint. Therefore, when the viewpoint of interest is whether the customer agrees or disagrees, not only an utterance set including an utterance of agreement by the customer but also an utterance set including an utterance of disagreement by the customer are determined as utterance sets having a viewpoint.

[0052] Fig. 6 is a diagram showing an example of dialogue information related to an utterance set having a viewpoint. The dialogue information 240 shown in Fig. 6 has a record for each utterance set determined to have a viewpoint, and each record is configured with the following items: user ID 241 of utterance A, utterance ID 242 of utterance A, utterance A 243, user ID 244 of utterance B, utterance ID 245 of utterance B, utterance B 246, and viewpoint 247.

[0053] In the dialogue information 240, the user ID 241 of utterance A, the utterance ID 242 of utterance A, and the utterance A243 are utterance information related to the answerer's "utterance A" that constitutes the utterance set. The utterance ID 242 of utterance A can be obtained from the utterance ID 231 of the dialogue structure data 230 (see FIG. 5). Then, by referring to the past utterance information group 220 using this utterance ID as a key (see FIG. 4), the user ID 241 can be obtained from the user ID 223, and the utterance A243 can be obtained from the text 224. Furthermore, the user ID 244 of utterance B, the utterance ID 245 of utterance B, and the utterance B246 are utterance information related to the asker's "utterance B" that constitutes the utterance set. Therefore, similarly to the utterance information regarding utterance A, the utterance ID 245 of utterance B can be obtained from the utterance ID 232 in the dialogue structure data 230, the user ID 244 from the user ID 223 in the past utterance information group 220, and the utterance B 246 from the text 224 in the past utterance information group 220. Furthermore, the viewpoint 247 indicates the content of the viewpoint that the utterance set has. The content of the viewpoint is, for example, "agreement" or "disagreement," and these contents are determined when the viewpoint determination unit 115 determines whether or not a viewpoint exists.

[0054] After step S108, the extraction result display unit 116 filters the dialogue information to be extracted stored in the extraction result storage unit 125 using the filter conditions specified by the user, and outputs the dialogue information related to the utterance set that meets the filter conditions as the extraction result (step S109).

[0055] 7 is a diagram showing an example of the display of the dialogue information extraction results. The display screen 250 shown in FIG. 7 is an example of a screen displayed on a terminal (which may be a display provided in the dialogue information extraction device 100) that can be operated by a user using, for example, a GUI, and is configured to have an input screen 251 and an output screen 252. The input screen 251 accepts filter conditions specified by the user and displays the specified contents. Note that the filter conditions can be arbitrarily specified by the user based on keywords, speaking users, viewpoints, etc., but the condition items are not limited to these. The output screen 252 displays the dialogue information extraction results based on the filter conditions.

[0056] 7, on the input screen 251, "contract" is specified in the keyword search, "User-B" is specified in the user search, and "disagree" is specified in the extraction viewpoint. In this case, the extraction result display unit 116 searches for records in which the utterance A243 or the utterance B245 contains the term "contract," the user ID 241 or the user ID 244 is "User-B," and the viewpoint 247 is "disagree," from the dialogue information 240 (see FIG. 6) stored in the extraction result storage unit 125. Note that the dialogue information extraction device 100 may be able to accept an input operation by the user on the input screen 251 at any timing before step S109 in FIG. 2. Then, in step S109 in FIG. 2, the extraction result display unit 116 acquires the contents of records that meet the above-mentioned various search conditions (filter conditions) from the dialogue information to be extracted stored in the extraction result storage unit 125, and displays the contents on the output screen 252.

[0057] The output screen 252 output as described above is an extracted dialogue content based on the viewpoint that the user focuses on, and the user can efficiently check the dialogue status by looking at the display screen 250.

[0058] It should be noted that the output format of the extraction results by the extraction result display unit 116 is not limited to the screen configuration of the display screen 250 in Fig. 7. Alternatively, for example, text data of utterances in a dialogue scenario may be displayed in the form of a real-time response chat in one area of ​​the display screen 250. In this case, it is possible to more clearly display in what kind of response the dialogue having the focused viewpoint occurred.

[0059] After step S109, the utterance information input unit 111 adds the utterance information input in step S101 to the past utterance information group of the current dialogue scenario stored in the past utterance information storage unit 123 (step S110), and ends the dialogue information extraction process. Note that step S110 for registering the utterance information may be executed at any timing after step S101.

[0060] By executing the dialogue information extraction process as described above, the dialogue information extraction device 100 estimates the dialogue structure and determines the viewpoints from a sequence of utterances by multiple speakers, extracts an utterance set including utterances having the viewpoints, and outputs dialogue information related to the utterance set. The dialogue information extraction device 100 according to this embodiment can extract information on a dialogue having a viewpoint of interest from a dialogue between multiple speakers, thereby enabling efficient confirmation of the dialogue situation.

[0061] Furthermore, by making a judgment regarding a viewpoint on a per-utterance set basis, the dialogue information extraction device 100 can accurately determine not only whether a viewpoint exists but also the content of the viewpoint (for example, agreement or disagreement), thereby improving the granularity of the extracted dialogue information.

[0062] Furthermore, the dialogue information extraction device 100 filters the extracted dialogue information using a filter condition specified by the user and outputs the filtered information, allowing the user to check the dialogue status more efficiently.

[0063] Furthermore, the dialogue information extraction device 100 sequentially executes the dialogue information extraction process shown in Fig. 2 every time an utterance is generated (every time text utterance information is input), thereby enabling online processing. At this time, if the screen display of the dialogue information extraction result is configured to be updated every time an utterance is generated (every time text utterance information is input), the user can check the dialogue status (for example, the status of viewpoints such as agreement or disagreement) in real time. The ability to extract information related to a specific viewpoint while continuously monitoring the utterance status online is useful for reducing the risk of failing to confirm consent, etc.

[0064] (2) Second embodiment 8 is a block diagram showing an example of the configuration of a dialogue information extraction device 101 according to the second embodiment of the present invention. The dialogue information extraction device 101 has a configuration in which a judgment target selection unit 117 is added to the dialogue information extraction device 100 shown in FIG. 1 in the first embodiment, and a description of the configuration common to the dialogue information extraction device 100 will be omitted.

[0065] The judgment target selection unit 117 has the function of selecting an utterance (a judgment target utterance) to be judged for dialogue information extraction from a series of dialogues (multiple utterances) in a dialogue scenario based on a predetermined judgment condition (which may be specified by the user at any time).

[0066] As the above-mentioned determination conditions, for example, a "start keyword utterance" for determining the start of the dialogue information extraction process and an "end keyword utterance" for determining the end of the dialogue information extraction process are set. At this time, the determination target selection unit 117 searches for the presence of a "start keyword utterance" or an "end keyword utterance" in the text data of utterances (utterance information) input to the dialogue information extraction device 101 as the dialogue scenario progresses, and selects the utterance from the utterance in which the start keyword utterance first appeared to the utterance in which the end keyword utterance first appeared as the determination target utterance. Furthermore, as another determination condition, for example, if the system is configured so that one of the speakers (such as an operator) presses a predetermined switch at the start of the dialogue section to be determined, the operation of this switch can be used as the determination condition.

[0067] Fig. 9 is a flowchart showing an example of the processing procedure of the dialogue information extraction processing in the second embodiment. In Fig. 9, the same processes as those in the dialogue information extraction processing in the first embodiment shown in Fig. 2 are denoted by the same reference numerals. The processing procedure of the dialogue information extraction processing shown in Fig. 9 will be described below, focusing on the differences from the dialogue information extraction processing in Fig. 2.

[0068] 9, first, the utterance information input unit 111 inputs utterance information (step S101). Next, the judgment target selection unit 117 judges whether or not the utterance information input in step S101 is an utterance to be judged based on a predetermined judgment condition (step S201).

[0069] If it is determined in step S201 that the utterance is not a target utterance to be determined (NO in step S201), there is no need to extract dialogue information, and therefore the dialogue information extraction process is terminated without performing any special process.

[0070] On the other hand, if it is determined in step S201 that the utterance is a target utterance to be determined (YES in step S201), the utterance information input unit 111 refers to the past utterance information group of the current dialogue scenario stored in the past utterance information storage unit 123, and checks whether or not there is past utterance information that satisfies the determination condition of the target utterance (step S202). In the second embodiment, only past utterance information that satisfies the determination condition of the target utterance is stored in the past utterance information storage unit 123. Therefore, in practice, in step S202, the utterance information input unit 111 only needs to check whether or not there is past utterance information in the past utterance information group of the current dialogue scenario.

[0071] If past utterance information satisfying the determination condition for the determination target utterance is present in step S202 (YES in step S202), the processes of steps S103 to S108 are performed, whereby an utterance set including an utterance having a viewpoint is determined as an extraction target for dialogue information, and dialogue information related to the utterance set to be extracted is stored in the extraction result storage unit 125. Next, the extraction result display unit 116 outputs, as the extraction result, dialogue information that meets the filter condition specified by the user from the dialogue information stored in the extraction result storage unit 125 (step S109). Finally, in step S110, the utterance information input unit 111 adds the utterance information input in step S101 to a group of past utterance information of the current dialogue scenario stored in the past utterance information storage unit 123 (step S110), and the dialogue information extraction process ends.

[0072] On the other hand, if there is no past utterance information that satisfies the determination condition of the utterance to be determined in step S202 (YES in step S202), no dialogue from which dialogue information is to be extracted has been formed in the dialogue scenario at this time. Therefore, the utterance information input unit 111 stores the utterance information input in step S101 in the past utterance information storage unit 123 as a past utterance information group of the corresponding dialogue scenario (step S110), and terminates the dialogue information extraction process. Note that if the utterance information input in step S101 relates to the first utterance of a new dialogue scenario, the utterance information input unit 111 may store the utterance information in the past utterance information storage unit 123 as a past utterance information group of the new dialogue scenario.

[0073] By performing the dialogue information extraction process as described above, the dialogue information extraction device 101 according to the second embodiment can extract dialogue information having a viewpoint after limiting the utterances to be judged, thereby reducing the processing load of the processing units (e.g., the dialogue structure estimation unit 112 and the viewpoint-specific dialogue extraction unit 113) and suppressing the capacity consumption of the storage units (e.g., the past utterance information storage unit 123 and the dialogue structure storage unit 124) compared to the first embodiment. In addition, by appropriately setting judgment conditions such as "start keyword utterance" and "end keyword utterance", it is possible to exclude utterances unrelated to the viewpoint, thereby achieving the effect of further improving the granularity of the extracted dialogue information.

[0074] (3) Hardware configuration Finally, an example of the hardware configuration of the conversation information extraction device 100, 101 according to the first or second embodiment will be described.

[0075] Fig. 10 is a block diagram showing an example of the hardware configuration of the dialogue information extraction devices 100, 101. As shown in Fig. 10, the dialogue information extraction devices 100, 101 are configured to include a CPU 11, a ROM 12, a RAM 13, an input / output interface 14, a storage device 15, a drive device 16, and a communication interface 17. The storage device 15 stores various data and a program 20 used by the dialogue information extraction devices 100, 101. The CPU 11 reads the program 20 into the RAM 13 and executes it, thereby realizing each processing unit of the dialogue information extraction devices 100, 101 (an utterance information input unit 111, a dialogue structure estimation unit 112, a viewpoint-specific dialogue extraction unit 113, an extraction result display unit 116, and a judgment target selection unit 117). [Explanation of symbols]

[0076] 11 CPU 12 ROM 13 RAM 14 Input / Output Interface 15 Storage device 16 Drive device 17 Communication Interface 20 Programs 100,101 Dialogue information extraction device 111 Speech information input unit 112 Dialogue Structure Estimation Unit 113 Perspective-based dialogue extraction unit 114 Utterance Set Generation Unit 115 Viewpoint Determination Section 116 Extraction result display area 117 Judgment target selection unit 121 Dialogue structure estimation model memory unit 122 Viewpoint judgment model memory unit 123 Past speech information storage unit 124 Dialogue Structure Memory Unit 125 Extraction result storage section 210 Speech Information 220 Past speech information set 230 Dialogue Structure Data 240 Dialogue Information 250 display screen

Claims

1. A dialogue information extraction device that analyzes a dialogue between multiple speakers based on a viewpoint to be noted, an utterance information input unit that inputs, for each utterance constituting the dialogue, utterance information including a text transcribed from the utterance; a dialogue structure estimation unit that estimates a dialogue structure of the dialogue from the utterance information of a plurality of utterances in the dialogue using a trained estimation model; a viewpoint-specific dialogue extraction unit that extracts a group of utterances having the viewpoint by determining whether or not the viewpoint exists in a group of utterances made up of utterances from a plurality of speakers, the group of utterances being generated based on the dialogue structure; Equipped with the dialogue structure estimation unit uses the estimation model to calculate a degree of relationship between two utterances by different speakers for a plurality of utterances in the dialogue, and estimates a combination of utterances having a dialogue relationship from the calculation result, thereby estimating the dialogue structure; The viewpoint-specific dialogue extraction unit an utterance set generation unit that generates an utterance set by selecting one or more utterance combinations from the utterance combinations estimated by the dialogue structure estimation unit as an utterance group consisting of utterances from the plurality of speakers; a viewpoint determination unit that determines, for each utterance set, whether or not the viewpoint is present in the utterance set using a determination model that has been trained in advance, and extracts an utterance set that includes the viewpoint from the determination result. A dialogue information extraction device characterized by:

2. an extraction result display unit that outputs dialogue information including the text of each utterance constituting the utterance group and the content of the viewpoint of the utterance group, for the utterance group extracted by the viewpoint-specific dialogue extraction unit.

2. The dialogue information extraction device according to claim 1.

3. The utterance set generation unit generates, as the utterance set, a combination of utterances that have a dialogue relationship with a probability equal to or greater than a specified value from among the combinations of utterances estimated by the dialogue structure estimation unit.

2. The dialogue information extraction device according to claim 1.

4. The utterance set generation unit generates, as the utterance sets, combinations of an utterance from the answering side and an utterance from the questioning side that occurred earlier than the utterance, from among the combinations of utterances estimated by the dialogue structure estimation unit, up to a predetermined number in order from the combination with the smallest time difference.

2. The dialogue information extraction device according to claim 1.

5. The extraction result display unit extracts a group of utterances that meets a filter condition designated by a user from the group of utterances having the viewpoint extracted by the viewpoint-specific dialogue extraction unit, and outputs the dialogue information about the extracted group of utterances.

3. The dialogue information extraction device according to claim 2.

6. Each time an utterance is made as the dialogue progresses, the utterance information input unit inputs the utterance information of the utterance; The input of the utterance information is a trigger for the processing by the dialogue structure estimation unit and the viewpoint-specific dialogue extraction unit, and the extraction result display unit updates the output of the dialogue information based on the results of the processing.

3. The dialogue information extraction device according to claim 2.

7. a determination target selection unit that determines, based on a predetermined determination condition, whether or not the utterance information to be input by the utterance information input unit is to be processed by the dialogue structure estimation unit and the viewpoint-specific dialogue extraction unit.

2. The dialogue information extraction device according to claim 1.

8. When the answerer's utterance expresses agreement or disagreement with the questioner's utterance, it is considered that the answerer has the above viewpoint.

2. The dialogue information extraction device according to claim 1.

9. A dialogue information extraction method for a dialogue information extraction device that analyzes a dialogue between multiple speakers based on a viewpoint to be focused on, comprising: an utterance information input step in which the dialogue information extraction device inputs utterance information including a text transcribed from each utterance constituting the dialogue; a dialogue structure estimation step in which the dialogue information extraction device estimates a dialogue structure of the dialogue from the utterance information of a plurality of utterances in the dialogue inputted in the utterance information input step, using a trained estimation model; a viewpoint-based dialogue extraction step in which the dialogue information extraction device extracts a group of utterances having the viewpoint by determining whether or not the viewpoint exists in a group of utterances made up of utterances from a plurality of speakers, the group of utterances being generated based on the estimation result of the dialogue structure estimation step; Equipped with the dialogue structure estimation step calculates a degree of relationship between two utterances by different speakers for a plurality of utterances in the dialogue using the estimation model, and estimates a combination of utterances having a dialogue relationship from the calculation result, thereby estimating the dialogue structure; The viewpoint-based dialogue extraction step includes: an utterance set generating step of selecting one or more utterance combinations from the utterance combinations estimated by the dialogue structure estimating step to generate an utterance set as a group of utterances made up of utterances from the plurality of speakers; a viewpoint determination step of determining, for each utterance set, whether or not the viewpoint is present in the utterance set using a determination model trained in advance, and extracting an utterance set having the viewpoint from the determination result. A dialogue information extraction method comprising:

Citation Information

Patent Citations

  • Dialogue action estimation method, dialogue action estimation device and program

    JP2020095732A

  • Call center conversational content display system, method, and program

    WO2019003395A1