Speech section extraction device, speech section extraction method, and speech section extraction program

The speech section extraction device accurately determines and extracts important speech sections by analyzing transitions and combinations of utterances, improving call analysis and sales effectiveness in contact centers.

JP7786470B2Active Publication Date: 2025-12-16NIPPON TELEGRAPH & TELEPHONE CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
JP2023564724
Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
Filing Date
2021-12-03
Publication Date
2025-12-16
Estimated Expiration
2041-12-03

Smart Images

  • Figure 0007786470000001
    Figure 0007786470000001
  • Figure 0007786470000002
    Figure 0007786470000002
  • Figure 0007786470000003
    Figure 0007786470000003
Patent Text Reader

Abstract

This speech segment extraction device comprises: a speech segment identification unit that identifies a speech segment which includes at least one speech item, from speech text data which includes the speech of two or more people; a speech segment type determination unit that determines a speech segment type for each of the identified speech segments; a speech type extraction unit that extracts, from the speech text data, a speech type for each speech item included in the speech text data; and a speech segment extraction unit that extracts an important speech segment from among the identified speech segments, on the basis of a combination and transition of the identified speech segment types, as well as a combination and transition of the extracted speech types.
Need to check novelty before this filing date? Find Prior Art

Description

[Technical Field]

[0001] The disclosed technology relates to a speech segment extraction device, a speech segment extraction method, and a speech segment extraction program. [Background technology]

[0002] In contact centers of companies and organizations, operators use telephone calls, text chat, etc. to handle a large volume of customer inquiries, propose products to customers, make sales, etc. While handling a large volume of calls every day, selling products effectively in response to customer needs and requests and providing accurate answers to inquiries leads to increased profits and improved customer satisfaction.

[0003] In order to effectively utilize opportunities for calls from and to customers, it is necessary to extract examples of excellent customer service from call data and share and analyze them within the company or among operators.

[0004] A good response is made up of a series of interactions, and is determined by taking into consideration multiple utterances, the transition of utterance sections, the structure, and the frequency and location of questions and answers. Such sections are designated as important utterance sections (hereinafter referred to as "important utterance sections") and are extracted.

[0005] For example, it is conceivable to determine important speech sections using keywords, etc. However, this requires manually checking and determining the sections before and after the utterances searched for using keywords, and it is not possible to narrow down and extract the desired speech section. There is also a method in which speech sections are defined in advance and then judged based on the similarity of words appearing within the utterance unit (see, for example, Patent Document 1). [Prior art documents] [Patent documents]

[0006] [Patent Document 1] WO2020 / 036190 issue Summary of the Invention [Problem to be solved by the invention]

[0007] According to the technology described in Patent Document 1, it is possible to determine the importance and superiority / inferiority of each section, but it is not possible to determine an utterance section while taking into consideration transitions of important utterance sections.

[0008] It is also possible to analyze each utterance made during a call and extract utterances that include a specific way of progressing or developing the utterance during the call. However, as shown in Figure 25, although it is possible to grasp specific utterance situations such as "needs hearing," "no needs," "question," and "answer," it is not possible to determine and extract a corresponding pattern for the next utterance section based on the answer or reaction from the other party.

[0009] Another possible method is to define speech intervals and determine the types of speech intervals (for example, "open-type business interval," "theme-type business interval," "end-type business interval," etc.), and then make a judgment based on a rule based on a combination of the types of speech intervals, as shown in Figure 25. However, the structure and frequency of speech within an interval are important for determining an important speech interval, and the judgment cannot be made based on the information about the speech interval alone.

[0010] In other words, conventional techniques are unable to determine section information for each utterance section and to determine and extract important utterance sections that are useful for call analysis in sales, etc. Furthermore, when analysis information on an utterance-by-utterance basis is used alone, it is unable to determine important utterance sections taking into account the transitions in each utterance section.

[0011] The disclosed technology has been made in consideration of the above points, and aims to provide a speech section extraction device, a speech section extraction method, and a speech section extraction program that can extract important speech sections by taking into account the speech sections and the respective combinations and transitions of utterances. [Means for solving the problem]

[0012] A first aspect of the present disclosure is a speech section extraction device comprising: a speech section identification unit that identifies a speech section including at least one utterance from speech text data including utterances from two or more people; a speech section type determination unit that determines the speech section type for each of the speech sections identified by the speech section identification unit; an utterance type extraction unit that extracts, from the utterance text data, an utterance type for each utterance included in the speech text data; and a speech section extraction unit that extracts important speech sections from the speech sections identified by the speech section identification unit based on combinations and transitions of speech section types determined by the speech section type determination unit and combinations and transitions of speech types extracted by the speech type extraction unit.

[0013] A second aspect of the present disclosure is a speech section extraction method, which identifies a speech section containing at least one utterance from speech text data containing utterances from two or more people, determines a speech section type for each of the identified speech sections, extracts from the speech text data a speech type for each utterance included in the speech text data, and extracts important speech sections from the identified speech sections based on combinations and transitions of the determined speech section types and combinations and transitions of the extracted utterance types.

[0014] A third aspect of the present disclosure is a speech section extraction program that causes a computer to execute the following steps: identify speech sections containing at least one utterance from speech text data containing utterances from two or more people; determine a speech section type for each of the identified speech sections; extract from the speech text data a speech type for each utterance included in the speech text data; and extract important speech sections from the identified speech sections based on combinations and transitions of the determined speech section types and combinations and transitions of the extracted speech types. [Effects of the Invention]

[0015] The disclosed technology has an advantage that it is possible to extract important speech sections by taking into consideration the speech sections and the respective combinations and transitions of speech sections. [Brief explanation of the drawings]

[0016] [Figure 1] FIG. 1 is a block diagram illustrating an example of a hardware configuration of an utterance period extraction device according to an embodiment. [Figure 2] FIG. 2 is a diagram illustrating terms used in the embodiment. [Figure 3] 1 is a block diagram showing an example of a functional configuration of an utterance period extraction device according to an embodiment. [Figure 4] 4 is a diagram illustrating an example of the configuration of a sentence input unit illustrated in FIG. 3. FIG. [Figure 5] 4 is a diagram illustrating an example of the configuration of an utterance section identification unit illustrated in FIG. 3. FIG. [Figure 6] 4 is a diagram illustrating an example of the configuration of a speech section type determination unit illustrated in FIG. 3. FIG. [Figure 7] 4 is a diagram illustrating an example of the configuration of an utterance type extraction unit illustrated in FIG. 3. FIG. [Figure 8] 4 is a diagram illustrating an example of the configuration of an utterance period extraction unit illustrated in FIG. 3. FIG. [Figure 9] 3A to 3C are diagrams illustrating an important speech section extraction process according to the first embodiment. [Figure 10] 10 is a flowchart showing an example of a processing flow by the speech segment extraction program according to the first embodiment. [Figure 11] 1 is a flowchart showing an example of the flow of an important speech section extraction process according to the first embodiment, illustrating an example of rule A of the speech section extraction rules. [Figure 12] 10 is a flowchart showing another example of the flow of the important speech section extraction process according to the first embodiment, illustrating an example of rule B of the speech section extraction rules. [Figure 13] FIG. 10 is a diagram illustrating an example of the configuration of an utterance section identification unit according to the second embodiment. [Figure 14] FIG. 10 is a diagram illustrating an example of the configuration of a speech section type determination unit according to the second embodiment. [Figure 15] FIG. 10 is a diagram illustrating an example of the configuration of an utterance type extraction unit according to the second embodiment. [Figure 16]FIG. 10 is a diagram illustrating an example of the configuration of an utterance period extraction unit according to the second embodiment. [Figure 17] 10A and 10B are diagrams illustrating an important speech section extraction process according to the second embodiment. [Figure 18] FIG. 10 is a diagram illustrating another important speech section extraction process according to the second embodiment. [Figure 19] 10 is a flowchart showing an example of the flow of speech section extraction processing according to the second embodiment, illustrating an example of rule C of the speech section extraction rules. [Figure 20] FIG. 11 is a diagram illustrating an example of the configuration of a speech section type determination unit according to the third embodiment. [Figure 21] FIG. 11 is a diagram illustrating an example of the configuration of an utterance type extraction unit according to the third embodiment. [Figure 22] FIG. 11 is a diagram illustrating an example of the configuration of an utterance period extraction unit according to the third embodiment. [Figure 23] 10A and 10B are diagrams illustrating an important speech section extraction process according to the third embodiment. [Figure 24] 13 is a flowchart showing an example of the flow of speech section extraction processing according to the third embodiment, illustrating an example of rule D of the speech section extraction rules. [Figure 25] FIG. 1 is a diagram illustrating the prior art. DETAILED DESCRIPTION OF THE INVENTION

[0017] An example of an embodiment of the disclosed technology will be described below with reference to the drawings. Note that in each drawing, the same or equivalent components and parts are given the same reference numerals. Also, the dimensional proportions in the drawings are exaggerated for the convenience of explanation and may differ from the actual proportions.

[0018] [First embodiment] The speech section extraction device of the first embodiment provides specific improvements over conventional methods of extracting important speech sections without considering the speech sections and the respective combinations and transitions of speech, and represents an advancement in the technical field of extracting important speech sections from speech data containing speech from two or more people.

[0019] The speech section extraction device of this embodiment identifies the speech section to be analyzed, determines the speech section type that represents the type of the identified speech section, extracts the speech type of each utterance unit, and extracts important speech sections from the identified speech sections based on the respective combinations and transitions of the speech section types and utterance types.

[0020] First, with reference to FIG. 1, the hardware configuration of a speech segment extraction device 10 according to this embodiment will be described.

[0021] FIG. 1 is a block diagram showing an example of the hardware configuration of an utterance segment extraction device 10 according to this embodiment.

[0022] 1, the speech segment extraction device 10 includes a CPU (Central Processing Unit) 11, a ROM (Read Only Memory) 12, a RAM (Random Access Memory) 13, a storage 14, an input unit 15, a display unit 16, and a communication interface (I / F) 17. Each component is connected to each other via a bus 18 so as to be able to communicate with each other.

[0023] The CPU 11 is a central processing unit that executes various programs and controls each component. That is, the CPU 11 reads a program from the ROM 12 or the storage 14 and executes the program using the RAM 13 as a work area. The CPU 11 controls the above components and performs various arithmetic processing in accordance with the program stored in the ROM 12 or the storage 14. In this embodiment, the ROM 12 or the storage 14 stores a speech section extraction program for executing speech section extraction processing.

[0024] The ROM 12 stores various programs and various data. The RAM 13 temporarily stores programs or data as a working area. The storage 14 is configured with an HDD (Hard Disk Drive) or SSD (Solid State Drive) and stores various programs including the operating system and various data.

[0025] The input unit 15 includes a pointing device such as a mouse and a keyboard, and is used to input various information to the device itself.

[0026] The display unit 16 is, for example, a liquid crystal display, and displays various information. The display unit 16 may also function as the input unit 15 by adopting a touch panel system.

[0027] The communication interface 17 is an interface for the device itself to communicate with other external devices. For this communication, a wired communication standard such as Ethernet (registered trademark) or FDDI (Fiber Distributed Data Interface) or a wireless communication standard such as 4G, 5G, or Wi-Fi (registered trademark) is used.

[0028] The speech segment extraction device 10 according to this embodiment is implemented as a general-purpose computer device such as a server computer or a personal computer (PC).

[0029] Here, the terms used in this embodiment will be explained with reference to FIG.

[0030] FIG. 2 is a diagram illustrating the terms used in this embodiment. As shown in FIG. 2, an utterance is input as a segment from voice recognition, text chat, or the like. The utterance text includes all exchanges during one conversation and represents a collection of all utterances during one call. The utterance type represents the type of each utterance. The utterance type is not affected by the utterances before and after it. Examples of utterance types include "needs hearing." The utterance type ID (Identification) is an ID assigned to each utterance to identify the utterance type. The utterance section represents a section including at least one utterance. The utterance section may be a collection of multiple consecutive utterances, in which case, one utterance section is composed of multiple utterances. The utterance section is composed of, for example, a scene such as a greeting, the meaning and content of the utterance, and the like, as a single unit. The utterance section may also be a single utterance, in which case, one utterance section is composed of one utterance. The speech section ID is an ID assigned to each utterance to identify the speech section. The speech section type indicates the type of each speech section. The speech section type includes, for example, a "theme-type sales section." The speech section type ID is an ID assigned to each utterance to identify the speech section type for each speech section.

[0031] Next, the functional configuration of the speech segment extraction device 10 will be described with reference to FIG.

[0032] FIG. 3 is a block diagram showing an example of the functional configuration of the speech segment extraction device 10 according to this embodiment.

[0033] 3, the speech section extraction device 10 includes, as functional components, a sentence input unit 101, a speech section identification unit 102, a speech section type determination unit 103, an utterance type extraction unit 104, a speech section extraction unit 105, and an output unit 106. Each functional component is realized by the CPU 11 reading out the speech section extraction program stored in the ROM 12 or the storage 14, expanding it in the RAM 13, and executing it.

[0034] Note that the utterance DB (Data Base) 20 storing utterance data and the extraction result DB 25 storing extraction result data may each be stored in the storage 14 or in an external, accessible storage device. Similarly, the utterance text DB 21 storing utterance text data, the utterance section DB 22 storing utterance section data, the utterance section type DB 23 storing utterance section type data, and the utterance type DB 24 storing utterance type data may each be stored in the storage 14 or in an external, accessible storage device. Note that, although the utterance text data, the utterance section data, the utterance section type data, and the utterance type data are each stored in different DBs in the example of Fig. 3, they may be stored in a single DB.

[0035] 4 to 8, the configuration of each functional unit (sentence input unit 101, speech section identification unit 102, speech section type determination unit 103, speech type extraction unit 104, speech section extraction unit 105, and output unit 106) shown in FIG. 3 will be specifically described.

[0036] The sentence input unit 101 shown in FIG. 4 acquires utterance data from the utterance DB 20, converts the acquired utterance data, and stores the resulting utterance text data in the utterance text DB 21. The utterance data is data including utterances from two or more people, and may be a character string or a voice. When the utterance data is a voice, the sentence input unit 101 converts the utterance into text by performing voice recognition and stores the converted text in the utterance text DB 21. When the utterance data is a character string, the sentence input unit 101 stores the character string as it is in the utterance text DB 21 since it has already been converted into text. For example, the utterances in the dialogue example shown in FIG. 2 described above are stored as voice in the utterance DB 20 as voice data. When the utterance data is input as voice, the sentence input unit 101 converts the utterance data into text by using voice recognition and stores the resulting utterance text data in the utterance text DB 21.

[0037] The speech section identification unit 102 shown in Fig. 5 acquires utterance text data from the utterance text DB 21, identifies utterance sections from the acquired utterance text data, and stores the acquired utterance section data in the utterance section DB 22. Specifically, when the utterance text data is input, the utterance section identification unit 102 identifies utterance sections using the utterance section identification model 30, and stores the acquired utterance section data in the utterance section DB 22. The utterance section identification model 30 is a trained model that receives utterance text data as input and outputs utterance section data. For example, a DNN (Deep Neural Network), which is a multi-layered neural network, is used as the utterance section identification model 30. The utterance section identification model 30 may be stored in the storage 14 or an external storage device. The speech section identification model 30 is generated by assigning teacher labels to utterances containing clue words indicating topic changes, such as "Well then" or "By the way," and performing machine learning using the utterance text data with the teacher labels as training data to generate a model for determining topic changes. This speech section identification model 30 is used to determine speech changes, and the utterance from one topic change to the next is identified as a speech section. In other words, the speech section identification model 30 identifies a topic change utterance, which is an utterance in which the topic in the utterance text has changed. The utterance from the opening utterance to the utterance immediately before the topic change utterance, the utterance from the topic change utterance to the topic immediately before the next topic change utterance, and the utterance from the next topic change utterance to the closing utterance are each considered to be a speech section, and an utterance section ID is assigned to each utterance.

[0038] The speech section type determination unit 103 shown in FIG. 6 acquires speech section data from the speech section DB 22, determines the speech section type of the acquired speech section data, and stores the obtained speech section type data in the speech section type DB 23. Specifically, when the speech section data is input, the speech section type determination unit 103 determines the speech section type of the speech section using the speech section type determination model 31, and stores the obtained speech section type data in the speech section type DB 23. The speech section type determination model 31 is a trained model that receives the speech section data as input and outputs the speech section type data. For example, a DNN is used for the speech section type determination model 31. The speech section type determination model 31 may be stored in the storage 14 or in an external storage device. For example, the following labels (type 1 to type 4) are defined as the speech section types. A model that determines these speech section types is generated in advance by machine learning using learning data with these labels. Using this speech section type determination model 31, the speech section type is determined for the input speech section, and the determination result of the speech section type is assigned to each speech section as a speech section type ID.

[0039] (Type 1) A section where services are not limited to specific topics or themes (hereinafter referred to as "open-type business section") (Type 2) A section where the customer is asked if they have any other topics or themes (hereinafter referred to as the "end-type sales section"). Specifically, this is a speech section that ends the conversation on a specific topic or theme, or a speech section where the customer is asked if they have any other needs. (Type 3) A section where calls are made on specific topics or themes, such as topics prepared in advance (hereinafter referred to as "theme-based business section"). (Type 4) No type

[0040] The utterance type extraction unit 104 shown in FIG. 7 acquires utterance text data from the utterance text DB 21, extracts the utterance type of each utterance included in the acquired utterance text data, and stores the obtained utterance type data in the utterance type DB 24. Specifically, when utterance section data is input, the utterance type extraction unit 104 extracts the utterance type of each utterance using the utterance type extraction model 32 and stores the obtained utterance type data in the utterance type DB 24. The utterance type extraction model 32 is a trained model that receives utterance text data as input and outputs utterance type data. For example, a DNN is used for the utterance type extraction model 32. The utterance type extraction model 32 may be stored in the storage 14 or in an external storage device. As the utterance type, for example, a type is assigned to each of the following: a customer service scene, an utterance related to a dialogue act, and an utterance related to a sales act. In the case of a customer service scene, for example, labels are defined for an utterance in a scene where the business is understood, an utterance regarding the business, etc. For utterances related to dialogue acts, for example, labels such as "question," "answer," and "explanation" are defined. For utterances related to sales activities, for example, labels such as "needs hearing," "needs present," "no need," "proposal," and "question" are defined. A model for extracting these utterance types is generated in advance by machine learning using utterance text data with these labels attached to each utterance as training data. This utterance type extraction model 32 is used to determine the utterance type of each utterance for the input utterance text, and the utterance type determination result is assigned to each utterance as an utterance type ID.

[0041] The speech section extraction unit 105 shown in Fig. 8 acquires speech section type data from the speech section type DB 23 and acquires speech type data from the speech type DB 24. The speech section extraction unit 105 then extracts important speech sections from the speech sections identified by the speech section identification unit 102 based on the combinations and transitions of speech section types determined by the speech section type determination unit 103 and the combinations and transitions of speech types extracted by the speech type extraction unit 104. Specifically, the speech section extraction unit 105 performs extraction using the speech section extraction rules 33. The speech section extraction rules 33 predetermine combinations and transitions of speech section types and combinations and transitions of speech types associated with important speech sections. The speech section extraction rules 33 extract multiple speech sections as important speech sections, for example, when the multiple speech sections include a speech type indicating a sales utterance made by an operator to a customer and do not include a combination of speech section types indicating multiple consecutive speech sections predeterminably designated as unimportant sections. This makes it possible to extract important speech segments with high accuracy when one speech segment contains multiple speeches.

[0042] The speech segment extraction rules 33 include, for example, the following rule A. When rule A is satisfied, the segment is determined to be important.

[0043] (Rule A) a1. There must be three or more consecutive speech sections. a2. The combination of speech section types (C1 to C4) specified as non-important sections below is not included. (C1) "Open-type operating section" → "Open-type operating section" → "Other" (C2) "Open-type operating section" → "Open-type operating section" → "Open-type operating section" (C3) "Open-type operating section" → "Open-type operating section" → "End-type operating section" (C4) "End-type business section" → "End-type business section" → "Other"

[0044] Furthermore, the speech segment extraction rules 33 include, for example, the following rule B. When rule B is satisfied, the segment is determined to be important.

[0045] (Rule B) b1. In the response scene of the utterance included in the speech section, the utterance type includes "response." b2. The speech types included in the speech section include sales-related speech such as "needs hearing." b3. If the customer responds to the operator's "needs hearing" utterance and the response is negative ("no needs"), the operator then makes a "suggestion" or "question." b4. There are a certain number of questions uttered. b5. Questions are often used in the first half of the speech section. b6. Asking multiple questions before making a suggestion.

[0046] The output unit 106 shown in FIG. 8 acquires the extraction result data extracted by the speech segment extraction unit 105 and stores the acquired extraction result data in the extraction result DB 25.

[0047] Next, the important speech section extraction process according to the first embodiment will be specifically described with reference to FIG.

[0048] 9 is a diagram illustrating the important utterance section extraction process according to the first embodiment. The utterance text shown in FIG. 9 includes an utterance by an operator and an utterance by a customer. The utterance text includes a plurality of utterance sections W1 to W4, and each of the utterance sections W1 to W4 includes a plurality of utterances.

[0049] As shown in Figure 9, the type of speech section W1 is an "open sales section," the type of speech section W2 is a "theme sales section," the type of speech section W3 is a "theme sales section," and the type of speech section W4 is an "end sales section." Furthermore, speech section W1 includes the utterances "response," "needs hearing," and "no needs," while speech section W2 includes the utterances "proposal" and "needs present." Speech section W3 includes the utterances "needs hearing" and "no needs," while speech section W4 includes the utterances "needs hearing" and "no needs."

[0050] In the example of FIG. 9, as described above, the utterance types include utterances indicating sales-related utterances made by an agent to a customer, and do not include combinations of utterance section types indicating multiple consecutive utterance sections previously designated as non-important sections (combinations C1 to C4 described above). A combination of multiple consecutive utterance sections that is designated as a non-important section is a combination that includes at least one of an "open-type sales section" and an "end-type sales section" among "open-type sales section," "theme-type sales section," and "end-type sales section," but does not include a "theme-type sales section." In other words, the utterance sections W1 to W3 are extracted as important (good) utterance sections because the utterance type is "response," they include sales-related utterances such as "needs hearing," and the combination of utterance section types is a transition from "open-type sales section" to "theme-type sales section" to "theme-type sales section." Furthermore, the speech sections W2 to W4 are extracted as important (good) speech sections because the speech type is "response" and includes sales-related speech such as "needs hearing," and the combination of speech section types is a transition from "theme-type sales section" to "theme-type sales section" to "end-type sales section."

[0051] Next, the operation of the speech segment extraction device 10 according to the first embodiment will be described with reference to FIG.

[0052] 10 is a flowchart showing an example of the flow of processing by the speech section extraction program according to the embodiment 1. The processing by the speech section extraction program is realized by the CPU 11 of the speech section extraction device 10 writing the speech section extraction program stored in the ROM 12 or the storage 14 into the RAM 13 and executing it.

[0053] In step S101 of FIG. 10, the CPU 11 accepts input of utterance data from the utterance DB 20, converts the accepted utterance data, and stores the resulting utterance text data in the utterance text DB 21.

[0054] In step S102, the CPU 11 acquires utterance text data from the utterance text DB 21, identifies an utterance section corresponding to the acquired utterance text data using the utterance section identification model 30, and stores the acquired utterance section data in the utterance section DB 22.

[0055] In step S103, the CPU 11 acquires speech section data from the speech section DB 22, determines the speech section type for the acquired speech section data using the speech section type determination model 31, and stores the acquired speech section type data in the speech section type DB 23.

[0056] In step S104, the CPU 11 acquires utterance text data from the utterance text DB 21, extracts the utterance type of each utterance contained in the acquired utterance text data using the utterance type extraction model 32, and stores the acquired utterance type data in the utterance type DB 24.

[0057] In step S105, the CPU 11 acquires speech section type data from the speech section type DB 23 and acquires speech type data from the speech type DB 24, and extracts important speech sections from the speech sections identified in step S102 based on the combination and transition of the speech section types determined in step S103 and the combination and transition of the speech types extracted in step S104. Specifically, the extraction is performed using the speech section extraction rule 33. A specific example of this important speech section extraction process will be described with reference to Figs. 11 and 12.

[0058] FIG. 11 is a flowchart showing an example of the flow of the important speech section extraction process according to the first embodiment, and shows an example of rule A of the speech section extraction rules 33.

[0059] In step S111, the CPU 11 acquires utterance section type data from the utterance section type DB 23 and acquires utterance type data from the utterance type DB 24.

[0060] In step S112, the CPU 11 determines whether or not there are three or more consecutive utterance periods based on the utterance period type data and utterance type data acquired in step S111. If it is determined that there are three or more consecutive periods (in the case of a positive determination), the process proceeds to step S113, and if it is determined that there are not three or more consecutive periods (in the case of a negative determination), the process proceeds to step S115.

[0061] In step S113, CPU 11 determines whether the combination of consecutive speech section types is a pre-designated combination of non-important sections (for example, the above-mentioned combinations of non-important sections C1 to C4). If it is determined that the combination is not a pre-designated combination of non-important sections (in the case of a negative determination), the process proceeds to step S114, and if it is determined that the combination is a pre-designated combination of non-important sections (in the case of a positive determination), the process proceeds to step S115.

[0062] In step S114, the CPU 11 determines that the continuous utterance section is an important utterance section, and returns to step S106 in FIG.

[0063] In step S115, the CPU 11 determines that the utterance section is not an important utterance section, and returns to step S106 in FIG.

[0064] Fig. 12 is a flowchart showing another example of the flow of the important speech section extraction process according to the first embodiment, and shows an example of rule B of the speech section extraction rules 33. The process of Fig. 12 may be executed as a process independent of the process of Fig. 11, or may be executed subsequent to the process of Fig. 11 for a speech section determined to be an important speech section in the process of Fig. 11.

[0065] In step S121, the CPU 11 acquires utterance section type data from the utterance section type DB 23 and acquires utterance type data from the utterance type DB 24.

[0066] In step S122, the CPU 11 determines whether the response scene of the utterance section is "response" from the utterance section type data and the utterance type data acquired in step S121. If it is determined that it is "response" (in the case of a positive determination), the process proceeds to step S123, and if it is determined that it is not "response" (in the case of a negative determination), the process proceeds to step S130.

[0067] In step S123, CPU 11 determines whether or not the utterance types include sales-related utterances (i.e., sales information). Sales-related utterances are utterances that have been assigned categories such as "needs hearing," "needs present," "no needs," and "proposals." If it is determined that there are sales-related utterances (sales information) (in the case of a positive determination), the process proceeds to step S124, and if it is determined that there are no sales-related utterances (sales information) (in the case of a negative determination), the process proceeds to step S130.

[0068] In step S124, the CPU 11 determines whether the sales-related utterance (sales information) is a "needs hearing." If it is determined to be a "needs hearing" (in the case of a positive determination), the process proceeds to step S125, and if it is determined not to be a "needs hearing" (in the case of a negative determination), the process proceeds to step S127.

[0069] In step S125, the CPU 11 determines whether the customer has a negative reaction ("no needs") after the "needs hearing." If it is determined that the customer has not a negative reaction ("no needs") (negative determination), the process proceeds to step S126, and if it is determined that the customer has a negative reaction ("no needs") (positive determination), the process proceeds to step S127.

[0070] In step S126, the CPU 11 determines that the utterance section is an important utterance section, and returns to step S106 in FIG.

[0071] On the other hand, in step S127, the CPU 11 determines whether the sales-related utterance (sales information) is a “proposal.” If it is determined to be a “proposal” (in the case of a positive determination), the process proceeds to step S128, and if it is determined to not be a “proposal” (in the case of a negative determination), the process proceeds to step S129.

[0072] In step S128, the CPU 11 determines whether or not a "question" or an "explanation" precedes the "proposal." If it is determined that a "question" or an "explanation" exists (in the case of a positive determination), the process proceeds to step S126, and if it is determined that a "question" or an "explanation" does not exist (in the case of a negative determination), the process proceeds to step S129.

[0073] In step S129, CPU 11 determines whether or not there is a certain number of "questions" in the speech section. If it is determined that there is a certain number of "questions" (in the case of a positive determination), the process proceeds to step S126, and if it is determined that there is not a certain number of "questions" (in the case of a negative determination), the process proceeds to step S130.

[0074] In step S130, the CPU 11 determines that the utterance section is not an important utterance section, and returns to step S106 in FIG.

[0075] Returning to step S106 in FIG. 10, the CPU 11 outputs the extraction result data obtained by extracting the important utterance section in step S105 to the extraction result DB 25, and ends the series of processes according to the utterance section extraction program.

[0076] As described above, according to this embodiment, when one speech section contains multiple utterances, important speech sections can be extracted with high accuracy by taking into consideration the combinations and transitions of each speech section type and utterance type.

[0077] In addition, by extracting and determining speech sections and using analysis information on a speech unit basis to determine the combinations and transitions between speech sections, it is possible to determine and extract important sections consisting of multiple speech sections based on the combinations and transitions of speech sections and the combinations and transitions of speech sections.

[0078] Furthermore, when determining important and good speech segments, important speech segments can be determined taking into consideration information obtained from individual speech segments.

[0079] Furthermore, by configuring the speech section type of the speech section to be independent of the analysis target or the target of use, and configuring the speech type for each utterance to be dependent on the analysis target or the target of use, or vice versa, the model can be replaced to suit the target to which it is applied.

[0080] In addition, it is possible to determine an effective sales method for sales calls.

[0081] [Second embodiment] Like the first embodiment, the speech section extraction device of the second embodiment provides specific improvements over conventional methods of extracting important speech sections without considering the speech sections and the respective combinations and transitions of the speech sections, and represents an advancement in the technical field of extracting important speech sections from speech data containing speech from two or more people.

[0082] In the first embodiment, a configuration in which one utterance section includes a plurality of utterances has been described, but in the second embodiment, a configuration in which one utterance section includes a single utterance will be described.

[0083] The speech segment extraction device according to the second embodiment (hereinafter referred to as speech segment extraction device 10A) has, as functional components, a sentence input unit 101, a speech segment identification unit 102A, a speech segment type determination unit 103A, an utterance type extraction unit 104A, a speech segment extraction unit 105A, and an output unit 106. Note that repeated explanations of the sentence input unit 101 and the output unit 106 will be omitted.

[0084] The configuration of each functional unit (utterance section identification unit 102A, utterance section type determination unit 103A, utterance type extraction unit 104A, and utterance section extraction unit 105A) according to the second embodiment will be specifically described with reference to FIGS. 13 to 16.

[0085] 13 acquires utterance text data from the utterance text DB 21, identifies an utterance section from the acquired utterance text data, and stores the obtained utterance section data in the utterance section DB 22. Specifically, when utterance text data is input, the utterance section identification unit 102 identifies a single utterance as an utterance section, and stores the obtained utterance section data in the utterance section DB 22.

[0086] The speech section type determination unit 103A shown in FIG. 14 acquires speech section data from the speech section DB 22, determines the speech section type of the acquired speech section data, and stores the obtained speech section type data in the speech section type DB 23. Specifically, when the speech section data is input, the speech section type determination unit 103A determines the speech section type of the speech section using the speech section type determination model 31, and stores the obtained speech section type data in the speech section type DB 23. As the speech section type, for example, labels representing basic dialogue acts in a conversation (e.g., "question," "explanation," "answer," "other," etc.) are defined. A model for determining these speech section types is generated in advance by machine learning using learning data with these labels. The speech section type determination unit 103A determines the speech section type of the input speech section using the speech section type determination model 31, and assigns the determination result of the speech section type to each speech section as an utterance section type ID.

[0087] The utterance type extraction unit 104A shown in FIG. 15 acquires utterance text data from the utterance text DB 21, extracts the utterance type of each utterance included in the acquired utterance text data, and stores the obtained utterance type data in the utterance type DB 24. Specifically, when utterance section data is input, the utterance type extraction unit 104A extracts the utterance type of each utterance using the utterance type extraction model 32, and stores the obtained utterance type data in the utterance type DB 24. As for the utterance type, for example, a type is assigned to each of the customer service scene, the utterance related to a dialogue act, and the utterance related to a sales activity. In the case of a customer service scene, for example, labels such as an utterance in a scene where the business is grasped and an utterance related to the business are defined. In the case of an utterance related to a dialogue act, for example, labels such as "question," "answer," and "explanation" are defined. In the case of an utterance related to a sales activity, for example, labels such as "needs hearing," "needs present," "no need," and "proposal" are defined. A model for extracting these utterance types is generated in advance by machine learning using speech text data in which these labels are attached to each utterance as training data. Using this utterance type extraction model 32, the utterance type of each utterance is determined for the input speech text, and the utterance type determination result is assigned to each utterance as an utterance type ID.

[0088] 16 acquires speech section type data from the speech section type DB 23 and acquires speech type data from the speech type DB 24. Then, the speech section extraction unit 105A extracts important speech sections from the speech sections identified by the speech section identification unit 102A based on the combinations and transitions of utterance section types determined by the speech section type determination unit 103A and the combinations and transitions of utterance types extracted by the speech type extraction unit 104A. Specifically, the speech section extraction unit 105A performs the extraction using the speech section extraction rules 33. In the speech section extraction rules 33, combinations and transitions of utterance section types and combinations and transitions of utterance types are predetermined in association with important utterance sections. The speech section extraction rule 33 extracts a group of multiple speech sections as an important speech section, for example, if the group does not include an utterance type ("no need") indicating an utterance expressing that the customer has no needs when the salesperson makes a sales pitch to the customer, and includes an utterance section type ("question") indicating an utterance section in which the operator asks the customer a question before an utterance type ("proposal") indicating an utterance in which the operator makes a suggestion to the customer. This makes it possible to extract important speech sections with high accuracy, even when one speech section includes a single utterance.

[0089] The speech segment extraction rules 33 include, for example, the following rule C. When rule C is satisfied, the segment is determined to be important.

[0090] (Rule C) -The "proposal" after the switch must include one or more "questions" and must not include "no needs."

[0091] Next, the important speech section extraction process according to the second embodiment will be specifically described with reference to FIGS.

[0092] Fig. 17 is a diagram illustrating the important utterance section extraction process according to the second embodiment. The utterance text shown in Fig. 17 includes an utterance by an operator and an utterance by a customer. The utterance text includes a plurality of utterance sections W11 to W14, and each of the utterance sections W11 to W14 includes a single utterance.

[0093] 17, the type of each of the utterance sections W11, W12, and W13 is "question," and the type of the utterance section W14 is "explanation / answer." Furthermore, these multiple utterance sections W11 to W14 are grouped together based on transition utterances, scene information, etc., and this group includes utterances of "suggestion" and "needs."

[0094] 17, as described above, the example does not include an utterance type ("no need") indicating an utterance expressing that the customer has no needs when the operator makes a sales pitch to the customer, and includes an utterance section type ("question") indicating an utterance section in which the operator asks a customer a question before an utterance type ("proposal") indicating an utterance in which the operator makes a suggestion to the customer. In other words, before the "proposal" after the utterance change, there is at least one "question" and there is no "no need," so the utterance is extracted as a group of important (good) utterance sections.

[0095] Fig. 18 is a diagram illustrating another important utterance section extraction process according to the second embodiment. The utterance text shown in Fig. 18 includes an utterance by an operator and an utterance by a customer, similar to the example in Fig. 17. The utterance text includes a plurality of utterance sections W21 and W22, and each of the utterance sections W21 and W22 includes a single utterance.

[0096] 18, the type of speech section W21 is "question," and the type of speech section W22 is "explanation / answer." These multiple speech sections W21 and W22 are grouped together based on transition utterances, scene information, etc., and this group includes the utterances "needs hearing," "no needs," "suggestion," and "no needs."

[0097] In the example of Figure 18, there is at least one "question" before the "proposal" after the change in speech, but since the "proposal" comes after "no need," it is extracted as a block of speech that is not important (good).

[0098] Next, the speech section extraction process according to the second embodiment will be described with reference to FIG.

[0099] FIG. 19 is a flowchart showing an example of the flow of the speech section extraction process according to the second embodiment, and shows an example of rule C of the speech section extraction rules 33.

[0100] In step S131, the CPU 11 acquires utterance section type data from the utterance section type DB 23 and acquires utterance type data from the utterance type DB 24.

[0101] In step S132, the CPU 11 determines whether the response scene of the utterance section is "response" from the utterance section type data and the utterance type data acquired in step S131. If it is determined that it is "response" (in the case of a positive determination), the process proceeds to step S133, and if it is determined that it is not "response" (in the case of a negative determination), the process proceeds to step S137.

[0102] In step S133, the CPU 11 determines whether or not the utterance types include sales-related utterances (i.e., sales information). Sales-related utterances are utterances that have been assigned categories such as "needs hearing," "needs present," "no needs," and "proposals." If it is determined that there are sales-related utterances (sales information) (in the case of a positive determination), the process proceeds to step S134, and if it is determined that there are no sales-related utterances (sales information) (in the case of a negative determination), the process proceeds to step S137.

[0103] In step S134, the CPU 11 determines whether or not the current state is within a switching interval. If it is determined that the current state is within a switching interval (if the determination is affirmative), the process proceeds to step S135, and if it is determined that the current state is not within a switching interval (if the determination is negative), the process proceeds to step S137.

[0104] In step S135, the CPU 11 determines whether or not there is an utterance that matches the rule within the switching section (for example, "There is a need" after "Proposal"). If it is determined that there is an utterance that matches the rule within the switching section (in the case of a positive determination), the process proceeds to step S136, and if it is determined that there is no utterance that matches the rule within the switching section (in the case of a positive determination), the process proceeds to step S137.

[0105] In step S136, the CPU 11 determines that the switching section is an important speech section, and returns to step S106 in FIG.

[0106] In step S137, the CPU 11 determines that the switching section is not an important speech section, and returns to step S106 in FIG.

[0107] As described above, according to this embodiment, when one speech section contains a single utterance, important speech sections can be extracted with high accuracy by taking into consideration the respective combinations and transitions of speech section types and utterance types.

[0108] [Third embodiment] The speech section extraction device of the third embodiment, like the first embodiment, provides specific improvements over conventional methods of extracting important speech sections without considering the speech sections and the respective combinations and transitions of speech, and represents an advancement in the technical field of extracting important speech sections from speech data containing speech from two or more people.

[0109] In the third embodiment, a check of compliance with a talk script will be described. A talk script is a dialogue script used when an operator responds to a customer in telephone sales, a contact center, etc.

[0110] The speech section extraction device according to the third embodiment (hereinafter referred to as speech section extraction device 10B) has, as functional components, a sentence input unit 101, a speech section identification unit 102, a speech section type determination unit 103B, an utterance type extraction unit 104B, an utterance section extraction unit 105B, and an output unit 106. Note that repeated explanations of the sentence input unit 101, the speech section identification unit 102, and the output unit 106 will be omitted.

[0111] The configuration of each functional unit (utterance section type determination unit 103B, utterance type extraction unit 104B, and utterance section extraction unit 105B) according to the third embodiment will be specifically described with reference to FIGS.

[0112] The speech section type determination unit 103B shown in FIG. 20 acquires speech section data from the speech section DB 22, determines the speech section type of the acquired speech section data, and stores the obtained speech section type data in the speech section type DB 23. Specifically, when the speech section data is input, the speech section type determination unit 103B determines the speech section type of the speech section using the speech section type determination model 31, and stores the obtained speech section type data in the speech section type DB 23. For example, the following labels (type 11 to type 13) are defined as the speech section types. A model for determining these speech section types is generated in advance by performing machine learning using learning data with these labels. The speech section type determination model 31 is used to determine the speech section type for the input speech section, and the determination result of the speech section type for each speech section is assigned as an utterance section type ID.

[0113] (Type 11) Section to confirm the customer's request (hereinafter referred to as "Request Content Confirmation Section") (Type 12) Section where the customer's environmental condition is confirmed (hereinafter referred to as "Customer Environmental Confirmation Section") (Type 13) Section corresponding to the request from the customer (hereinafter referred to as "request content response section")

[0114] The utterance type extraction unit 104B shown in FIG. 21 acquires utterance text data from the utterance text DB 21, extracts the utterance type of each utterance included in the acquired utterance text data, and stores the obtained utterance type data in the utterance type DB 24. Specifically, when utterance section data is input, the utterance type extraction unit 104B extracts the utterance type of each utterance using the utterance type extraction model 32, and stores the obtained utterance type data in the utterance type DB 24. As the utterance type, for example, a type is assigned to each of the customer service scene, the utterance related to a dialogue act, and the utterance related to a sales activity. In the case of a customer service scene, for example, labels such as an utterance in a scene where the business is grasped and an utterance related to the business are defined. In the case of an utterance related to a dialogue act, for example, labels such as "question," "answer," and "explanation" are defined. In the case of an utterance related to a sales activity, for example, labels such as "needs hearing," "needs present," "no need," and "proposal" are defined. A model for extracting these utterance types is generated in advance by machine learning using speech text data in which these labels are attached to each utterance as training data. Using this utterance type extraction model 32, the utterance type of each utterance is determined for the input speech text, and the utterance type determination result is assigned to each utterance as an utterance type ID.

[0115] 22 acquires speech section type data from the speech section type DB 23 and acquires speech type data from the speech type DB 24. Then, the speech section extraction unit 105B extracts important speech sections from the speech sections identified by the speech section identification unit 102, based on the combinations and transitions of speech section types determined by the speech section type determination unit 103B and the combinations and transitions of speech types extracted by the speech type extraction unit 104B. Specifically, the speech section extraction unit 105B extracts the important speech sections using the speech section extraction rules 33. In the speech section extraction rules 33, the combinations and transitions of speech section types and the combinations and transitions of utterance types are predetermined in association with important speech sections. The speech section extraction rule 33 extracts multiple speech sections as important speech sections, for example, when the multiple speech sections include a combination of speech section types indicating multiple consecutive speech sections that have been designated as important sections in advance, and each of the multiple consecutive speech sections includes a combination of speech types indicating the pre-designated utterances. This allows for accurate extraction of important speech sections, even when one speech section includes multiple utterances and is applied to checking compliance with a talk script.

[0116] The speech segment extraction rules 33 include, for example, the following rule D. When rule D is satisfied, the segment is determined to be important.

[0117] (Rule D) Each speech section includes necessary utterances in no particular order, and the speech sections transition in the specified order. For example, the speech section (type 11) confirming the purpose of the call (request content) includes utterances asking about the schedule and request content. Also, the speech section (type 12) confirming the customer environment includes utterances from the operator such as the network status and device information to be used, as well as utterances asking about requests. Also, the speech section (type 13) responding to request content includes utterances to reconfirm the requested content, the location, schedule, and contact information.

[0118] Next, the important speech section extraction process according to the third embodiment will be specifically described with reference to FIG.

[0119] Fig. 23 is a diagram illustrating the important utterance section extraction process according to the third embodiment. The spoken text (not shown) shown in Fig. 23 includes, for example, an utterance by an operator and an utterance by a customer. The spoken text includes a plurality of utterance sections W31 to W33, and each of the utterance sections W31 to W33 includes a plurality of utterances.

[0120] 23, the type of speech section W31 is a “request content confirmation section,” the type of speech section W32 is a “customer environment confirmation section,” and the type of speech section W33 is a “request content response section.” Furthermore, speech section W31 includes utterances of “schedule” and “request content,” speech section W32 includes utterances of “network,” “number of devices,” and “needs hearing,” and speech section W33 includes utterances of “request content,” “location,” “schedule,” and “contact information.”

[0121] 23, as described above, the important section includes a combination of speech section types indicating a plurality of consecutive speech sections designated in advance, and each of the consecutive speech sections includes a combination of speech types indicating a pre-designated speech. That is, the speech sections W31 to W33 are a transition of the speech section type combinations of "request content confirmation section" → "customer environment confirmation section" → "request content response section." The "request content confirmation section" includes a combination of speeches of "schedule" and "request content" (in no particular order), the "customer environment confirmation section" includes a combination of speeches of "network," "number of devices," and "needs interview" (in no particular order), and the "request content response section" includes a combination of speeches of "request content," "location," "schedule," and "contact information" (in no particular order). Therefore, these sections are extracted as important (good) speech sections.

[0122] Next, the speech section extraction process according to the third embodiment will be described with reference to FIG.

[0123] FIG. 24 is a flowchart showing an example of the flow of the speech section extraction process according to the third embodiment, and shows an example of rule D of the speech section extraction rules 33.

[0124] In step S141, the CPU 11 acquires utterance section type data from the utterance section type DB 23 and acquires utterance type data from the utterance type DB 24.

[0125] In step S142, the CPU 11 determines whether the specified utterance is included in the utterance section based on the utterance section type data and the utterance type data acquired in step S141. If it is determined that the specified utterance is included (in the case of a positive determination), the process proceeds to step S143, and if it is determined that the specified utterance is not included (in the case of a negative determination), the process proceeds to step S147.

[0126] In step S143, CPU 11 determines whether the utterance section type is a "request content confirmation section" and whether the utterance type includes "schedule" and "request content." If it is determined that the utterance section type is a "request content confirmation section" and the utterance type includes "schedule" and "request content" (in the case of a positive determination), the process proceeds to step S146, and if it is determined that the utterance section type is a "request content confirmation section" and the utterance type does not include "schedule" and "request content" (in the case of a negative determination), the process proceeds to step S144.

[0127] In step S144, CPU 11 determines whether the speech section type is the "customer environment confirmation section" and whether the speech includes designated keywords such as "network," "number of devices," etc. If it is determined that the speech section type is the "customer environment confirmation section" and the speech includes designated keywords (in the case of a positive determination), the process proceeds to step S146, and if it is determined that the speech section type is the "customer environment confirmation section" and the speech type does not include designated keywords (in the case of a negative determination), the process proceeds to step S145.

[0128] In step S145, the CPU 11 determines whether the utterance section type is a "request content corresponding section" and whether the utterance includes designated keywords such as "repeated confirmation," "location," "date," etc. If it is determined that the utterance section type is a "request content corresponding section" and the utterance includes designated keywords (in the case of a positive determination), the process proceeds to step S146, and if it is determined that the utterance section type is a "request content corresponding section" and the utterance type does not include designated keywords (in the case of a negative determination), the process proceeds to step S147.

[0129] In step S146, the CPU 11 determines that the utterance section is an important utterance section, and returns to step S106 in FIG. 10 described above.

[0130] In step S147, the CPU 11 determines that the utterance section is not an important utterance section, and returns to step S106 in FIG. 10 described above.

[0131] Thus, according to this embodiment, even when a single speech section contains multiple utterances and is applied to checking compliance with a talk script, important speech sections can be extracted with high accuracy by taking into account the respective combinations and transitions of speech section types and speech types.

[0132] The speech section extraction process executed by the CPU 11 after reading the speech section extraction program in the above embodiment may be executed by various processors other than the CPU 11. Examples of processors in this case include dedicated electrical circuits, such as programmable logic devices (PLDs) (such as field-programmable gate arrays (FPGAs)) whose circuit configuration can be changed after manufacture, and application-specific integrated circuits (ASICs) that are processors with circuit configurations specifically designed to execute specific processes. Furthermore, the speech section extraction process may be executed by one of these various processors, or by a combination of two or more processors of the same or different types (e.g., multiple FPGAs, or a combination of a CPU and an FPGA). Furthermore, the hardware structure of these various processors is, more specifically, an electrical circuit that combines circuit elements such as semiconductor devices.

[0133] In the above embodiment, the speech section extraction program is described as being stored (also referred to as "installed") in advance in the ROM 12 or the storage 14, but the present invention is not limited to this. The speech section extraction program may be provided in a form stored in a non-transitory storage medium such as a CD-ROM (Compact Disk Read Only Memory), a DVD-ROM (Digital Versatile Disk Read Only Memory), or a USB (Universal Serial Bus) memory. The speech section extraction program may also be downloaded from an external device via a network.

[0134] All publications, patent applications, and technical standards mentioned in this specification are herein incorporated by reference to the same extent as if each individual publication, patent application, or technical standard was specifically and individually indicated to be incorporated by reference.

[0135] The following additional notes are provided regarding the above-described embodiments.

[0136] (Additional note 1) Memory and at least one processor coupled to said memory; Including, The processor: Identifying a speech section including at least one utterance from speech text data including utterances by two or more people; determining a speech section type for each of the identified speech sections; extracting an utterance type for each utterance included in the utterance text data from the utterance text data; Extracting important speech sections from the identified speech sections based on the determined combinations and transitions of the speech section types and the extracted combinations and transitions of the speech types. The speech segment extraction device is configured as follows.

[0137] (Additional note 2) A non-transitory storage medium storing a program executable by a computer to execute a speech segment extraction process, The speech section extraction process includes: Identifying a speech section including at least one utterance from speech text data including utterances by two or more people; determining a speech section type for each of the identified speech sections; extracting an utterance type for each utterance included in the utterance text data from the utterance text data; Extracting important speech sections from the identified speech sections based on the determined combinations and transitions of the speech section types and the extracted combinations and transitions of the speech types. Non-transitory storage medium. [Explanation of symbols]

[0138] 10 Speech segment extraction device 11 CPU 12 ROM 13 RAM 14. Storage 15 Input section 16 Display 17 Communication I / F 18 Bus 20 Speech DB 21 Speech Text DB 22 Speech Segment DB 23 Speech segment type DB 24 Utterance type DB 25 Extraction result DB 30 Speech segment identification model 31 Speech segment type determination model 32 Utterance type extraction model 33 Speech segment extraction rules 101 Sentence input section 102, 102A Speech section identification unit 103, 103A, 103B Speech section type determination unit 104, 104A, 104B Speech type extraction unit 105, 105A, 105B Speech section extraction unit 106 Output section

Claims

1. a speech section identification unit that identifies a speech section including at least one utterance from speech text data including utterances by two or more people; a speech section type determination unit that determines a speech section type for each of the speech sections identified by the speech section identification unit; an utterance type extraction unit that extracts an utterance type for each utterance included in the utterance text data from the utterance text data; an utterance section extraction unit that extracts important utterance sections from the utterance sections identified by the utterance section identification unit by determining whether or not a combination and transition of utterance section types determined by the utterance section type determination unit and a combination and transition of utterance types extracted by the utterance type extraction unit match a speech section extraction rule that is a rule for determining whether or not an utterance section is an important utterance section that has been determined in advance; A speech segment extraction device comprising:

2. the spoken text data includes an operator's utterance and a customer's utterance; the speech section includes a plurality of speech sections each including a plurality of utterances, The speech section extraction rule extracts the plurality of speech sections as the important speech sections when the speech section extraction rule includes a speech type indicating a sales-related speech made by the operator to the customer and does not include a combination of speech section types indicating a plurality of consecutive speech sections designated in advance as a non-important section. The speech segment extraction device according to claim 1 .

3. The combination of a plurality of consecutive utterance sections designated in advance as the non-important section is a combination that includes at least one of the open-type business section and the end-type business section among the open-type business section, the theme-type business section, and the end-type business section, and does not include the theme-type business section. The speech segment extraction device according to claim 2 .

4. the spoken text data includes an operator's utterance and a customer's utterance; the speech section is a plurality of speech sections each including one utterance, The speech section extraction rule extracts a group including the plurality of utterance sections as the important utterance section when the utterance section extraction rule does not include an utterance type indicating an utterance expressing that the customer does not have a need when the operator makes a sales pitch to the customer, and includes an utterance section type indicating an utterance section in which the operator asks the customer a question before an utterance type indicating an utterance in which the operator makes a suggestion to the customer. The speech segment extraction device according to claim 1 .

5. the speech section includes a plurality of speech sections each including a plurality of utterances, The speech section extraction rule includes a combination of speech section types indicating a plurality of consecutive speech sections designated in advance as important sections, and when each of the consecutive speech sections includes a combination of speech types indicating a pre-designated utterance, the plurality of speech sections are extracted as the important speech sections. The speech segment extraction device according to claim 1 .

6. Identifying a speech section including at least one utterance from speech text data including utterances by two or more people; determining a speech section type for each of the identified speech sections; extracting an utterance type for each utterance included in the utterance text data from the utterance text data; extracting important speech sections from the identified speech sections by determining whether or not the determined combination and transition of speech section types and the extracted combination and transition of utterance types match a speech section extraction rule, which is a rule for determining whether or not an utterance section is an important utterance section that has been determined in advance; A computer-implemented speech segment extraction method.

7. Identifying a speech section including at least one utterance from speech text data including utterances by two or more people; determining a speech section type for each of the identified speech sections; extracting an utterance type for each utterance included in the utterance text data from the utterance text data; extracting important speech sections from the identified speech sections by determining whether or not the determined combination and transition of speech section types and the extracted combination and transition of utterance types match a speech section extraction rule, which is a rule for determining whether or not an utterance section is an important utterance section that has been determined in advance; A speech segment extraction program to be executed by a computer.

Citation Information

Patent Citations

  • Business section extracting method for contact center, device therefor, and program

    JP2011172163A

  • Talk script extraction device, method, and program

    JP2014106551A

  • Dialog log analyzer, dialog log analysis method, and program

    JP2018045639A

  • Major point extraction device, major point extraction method, and program

    WO2020036190A1