Information processing device, information processing method, and computer program

The summary generation system allows real-time summarization of conversations by determining multiple candidate segments and enabling users to select and adjust summary lengths, enhancing efficiency and information retention.

WO2025164110A1PCT designated stage Publication Date: 2025-08-07SONY GROUP CORP
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
PCT/JP2024/044269
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-01-30
Filing Date
2024-12-13
Publication Date
2025-08-07

AI Technical Summary

Technical Problem

Existing speech recognition technologies struggle to allow users to accurately summarize conversations in real-time while maintaining focus on the dialogue, particularly in contexts like contact centers where operators need to handle multiple tasks simultaneously.

Method used

A summary generation system that determines multiple candidate summary sections with different combinations of speech segments, allowing users to select and adjust the length of summarized intervals using a graphical user interface, and generates summaries based on user-selected segments.

Benefits of technology

Enables real-time summarization of conversations, allowing users to efficiently select and adjust summary lengths, improving data compression and retaining detailed information without distracting from other tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure JP2024044269_07082025_PF_FP_ABST
    Figure JP2024044269_07082025_PF_FP_ABST
Patent Text Reader

Abstract

Provided is an information processing device that executes processing for generating a summary sentence for each section selected by a user in voice recognition data. The information processing device comprises: an acquisition unit that acquires a plurality of pieces of voice section data comprising text that has been voice-recognized for each voice section; a determination unit that determines, from the plurality of pieces of voice section data, candidates for a plurality of summary sections in which combinations of voice section data are different; a reception unit that receives a summary section selected by a user from among the candidates for the plurality of summary sections; and a generation unit that generates a summary sentence of the summary section. The information processing device determines the summary section received by the reception unit and the summary sentence of the summary section.
Need to check novelty before this filing date? Find Prior Art

Description

Information processing device, information processing method, and computer program

[0001] The technology disclosed in this specification (hereinafter referred to as "the present disclosure") relates to an information processing device, an information processing method, and a computer program that perform processing related to voice recognition data.

[0002] Speech recognition technology that recognizes speech and converts it into text is widely used, and technology has also been developed that divides a speech period into multiple segments and summarizes the content of the speech for each segment. For example, a content summarization system (see Patent Document 1) that summarizes the text of a speech segment that a user indicates as important in a speech, and a summary creation device (see Patent Document 2) that displays selectable text converted by speech recognition of each of multiple people in a conversation and generates a summary consisting of text selected by the user from the displayed text have been proposed.

[0003] WO2008 / 050649 JP 2022-25665 A

[0004] Alec Radford, Jong Wook Kim, Tao Xu, Greg Brockman, Christine McLeavey, Ilya Sutskever "Robust Speech Recognition via Large-scale Weak Supervision"(arXiv:2212.04356v1 [eess.AS] 6 Dec 2022)GPT-4 Technical Report (arXiv:2303.08774v4 [cs.CL] 19 Dec 2023)

[0005] An object of the present disclosure is to provide an information processing device, an information processing method, and a computer program that perform processing to generate a summary sentence from speech recognition data.

[0006] The present disclosure has been made in consideration of the above-mentioned problems, and a first aspect thereof is an information processing device comprising: an acquisition unit that acquires a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination unit that determines a plurality of candidate summary segments having different combinations of speech segment data from the plurality of speech segment data; a reception unit that accepts a summary segment selected by a user from the plurality of candidate summary segments; and a generation unit that generates a summary sentence for the summary segment, wherein the summary segment and the summary sentence for the summary segment accepted by the reception unit are determined.

[0007] The information processing device according to a first aspect further includes a display control unit that causes information about the plurality of candidate summary segments to be selectably displayed in response to a user operation received by the receiving unit, the display control unit causes a summary sentence generated by the generating unit for any one of the plurality of candidate summary segments to be selectably displayed, and the receiving unit receives a selection of the corresponding candidate summary segment based on a predetermined user operation on the displayed summary sentence.

[0008] The display control unit may display, near the displayed summary, whether or not there are at least one of shorter and longer summary segment candidates. The display control unit may also display, near the displayed summary, whether or not there are new summary segment candidates currently being recognized or summarized.

[0009] A second aspect of the present disclosure is an information processing method including: an acquisition step of acquiring a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination step of determining a plurality of candidate summary segments having different combinations of speech segment data from the plurality of speech segment data; a reception step of accepting a summary segment selected by a user from the plurality of candidate summary segments; and a generation step of generating a summary sentence for the summary segment, wherein the summary segment and the summary sentence for the summary segment accepted in the reception step are determined.

[0010] Furthermore, a third aspect of the present disclosure is a computer program written in a computer-readable format so that a computer functions as: an acquisition unit that acquires a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination unit that determines a plurality of summary segment candidates each having a different combination of speech segment data from the plurality of speech segment data; a reception unit that accepts a summary segment selected by a user from the plurality of summary segment candidates; and a generation unit that generates a summary sentence for the summary segment; and the computer program determines the summary segment and the summary sentence for the summary segment accepted by the reception unit.

[0011] A computer program according to a third aspect of the present disclosure defines a computer program written in a computer-readable format to perform predetermined processing on a computer. The computer program can be provided to a computer capable of executing various program codes in a computer-readable format via a storage medium or communication medium, such as an optical disk, a magnetic disk, or a semiconductor memory, or a communication medium such as a network. By installing the computer program according to the third aspect of the present disclosure on a computer via any of these media, a cooperative effect is exerted on the computer, and the same effects as those of the information processing device according to the first aspect of the present disclosure can be obtained.

[0012] According to the present disclosure, it is possible to provide an information processing device, an information processing method, and a computer program that perform processing to generate a summary sentence for each section selected by a user from speech recognition data.

[0013] It should be noted that the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited to these. Furthermore, the present disclosure may also bring about additional effects in addition to the effects described above.

[0014] Further objects, features, and advantages of the present disclosure will become apparent from the following detailed description based on the embodiments and accompanying drawings.

[0015] FIG. 1 is a diagram showing an example of the basic configuration of a summary generation system 100. FIG. 2 is a diagram showing an example of a dialogue to be summarized. FIG. 3 is a diagram showing examples of speech section data extracted and speech-recognized by the summary generation system 100, and examples of summary section candidates sequentially generated. FIG. 4 is a diagram showing an example of a screen configuration for displaying a summary of text data and intermediate summarization results. FIG. 5 is a diagram showing an example of operation of the summary generation system 100 to display intermediate summarization results. FIG. 6 is a diagram showing timing settings for displaying intermediate summarization progress. FIG. 7 is a diagram showing an example of displaying a summary and intermediate summarization results. FIG. 8 is a diagram showing an example of displaying a summary and intermediate summarization results. FIG. 9 is a diagram showing an example of displaying a summary and intermediate summarization results. FIG. 10 is a diagram showing an example of displaying a summary and intermediate summarization results. FIG. 11 is a diagram showing an example of displaying a summary and intermediate summarization results. FIG. 12 is a diagram showing a detailed configuration example of the summary generation system 100. FIG. 13 is a flowchart showing an example of operation of the summary generation system 1200. FIG. 14 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 15 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 16 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 17 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 18 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 19 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 20 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 21 is a diagram showing the operation and screen transitions of the summary generation system 100 in chronological order. FIG. 22 is a diagram showing an example of the hardware configuration of an information processing device 2000.

[0016] Hereinafter, embodiments of the present disclosure will be described in the following order with reference to the drawings.

[0017] A. Overview B. Basic configuration C. Example of operation C-1. Example of basic operation C-2. Automatic selection of summary sections C-3. Method of using summary sentences selected through comparative analysis C-4. Function for automatically determining summary sentences D. Specific configuration E. Example of configuration of information processing device

[0018] A. Overview After a conversation, such as a meeting or telephone conversation, ends, it is relatively easy for a user to operate a UI (User Interface) to select a section to summarize while viewing the text data generated by speech recognition on a screen. However, once the conversation has already concluded, it can be difficult to rely on memory to identify the information that the user wanted to retain (in the summary). Furthermore, when summarizing at the end, the only way to confirm the flow of the conversation up to that point is to review the entire speech recognition results, which is redundant because it includes all of the speech.

[0019] On the other hand, if it were possible to sequentially specify summarization sections in real time during an utterance, there would be no need to rely on memory to confirm information that the user wants to keep, but there is a risk that the user would become distracted by specifying summarization sections and be unable to keep up with the utterance.For example, a contact center operator must sequentially specify summarization sections while dealing with customers at the counter or over the phone, which raises concerns that they may not be able to specify summarization sections accurately, or that they may be distracted by specifying summarization sections and neglect to deal with the customer.

[0020] Therefore, the present disclosure proposes a summary generation technology that allows a user to control in real time the information that they want to keep in the summary while checking the flow of the dialogue up to that point during the speaking period.

[0021] The summary generation technology disclosed herein determines multiple candidate summary sections with different combinations of speech section data from multiple speech section data consisting of text obtained by speech recognition for each speech section, and generates a summary based on the speech section data included in a summary section selected by a user from among these multiple candidate summary sections.

[0022] Furthermore, the summary generation technology according to the present disclosure performs display control using a GUI (Graphical User Interface) screen or the like to display multiple selectable summary section candidates. The multiple summary section candidates presented on the screen differ in the number of speech sections to be summarized (in other words, the lengths of the speech sections to be summarized). Therefore, selecting multiple summary section candidates on the screen corresponds to adjusting the sections to be summarized, such as lengthening or shortening them.

[0023] The present disclosure can be applied to a dialogue between two or more speakers, such as a conference or telephone conversation, and performs real-time speech recognition on the content of the dialogue and then summarizes the speech recognition results in real time. In this process, the speech interval to be summarized can be lengthened or shortened to present multiple candidate summary intervals. A user who is one of the speakers can switch the interval length of the speech interval to be summarized, check the summary sentences generated from each of the multiple candidate summary intervals, and select an appropriate summary interval and summary sentence.

[0024] B. Basic Configuration FIG. 1 shows a schematic diagram of a basic configuration example of a summary generation system 100 to which the present disclosure is applied, which generates a summary of a person's utterance. The summary generation system 100 processes a dialogue between two or more speakers, but of course, it may also process the utterance of a single speaker. The summary generation system 100 can be configured using a general information processing device such as a personal computer, and can be equipped with additional devices such as a microphone or a reception unit (described below) as needed. The summary generation system 100 may be configured using a single information processing device, or may be configured by linking two or more information processing devices. Furthermore, at least some of the functions of the summary generation system 100 may be located on the cloud.

[0025] The summary generation system 100 shown in Fig. 1 includes a speech input unit 101, a speech segment detection unit 102, a speech recognition unit 103, a recognition result memory 104, a summary segment candidate determination unit 105, a summary generation unit 106, a display control unit 107, a display unit 108, a reception unit 109, and an input control unit 110. Each unit will be described below.

[0026] The voice input unit 101 inputs an acoustic signal containing the voice uttered by each speaker. The voice input unit 101 is, for example, a microphone installed for each speaker, a microphone built into a terminal, a smartphone, or another telephone. When multiple speakers are gathered in one conference room, the voice input unit 101 may be a single microphone. The speakers are, for example, a combination of a contact center operator and an inquirer who has made an inquiry to the contact center via telephone or the web.

[0027] The voice activity detection unit 102 extracts voice activity segments that appear to be voice (human voice) from the audio signal captured by the audio input unit 101. The voice activity detection unit 102 divides the audio signal into voice activity segments based on audio breaks, speaker changes, etc. In addition to determining whether or not a signal is a human voice, the voice activity detection unit 102 may also separate the signal into time segments and extract the voice activity segments.

[0028] The speech recognition unit 103 performs speech recognition processing for each speech segment detected by the speech segment detection unit 102 and converts the resulting data into text. In this specification, text data converted from the speech of a speech segment is also referred to as "speech segment data." The speech recognition unit 103 performs speech recognition processing using a trained model. For example, a speech recognition model such as Whisper (see Non-Patent Document 1) can be applied to the speech recognition unit 103.

[0029] Each speech segment data generated by the speech recognition unit 103 is temporarily stored in the recognition result memory 104. "Temporarily" means until a summary segment is determined by subsequent processing. The speech segment data after the summary segment is determined is not used for generating a summary sentence, and is therefore discarded from the recognition result memory 104. Of course, if there are no limitations on memory capacity, the speech segment data after the summary segment is determined may continue to be held in the recognition result memory 104. The recognition result memory 104 is configured, for example, by a main storage device.

[0030] The summary section candidate determination unit 105 determines multiple summary section candidates each having a different combination of speech section data from multiple speech section data sequentially output from the speech recognition unit 103 (or temporarily stored in the recognition result memory 104). The speech section detection unit 102 and the speech recognition unit 103 acquire speech section data in time series from the input acoustic signal. The summary section candidate determination unit 105 sets the first unprocessed speech section data as the summarization start point, and combines one, two, three, ..., N pieces of speech section data from that summarization start point to determine a summary section candidate consisting of the first speech section data and multiple summary section candidates each consisting of two or more consecutive speech section data including the first speech section data. Thus, the summary section candidate determination unit 105 determines multiple summary section candidates each having a different length for the section to be summarized when speech recognition processing is performed on the input acoustic signal.

[0031] The summary generation unit 106 generates a summary using one or more pieces of speech section data included in each of a plurality of summary section candidates with different speech section lengths determined by the summary section candidate determination unit 105. As described above, when the summary section candidate determination unit 105 determines a plurality of summary section candidates by combining 1, 2, 3, ..., N, ... pieces from the summarization start point, the summary generation unit 106 generates each summary using a summary section candidate consisting of the first piece of speech section data and a plurality of summary section candidates each combining two or more consecutive pieces of speech section data.

[0032] The summary generation unit 106 uses a trained model (hereinafter also referred to as a "summary generation model") to generate a summary of the text data of each candidate summary section. A large-scale language model (LLM) such as GPT-4 (see Non-Patent Document 2) may be used to generate the summary. When generating sentences probabilistically using a trained model, multiple summary candidates can be generated from a single piece of text data.

[0033] The display control unit 107 controls the display unit 108 to display the summary generated by the summary generation unit 106 from the text data of each candidate summary section. The summary generation unit 106 generates a summary from each of the multiple candidate summary sections. The display control unit 107 causes the display unit 108 to display one of the multiple summaries generated by the summary generation unit 106 in accordance with a display switching instruction received by the receiving unit 109.

[0034] In this embodiment, the display control unit 107 displays the summary results of the candidate summary intervals and the intermediate results of the summarization using only one line on the screen of the display unit 108 (in other words, the summaries of two or more summary intervals are not displayed simultaneously), and sequentially switches the target summary interval in response to a simple user operation received by the reception unit 109. This is because, for example, in a use case in which a contact center operator specifies a summary interval while simultaneously dealing with a customer, the operator does not have the time to visually check all of the multiple summaries displayed in the list, and there is a risk that an optical illusion could cause the operator to mistake one summary interval for one generated from another summary interval. Note that the display control unit 107 may also display other reference information on the display unit 109 in addition to the display of the summary of one summary interval. Examples of reference information include the original text data of the summary interval whose summary is currently being displayed, the current full summary result, and the like.

[0035] The display unit 108 is a device that presents visual information to the user using a screen, and mainly displays the processing results (including intermediate processing results) of the summary generation system 100 on the screen. The reception unit 109 is a device that receives instructions from the user. The user here refers to any speaker who is the target of inputting utterances into the voice input unit 101, and is assumed to be, for example, an operator who handles customers at a contact center. Therefore, it is preferable that the screen configuration of the display unit 108 and the input operation of the reception unit 109 be simple so that the operator does not make any mistakes while handling customers.

[0036] The display unit 108 is, for example, a liquid crystal display (LCD) or an organic electroluminescence (EL) display. The reception unit 109 is composed of a typical computer input device such as a keyboard. This embodiment assumes a use case in which, for example, a contact center operator specifies a summary section while assisting a customer. The system allows instructions such as switching between summary sections, selecting a summary section, and discarding a summary section to be performed simply by simple key input operations using special keys such as the Tab key, Esc key, and Ctrl key. The input control unit 110 interprets user operations (such as key inputs) received by the reception unit 109 and outputs the interpretation results to related functional modules such as the summary section candidate determination unit 105 and the display control unit 107.

[0037] The summary generation system 100 outputs a summary for each summary section that is sequentially determined in accordance with user operations. The summary generation system 100 may store each output summary in a local mass storage device (not shown), or may archive the output summary in a storage location (not shown) operated by, for example, a contact center.

[0038] C. Operational Examples C-1. Basic Operational Examples Next, a description will be given of the operation of the summary generation system 100. In the following, a conversation between an operator and a customer at a contact center will be taken as an example.

[0039] First, as shown in Fig. 2, let us assume that acoustic signals containing inquiries from a customer, such as "My PC has been malfunctioning since yesterday," "I'm having trouble," and "What should I do?", and responses from an operator, such as "We apologize for the inconvenience," and "Please tell us your specific symptoms," are input sequentially to the speech input unit 101. Fig. 3 shows an example of speech section data extracted and speech-recognized, and summary section candidates that are sequentially generated when the dialogue shown in Fig. 2 is input to the summary generation system 100.

[0040] An audio signal containing continuous speech from each speaker is input to the audio input unit 101. The audio activity detection unit 102 extracts audio activity segments that appear to be audio (human voices) from the audio signal captured by the audio input unit 101. The audio activity detection unit 102 divides the audio signal into audio activity segments based on audio segments, speaker changes, etc. The audio recognition unit 103 then performs audio recognition processing for each audio activity segment detected by the audio activity detection unit 102 and converts it into text.

[0041] In the example dialogue shown in FIG. 2 , an acoustic signal including one continuous speech "My PC has been malfunctioning since yesterday. I'm in trouble. What should I do? I'm sorry for the inconvenience. Please tell me the specific symptoms." is input to the speech input unit 101. The speech segment detection unit 102 sequentially detects speech segments corresponding to each of the utterances "Since yesterday," "My PC has been malfunctioning," "I'm in trouble," "What should I do?", "I'm sorry for the inconvenience," and "Please tell me the specific symptoms." The speech recognition unit 103 then sequentially generates speech segment data 1 to 6 consisting of the speech recognition results for each speech segment (see FIG. 3A). The recognition result memory 104 temporarily stores the speech recognition results (speech segment data) for each speech segment, as shown in Table 1 below.

[0042]

[0043] The summary segment candidate determination unit 105 can acquire the speech segment data 1 to 6 generated by the speech recognition unit 103 via the recognition result memory 104. Alternatively, the summary segment candidate determination unit 105 may temporarily store the speech segment data sequentially output from the speech recognition unit 103 in its own local memory.

[0044] The summary segment candidate determination unit 105 then determines multiple summary segment candidates each having a different combination of speech segment data from multiple time-series speech segment data as shown in Fig. 3(A). Specifically, as shown in Fig. 3(B), the summary segment candidate determination unit 105 sets speech segment data 1 (i.e., the first unprocessed speech segment data) as the summarization start point, and generates summary segment candidate 1 consisting of one piece of speech segment data at the summarization start point, and summary segment candidates 2 to 6 each consisting of two or more consecutive pieces of speech segment data including the first speech segment data, each time new speech segment data is acquired. Specifically, when speech segment data 1 (the speech recognition result of speech segment 1) is input, the summary segment candidate determination unit 105 determines the speech segment candidate as summary segment candidate 1. When speech segment data 2 is subsequently input, the summary segment candidate determination unit 105 determines the speech segment candidate as summary segment candidate 2 by concatenating the speech segment data 2 with the speech segment data 1. Thereafter, each time subsequent speech segment data is input, the summary segment candidate determination unit 105 determines a new summary segment candidate by concatenating the end of the previous summary segment candidate. The summary generation unit 106 generates summaries for each of the summary segment candidates determined by the summary segment candidate determination unit 105. Next, the speech segment data constituting each of the summary segment candidates 1 to 6 at the time when summary segment candidate 6 was generated is summarized in Table 2 below.

[0045]

[0046] The summary generation unit 106 sequentially generates summaries from the speech segment data (speech recognition text) corresponding to each of the summary segment candidates 1 to 6 determined by the summary segment candidate determination unit 105. Note that the summary generation unit 106 generates sentences probabilistically using a trained model (e.g., a large-scale language model), and therefore can generate multiple summary candidates from a single piece of text data.

[0047] The display control unit 107 displays the summary generated by the summary generation unit 106 from the text data corresponding to each summary section candidate, or the intermediate results of the summarization, on the display unit 108. Details of the method for displaying the generated summary and the intermediate results of the summarization will be described later.

[0048] The user can issue instructions via the receiving unit 109 to switch between summary intervals, select a summary interval, discard a summary interval, etc. For example, when a keyboard is used as the receiving unit 109, the user can confirm a summary interval candidate as a summary interval and adopt the summary sentence of that summary interval by pressing the Tab key after selecting the summary interval candidate. Also, the user can issue instructions to discard the summary up to that summary interval candidate (determining that the information up to that interval is unnecessary) by pressing the Esc key after selecting the summary interval candidate.

[0049] 3B, when the Tab key is pressed with summary section candidate 3 selected, summary section candidate 3 is confirmed as the summary section, and the text data "My PC has been malfunctioning since yesterday. I'm having trouble." corresponding to the corresponding speech section data 1 to 3 can be used as the target for summarization. On the other hand, when the Esc key is pressed with summary section candidate 3 selected, the summaries up to summary section candidate 3 are discarded.

[0050] If either the Tab or Esc key is pressed while summary segment candidate 3 is selected, summary segment candidate 3, i.e., the segment from speech segment data 1 to 3, has been processed or no longer requires processing. Therefore, the summary segment candidate determination unit 105 temporarily discards summary segment candidates 4 to 6, which include processed speech segment data 1 to 3, and, as shown in FIG. 3C , sets the unprocessed first speech segment data 4 as a new summarization start point. Then, as shown in FIG. 3C , the summary segment candidate determination unit 105 generates new summary segment candidate 7, which consists only of speech segment data 4, and summary segment candidate 8 and summary segment candidate 9, which combine two and three consecutive speech segment data pieces including speech segment data 4, respectively. The summary generation unit 106 then generates summaries for each summary segment candidate. The speech segment data constituting each of summary segment candidates 7 to 9 at the time when they are newly generated is summarized in Table 3 below.

[0051]

[0052] Furthermore, speech segment data 1 to 3 that have been processed or no longer require processing are discarded from the recognition result memory 104, as shown in Table 4 below. In Table 4, for ease of understanding, the records of speech segment data 1 to 3 that are to be discarded are shown shaded in gray, but in reality, these records are discarded from the memory area of ​​the recognition result memory 104.

[0053]

[0054] Although not shown in the figure, each time new speech section data is acquired after speech section data 6, the summary section candidate determination unit 105 sequentially generates summary section candidates by combining three or more consecutive speech section data including speech section data 4.

[0055] Depending on the results of speech recognition, different summary segments may produce the same summary. In the example shown in Figure 3, the text data "We apologize for the inconvenience" spoken by the operator in speech segment data 5 does not contain specific information, so it is expected that the same summary will be generated from summary segment candidate 4 and summary segment candidate 5. In cases where different summary segments produce the same summary, rather than repeatedly outputting the same summary in each of the consecutive summary segment candidates 4 and 5, summary segment candidate 4 and summary segment candidate 5 may be combined into a single summary segment candidate (for example, summary segment candidates 4 and 5 in Figure 3 may be excluded, and speech segment data 1 to 5 may be newly combined into "summary segment candidate 4"). Furthermore, speech segments containing only meaningless speech, such as fillers, are expected to produce an empty summary result, so they may be excluded from the summary segment candidates.

[0056] 4 shows an example of a screen configuration in which the display control unit 107 displays on the display unit 108 a summary or an intermediate result of the summary generated by the summary generation unit 106 from text data corresponding to each summary segment candidate. The screen shown in FIG. 4 is configured as a window 400 assigned to an application (for example, a "summary generation application (tentative name)" that summarizes text resulting from speech recognition). This window 400 includes a summary candidate display field indicated by the reference numeral 410. Within the summary candidate display field 410, the summary result of the summary segment candidate or the intermediate result of the summary is displayed in one line.

[0057] Figure 4 shows an example display of summary results and intermediate summary results for summary segment candidates when the dialogue shown in Figures 2 and 3 is input. More specifically, this is an example display when speech recognition processing and summary generation processing have been completed up to summary segment candidate 3, and speech recognition or summary generation is currently being processed for summary segment candidate 4 and beyond (or speech segment data 4 and beyond), and summary segment candidate 3 has been selected. As indicated by reference numeral 411, the summary candidate display field 410 displays the text "My PC is not working properly," which is the summary result of selected summary segment candidate 3.

[0058] Each time new subsequent speech segment data is input, the summary segment candidate determination unit 105 determines a new summary segment candidate by concatenating it to the end of the immediately preceding summary segment candidate. Any method can be used to select a segment to be selected from the multiple determined summary segment candidates. For example, the segment to be selected may be determined based on user input. Alternatively, the summary segment candidate determined last (i.e., the most recent) by the summary segment candidate determination unit 105 may be selected by default. Figure 4 illustrates an example of how the summary segment candidate determination unit 105 determines a new summary segment candidate 3 by concatenating speech segment data 3 to the end of the immediately preceding summary segment candidate 2 in response to input of speech segment data 3, and how the summary segment candidate 3 is selected and the summary generated for that segment is displayed in the summary candidate display field 410.

[0059] Furthermore, in the summary candidate display field 410, an ellipsis "..." is placed to the left of the text "My PC is not working properly," as indicated by reference numeral 412, thereby indicating the presence of a summary segment candidate consisting of speech segment data shorter than the selected summary segment candidate 3. Furthermore, an ellipsis "..." is placed to the right of the text "My PC is not working properly," as indicated by reference numeral 413, thereby indicating the presence of a summary segment candidate consisting of speech segment data longer than the selected summary segment candidate 3. Furthermore, at the right end of the summary candidate display field 410, as indicated by reference numeral 414, a character string "[Summarizing...]" is placed, indicating that speech recognition processing or summarization processing of the input acoustic signal is ongoing. By displaying the character string "[Summarizing...]" 414, it is possible to indicate the possibility that new speech segment data will be added in the future, increasing the number of summary segment candidates.

[0060] Generally, by increasing the range of speech segments to be summarized, the summarization rate can be increased and information can be compressed, while by shortening the range of speech segments to be summarized, the summarization rate can be decreased and more detailed information can be retained in the summary. In this embodiment, while dealing with a customer, a contact center operator looks at the summary candidate display field 410 and compares and considers the multiple summaries created for each summary segment, while confirming whether the summary segment 411 of the selected summary segment candidate is accurate. Note that the "summarization rate" is a value that represents the degree of data compression achieved by summarizing text data, and is defined here as the length of the summary segment generated from the summary segment relative to the length of the original speech-recognized text of the summary segment to be summarized.

[0061] For example, if the operator feels that important information is missing from the summary sentence 411 displayed in the summary candidate display field 410, the operator can use the BackSpace key or the like to return to a shorter summary segment candidate by lowering the summarization rate and readjusting the summary sentence so that more detailed information remains. Furthermore, the operator can understand from the display of the ellipsis symbol "..." 413 on the right that there is a summary segment candidate consisting of speech segment data longer than the currently selected summary segment candidate (in other words, the summarization rate can be improved by extending the summary segment). Therefore, if the operator feels that the summary sentence 411 displayed in the summary candidate display field 410 contains important information and that the summarization rate can be further improved, the operator can wait until the summary sentence generated for the next (longer) summary segment candidate is displayed, thereby improving the summarization rate. Furthermore, the operator can understand from the display of the character string "[Summarizing...]" 414 that speech recognition processing or summarization processing is ongoing, and can wait until the summary sentence generated for the next (longer) summary segment candidate is displayed.

[0062] When the processing time has elapsed and the summarization of the subsequent summary section candidate that was being summarized is completed, the summary candidate display field 410 switches to display the summary sentence of the next summary section candidate. As with the above, when the display switches to the next summary sentence, an ellipsis symbol "..." is displayed on the left side, indicating that there is a summary section candidate consisting of short speech section data, and if more detailed information is desired, the user can use the BackSpace key or the like to go back to the summary sentence of the previous summary section candidate.

[0063] While dealing with a customer, a contact center operator constantly monitors the summary candidate display field 410, compares and considers the multiple summaries created successively for each summary interval, and when it is determined that an appropriate summary has appeared, the operator can adopt the summary for the relevant summary interval by, for example, pressing the Tab key.Furthermore, when it is determined that the summary for the interval from the start of the summary up to that point is unnecessary, the contact center operator can press, for example, the ESC key to discard the summary up to that summary interval candidate (the information up to that interval is unnecessary).

[0064] Additionally, in the space below the window 400 shown in Fig. 4, the speech recognition result indicated by reference numeral 421 and the summarization result indicated by reference numeral 422 are displayed together as reference information for the summary generation process. These speech recognition results and summarization results are displayed in pairs with the start time (Start) and end time (End) of the speech section or the summarization section, respectively. The operator may select the generated summaries by referring to the speech recognition result 421 and the summarization result 422 as well as the display contents of the summary candidate display field 410. Note that the display of the speech recognition result and the summarization result as shown in Fig. 4 is not essential to realizing the present disclosure.

[0065] In other words, even while a contact center operator is busy dealing with a customer, they can manually select and select appropriate summaries one after another based on only one line of information displayed in the summary candidate display field 410. In the background (i.e., inside the summary generation system 100), a summary segment candidate is added and a summary is generated for each new speech segment data generated. The operator does not need to know which summary segment candidate the summary displayed in the summary candidate display field 410 was generated from (e.g., without grasping the overall processing status within the summary generation system 100 as shown in FIG. 3 ), but can adjust the length of the summary segment and sequentially obtain appropriate summaries simply by operating the special keys for the currently displayed summary on the screen as shown in FIG. 4 .

[0066] Note that the summary generation unit 106 generates sentences probabilistically using a trained model, and therefore may generate multiple summary candidates from a single piece of text data. In such cases, the summary candidate display field 410 may provide a UI that enables the user to compare and consider multiple summary candidates generated from the same summary section.

[0067] In the display example shown in Fig. 4, as described above, the display control unit 107 uses the summary candidate display field 410 to display the details of the ongoing speech recognition and summary processing, so that it is clear that these processes are in progress. However, because the speech recognition and summary generation processes are performed on speech segments for which utterance has been completed, it is necessary to wait until the results are obtained during speech (i.e., speech segments for which utterance has not yet been completed). Therefore, as an improved example, supplemental information consisting of intermediate speech recognition results and provisional summary results based on the intermediate results may be sequentially displayed separately from the summary candidates.

[0068] An example of the operation of the summary generation system 100 for sequentially displaying intermediate results of speech recognition and summary generation will be described in detail below, along with another example dialogue shown in Figure 5. Figure 5 shows an example dialogue including a customer utterance, "My PC is malfunctioning and I'm having trouble," an operator utterance, "What are the specific symptoms?", and a customer utterance, "My battery is malfunctioning and runs out quickly."

[0069] In this case, the speech input unit 101 receives successive acoustic signals containing three speech utterances: "My PC is malfunctioning and I'm having trouble.", "What are the specific symptoms?", and "My battery is malfunctioning and it runs out quickly." The resulting text data, "My PC is malfunctioning and I'm having trouble. What are the specific symptoms? My battery is malfunctioning and it runs out quickly.", is used as the target for summarization. As in the above, the summary segment candidate determination unit 105 determines multiple summary segment candidates with different combinations of speech segment data, and the summary generation unit 106 sequentially generates summaries from the text data corresponding to each summary segment candidate. The display control unit 107 then displays the summaries generated by the summary generation unit 106 from the text data corresponding to each summary segment candidate and the intermediate results of the summaries on the display unit 108.

[0070] As shown on the time axis in Fig. 6, five timings T1 to T5 are set for displaying the intermediate results of speech recognition and summary sentences in the example dialogue shown in Fig. 5. The definitions of the timings T1 to T5 are as follows.

[0071] T1: The timing when the second utterance ends. T2: In the first half of the third utterance, there are no provisional speech recognition results. T3: In the middle of the third utterance, provisional speech recognition and summary results are output. T4: In the second half of the third utterance, the provisional speech recognition and summary results are updated. T5: The third utterance is completed, and the summary sentence of the longer summary section candidate is completed.

[0072] 7 to 11 show examples of displaying speech recognition result sentences and intermediate results of summarization using the summary candidate display field 410 (see FIG. 4) at each of the timings T1 to T5 shown in FIG. 6, respectively.

[0073] 7 shows an example of the display in the summary candidate display field 410 at timing T1. Timing T1 marks the state in which speech recognition and summary generation have been completed up to the second utterance. Therefore, the generated summary text "My PC is not working" is displayed. Furthermore, placing an ellipsis "..." to the left of the text "My PC is not working" indicates the existence of a shorter summary section candidate, and placing an ellipsis "..." to the right of the text "My PC is not working" indicates the existence of an even longer summary section candidate.

[0074] FIG. 8 shows an example of the display of the summary candidate display field 410 at timing T2. At timing T2, a third newly input speech is received, and speech recognition and summarization are in progress. The text of the already completed summary, "My PC is malfunctioning," is displayed, and an ellipsis "..." indicating the presence of a shorter summary sentence candidate and an ellipsis "..." indicating the presence of an even longer summary segment candidate are placed on either side of the text, similar to FIG. 7 . An animated icon, indicated by reference numeral 801, is placed at the end of the line to indicate that speech recognition and summarization are in progress. Note that, in the example shown in FIG. 8 , the animated icon 801 is a rotating icon, but other animated icons (e.g., an hourglass icon) may be used as long as they clearly indicate that processing is in progress.

[0075] FIG. 9 shows an example of the display of the summary candidate display field 410 at timing T3. At timing T3, speech recognition of the third utterance progresses, and a provisional summary result is displayed based on the provisional speech recognition results. Similar to FIG. 8 , an ellipsis symbol "..." is placed on the left and right of the already completed summary sentence text "The PC is malfunctioning" and the text "The PC is malfunctioning," and an animated icon 801 is placed to the right of the right-hand ellipsis symbol "...." At timing T3, as indicated by reference numeral 901, a provisional summary result "Battery is malfunctioning" based on the provisional speech recognition results is displayed at the end of the line. The summary result "Battery is malfunctioning" is displayed in italics to indicate that it is provisional (i.e., unconfirmed) and may change when confirmed, i.e., to distinguish it from a confirmed summary result.

[0076] FIG. 10 shows an example of the display of the summary candidate display field 410 at timing T4. At timing T4, speech recognition of the third utterance progresses further, and the provisional summary result changes as the speech recognition process progresses. Similar to FIGS. 8 and 9 , an ellipsis (...) is placed on the left and right of the already completed summary sentence text "The PC is malfunctioning" and the text "The PC is malfunctioning," and an animated icon 801 is placed to the right of the right-hand ellipsis (...). At timing T4, as the speech recognition process progresses and the provisional summary result based on the speech recognition results changes, the provisional summary result displayed at the end of the line also changes from "The battery is malfunctioning" to "The battery runs out quickly," as indicated by reference numeral 1001. The summary result "The battery runs out quickly" is displayed in italics to distinguish it from the final summary result because it is provisional (i.e., unconfirmed).

[0077] 11 shows an example of the display in the summary candidate display field 410 at timing T5. At timing T5, the third utterance is completed, completing the summary sentence for the longer summary section candidate. The text of the already completed summary sentence, "My PC is not working properly," and the ellipsis "..." are placed on either side of the text "My PC is not working properly," as in FIGS. 8 to 10. Because the utterance is completed at timing T5, the animated icon indicating processing and the display of the provisional summary result disappear.

[0078] As shown in Figures 7 to 11, the summary generation system 100 can sequentially display supplementary information consisting of intermediate speech recognition results and provisional summary results based on the intermediate results, thereby indicating that processing is in progress and what kind of summary results are likely to be output under the current circumstances.

[0079] The summary generation system 100 divides the input audio signal into speech segments based on speech breaks, speaker changes, etc. (see FIGS. 3A and 6 ), generates new summary segment candidates for each new speech segment (see FIG. 3B ), and gradually creates summaries from the start of the summarization and displays them in a selectable form on the display unit 108 (see FIGS. 4 and 7 to 11 ). Thus, a user (e.g., a contact center operator) can compare and consider the multiple summaries created sequentially for each summary segment, and confirm the selection of a summary segment for the corresponding summary segment when an appropriate summary segment has emerged. The summary generation system 100 then repeatedly creates summary segment candidates and summaries, starting from the speech segment immediately following the summary segment for which a summary segment has been confirmed. Therefore, the user can continue to sequentially select an appropriate summary segment for each summary segment while speech input continues (e.g., while the customer interaction continues).

[0080] The operation and screen transitions of the summary generation system 100 will be explained in chronological order with reference to Figures 14 to 21. However, for the sake of convenience, in Figures 14 to 21, it is assumed that the most recent summary is selected by default and displayed in the summary candidate display field 410.

[0081] 14 shows the operation and display screen at the time when the speech segment detection unit 102 detects speech segment 1, the speech recognition unit 103 performs speech recognition processing on speech segment 1 to generate text 1, the summary generation unit 106 summarizes text 1 of summary segment candidate 1 to generate summary sentence 1, and then summarizes summary sentence 2 of the summary segment candidate including text 2 obtained by speech recognition for the next speech segment 2. At this point, the summary candidate display field 410 displays the completed summary sentence 1, and to the right of summary sentence 1 is displayed the character string "[summarizing...]" indicating that summary sentence 2 of the next summary segment candidate is currently being summarized. It is assumed that the summary sentence currently displayed in the summary candidate display field 410 and its corresponding summary segment candidate are in a selected state (the same applies hereinafter).

[0082] 15 shows the operation and display screen at the time when the speech recognition unit 103 has generated text 2 up to speech segment 2 detected by the speech segment detection unit 102, the summary generation unit 106 has generated summary sentence 2 by summarizing summary segment candidate 2 including text 1 and text 2, and is further summarizing summary sentence 3 of the summary segment candidate including text 3 obtained by speech recognition for the next speech segment 3. At this point, the completed summary sentence 2 is displayed in the summary candidate display field 410. Also, an ellipsis "..." is placed to the left of summary sentence 2, indicating that a shorter summary segment candidate is available, and a character string "[summarizing...]" is placed to the right of summary sentence 2, indicating that summary sentence 3 of the next summary segment candidate is currently being summarized.

[0083] 16 shows the operation and display screen at the time when the speech recognition unit 103 has generated text 3 up to speech segment 3 detected by the speech segment detection unit 102, the summary generation unit 106 has generated summary sentence 3 by summarizing summary segment candidate 3 including texts 1 to 3, and is currently summarizing summary sentence 4 of the summary segment candidate including text 4 obtained by speech recognition for the next speech segment 4. At this point, the summary candidate display field 410 displays the completed summary sentence 3. Also, to the left of summary sentence 3 is placed an ellipsis "..." indicating that a shorter summary segment candidate is available, and to the right of summary sentence 3 is placed the character string "[summarizing...]" indicating that summary sentence 4 of the next summary segment candidate is currently being summarized.

[0084] 17 shows the operation and display screen at the time when the speech recognition unit 103 has generated text 4 up to detected speech segment 4, the summary generation unit 106 has generated summary sentence 4 by summarizing summary segment candidate 4 including texts 1 to 3, and is further summarizing summary sentence 5 of the summary segment candidate including text 5 obtained by speech recognition for the next speech segment 5. At this point, the completed summary sentence 4 is displayed in the summary candidate display field 410. Also, an ellipsis "..." is placed to the left of summary sentence 4, indicating that a shorter summary segment candidate is available, and a character string "[summarizing...]" is placed to the right of summary sentence 4, indicating that summary sentence 5 of the next summary segment candidate is currently being summarized.

[0085] 18 shows the state in which the accepting unit 109 accepts the user's pressing of the Tab key or Esc key when summary sentence 4 of summary segment candidate 4 is displayed in the summary candidate display field 410 in a selected state as shown in FIG. 17. By pressing the Tab key, the user confirms the selection of the selected summary segment candidate 4. The user can also instruct the discarding of the summaries of summary segment candidates 1 to 4 by pressing the Esc key. In either case, when the Tab key or Esc key is pressed on selected summary sentence 4, it becomes unnecessary to retain summary sentences 1 to 4 and texts 1 to 4 that constitute summary segment candidate 4, and so they are discarded from the recognition result memory 104, as shown in FIG. 19.

[0086] 19, after discarding texts 1 to 4 of speech segments 1 to 4 that are no longer needed, speech segment 5 becomes the first unprocessed speech segment. Therefore, as shown in FIG. 20, the speech segment detection unit 102 detects speech segment 5, the speech recognition unit 103 performs speech recognition processing on speech segment 1 to generate text 5, the summary generation unit 106 summarizes text 5 of summary segment candidate 5 to generate summary sentence 5, and then summarizes summary sentence 6 of the summary segment candidate including text 6 obtained by speech recognition for the next speech segment 6. The summary candidate display field 410 then displays summary sentence 5 as the newly selected one, and displays the character string "[summarizing...]" to the right of summary sentence 5, indicating that summary sentence 6 of the next summary segment candidate is currently being summarized.

[0087] 21 shows the operation and display screen at the time when the speech recognition unit 103 has generated text 6 up to speech segment 6 detected by the speech segment detection unit 102, the summary generation unit 106 has generated summary sentence 6 by summarizing summary segment candidate 6 including text 5 and text 6, and is further summarizing summary sentence 7 of the summary segment candidate including text 7 obtained by speech recognition for the next speech segment 7. At this point, the completed summary sentence 6 is displayed in the summary candidate display field 410. Also, an ellipsis "..." is placed to the left of summary sentence 6, indicating that a shorter summary segment candidate is available, and a character string "[summarizing...]" is placed to the right of summary sentence 6, indicating that summary sentence 7 of the next summary segment candidate is currently being summarized.

[0088] Thereafter, the same process as described above is repeated each time a new subsequent voice section is input, and each time the Tab key or Esc key is pressed.

[0089] C-2. Automatic Selection of Summary Sections In the basic operation described in Section C-1 above, the user can sequentially select a better summary sentence in an appropriate summary section through simple manual operations using special keys such as the Tab key or Esc key. The summary generation system 100 may further include an optional function for automatic selection of summary sections. For example, predetermined criteria may be set in advance, and when an event that meets these criteria occurs, the function may be activated to automatically select a summary section without relying on user operation. By using the automatic selection function, for example, a contact center operator is relieved of the burden of manual operations of comparing and considering the summaries generated sequentially for each candidate summary section and selecting a summary section, allowing them to focus more on responding to customers.

[0090] Examples of criteria for activating the automatic summary segment selection function include a method of detecting summary segment boundaries using context understanding and a method of determining summary segment boundaries based on a threshold such as elapsed time. The former method of dividing summary segments using context understanding includes, for example, dividing a "question and its answer," a "series of utterances until the speaker changes," or an "explanation of a series of operations or procedures" into a single summary segment. A trained model can be applied to context understanding. The latter method of dividing summary segments using elapsed time as a threshold includes dividing summary segments based on the duration of the speech segment. Each case will be explained further below.

[0091] Questions and their answers: In the case of a question-and-answer exchange between an operator and a customer, such as "Are you currently under contract?" and "No, it only last month," automatic selection of summary sections can be used to combine these to generate the summary sentence "The contract last month." If the question and answer section between the operator and the customer are identified through conversation analysis, by automatically expanding the summary section to include both the voice section data for the question and the voice section data for the answer, rather than separating them, a summary sentence with a high summary rate can be created that includes information on the question and the answer.

[0092] A series of utterances leading up to a speaker change: In situations where a specific speaker is explaining a situation, rather than a two-way dialogue such as a question and answer, a specific speaker may speak continuously. In multi-channel recording, where a channel is recorded for each speaker, it is possible to separate a series of utterances by a specific speaker in each channel. In other cases, it is also possible to separate a series of utterances by a specific speaker using speaker identification. Therefore, by automatically expanding the summary section to include the series of utterances leading up to the speaker change, it is possible to create a summary that includes the situation explanation and has a high summarization rate.

[0093] Explanation of a series of operations or procedures: In contact centers, operators may provide customers with guidance in accordance with a specific manual, such as instructions on how to operate or set up the system, checking the status, and procedures. In such cases, it is sufficient to have a record that all the necessary information has been provided. By automatically adjusting the length of the summary intervals and determining which parts of the manual have been explained without any omissions or omissions based on the degree of match between the text data obtained by speech recognition of the operator's speech and the explanations in the manual, accurate summaries with a high summarization rate can be created without relying on manual operation by the operator.

[0094] Duration of speech interval: When a contact center operator is assisting a customer, the speech is interrupted when they need to do some research or when the customer is put on hold. By expanding the summary interval to include intervals where there is no speech for a specified period of time (for example, 10 seconds), it is possible to create an accurate summary based on the time period.

[0095] C-3. Method of Using the Summary Selected by Comparison and Analysis As described in Section C-1 above, the summary generated for each summary section candidate is manually compared and the selected summary is considered to be a summary that contains more necessary information than the other summary section candidates. It is considered that the user selected the manually selected summary section candidate because, compared to the summary section candidates being compared in particular, the summary contains characteristic words and phrases that retain important information. Therefore, the words and phrases remaining in the summary selected by the user can be used for other purposes (i.e., purposes other than creating a summary).

[0096] The summary generation system 100 can automatically extract important words and phrases that contributed to the user's selection by comparing the summary sentence selected by the user with the summary sentences used as comparison targets. For example, if, while displaying the summary sentence of a certain summary section candidate in the summary candidate display field 410 shown in Figure 4, the user moves back to a shorter summary section candidate by moving in the direction of the ellipsis "..." 412 on the left and selects a summary sentence, the summary sentence displayed before moving in the direction of the ellipsis "..." 412 on the left is the summary sentence that the user compares. Similarly, if the user moves to a longer summary section candidate indicated by the ellipsis "..." 413 on the right and selects a summary sentence, the summary sentence displayed by moving to the ellipsis "..." 413 on the right is the summary sentence that the user compares.

[0097] Specific uses of important words remaining in the summary selected by the user include recording minutes and telephone conversations, searching for and presenting related materials, and providing information to managers. Furthermore, the summary selected through comparative analysis can also be used as training data when building a summary generation model. The following provides additional information on each use.

[0098] Use in recording minutes and telephone responses: When recording meeting minutes or telephone responses, priority is given to preserving important words and phrases extracted based on differences with a comparison summary. A summary generated through real-time speech recognition is like a memorandum taken during a meeting or a call. After the meeting or telephone response ends, the summary generation system 100 re-summarizes the summaries for each summary section sequentially generated in real time, thereby creating a new, official (or recordable) copy of the minutes of the entire meeting or the entire telephone response history. By prioritizing preserving important words and phrases extracted based on differences with a comparison summary, a meaningful and accurate overall summary can be generated.

[0099] Use in searching and presenting related materials: For example, a contact center operator may want to view materials related to a customer's inquiry while answering the phone. A typical information search is performed by manually entering or speaking words. In contrast, the summary generation system 100 can automatically search for related materials using important words remaining in the summary selected by the user as search terms, and present the materials.

[0100] Alerts to managers: Contact center managers can set keywords to be monitored in advance. If an important phrase remaining in a summary selected by an operator while interacting with a customer matches a keyword, an alert is sent to the manager. This allows managers to quickly become aware of situations that require monitoring.

[0101] Use as training data: The summary sentences selected through comparison and consideration can also be used as training data when constructing a summary generation model. In the summary generation system 100, the user reviews and judges each of the multiple summary sentences of the candidate summary sections automatically generated from the input acoustic signal, and selects the summary sentence that they deem good. Therefore, based on the summary sentences selected by the user, an intermediate model (a reward model that infers the quality of the summary sentence) can be trained, and the original summary generation model can be trained (reinforced learning with human feedback: RLHF) to match the user's judgment.

[0102] C-4. Automatic Summary Confirmation Function As a basic operation, the summary generation system 100 allows the user to manually select summary intervals, as described in section C-1 above, and outputs the summary generated for each summary interval. Furthermore, as an extended function, the summary generation system 100 can provide a function for automatically selecting summary intervals based on predetermined criteria, as described in section C-2 above. The summary generation system 100 may also provide an automatic summary confirmation function.

[0103] For example, while in a meeting or on the phone with a client, a user manually selects a summary section one by one while checking the summary sentences for each candidate summary section. However, there are cases where the user is unable to perform the manual operation while concentrating on the meeting or phone call. Furthermore, retroactively determining a summary sentence requires the user to recall the situation at that time, which is difficult. Therefore, the summary generation system 100 may, for example, preset a predetermined event and automatically determine a summary sentence when this event occurs. Furthermore, the summary generation system 100 may automatically determine a summary sentence while expanding the summary section in combination with the automatic summary section selection function.

[0104] Examples of events that trigger the automatic summary confirmation function include an increase in unconfirmed summary candidates, the passage of time, the status of a call or meeting, etc. Below, we will provide additional information on each case.

[0105] Increase in unconfirmed summary candidates: In the summary generation system 100, a summary is generated for each summary section candidate determined by the summary section candidate determination unit 105, and the user selects one of the summary sections based on the results of comparing and examining the summary sections. The accumulation of many unconfirmed summary section candidates also indicates that the overall system processing is stagnating. The summary generation system 100 may automatically confirm a summary when the number of unconfirmed candidates exceeds a predetermined number. For example, when 10 unconfirmed speech section data have accumulated, the first five pieces of speech section data may be automatically determined as summary sections, and the summaries generated for each summary section may be confirmed. Specifically, of the 10 pieces of unconfirmed speech section data, the first three pieces of speech section data and the fourth and fifth pieces of speech section data may be determined as summary sections, and the summaries generated for each of these summary sections may be confirmed. The sixth and subsequent pieces of speech section data may then be left to the user's manual selection of the summary section.

[0106] Time lapse: A speech section that has remained unprocessed for a predetermined time (for example, a dozen seconds to several tens of seconds) is automatically determined to be a summary section, preventing processing in the entire system from stalling.

[0107] Call or conference status: When the target of a summary is a telephone call or web conference, information about the status, such as on hold or end, can be acquired. The summary section candidate determination unit 105 may determine a summary section based on the on hold or end status without relying on input to the reception unit 109, and may automatically determine the summary sentence for that summary section.

[0108] D. Specific Configuration Section B above described a basic and schematic configuration example of the summary generation system 100 that realizes the present disclosure. Section D describes a detailed configuration example of the summary generation system 100 that can realize the basic operations and each of the extended functions described in Section C above.

[0109] 12 shows a detailed configuration example of the summary generation system 1200. The summary generation system 1200 may be configured using a single information processing device, or may be configured by linking two or more information processing devices. At least some of the functions of the summary generation system 1200 may be located on the cloud. The summary generation system 1200 is intended to summarize, for example, the content of conversations between two parties in a use case in which a contact center operator is answering a customer's phone call while referring to a manual.

[0110] The summary generation system 1200 shown in Figure 12 includes a speech segment detection unit 1201, a speech recognition unit 1202, a recognition result memory 1203, a summary processing unit 1204, a display control unit 1205, a display unit 1206, an operation reception unit 1207, a display / input control unit 1208, a speaker recognition unit 1209, a viewing / reading portion identification unit 1210, an automatic candidate selection unit 1211, an automatic candidate determination unit 1212, and an operation recording unit 1213. However, components that are substantially the same as those included in the summary generation system 100 shown in Figure 1 are given the same names. At least some of these components are realized by executing a predetermined program on an information processing device such as a computer. Each unit will be described below.

[0111] The voice activity detection unit 1201 extracts voice activity segments that appear to be voice (human voice) from the audio signal captured from the audio input unit (not shown in FIG. 12). The voice activity detection unit 1201 divides the audio signal into voice activity segments based on audio breaks, speaker changes, etc. In addition to determining whether or not a signal is a human voice, the voice activity detection unit 1201 may also divide the signal into time segments and extract the voice activity segments.

[0112] The speech recognition unit 1202 uses a speech recognition model to perform speech recognition processing and convert into text for each speech segment detected by the speech segment detection unit 1201. In Fig. 12, the speech recognition unit 1202 outputs "confirmed text" for a speech segment for which speech recognition processing has been confirmed, and "tentative text" for a speech segment in the middle of speech recognition processing.

[0113] The recognition result memory 1203 temporarily stores the speech recognition results (confirmed text) for each speech segment that are sequentially output from the speech recognition processing unit 1202 .

[0114] The text data for each speech segment stored in the recognition result memory 1203 is read out for each candidate summarization segment into the subsequent summarization processor 1204. Basically, the first unprocessed speech segment is set as the summarization start point, and multiple candidate summarization segments are determined by combining one, two, three, ..., N, ... speech segments from the summarization start point, and text data is read out for each candidate summarization segment, and a summary is created in the summarization processor 1204. For example, if there are four speech segments 1 to 4, the text data of the candidate summarization segment consisting of only speech segment 1, the text data of the candidate summarization segment consisting of speech segments 1 and 2, the text data of the candidate summarization segment consisting of speech segments 1 to 3, and the text data of the candidate summarization segment consisting of speech segments 1 to 4 are read out in this order from the recognition result memory 1203, and a summary is created in the summarization processor 1204 for each of them.

[0115] Furthermore, when the Tab key or Esc key is pressed on a summary created from text data of a candidate summary segment, the text data of the speech segments prior to that point is no longer required to be retained and is therefore discarded from the recognition result memory 1203. The first speech segment after the discarding is then set as the summarization start point, and multiple summary segment candidates are re-determined by combining one, two, three, ..., N, ... speech segments from that summarization start point, and text data for each candidate summary segment is read into the summarization processor 1204. For example, in the case where there are four speech segments 1 to 4, when the Tab key or Esc key is pressed on a summary created from a candidate summary segment consisting of speech segments 1 and 2, the text data up to speech segment 2 is discarded from the recognition result memory 1203, and then the text data of the candidate summary segment consisting only of the immediately following speech segment 3 and the text data of the candidate summary segment consisting of speech segments 3 and 4 are read again from the recognition result memory 1203, and summarizations are created in the summarization processor 1204 for each segment.

[0116] The summary processor 1204 uses a summary generation model to generate a summary of the input text data. The summary generation model is, for example, a large-scale language model (LLM) such as GPT-4 (see Non-Patent Document 2). In FIG. 12, the summary processor 1204 targets for summarization "provisional text" for a speech segment during the speech recognition process and "confirmed text" for a speech segment for which the speech recognition process has been finalized, both of which are output from the speech recognition unit 1202. The summary processor 1204 outputs a provisional summary result based on the intermediate results of the speech recognition from the provisional text. The summary processor 1204 also generates a definitive summary from the finalized text for each candidate summary segment.

[0117] When speech is detected by the speech activity detection unit 1201, the display control unit 1205 causes the display unit 1206 to display a provisional summary result based on the intermediate results of speech recognition. At this time, as described with reference to Figures 7 to 11, the display control unit 1205 displays the provisional summary result in a manner that makes it easy to visually understand that speech recognition and summarization processing are in progress and that the presented summary sentence is a provisional summary result. Furthermore, the display control unit 1205 may output the provisional summary result received from the summary processing unit 1204 to an editing function such as "memo pad."

[0118] The display / input control unit 1208 causes the display unit 1206 to display the summary generated by the summarization processing unit 1204 from the confirmed text for each candidate summarization segment. In this case, the display / input control unit 1208 causes the summary candidate display field 410 to display the summary results of the candidate summarization segments, the intermediate results of the summarization, the presence or absence of shorter and longer candidate summarization segments, the fact that speech recognition processing or summarization processing is in progress, and the like, in a manner that makes it easy to visually understand, as described with reference to Fig. 4. The display / input control unit 1208 may output the confirmed summary results received from the summarization processing unit 1204 to an editing function such as "Memopad."

[0119] Furthermore, the display / input control unit 1208 switches the display of summary section candidates and summary sentences in the summary candidate display field 410 in response to user operations accepted by the operation acceptance unit 1207 (including operations based on the determination results of the automatic candidate confirmation units 1211 and 1212). When the Tab key or Esc key is pressed on a summary section candidate and summary sentence currently being displayed (i.e., selected) in the summary candidate display field 410, the display / input control unit 1208 discards the summary section candidate and summary sentence and waits until a new summary section candidate and summary sentence are output from the summary processing unit 1204. Furthermore, the Tab key instructs the confirmation of the selection of the summary section candidate and summary sentence currently being displayed (i.e., selected) in the summary candidate display field 410, and therefore outputs the summary section candidate and summary sentence as the confirmation results. The output destination may be, for example, an output file that saves the output dataset of the summary sentence generation system 1200.

[0120] The operation accepting unit 1207 is configured with an easily operable input device such as a keyboard, and notifies the display / input control unit 1208 of operations such as Tab key and Esc key operations and display switching instructions (e.g., an instruction to move to a digest segment candidate that is shorter or longer than the currently displayed digest segment candidate). The operation accepting unit 1207 may also accept instructions from the user, such as discarding past recognition results or re-summarizing. In addition to user operations, the operation accepting unit 1207 also reflects the results of automatic selection of digest segments by the automatic candidate selection unit 1211 and automatic determination of digest segments by the automatic candidate determination unit 1212 as user operations.

[0121] The speaker recognition unit 1209 identifies, based on the voice, the speaker of each voice segment detected by the voice segment detection unit 1201. However, in the case of one-to-one telephone conversation or multi-channel voice consisting of a channel for each speaker, the speaker recognition unit 1209 may identify the speaker by channel.

[0122] The viewing / reading portion specifying unit 1210 specifies the portion in the manual that is being read aloud by the contact center operator from the recognition result (confirmed text) of the voice recognition unit 1202. The viewing / reading portion specifying unit 1210 may specify the portion to be read aloud by using information such as an operation log of the manual (for example, the page that is open in a PDF file being viewed on a PC) in addition to the voice recognition result.

[0123] The automatic candidate selection unit 1211 automatically selects a summary section from the recognition results (confirmed text) sequentially output by the speech recognition unit 1202 for each speech section. While the basic operation of the summary generation system 1200 is for the user to manually select a summary section, the automatic candidate selection unit 1211 automatically selects a summary section that meets predetermined criteria. For example, the automatic candidate selection unit 1211 may extract a "question and its answer," a "series of utterances until the speaker switches," or an "explanation of a series of operations or procedures" from the confirmed text and automatically select them as a single summary section (see Section C-2 above). In this case, the automatic candidate selection unit 1211 can select the section of the "explanation of a series of operations or procedures" based on the reading section identified by the viewing / reading section identification unit 1210. The automatic candidate selection unit 1211 may also automatically select a summary section that is expanded to include a section without speech for a predetermined period of time (e.g., 10 seconds) or more. The automatic candidate selection unit 1211 can automatically widen the summarization interval while leaving important information intact, thereby improving the summarization rate.

[0124] The automatic candidate determination unit 1212 automatically determines one of a plurality of summary interval candidates and automatically selects a summary for that summary interval. The basic operation of the summary generation system 1200 is for a user to manually select a summary interval and output a summary generated for each summary interval, but the automatic candidate determination unit 1212 automatically determines a summary interval and automatically determines a summary when a predetermined event occurs. The automatic candidate determination unit 1212 may operate in parallel with the automatic candidate selection unit 1211. For example, the automatic candidate determination unit 1212 automatically determines a summary interval based on information such as an increase in undetermined summary interval candidates, the passage of time, and the status of calls and meetings (see Section C-4 above).

[0125] The operation recording unit 1213 records output data and intermediate data during processing of the summary generation system 1200, such as the summary intervals and summary content compared and considered in accordance with user operations accepted by the operation accepting unit 1207, and the finalized summary intervals and summary content. The information recorded by the operation recording unit 1213 can be used for subsequent model learning (e.g., reinforcement learning of the summary generation model).

[0126] FIG. 13 shows an example of the operation of the summary generation system 1200 in the form of a flowchart.

[0127] The voice activity detection unit 1201 extracts a voice activity that appears to be a voice (human voice) from the acoustic signal received from the voice input unit (not shown in FIG. 12) (step S1301).

[0128] Next, the speech recognition unit 1202 uses a speech recognition model to perform speech recognition processing for each speech segment detected by the speech segment detection unit 1201, and converts the speech segment into text (step S1302).

[0129] Next, each time text data of a subsequent new speech segment is input, the summarization processor 1204 determines a new summary segment candidate by concatenating it with the end of the immediately preceding summary segment candidate, and generates a summary for each summary segment (step S1303).The summary segment candidate determined by the summary segment candidate determiner 105 is then selected by default, and the summary is displayed on the display unit 1206 (for example, in the summary candidate display field 410 shown in FIG. 4) (step S1304).

[0130] Thereafter, when the user switches the selection of the summary section candidate via the operation reception unit 1207, or when the automatic candidate selection unit 1211 automatically selects a summary section (Yes in step S1305), the process switches to the selected summary section candidate (step S1311), and the process returns to step S1304 to display the summary sentence generated from the selected summary section candidate.

[0131] If the user instructs the operation accepting unit 1207 (for example, by using the Esc key) to discard the selected summary segment candidate (Yes in step S1306), all summary segment candidates up to the instructed summary segment candidate (and the speech segment data constituting the summary segment candidate) are discarded (step S1310).Then, the process returns to step S1303, where a new summary segment candidate is determined from the speech segment data remaining after the discarding, and the same process as above is repeated.

[0132] Furthermore, when the user confirms the selected summary section candidate via the operation reception unit 1207 (for example, by using the Tab key) or when the automatic candidate confirmation unit 1212 automatically confirms the summary section (Yes in step S1307), the summary section candidate and the summary sentence are output as confirmation results (step S1308). The output destination may be, for example, an output file that saves the output dataset of the summary sentence generation system 1200.

[0133] On the other hand, if the selected summary section candidate is not confirmed (No in step S1307), the process returns to step S1303, and the same process as above is repeated for the newly determined summary section candidate by adding the newly input speech section data.

[0134] The above process continues until a predetermined termination condition is met (for example, until the contact center operator finishes dealing with the customer) (No in step S1309), at which point the process returns to step S1303, where new summary section candidates are determined from the speech section data remaining after the discarding, and the same process as above is repeated. When returning to step S1303, all summary section candidates (and the speech section data constituting the summary section candidates) up to the summary section candidate whose selection has been confirmed are discarded (step S1310).

[0135] E. Configuration Example of Information Processing Device In this section E, the configuration of an information processing device that can be used for the summary generation systems 100 and 1200 will be described.

[0136] 22 shows an example of the hardware configuration of an information processing device 2000. This information processing device 2000 includes a CPU (Central Processing Unit) 2001, a ROM (Read Only Memory) 2002, a RAM (Random Access Memory) 2003, a host bus 2004, a bridge 2005, an expansion bus 2006, an interface unit 2007, an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013. The information processing device 2000 is configured, for example, by an information terminal such as a personal computer, a tablet, or a smartphone.

[0137] The CPU 2001 controls the overall operation of the information processing device 2000 in accordance with various programs. When performing processes with a high computational load on the information processing device 2000 (for example, processes related to model learning such as a speech recognition model or a summary generation model), it is desirable that the CPU 2001 be a multi-core CPU (for example, Apple M1 Max, etc.), or that the information processing device 2000 further be equipped with a multi-core processor (for example, NVIDIA R6000, etc.) such as a GPU (Graphics Processing Unit) or GPGPU (General-purpose computing on graphics processing units) in addition to the CPU 2001. However, for convenience, these will be collectively referred to simply as the CPU 2001 below.

[0138] The ROM 2002 stores in a nonvolatile manner programs (such as a basic input / output system) and calculation parameters used by the CPU 2001. The RAM 2003 is used to load programs to be executed by the CPU 2001 and to temporarily store parameters such as working data that change as appropriate during program execution. Programs loaded into the RAM 2003 and executed by the CPU 2001 include, for example, various application programs and an operating system (OS).

[0139] The CPU 2001, ROM 2002, and RAM 2003 are interconnected by a host bus 2004, which includes a CPU bus and other components. The CPU 2001 executes various application programs in an execution environment provided by an OS through the cooperative operation of the ROM 2002 and RAM 2003, thereby enabling various functions and services. If the information processing device 2000 is a personal computer, the OS may be, for example, Microsoft Windows (registered trademark), Unix (registered trademark), or a successor OS. Furthermore, application programs executed on the information processing device 2000 include a program that executes a summary generation process that generates a summary from the content of utterances, such as those made during telephone conversations. Furthermore, the information processing device 2000 may include application programs used in the summary generation process, such as a speech recognition model or a summary generation model. At least some of the application programs may be computer programs provided as libraries.

[0140] The host bus 2004 is connected to an expansion bus 2006 via a bridge 2005. The expansion bus 2006 is, for example, a PCI (Peripheral Component Interconnect) bus or PCI Express, and the bridge 2005 is based on the PCI standard. However, the information processing device 2000 does not need to be configured so that the circuit components are separated by the host bus 2004, bridge 2005, and expansion bus 2006, and may be implemented so that almost all circuit components are interconnected by a single bus (not shown).

[0141] The interface unit 2007 connects peripheral devices such as an input unit 2008, an output unit 2009, a storage unit 2010, a drive 2011, and a communication unit 2013 in accordance with the standards of the expansion bus 2006. However, not all of the peripheral devices shown in Fig. 22 are necessarily required, and the information processing device 2000 may further include peripheral devices not shown. Furthermore, the peripheral devices may be built into the main body of the information processing device 2000, or some of the peripheral devices may be externally connected to the main body of the information processing device 2000.

[0142] The input unit 2008 is composed of an input control circuit that generates an input signal based on user input and outputs it to the CPU 2001. When the information processing device 2000 is a personal computer, the input unit 2008 may include a keyboard, mouse, and touch panel, as well as a camera and microphone used for remote conferences and face-to-face customer service. The output unit 2009 includes, for example, a liquid crystal display (LCD), an organic electroluminescence (EL) display, or an LED (light-emitting diode), as well as a sound output device such as a speaker. The input unit 2008 and the output unit 2009 are used to display a GUI screen (see, for example, FIGS. 4 and 7 to 11) during the execution of the summary generation process according to the present disclosure, and to accept user operations (such as input of the Tab key or Esc key).

[0143] The storage unit 2010 stores files such as programs (applications, OS, etc.) executed by the CPU 2001 and various data. The storage unit 2010 is configured with a large-capacity storage device such as an SSD (Solid State Drive) or an HDD (Hard Disk Drive), but may also include an external storage device.

[0144] The removable storage medium 2012 is a storage medium configured as a cartridge, such as a microSD card. The drive 2011 performs read and write operations on the loaded removable storage medium 113. The drive 2011 outputs data read from the removable storage medium 2012 to the RAM 2003 or the storage unit 2010, and writes data on the RAM 2003 or the storage unit 2010 to the removable storage medium 2012.

[0145] The communication unit 2013 is a device that performs wireless communication via Wi-Fi (registered trademark), Bluetooth (registered trademark), or cellular communication networks such as 4G and 5G. The communication unit 2013 may also include terminals such as a Universal Serial Bus (USB) or a High-Definition Multimedia Interface (HDMI) (registered trademark), and may further include a function for performing HDMI (registered trademark) communication with USB devices such as scanners and printers, displays, and the like. Programs executed on the information processing device 2000 are installed from an external device, for example, via the communication unit 2013. An acoustic signal that is the subject of the summary generation process according to the present disclosure is captured, for example, via the communication unit 2013.

[0146] The present disclosure has been described in detail above with reference to specific embodiments. However, the present disclosure should not be construed as being limited to the above-described embodiments, and it is obvious that those skilled in the art can modify or substitute the embodiments without departing from the spirit of the present disclosure. Furthermore, the effects described in this specification are merely examples, and the effects brought about by the present disclosure are not limited thereto, and additional effects not described in this specification may exist.

[0147] The present disclosure can be applied to, for example, a contact center, and can create summaries of utterances made during customer service. While serving a customer, an operator can issue instructions for dividing the speech recognition results of the utterance into summary segments of appropriate length with simple and minimal input operations. Therefore, the operator can obtain high-quality summaries for each summary segment without missing any important points while continuing to serve the customer.

[0148] The scope of application of the present disclosure is not limited to contact centers. The present disclosure can also be applied to generating summaries of utterances between multiple speakers in conferences and various other situations. According to the present disclosure, one of the speakers participating in a dialogue can specify a summary section with simple and minimal input operations while participating in the dialogue. Of course, a person other than a participant in the dialogue (such as an observer of the dialogue) may also specify a summary section.

[0149] In short, the present disclosure has been described in the form of examples, and the contents of the specification should not be interpreted as limiting. To determine the gist of the present disclosure, the claims should be taken into consideration.

[0150] The series of processes described in this specification can be executed by hardware, software, or a configuration that combines hardware and software. When executing processes by software, a program recording a processing sequence related to realizing the present disclosure is installed in memory in a computer incorporated in dedicated hardware and executed. It is also possible to install the program in a general-purpose computer capable of executing various processes and execute the processes related to realizing the present disclosure.

[0151] The program can be stored in advance on a recording medium installed in the computer, such as a HDD, SSD, or ROM. Alternatively, the program can be temporarily or permanently stored on a removable recording medium such as a flexible disk, CD-ROM (Compact Disc Read Only Memory), MO (Magneto Optical) disk, DVD (Digital Versatile Disc), BD (Blu-Ray Disc (registered trademark)), magnetic disk, or USB (Universal Serial Bus) memory. Using such a removable recording medium, a program related to the realization of the present disclosure can be provided as so-called package software.

[0152] The program may also be transferred wirelessly or via a wire from a download site to a computer via a network such as a wide area network (WAN) typified by cellular, a local area network (LAN), the Internet, etc. The computer can receive the program transferred in this manner and install it in a large-capacity storage device such as an HDD or SSD within the computer.

[0153] The present disclosure may also be configured as follows.

[0154] (1) An information processing device comprising: an acquisition unit that acquires a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination unit that determines a plurality of candidate summary segments having different combinations of speech segment data from the plurality of speech segment data; a reception unit that receives a summary segment selected by a user from the plurality of candidate summary segments; and a generation unit that generates a summary sentence for the summary segment; and the information processing device determines the summary segment and the summary sentence for the summary segment received by the reception unit.

[0155] (2) The information processing device according to (1), further comprising: a display control unit that displays information about the plurality of summary section candidates in a selectable manner in response to a user operation received by the receiving unit.

[0156] (3) The information processing device according to (2), wherein the display control unit displays a selectable summary generated by the generation unit for any one of the plurality of summary section candidates.

[0157] (3-1) The information processing device according to (3), wherein the display control unit displays a provisional summary generated by the generation unit for the last determined (or latest) candidate summary section.

[0158] (4) The information processing device according to (3), wherein the accepting unit accepts a selection of a corresponding candidate summary section based on a predetermined operation by a user on the displayed summary sentence.

[0159] (5) The information processing device described in any one of (3) or (4) above, wherein the display control unit displays, near the displayed summary sentence, whether or not there is at least one of a candidate for a shorter summary section and a candidate for a longer summary section.

[0160] (6) The information processing device according to any one of (3) to (5), wherein the display control unit displays, near the displayed summary sentence, whether or not there are candidates for new summary sections during speech recognition processing or summarization processing.

[0161] (7) The information processing device according to any one of (3) to (6), wherein the display control unit sequentially displays an intermediate result of the speech recognition and a summary result based on the intermediate result of the speech recognition.

[0162] (8) The information processing device according to any one of (1) to (7), wherein the information processing device selects a summary section that satisfies a predetermined criterion regardless of whether the accepting unit has accepted a selection by a user.

[0163] (8-1) The information processing device described in (8) above, wherein the summary section that satisfies the predetermined criteria includes at least one of a section of a question and its answer, a section of a series of utterances until the speaker switches, a section of a series of explanations regarding a predetermined matter, and a section until the audio is interrupted for a predetermined period of time or more.

[0164] (9) The information processing device according to any one of (1) to (8), further comprising: a processing unit that performs processing using the determined summary sentence.

[0165] (10) The information processing device according to (9), wherein the processing unit processes words and phrases extracted based on a difference between the summary sentence of the summary section received by the receiving unit and the summary sentence compared and considered by the user.

[0166] (10-1) The information processing device described in (10) above, wherein the processing unit performs at least one of creating minutes or records of telephone conversations including the phrase, searching for related documents based on the phrase, and processing based on the results of comparing the phrase with pre-set keywords.

[0167] (11) The information processing device according to any one of (1) to (10), wherein the summary is determined based on the occurrence of a predetermined event, regardless of whether the accepting unit has accepted a selection by the user.

[0168] (11-1) The information processing device according to (11), wherein the predetermined event includes at least one of an increase in the number of undetermined summary section candidates, the passage of time, and a state of a summary target.

[0169] (12) The information processing device described in any one of (1) to (11) above, wherein the determination unit determines, from the plurality of speech section data acquired in time series, a candidate summary section consisting of one speech section data from the beginning, and a candidate summary section consisting of two or more consecutive speech section data from the beginning.

[0170] (13) The information processing device described in (12) above, wherein after any of the summary section candidates is selected via the reception unit, the determination unit generates a new summary section candidate from speech section data immediately after the last speech section data included in the selected summary section.

[0171] (14) An information processing method comprising: an acquiring step of acquiring a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determining step of determining a plurality of summary segment candidates each having a different combination of speech segment data from the plurality of speech segment data; a receiving step of receiving a summary segment selected by a user from the plurality of summary segment candidates; and a generating step of generating a summary sentence for the summary segment, wherein the summary segment and the summary sentence for the summary segment received in the receiving step are determined.

[0172] (15) A computer program written in a computer-readable format, in which a computer functions as: an acquisition unit that acquires a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination unit that determines a plurality of summary segment candidates each having a different combination of speech segment data from the plurality of speech segment data; a reception unit that receives a summary segment selected by a user from the plurality of summary segment candidates; and a generation unit that generates a summary sentence for the summary segment; and the computer program determines the summary segment and the summary sentence for the summary segment received by the reception unit.

[0173] 100... Summary generation system, 101... Speech input unit, 102... Speech interval detection unit, 103... Speech recognition unit, 104... Recognition result memory, 105... Summary interval candidate determination unit, 106... Summary generation unit, 107... Display control unit, 108... Display unit, 109... Reception unit, 110... Input control unit, 1200... Summary generation system, 1201... Speech interval detection unit, 1202... Speech recognition unit, 1203... Recognition result memory, 1204... Summary processing unit, 1205... Display control unit, 1206... Display unit, 1207... Operation reception unit, 1208... Display / input control unit, 1209... Speaker recognition unit, 1210... Viewing / reading part identification unit, 1211... Automatic candidate selection unit, 1212... Automatic candidate determination unit, 1213... Operation recording unit, 2000... Information processing device, 2001... CPU, 2002... ROM 2003...RAM, 2004...host bus, 2005...bridge, 2006...extension bus, 2007...interface section, 2008...input section, 2009...output section, 2010...storage section, 2011...drive, 2012...removable recording medium, 2013...communication section

Claims

1. An information processing device comprising: an acquisition unit that acquires a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination unit that determines a plurality of candidate summary segments having different combinations of speech segment data from the plurality of speech segment data; a reception unit that accepts a summary segment selected by a user from the plurality of candidate summary segments; and a generation unit that generates a summary sentence for the summary segment; and the information processing device determines the summary segment and the summary sentence for the summary segment accepted by the reception unit.

2. The information processing device according to claim 1, further comprising a display control unit that displays information about the plurality of summary section candidates in a selectable manner in response to a user operation received by the receiving unit.

3. The information processing device according to claim 2, wherein the display control unit displays a selectable summary generated by the generation unit for any one of the plurality of summary section candidates.

4. The information processing device according to claim 3, wherein the accepting unit accepts a selection of a corresponding candidate summary section based on a predetermined user operation on the displayed summary sentence.

5. The information processing device according to claim 3, wherein the display control unit displays, near the displayed summary, whether or not there are at least one of shorter summary section candidates and longer summary section candidates.

6. The information processing device according to claim 3, wherein the display control unit displays, near the displayed summary, whether or not there are candidates for new summary sections during speech recognition processing or summarization processing.

7. The information processing device according to claim 3, wherein the display control unit sequentially displays an intermediate result of the speech recognition and a summary result based on the intermediate result of the speech recognition.

8. The information processing device according to claim 1, wherein the accepting unit selects a digest section that satisfies a predetermined criterion regardless of whether the accepting unit has accepted a selection by the user.

9. The information processing device according to claim 1, further comprising a processing unit that performs processing using the determined summary sentence.

10. The information processing device according to claim 9, wherein the processing unit processes words and phrases extracted based on the difference between the summary sentence of the summary section received by the receiving unit and the summary sentence compared and considered by the user.

11. The information processing device according to claim 1, wherein the summary is determined based on the occurrence of a predetermined event, regardless of whether the accepting unit has accepted a selection by the user.

12. The information processing device according to claim 1, wherein the determination unit determines, from the plurality of speech section data acquired in time series, a candidate summary section consisting of one speech section data from the beginning, and a candidate summary section consisting of two or more consecutive speech section data from the beginning.

13. The information processing device according to claim 12, wherein after any one of the summary section candidates is selected via the reception unit, the determination unit generates a new summary section candidate from the speech section data immediately following the last speech section data included in the selected summary section.

14. An information processing method comprising: an acquisition step of acquiring a plurality of speech segment data consisting of text obtained by speech recognition for each speech segment; a determination step of determining a plurality of candidate summary segments having different combinations of speech segment data from the plurality of speech segment data; a reception step of accepting a summary segment selected by a user from the plurality of candidate summary segments; and a generation step of generating a summary sentence for the summary segment, wherein the summary segment accepted in the reception step and the summary sentence for the summary segment are determined.

15. A computer program written in a computer-readable format, in which a computer functions as: an acquisition unit that acquires multiple speech segment data consisting of text obtained by speech recognition for each speech segment; a determination unit that determines multiple candidate summary segments having different combinations of speech segment data from the multiple speech segment data; a reception unit that accepts a summary segment selected by a user from the multiple candidate summary segments; and a generation unit that generates a summary sentence for the summary segment; and the computer program determines the summary segment and the summary sentence for the summary segment accepted by the reception unit.

Citation Information

Patent Citations

  • Information processing program, information processing method, and information processing device

    JP2021036292A

  • Summary sentence generation device, summary sentence generation method, and program

    JP2022025665A

  • Program, information processing device, information processing system, information processing method, and information processing terminal

    JP2023168692A

  • Systems and methods for sentence based interactive topic-based text summarization

    US20040117725A1