Information processing device, information processing method, and information processing program

The information processing apparatus and method address the issue of inadequate summary generation by estimating partial dialogue histories and generating topic-specific summary documents, improving the accuracy and relevance of the summary.

JP2026091538APending Publication Date: 2026-06-04NEC CORP

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
NEC CORP
Filing Date
2024-11-25
Publication Date
2026-06-04

AI Technical Summary

Technical Problem

Existing techniques for generating summary data from interaction histories fail to adequately capture multiple viewpoints, leading to inappropriate summary documents.

Method used

An information processing apparatus and method that estimates partial dialogue histories related to multiple topics, generates partial summary sentences, and constructs a summary document based on these sentences for each topic, utilizing large language models and named entity recognition.

Benefits of technology

Generates more appropriate summary documents by considering multiple topics in the dialogue history, enhancing the accuracy and relevance of the summary.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 2026091538000001_ABST
    Figure 2026091538000001_ABST
Patent Text Reader

Abstract

This technology provides the ability to generate more appropriate summary documents from dialogue history. [Solution] The information processing device includes an estimation unit that estimates a partial dialogue history related to a topic for each of several topics in the dialogue history, a summarization unit that generates a partial summary from the partial dialogue history, and a generation unit that generates a summary document of the dialogue history based on the partial summary for each of the several topics.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to an information processing apparatus, an information processing method, and an information processing program.

Background Art

[0002] Patent Document 1 describes a technique for generating summary data from an interaction history. This technique extracts the statement (utterance) with the highest score in the interaction history as an important sentence, and repeatedly adds scores to the statements included in the block containing the important sentence and the neighboring blocks, to generate summary data consisting of the extracted important sentences.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the technique described in Patent Document 1, summary data is simply generated from important sentences with high scores, and there is a problem that, for example, appropriate summary data may not be generated from an interaction history with multiple viewpoints that should be emphasized. The present disclosure has been made in view of the above problems, and an exemplary object thereof is to provide a technique for generating a more appropriate summary document from an interaction history.

Means for Solving the Problems

[0005] An information processing apparatus according to an exemplary aspect of the present disclosure includes: an estimation unit that estimates, for each of a plurality of topics in an interaction history, a partial interaction history related to the topic in the interaction history; a summarization unit that generates a partial summary sentence from the partial interaction history; and a generation unit that generates a summary document of the interaction history based on the partial summary sentences for each of the plurality of topics.

[0006] An information processing method relating to an illustrative aspect of this disclosure includes: an estimation process in which at least one processor estimates a partial dialogue history related to a topic in the dialogue history for each of a plurality of topics in the dialogue history; a summarization process in which the at least one processor generates a partial summary sentence from the partial dialogue history; and a generation process in which the at least one processor generates a summary document of the dialogue history based on the partial summary sentence for each of the plurality of topics.

[0007] An illustrative aspect of the present disclosure is an information processing program that causes a computer to function as an information processing device, wherein the computer functions as: estimation means for estimating a partial dialogue history related to a topic in the dialogue history for each of a plurality of topics in the dialogue history; summarization means for generating a partial summary sentence from the partial dialogue history; and generation means for generating a summary document of the dialogue history based on the partial summary sentence for each of the plurality of topics. [Effects of the Invention]

[0008] One illustrative aspect of this disclosure is that it can provide a technology for generating more appropriate summary documents from dialogue history. [Brief explanation of the drawing]

[0009] [Figure 1] This is a block diagram showing the configuration of the information processing device related to this disclosure. [Figure 2] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 3] This diagram schematically shows an overview of the information processing system related to this disclosure. [Figure 4] This is a block diagram showing the configuration of the information processing system related to this disclosure. [Figure 5] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 6]This diagram schematically illustrates the estimation process in an example of application of this disclosure. [Figure 7] This figure schematically illustrates the first extraction process in an example of application of this disclosure. [Figure 8] This diagram schematically illustrates an example of relevant information in an application example of this disclosure. [Figure 9] This figure schematically illustrates an example of a medical report in an application of this disclosure. [Figure 10] This diagram schematically shows an overview of the information processing system related to this disclosure. [Figure 11] This is a block diagram showing the configuration of the information processing system related to this disclosure. [Figure 12] This is a flowchart showing the flow of the information processing method related to this disclosure. [Figure 13] This is a block diagram showing the configuration of the information processing system related to this disclosure. [Figure 14] This block diagram shows the hardware configuration of the computer that functions as each device related to this disclosure. [Modes for carrying out the invention]

[0010] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining some or all of the technologies (things or methods) employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technologies employed in each of the exemplary embodiments shown below may also be included in the scope of the present invention. In addition, the effects mentioned in each of the exemplary embodiments shown below are examples of effects that can be expected in that exemplary embodiment and do not define the scope of the present invention. That is, embodiments that do not produce the effects mentioned in each of the exemplary embodiments shown below may also be included in the scope of the present invention.

[0011] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form of each of the exemplary embodiments described later. Note that the scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can be adopted in other exemplary embodiments included in the present disclosure as long as there are no particular technical obstacles. Also, each technology shown in the drawings referred to for explaining this exemplary embodiment can be adopted in other exemplary embodiments included in the present disclosure as long as there are no particular technical obstacles.

[0012] (Configuration of Information Processing Apparatus 1) The configuration of the information processing apparatus 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the information processing apparatus 1. As shown in FIG. 1, the information processing apparatus 1 includes an estimation unit 11, a summarization unit 12, and a generation unit 13. The estimation unit 11 is an example of a configuration that realizes estimation means. The summarization unit 12 is an example of a configuration that realizes summarization means. The generation unit 13 is an example of a configuration that realizes generation means.

[0013] The estimation unit 11 estimates, for each of a plurality of topics in the dialogue history, a partial dialogue history related to the topic in the dialogue history. Here, the dialogue history indicates a sequence in which utterances in a dialogue conducted among a plurality of speakers are arranged in the order in which they were uttered. For example, the dialogue history is represented by text data indicating natural language sentences. For example, such text data may be data obtained by converting voice data indicating a dialogue into text data, but is not limited thereto.

[0014] Also, it is desirable that the partial conversation history be generated to include consecutive utterances in the original conversation history. However, the partial conversation history may not require all the utterances constituting the partial conversation history to be consecutive in the original conversation history, and a lower limit number of consecutive utterances may be defined. For example, when utterances 1 to 10 are included in the conversation history in this order and the number of consecutive utterances to be considered is 3, it may be acceptable to estimate that utterances 1 to 3 and 7 to 10 are partial conversation histories related to a certain topic, as they include at least three consecutive utterances.

[0015] For example, the estimation unit 11 may use a large language model to generate partial conversation histories for each of a plurality of topics from the conversation history. In this case, for example, the estimation unit 11 may generate partial conversation histories for each of the plurality of topics by inputting a prompt including the conversation history and generation instructions for the partial conversation histories of each of the plurality of topics to the large language model.

[0016] Also, for example, the estimation unit 11 may estimate which of a plurality of topics each utterance included in the conversation history belongs to, and generate a partial conversation history by grouping together utterances with the same estimated topic. For example, when a plurality of topics are not predefined, a large language model may be used in the process of estimating the topic of each utterance. Also, for example, when a plurality of topics are predefined, a classification model or a large language model may be used in the process of estimating the topic of each utterance. However, the method of generating the partial conversation history is not limited to the examples described above.

[0017] The summarization unit 12 generates a partial summary sentence from the partial conversation history. For example, the summarization unit 12 may use a large language model to generate a partial summary sentence from the partial conversation history. In this case, for example, the summarization unit 12 may generate a partial summary sentence by inputting a prompt including the partial conversation history and its summarization instruction to the large language model. However, the method of generating the partial summary sentence is not limited to the examples described above, and other known methods may be adopted.

[0018] Furthermore, when a large-scale language model is used in either or both of the estimation unit 11 and the summarization unit 12, the large-scale language model may be a general-purpose large-scale language model or a large-scale language model that has been fine-tuned using training data in a field related to dialogue history. Also, when a large-scale language model is used in both the estimation unit 11 and the summarization unit 12, the same large-scale language model may be used, or different large-scale language models may be used.

[0019] The generation unit 13 generates a summary document of the dialogue history based on partial summary sentences for each of the multiple topics. For example, the generation unit 13 may generate a summary document of the dialogue history by combining partial summary sentences for each of the multiple topics. Alternatively, for example, the generation unit 13 may construct the summary document into multiple sections. In this case, the generation unit 13 may place a natural language sentence indicating the topic relevant to that section from among the multiple topics as the title of each section, and place a partial summary sentence about that topic as the content of that section. Furthermore, the generation unit 13 may use a large-scale language model to generate a summary document from partial summary sentences for each of the multiple topics. However, the method for generating the summary document is not limited to the examples described above.

[0020] (Effects of Information Processing Device 1) As described above, the information processing device 1 employs a configuration that includes an estimation unit 11 that estimates the partial dialogue history related to each of the multiple topics in the dialogue history, a summarization unit 12 that generates a partial summary from the partial dialogue history, and a generation unit 13 that generates a summary document of the dialogue history based on the partial summary for each of the multiple topics. Therefore, the information processing device 1 has the effect of being able to generate a more appropriate summary document from the dialogue history because it takes into account multiple topics in the dialogue history.

[0021] (Information processing method S1 flow) The flow of the information processing method S1 will be explained with reference to Figure 2. For example, if the information processing device 1 is equipped with at least one processor, the information processing device 1 executes the information processing method S1. Figure 2 is a flowchart showing the flow of the information processing method S1. As shown in Figure 2, the information processing method S1 includes an estimation process S11, a summarization process S12, and a generation process S13.

[0022] In estimation process S11, at least one processor (e.g., estimation unit 11) estimates the partial dialogue history related to each of the multiple topics in the dialogue history. For example, the details of estimation process S11 will be explained in the same way as the estimation unit 11, so a detailed explanation will not be repeated.

[0023] In summarization processing S12, at least one processor (e.g., summarization unit 12) generates a partial summary from the partial dialogue history. For example, the details of summarization processing S12 are described in the same way as the summarization unit 12, so a detailed explanation will not be repeated.

[0024] In generation process S13, at least one processor (e.g., generation unit 13) generates a summary document of the dialogue history based on partial summary sentences for each of the multiple topics. For example, the details of generation process S13 are described in the same way as those of generation unit 13, so a detailed explanation will not be repeated.

[0025] (Effects of information processing methods) As described above, the information processing method S1 employs a configuration that includes: an estimation process S11 in which at least one processor estimates the partial dialogue history related to each of the multiple topics in the dialogue history; a summarization process S12 in which at least one processor generates a partial summary sentence from the partial dialogue history; and a generation process S13 in which at least one processor generates a summary document of the dialogue history based on the partial summary sentences for each of the multiple topics. Therefore, the same effects as the information processing device 1 can be obtained with the information processing method S1.

[0026] [Second exemplary embodiment] A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same function as those described in the above-described exemplary embodiment are denoted by the same reference numerals, and their descriptions are omitted as appropriate. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems arise. Furthermore, each technology shown in the drawings referenced to describe this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical problems arise.

[0027] (Overview of Information Processing System 100A) Information processing system 100A is a system that generates summary documents in a predetermined format from dialogue history. For example, information processing system 100A estimates multiple topics from the dialogue history and estimates partial dialogue history related to each topic by estimating which of the multiple topics each utterance in the dialogue history relates to. Furthermore, information processing system 100A extracts keywords from the partial dialogue history related to each topic, generates partial summary sentences related to each topic based on the extracted keywords and the partial dialogue history, and generates a summary document in a predetermined format based on each generated partial summary sentence.

[0028] Here, the summary document generated by the information processing system 100A is a document conforming to a predetermined format consisting of multiple sections. The predetermined format indicates, for example, the multiple sections to be included in the summary document. Furthermore, the predetermined format may also indicate the hierarchical structure of the sections, the order of the sections, etc., but is not limited to this. Examples of documents conforming to the predetermined format include medical reports, call center call records, and financial institution counter reception records. For example, a treatment explanation document, which is a type of medical report, may have a predetermined format consisting of multiple sections such as "Disease Name and Condition" and "Purpose of Surgery." Therefore, such a medical report is an example of a document conforming to the predetermined format. Note that documents conforming to the predetermined format are not limited to the examples given above.

[0029] Figure 3 is a schematic diagram illustrating the overview of the information processing system 100A. As shown in Figure 3, multiple topics T1, T2, ... are estimated from the dialogue history using the large-scale language model LLM1. Furthermore, for each utterance that constitutes the dialogue history, the large-scale language model LLM2 is used to estimate which of the multiple topics T1, T2, ... it relates to. By grouping utterances that belong to the same estimated topic, partial dialogue histories for topic T1, partial dialogue histories for topic T2, ... are generated. In addition, one or more keywords are extracted from each partial dialogue history using the named entity recognition model NER1. Furthermore, partial summary sentences are generated from each partial dialogue history and the one or more keywords extracted from it using the large-scale language model LLM3. Finally, a summary document of the entire dialogue history is generated based on the partial summary sentences for each of the multiple topics generated in this way.

[0030] (Configuration of Information Processing System 100A) The configuration of the information processing system 100A will be described with reference to Figure 4. Figure 4 is a block diagram showing the configuration of the information processing system 100A. As shown in Figure 4, the information processing system 100A includes an information processing device 1A, a model storage device 2, an interaction history database 3, an input device 4, and a display device 5. The information processing device 1A is communicated with the model storage device 2, the interaction history database 3, the input device 4, and the display device 5 via a network or peripheral device connection interface. Some or all of the information stored in the model storage device 2 and the interaction history database 3 may also be stored in the storage unit 120 of the information processing device 1A. In addition, one or both of the input device 4 and the display device 5 may be built into the information processing device 1A instead of being connected to it. Furthermore, the input device 4 and the display device 5 may be connected to or built into a user terminal (not shown), and the user terminal may be communicated with the information processing device 1A via a network. Although Figure 4 shows one model storage device 2, one dialogue history database 3, one input device 4, and one display device 5, the information processing system 100A may include multiple instances of some or all of these devices.

[0031] (Model Memory 2) Model memory device 2 stores the large-scale language models LLM1 to LLM3 and the named entity recognition model NER1. The large-scale language models LLM1 to LLM3 are deep learning models generated to perform natural language processing tasks. For example, the large-scale language models LLM1 to LLM3 are models that perform a text generation task, taking natural language sentence prompts as input and outputting generated natural language sentences. The large-scale language models LLM1 to LLM3 may be fine-tuned versions of general-purpose large-scale language models, or they may be general-purpose large-scale language models. If at least one of the large-scale language models LLM1 to LLM3 is a general-purpose large-scale language model, in-context learning may be performed using that large-scale language model. Furthermore, at least two of the large-scale language models LLM1 to LLM3 may be the same model or different models.

[0032] The NER1 named entity recognition model is a model that outputs words or sequences of words and labels as named entities in an input natural language sentence. For example, the NER1 named entity recognition model may be a general-purpose model, or it may be a fine-tuned model that can recognize named entities with labels specific to a field related to dialogue history.

[0033] (Dialogue History Database 3) The dialogue history database 3 stores the dialogue history. The dialogue history is text data represented by natural language sentences that show a conversation between multiple speakers. Alternatively, for example, the dialogue history may be stored as a sequence of text data for each utterance, with speaker identification information associated with each utterance. Alternatively, for example, the dialogue history may be converted from audio data in which a conversation between multiple speakers has been recorded.

[0034] (Input device 4 and display device 5) The input device 4 is configured to receive input to the information processing device 1A, and may include, for example, an input device such as a keyboard, mouse, touch panel, camera, or microphone. The display device 5 is configured to display the screen output from the information processing device 1A, and may include, for example, a display. The input device 4 and the display device 5 may also be integrally formed as a touch panel or the like.

[0035] (Configuration of Information Processing Device 1A) As shown in Figure 4, the information processing device 1A includes a control unit 110 and a storage unit 120. The control unit 110 controls all parts of the information processing device 1A. The storage unit 120 stores various data and programs that the control unit 110 references.

[0036] The control unit 110 includes the estimation unit 11, summarization unit 12, and generation unit 13 of the information processing device 1, as well as a first extraction unit 14. The first extraction unit 14 is an example of a configuration that realizes the first extraction means.

[0037] In addition to being configured similarly to Exemplary Embodiment 1, the estimation unit 11 is configured as follows: The estimation unit 11 estimates multiple topics based on the dialogue history and estimates a partial dialogue history for each of the estimated topics. This allows for a more appropriate estimation of multiple topics for which a partial dialogue history should be generated. Furthermore, the estimation unit 11 estimates which of the multiple topics each utterance included in the dialogue history relates to, and estimates utterances with the same estimated topic as a partial dialogue history. This allows for a more appropriate generation of a partial dialogue history from the dialogue history.

[0038] For example, the estimation unit 11 uses the large-scale language model LLM1 to estimate multiple topics from the dialogue history. For example, the estimation unit 11 may estimate multiple topics to be output by inputting a prompt to the large-scale language model LLM1 that includes the dialogue history and instructions for estimating multiple topics.

[0039] Furthermore, for example, the estimation unit 11 uses the large-scale language model LLM2 to estimate which of the multiple topics each utterance included in the dialogue history is related to. For example, the estimation unit 11 may estimate the topic of an utterance by inputting a prompt to the large-scale language model LLM2 that includes an utterance and an instruction to estimate which of the multiple topics it relates to.

[0040] Alternatively, for example, the estimation unit 11 may generate partial dialogue histories for each of multiple topics by performing a process to add each utterance included in the dialogue history to the partial dialogue history of the estimated topic, in the order in which the utterances are included in the dialogue history.

[0041] The first extraction unit 14 extracts keywords from the partial dialogue history. For example, the first extraction unit 14 may use the named entity recognition model NER1 to extract keywords from the partial dialogue history. In other words, named entities output by inputting the partial dialogue history into the named entity recognition model NER1 are extracted as keywords in that partial dialogue history. The first extraction unit 14 extracts keywords for the partial dialogue history for each of the multiple topics.

[0042] In addition to being configured similarly to Exemplary Embodiment 1, the summarization unit 12 is configured as follows: The summarization unit 12 generates a partial summary sentence based on the partial dialogue history and keywords. For example, the summarization unit 12 may use the large-scale language model LLM3 to generate a partial summary sentence based on the partial dialogue history and keywords. In other words, the text output when a prompt including the partial dialogue history, keywords, and instructions to summarize the partial dialogue history based on those keywords is input to the large-scale language model LLM3 may be obtained as a partial summary sentence. This generates a more appropriate partial summary sentence compared to simply inputting the partial dialogue history into the large-scale language model.

[0043] The generation unit 13 is configured in the same manner as in Exemplary Embodiment 1, and in addition, it is configured as follows: The generation unit 13 generates a summary document by referring to relational information that shows the relationship between at least one of a plurality of topics and at least one of a plurality of sections. The plurality of sections are the plurality of sections that should be included in the summary document, as indicated by a predetermined format defined for the summary document. For example, the relational information may be information in which one or more topics from a plurality of topics are associated with each of the plurality of sections. For example, the relationship between sections and topics is not limited to one-to-one relationships; a single section may be associated with multiple topics, or a single topic may be associated with multiple sections.

[0044] For example, the generation unit 13 may include a title defined for each section and one or more partial summaries of topics associated with that section in the summary document as the section itself. This makes it possible to generate an appropriate summary document in a predetermined format.

[0045] (Information processing method S1A flow) The information processing system 100A, configured as described above, executes the information processing method S1A. Figure 5 is a flowchart showing the flow of the information processing method S1A. As shown in Figure 5, the information processing method S1A includes steps S101 to S107.

[0046] In step S101, the control unit 110 of the information processing device 1A acquires the dialogue history from the dialogue history database 3.

[0047] Steps S102 to S104 are an example of the estimation process. In step S102, the estimation unit 11 uses the large-scale language model LLM1 to estimate multiple topics in the dialogue history.

[0048] Steps S103 to S104 are processes that are executed for each utterance included in the dialogue history, in the order in which they appear in the dialogue history.

[0049] In step S103, the estimation unit 11 uses the large-scale language model LLM2 to estimate which of the multiple topics estimated in step S102 is related to the utterance in question.

[0050] In step S104, the estimation unit 11 adds the utterance to the partial dialogue history for the estimated topic.

[0051] Once steps S103-S104 are completed for all utterances included in the dialogue history, the next steps S105-S106 are executed. Steps S105-S106 are processes that are executed for each of the multiple partial dialogue histories.

[0052] In step S105, the first extraction unit 14 uses the named entity recognition model NER1 to extract keywords from the relevant partial dialogue history.

[0053] Step S106 is an example of summarization processing. The summarization unit 12 uses the large-scale language model LLM2 to generate a partial summary based on the partial dialogue history and the keywords.

[0054] Once steps S105 to S106 are completed for all partial dialogue history, the next step S107 is executed.

[0055] Step S107 is an example of the generation process. In step S107, the generation unit 13 refers to the relevant information and generates a summary document based on multiple partial dialogue histories.

[0056] This concludes information processing method S1A.

[0057] (Examples of application) As an example of the application of the information processing system 100A, we will describe an example where the summary document is a medical report. We will also describe an example where the prescribed format is a format defined as a treatment explanation document. In this application example, in step S101, the history of conversations between the doctor, the patient, and their family is acquired.

[0058] In steps S102 to S104, multiple partial dialogue histories are estimated from the dialogue history. Figure 6 schematically shows the estimation process in this application example. As shown in Figure 6, in this application example, in step S102, eight topics present in the dialogue history are estimated: T1 "Disease name / condition", T2 "Treatment purpose / alternative treatment", T3 "Surgical procedure", T4 "Post-surgery to discharge", T5 "Complications", T6 "Withdrawal of consent / SO", T7 "Question", and T8 "Answer". Note that the estimation process for topics T1 to T8 may involve pre-defined topics or topics that are not pre-defined.

[0059] As shown in Figure 6, the dialogue history includes utterances Q1 to Q9, ... in this order. Each utterance is associated with a speaker such as "Doctor A," "Patient B," "Patient's Family C," etc. In step S103, it is estimated which of topics T1 to T8 each utterance Q1 to Q9 relates to. In a table where the topic is the column item TS and the utterance is the row item QS, a value of 1 in a cell indicates that the topic has been estimated for that utterance. For example, utterances Q1 and Q2, "XX disease is a type of YY," are estimated to be related to topic T1, "Disease name / pathology." For example, utterances Q7 and Q8, "MRI and ultrasound, etc., blood tests, etc.," are estimated to be related to topic T3, "Surgical procedures." Also, utterance Q9, "Are there any other treatment options?", is estimated to be related to topic T2, "Treatment objectives / alternative treatments." Furthermore, utterances Q3-Q6, for which no topic has been inferred, may be included in the partial dialogue history for the same topic as the preceding or succeeding topic, or they may not be included in any partial dialogue history.

[0060] In step S104, utterances Q1 to Q9, ... are each added to the partial dialogue history for the estimated topic among topics T1 to T8. This generates a partial dialogue history for each of topics T1 to T8.

[0061] In step S105, the first extraction unit 14 extracts keywords from the partial dialogue history for each of the topics T1 to T8. Figure 7 schematically shows the first extraction process in this application example. In Figure 7, the partial dialogue history 71 shows, for example, the partial dialogue history for topic T5 "complications". The keyword 72 is extracted as a named entity included in the partial dialogue history 71 by the named entity recognition model NER1. The named entity recognition model NER1 is a model that has been fine-tuned using training data in the medical field. As training data, for example, examples of dialogue history between doctors and patients and their families that have been accumulated in the past may be used.

[0062] In step S106, for each of topics T1 to T8, a partial summary is generated based on the partial dialogue history 71 and keywords 72.

[0063] In step S107, relevant information is referenced to generate the medical report. Figure 8 schematically shows an example of relevant information in this application. As shown in Figure 8, in this application, the medical report as a summary document is specified to include the following sections in the format of a treatment explanation document: Section P1 "Your Diagnosis and Condition", Section P2 "Purpose, Necessity, Effectiveness and Alternative Treatments of Surgery", Section P3 "Details and Precautions of Surgery", Section P4 "Specific Wishes of the Patient", Section P5 "Handling of Removed Organs", Section P6 "Other Options", and Section P7 "If You Withdraw Your Consent to Treatment".

[0064] In Figure 8, in a table where topics are represented by the column item TS and sections by the row item PS, a value of 1 in a cell indicates that the topic is associated with that section. For example, section P1, "Your Diagnosis and Condition," is associated with topic T1, "Diagnosis and Condition." Thus, the relationship between sections and topics may be one-to-one. Also, for example, section P3, "Surgical Procedure and Precautions," is associated with topics T3, "Surgical Procedure," T4, "Post-Surgery to Discharge," and T5, "Complications." Thus, the relationship between sections and topics may be one-to-many. Furthermore, for example, topic T2, "Treatment Objectives and Alternative Therapies," is associated with all of the following: section P2, "Surgical Objectives, Necessity, Effectiveness, and Alternative Therapies," section P6, "Other Options," and section P7, "Withdrawal of Treatment Consent." Thus, the relationship between sections and topics may be many-to-one.

[0065] This relationship information may be predetermined or estimated using an estimation model. For example, such an estimation model may be a machine learning model that takes each of the topics T1 to T8 as input and outputs the relevant sections from sections P1 to P7.

[0066] Furthermore, in step S107, a medical report is generated based on the partial summaries of topics T1 to T8, referencing the aforementioned relationship information. Figure 9 schematically shows an example of a medical report in this application example.

[0067] In Figure 9, medical report D1 includes multiple sections P1, P2, P3, ... Section P1 includes the title "Your Diagnosis and Condition" and a partial summary of topic T1 "Diagnosis and Condition" associated with section P1 in the related information. Section P2 includes the title "Purpose, Necessity, Effectiveness, and Alternative Treatments of Surgery" and a partial summary of topic T2 "Treatment Objectives and Alternative Treatments" associated with section P2 in the related information. Section P3 includes the title "Details of the Surgery and Precautions" and partial summaries of topics T3 "Details of the Surgery," T4 "Post-Surgery to Discharge," and T5 "Complications" associated with section P3 in the related information.

[0068] Thus, in this application example, a medical report in a predetermined format, such as a treatment explanation document, can be generated from the history of conversations between the doctor, the patient, and their family.

[0069] (Effects of Information Processing System 100A) As described above, in the information processing system 100A, the summary document is a document conforming to a predetermined format consisting of multiple sections, and the generation unit 13 is configured to generate the summary document by referring to relational information that shows the relationship between at least one of the multiple topics and at least one of the multiple sections. Therefore, in addition to the effects performed by the information processing device 1, the information processing system 100A can also provide the effect of generating a summary document conforming to a predetermined format from the dialogue history.

[0070] Furthermore, the information processing system 100A is further equipped with a first extraction unit 14 that extracts keywords from partial dialogue history, and the summarization unit 12 generates partial summary sentences based on the partial dialogue history and keywords. As a result, in addition to the effects achieved by the information processing device 1, the information processing system 100A can generate partial summary sentences that more appropriately summarize the partial dialogue history for each topic.

[0071] Furthermore, in the information processing system 100A, the estimation unit 11 is configured to estimate multiple topics based on the dialogue history and to estimate a partial dialogue history for each of the estimated topics. Therefore, in addition to the effects achieved by the information processing device 1, the information processing system 100A can generate a more appropriate summary document by taking into account the topics that actually exist in the dialogue history.

[0072] Furthermore, in the information processing system 100A, the estimation unit 11 estimates which of several topics each utterance included in the dialogue history relates to, and estimates utterances of the same estimated topic as partial dialogue history. Therefore, according to the information processing system 100A, in addition to the effects performed by the information processing device, the effect of being able to generate partial dialogue history related to each topic more appropriately from the dialogue history can be obtained.

[0073] [Third Exemplary Embodiment] A third exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same function as those described in the above-described exemplary embodiments are denoted by the same reference numerals, and their descriptions are omitted as appropriate. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technology shown in the drawings referenced to describe this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical hindrance occurs.

[0074] (Overview of Information Processing System 100B) Information processing system 100B is a modified version of information processing system 100A that estimates partial dialogue history with greater accuracy. Figure 10 is a schematic diagram showing the overview of information processing system 100B. Figure 10 is almost the same as the overview of information processing system 100A shown in Figure 3, but differs in that, in addition to dialogue history, keywords extracted by the named entity recognition model NER2 are input to the large-scale language model LLM1 for estimating multiple topics T1, T2, ... Also differs in that, in addition to utterances, keywords extracted by the named entity recognition model NER3 are input to the large-scale language model LLM2 for estimating topics related to utterances. Other points are as explained in Figure 3, so a detailed explanation will not be repeated.

[0075] (Configuration of Information Processing System 100B) The configuration of the information processing system 100B will be explained with reference to Figure 11. Figure 11 is a block diagram showing the configuration of the information processing system 100B. As shown in Figure 11, the information processing system 100B is configured in almost the same way as the information processing system 100A shown in Figure 4, but differs in that it includes an information processing device 1B instead of an information processing device 1A. It also differs in that the model storage device 2 further stores named entity recognition models NER2 and NER3. Details of named entity recognition models NER2 and NER3 will be explained in the same way as named entity recognition model NER1. At least two of the named entity recognition models NER1, NER2, and NER3 may be different models or the same model. If they are different models, the at least two models may be models that have been fine-tuned using the same general-purpose named entity recognition model with different training data, or different general-purpose named entity recognition models may be used.

[0076] The information processing device 1B has the same configuration as the information processing device 1A, but further includes a second extraction unit 15 and a third extraction unit 16 in the control unit 110. The second extraction unit 15 is an example of a configuration that realizes a second extraction means. The third extraction unit 16 is an example of a configuration that realizes a third extraction means.

[0077] The second extraction unit 15 extracts keywords from the dialogue history. For example, the second extraction unit 15 may use the named entity recognition model NER2 to extract keywords from the dialogue history. In other words, named entities output by inputting the dialogue history into the named entity recognition model NER2 are extracted as keywords in that dialogue history.

[0078] The third extraction unit 16 extracts keywords from each utterance included in the dialogue history. For example, the third extraction unit 16 may use the named entity recognition model NER3 to extract keywords from each utterance. In other words, the named entity output by inputting a certain utterance into the named entity recognition model NER3 is extracted as the keyword in that utterance.

[0079] The estimation unit 11 is configured in the same manner as in the exemplary embodiment 2, and in addition, it is configured as follows: The estimation unit 11 estimates multiple topics in the dialogue history based on the dialogue history and keywords extracted from the dialogue history. For example, the estimation unit 11 may estimate multiple topics to be output by inputting a prompt to the large-scale language model LLM1 that includes the dialogue history, keywords extracted from the dialogue history, and instructions for estimating multiple topics.

[0080] Furthermore, the estimation unit 11 estimates which of the multiple topics an utterance relates to, based on each utterance included in the dialogue history and the keywords extracted from that utterance. For example, the estimation unit 11 may estimate the topic of an utterance by inputting a prompt to the large-scale language model LLM2 that includes an utterance, keywords extracted from that utterance, and an instruction to estimate which of the multiple topics it relates to.

[0081] (Information processing method S1B flow) The information processing system 100B, configured as described above, executes the information processing method S1B. Figure 12 is a flowchart showing the flow of the information processing method S1B. As shown in Figure 12, the information processing method S1B includes almost the same steps as the information processing method S1A shown in Figure 5, but steps S102B-1 and S102B-2 are added instead of step S102, and steps S103B-1 and S103B-2 are added instead of step S103. The other points are as explained in Figure 5, so a detailed explanation will not be repeated.

[0082] Step S102B-1 is an example of the second extraction process. In step S102B-1, the second extraction unit 15 extracts keywords from the dialogue history using the named entity recognition model NER2.

[0083] In step S102B-2, the estimation unit 11 uses the large-scale language model LLM1 to estimate multiple topics based on the dialogue history and keywords extracted from the dialogue history.

[0084] Step S103B-1 is an example of the third extraction process. In step S103B-1, the third extraction unit 16 uses the named entity recognition model NER3 to extract keywords from each utterance included in the dialogue history.

[0085] In step S103B-2, the estimation unit 11 uses the large-scale language model LLM2 to estimate the topic related to each utterance based on the utterance and the keywords extracted from it.

[0086] (Effects of Information Processing System 100B) As described above, the information processing system 100B further includes a second extraction unit 15 that extracts keywords from the dialogue history, and the estimation unit 11 is configured to estimate multiple topics based on the dialogue history and the keywords. Therefore, in addition to the effects achieved by the information processing system 100A, the information processing system 100B can more appropriately estimate multiple topics for generating a summary document.

[0087] Furthermore, the information processing system 100B further includes a third extraction unit 16 that extracts keywords from each utterance included in the dialogue history, and the estimation unit 11 is configured to estimate which of multiple topics the utterance relates to based on the utterance and the keyword. Therefore, in addition to the effects achieved by the information processing system 100A, the information processing system 100B can more appropriately estimate partial dialogue histories that are suitable for each of the multiple topics for generating a summary document.

[0088] [Fourth exemplary embodiment] A fourth exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same function as those described in the above-described exemplary embodiments are denoted by the same reference numerals, and their descriptions are omitted as appropriate. The scope of application of each technology adopted in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technology adopted in this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technology shown in the drawings referenced to describe this exemplary embodiment can also be adopted in other exemplary embodiments included in this disclosure, to the extent that no particular technical hindrance occurs.

[0089] (Overview of Information Processing System 100C) Information processing system 100C is a modified version of information processing system 100A that generates a summary document from a dialogue history acquired based on voice data input via a voice input device.

[0090] (Configuration of Information Processing System 100C) The configuration of the information processing system 100C will be explained with reference to Figure 13. Figure 13 is a block diagram showing the configuration of the information processing system 100C. As shown in Figure 13, the information processing system 100C is configured in almost the same way as the information processing system 100A shown in Figure 4, but differs in that it includes an information processing device 1C instead of an information processing device 1A, and also includes a voice input device 6. Furthermore, unlike the information processing system 100A, the information processing system 100C does not necessarily have to include a dialogue history database 3.

[0091] The voice input device 6 may be, for example, a microphone. The voice input device 6 may also be connected to the information processing device 1C via an input / output interface, or it may be built into the device. Furthermore, the voice input device 6 may be connected to or built into a user terminal (not shown), and the user terminal may be connected to the information processing device 1C via a network for communication.

[0092] The information processing device 1C has the same configuration as the information processing device 1A, but in addition, the control unit 110 further includes an acquisition unit 17 and a display control unit 18. The acquisition unit 17 is an example of a configuration that realizes an acquisition means. The display control unit 18 is an example of a configuration that realizes a display control means.

[0093] The acquisition unit 17 acquires the dialogue history based on the voice data input via the voice input device 6. For example, the acquisition unit 17 may acquire text data converted from the voice data using speech recognition technology as the dialogue history. Alternatively, for example, the acquisition unit 17 may acquire a sequence of text data in utterance units as the dialogue history using technology to identify the speaker in the voice data. Furthermore, for example, the acquisition unit 17 may update the dialogue history based on voice data continuously input for an ongoing dialogue. In other words, the acquisition unit 17 may acquire the dialogue history in real time.

[0094] For example, the estimation unit 11, summarization unit 12, generation unit 13, and first extraction unit 14 may update the summary document by functioning again when the dialogue history is updated. For example, the estimation unit 11, summarization unit 12, generation unit 13, and first extraction unit 14 may function at predetermined intervals (e.g., every minute), whenever the dialogue history increases by a predetermined amount, or whenever a new topic is added to the dialogue history.

[0095] The display control unit 18 displays the summary document on the display device 5. For example, if the summary document is updated in response to an update in the dialogue history, the display control unit 18 may display the updated summary document on the display device 5.

[0096] (Examples of application) The information processing system 100C can be used to generate medical reports (an example of a summary document) in real time during conversations between a doctor and a patient and their family. In this application example, for example, a doctor can check the medical report displayed on the display device 5 while conversing with the patient and their family. Furthermore, if the displayed medical report is missing necessary sections, the doctor can continue the conversation by bringing up topics related to the missing sections.

[0097] (Effects of Information Processing System 100C) As described above, the information processing system 100C further includes an acquisition unit 17 that acquires dialogue history based on voice data input via the voice input device 6, and a display control unit 18 that displays the summary document on the display device 5. Therefore, in addition to the effects of the information processing system 100A, the information processing system 100C provides the effect that the user can generate an appropriate summary document by having a dialogue between multiple speakers as input into the voice input device 6. Furthermore, when the acquisition of dialogue history and the generation of the summary document are performed in real time, at least one of the multiple speakers can check the summary document while continuing the dialogue. Furthermore, at least one of the multiple speakers can continue the dialogue so that the displayed summary document approaches the desired content.

[0098] [Variation] The information processing system 100C according to the exemplary embodiment 4 described above may include an information processing device 1 or 1B modified to include an acquisition unit 17 and a display control unit 18 instead of the information processing device 1C.

[0099] Furthermore, the information processing device 1B according to the exemplary embodiment 3 described above does not necessarily have to include all of the first extraction unit 14, the second extraction unit 15, and the third extraction unit 16, and may be modified to include at least one of them.

[0100] For example, the information processing device 1B may include a first extraction unit 14 and a second extraction unit 15, but may not include a third extraction unit 16. In this case, the estimation unit 11 estimates the topic related to each utterance included in the dialogue history by inputting it into the large-scale language model LLM2.

[0101] Furthermore, for example, the information processing device 1B may include a first extraction unit 14 and a third extraction unit 16, but may not include a second extraction unit 15. In this case, the estimation unit 11 estimates multiple topics in the dialogue history by inputting the dialogue history into the large-scale language model LLM1.

[0102] Furthermore, for example, the information processing device 1B may include a third extraction unit 16 but not the first extraction unit 14 and the second extraction unit 15. In this case, the estimation unit 11 estimates multiple topics in the dialogue history by inputting the dialogue history into the large-scale language model LLM1. The summarization unit 12 generates a partial summary sentence by inputting the partial dialogue history into the large-scale language model LLM3.

[0103] Furthermore, in the exemplary embodiments 2 to 4 described above, the summary document does not necessarily have to be a document conforming to a predetermined format. Also, each exemplary embodiment can be applied to generate summary documents from dialogues between service providers and service recipients, not limited to the medical field. Such fields include, but are not limited to, call centers and financial institutions.

[0104] [Examples of implementation using software] Some or all of the functions of the devices constituting the information processing device 1 and the information processing systems 100A, 100B, and 100C (hereinafter also referred to as "the above devices") may be implemented by hardware such as integrated circuits (IC chips) or by software.

[0105] In the latter case, each of the above devices is implemented, for example, by a computer that executes instructions for a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as Computer C) is shown in Figure 14. Figure 14 is a block diagram showing the hardware configuration of Computer C, which functions as each of the above devices.

[0106] Computer C comprises at least one processor C1 and at least one memory C2. Memory C2 stores a program P that causes computer C to operate as each of the above-mentioned devices. In computer C, processor C1 reads program P from memory C2 and executes it, thereby realizing each of the above-mentioned devices.

[0107] For processor C1, for example, a CPU (Central Processing Unit), GPU (Graphic Processing Unit), DSP (Digital Signal Processor), MPU (Micro Processing Unit), FPU (Floating Point Number Processing Unit), PPU (Physics Processing Unit), TPU (Tensor Processing Unit), quantum processor, microcontroller, or a combination thereof can be used. For memory C2, for example, flash memory, HDD (Hard Disk Drive), SSD (Solid State Drive), or a combination thereof can be used.

[0108] Computer C may also be equipped with RAM (Random Access Memory) for loading program P at runtime and for temporarily storing various data. Furthermore, computer C may be equipped with communication interfaces for sending and receiving data with other devices. Additionally, computer C may be equipped with input / output interfaces for connecting input / output devices such as keyboards, mice, displays, and printers.

[0109] Furthermore, program P can be recorded on a non-temporary, tangible recording medium M that is readable by computer C. Such a recording medium M could be, for example, tape, disk, card, semiconductor memory, or programmable logic circuitry. Computer C can acquire program P via such a recording medium M. Program P can also be transmitted via a transmission medium. Such a transmission medium could be, for example, a communication network or broadcast waves. Computer C can also acquire program P via such a transmission medium.

[0110] Furthermore, each of the above functions of each of the above devices may be implemented by a single processor in a single computer, by multiple processors in a single computer working together, or by multiple processors in each of multiple computers working together. In addition, the programs for implementing each of the above functions in each of the above devices may be stored in a single memory in a single computer, distributed and stored in multiple memories in a single computer, or distributed and stored in multiple memories in each of multiple computers.

[0111] [Additional Note A] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0112] (Note A1) For each of the multiple topics in the dialogue history, estimation means for estimating the partial dialogue history related to that topic from the dialogue history, A summarization means for generating a partial summary from the aforementioned partial dialogue history, A generation means for generating a summary document of the dialogue history based on the partial summary sentences for each of the aforementioned multiple topics, It is equipped with Information processing device.

[0113] (Appendix A2) The aforementioned summary document is a document in a predetermined format consisting of multiple sections, The generation means generates the summary document by referring to relational information that shows the relationships between at least one of the plurality of topics and at least one of the plurality of sections. The information processing device described in Appendix A1.

[0114] (Note A3) The system further comprises a first extraction means for extracting keywords from the aforementioned partial dialogue history, The summarization means generates the partial summary sentence based on the partial dialogue history and the keywords. The information processing device described in Appendix A1 or A2.

[0115] (Note A4) The estimation means estimates the plurality of topics based on the dialogue history, and estimates the partial dialogue history for each of the estimated plurality of topics. An information processing device as described in any one of the appendices A1 to A3.

[0116] (Note A5) The system further comprises a second extraction means for extracting keywords from the aforementioned dialogue history, The estimation means estimates the plurality of topics based on the dialogue history and the keywords. The information processing device described in Appendix A4.

[0117] (Note A6) The estimation means estimates which of the multiple topics each utterance included in the dialogue history relates to, and estimates utterances of the same estimated topic as the partial dialogue history. An information processing device as described in any one of the appendices A1 to A5.

[0118] (Note A7) The system further comprises a third extraction means for extracting keywords from each of the aforementioned utterances, The estimation means estimates which of the multiple topics the utterance relates to, based on the utterance and the keywords. The information processing device described in Appendix A6.

[0119] (Note A8) An acquisition means for acquiring the dialogue history based on voice data input via a voice input device, Display control means for displaying the summary document on a display device, An information processing device described in any one of the appendices A1 to A7, further comprising the above.

[0120] [Additional Note B] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0121] (Note B1) At least one processor performs an estimation process to estimate, for each of several topics in the dialogue history, a portion of the dialogue history related to that topic. The at least one processor performs a summarization process that generates a partial summary sentence from the partial dialogue history, The at least one processor performs a generation process to generate a summary document of the dialogue history based on the partial summary sentences for each of the plurality of topics, Includes, Information processing methods.

[0122] (Note B2) The aforementioned summary document is a document in a predetermined format consisting of multiple sections, In the generation process, the at least one processor generates the summary document by referring to relational information that shows the relationship between at least one of the plurality of topics and at least one of the plurality of sections. The information processing method described in Appendix B1.

[0123] (Note B3) The at least one processor further includes a first extraction process for extracting keywords from the partial dialogue history, In the summarization process, the at least one processor generates the partial summary sentence based on the partial dialogue history and the keywords. The information processing method described in Appendix B1 or B2.

[0124] (Note B4) In the estimation process, the at least one processor estimates the plurality of topics based on the dialogue history, and estimates the partial dialogue history for each of the estimated plurality of topics. The information processing method described in any one of the appendices B1 to B3.

[0125] (Note B5) The at least one processor further includes a second extraction process for extracting keywords from the dialogue history, In the estimation process, the at least one processor estimates the plurality of topics based on the dialogue history and the keywords. The information processing method described in Appendix B4.

[0126] (Note B6) In the estimation process, the at least one processor estimates which of the multiple topics each utterance included in the dialogue history relates to, and estimates utterances whose estimated topics are the same as the partial dialogue history. The information processing method described in any one of the appendices B1 through B5.

[0127] (Note B7) The at least one processor further includes a third extraction process for extracting keywords from each of the utterances, In the estimation process, the at least one processor estimates which of the plurality of topics the utterance relates to based on the utterance and the keywords. The information processing method described in Appendix B6.

[0128] (Note B8) The at least one processor performs an acquisition process to acquire the dialogue history based on voice data input via a voice input device, The at least one processor performs display control processing to display the summary document on a display device, An information processing method described in any one of the appendices B1 to B7, which further includes the above.

[0129] [Additional Note C] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0130] (Note C1) A program that makes a computer function as an information processing device. The aforementioned computer, For each of the multiple topics in the dialogue history, estimation means for estimating the partial dialogue history related to that topic from the dialogue history, A summarization means for generating a partial summary from the aforementioned partial dialogue history, A generation means for generating a summary document of the dialogue history based on the partial summary sentences for each of the aforementioned multiple topics, To make it function as Information processing program.

[0131] (Note C2) The aforementioned summary document is a document in a predetermined format consisting of multiple sections, The generation means generates the summary document by referring to relational information that shows the relationships between at least one of the plurality of topics and at least one of the plurality of sections. The information processing program described in Appendix C1.

[0132] (Note C3) The aforementioned computer, This is further configured to function as a first extraction means for extracting keywords from the aforementioned partial dialogue history. The summarization means generates the partial summary sentence based on the partial dialogue history and the keywords. The information processing program described in Appendix C1 or C2.

[0133] (Note C4) The estimation means estimates the plurality of topics based on the dialogue history, and estimates the partial dialogue history for each of the estimated plurality of topics. An information processing program described in any one of the appendices C1 to C3.

[0134] (Note C5) The aforementioned computer, This further functions as a second extraction means for extracting keywords from the aforementioned dialogue history. The estimation means estimates the plurality of topics based on the dialogue history and the keywords. The information processing program described in Appendix C4.

[0135] (Appendix C6) The estimation means estimates which of the multiple topics each utterance included in the dialogue history relates to, and estimates utterances of the same estimated topic as the partial dialogue history. An information processing program described in any one of the appendices C1 to C5.

[0136] (Note C7) The aforementioned computer, This further functions as a third extraction means for extracting keywords from each of the aforementioned utterances. The estimation means estimates which of the multiple topics the utterance relates to, based on the utterance and the keywords. The information processing program described in Appendix C6.

[0137] (Note C8) The aforementioned computer, An acquisition means for acquiring the dialogue history based on voice data input via a voice input device, Display control means for displaying the summary document on a display device, An information processing program described in any one of the appendices C1 to C7 to further enable this function.

[0138] [Additional Note D] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0139] (Note D1) It comprises at least one processor, and the at least one processor is For each of the multiple topics in the dialogue history, an estimation process is performed to estimate the partial dialogue history related to that topic from the dialogue history, A summarization process that generates a partial summary from the aforementioned partial dialogue history, A generation process that generates a summary document of the dialogue history based on the partial summary sentences for each of the aforementioned multiple topics, Execute Information processing device.

[0140] The information processing device may further include memory. The memory may also store a program for causing at least one processor to perform each of the aforementioned processes.

[0141] (Note D2) The aforementioned summary document is a document in a predetermined format consisting of multiple sections, In the generation process, the at least one processor generates the summary document by referring to relational information that shows the relationship between at least one of the plurality of topics and at least one of the plurality of sections. The information processing device described in Appendix D1.

[0142] (Note D3) The aforementioned at least one processor, A first extraction process is further performed to extract keywords from the aforementioned partial dialogue history. In the summarization process, the at least one processor generates the partial summary sentence based on the partial dialogue history and the keywords. The information processing device described in Appendix D1 or D2.

[0143] (Note D4) In the estimation process, the at least one processor estimates the plurality of topics based on the dialogue history, and estimates the partial dialogue history for each of the estimated plurality of topics. An information processing device as described in any one of the appendices D1 to D3.

[0144] (Note D5) The aforementioned at least one processor, A second extraction process is then performed to extract keywords from the aforementioned dialogue history. In the estimation process, the at least one processor estimates the plurality of topics based on the dialogue history and the keywords. The information processing device described in Appendix D4.

[0145] (Note D6) In the estimation process, the at least one processor estimates which of the multiple topics each utterance included in the dialogue history relates to, and estimates utterances whose estimated topics are the same as the partial dialogue history. An information processing device as described in any one of the appendices D1 to D5.

[0146] (Note D7) The aforementioned at least one processor, A third extraction process is then performed to extract keywords from each of the aforementioned utterances. In the estimation process, the at least one processor estimates which of the plurality of topics the utterance relates to based on the utterance and the keywords. The information processing device described in Appendix D6.

[0147] (Note D8) The aforementioned at least one processor, An acquisition process that acquires the dialogue history based on voice data input via a voice input device, Display control processing for displaying the summary document on a display device, An information processing device described in any one of the appendices D1 to D7 that further performs the following.

[0148] [Additional Note E] This disclosure includes the technologies described in the following appendices. However, the present invention is not limited to the technologies described in the following appendices, and various modifications are possible within the scope of the claims.

[0149] (Note E1) A program that makes a computer function as an information processing device. To the aforementioned computer, For each of the multiple topics in the dialogue history, an estimation process is performed to estimate the partial dialogue history related to that topic from the dialogue history, A summarization process that generates a partial summary from the aforementioned partial dialogue history, A generation process that generates a summary document of the dialogue history based on the partial summary sentences for each of the aforementioned multiple topics, To execute A non-temporary recording medium on which an information processing program is stored. [Explanation of symbols]

[0150] 100A, 100B, 100C Information Processing Systems 1, 1A, 1B, 1C Information Processing Devices Q1, Q3, Q7, Q9 Utterances 2 Model Storage Devices 3. Dialogue History Database 4 Input devices 5 Display device 6. Voice input device 11 Estimation part 12 Summary Section 13 Generation part 14 1st extraction part 15 Second extraction part 16 Third extraction part 17 Acquisition Department 18 Display Control Unit 110 Control Unit 120 Storage section C1 Processor C2 Memory

Claims

1. For each of the multiple topics in the dialogue history, estimation means for estimating the partial dialogue history related to that topic from the dialogue history, A summarization means for generating a partial summary from the aforementioned partial dialogue history, A generation means for generating a summary document of the dialogue history based on the partial summary sentences for each of the aforementioned multiple topics, It is equipped with Information processing device.

2. The aforementioned summary document is a document in a predetermined format consisting of multiple sections, The generation means generates the summary document by referring to relational information that shows the relationships between at least one of the plurality of topics and at least one of the plurality of sections. The information processing apparatus according to claim 1.

3. The system further comprises a first extraction means for extracting keywords from the aforementioned partial dialogue history, The summarization means generates the partial summary sentence based on the partial dialogue history and the keywords. The information processing apparatus according to claim 1 or 2.

4. The estimation means estimates the plurality of topics based on the dialogue history, and estimates the partial dialogue history for each of the estimated plurality of topics. The information processing apparatus according to claim 1 or 2.

5. The system further comprises a second extraction means for extracting keywords from the aforementioned dialogue history, The estimation means estimates the plurality of topics based on the dialogue history and the keywords. The information processing apparatus according to claim 4.

6. The estimation means estimates which of the multiple topics each utterance included in the dialogue history relates to, and estimates utterances of the same estimated topic as the partial dialogue history. The information processing apparatus according to claim 1 or 2.

7. The system further comprises a third extraction means for extracting keywords from each of the aforementioned utterances, The estimation means estimates which of the multiple topics the utterance relates to, based on the utterance and the keywords. The information processing apparatus according to claim 6.

8. An acquisition means for acquiring the dialogue history based on voice data input via a voice input device, Display control means for displaying the summary document on a display device, The information processing apparatus according to claim 1 or 2, further comprising:

9. At least one processor performs an estimation process to estimate, for each of several topics in the dialogue history, a partial dialogue history from the dialogue history that is relevant to that topic. The at least one processor performs a summarization process that generates a partial summary sentence from the partial dialogue history, The at least one processor performs a generation process to generate a summary document of the dialogue history based on the partial summary sentences for each of the plurality of topics, Includes, Information processing methods.

10. An information processing program that enables a computer to function as an information processing device, The aforementioned computer, For each of the multiple topics in the dialogue history, estimation means for estimating the partial dialogue history related to that topic from the dialogue history, A summarization means for generating a partial summary from the aforementioned partial dialogue history, A generation means for generating a summary document of the dialogue history based on the partial summary sentences for each of the aforementioned multiple topics, To make it function as Information processing program.