Document generation device, document generation method, and document generation program
Patent Information
- Application Number
- PCT/JP2023/039123
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-10-30
- Publication Date
- 2025-05-08
AI Technical Summary
It is difficult to generate easy-to-understand documents in the prior art, especially when the readers of the documents have different professional backgrounds and knowledge levels.
By obtaining the attributes of the document recipient, extracting the keyword sets in the document, and using a large-scale language model to generate easy-to-understand documents. The system includes steps such as attribute acquisition, keyword collection acquisition, prompt generation and document generation.
It realizes the generation of easy-to-understand documents based on the attributes of the recipient, improves the readability and understanding of the documents, and is suitable for readers of different professional backgrounds.
Smart Images

Figure JP2023039123_08052025_PF_FP_ABST
Abstract
Description
Document generation device, document generation method, and document generation program
[0001] The present disclosure relates to document generation techniques.
[0002] One application involves generating a target document by referencing a document containing an information source. One example of such an application is generating a referral letter (hereinafter also referred to as a medical information report) for referring a patient to another medical institution by referencing the patient's medical records. It is also conceivable to insert annotations to make the generated medical information report easier to understand. For example, Patent Literature 1 describes a technology for recognizing disease names in text and inserting annotations corresponding to the recognized disease names into the text.
[0003] Japan Special Table No. 2021-522569
[0004] Here, the target document to be generated by referencing a document including an information source is required to be easy to understand for the reader. For example, the ease of understanding of the same document may vary depending on the reader's expertise, experience, knowledge, etc. In the example of the medical information report mentioned above, the ease of understanding of the same medical information report may vary depending on whether the reader is a doctor, a nurse, a care manager, etc.
[0005] However, the technology described in Patent Document 1 has the problem that annotations are inserted into text regardless of the reader, so even if the text has annotations inserted, it may not be easy to understand depending on the reader.
[0006] The present disclosure has been made in consideration of the above-mentioned problems, and one exemplary purpose thereof is to provide a technology for generating a target document that is easy to understand for a reader by referring to a document including an information source.
[0007] A document generation device according to an exemplary aspect of the present disclosure includes an attribute acquisition means for acquiring attributes of a recipient of a second document, which is a document to be generated by referring to a first document; a keyword set acquisition means for acquiring a set of keywords included in the first document; a prompt generation means for generating examples of the keyword set, examples of the second document corresponding to the examples of the keyword set and which correspond to the attributes of the recipient, and a prompt including the set of keywords; and a document generation means for generating the second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
[0008] A document generation method according to an exemplary aspect of the present disclosure includes: an attribute acquisition process in which at least one processor acquires attributes of a recipient of a second document, which is a document to be generated by referring to a first document; a keyword set acquisition process in which the at least one processor acquires a set of keywords included in the first document; a prompt generation process in which the at least one processor generates examples of the keyword set, examples of the second document corresponding to the examples of the keyword set, the examples of the second document corresponding to the attributes of the recipient, and a prompt including the set of keywords; and a document generation process in which the at least one processor generates the second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
[0009] A document generation program according to an exemplary aspect of the present disclosure is a program that causes a computer to function as a document generation device, and causes the computer to function as an attribute acquisition means that acquires attributes of a recipient of a second document that is a document to be generated by referencing a first document; a keyword set acquisition means that acquires a set of keywords contained in the first document; a prompt generation means that generates examples of the keyword set, examples of the second document that correspond to the examples of the keyword set and that correspond to the attributes of the recipient, and a prompt that includes the set of keywords; and a document generation means that generates the second document that corresponds to the set of keywords by inputting the prompt into a large-scale language model.
[0010] According to an exemplary aspect of the present disclosure, an exemplary effect is provided in that a technology can be provided that references a document including an information source and generates a target document that is easy to understand for the reader.
[0011] 1 is a block diagram showing the configuration of a document generation device according to the present disclosure. FIG. 2 is a flow diagram showing the flow of a document generation method according to the present disclosure. FIG. 3 is a block diagram showing the configuration of a document generation device according to the present disclosure. FIG. 4 is a diagram showing a specific example of a prompt template according to the present disclosure. FIG. 5 is a flow diagram showing the flow of a document generation method according to the present disclosure. FIG. 6 is a diagram showing a specific example of a first document according to the present disclosure. FIG. 7 is a diagram showing a specific example of a keyword set according to the present disclosure. FIG. 8 is a diagram showing a specific example of a keyword set according to the present disclosure. FIG. 9 is a diagram showing a specific example of a second document according to the present disclosure. FIG. 10 is a diagram showing a specific example of a second document according to the present disclosure. FIG. 11 is a diagram showing a specific example of an annotation list according to the present disclosure. FIG. 12 is a diagram showing another specific example of an annotation list according to the present disclosure. FIG. 13 is a diagram showing another specific example of a second document according to the present disclosure. FIG. 14 is a block diagram showing an example of a computer functioning as a document generation device according to the present disclosure.
[0012] The following are examples of embodiments of the present invention. However, the present invention is not limited to the exemplary embodiments shown below, and various modifications are possible within the scope of the claims. For example, embodiments obtained by appropriately combining the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, embodiments obtained by appropriately omitting some of the technical means employed in the exemplary embodiments shown below may also be included in the scope of the present invention. Furthermore, the effects mentioned in the exemplary embodiments shown below are examples of effects expected in the exemplary embodiments, and do not define the scope of the present invention. In other words, embodiments that do not exhibit the effects mentioned in the exemplary embodiments shown below may also be included in the scope of the present invention.
[0013] [First Exemplary Embodiment] A first exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. This exemplary embodiment is a basic form for each of the exemplary embodiments described below. Note that the scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise. Furthermore, each technical means shown in the drawings referenced to explain this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical obstacles arise.
[0014] (Configuration of Document Generation Device) The configuration of the document generation device 1 will be described with reference to FIG. 1. FIG. 1 is a block diagram showing the configuration of the document generation device 1. As shown in FIG. 1, the document generation device 1 includes an attribute acquisition unit 11, a keyword set acquisition unit 12, a prompt generation unit 13, and a document generation unit 14. The attribute acquisition unit 11 acquires attributes of a recipient of a second document, which is a document to be generated by referring to a first document. The keyword set acquisition unit 12 acquires a set of keywords included in the first document. The prompt generation unit 13 generates a prompt including an example of the keyword set, an example of a second document corresponding to the example of the keyword set and corresponding to the attribute of the recipient, and the set of keywords. The document generation unit 14 generates a second document corresponding to the keyword set by inputting the prompt into a large-scale language model.
[0015] Here, the document generation device 1 is a device that generates a second document by referencing a first document. The first document includes an information source for generating the second document. The second document is a target document to be generated by referencing the first document. The recipient of the second document refers to a reader who receives the second document and reads its contents.
[0016] For example, the attribute acquisition unit 11 may acquire the attributes of the recipient of the second document based on input information input from an input device (not shown). Alternatively, the attribute acquisition unit 11 may acquire the attributes of the recipient of the second document by referring to the first document. For example, attributes related to the content described in the first document may be acquired as the attributes of the recipient of the second document.
[0017] (Effects of the Document Generation Device) As described above, the document generation device 1 employs a configuration including the above-mentioned attribute acquisition unit 11, keyword set acquisition unit 12, prompt generation unit 13, and document generation unit 14. Therefore, the document generation device 1 has the effect of being able to refer to a first document including an information source and generate a desired second document that is easy to understand for the reader.
[0018] (Example of Implementation by a Program) When the document generation device 1 is configured by a computer having a processor and a memory, the memory stores a document generation program. The document generation program is a program that causes a computer to function as the document generation device 1, and causes the computer to function as an attribute acquisition unit 11 that acquires attributes of a recipient of a second document that is a document to be generated by referring to a first document, a keyword set acquisition unit 12 that acquires a set of keywords included in the first document, a prompt generation unit 13 that generates examples of the keyword set, examples of second documents that correspond to the examples of the keyword set and that correspond to the attributes of the recipient, and a prompt that includes the set of keywords, and a document generation unit 14 that generates the second document that corresponds to the set of keywords by inputting the prompt into a large-scale language model.
[0019] (Flow of Document Generation Method) The flow of document generation method S1 will be described with reference to FIG. 2. FIG. 2 is a flowchart showing the flow of document generation method S1. As shown in FIG. 2, document generation method S1 includes attribute acquisition processing S11, keyword acquisition processing S12, prompt generation processing S13, and document generation processing S14. In attribute acquisition processing S11, at least one processor acquires attributes of a recipient of a second document, which is a document to be generated by referring to a first document. In keyword acquisition processing S12, at least one processor acquires a set of keywords included in the first document. In prompt generation processing S13, at least one processor generates a prompt including an example of the keyword set, an example of a second document corresponding to the example of the keyword set and corresponding to the attribute of the recipient, and the set of keywords. In document generation processing S14, at least one processor generates a second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
[0020] (Effects of the Document Generation Method) As described above, the document generation method S1 employs a configuration including the attribute acquisition process S11, keyword acquisition process S12, prompt generation process S13, and document generation process S14. Therefore, the document generation method S1 has the effect of being able to refer to a first document including an information source and generate a desired second document that is easy to understand for the reader.
[0021] Second Exemplary Embodiment A second exemplary embodiment, which is an example of an embodiment of the present invention, will be described in detail with reference to the drawings. Components having the same functions as those described in the above exemplary embodiment will be denoted by the same reference numerals, and their description will be omitted as appropriate. The scope of application of each technical means employed in this exemplary embodiment is not limited to this exemplary embodiment. That is, each technical means employed in this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs. Furthermore, each technical means shown in each drawing referenced to describe this exemplary embodiment can also be employed in other exemplary embodiments included in the present disclosure, to the extent that no particular technical hindrance occurs.
[0022] The document generation device 1A is a device that generates a second document by referencing a first document, similar to the document generation device 1. In this exemplary embodiment, the document generation device 1A is applied to the medical field, and an example in which the first document is a medical record and the second document is a medical information report will be mainly described.
[0023] (Configuration of Document Generation Device) The configuration of document generation device 1A will be described with reference to Fig. 3. Fig. 3 is a block diagram showing the configuration of document generation device 1A. In addition to the attribute acquisition unit 11, keyword set acquisition unit 12, prompt generation unit 13, and document generation unit 14 that are included in document generation device 1, document generation device 1A also includes an annotation list acquisition unit 15 and a user interface unit 16.
[0024] The document generation device 1A is also connected to a first document storage device 21, a keyword extraction model storage device 22, a score calculation model storage device 23, an editing result storage device 24, a template storage device 25, a case information storage device 26, a necessity information storage device 27, an annotation dictionary storage device 28, and a large-scale language model storage device 29. Some or all of these storage devices 21 to 29 may be connected to the document generation device 1A as peripheral devices, or may be connected via a network. These storage devices 21 to 29 do not all have to be physically different storage devices, and at least two of them may be physically the same storage device. Some or all of these storage devices 21 to 29 may be included in the document generation device 1A instead of being connected to the document generation device 1A.
[0025] The user interface unit 16 acquires information input by a user via an input device (not shown). Hereinafter, the information acquired by the user interface unit 16 will also be referred to as user input information, etc. The user interface unit 16 also displays the information generated by the document generation device 1A on a display device (not shown).
[0026] The attribute acquisition unit 11 is configured similarly to the first exemplary embodiment, and is also configured as follows: The attribute acquisition unit 11 may acquire attributes of the recipient of the second document based on user-input information. The acquired recipient attributes are any of a plurality of preset attributes. When the second document is a medical information report, the plurality of preset recipient attributes include doctor, nurse, care manager, etc. Note that the attribute acquisition unit 11 may acquire the attributes of the recipient of the second document by referring to the first sentence; however, the following description will focus on an example in which the attributes are acquired based on user-input information.
[0027] The keyword set acquisition unit 12 is configured similarly to that of the first exemplary embodiment, and is further configured as follows: The keyword set acquisition unit 12 includes a first document acquisition unit 121, a keyword extraction unit 122, a score calculation unit 123, and a keyword set editing unit 124.
[0028] The first document acquisition unit 121 acquires a first document to be referenced in order to generate a second document from the first document storage device 21. The first document storage device 21 stores one or more first documents. For example, if the second document to be generated is a medical information report, the first document storage device 21 may store a medical record (first document) in association with a patient ID and a date and time.
[0029] For example, the first document acquisition unit 121 may search the first document storage device 21 for one or more medical records corresponding to the patient ID included in the user input information and acquire the one or more medical records. Furthermore, if the user input information further includes a period, the first document acquisition unit 121 may search the first document storage device for one or more medical records corresponding to the period among the one or more medical records corresponding to the patient ID and acquire the one or more medical records. In this case, the first document is composed of a chronological order of one or more medical records.
[0030] The keyword extraction unit 122 extracts one or more keywords from the first document. For example, the keyword extraction unit 122 may extract one or more keywords from the first document according to a summarization rate specified by a user. A keyword set is acquired based on the extracted one or more keywords. A keyword extraction model is used in the process of extracting one or more keywords from the first document.
[0031] The keyword extraction model storage device 22 stores a keyword extraction model. The keyword extraction model is a model that extracts and outputs keywords from an input document. For example, a known technique for extracting named entities from natural language text can be applied as the keyword extraction model. The keyword extraction model may also be a model that extracts named entities and factual attributes from natural language text. The factual attributes are descriptions that indicate the factuality of the named entities (whether they actually happened, the likelihood of them happening, etc.). The keyword extraction model may also be configured to extract named entities of a specific category. Specific examples of specific categories include, but are not limited to, dates and times, people's names, organization names, and disease names.
[0032] An example of a technology applicable as a keyword extraction model will be described. For example, extraction of keyword (proper name or event) expressions from natural language sentences is often solved by BIO-style sequence labeling, in which an input token (or morpheme) sequence is assigned one of the following labels: "B" indicating the beginning of a span corresponding to the keyword expression, "I" indicating the middle of the span, or "O" indicating the outside of the span. A model that realizes this sequence labeling is often BERT (Bidirectional Encoder Representations from Transformers) or its derivative model. By performing supervised fine tuning (FST) using supervised data from BIO-style sequence labeling based on a BERT model or its derivative model pre-trained on a large corpus, a keyword extraction (proper name extraction, event detection) model with high performance for a target domain can be realized. However, technologies applicable as keyword extraction models are not limited to the above-mentioned examples.
[0033] The score calculation unit 123 assigns a score to each of one or more keywords extracted from the first document according to the attributes of the recipient. A keyword set is acquired based on the assigned scores. A score calculation model according to the attributes of the recipient is used in the process of assigning scores to one or more keywords extracted from the first document.
[0034] The score calculation model storage device 23 stores a score calculation model for each attribute of the recipient. The score calculation model is a model for calculating a score to be assigned to each keyword output from the keyword extraction model. For example, the score calculation model may be a classifier that predicts whether a keyword included in a first document is used in the corresponding second document. In this case, the likelihood output from the classifier may be used as the score. The score calculation model for each attribute of the recipient may be obtained by training the classifier described above using a combination of examples from the first document and examples from the second document corresponding to the attribute. Known technologies can be applied to such a classifier.
[0035] An example of a technology applicable as a score calculation model will be described. For example, a score for a keyword extracted from a first document can be obtained through a sequence labeling process in which the keyword is assigned a label "1" if the keyword is used in a second document, or a label "0" if the keyword is not used in the second document. A BERT model or a derivative model thereof is often used as a model for realizing such sequence labeling. Based on a BERT model or a derivative model thereof pre-trained on a large corpus, SFT is performed using supervised binary sequence labeling data indicating whether a keyword from a first document appears in a second document. When new input keywords are input to a BERT model or a derivative model thereof trained by SFT, an estimated probability that each keyword will appear in the second document is obtained. A score calculation model for a keyword can be realized by using this estimated probability as a score for the keyword. However, technologies applicable as a score calculation model are not limited to the above example.
[0036] In addition, the score calculation unit 123 does not necessarily have to assign a score to each keyword according to the attributes of the recipient. In that case, it may use a single score calculation model to assign a score to each of one or more keywords extracted from the first document.
[0037] The keyword set editing unit 124 acquires a keyword set based on the user's editing results for one or more keywords extracted from the first document. For example, the keyword set editing unit 124 selects one or more keywords based on their scores from among the one or more keywords extracted from the first document according to the summarization rate, and sets the selected keywords as a provisional keyword set. Examples of selection based on scores include, but are not limited to, selecting one or more keywords with scores equal to or greater than a threshold, or selecting a predetermined number or a predetermined percentage of scores in descending order of score.
[0038] The keyword set editing unit 124 also accepts user edits to the tentative keyword set by displaying the tentative keyword set on the display device in an editable manner via the user interface unit 16. The keyword set editing unit 124 also acquires the results of editing by the user via the user interface unit 16 and finalizes the keyword set based on the results of editing. User edits to the tentative keyword set may include deleting or adding keywords, or modifying the keywords themselves. The edits may also include changing the summarization rate when extracting one or more keywords from the first document.
[0039] The edited result storage device 24 stores the edited result by the keyword set editor 124, that is, the confirmed keyword set. The prompt generator 13 references this keyword set.
[0040] The annotation list acquisition unit 15 acquires an annotation list including at least one term used in a keyword set included in the first document and an annotation for that term. A "term used in the keyword set" may be a keyword itself included in the keyword set, or a part of the keyword. For example, the annotation list acquisition unit 15 may acquire the annotation list by searching an annotation dictionary for a term used in the keyword set.
[0041] The annotation dictionary storage device 28 stores an annotation dictionary. The annotation dictionary includes one or more pairs of terms that may be included in the first document and annotation sentences corresponding to the terms. Terms included in the annotation dictionary are also referred to as terms registered in the annotation dictionary. It is desirable that the terms registered in the annotation dictionary are terms that correspond to the field of the first document. For example, if the first document is a medical record, the annotation dictionary may include registered terms such as technical terms and abbreviations in the medical field, local terms and abbreviations used in the specific medical institution where the medical record is recorded, and the like.
[0042] Furthermore, the annotation list acquisition unit 15 may select, from one or more terms used in a keyword set included in the first document, terms to be included in the annotation list in accordance with the attributes of the recipient. For example, the annotation list acquisition unit 15 may select terms in accordance with the attributes of the recipient by referring to the necessity information storage device 27.
[0043] The necessity information storage device 27 stores necessity information for each term. The necessity information is information indicating whether an annotation is necessary for the term. For example, the necessity information may be information in which necessity is associated with a combination of a term and an attribute of a recipient.
[0044] The prompt generation unit 13 is configured similarly to the first exemplary embodiment, but is also configured as follows: The prompt generation unit 13 generates a prompt consisting of in-context learning data and a command sentence. The prompt generation unit 13 also includes an annotation list in the prompt. Here, the in-context learning data is learning data used for in-context learning (ICL) using a large-scale language model. The in-context learning data includes examples of a keyword set and examples of second documents corresponding to the examples of the keyword set, the second document examples being in accordance with the attributes of the recipient. The structure of the prompt may be defined by a template. The template storage device 25 stores prompt templates.
[0045] 4 is a diagram illustrating a specific example of a template stored in the template storage device 25. The template 401 shown in FIG. 4 defines the structure of a prompt. The template 401 is composed of in-context learning data, a prompter, and an answer.
[0046] The in-context training data includes prompter examples and answer examples. The prompter examples are the information in the in-context training data from the start tag "" to the end tag "". The answer examples are the information in the in-context training data from the start tag "" to the end tag "". By including the in-context training data in the prompts, a large-scale language model can be trained in-context to generate answers to the prompt.
[0047] A prompter is information from the start tag "" to the end tag "" and indicates a query to the large-scale language model. An answer includes the start tag "". By inputting a prompt that ends with the answer start tag "" into the large-scale language model, it is possible to request an answer corresponding to the prompter from the large-scale language model.
[0048] In the example of FIG. 4, character strings enclosed in curly brackets "{}" (for example, "{command statement}") indicate embedded information that may differ for each generated prompt.
[0049] The {command} contains information about the task to be instructed to the large-scale language model. The {command} may include, for example, task content, conditions, auxiliary information, etc. Details of the {command} are specified, for example, by the command 402. The {sender} in the command 402 contains information about the sender of the second document (e.g., the creator of the medical information report) (e.g., the sender's institution, occupation, name, etc.). The {recipient} contains the recipient of the second document (e.g., the institution, occupation, name, etc. of the person to whom the medical information report is to be sent). The {purpose} contains the purpose of providing the medical information report (e.g., the medical treatment requested from the recipient, etc.). The {document name} contains the name of the second document (e.g., medical information report, etc.). The {constraint} contains constraints for the task to be executed by the large-scale language model (e.g., a character limit of 400 characters or less). The {annotation list} contains an annotation list.
[0050] An example of information to be embedded in the above-mentioned {command sentence} is embedded in the {command sentence} included in the in-context learning data. An example of information to be embedded in the above-mentioned {keyword set} is embedded in the {keyword set}. An example of a second document corresponding to the {command sentence case} and the {keyword set case} is embedded in the {second document case}. The prompt generation unit 13 acquires information to be embedded in the prompt as in-context learning data from the case information storage device 26.
[0051] The case information storage device 26 stores one or more pieces of case information. For example, the case information includes a first document case and a second document case. The case information is also associated with the attributes of a recipient who received the second document case. For example, the prompt generation unit 13 may embed a second document case included in the case information associated with the recipient's attributes in {second document case}. The prompt generation unit 13 may also embed a set of keywords extracted by the keyword extraction unit 122 from the first document case included in the case information in {keyword set case}. The prompt generation unit 13 may also extract the sender, recipient, purpose, document name, constraints, annotation list, etc. from the second document case, and embed the extracted information-embedded command statement 402 in {command statement case}.
[0052] For example, a second document example included in case information associated with a certain attribute (e.g., doctor) does not include annotations and has brief content. A second document example included in case information associated with another attribute (e.g., other than doctor) includes annotations and has detailed content. Thus, the amount or level of detail of annotations in the second document example may vary depending on the level of expertise, experience, or knowledge assumed for the associated attribute.
[0053] Such case information may include, as respective examples, a second document previously generated by the document generation device 1A and a first document corresponding to the second document. Furthermore, the case information may include, as respective examples, a second document manually created by a creator (e.g., a doctor) without using the document generation device 1A and a first document corresponding to the second document. Furthermore, the case information may include an example of a fictitious second document generated for in-context learning and an example of a fictitious first document corresponding to the example of the second document.
[0054] The case information may include examples of command sentences, examples of keyword sets, and examples of the second document, instead of examples of the first document and examples of the second document. In this case, the prompt generation unit 13 can embed the examples of command sentences, examples of keyword sets, and examples of the second document included in the case information associated with the attributes of the recipient directly into the corresponding locations of the in-context learning data.
[0055] The document generation unit 14 is configured similarly to the first exemplary embodiment, and is also configured as follows: The document generation unit 14 may generate a second document including annotation sentences by inputting a prompt including an annotation list to the large-scale language model. The document generation unit 14 also displays the second document obtained from the large-scale language model on a display device via the user interface unit 16.
[0056] The large-scale language model storage device 29 stores a large-scale language model. The large-scale language model may be a general-purpose large-scale language model or a large-scale language model fine-tuned for the domain of the first document. When a prompt indicating a user's request is input, the large-scale language model generates and outputs a natural language sentence in response to the request. The prompt is as described above.
[0057] (Flow of Document Generation Method) The document generation device 1A configured as above executes a document generation method S1A. The flow of the document generation method S1A will be described with reference to Fig. 5. Fig. 5 is a flow diagram illustrating the flow of the document generation method S1A. As shown in Fig. 5, the document generation method S1A includes steps S101 to S110.
[0058] In step S101, the user interface unit 16 acquires user-input information. The user-input information includes at least information about the recipient of the second document. The user-input information may also include information for identifying the first document to be referenced (e.g., patient ID, time period, etc.), the sender of the second document, purpose, document name, constraints, summarization rate, etc.
[0059] Step S102 is an example of attribute acquisition processing. In step S102, the attribute acquisition unit 11 acquires the attributes of the recipient based on user-input information. For example, if the user-input information includes a character string related to the recipient, such as "Recipient = Organization: A Hospital, Department AA, Occupation: Doctor, Name: BB," the attribute acquisition unit 11 may acquire the attribute "Doctor" of the recipient based on the character string related to the recipient. Furthermore, the attribute acquisition unit 11 may acquire the attributes of the recipient based on setting information recorded in a setting file or the like, rather than based on the user-input information.
[0060] Steps S103 to S106 are an example of a keyword set acquisition process. In step S103, the primary document acquisition unit 121 acquires a primary document to be referenced from the primary document storage device 21. For example, the primary document acquisition unit 121 acquires the primary document by searching for the relevant primary document from the primary document storage device 21 based on information for identifying the primary document included in the user input information.
[0061] FIG. 6 is a diagram showing a specific example of the first document acquired in step S103. The first document shown in FIG. 6 is a time series of medical records. For example, if the user-input information includes information such as "patient ID=X, period=ymd1 to ymd2," the time series of medical records will consist of medical records associated with the patient ID "X" and the dates and times included in the period "ymd1 to ymd2." The natural language sentences included in the medical records indicate the details of medical treatment entered by the doctor.
[0062] In step S104, the keyword extraction unit 122 uses the keyword extraction model to extract one or more keywords from the first document according to the summarization rate. If the user-input information includes a summarization rate, the summarization rate is used. If the user-input information does not include a summarization rate, a default summarization rate may be used.
[0063] In step S105, the score calculation unit 123 assigns a score to each keyword using a score calculation model according to the attributes of the recipient.
[0064] In step S106, the keyword set editing unit 124 selects one or more keywords from the one or more keywords extracted from the first document based on the scores, and sets the selected keywords as a tentative keyword set. The keyword set editing unit 124 also accepts user edits to the tentative keyword set. The keyword set editing unit 124 also finalizes the edited result as a keyword set and stores it in the edited result storage device 24.
[0065] For example, suppose the provisional keyword set is extracted according to the summarization rate input by the user in step S101. If the user edits this provisional keyword set to reduce the summarization rate, the number of keywords included in the provisional keyword set will decrease. Furthermore, suppose the provisional keyword set includes, for example, five keywords, Keyword 1 to Keyword 5. Suppose the user edits the provisional keyword set by deleting Keyword 2, modifying the expression of Keyword 3, and adding Keyword 6. In this case, based on the editing results, a keyword set consisting of Keyword 1, Keyword 3 with the modified expression, Keyword 4, Keyword 5, and Keyword 6 is finalized.
[0066] Fig. 7 shows a specific example of a set of keywords extracted from the time series of medical records shown in Fig. 6. In this example, pairs of keywords indicating dates and keywords contained in the medical records corresponding to those dates are extracted as keywords.
[0067] Step S107 is an example of an annotation list acquisition process. In step S107, the annotation list acquisition unit 15 generates an annotation list including at least one term used in the keyword set and a comment for that term, by referring to necessity information. Specifically, the annotation list acquisition unit 15 searches for terms registered in the annotation dictionary in the keyword set. The annotation list acquisition unit 15 also acquires necessity information for each of the searched terms according to the recipient's attributes by referring to the necessity information storage device 27. The annotation list acquisition unit 15 also generates an annotation list including terms whose necessity information is "necessary" and comment. Note that the annotation list acquisition unit 15 does not include terms whose necessity information is "no" in the annotation list.
[0068] For example, in the keyword set of FIG. 7, seven terms, "hoarseness," "laryngeal cancer," "posterior," "cT1a," "T1aN0M0," "alloid," and "Gy," are registered in the annotation dictionary. The terms used in the keyword set may be keywords themselves, such as the term "hoarseness," or parts of keywords, such as the term "laryngeal cancer." Furthermore, assume that the recipient's attribute is "physician," and the necessity information for these seven terms corresponding to "physician" is all "no." Thus, depending on the recipient's attribute and the necessity information, an annotation list may not be generated. Here, the following description will be continued assuming that an annotation list has not been generated.
[0069] It is also possible that the necessity information for any term regarding a specific attribute of the recipient (in this example, "doctor") is "no." In such a case, in step S107, it is determined whether the attribute of the recipient is a specific attribute, and if it is a specific attribute, the process of referencing the necessity information may be omitted so that an annotation list is not generated.
[0070] Step S108 is an example of a prompt generation process. In step S108, the prompt generation unit 13 generates a prompt using a keyword set example, a second document example corresponding to the keyword set example and corresponding to the recipient's attribute, and a keyword set included in the first set. Specifically, the prompt generation unit 13 obtains, from the case information storage device 26, case information associated with the recipient's attribute (here, for example, "doctor").
[0071] Furthermore, the prompt generation unit 13 generates an example of an imperative sentence by referring to an example of a second document included in the case information associated with the "doctor." Alternatively, if the case information includes an example of an imperative sentence itself, the prompt generation unit 13 may acquire the example of the imperative sentence itself. The prompt generation unit 13 embeds such an example of an imperative sentence in {imperative sentence examples} constituting the in-context learning data of the template shown in FIG. 4.
[0072] The prompt generation unit 13 also extracts keyword set examples from the first document examples included in the case information using the keyword set acquisition unit 12. Alternatively, if the case information includes a keyword set example itself, the prompt generation unit 13 may acquire the keyword set example itself. Figure 8 shows a specific example of a keyword set example acquired in this manner. The prompt generation unit 13 embeds such keyword set examples in {keyword set examples} that constitute the in-context learning data of the template shown in Figure 4.
[0073] 9 is a diagram showing a specific example of a second document example included in the case information associated with the "doctor." The second document example (medical information report example) shown in FIG. 9 is intended for a recipient with the "doctor" attribute, and therefore does not include annotations and has simple content. The prompt generation unit 13 embeds such a second document example in {second document example} constituting the in-context learning data of the template shown in FIG. 4.
[0074] The prompt generation unit 13 also extracts the sender, recipient, purpose, document name, and constraints related to the second document to be generated from the user input information acquired in step S101, and generates a command statement by embedding these in the corresponding locations of the command statement 402 shown in Fig. 4. Note that if an annotation list has not been generated, the prompt generation unit 13 does not embed the annotation list in the {annotation list} of the command statement 402. The prompt generation unit 13 also embeds the command statement generated in this manner in the {command statement} that constitutes the prompter of the template shown in Fig. 4.
[0075] The prompt generation unit 13 also embeds the keyword set stored in the edited result storage device 24 into the {keyword set} that constitutes the prompter of the template shown in Fig. 4. In this way, the prompt generation unit 13 generates a prompt by embedding necessary information in each embedding location included in the template.
[0076] Step S109 is an example of a document generation process. In step S109, the document generation unit 14 inputs the generated prompt into the large-scale language model to obtain a second document output from the large-scale language model. Specifically, the second document is output from the large-scale language model as a response to the input prompt. Here, a set of keywords contained in the first document is embedded in the prompter in the input prompt. Therefore, the second document output as a response to such a prompter has content corresponding to the set of keywords contained in the first document. Furthermore, examples of the second document corresponding to the attributes of the recipient are embedded in the in-context learning data in the input prompt. Therefore, the second document output as a response through in-context learning that references such in-context learning data has content corresponding to the attributes of the recipient.
[0077] FIG. 10 is a diagram showing a specific example of the second document generated in step S109. The specific example of the second document is a medical information report for a doctor generated with reference to the time series of medical records ( FIG. 6 ). The medical information report was output from the large-scale language model as a response corresponding to a set of keywords ( FIG. 7 ) acquired from the time series of medical records. Furthermore, the medical information report has concise content without annotations, in accordance with the recipient attribute "doctor" entered by the user. This is because the large-scale language model generated the medical information report with reference to examples of the keyword set ( FIG. 8 ) and examples of medical information reports corresponding to the recipient attribute "doctor" ( FIG. 8 ) as in-context training data.
[0078] In step S110, the user interface unit 16 displays the generated second document on the display device, and the document generation method S1A is now complete.
[0079] (Example of Generating Second Document Including Annotation) In the above-described document generation method S1A, an example of generating a second document that does not include an annotation for a recipient with an attribute such as "doctor" has been mainly described. Below, an example of generating a second document that includes an annotation for a recipient with an attribute such as "care manager" will be described.
[0080] In this case, the processes from steps S101 to S106 are executed as described above, except that the user input information acquired in step S101 includes an attribute of the recipient (e.g., a care manager) for which it is desirable to insert an annotation.
[0081] In step S107, the annotation list acquisition unit 15 selects terms whose necessity information is "necessary" from among the terms registered in the annotation dictionary used in the keyword set. The annotation list acquisition unit 15 also generates an annotation list including the selected terms and their annotations. Figures 11 and 12 show specific examples of the annotation list generated in step S107.
[0082] For example, suppose the attribute of the recipient is "care manager." Furthermore, suppose that the necessity information for "care manager" indicates "needed" for each of the seven terms used in the keyword set of FIG. 7 , namely, "hoarseness," "laryngeal cancer," "posterior," "cT1a," "T1aN0M0," "alloyd," and "Gy." In this case, the annotation list shown in FIG. 11 is generated. This annotation list includes the seven terms for which the necessity information for "care manager" indicates "needed" and their annotations.
[0083] For example, suppose the attribute of the recipient is "nurse." Furthermore, of the seven terms used in the keyword set of FIG. 7 , for four terms, "hoarseness," "laryngeal cancer," "posterior," and "alloid," the necessity information for "nurse" all indicates "no." Furthermore, for three terms, "cT1a," "T1aN0M0," and "Gy," the necessity information for "nurse" all indicates "needed." In this case, the annotation list shown in FIG. 12 is generated. This annotation list includes the three terms whose necessity information for "nurse" indicates "needed," along with their annotations.
[0084] The prompt generation process in step S108 is explained in much the same way as above, except for the following two points.
[0085] The first difference is that the second document example is different. The second document example associated with the attributes of the recipient for whom insertion of a note is desirable includes the note. The prompt generation unit 13 embeds the second document example including the note in {second document example} of the template shown in FIG. 4. FIG. 13 is a diagram showing a specific example of a second document example including a note (example of a medical information report). This medical information report includes the note (underlined portion). In other words, this medical information report is more detailed than the example of a medical information report for a "doctor" shown in FIG. 9.
[0086] The second point is that the prompt further includes an annotation list. The prompt generator 13 embeds the annotation list generated in step S107 in the {annotation list} in the command statement 402 to be embedded in the {command statement} of the template shown in Fig. 4. Here, it is assumed that the annotation list shown in Fig. 11 has been embedded, for example.
[0087] The document generation process in step S109 is similar to that described above, except for the second document that is output in step S109, which includes annotations.
[0088] FIG. 14 is a diagram showing an example of a second document including annotations generated in step S109. The example of the second document is a medical information report for a care manager generated with reference to the time series of medical records ( FIG. 6 ). The medical information report is output from the large-scale language model as a response corresponding to a set of keywords ( FIG. 7 ) acquired from the time series of medical records. The medical information report also includes annotations (underlined) corresponding to the recipient's attribute "care manager." In other words, the medical information report is more detailed than the medical information report for a "doctor" shown in FIG. 10 . This is because the large-scale language model generated the medical information report by referring to the example keyword set ( FIG. 8 ) and the example medical information report for "care manager" ( FIG. 13 ) as in-context training data, and by referring to the annotation list corresponding to the "care manager."
[0089] In this way, when a medical record is used as the first document and a medical information report is used as the second document, the document generation device 1A can generate a medical information report to be sent from one medical professional to another with content that is easy to understand according to the recipient's occupation. Therefore, for example, when medical care for a patient is provided comprehensively by medical professionals from multiple occupations, it can support collaboration between medical professionals with different occupations and different specialties, experience, knowledge, etc.
[0090] (Modification) The exemplary embodiment described above is not limited to the medical field and can be applied to other fields as well. For example, when applied to a specific business field, the first document may be a business record and the second document may be a business handover document. In this case, the annotation dictionary may be a dictionary appropriate for the business field. The large-scale language model may be a model in which a general-purpose large-scale language model is fine-tuned for the business field. For example, when applied to the education field, the first document may be a learning instruction record and the second document may be a teaching report. In this case, the annotation dictionary may be a dictionary appropriate for the education field. The large-scale language model may be a model in which a general-purpose large-scale language model is fine-tuned for the education field.
[0091] Furthermore, in the exemplary embodiment described above, the recipient's attribute is not limited to the recipient's occupation. For example, the recipient's attribute may be the organization to which the recipient belongs. In this case, by using case information associated with the organization to which the recipient belongs, a second document corresponding to the organization to which the recipient belongs can be generated. For example, the recipient's attribute may be the recipient's age group. In this case, by using case information associated with the recipient's age group, a second document corresponding to the recipient's age group can be generated. Furthermore, the recipient's attribute may be whether or not elaboration is required. In this case, if the acquired recipient's attribute is "elaboration required," a detailed second document can be generated by using case information associated with the attribute "elaboration required." In this case, if the acquired recipient's attribute is "elaboration not required," a concise second document can be generated by using case information associated with the attribute "elaboration not required." However, the recipient's attribute is not limited to the above example.
[0092] (Effects of the Exemplary Embodiment) As described above, the document generation device 1A further includes the annotation list acquisition unit 15 that acquires an annotation list including at least one term used in a keyword set and an annotation for that term, the prompt generation unit 13 further includes the annotation list in a prompt, and the document generation unit 14 generates a second document including an annotation by inputting the prompt into a large-scale language model. Therefore, in addition to the effects achieved by the document generation device 1, the document generation device 1A can generate a second document including an annotation, thereby eliminating ambiguity of terms included in the second document and further improving the understandability of the second document.
[0093] Furthermore, the document generation device 1A employs a configuration in which the annotation list acquisition unit 15 selects a term to be included in the annotation list from one or more terms used in the keyword set according to the attributes of the recipient. Therefore, in addition to the effects achieved by the document generation device 1, the document generation device 1A can generate a second document including annotations for terms that require explanation according to the attributes of the recipient, thereby further improving the ease of understanding of the second document.
[0094] Furthermore, the document generation device 1A employs a configuration in which the keyword set acquisition unit 12 acquires a keyword set based on the user's editing results for one or more keywords extracted from the first document. Therefore, in addition to the effects achieved by the document generation device 1, the document generation device 1A can generate a second document that reflects the user's intentions by using the keyword set edited by the user. As a result, it is possible to reduce the user's work of correcting the second document and the user's work of reviewing the first document to correct the second document.
[0095] Furthermore, the document generation device 1A employs a configuration in which the keyword set acquisition unit 12 acquires a keyword set based on one or more keywords extracted from the first document according to the summarization rate specified by the user. Therefore, in addition to the effects achieved by the document generation device 1, the document generation device 1A can generate a second document that reflects the degree of summarization desired by the user by using a keyword set based on the summarization rate specified by the user.
[0096] Furthermore, the document generation device 1A employs a configuration in which the keyword set acquisition unit 12 assigns a score according to the attributes of the recipient to each of one or more keywords extracted from the first document, and acquires a keyword set based on the assigned scores. Therefore, in addition to the effects achieved by the document generation device 1, the document generation device 1A can use a keyword set according to the attributes of the recipient, and can generate a second document including keywords according to the attributes of the recipient.
[0097] [Software Implementation Example] Some or all of the functions of the document generation devices 1, 1A (hereinafter also referred to as "each of the above devices") may be realized by hardware such as an integrated circuit (IC chip), or by software.
[0098] In the latter case, each of the above devices is realized by, for example, a computer that executes instructions of a program, which is software that realizes each function. An example of such a computer (hereinafter referred to as computer C) is shown in Figure 15. Figure 15 is a block diagram showing the hardware configuration of computer C that functions as each of the above devices.
[0099] The computer C includes at least one processor C1 and at least one memory C2. The memory C2 stores a program P for causing the computer C to function as each of the above-mentioned devices. In the computer C, the processor C1 reads and executes the program P from the memory C2, thereby realizing the functions of each of the above-mentioned devices.
[0100] The processor C1 may be, for example, a central processing unit (CPU), a graphics processing unit (GPU), a digital signal processor (DSP), a micro processing unit (MPU), a floating point number processing unit (FPU), a physics processing unit (PPU), a tensor processing unit (TPU), a quantum processor, a microcontroller, or a combination thereof. The memory C2 may be, for example, a flash memory, a hard disk drive (HDD), a solid state drive (SSD), or a combination thereof.
[0101] The computer C may further include a RAM (Random Access Memory) for expanding the program P during execution and for temporarily storing various data. The computer C may also include a communication interface for transmitting and receiving data to and from other devices. The computer C may also include an input / output interface for connecting input / output devices such as a keyboard, a mouse, a display, and a printer.
[0102] The program P can also be recorded on a non-transitory, tangible recording medium M that can be read by the computer C. Such a recording medium M can be, for example, a tape, a disk, a card, a semiconductor memory, or a programmable logic circuit. The computer C can acquire the program P via such a recording medium M. The program P can also be transmitted via a transmission medium. Such a transmission medium can be, for example, a communication network or broadcast waves. The computer C can also acquire the program P via such a transmission medium.
[0103] [Appendix A] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0104] (Appendix A1) A document generation device comprising: an attribute acquisition means for acquiring attributes of a recipient of a second document that is a document to be generated by referring to a first document; a keyword set acquisition means for acquiring a set of keywords included in the first document; a prompt generation means for generating examples of the keyword set, examples of the second document corresponding to the examples of the keyword set, the second document examples corresponding to the attributes of the recipient, and a prompt including the set of keywords; and a document generation means for generating the second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
[0105] (Appendix A2) The document generation device described in Appendix A1, further comprising an annotation list acquisition means for acquiring an annotation list including at least one term used in the keyword set and an annotation sentence for the term, wherein the prompt generation means further includes the annotation list in the prompt, and the document generation means generates the second document including the annotation sentence by inputting the prompt into the large-scale language model.
[0106] (Appendix A3) The document generation device according to appendix A2, wherein the annotation list acquisition means selects a term to be included in the annotation list from one or more terms used in the keyword set according to an attribute of the recipient.
[0107] (Appendix A4) The document generation device according to any one of appendices A1 to A3, wherein the keyword set acquisition means acquires the keyword set based on a result of editing by a user on one or more keywords extracted from the first document.
[0108] (Appendix A5) The document generation device according to any one of Appendices A1 to A4, wherein the keyword set acquisition means acquires the keyword set based on one or more keywords extracted from the first document according to a summarization rate designated by a user.
[0109] (Appendix A6) The document generation device according to any one of Appendices A1 to A5, wherein the keyword set acquisition means assigns a score according to attributes of the recipient to each of one or more keywords extracted from the first document, and acquires the keyword set based on the assigned scores.
[0110] [Appendix B] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0111] (Appendix B1) A document generation method comprising: an attribute acquisition process in which at least one processor acquires attributes of a recipient of a second document, which is a document to be generated by referring to a first document; a keyword set acquisition process in which the at least one processor acquires a set of keywords included in the first document; a prompt generation process in which the at least one processor generates examples of the keyword set, examples of the second document corresponding to the examples of the keyword set, the second document examples corresponding to the attributes of the recipient, and a prompt including the set of keywords; and a document generation process in which the at least one processor generates the second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
[0112] (Appendix B2) The document generation method described in Appendix B1, further comprising an annotation list acquisition process in which the at least one processor acquires an annotation list including at least one term used in the keyword set and an annotation sentence for the term; in the prompt generation process, the at least one processor further includes the annotation list in the prompt; and in the document generation process, the at least one processor generates the second document including the annotation sentence by inputting the prompt into the large-scale language model.
[0113] (Appendix B3) The document generation method described in Appendix B2, wherein, in the annotation list acquisition process, the at least one processor selects terms to be included in the annotation list from one or more terms used in the keyword set according to attributes of the recipient.
[0114] (Appendix B4) The document generation method according to any one of Appendices B1 to B3, wherein in the keyword set acquisition process, the at least one processor acquires the keyword set based on a user's editing results for one or more keywords extracted from the first document.
[0115] (Appendix B5) A document generation method according to any one of Appendices B1 to B4, wherein in the keyword set acquisition process, the at least one processor acquires the keyword set based on one or more keywords extracted from the first document according to a summarization rate specified by a user.
[0116] (Appendix B6) A document generation method described in any one of Appendices B1 to B5, wherein in the keyword set acquisition process, the at least one processor assigns a score to each of one or more keywords extracted from the first document according to the attributes of the recipient, and acquires the keyword set based on the assigned scores.
[0117] [Appendix C] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0118] (Appendix C1) A program that causes a computer to function as a document generation device, causing the computer to function as: attribute acquisition means that acquires attributes of a recipient of a second document that is a document to be generated by referring to a first document; keyword set acquisition means that acquires a set of keywords included in the first document; prompt generation means that generates examples of the keyword set, examples of the second document that correspond to the examples of the keyword set and that correspond to the attributes of the recipient, and a prompt that includes the set of keywords; and document generation means that generates the second document that corresponds to the set of keywords by inputting the prompt into a large-scale language model.
[0119] (Appendix C2) The document generation program described in Appendix C1, further causing the computer to function as annotation list acquisition means for acquiring an annotation list including at least one term used in the keyword set and an annotation sentence for the term, the prompt generation means further including the annotation list in the prompt, and the document generation means generating the second document including the annotation sentence by inputting the prompt into the large-scale language model.
[0120] (Supplementary Note C3) The document generation program according to Supplementary Note C2, wherein the annotation list acquisition means selects a term to be included in the annotation list from one or more terms used in the keyword set according to an attribute of the recipient.
[0121] (Appendix C4) The document generation program according to any one of appendices C1 to C3, wherein the keyword set acquisition means acquires the keyword set based on a result of editing by a user on one or more keywords extracted from the first document.
[0122] (Appendix C5) The document generation program according to any one of Appendices C1 to C4, wherein the keyword set acquisition means acquires the keyword set based on one or more keywords extracted from the first document according to a summarization rate specified by a user.
[0123] (Appendix C6) The document generation program according to any one of Appendices C1 to C5, wherein the keyword set acquisition means assigns a score according to the attributes of the recipient to each of one or more keywords extracted from the first document, and acquires the keyword set based on the assigned scores.
[0124] [Appendix D] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0125] (Appendix D1) A document generation device comprising at least one processor, the at least one processor executing: an attribute acquisition process for acquiring attributes of a recipient of a second document that is a document to be generated by referring to a first document; a keyword set acquisition process for acquiring a set of keywords included in the first document; a prompt generation process for generating examples of the keyword set, examples of the second document that correspond to the examples of the keyword set and that correspond to the attributes of the recipient, and a prompt including the set of keywords; and a document generation process for generating the second document that corresponds to the set of keywords by inputting the prompt into a large-scale language model.
[0126] The document generation device may further include a memory, and the memory may store a program for causing the at least one processor to execute each of the processes.
[0127] (Appendix D2) The document generation device described in Appendix D1, wherein the at least one processor further executes an annotation list acquisition process to acquire an annotation list including at least one term used in the keyword set and an annotation sentence for the term; in the prompt generation process, the at least one processor further includes the annotation list in the prompt; and in the document generation process, the at least one processor generates the second document including the annotation sentence by inputting the prompt into the large-scale language model.
[0128] (Appendix D3) The document generation device according to Appendix D2, wherein in the annotation list acquisition process, the at least one processor selects terms to be included in the annotation list from one or more terms used in the keyword set according to attributes of the recipient.
[0129] (Appendix D4) The document generation device according to any one of appendices D1 to D3, wherein in the keyword set acquisition process, the at least one processor acquires the keyword set based on a user's editing results for one or more keywords extracted from the first document.
[0130] (Appendix D5) The document generation device described in any one of Appendices D1 to D4, wherein in the keyword set acquisition process, the at least one processor acquires the keyword set based on one or more keywords extracted from the first document according to a summarization rate specified by a user.
[0131] (Appendix D6) The document generation device described in any one of Appendices D1 to D5, wherein in the keyword set acquisition process, the at least one processor assigns a score according to the attributes of the recipient to each of one or more keywords extracted from the first document, and acquires the keyword set based on the assigned scores.
[0132] [Appendix E] This disclosure includes the techniques described in the following appendices. However, the present invention is not limited to the techniques described in the following appendices, and various modifications are possible within the scope of the claims.
[0133] (Appendix E1) A non-transitory recording medium having recorded thereon a program that causes a computer to function as a document generation device, the program causing the computer to execute: an attribute acquisition process that acquires attributes of a recipient of a second document that is a document to be generated by referring to a first document; a keyword set acquisition process that acquires a set of keywords included in the first document; a prompt generation process that generates examples of the keyword set, examples of the second document that correspond to the examples of the keyword set and that correspond to the attributes of the recipient, and a prompt that includes the set of keywords; and a document generation process that generates the second document that corresponds to the set of keywords by inputting the prompt into a large-scale language model.
[0134] 1, 1A Document generation device 11 Attribute acquisition unit 12 Keyword set acquisition unit 13 Prompt generation unit 14 Document generation unit 15 Annotation list acquisition unit 16 User interface unit 21 First document storage unit 22 Keyword extraction model storage unit 23 Score calculation model storage unit 24 Editing result storage unit 25 Template storage unit 26 Case information storage unit 27 Necessity information storage unit 28 Annotation dictionary storage unit 29 Large-scale language model storage unit 121 First document acquisition unit 122 Keyword extraction unit 123 Score calculation unit 124 Keyword set editing unit C1 Processor C2 Memory
Claims
1. A document generation device comprising: an attribute acquisition means for acquiring attributes of a recipient of a second document which is a document to be generated by referring to a first document; a keyword set acquisition means for acquiring a set of keywords contained in the first document; a prompt generation means for generating examples of the keyword set, examples of the second document which correspond to the examples of the keyword set and correspond to the attributes of the recipient, and a prompt including the set of keywords; and a document generation means for generating the second document which corresponds to the set of keywords by inputting the prompt into a large-scale language model.
2. The document generation device of claim 1, further comprising an annotation list acquisition means for acquiring an annotation list including at least one term used in the keyword set and an annotation sentence for the term, wherein the prompt generation means further includes the annotation list in the prompt, and the document generation means generates the second document including the annotation sentence by inputting the prompt to the large-scale language model.
3. The document generation device according to claim 2, wherein the annotation list acquisition means selects, from among one or more terms used in the keyword set, terms to be included in the annotation list in accordance with attributes of the recipient.
4. The document generation device according to claim 1, wherein the keyword set acquisition means acquires the keyword set based on a result of editing by a user on one or more keywords extracted from the first document.
5. A document generation device according to any one of claims 1 to 4, wherein the keyword set acquisition means acquires the keyword set based on one or more keywords extracted from the first document according to a summarization rate designated by a user.
6. A document generation device according to any one of claims 1 to 5, wherein the keyword set acquisition means assigns a score to each of one or more keywords extracted from the first document according to the attributes of the recipient, and acquires the keyword set based on the assigned scores.
7. A document generation method comprising: an attribute acquisition process in which at least one processor acquires attributes of a recipient of a second document to be generated by referring to a first document; a keyword set acquisition process in which the at least one processor acquires a set of keywords contained in the first document; a prompt generation process in which the at least one processor generates an example of the keyword set, an example of the second document corresponding to the example of the keyword set and corresponding to the attributes of the recipient, and a prompt including the set of keywords; and a document generation process in which the at least one processor generates the second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
8. A program that causes a computer to function as a document generation device, the document generation program causing the computer to function as: attribute acquisition means for acquiring attributes of a recipient of a second document that is a document to be generated by referring to a first document; keyword set acquisition means for acquiring a set of keywords contained in the first document; prompt generation means for generating examples of the keyword set, examples of the second document corresponding to the examples of the keyword set and corresponding to the attributes of the recipient, and a prompt including the set of keywords; and document generation means for generating the second document corresponding to the set of keywords by inputting the prompt into a large-scale language model.
Citation Information
Patent Citations
Information processing system, information processing apparatus, and control program
JP2022094035A