Information processing apparatus, information processing system, information processing method, and program
The information processing device generates new documents from multiple existing data without templates by using a trained model and summary-based table of contents, addressing the limitation of requiring pre-prepared templates in existing techniques.
Patent Information
- Application Number
- JP2023199309
- Authority / Receiving Office
- JP · JP
- Patent Type
- Applications
- Current Assignee / Owner
- Filing Date
- 2023-11-24
- Publication Date
- 2025-06-05
AI Technical Summary
Existing document generation techniques require a pre-prepared template, limiting the ability to generate new documents from multiple existing document data without template usage.
An information processing device with units for acquiring a generation policy, document data, generating summaries, calculating similarities, selecting summary sentences, and using a trained model to generate new document data based on a table of contents derived from selected summary sentences.
Enables the generation of new documents from multiple existing document data without using a template, improving accuracy and focusing on necessary information by leveraging a trained model and summary-based table of contents.
Smart Images

Figure 2025085433000001_ABST
Abstract
Description
[Technical field]
[0001] The present invention relates to an information processing device, an information processing system, an information processing method, and a program. [Background technology]
[0002] There is a technique for generating a different document based on an existing document.
[0003] For example, Patent Document 1 discloses a user-customized automatic document creation method, characterized in that it includes a step in which a target for automatic document creation is input by a user, the target for automatic document creation including a plurality of data items, and a plurality of tags that are matched one-to-one with the plurality of data items are defined, and a step in which a template file in which a creation format for the user-customized document is set is input by the user, the template file including at least one tag, and each tag in the template file is replaced with a data item that is matched one-to-one with each tag in the target for automatic document creation, thereby generating the user-customized document. Summary of the Invention [Problem to be solved by the invention]
[0004] However, in order for a user to generate a document that he or she desires using the technique of Patent Document 1, a template or the like must be prepared in advance.
[0005] The present invention has been made in consideration of the above-mentioned points, and has an object to make it possible to generate a new document from multiple existing document data without using a template. [Means for solving the problem]
[0006] In order to solve the above problem, an information processing device has a generation policy acquisition unit that acquires a character string indicating a generation policy for new document data based on a plurality of existing document data, a document acquisition unit that acquires the plurality of existing document data, a summarization unit that generates a summary sentence for each of the existing document data, a similarity calculation unit that calculates a similarity between each of the plurality of summary sentences and the character string, a summary selection unit that selects some summary sentences from the plurality of summary sentences based on the similarity between each of the plurality of summary sentences, and a document generation unit that causes a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the some summary sentences. Effect of the Invention
[0007] It is possible to generate a new document from multiple existing document data without using a template. [Brief description of the drawings]
[0008] [Figure 1] FIG. 1 illustrates an example of a configuration of an information processing system according to a first embodiment. [Diagram 2] 1 is a diagram illustrating an example of a hardware configuration of an information processing device 10 according to a first embodiment. [Diagram 3] 1 is a diagram illustrating an example of a functional configuration of an information processing device 10 according to a first embodiment. [Figure 4] 1 is a flowchart illustrating an example of a processing procedure executed by information processing device 10 according to the first embodiment. [Diagram 5] FIG. 2 is a diagram showing a first display example of a UI screen in the first embodiment. [Figure 6] FIG. 4 is a diagram showing an example of a prompt for generating a target document in the first embodiment. [Figure 7] FIG. 2 is a diagram showing an example of the contents of an original document in the first embodiment. [Figure 8] FIG. 13 is a diagram showing a second display example of the UI screen in the first embodiment. [Figure 9]FIG. 13 is a diagram showing a third display example of the UI screen in the first embodiment. [Figure 10] FIG. 11 is a diagram showing an example of a screen displaying a directory structure in the second embodiment. [Figure 11] FIG. 13 is a diagram illustrating an example of the configuration of an input directory. [Figure 12] FIG. 11 is a diagram showing an example of a prompt for generating a target document in the second embodiment. [Figure 13] FIG. 13 is a diagram for explaining a workspace in the third embodiment. [Figure 14] FIG. 13 is a diagram illustrating an example of a functional configuration of an information processing device 10 according to a third embodiment. [Figure 15] 13 is a diagram illustrating an example of the configuration of a workspace storage unit 132. FIG. [Figure 16] 13 is a flowchart illustrating an example of a processing procedure executed by an information processing device 10 according to the third embodiment. [Figure 17] FIG. 13 is a diagram illustrating a display example of a UI screen in the third embodiment. [Figure 18] FIG. 13 is a diagram illustrating a display example of a workspace list screen. [Figure 19] FIG. 13 is a diagram illustrating an example of a workspace details screen. [Figure 20] FIG. 13 is a diagram illustrating an example of a functional configuration of an information processing device 10 according to a fourth embodiment. [Figure 21] FIG. 13 is a diagram showing an example of a template according to the fourth embodiment. [Figure 22] 13 is a flowchart illustrating an example of a processing procedure executed by an information processing device 10 according to the fourth embodiment. [Figure 23] FIG. 13 is a diagram illustrating a display example of a UI screen in the fourth embodiment. [Figure 24] A figure showing an example of input and output of the trained model 121 when identifying necessary information. [Diagram 25] A figure showing an example of input and output of the trained model 121 when collecting necessary information. [Figure 26]FIG. 13 is a diagram showing an example of the contents of an original document in the fourth embodiment. [Figure 27] FIG. 13 is a diagram showing an example of a target document in the fourth embodiment. DETAILED DESCRIPTION OF THE PREFERRED EMBODIMENTS
[0009] Hereinafter, an embodiment of the present invention will be described with reference to the drawings. Fig. 1 is a diagram showing an example of the configuration of an information processing system in a first embodiment. As shown in Fig. 1, the information processing system includes one or more user terminals 20 and an information processing device 10.
[0010] The user terminal 20 is a terminal used by a user to input an instruction to generate one new piece of document data from a plurality of existing document data. In the instruction to generate the new piece of document data, the user specifies information indicating what kind of document data is desired to be generated from the plurality of existing document data. For example, a PC (Personal Computer), a tablet terminal, a smartphone, or the like may be used as the user terminal 20.
[0011] The information processing device 10 is one or more computers that executes generation of document data in response to a generation instruction received as an input by a user terminal 20 .
[0012] The information processing device 10 may also function as the user terminal 20.
[0013] Fig. 2 is a diagram showing an example of a hardware configuration of the information processing device 10 in the first embodiment. The information processing device 10 in Fig. 2 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, which are all connected to each other via a bus B.
[0014] A program for implementing processing in the information processing device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 via the drive device 100 into the auxiliary storage device 102. However, the program does not necessarily have to be installed from the recording medium 101, but may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program as well as necessary files, data, and the like.
[0015] When an instruction to start a program is received, the memory device 103 reads out the program from the auxiliary storage device 102 and stores it. The processor 104 is a CPU or a GPU (Graphics Processing Unit), or a CPU and a GPU, and executes functions related to the information processing device 10 in accordance with the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.
[0016] FIG. 3 is a diagram showing an example of a functional configuration of the information processing device 10 in the first embodiment. In FIG. 3, the information processing device 10 includes an instruction receiving unit 111, a generation policy acquisition unit 112, a document acquisition unit 113, a summary unit 114, a similarity calculation unit 115, a summary selection unit 116, a document generation unit 117, a table of contents generation unit 118, a document correction unit 119, a display control unit 120, and a trained model 121. Each of these units is realized by a process in which one or more programs installed in the information processing device 10 are executed by the processor 104. The information processing device 10 also uses a document storage unit 131. The document storage unit 131 can be realized by using, for example, the auxiliary storage device 102, or a storage device connectable to the information processing device 10 via a network.
[0017] The document storage unit 131 stores each of the multiple document data in a file.
[0018] The instruction receiving unit 111 receives instructions from the user inputted at the user terminal 20. For example, the instruction receiving unit 111 receives an instruction to generate new (user-desired) document data (second data) based on a plurality of existing document data (hereinafter simply referred to as a "generation instruction"), an instruction to modify the new document data (hereinafter simply referred to as a "modification instruction"), etc. Hereinafter, the new document data to be generated is referred to as a "target document," and each of the plurality of existing document data from which the target document is generated is referred to as an "original document."
[0019] The generation policy acquisition unit 112 acquires from the generation instruction a character string (hereinafter simply referred to as the "generation policy") that indicates the generation policy of a target document based on multiple source documents (what kind of document is desired to be generated) contained in the generation instruction.
[0020] The document acquisition unit 113 acquires the original documents based on the identification information (such as file names) of the original documents specified in the generation instruction. For example, the document acquisition unit 113 acquires the original documents related to the identification information from the document storage unit 131.
[0021] The summarization unit 114 generates a summary sentence for each of the original documents. The summarization unit 114 may cause the trained model 121 to generate the summary sentence for the original document.
[0022] The similarity calculation unit 115 calculates the similarity between each of the multiple summarized sentences generated for each original document and the generation policy. The method of calculating the similarity is not limited to a predetermined one as long as it is possible to evaluate the similarity between texts (character strings). For example, the similarity calculation unit 115 may convert each original document and the generation policy into a vector (distributed representation or embedded representation) using the trained model 121, and calculate the similarity between each vector of each original document and the vector of the generation policy. In this case, cosine similarity may be used as the similarity. When cosine similarity is used, the cosine similarity between vector a and vector b can be calculated based on the following formula.
[0023]
number
[0024] The table of contents generation unit 118 aligns (rearranges) some of the summary sentences selected by the summary selection unit 116 in the trained model 121 so that they make sense, and obtains a table of contents for the target document by having the trained model 121 convert each of the aligned summary sentences into a string (wording) suitable for a table of contents (appropriate for a table of contents).
[0025] The document generation unit 117 causes the trained model 121 to generate a target document based on a plurality of original documents in accordance with a table of contents based on a portion of the summary sentences selected by the summary selection unit 116. The trained model 121 generates a target document from only the contents of the original documents, and modifies the wording so that the target document makes sense.
[0026] The document correction unit 119 causes the trained model 121 to perform corrections on the target document in accordance with correction instructions from the user.
[0027] The display control unit 120 controls the display of a screen (hereinafter referred to as a "UI screen") to be displayed on the user terminal 20. The UI screen is a screen for receiving creation instructions and correction instructions from a user and displaying a target document. Screen data of the UI screen may be, for example, web-based data generated using HTML (HyperText Markup Language) or the like, or may be data based on other specifications.
[0028] The trained model 121 is a large-scale language model of a trained generative AI that responds with a character string in response to an input of a character string. As the trained model 121, a known model such as ChatGPT (GPT-4 (https: / / openai.com / research / gpt-4)) that receives a text as an input and generates and outputs a text from the occurrence probability of a word can be used. GPT is realized by a transformer (https: / / arxiv.org / abs / 1706.03762) that models a language by the occurrence probability of a word. By using this language model, a sentence vector or a summary can be obtained for a specified sentence. The trained model 121 may be included in a computer other than the information processing device 10.
[0029] The following describes the processing procedure executed by the information processing device 10. Fig. 4 is a flowchart for explaining an example of the processing procedure executed by the information processing device 10 in the first embodiment.
[0030] In step S100, the instruction receiving unit 111 waits for an instruction from the user to be input via a UI screen displayed on the user terminal 20.
[0031] Fig. 5 is a diagram showing a first display example of a UI screen in the first embodiment. As shown in Fig. 5, a UI screen 510 includes a dialogue area 511, a target document area 512, an input file area 513, an output file area 514, and the like.
[0032] The dialogue area 511 is an area where messages from the information processing device 10 to the user and messages (instructions) from the user to the information processing device 10 are displayed. For convenience, Fig. 5 shows a series of messages up to the generation of the target document, but at the time of step S100, only the message ms1 is displayed.
[0033] The generated target document is displayed in the target document area 512. Although the target document is shown in Fig. 5 for convenience, the target document is not included in the target document area 512 at the time of step S100.
[0034] The input file area 513 is an area where the file names of the files (hereinafter referred to as "input files") that store the original documents are displayed. Although the file names are included in Fig. 5 for convenience, the input file area 513 does not include the file names at the time of step S100.
[0035] The output file area 514 is an area where the file name of the file that stores the target document (hereinafter referred to as the "output file") is displayed. Although the file name is included in Fig. 5 for convenience, the output file area 514 does not include the file name at the time of step S100.
[0036] When the user inputs the file name of each original document (hereinafter referred to as the “input file name”) into the input file area 513, a message mu1 including each input file name is displayed in the dialogue area 511.
[0037] When the user further inputs a generation policy as a message mu2, the user terminal 20 transmits to the information processing device 10 a generation instruction including each input file name and the generation policy.
[0038] When the instruction receiving unit 111 receives the generation instruction (Yes in S101), the generation policy acquisition unit 112 acquires the generation policy from the generation instruction (S102). In the example of the message mu2 in Fig. 5, the generation policy is "Create explanatory materials for the sales area in slide format from the input file." The generation policy indicates that the user desires materials to explain to the sales area what kind of product it is as the target document.
[0039] Next, the document acquisition unit 113 acquires the original documents corresponding to each input file name included in the generation instruction from the document storage unit 131 (S103). In the example of message mu1 in Fig. 5, four original documents with file names "3C analysis results", "Product concept", "Sample user's comments", and "Return on investment estimate" are acquired.
[0040] Next, the summarization unit 114 generates a summary sentence for each original document (S104). For example, the summarization unit 114 causes the trained model 121 to generate a summary sentence for each original document. The number of sentences in each summary sentence may be one, or two or more.
[0041] Next, the similarity calculation unit 115 calculates the similarity between each of the multiple summaries generated for each original document and the generation policy (S105).
[0042] Next, the summary selection unit 116 selects the summaries having the top N similarities from among the plurality of summaries as summaries for generating a table of contents (S106).
[0043] Next, the table of contents generation unit 118 causes the trained model 121 to generate a table of contents based on the summary sentences selected by the summary selection unit 116 (hereinafter referred to as "selected summary sentences") (S107). Specifically, the table of contents generation unit 118 rearranges (switches) the summary sentences in the trained model 121 so that the story is understandable, and causes the trained model 121 to convert each of the rearranged summary sentences into a character string (wording) suitable for the table of contents (suitable for the table of contents).
[0044] Next, the document generation unit 117 generates a prompt for causing the trained model 121 to generate a target document based on multiple original documents (all original documents) according to the table of contents generated by the table of contents generation unit 118 in the trained model 121 (S108). The prompt refers to text that is input to the trained model 121.
[0045] Fig. 6 is a diagram showing an example of a prompt for generating a target document in the first embodiment. In the prompt shown in Fig. 6, the part other than the part enclosed in {} is a pre-prepared fixed phrase (template). Therefore, the document generating unit 117 generates a prompt corresponding to the current generation instruction by replacing the part enclosed in {} with a character string corresponding to the current generation instruction.
[0046] Specifically, the document generation unit 117 replaces {generation policy} with the generation policy acquired by the generation policy acquisition unit 112, and replaces {table of contents} with the table of contents generated by the table of contents generation unit 118 using the trained model 121. The document generation unit 117 also replaces {file name} with the file name of the original document, and replaces {text} with text extracted from the original document, for each original document.
[0047] Fig. 7 is a diagram showing an example of the contents of an original document in the first embodiment. In Fig. 7, (1) shows the contents of "3C analysis results", (2) shows the contents of "product concept", (3) shows the contents of "sample user feedback", and (4) shows the contents of "return on investment estimate".
[0048] In this case, the document generation unit 117 matches the {file name} and {text} in the prompt in Figure 6 with "3C analysis results" and the text extracted from Figure 7(1), "product concept" and the text extracted from Figure 7(2), "sample user feedback" and the text extracted from Figure 7(3), and "return on investment estimate" and the text extracted from Figure 7(4).
[0049] Next, the document generation unit 117 inputs the generated prompt to the trained model 121, thereby causing the trained model 121 to generate a document based on the prompt (S109). The document generation unit 117 acquires, as a target document, a document (text) output from the trained model 121 to which the prompt has been input. The target document is generated, for example, as a file.
[0050] Next, the display control unit 120 executes a process for outputting the target document to the target document area 512 of the UI screen 510 (S110). For example, the display control unit 120 transmits data including the target document and the file name of the file storing the target document to the user terminal 20. Based on the data, the user terminal 20 displays the target document in the target document area 512 of the UI screen 510, and also displays the file name of the target document in the output file area 514. The user terminal 20 also displays a message ms3 in the dialogue area 511 of the UI screen 510. As a result, the UI screen 510 becomes the state shown in FIG. 5.
[0051] Then, the process returns to step S100.
[0052] If the user wishes to modify the generated target document, the user can modify the target document by inputting a modification instruction into the dialogue area 511.
[0053] Fig. 8 is a diagram showing a second display example of the UI screen in the first embodiment. In Fig. 8, the same parts as in Fig. 5 are given the same reference numerals and their explanations are omitted. Fig. 8 shows a display example of the UI screen 510 at the time when the correction of the target document is completed.
[0054] When correcting the target document, the user inputs a message indicating a correction instruction. In FIG. 8, message mu3 corresponds to the message indicating a correction instruction ("The sales area will be interesting to see how it differs from other companies, so please add that information as well"). At this point in time, the state of output file area 514 is as shown in FIG. 5. User terminal 20 transmits the correction instruction (message mu3) to information processing device 10.
[0055] When the instruction receiving unit 111 of the information processing device 10 receives the correction instruction (Yes in S111), the document correction unit 119 inputs the correction instruction to the trained model 121, thereby causing the trained model 121 to execute correction of the target document in accordance with the correction instruction (S112). Note that since the trained model 121 stores the context up to that point, correction can be performed without inputting the target document again.
[0056] Next, the display control unit 120 identifies the part of the target document after correction that has been changed from the target document before correction (hereinafter, referred to as the "changed part") by comparing the entire text of the target document before correction with the entire text of the target document after correction (S113). Next, the display control unit 120 executes a process for outputting the target document after correction to the target document area 512 of the UI screen 510 (S114). For example, the display control unit 120 transmits data including the target document after correction and information indicating the changed part to the user terminal 20. The user terminal 20 displays the target document after correction in the target document area 512 of the UI screen 510 based on the data, and makes the display mode of the changed part different from other parts based on the information indicating the changed part. For example, the changed part may be highlighted. This allows the user to easily visually recognize the changed part. That is, the display control unit 120 transmits the information indicating the changed part to the user terminal 20, thereby making the display mode of the changed part different from other parts. As a result, the UI screen 510 will be as shown in FIG.
[0057] The user can subsequently instruct the generation of a new target document.
[0058] Fig. 9 is a diagram showing a third display example of the UI screen in the first embodiment. In Fig. 9, the same parts as in Fig. 5 are given the same reference numerals, and their description will be omitted. Fig. 9 shows an example in which a message mu4 indicating a new creation policy is input in the dialogue area 511, a target document according to the creation policy is displayed in the target document area 512, and the file name of the target document is added to the output file area 514. In this case, steps S102 to S110 in Fig. 4 are executed again to generate and display a new target document. Note that although the original document is the same in Fig. 9, the original document may be changed.
[0059] As described above, according to the first embodiment, a summary sentence having a relatively high similarity to the generation policy is selected from the summaries of multiple original documents, and document data according to the table of contents based on the selected summary sentences is generated as a target document. Therefore, it is possible to generate a new document from multiple existing document data without using a template.
[0060] In addition, by having the trained model 121 generate a summary of the original document, generate a table of contents based on the summary, and generate a target document based on the table of contents in a step-by-step manner, it is expected that the accuracy of the generated target document will be improved.
[0061] In addition, by limiting the summaries for which the table of contents is generated to those similar to the generation policy, it is possible to generate a target document by focusing on only the necessary information. In addition, it is possible to generate a target document that takes into account the contents of each source document, without prioritizing source documents with a large amount of information among all source documents.
[0062] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Therefore, the points not specifically mentioned may be the same as the first embodiment.
[0063] In the first embodiment, an example was shown in which the user selects an original document on a document data unit (file unit) basis, but in the second embodiment, an example is described in which the user selects one or more storage areas from among a plurality of storage areas that configure a hierarchical structure and store document data. In the second embodiment, a directory or folder (hereinafter, collectively referred to as "directory") is an example of such a storage area.
[0064] When making an input to the input file area 513 of the UI screen 510 shown in FIG. 5, the user selects, for example, any directory from a screen that displays the directory structure of the document storage unit 131.
[0065] Fig. 10 is a diagram showing an example of a screen displaying a directory structure in the second embodiment. In screen 520 in Fig. 10, area 521 is an area showing the hierarchical structure of directories in a tree format. Area 522 is an area showing directories or files immediately below the directory selected in area 521 in a list format.
[0066] 10 shows an example in which "Directory 1-1" is selected and a command is given to generate a new file (target document) from the contents of the directory. Directly below "Directory 1-1" are three directories named "Environmental Analysis," "Consideration," and "Various Questionnaires." Therefore, in this case, the names of these three directories are displayed in the input file area 513.
[0067] In this state, when a message mu2 (creation instruction) is input, the user terminal 20 transmits a creation instruction including these directory names and the creation policy to the information processing device 10.
[0068] When the instruction receiving unit 111 receives the generation instruction (Yes in S101 in Fig. 4), the process basically follows the procedure described in Fig. 4. However, in step S103, the document acquiring unit 113 acquires, as source documents, files stored in directories (hereinafter referred to as "input directories") related to each directory included in the generation instruction from the document storage unit 131 (S103).
[0069] Fig. 11 is a diagram showing an example of the configuration of an input directory. Fig. 11 shows the file names of files stored in each input directory (belonging to each input directory). In the example of Fig. 11, two files, a file "3C analysis" and a file "SWOT", are stored in the directory "Environmental analysis", a file "Product concept" is stored in the directory "Consideration", and a file "Market questionnaire survey results" is stored in the directory "Various questionnaires".
[0070] In this case, the document acquisition unit 113 acquires the document data stored in these four files as the original documents.
[0071] In step S108, the document generation unit 117 generates a prompt to cause the trained model 121 to generate a target document based on the directory storing each original document. Specifically, in order to realize generation of a target document based on the directory storing each original document, the document generation unit 117 replaces the {file name} portion of the prompt in Fig. 6 with a file name that also includes the directory name of the parent directory of each original document. As a result, the {file name} and {text} portions of each original document in the prompt in Fig. 6 are replaced as follows:
[0072] Fig. 12 is a diagram showing an example of a prompt for generating a target document in the second embodiment. In the prompt in Fig. 12, parts that may be the same as those in Fig. 6 are omitted.
[0073] 12, the {file name} portion is a path name from the input directory to which each original document belongs, such as "Environmental analysis / 3C analysis," "Environmental analysis / SWOT," and "Various surveys / market survey results." By inputting such a prompt to the trained model 121, it becomes possible to generate original documents that also take into account the directory structure to which each original document belongs.
[0074] As described above, according to the second embodiment, a target document can be generated from a summary that takes into account the structure of a directory, which is a storage area for document data, rather than from document data with a large amount of text.
[0075] The second embodiment may be combined with the first embodiment, that is, the source document may be selected in a state where the directory unit and the file unit are mixed.
[0076] Next, a third embodiment will be described. In the third embodiment, differences from the first embodiment will be described. Therefore, the points not specifically mentioned may be the same as the first embodiment.
[0077] In the third embodiment, as in the second embodiment, an example will be described in which a user selects one or more storage areas from among a plurality of storage areas that configure a hierarchical structure and store document data, although in the third embodiment, a workspace is an example of such a storage area.
[0078] FIG. 13 is a diagram for explaining a workspace in the third embodiment.
[0079] Generally, each user classifies and manages files by directories on a file system. In Fig. 13, the files managed in the file system are expressed as real data. When the real data is stored in an environment accessible to multiple users by a file sharing server or the like, each user will refer to the real data to carry out their work.
[0080] However, the information (knowledge) that each user wants to collect varies depending on the job they are responsible for and their field of interest, and knowledge cannot be obtained efficiently from the actual data (unified classification state).
[0081] Therefore, in the third embodiment, a workspace is defined as a virtual space (storage area) that manages the results of information collection by each user or the edited data thereof. In other words, a workspace is a unit that manages information indicating the results of collection of document data in the past (search results) and a collection of results of editing the collected results. Each user can efficiently acquire knowledge by referring to workspaces created by other users who share the same tasks, interests, etc. Furthermore, the quality and quantity of the workspaces will improve the more they are used, improving user convenience.
[0082] As shown in FIG. 13, a name (workspace name) is given to a workspace (a collection result of document data, etc.) by a user. One workspace contains one or more pieces of document data collected based on a viewpoint corresponding to the workspace. The one or more pieces of document data in one workspace are classified into one or more groups based on their mutual similarity. The groups are called "classes." Classification into classes is performed by converting each piece of document data belonging to the workspace into a vector using a trained model 121 or the like, and clustering the vector. In addition, one or more words determined to be relatively important in a set of document data belonging to the class using TF-IDF or the like are assigned to each class as a label for the class.
[0083] In Fig. 13, the units enclosed by dashed lines are classes. Therefore, one workspace contains one or more classes, and one class contains one or more document data (files). As a result, a hierarchical structure of workspace-class-document data is formed.
[0084] Fig. 14 is a diagram showing an example of a functional configuration of an information processing device 10 in the third embodiment. In Fig. 14, the same components as those in Fig. 3 are given the same reference numerals, and the description thereof will be omitted.
[0085] 14, the information processing device 10 further includes a document management unit 122. The document management unit 122 is realized by a process in which one or more programs installed in the information processing device 10 are caused to be executed by the processor 104.
[0086] The information processing device 10 also uses a workspace storage unit 132. The workspace storage unit 132 can be realized using, for example, the auxiliary storage device 102, or a storage device connectable to the information processing device 10 via a network. The storage device that functions as the workspace storage unit 132 may be the same as or different from the storage device that functions as the document storage unit 131.
[0087] Fig. 15 is a diagram showing an example of the configuration of the workspace storage unit 132. As shown in Fig. 15, the workspace storage unit 132 stores, for each workspace, a workspace including a workspace ID, a workspace name, a label, a creator, an updater, a query, a number of uses, an evaluation score, an associated data ID, an associated data path, an associated class label, and the like.
[0088] The workspace ID is identification information of the workspace, and is given, for example, when the workspace is generated. The workspace name is the name of the workspace input by the user, as described above. The creator is identification information (user ID or name, etc.) of the creator of the workspace. The updater is identification information (user ID or name, etc.) of the person who updated the workspace when it was updated. The query is a query (search condition) input in collecting document data that is the source of the workspace. Therefore, the query can be said to be information indicating what viewpoint the workspace is based on as a collection of document data. The number of uses is the number of times the workspace has been used (referenced). The evaluation score is an evaluation value input by a user who has referred to the workspace. For example, the evaluation score is the average value of the numerical values in a five-point evaluation. The belonging data ID is an ID given to each document data belonging to the workspace. The belonging data path is the file path of each document data. The belonging class label is a label for the class to which each document data belongs. The same belonging class label is saved for document data classified into the same belonging class in the workspace.
[0089] The document management unit 122 searches for a workspace from the workspace storage unit 132, registers (stores) document data in the workspace, and so on.
[0090] Fig. 16 is a flowchart for explaining an example of a processing procedure executed by information processing device 10 in the third embodiment. In Fig. 16, the same steps as those in Fig. 4 are given the same step numbers, and their explanations will be omitted as appropriate. In Fig. 16, the document generation process in step S121 corresponds to steps S102 to S110 in Fig. 4, and the document correction process in step S122 corresponds to steps S112 to S114 in Fig. 4.
[0091] In step S100, the instruction receiving unit 111 waits for an instruction from the user to be input via the UI screen 510 displayed on the user terminal 20.
[0092] Fig. 17 is a diagram showing a display example of a UI screen in the third embodiment. In Fig. 17, the same parts as in Fig. 5 are given the same reference numerals, and the description thereof will be omitted.
[0093] The configuration of the UI screen 510 in Fig. 17 is the same as that in Fig. 5, but the display contents of each area are different. Specifically, in the input file area 513, document data belonging to each class in a workspace selected by the user (hereinafter referred to as the "target workspace") is displayed as the original document. In the input file area 513, "Label 1" and "Label 2" are the labels of each class in the target workspace. "File X" (X is 1, 2, or 5) is the name of the document data belonging to each class. Also, the message ms11 in the dialogue area 511 indicates that the target document is to be generated from the workspace.
[0094] The target workspace may be selected, for example, via a screen such as the one shown below.
[0095] 18 is a diagram showing a display example of a workspace list screen 540. As shown in FIG.
[0096] The search condition display area 541 is an area for displaying search conditions for the workspace, and includes a query display area 5412. The query display area 5412 is an area for displaying a query as a search condition for the workspace.
[0097] The search result display area 542 is an area where a list of workspaces searched for based on the query is displayed.
[0098] For example, when a user inputs a query in the query display area 5412 and presses the execute button 5413, the user terminal 20 transmits a search request for workspaces including the query to the information processing device 10. The document management unit 122 of the information processing device 10 searches the workspace storage unit 132 for workspaces that match the query. The display control unit 120 generates screen data for a workspace list screen 540 based on the search results, and transmits the screen data to the user terminal 20. The user terminal 20 displays the workspace list screen 540 based on the screen data. The user can check a list of workspaces that match the search criteria by referring to the workspace list screen 540.
[0099] When a detail button 543 corresponding to any workspace is pressed on the workspace list screen 540, the user terminal 20 transmits a request corresponding to the detail button 543 to the information processing device 10. The request includes, for example, the workspace ID of the workspace.
[0100] When the document management unit 122 of the information processing device 10 receives the request, it acquires from the workspace storage unit 132 more detailed information than that displayed on the workspace list screen 540 for the workspace (hereinafter referred to as the "target workspace") related to the workspace ID (hereinafter referred to as the "target workspace ID") included in the request. The display control unit 120 generates screen data for a screen that displays the information (hereinafter referred to as the "workspace details screen"). When the display control unit 120 transmits the screen data to the user terminal 20, the user terminal 20 displays the workspace details screen based on the screen data.
[0101] Fig. 19 is a diagram showing a display example of a workspace details screen 550. As shown in Fig. 19, a workspace details screen 550 includes a basic information display area 551, a configuration display area 552, an associated document display area 553, a new document generation button 554, and the like.
[0102] Basic information display area 551 is an area including the contents displayed on workspace list screen 540 for the workspace to be displayed, an edit button 5511, and an evaluation button 5512.
[0103] The configuration display area 552 is an area including information indicating the relationship between a document data group belonging to a workspace to be displayed, which can be specified based on the class label and the data ID of the workspace, and the class into which the document data group is classified. In Fig. 19, an example in which three classes belong to the workspace is shown.
[0104] The belonging document display area 553 is an area including a list of document data belonging to the class (hereinafter referred to as the "target class") selected in the configuration display area 552. In Fig. 19, "No access right" is displayed for the third document data. "No access right" indicates that the logged-in user does not have the right to access the document data.
[0105] When the user presses the Create New Document button 554 on the workspace details screen 550 (FIG. 19), a UI screen 510 (FIG. 17) including each document data belonging to each class in the workspace in the input file area 513 is displayed on the user terminal 20. In this case, a target document can be generated using multiple document data belonging to the workspace as source documents.
[0106] Alternatively, the user may select one or more workspaces from which the target document is to be generated, from among the workspaces displayed in search result display area 542 on workspace list screen 540 (FIG. 18), and press new document generation button 544. When new document generation button 544 is pressed, UI screen 510 (FIG. 17) including each document data belonging to each selected workspace in input file area 513 is displayed on user terminal 20. In this case, a target document can be generated using the document data belonging to each of the multiple workspaces as a source document.
[0107] Returning to Fig. 17, when the user inputs a creation policy as a message mu11, the user terminal 20 transmits to the information processing device 10 a creation instruction including the creation policy and, for each original document displayed in the input file area 513, the workspace ID of the workspace to which the original document belongs and the belonging data ID corresponding to the original document.
[0108] The information processing device 10 executes step S121 in FIG. 16 (steps S102 to S110 in FIG. 4) in response to the generation instruction. However, in step S103, the document acquisition unit 113 acquires each original document based on the workspace ID and the belonging data ID included in the generation instruction for each original document. Specifically, the document acquisition unit 113 acquires the original document (document data) from the document storage unit 131 based on the belonging data path stored in the workspace storage unit 132 (FIG. 15) in association with the workspace ID and the belonging data ID of a certain original document. The rest is the same as the processing procedure in FIG. 4. As a result, the generated target document is displayed in the target document area 512 of the UI screen 510 (FIG. 17), and the message ms12 is displayed in the dialogue area 511.
[0109] Here, when the user inputs an instruction to register the target document to a specific class in a specific workspace, as indicated by message mu12, the user terminal 20 transmits an instruction to register the target document to the workspace, including the workspace ID of the workspace and the label name of the class, to the information processing device 10.
[0110] When the instruction receiving unit 111 receives the registration instruction (Yes in S133 in FIG. 16), the document management unit 122 determines whether the target document has already been saved in the document storage unit 131 (whether an instruction to save the target document has been issued) (S134). If the target document has not been saved (No in S134), the display control unit 120 transmits a message to the user terminal 20 to prompt the user to save the target document (S136). When the user terminal 20 receives the message, it displays the message in the dialogue area 511 of the UI screen 510. In FIG. 17, the message ms13 corresponds to the message.
[0111] When the user inputs the path name of the destination where the target document is to be saved as shown in message mu13, the user terminal 20 transmits to the information processing device 10 an instruction to save the target document, including the path name.
[0112] When the instruction receiving section 111 receives the save instruction (Yes in S131), the document management section 122 saves a file containing the target document in the folder corresponding to the path name in the document storage section 131 (S132).
[0113] Thereafter, when the user again inputs a message similar to message mu12 and the instruction receiving unit 111 receives an instruction to register the target document to the workspace (Yes in S133), since the target document has already been saved (Yes in S134), the document management unit 122 registers (stores) the target document under the workspace ID and class specified in the registration instruction (S135). Specifically, the document management unit 122 adds a row including the target document's belonging data ID, belonging data path, and belonging class label to the record corresponding to the workspace in the workspace storage unit 132 (FIG. 15).
[0114] As described above, according to the third embodiment, a target document can be generated by inputting a workspace. A workspace is a collection of document data that is organized according to the work or field of interest of the creator of the workspace. Therefore, a target document can be generated based on such a collection of document data.
[0115] Furthermore, a target document can be generated from a summary that takes into account the structure of a workspace, which is a storage area for document data, rather than from document data with a large amount of text.
[0116] The third embodiment may be implemented in combination with the first and second embodiments.
[0117] Next, a fourth embodiment will be described. In the fourth embodiment, differences from the first embodiment will be described. Therefore, the points not specifically mentioned may be the same as the first embodiment.
[0118] In the fourth embodiment, a template is input instead of a generation policy, and a target document is generated according to the template based on the source document. That is, in the fourth embodiment, a template refers to a model of a target document.
[0119] Fig. 20 is a diagram showing an example of a functional configuration of an information processing device 10 in the fourth embodiment. In Fig. 20, the same reference numerals are given to the same or corresponding parts as in Fig. 3, and the description thereof will be omitted as appropriate.
[0120] 20, the information processing device 10 does not have the generation policy acquisition unit 112, the summarization unit 114, the similarity calculation unit 115, the summary selection unit 116, and the table of contents generation unit 118, but instead has a template acquisition unit 123, a necessary information identification unit 124, and a necessary information collection unit 125. Each of these units is realized by a process in which one or more programs installed in the information processing device 10 are executed by the processor 104.
[0121] The template acquisition unit 123 acquires a template based on a generation instruction.
[0122] Fig. 21 is a diagram showing an example of a template in the fourth embodiment. In Fig. 21, the part other than the part enclosed in {} is a fixed phrase (model) prepared in advance, and the part enclosed in {} is a part generated based on the original document.
[0123] The necessary information identification unit 124 uses the trained model 121 to identify information that needs to be collected (hereinafter referred to as "necessary information") in order to generate a target document based on a template (i.e., to fill in the {}).
[0124] The necessary information collection unit 125 uses the trained model 121 to collect necessary information from the original document.
[0125] In the fourth embodiment, the document generation unit 117 causes the trained model 121 to generate a target document based on necessary information and a template.
[0126] Fig. 22 is a flowchart for explaining an example of a processing procedure executed by the information processing device 10 in the fourth embodiment. In Fig. 22, the same steps as those in Fig. 4 are given the same step numbers, and their explanations are omitted as appropriate. In Fig. 22, step S102 in Fig. 4 is replaced with S102a. Also, steps S104 to S109 in Fig. 4 are replaced with S141 to S143.
[0127] In step S102a, the template acquisition unit 123 acquires a template based on the generation instruction.
[0128] The template is specified, for example, on a UI screen 510 as shown in FIG.
[0129] Fig. 23 is a diagram showing a display example of a UI screen in the fourth embodiment. In Fig. 23, the same or corresponding parts as those in Fig. 8 are given the same reference numerals, and their explanation will be omitted as appropriate. Fig. 23 shows a state in which the target document has been corrected after it has been generated. Therefore, the dialogue area 511 in Fig. 23 shows the history of dialogue between the user and the information processing device 10 during the process of generating and correcting the target document.
[0130] The input file area 513 specifies the file names of the original document and the template. In this state, when the user inputs a message such as that shown in message mu22, the user terminal 20 transmits a generation instruction to the information processing device 10. The message mu22 indicates that an "Expense Reduction Utilization Application Form" is to be generated as the target document, and that the template is an "Expense Reduction Utilization Application Form Template." Therefore, the user terminal 20 transmits to the information processing device 10 the names of each file contained in the input file area 513, and a generation instruction indicating that the "Expense Reduction Utilization Application Form" is a template. In step S102a, the template acquisition unit 123 acquires the template stored in a file related to the file name of the template included in such a generation instruction ("Expense Reduction Utilization Application Form Template").
[0131] In step S141, the necessary information identifying unit 124 causes the trained model 121 to identify the necessary information based on the template acquired by the template acquiring unit 123.
[0132] FIG. 24 is a diagram showing an example of input and output of the trained model 121 in identifying necessary information. In FIG. 24, (1) shows a prompt that the necessary information identifying unit 124 inputs to the trained model 121 to identify necessary information. In the prompt, the {template} portion is filled in with the entire text of the template shown in FIG. 21. The portion other than {template} is a fixed phrase. The necessary information identifying unit 124 generates such a prompt and inputs the prompt to the trained model 121.
[0133] In FIG. 24, (2) is the output from the trained model 121 in response to the prompt, and indicates the result of identifying the necessary information. The trained model 121 identifies the necessary information from the template and outputs the result. As is clear from FIG. 21, a character string indicating the content of the information to be written in the {} is written within the {} of the template. The trained model 121 can identify the necessary information based on the character string as well.
[0134] Furthermore, by having the trained model 121 identify the necessary information in the template, even when a template is specified, advance preparation such as tags can be eliminated.
[0135] Next, the necessary information collecting unit 125 causes the trained model 121 to collect the necessary information identified by the necessary information identifying unit 124 from the original document (S142).
[0136] FIG. 25 is a diagram showing an example of input and output of the trained model 121 in collecting necessary information. In FIG. 25, (1) shows a prompt that the necessary information collecting unit 125 inputs to the trained model 121 to collect necessary information. In the prompt, the {file name} and {text} parts are filled in with the file name of the original document and the text extracted from the original document for each original document. The other parts are fixed phrases. The necessary information collecting unit 125 generates such a prompt and inputs the prompt to the trained model 121.
[0137] In Figure 25, (2) is the output from the trained model 121 for the prompt, and shows the necessary information collected from the original document.
[0138] It should be noted that Fig. 25 is based on the case where the original document includes the contents shown in Fig. 26. Fig. 26 is a diagram showing an example of the contents of the original document in the fourth embodiment.
[0139] In FIG. 26, (1) shows the contents of the first original document, "X Company's Quotation Document," and (2) shows the contents of the second original document, "Budget Control Number."
[0140] Next, the document generation unit 117 causes the trained model 121 to generate a target document based on a plurality of original documents (all original documents) according to the template (S143). For example, when the document generation unit 117 inputs a prompt such as "Please fill in the template provided earlier with the collected information to complete the document" to the trained model 121, the trained model 121 outputs the target document.
[0141] Fig. 27 is a diagram showing an example of a target document in the fourth embodiment. Fig. 27 shows a target document generated by applying information based on the necessary information shown in Fig. 25(2) to the portion in {} of the template shown in Fig. 21.
[0142] The other processing procedures may be similar to those in the first embodiment.
[0143] The fourth embodiment may be implemented in combination with each of the above embodiments.
[0144] As described above, according to the fourth embodiment, a document can be generated according to a template specified by a user. In this case, by having the trained model 121 execute step-by-step processes of identifying necessary information, collecting necessary information, and generating a target document, it is expected that the accuracy of the generated target document with respect to the contents desired by the user can be improved.
[0145] Also, unlike Patent Document 1, there is no need to prepare tags in advance.
[0146] Each function of the above-described embodiment can be realized by one or more processing circuits. In this specification, the term "processing circuit" includes a processor programmed to execute each function by software, such as a processor implemented by an electronic circuit, and a device such as an ASIC (Application Specific Integrated Circuit), a DSP (digital signal processor), an FPGA (field programmable gate array), or a conventional circuit module designed to execute each function described above.
[0147] Additionally, the devices described in the example are merely representative of one of several computing environments for implementing the embodiments disclosed herein.
[0148] In one embodiment, information processing apparatus 10 includes multiple computing devices, such as a server cluster, configured to communicate with each other over any type of communications link, including a network, shared memory, etc., and may perform the processes disclosed herein.
[0149] Although the embodiment of the present invention has been described in detail above, the present invention is not limited to such specific embodiment, and various modifications and variations are possible within the scope of the gist of the present invention described in the claims.
[0150] For example, aspects of the present invention are as follows. <1> a generation policy acquisition unit for acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition unit for acquiring the plurality of existing document data; a summarizing unit for generating a summary of each of the existing document data; a similarity calculation unit that calculates a similarity between each of the plurality of abstract sentences and the character string; a summary selection unit that selects a portion of the plurality of summary sentences based on the degree of similarity between the plurality of summary sentences; a document generation unit that causes a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; 13. An information processing device comprising: <2> a table of contents generation unit that aligns the part of the summary sentences to a trained model and causes the trained model to convert the aligned summary sentences into character strings suitable for a table of contents; characterized in that <1> The information processing device described. <3> a document correction unit that causes the trained model to perform corrections on the new document data in accordance with a correction instruction from a user; a display control unit that changes a display mode of a portion of the new document data after correction that has been changed from the new document data before correction from a display mode of other portions; characterized in that <1> or <2> The information processing device described. <4> the document acquisition unit acquires the existing document data from each of one or more storage areas designated by a user among a plurality of storage areas that store document data in a hierarchical structure; Characterized by <1> ~ <3> 13. The information processing device according to claim 12 . <5> a document management unit for storing the new document data in the storage area designated by a user; characterized in that <4> The information processing device described. <6> The document generation unit causes the trained model to generate the new document data based on the storage area that stores the plurality of existing document data. Characterized by <4> or <5> The information processing device described. <7> a template acquisition unit for acquiring a template for new document data based on a plurality of existing document data; a document acquisition unit for acquiring the plurality of existing document data; a necessary information identification unit that identifies, using a trained model, information that needs to be collected in order to generate the new document data using the template; A necessary information collection unit that collects the information from the plurality of existing document data using the trained model; a document generation unit that causes the trained model to generate the new document data based on the template and the information; 13. An information processing device comprising: <8> a generation policy acquisition unit for acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition unit for acquiring the plurality of existing document data; a summarizing unit for generating a summary of each of the existing document data; a similarity calculation unit that calculates a similarity between each of the plurality of abstract sentences and the character string; a summary selection unit that selects a portion of the plurality of summary sentences based on the degree of similarity between the plurality of summary sentences; a document generation unit that causes a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; An information processing system comprising: <9> A generation policy acquisition step of acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition step for acquiring the plurality of existing document data; a summarization step for generating a summary of each of the existing document data; a similarity calculation step of calculating a similarity between each of the plurality of abstract sentences and the character string; a summary selection step of selecting a portion of the plurality of summary sentences based on the similarity between the plurality of summary sentences; a document generation step of causing a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; An information processing method characterized by being executed by a computer. <10> A generation policy acquisition step of acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition step for acquiring the plurality of existing document data; a summarization step for generating a summary of each of the existing document data; a similarity calculation step of calculating a similarity between each of the plurality of abstract sentences and the character string; a summary selection step of selecting a portion of the plurality of summary sentences based on the similarity between the plurality of summary sentences; a document generation step of causing a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; A program characterized by causing a computer to execute the above. [Explanation of symbols]
[0151] 10. Information processing device 20 User terminal 100 Drive device 101 Recording media 102 Auxiliary storage 103 Memory device 104 processors 105 Interface device 111 Instructions Reception Department 112 Generation policy acquisition unit 113 Document Acquisition Department 114 Summary 115 Similarity calculation part 116 Summary Selection Section 117 Document Generation Unit 118 Table of Contents Generator 119 Document Correction Section 120 Display control unit 121 trained models 122 Document Management Department 123 Template Acquisition Department 124 Necessary information identification department 125 Necessary Information Collection Department 131 Document storage unit 132 Workspace storage B Bus [Prior art documents] [Patent documents]
[0152] [Patent Document 1] Special Publication No. 2022-547895
Claims
1. a generation policy acquisition unit for acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition unit for acquiring the plurality of existing document data; a summarizing unit for generating a summary of each of the existing document data; a similarity calculation unit that calculates a similarity between each of the plurality of abstract sentences and the character string; a summary selection unit that selects a portion of the plurality of summary sentences based on the degree of similarity between the plurality of summary sentences; a document generation unit that causes a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; 13. An information processing device comprising:
2. a table of contents generation unit that aligns the part of the summary sentences to a trained model and causes the trained model to convert the aligned summary sentences into character strings suitable for a table of contents; 2. The information processing apparatus according to claim 1, further comprising:
3. a document correction unit that causes the trained model to perform corrections on the new document data in accordance with a correction instruction from a user; a display control unit that changes a display mode of a portion of the new document data after correction that has been changed from the new document data before correction from a display mode of other portions; 2. The information processing apparatus according to claim 1, further comprising:
4. the document acquisition unit acquires the existing document data from each of one or more storage areas designated by a user among a plurality of storage areas that store document data in a hierarchical structure; 2. The information processing apparatus according to claim 1,
5. a document management unit for storing the new document data in the storage area designated by a user; 5. The information processing apparatus according to claim 4, further comprising:
6. The document generation unit causes the trained model to generate the new document data based on the storage area that stores the plurality of existing document data.
5. The information processing apparatus according to claim 4.
7. a template acquisition unit for acquiring a template for new document data based on a plurality of existing document data; a document acquisition unit for acquiring the plurality of existing document data; a necessary information identification unit that identifies, using a trained model, information that needs to be collected in order to generate the new document data using the template; A necessary information collecting unit for collecting information from the plurality of existing document data using the trained model; a document generation unit that causes the trained model to generate the new document data based on the template and the information; 13. An information processing device comprising:
8. a generation policy acquisition unit for acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition unit for acquiring the plurality of existing document data; a summarizing unit for generating a summary of each of the existing document data; a similarity calculation unit that calculates a similarity between each of the plurality of abstract sentences and the character string; a summary selection unit that selects a portion of the plurality of summary sentences based on the degree of similarity between the plurality of summary sentences; a document generation unit that causes a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; An information processing system comprising:
9. A generation policy acquisition step of acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition step for acquiring the plurality of existing document data; a summarization step for generating a summary of each of the existing document data; a similarity calculation step of calculating a similarity between each of the plurality of abstract sentences and the character string; a summary selection step of selecting a portion of the plurality of summary sentences based on the similarity between the plurality of summary sentences; a document generation step of causing a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; An information processing method characterized by being executed by a computer.
10. A generation policy acquisition step of acquiring a character string indicating a generation policy of new document data based on a plurality of existing document data; a document acquisition step for acquiring the plurality of existing document data; a summarization step for generating a summary of each of the existing document data; a similarity calculation step of calculating a similarity between each of the plurality of abstract sentences and the character string; a summary selection step of selecting a portion of the plurality of summary sentences based on the similarity between the plurality of summary sentences; a document generation step of causing a trained model to generate the new document data based on the plurality of existing document data in accordance with a table of contents based on the portion of the summary sentences; A program characterized by causing a computer to execute the above.
Citation Information
Patent Citations
User-customized automatic document creation method, device therefor, and server
JP2022547895A