Text generation method, text generation device

The text generation method and device address the limitation of existing technologies by generating summaries based on specified logical structures, ensuring high accuracy and adaptability across various content types through user-selectable templates and optimization, enhancing the flexibility and relevance of generated text.

JP7875796B2Active Publication Date: 2026-06-18HITACHI LTD
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
JP · JP
Patent Type
Patents
Current Assignee / Owner
HITACHI LTD
Filing Date
2022-12-14
Publication Date
2026-06-18

AI Technical Summary

Technical Problem

Existing text generation technologies, such as those described in Patent Document 1, are unable to generate text based on a specified logical structure, limiting their applicability and accuracy in generating summaries for diverse content types.

Method used

A text generation method and device that utilizes a pre-trained model to process input text and logical structures, incorporating a language information acquisition process, logical text acquisition, and text output process to generate summaries that adhere to a specified logical structure, using models like GPT, GPT-2, GPT-3, T5, or BART, and employing a stack-based algorithm to convert logical templates into logical texts.

Benefits of technology

Enables high-accuracy text generation that reflects the logical structure and author's intent, allowing for customizable summaries that can handle diverse content types and logical structures, including sustainability reports and academic papers, with user-selectable templates and optimization processes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 0007875796000001
    Figure 0007875796000001
  • Figure 0007875796000002
    Figure 0007875796000002
  • Figure 0007875796000003
    Figure 0007875796000003
Patent Text Reader

Abstract

To generate a text based on a designated logical structure.SOLUTION: The present invention is directed to a text generating method executed by a computer having a storage unit for storing a text generation model previously trained. The method has a language information acquiring process of acquiring language information indicating a processing-target sentence, a logical text acquiring process of acquiring a logical text, and a text output process of inputting the processing target sentence and the logical text into the text generation model to obtain an output text. The text generation model uses, as an input, the logical text and the processing target sentence which indicate a specified logical structure as logical structure between sentences and uses, as an output, text obtained from the processing target sentence and changed in accordance with the specified logical structure.SELECTED DRAWING: Figure 1
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a text generation method and a text generation device.

Background Art

[0002] As represented by automatic summarization, a technology for converting a given text into an appropriate text for a reader, that is, a text generation technology, is known. With the development of deep learning in recent years, the accuracy of text generation technologies such as automatic summarization and machine translation has been improving. In particular, a text generation language model, which is a language model equipped with a text generation function, has emerged, and the accuracy of text generation has been further improved. Along with the improvement in the accuracy of text generation, text generation technologies have begun to be applied to generation targets that have not been addressed before. For example, automatic summarization has often been targeted at generating news headlines, but in recent years, it has been targeted at summarizing papers and patent documents, and it is expected that the targets will further expand in the future, such as reports issued by companies. Patent Document 1 discloses a computer-implemented method for translating an input text in a first language into an output text in a second language, the method comprising the steps of generating an input logical form based on the input text, selecting a set of one or more transformation mappings that match at least a portion of the input logical form based on a predefined measure, combining the set of transformation mappings with a target logical form, and generating an output text based on the target logical form.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] The invention described in Patent Document 1 cannot generate text based on a specified logical structure. [Means for solving the problem]

[0005] A first aspect of the present invention is a text generation method executed by a computer having a memory unit that stores a pre-trained text generation model, wherein the text generation model takes a logical text indicating a specified logical structure which is the logical structure between sentences and a sentence to be processed as input, outputs an output text which is text according to the specified logical structure from the sentence to be processed, and includes a language information acquisition process to acquire language information indicating the sentence to be processed, a logical text acquisition process to acquire the logical text, and a text output process to input the sentence to be processed and the logical text to the text generation model and obtain the output text. The text output process includes an optimization process that optimizes the logical text for the text generation model, wherein the logical text optimized by the optimization process is input to the text generation model. . A text generation apparatus according to a second aspect of the present invention includes: a language information acquisition unit that acquires language information indicating a document to be processed; a logical text acquisition unit that acquires logical text which is a string indicating a specified logical structure which is the logical structure between documents; a storage unit that takes the document to be processed and the logical text as input and stores a text generation model trained to generate text from the document to be processed according to the specified logical structure; and a text generation unit that uses the text generation model to acquire text from the document to be processed according to the specified logical structure. The system comprises a logical text modification unit that optimizes the logical text for the text generation model, and the text generation unit obtains the text by inputting the logical text optimized by the logical text modification unit and the text to be processed. . [Effects of the Invention]

[0006] According to the present invention, text can be generated based on a specified logical structure. [Brief explanation of the drawing]

[0007] [Figure 1] Functional configuration diagram of the text generation device in the first embodiment [Figure 2] Hardware configuration diagram of the text generation device [Figure 3] Diagram showing the operation of the text generation device during the learning phase. [Figure 4] This figure shows the operation of the inference phase of the text generation device in the first embodiment. [Figure 5] Diagram showing specific examples of summaries, logical templates, and logical texts. [Figure 6] This figure shows the operation of the inference phase of the text generation device in the second embodiment. [Figure 7] This figure shows the operation of the inference phase of the text generation device in the third embodiment. [Figure 8] This figure shows the operation of the inference phase of the text generation device in the fourth embodiment. [Figure 9] This figure shows the operation of the inference phase of the text generation device in a modified example of the fourth embodiment. [Figure 10] This figure shows the operation of the inference phase of the text generation device in the fifth embodiment. [Figure 11] This figure shows the operation of the inference phase of the text generation device in the sixth embodiment. [Figure 12] Conceptual diagram showing input and output to the text generation unit. [Modes for carrying out the invention]

[0008] —First Embodiment— A first embodiment of the text generation method and text generation device will be described below with reference to Figures 1 to 5.

[0009] Figure 1 is a functional configuration diagram of the text generation device 1. The text generation device 1 comprises a language information acquisition unit 11, a logical template generation unit 12, a logical text generation unit 13, a text generation unit 14, a logical template acquisition unit 15, and a logical text acquisition unit 16. The storage unit 6 provided in the text generation device 1 stores the input text 61, a logical template 62, a logical text 63, a text generation model 64, and a summary text 65. However, not all of the data stored in the storage unit 6 in Figure 1 is stored before the text generation device 1 starts operating. Hereafter, the person using the text generation device 1 will be referred to as the "user". The user may be the same person, or it may be a different person depending on the situation.

[0010] The input text 61 is a pre-generated text containing multiple sentences. The input text 61 is always created outside of the text generation device 1. Since the input text 61 is the text that the text generation device 1 processes, it is also called the "processing target text". In this embodiment, the input text 61 is described as text composed of strings, but the input text 61 can use various types of data such as audio, documents, and drawings. In this case, the documents may be electronic documents or physically recorded documents such as paper or boards. Similarly, the drawings may be electronic drawings or physically recorded documents such as paper or boards. Furthermore, documents and drawings may contain various information such as strings, figures, and photographs. The language of the input text 61 is not particularly limited, and various languages ​​can be used. Also, the strings may include not only language but also figures and symbols.

[0011] The logical template 62 is a graph that shows the logical structure between sentences. Here, "graph" is a concept that encompasses tree structures, representing sentences as nodes and the logical relationships between sentences as edges. It is a directed or undirected graph with no restrictions on the number of root nodes or the presence or absence of loops. The logical template 62 may be generated by the text generation device 1, as described later, or it may be obtained from an external source. Because the logical template 62 is a graph, it is easy for humans to understand.

[0012] The logical text 63 is a character string obtained by converting a logical template 62, which is a graph, by a predetermined method. The logical text 63 indicates the logical structure between sentences, similar to the logical template 62. The logical text 63 in the present embodiment is created using the logical template 62 in the text generation device 1.

[0013] The text generation model 64 is a language model that takes an input sentence 61 and a logical text 63 as inputs and generates a summary sentence 65. The parameters of the text generation model 64 are adjusted in the learning phase, as will be described later. The summary sentence 65 is text generated based on the input sentence 61. Here, it is called "summary sentence" for convenience, but it does not necessarily have to be a summary. The summary sentence 65 is generated outside the text generation device 1 in the learning phase and is generated by the text generation model 64 in the inference phase.

[0014] There are a plurality of input sentences 61 and summary sentences 65 respectively, and there is a summary sentence 65 corresponding to each input sentence 61. The summary sentence 65 corresponds to the input sentence 61, but it is a conceptual existence and there is no clear correct answer. For example, in the inference phase, the text generation model 64 generates a first summary sentence 65-1 based on the first input sentence 61-1 and the first logical text 63-1. If the first logical text 63-1 is changed, the generated first summary sentence 65-1 will also be changed.

[0015] The language information acquisition unit 11 reads the input sentence 61 from outside the text generation device 1 or from the storage unit 6. The logical template generation unit 12 acquires the summary sentence 65 from outside the text generation device 1 or from the storage unit 6 and generates the logical template 62 based on the summary sentence 65. The logical text generation unit 13 acquires the logical template 62 from the storage unit 6 or the logical template generation unit 12 and generates the logical text 63 based on the logical template 62.

[0016] In the inference phase, the text generation unit 14 acquires the input text 61 and logical text 63, and generates a summary text 65 using the text generation model 64. In the learning phase, the text generation unit 14 acquires the input text 61, logical text 63, and summary text 65, and updates the parameters of the text generation model 64. The logical template acquisition unit 15 acquires the logical template 62 from outside the text generation device 1 or from the storage unit 6 and outputs it to the logical text acquisition unit 16. The logical text acquisition unit 16 acquires the logical text 63 generated by the logical text generation unit 13 and outputs it to the text generation unit 14.

[0017] The text generation unit 14 can generate text with high accuracy by using a pre-trained language model employing a Transformer architecture such as GPT, GPT-2, GPT-3, T5, and BART. However, the type of text generation model 64 is not limited to these, and LSTM, Transformer, ELMo, or BERT can also be used. As a method of inputting the input sentence 61 and logical text 63 to the text generation model 64, for example, it is conceivable to insert special characters between the input sentence 61 and the logical text 63 and input them together. However, the method of combination and input may take any form as long as it does not depart from the spirit of the invention.

[0018] The output unit 17 outputs the summary text 65 output by the text generation unit 14 to the outside of the text generation device 1. The output of the summary text 65 by the output unit 17 can take various forms, and may be output directly as text, images, or audio, or indirectly as text, images, or audio via communication.

[0019] Figure 2 is a hardware diagram of the text generation device 1. The text generation device 1 includes a CPU 41 which is a central processing unit, a ROM 42 which is a read-only storage device, a RAM 43 which is a read-write storage device, an input device 44 and an output device 45 which are user interfaces, and a communication device 46. The CPU 41 performs the various calculations mentioned above by loading the program stored in the ROM 42 into the RAM 43 and executing it.

[0020] The text generation device 1 may be implemented using a rewritable logic circuit such as an FPGA (Field Programmable Gate Array) or an application-specific integrated circuit such as an ASIC (Application Specific Integrated Circuit) instead of the combination of CPU 41, ROM 42, and RAM 43. Alternatively, the text generation device 1 may be implemented using a different configuration, such as a combination of CPU 41, ROM 42, RAM 43 and FPGA, instead of the combination of CPU 41, ROM 42, and RAM 43.

[0021] The input device 44 is, for example, a keyboard, mouse, or pressure-sensitive panel, and accepts input from the user. The output device 45 is a display or speaker, and presents information to the user. The communication device 46 enables wired or wireless communication.

[0022] Figure 3 shows the operation of the text generation device 1 during the learning phase. The data stored in the memory unit 6 is the same as in Figure 1. However, data that is not read from the memory unit 6 during the learning phase is indicated by a dashed line. In other words, data indicated by a dashed line does not need to be stored in the memory unit 6 at the start of the learning phase. For example, the logical template 62 used is one generated by the logical template generation unit 12, and the logical text 63 used is one generated by the logical text generation unit 13. Also in Figure 3, information input from outside the text generation device 1 is indicated by a dashed arrow.

[0023] The language information acquisition unit 11 acquires the input text 61 from an external source or the storage unit 6 and outputs it to the text generation unit 14. The logical template generation unit 12 acquires the summary text 65 from an external source or the storage unit 6, generates a logical template 62, and outputs it to the logical template acquisition unit 15. However, the input text 61 acquired by the language information acquisition unit 11 and the summary text 65 acquired by the logical template generation unit 12 are in a corresponding relationship. The logical template acquisition unit 15 acquires the logical template 62 generated by the logical template generation unit 12 and outputs it to the logical text generation unit 13. The logical text generation unit 13 acquires the logical template 62 output by the logical template acquisition unit 15, generates logical text 63, and outputs it to the logical text acquisition unit 16.

[0024] The logical text acquisition unit 16 acquires the logical text 63 generated by the logical text generation unit 13 and outputs it to the text generation unit 14. The text generation unit 14 uses the logical text 63, the input document 61, and the summary document 65 to update the parameters of the text generation model 64, i.e., to perform learning processing. In the learning phase, the logical text 63, the input document 61, and the summary document 65 input to the text generation model 64 correspond to each other, and the parameters are updated so that the summary document 65 is output for each input of the logical text 63 and the input document 61.

[0025] Figure 4 shows the operation of the text generation device 1 during the inference phase. The data stored in the memory unit 6 is the same as in Figure 1. However, data that is not read from the memory unit 6 during the inference phase is indicated by a dashed line. Specifically, the logical text 63 and the summary text 65 do not need to be stored in the memory unit 6 at the start of the inference phase. Also in Figure 4, information input from outside the text generation device 1 is indicated by a dashed arrow.

[0026] The language information acquisition unit 11 acquires the input text 61 from an external source or the storage unit 6 and outputs it to the text generation unit 14. The logical template acquisition unit 15 acquires the logical template 62 from an external source or the storage unit 6 and outputs it to the logical text generation unit 13. The logical text generation unit 13 acquires the logical template 62 output by the logical template acquisition unit 15, generates the logical text 63, and outputs it to the logical text acquisition unit 16. The logical text acquisition unit 16 acquires the logical text 63 generated by the logical text generation unit 13 and outputs it to the text generation unit 14. The text generation unit 14 inputs the logical text 63 and the input text 61 into the text generation model 64 to obtain the output summary text 65.

[0027] Figure 5 shows specific examples of summary sentences 65, logical templates 62, and logical texts 63. Figure 5 shows two sets of specific examples of summary sentences 65, logical templates 62, and logical texts 63. The summary sentences 65, logical templates 62, and logical texts 63 shown in the left half of Figure 5 correspond to each other. Similarly, the summary sentences 65, logical templates 62, and logical texts 63 shown in the right half of Figure 5 also correspond to each other.

[0028] Summary 65 in Example 1 is a summary of a paper and is expressed in text. In this example, the text is divided into sentences, and the order of each sentence is indicated at the beginning of each sentence. The first sentence describes the paper's proposal, the second sentence describes the background leading to the implementation, the third sentence describes the implementation, and the fourth sentence describes the results.

[0029] The logical template 62 in Example 1 is a graph consisting of four nodes and three edges, specifically a tree structure. Each node stores a statement label. The "Proposal" statement in Example 1 states the main argument of the entire text and is represented as the root node in the logical template 62. In Example 1, nodes other than the root node are represented as child nodes that support the argument or implementation. For example, the third sentence in the summary sentence 65 of Example 1 is labeled as an implementation and logically supports the first sentence, which is labeled as an argument.

[0030] Summary 65 in Example 2 is a summary of a sustainability report issued by a company. In Example 2, the text of summary 65 is also divided into sentences, with the order of each sentence indicated at the beginning. The first sentence makes an environmental claim, the second sentence provides a quantitative explanation for the claim in the first sentence, and the third sentence makes an environmental claim that differs from the first sentence.

[0031] When the logical structure of this summary 65 is analyzed using a certain statistical or rule-based method, such as argumentative structure analysis, the logical template 62 of Example 2 is obtained. This logical template 62 is a graph in which the logical structure is composed of three nodes and one edge. In the logical template 62 shown in Figure 5, the unit of a node is a sentence, but the unit of a node may also be a phrase or a sentence, as long as it is a valid element that constitutes the logical structure. Furthermore, the logical template 62 only needs to have a graph structure, and may be a tree structure with a single root node as in Example 1, or it may have independent nodes as in "3. Environmental Claim" in Example 2.

[0032] The logical text generation unit 13 generates logical text 63 using an algorithm that converts the logical template 62 into a text representation. This algorithm is implemented, for example, by representing the nodes and edges of the logical template 62 using a stack. Hereinafter, this stack-based algorithm will be referred to as the "stack algorithm".

[0033] The stack algorithm includes two operations: a PUSH operation to insert a node into the stack, and a POP operation to remove a node and assign an edge to it. To process the logical template 62 shown in Example 1, we first perform a node search of the graph contained in the logical template 62. The node search starts from the root node where no outflow edges exist. In subsequent searches, unless otherwise specified, the order in which statements appear will be the priority of the search.

[0034] The stack algorithm will be explained with reference to the logical text 63 shown in Figure 5. First, the stack and logical text 63 are initialized to empty. In Example 1, the root node of logical template 62 is the first "proposal" statement, so the start of logical text 63 is the proposal node, which is "1 <proposal> <node>The following is added to logical text 63. Additionally, the "Proposal" node is added to the stack by a PUSH operation. Here, the numbers at the beginning of each line in logical text 63 indicate the order in which the statements of the PUSHed node appear, and <> are special symbols representing labels. <node>is a special label that indicates the preceding string is the label of a node. By using such special symbols and special labels, it becomes possible to teach the text generation language model that the logical text 63 explicitly indicates a logical structure.

[0035] Next, among the child nodes of the root node of logical template 62 in Example 1, the node with the earliest statement appearance order is the "implementation" of the third statement. The same process as the previous PUSH operation is performed. Next, the same process as the aforementioned PUSH operation is performed on the "background" of the second statement. At this point, the stack contains "proposal," "implementation," and "background." Here, since the "background" of the second statement has no child nodes, a POP operation is performed. This POP operation retrieves "background" and "implementation" in order.

[0036] The two nodes extracted in this order are given an edge due to the argumentative relationship of "additional information," and therefore edge information is added to logical text 63. Specifically, "<additional information> <edge>Add the following to logical text 63. Here, <edge>is a special token indicating that the preceding string is the label of an edge. Due to the aforementioned POP operation, only "Proposal" remains in the stack, and we continue to search for unexplored child nodes of Proposal. Finally, the fourth sentence, "Result," is processed in the same way as the aforementioned PUSH operation, and this POP operation adds the "Support" edge to the logical text. Finally, when there are no more explored nodes, the conversion of logical template 62 to logical text 63 is completed.

[0037] According to the first embodiment described above, the following effects and advantages can be obtained. (1) The text generation method performed by the text generation device 1, which includes a storage unit 6 that stores a pre-trained text generation model 64, is as follows: The text generation model 64 takes a logical text 63 that indicates a specified logical structure, which is the logical structure between sentences, and an input sentence 61 as input, and outputs a summary sentence 65, which is text that conforms to the specified logical structure from the input sentence 61. The text generation method performed by the text generation device 1 includes a language information acquisition process performed by a language information acquisition unit 11 that acquires language information indicating the input sentence 61, a logical text acquisition process performed by a logical text acquisition unit 16 that acquires the logical text 63, and a text output process performed by a text generation unit 14 that inputs the input sentence 61 and the logical text 63 to the text generation model 64 and obtains a summary sentence 65. Therefore, the text generation method performed by the text generation device 1 can generate text based on the logical structure indicated by the logical text 63, i.e., a summary sentence 65, with high accuracy. Specifically, it is as follows.

[0038] For example, by inputting sustainability reports with different logical structures for each company as input text 61, and inputting logical text 63 with a logical structure that is easy for the reader to understand, the text generation unit 14 can obtain summaries 65 of multiple sustainability reports that are easy to compare. Furthermore, it can handle cases where the preferred logical structure differs depending on the technical field, even for the same academic paper, or where the type of input text 61 and the purpose of the summary 65 differ. In addition, by inputting logical text 63 that reflects the intentions of the author of the original input text 61, it is possible to generate a summary 65 that reflects the author's intentions not only in content but also in logical structure.

[0039] (Variation 1) In the first embodiment described above, the input text 61 was string information. However, the input text 61 may be any language information indicating the text to be processed by the text generation model 64, and may be data other than strings, such as audio data. In this case, the language information acquisition unit 11 may have a speech recognition function to convert audio data into strings, or the language information acquisition unit 11 may not have a speech recognition function if the text generation unit 14 supports audio input.

[0040] Furthermore, if the text generation unit 14 supports voice input, the logical template 62 may be stored in the storage unit 6 as voice data and directly input to the text generation unit 14, or the logical template 62 may be stored in the storage unit 6 as text data and the logical template acquisition unit 15 or logical text acquisition unit 16 may be equipped with a voice output function. In addition, the text generation device 1 may support voice input by combining a first language model that takes logical text 63, which is a string, as input, and a second language model that takes input sentence 61, which is voice data, as input.

[0041] (Modification 2) In the first embodiment described above, the text generation device 1 performed both learning and inference. However, the text generation device 1 only needs to be able to perform at least the inference phase and does not need to perform the learning phase. In this case, the text generation device 1 does not need to include a logic template generation unit 12 that is used only in the learning phase.

[0042] (Variation 3) In the first embodiment described above, the text generation device 1 stored a logical template 62 in the memory unit 6 during the inference phase. However, the memory unit 6 may also store logical text 63, in which case the text generation device 1 does not need to include a logical text generation unit 13.

[0043] (Modification 4) The logical text 63 shown in Figure 5 is an example of a stack algorithm, and the algorithm does not necessarily have to be the same. For example, the order of nodes and edges in the logical text 63 can be changed, or the numerical order can be omitted. Alternatively, only node labels or edge labels may be included in the logical text 63. Such modifications will change the performance of the text generation.

[0044] For example, a logical text that retains its graphical representational power while minimizing unnecessary information is preferable for a text generation language model. Furthermore, it is conceivable to generate logical text 63 while referencing the input media. In this case, for example, the logical structure of the input media can be analyzed and combined when generating logical text 63. Alternatively, important key phrases or sentences from the input media can be incorporated into the logical text 63.

[0045] —Second Embodiment— Referring to Figure 6, a second embodiment of the text generation method and text generation apparatus will be described. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically described are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that the user can select a logical template in the inference phase.

[0046] Figure 6 shows the operation of the text generation device 1A during the inference phase of the second embodiment. In addition to the configuration of the first embodiment, the text generation device 1A further includes a template selection unit 102. In this embodiment, the logical template acquisition unit 15 acquires two or more logical templates 62. In this embodiment, two or more logical templates 62 may be stored in the storage unit 6, or one logical template 62 may be stored in the storage unit 6 and one or more logical templates 62 may be input from an external source.

[0047] The logical template acquisition unit 15 acquires multiple logical templates 62, which are then input to the template selection unit 122. The template selection unit 122 selects one logical template 62 based on the external input 123 and outputs it to the logical text generation unit 13. For example, the template selection unit 122 presents a list of the input logical templates 62 to the user via the output device 45 and accepts the user's selection as the external input 123.

[0048] According to the second embodiment described above, the following effects and advantages can be obtained. (2) In the text output process, the logical text 63 input to the text generation model is determined by the user's selection. Therefore, the user can select from several logical structures for the output summary 65.

[0049] —Third Embodiment— Referring to Figure 7, a third embodiment of the text generation method and text generation apparatus will be described. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically described are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that the user can select a logical template in the inference phase.

[0050] Figure 7 shows the operation of the text generation device 1B during the inference phase of the third embodiment. In the text generation device 1B, the logical template generation unit 12 is used not only in the learning phase but also in the inference phase. In this embodiment, the user inputs a sample summary sentence 131 during the inference phase. This sample summary sentence 131 is input to the logical template generation unit 12 to generate a logical template 62. This logical template 62 is converted into logical text 63 by the logical text generation unit 13. The processing after the generation of the logical text 63 is the same as in the first embodiment, so the explanation is omitted. The processing of the logical template generation unit 12 and the logical text generation unit 13 in this embodiment generates logical text 63 corresponding to the sample summary sentence 131 input by the user, so it can also be called "user-initiated logical text generation processing".

[0051] According to the third embodiment described above, the following effects and advantages can be obtained. (3) The text generation method performed by the text generation device 1 includes a user-initiated logical text generation process performed by the logical text generation unit 13, which generates logical text 63 from text input by the user. The logical text 63 input to the text generation model 64 is the logical text 63 generated by the user-initiated logical text generation process. Therefore, the user is not limited to logical structures that have been created in advance and stored in the memory unit 6, but can have the text generation unit 14 generate summary sentences 65 with various logical structures of their own choosing.

[0052] —Fourth Embodiment— Referring to Figure 8, a fourth embodiment of the text generation method and text generation apparatus will be described. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically described are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that the user can select a logical template in the inference phase.

[0053] Figure 8 shows the operation of the text generation device 1C during the inference phase of the fourth embodiment. In addition to the configuration of the first embodiment, the text generation device 1C includes a logical text modification unit 141. The logical text modification unit 141 optimizes the logical text 63 for the text generation model 64. The logical text modification unit 141 then outputs the optimized logical text 63 to the text generation unit 14. The processing after the logical text 63 is generated is the same as in the first embodiment, so the explanation is omitted.

[0054] The optimization of logical text 63 in the logical text modification unit 141 includes, for example, simplification and reordering. Specifically, the logical text 63 "1 <Proposal> <node>Among the following, <node>The symbol " is a special character that indicates the preceding string represents a node, but it can be omitted as the graph information can be retained even without this special character. Also, in the logical text 63, "1" indicates the order in which the sentence corresponding to that node appears, but if the sentences in the summary text 65 have order universality, this order can be omitted. Also, the " shown in Examples 1 and 2 of Figure 5 <edge>You may remove the ''' and exclude edge information. Furthermore, you may swap the order in which node labels and edge labels are listed.

[0055] According to the fourth embodiment described above, the following effects and advantages can be obtained. (4) The processing of the text generation device 1C includes an optimization process that optimizes the logical text 63 for the text generation model 64. In the text output process performed by the text generation unit 14, the logical text 63 optimized by the optimization process is input to the text generation model 64. Therefore, the logical text 63 can be optimized to match the text generation model 64 being used.

[0056] (Modified version of the fourth embodiment) Figure 9 shows the operation of the text generation device 1D during the inference phase in a modified version of the fourth embodiment. In this modified version, two text generation models 64, the first model 64-1 and the second model 64-2, are stored in the memory unit 6. The text generation unit 14 can switch between using the first model 64-1 and the second model 64-2. The first model 64-1 and the second model 64-2 are capable of processing different types of input text appropriately. For example, the first model 64-1 is adept at processing written language such as academic papers and newspaper articles, i.e., literary language, while the second model 64-2 is adept at processing spoken language such as meeting minutes, i.e., alternating speech.

[0057] The logical text modification unit 141 and the text generation unit 14 receive a data type signal 142 from an external source. The data type signal 142 is a signal that specifies the text generation model 64 to be used by the text generation unit 14, for example, a signal that indicates whether the input text 61 is colloquial or written. The text generation unit 14 switches the text generation model 64 to be used to either the first model 64-1 or the second model 64-2, according to the received data type signal 142. The logical text modification unit 141 changes the optimization of the logical text 63 according to the data type signal 142. Specifically, if the data type signal 142 indicates the use of the first model 64-1, the logical text 63 is optimized for the first model 64-1, and if the data type signal 142 indicates the use of the second model 64-2, the logical text 63 is optimized for the second model 64-2.

[0058] This modified version yields the following effects: (5) Multiple text generation models 64 are stored in the memory unit 6. In the optimization process, the logical text 63 is optimized for the text generation model 64 used in the text output process. Therefore, the optimization method for the logical text 63 can be switched in accordance with the switching of the text generation model 64.

[0059] —Fifth Embodiment— Referring to Figure 10, a fifth embodiment of the text generation method and text generation apparatus will be described. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically described are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that it terminates processing when the logical text is inappropriate.

[0060] Figure 10 shows the operation of the text generation device 1D during the inference phase of the fifth embodiment. In addition to the configuration of the first embodiment, the text generation device 1D further includes a logical text evaluation unit 151. Furthermore, in this embodiment, the logical template 62M acquired by the logical template acquisition unit 15 may be inappropriate, unlike in other embodiments. An inappropriate logical template is one that has a significantly low degree of similarity to past logical templates or one whose logic is flawed. The processing of the logical text generation unit 13 is the same as in the first embodiment, and the logical text evaluation unit 151 evaluates the logical text 63 output by the logical text generation unit 13.

[0061] The logical text evaluation unit 151 may, for example, use machine learning or statistical information, or perform structure verification using rules. Furthermore, the logical text evaluation unit 151 may input the input logical text to the text generation unit 103 or to a text generation language model different from the text generation unit 103 and evaluate it as follows: That is, it may calculate the likelihood of the words output by the text generation language model, obtain a likelihood score for the logical text, and consider it inappropriate if this score is smaller than a certain threshold. If the logical text evaluation unit 151 determines that the logical text 63 is inappropriate, it outputs a processing abort command 152 to the text generation unit 14, and if it determines that the logical text 63 is appropriate, it outputs the logical text 63 to the text generation unit 14.

[0062] According to the fifth embodiment described above, the following effects and advantages can be obtained. (6) The text generation method executed by the text generation device 1D includes an evaluation process that evaluates the logical text 63 obtained by the logical text acquisition process and stops the text output process if the evaluation value is lower than a predetermined threshold. This makes it possible to suppress the output of low-quality summary texts 65.

[0063] —Sixth Embodiment— A sixth embodiment of the text generation method and text generation apparatus will be described with reference to Figures 11 and 12. In the following description, the same reference numerals are used for components that are the same as in the first embodiment, and the differences will be mainly explained. Points that are not specifically described are the same as in the first embodiment. This embodiment differs from the first embodiment mainly in that the user inputs a logical template multiple times to create a summary text.

[0064] Figure 11 shows the operation of the text generation device 1E during the inference phase of the sixth embodiment. In this embodiment, the logic template 62 input from an external source is modified, and the generation of the summary text 65 by the text generation unit 14 is repeated many times. Note that in this embodiment, modification of the input text 61 is not assumed.

[0065] The text generation device 1E further includes an editing continuation unit 161 in addition to the configuration of the first embodiment. The editing continuation unit 161 includes an output buffer 162 which stores the summary text 65 output by the text generation unit 14. The editing continuation unit 161 also determines whether or not to continue editing based on user input, and if editing is not to be continued, it outputs all the data accumulated in the output buffer 162 as the summary text 65. If editing is to be continued, the editing continuation unit 161 outputs all the data accumulated in the output buffer 162 as additional language information 163 to the language information acquisition unit 11.

[0066] Figure 12 is a conceptual diagram showing the input and output to the text generation unit 14. In Figure 12, time progresses from the top to the bottom of the diagram, and the text generation unit 14 operated three times. In the first operation of the text generation unit 14, the input text 61 and "logical text 1" were input, and "summary text 1" was stored in the output buffer 162. Subsequently, the user selected to continue editing, so the second process was performed, and "logical text 2" was generated based on different user input than before. In addition, the input to the text generation unit 14 in the second time included not only the input text 61 but also "summary text 1", which was the output from the first time.

[0067] The second processing by the text generation unit 14 generates "Summary 2" and adds it to the output buffer 162. At this point, the output buffer 162 contains both the first output, "Summary 1," and the second output, "Summary 2." In the third operation of the text generation unit 14, the input text 61, "Summary 1," "Summary 2," and "Logical Text 3" are input, and "Summary 3" is output and added to the output buffer 162. After that, if the user selects to end editing, the editing continuation unit 161 outputs all the data stored in the output buffer 162, namely "Summary 1," "Summary 2," and "Summary 3," as Summary 65.

[0068] According to the sixth embodiment described above, the following effects and advantages can be obtained. (7) The processing of the text generation device 1E includes a first decision process that determines the first logical text 63 based on user input, a presentation process that presents the summary text 65 obtained by the text output process to the user for the first time, and a second decision process that determines the second logical text 63 based on user input after the presentation process. The processing of the text generation device 1C further includes a second text output process that inputs the logical text 63 determined by the second decision process and the text obtained by combining the input document 61 and the summary text 65 output by the text output process executed immediately before into the text generation model 64 to obtain a new summary text 65, which is "summary text 2". When the user inputs after the second text output process, such as the third input, the text generation device 1E repeats the second decision process and the second text output process to update the output text. As a result, the entire summary text 65 can be constructed while confirming the summary text 65 corresponding to the logical text 63 input by the user one by one.

[0069] In the embodiments and modifications described above, the configuration of the functional blocks is merely an example. Several functional configurations shown as separate functional blocks may be integrated, or a configuration represented in one functional block diagram may be divided into two or more functions. Furthermore, some of the functions of one functional block may be provided by other functional blocks.

[0070] In the embodiments and modifications described above, the program is stored in a ROM (not shown), but the program may also be stored in a non-volatile storage device (not shown). Furthermore, the text generation device 1 may have an input / output interface (not shown), and the program may be read from another device via the input / output interface and a medium available to the text generation device 1 when needed. Here, "medium" refers to, for example, a storage medium detachable from the input / output interface, or a communication medium, i.e., a wired, wireless, or optical network, or a carrier wave or digital signal propagating through such a network. Also, some or all of the functions realized by the program may be realized by hardware circuits or FPGAs.

[0071] The embodiments and modifications described above may be combined in any way. Although various embodiments and modifications have been described above, the present invention is not limited to these. Other embodiments that can be conceivable within the scope of the technical idea of ​​the present invention are also included within the scope of the present invention. [Explanation of symbols]

[0072] 1~1E: Text generation device 6: Storage section 11: Language Information Acquisition Unit 12: Logical template generation unit 13: Logical Text Generation Unit 14: Text generation unit 15: Logical template acquisition unit 16: Logical Text Acquisition Unit 61: Input text 62: Logical Templates 63: Logical Text 64: Text generation model 65: Summary 102: Template Selection Section 103: Text generation unit 122: Template Selection Section 141: Logical Text Correction Section 142: Data type signal 151: Logical Text Evaluation Unit< / edge> < / node> < / node> < / edge> < / edge> < / node> < / node>

Claims

1. A text generation method performed by a computer having a memory unit that stores a pre-trained text generation model, The text generation model takes a logical text that indicates a specified logical structure, which is the logical structure between sentences, and a sentence to be processed as input, and outputs output text that is text according to the specified logical structure from the sentence to be processed. A language information acquisition process that acquires language information indicating the document to be processed, A logical text acquisition process that acquires the aforementioned logical text, A text output process that inputs the text to be processed and the logical text to the text generation model and obtains the output text, The process includes an optimization process for optimizing the logical text for the text generation model, A text generation method in which, in the text output process, the logical text optimized by the optimization process is input to the text generation model.

2. A text generation method according to claim 1, A text generation method wherein the logical text input to the text generation model in the text output process is determined by user selection.

3. A text generation method according to claim 1, The process further includes a user-initiated logical text generation process that generates the aforementioned logical text from text entered by the user, A text generation method wherein the logical text input to the text generation model in the text output process is the logical text generated by the user-initiated logical text generation process.

4. A text generation method according to claim 1, The memory unit stores a plurality of the text generation models. A text generation method in which the logical text is optimized for the text generation model used in the text output process during the optimization process.

5. A text generation method according to claim 1, A text generation method further comprising an evaluation process that evaluates the logical text obtained by the logical text acquisition process, and stops the text output process if the evaluation value is lower than a predetermined threshold.

6. A text generation method according to claim 1, A first decision process that determines the logical text based on user input, A presentation process that presents the output text obtained by the text output process to the user, A second decision process that determines the logical text based on user input after the presentation process, The method further includes a second text output process that inputs the logical text determined by the second decision process, and the text obtained by combining the text to be processed and the output text output by the text output process executed immediately before, into the text generation model to obtain a new output text, A text generation method that, upon receiving input from the user after the second text output process, repeats the second decision process and the second text output process to update the output text.

7. A language information acquisition unit that acquires language information indicating the text to be processed, A logical text acquisition unit that acquires logical text, which is a string that indicates a specified logical structure, which is the logical structure between sentences, A storage unit that takes the aforementioned text to be processed and the aforementioned logical text as input and stores a text generation model trained to generate text according to the specified logical structure from the text to be processed, A text generation unit that uses the text generation model to obtain text from the document to be processed according to the specified logical structure, The system comprises a logical text modification unit that optimizes the logical text for the text generation model, The text generation unit is a text generation device that obtains text by inputting the logical text optimized by the logical text correction unit and the document to be processed.