Abstract learning support device, abstract learning support method, and program
The summary learning support device enhances the efficiency of query-dependent generation models by calculating and selecting appropriate input parameters, addressing the lack of sufficient learning data and improving summarization performance.
Patent Information
- Application Number
- JP2023543588
- Authority / Receiving Office
- JP · JP
- Patent Type
- Patents
- Current Assignee / Owner
- Filing Date
- 2021-08-26
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2041-08-26
AI Technical Summary
Existing query-dependent generation models lack sufficient learning data that include additional input parameters, hindering their efficiency in generating summary sentences.
A summary learning support device and method that calculates and selects appropriate input parameters using a score-based approach to enhance the learning process for summary generation models, generating pseudo queries from query-free learning data.
Improves the efficiency of learning for summarization tasks requiring additional input parameters by effectively utilizing score-based parameter selection and pseudo query generation.
Smart Images

Figure 0007700862000001 
Figure 0007700862000002 
Figure 0007700862000003
Abstract
Description
Technical Field
[0001] The present invention relates to a summarization learning support device, a summarization learning support method, and a program.
Background Art
[0002] As learning data for a model that generates a summary sentence using a neural network, a pair of a source text to be summarized and summary data that is a correct summary result is common.
[0003] On the other hand, there is a model that requires input parameters other than the source text (hereinafter referred to as "query") (for example, Non-Patent Document 1). According to such a model, a summary sentence according to the query can be generated. In such a model, a set of parameters such as the source text, the query, and the summary data is used as learning data.
[0004] On the other hand, as methods for generating a summary sentence, there are an extraction type and a generation type. The extraction type is a method in which a part included in the source text is directly extracted. The generation type is a method in which summary data is generated based on words included in the source text. Hereinafter, a model that requires a query as input and generates summary data by the generation type is referred to as a "query-dependent generation type model".
Prior Art Documents
Non-Patent Documents
[0005]
Non-Patent Document 1
Summary of the Invention
Problems to be Solved by the Invention
[0006] Although there are a large number of learning data composed of pairs of source text and summary data, for learning query-dependent generation models, learning data including additional input parameters other than the source text is insufficient.
[0007] The present invention has been made in view of the above points, and an object thereof is to improve the efficiency of learning of summaries that require additional input parameters.
Means for Solving the Problems
[0008] Therefore, in order to solve the above problems, a summary learning support device includes a calculation unit that calculates, for a plurality of character strings, a score representing the appropriateness as an input parameter added when summarizing a first document based on a predetermined model, and a selection unit that selects, based on the score, a partial character string group from among the plurality of character strings as the input parameter constituting the learning data of a summary generation model that generates a summary of the document. The score is a score calculated for each of the output candidate strings for the model to select the output target string from among the output candidate strings when a second document, which is a summary of the first document, is input to a model that has learned the correspondence relationship between the main body of the document and the character string group constituting the title of the document .
Effects of the Invention
[0009] It is possible to improve the efficiency of learning of summaries that require additional input parameters.
Brief Description of the Drawings
[0010]
Figure 1
Figure 2
Figure 3
Figure 4
Figure 5
Figure 6
Figure 7
Figure 8
Embodiments for Carrying Out the Invention
[0011] Hereinafter, embodiments of the present invention will be described with reference to the drawings. FIG. 1 is a diagram showing a hardware configuration example of the summary generation device 10 in the first embodiment. The summary generation device 10 in FIG. 1 includes a drive device 100, an auxiliary storage device 102, a memory device 103, a processor 104, and an interface device 105, etc., which are mutually connected by a bus B.
[0012] The program for realizing the processing in the summary generation device 10 is provided by a recording medium 101 such as a CD-ROM. When the recording medium 101 storing the program is set in the drive device 100, the program is installed from the recording medium 101 to the auxiliary storage device 102 via the drive device 100. However, the installation of the program does not necessarily have to be performed from the recording medium 101, and it may be downloaded from another computer via a network. The auxiliary storage device 102 stores the installed program and also stores necessary files, data, etc.
[0013] When an activation instruction for the program is given, the memory device 103 reads out and stores the program from the auxiliary storage device 102. The processor 104 is a CPU, a GPU (Graphics Processing Unit), or both a CPU and a GPU, and executes the functions related to the summary generation device 10 according to the program stored in the memory device 103. The interface device 105 is used as an interface for connecting to a network.
[0014] FIG. 2 is a diagram showing a functional configuration example of the summary generation device 10 in the first embodiment. In FIG. 2, the summary generation device 10 includes a learning data generation unit 11 with a query, a summary learning unit 12, and a summary unit 13. Each of these units is realized by the processing executed by the processor 104 by one or more programs installed in the summary generation device 10.
[0015] The query-with-learning-data generation unit 11 generates query-with-learning data based on each piece of query-free learning data included in the query-free learning data group given as input. One piece of query-with-learning data is generated for one piece of query-free learning data. Therefore, for a query-free learning data group that is a set of a plurality of pieces of query-free learning data, a query-with-learning data group that is a set of a plurality of pieces of query-with-learning data is generated. Both the query-free learning data and the query-with-learning data are data used as learning data for a model such as a neural network that generates a summary of a document (hereinafter referred to as a "summary generation model"). The query-free learning data differs from the query-with-learning data in that it does not include a query as a component. A query refers to text (a character string) that is input to the summary generation model together with the document to be summarized as additional information regarding the summary. For example, the focus of the summary may be the query.
[0016] The query-free learning data is learning data composed of a pair of two text data, {source text, summary text}. The source text refers to the text data of the document to be summarized. The summary text refers to the text data indicating the correct answer of the result of summarizing the source text.
[0017] On the other hand, the query-with-learning data is learning data composed of a set of three text data, {source text, query, summary text}.
[0018] The summary learning unit 12 performs learning of the summary generation model using the query-with-learning data.
[0019] When the summary unit 13 receives the source text to be summarized and an input such as a query for the source text, it inputs the source text and the query to the learned summary generation model, thereby causing the summary generation model to generate a summary corresponding to the query for the source text.
[0020] The learning data generation unit 11 with queries will be described in more detail. FIG. 3 is a diagram showing a configuration example of the learning data generation unit 11 with queries in the first embodiment. In FIG. 3, the learning data generation unit 11 with queries includes an importance calculation unit 111, a query selection unit 112, and a query addition unit 113. The functions of these units will be described in detail with reference to FIG. 4.
[0021] FIG. 4 is a flowchart for explaining an example of the processing procedure of the generation process of the learning data with queries in the first embodiment.
[0022] In step S101, the importance calculation unit 111 generates, for each query-free learning data (a pair of source text and summary text) included in the query-free learning data group, a document (hereinafter referred to as the "extraction source document") that serves as the extraction source of the character string that is a candidate for the query. Therefore, N extraction source documents are generated from N query-free learning data.
[0023] For example, the importance calculation unit 111 generates, as the extraction source document based on the query-free learning data, any one of the following (a) to (d) of the query-free learning data. (a) A document obtained by combining the source text and the summary text (a document including both the source text and the summary text) (b) Only the summary text (c) Only the source text (d) A document obtained by combining any one of (a) to (c) with other attached information text (for example, the title of the source text, etc.) Subsequently, based on a predetermined model, the importance calculation unit 111 calculates the importance in these source document groups for each character string (e.g., word) of a predetermined unit that constitutes each source document, as an example of a score representing the appropriateness as a query (added input parameter) used in summarizing the document (S102). For example, the importance calculation unit 111 uses a TF-IDF calculation model as the predetermined model. In this case, the importance calculation unit 111 calculates the TF-IDF of each word as the importance. The TF-IDF of each word included in the document group can be calculated using a known method. In this embodiment, the "parameter" in the input parameter is clearly distinguished from, for example, the learning parameters of a model such as a neural network. The input parameter is data given as an input to the model, while the learning parameter is data whose value changes according to the learning of the model. As a general example, the input parameter is given as text data or the like, while the learning parameter is represented as a set of numerical data or the like.
[0024] Subsequently, for each source document, the query selection unit 112 selects K character strings in descending order of importance from among the character strings (words) of a predetermined unit that constitute the source document, as queries corresponding to the query-free learning data corresponding to the source document (S103). Note that the value of K (K >= 0) may be randomly selected for each source document or may be the same for all source documents. Also, when selecting a query from each source document, the query selection unit 112 may select only the words included in the summary text of the source document as the query. By doing so, it is possible to facilitate learning such that the specified query is included in the summary for the summary generation model.
[0025] Subsequently, for each piece of query-free learning data, the query addition unit 113 generates query-containing learning data by adding K words selected from the source document based on the query-free learning data as queries to the query-free learning data (S104). Therefore, the generated query-containing learning data includes the source text and summary text included in the query-free learning data, and K queries (query sequence) extracted from the query-free learning data.
[0026] As described above, according to the first embodiment, pseudo queries can be generated from learning data that does not include queries. Therefore, it is possible to improve the efficiency of learning for summarization that requires additional input parameters.
[0027] Next, a second embodiment will be described. In the second embodiment, differences from the first embodiment will be described. Regarding points not particularly mentioned in the second embodiment, they may be the same as those in the first embodiment.
[0028] In the second embodiment, the configuration of the query-containing learning data generation unit 11 and the processing procedure executed by the query-containing learning data generation unit 11 are different from those in the first embodiment.
[0029] FIG. 5 is a diagram showing a configuration example of the learning data generation unit 11 with queries in the second embodiment. In FIG. 5, parts that are the same as or corresponding to those in FIG. 3 are denoted by the same reference numerals. In the second embodiment, the learning data generation unit 11 with queries has a query generation model learning unit 114 and a query candidate generation unit 115 instead of the importance calculation unit 111. The query generation model learning unit 114 learns a model (hereinafter referred to as "query generation model") that generates one or more queries from the learning data without queries. The query generation model is configured by, for example, a neural network or the like. The query generation model learning unit 114 takes as input a learning document group that is the source of the learning data of the query generation model. The learning document group refers to a set of a plurality of learning documents. A learning document refers to text-form document data including an encyclopedia published on the Internet such as wikipedia or a newspaper, including a title (heading) and a text.
[0030] The query candidate generation unit 115 generates (outputs) query candidates based on the learned query generation model.
[0031] FIG. 6 is a flowchart for explaining an example of the processing procedure of generating learning data with queries in the second embodiment.
[0032] In step S201, the query generation model learning unit 114 generates learning data for the query generation model for each learning document included in the learning document group. Specifically, the query generation model learning unit 114 decomposes (divides) the title of each learning document into character strings (for example, words) of a predetermined unit. Therefore, for example, for each learning document, a word sequence (hereinafter simply referred to as "word sequence") constituting the title is generated. At this time, the query generation model learning unit 114 may generate a word sequence from which stop words are removed. The query generation model learning unit 114 generates, for each learning document, a pair of the text (paragraph text) of the learning document and the word sequence generated from the title corresponding to the text as learning data.
[0033] Subsequently, the query generation model learning unit 114 performs learning of the query generation model using the learning data group generated in step S201 (S202). Specifically, the query generation model learning unit 114 causes the query generation model to learn the correspondence between the main text and the word sequence in the case where the main text of each learning data is used as input and the word sequence of the title is used as output. Therefore, the query generation model is learned to output a word sequence related to the title of a document when the main text of the document is input. Note that the query generation model may be configured by, for example, a known encoder-decoder model or may be configured by another known text generation model.
[0034] Subsequently, for each query-free learning data, the query candidate generation unit 115 inputs the summary text of the query-free learning data into the learned query generation model, and generates a string group (word sequence) output by the query generation model as a query candidate sequence corresponding to the query-free learning data (S203).
[0035] At this time, if the query generation model is an encoder-decoder model, the query generation model sequentially outputs each word constituting the word sequence in response to the input of the query-free learning data. In the sequential output of words, the query generation model calculates a score for selecting an output target from among the output candidates for each of the D words constituting its vocabulary (the set of words that are output candidates of the query generation model), and outputs the word with the maximum score. In the second embodiment, the score corresponds to an example of a score indicating the appropriateness as a query (an additional input parameter) used in summarizing a document.
[0036] Subsequently, for each piece of query-free training data, the query selection unit 112 selects one or more words (query sequence) to be used as a query from among the query candidate sequences generated by the query candidate generation unit 115 for the query-free training data (S204). At this time, the query selection unit 112 may select all of the query candidate sequences as the query sequence, or may select a part of the query candidate sequences as the query sequence. When selecting a part of the query candidate sequences as the query sequence, the query selection unit 112 may select the words from the first to the K-th in the query candidate sequence as the query sequence. That is, among the outputs of words sequentially performed by the query generation model, the words up to the K-th may be selected as the query. Alternatively, in step S203, the number of sequential word outputs from the query generation model may be suppressed to K times. In this case, the query candidate sequence will be composed of K words. Therefore, in this case, in step S204, all of the query candidate sequences may be selected as the query sequence.
[0037] Subsequently, for each piece of query-free training data, the query addition unit 113 generates query-with-training data by adding the selected query sequence for the query-free training data to the query-free training data (S205).
[0038] As described above, according to the second embodiment, the same effects as those of the first embodiment can be obtained.
[0039] Next, as a third embodiment, a first example regarding the training of a summary generation model using query-with-training data and the generation of a summary using the trained summary generation model will be described. Note that the third embodiment is applicable to both the first embodiment and the second embodiment.
[0040] FIG. 7 is a diagram for explaining the training of the summary generation model and the generation of a summary in the third embodiment. In FIG. 7, the summary unit 13 includes a content selection unit 131, an encoder 132, and a decoder 133. These units constitute a summary generation model.
[0041] During the learning of the summary generation model, the summary learning unit 12 inputs the source text and query sequence of the learning data into the summary unit 13 for each piece of learning data (source text, query sequence, summary text) included in the learning data group with queries.
[0042] The content selection unit 131 is a model (e.g., a neural network) that calculates the importance for each string (e.g., word) constituting the text (hereinafter referred to as "combined text") obtained by combining the source text and the query sequence. The content selection unit 131 may be configured by fine-tuning a pre-trained model such as BERT or MASS. For BERT, for example, it is detailed in "https: / / arxiv.org / abs / 1810.04805" etc. For MASS, for example, it is detailed in "https: / / arxiv.org / abs / 1905.02450" etc.
[0043] The content selection unit 131 extracts N word sequences (important word sequences) from the combined text in descending order of importance, and inputs the important word sequences, the source text given as input, and the query sequence into the encoder 132. At this time, the content selection unit 131 combines the query sequence, the important word sequence, and the source text with special tokens such as [SEP] as "query sequence [SEP] important word sequence [SEP] source text". The value of N may be input to the content selection unit 131 together with queries and the like.
[0044] The encoder 132 and the decoder 133 are, for example, known encoder-decoder models (neural networks) such as BERT or MASS.
[0045] The encoder 132 encodes the input text. The decoder 133 generates and outputs a summary text based on the encoding result.
[0046] The summarization learning unit 12 updates the learning parameters of the encoder 132 and the decoder 133 based on the comparison between the summary text included in the learning data and the summary text output by the decoder 133. Note that the comparison and the update of the learning parameters may be performed based on known techniques.
[0047] When the learning is completed, the summarization unit 13 functions as a learned summary generation model that takes a query sequence and an input text as inputs and outputs a summary text.
[0048] Note that the summarization unit 13 in FIG. 7 may be configured using the technique disclosed in International Publication No. 2021 / 064907.
[0049] Next, as a fourth embodiment, a second example regarding the learning of a summary generation model using learning data with a query and the generation of a summary using the learned summary generation model will be described. Note that the fourth embodiment is applicable to both the first embodiment and the second embodiment.
[0050] FIG. 8 is a diagram for explaining the learning of the summary generation model and the generation of a summary in the fourth embodiment. In FIG. 8, the same reference numerals are given to the same or corresponding parts as in FIG. 7. In FIG. 8, the summarization unit 13 includes an encoder 132 and a decoder 133. These units constitute a summary generation model. That is, the summary generation model of the fourth embodiment does not have a content selection unit 131.
[0051] During the learning of the summary generation model, the summarization learning unit 12 inputs the source text and the query sequence of the learning data to the summarization unit 13 for each piece of learning data (source text, query sequence, summary text) included in the learning data group with a query. At this time, the summarization learning unit 12 combines the query sequence and the source text with a special token such as [SEP] as in "query sequence [SEP] source text".
[0052] The encoder 132 encodes the input text. The decoder 133 generates and outputs a summary text based on the encoding result.
[0053] The summarization learning unit 12 updates the learning parameters of the encoder 132 and the decoder 133 based on the comparison between the summary text included in the learning data and the summary text output by the decoder 133. Note that the comparison and the update of the learning parameters may be performed based on known techniques.
[0054] When the learning is completed, the summarization unit 13 functions as a learned summary generation model that takes a query sequence and an input text as inputs and outputs a summary text.
[0055] Note that in the fourth and fifth embodiments, the summarization unit 13 may be configured based on a sentence generation model other than the encoder-decoder model.
[0056] Regarding the above embodiments, the following additional remarks are further disclosed.
[0057] (Supplementary Note 1) A memory, At least one processor connected to the memory, Including, The processor, For a plurality of character strings, calculates a score representing the appropriateness as an input parameter added during the summarization of the first document based on a predetermined model, Based on the score, selects a partial character string group from the plurality of character strings as the input parameter constituting the learning data of the summary generation model for generating the summary of the document, A summarization learning support device characterized by the above.
[0058] (Supplementary Note 2) For a plurality of character strings, calculates a score representing the appropriateness as an input parameter added during the summarization of the first document based on a predetermined model, Based on the above scores, select some of the plurality of character strings as the input parameters that constitute the training data of the summary generation model for generating a summary of the document. A recording medium storing a program for causing a computer to execute processing.
[0059] In each of the above embodiments, the summary generation device 10 is an example of a summary learning support device. The importance calculation unit 111 or the query candidate generation unit 115 (query generation model) is an example of a calculation unit. The query selection unit 112 is an example of a selection unit. The summary learning unit 12 is an example of a learning unit.
[0060] Although the embodiments of the present invention have been described in detail above, the present invention is not limited to such specific embodiments, and various modifications and changes are possible within the scope of the gist of the present invention described in the claims.
Explanation of Signs
[0061] 10 Summary generation device 11 Query-with-learning data generation unit 12 Summary learning unit 13 Summary unit 100 Drive device 101 Recording medium 102 Auxiliary storage device 103 Memory device 104 Processor 105 Interface device 111 Importance calculation unit 112 Query selection unit 113 Query addition unit 114 Query generation model learning unit 115 Query candidate generation unit 131 Content selection unit 132 Encoder 133 Decoder B Bus
Claims
1. A calculation unit that calculates a score representing the appropriateness as an input parameter added when summarizing a first document based on a predetermined model for a plurality of character strings; A selection unit that selects, based on the score, a partial character string group from among the plurality of character strings as the input parameter constituting the learning data of a summary generation model that generates a summary of a document; having; The score is a score calculated for each of the output candidate character strings for the model to select an output target character string from among the output candidate character strings when a second document, which is a summary of the first document, is input to a model that has learned the correspondence relationship between the main text of the document and the character string group constituting the title of the document. A summary learning support device characterized by the above.
2. The score is the importance in a third document including either one or both of the first document and a second document that is a summary of the first document, for each of the plurality of character strings constituting the third document. The summary learning support device according to claim 1, characterized by the above.
3. A learning unit that learns the summary generation model using learning data including the first document, the character string group, and a second document that is a summary of the first document; The summary learning support device according to claim 1 or 2, characterized by having the above.
4. A summary unit that inputs a certain document and a character string related to the summary of the certain document to the summary generation model learned by the learning unit, and generates a summary of the certain document; The summary learning support device according to claim 3, characterized by having the above.
5. A calculation procedure for calculating a score representing the appropriateness as an input parameter added when summarizing a first document based on a predetermined model for a plurality of character strings; A selection procedure for selecting, based on the score, a partial character string group from among the plurality of character strings as the input parameter constituting the learning data of a summary generation model that generates a summary of a document; The computer executes; The score is a score calculated for each of the output candidate character strings for the model to select an output target character string from among the output candidate character strings when a second document, which is a summary of the first document, is input to a model that has learned the correspondence relationship between the main text of the document and the character string group constituting the title of the document. A summary learning support method characterized by the above.
6. A program that causes a computer to function as the summary learning support device according to any one of claims 1 to 4.
Citation Information
Patent Citations
Summary learning method, summary learning device, and program
WO2021124489A1