Summary generation system, summary generation method, and summary generation program

The summary generation system addresses the issue of low-quality summaries by using an evaluation processing unit to assess and improve the summary text, resulting in high-quality summaries that meet predetermined criteria.

JP2025086698APending Publication Date: 2025-06-09HITACHI SOFTWARE ENG

Patent Information

Application Number
JP2023200897
Authority / Receiving Office
JP · JP
Patent Type
Applications
Current Assignee / Owner
Filing Date
2023-11-28
Publication Date
2025-06-09

AI Technical Summary

Technical Problem

Existing methods for generating text summaries, such as those described in Patent Document 1, often fail to produce high-quality summaries when there is no implicative relationship between the summary sentence and the input sentences, leading to regenerations that may not improve summary quality.

Method used

A summary generation system that includes a summary text processing unit and an evaluation processing unit. The evaluation unit determines the quality of the summary text based on predetermined conditions and creates supplementary information to improve quality, ensuring that the generated summary meets the quality criteria.

Benefits of technology

The system effectively generates high-quality summaries by iteratively improving the summary text based on quality evaluations, ensuring that the final summary meets predetermined quality conditions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure 00000000_0000_ABST
    Figure 00000000_0000_ABST
Patent Text Reader

Abstract

To provide a technique capable of generating a high-quality summary.SOLUTION: A summary generation system comprises: a summary text processing unit which generates a summary text, which is the summarization of an original text; and an evaluation processing unit which determines whether the quality of the summary text satisfies a predetermined quality condition, and if the quality does not satisfy the quality condition, creates supplementary information to improve the quality, generates a summary text using the supplementary information, and evaluates the quality of the generated summary text.SELECTED DRAWING: Figure 3
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to a technique for generating a summary of a text.

Background Art

[0002] Natural language generation artificial intelligence is utilized to create summaries from long texts or large amounts of text. Patent Document 1 discloses a method for generating a summary related to a specific category. In the method of Patent Document 1, first, a second set of sentences belonging to a first category is extracted from a first set of sentences by a category analyzer. Next, the sentences in the second set of sentences are clustered so as to collect sentences having similar expressions and meanings, and a third set of sentences is extracted from the set of clusters with high scores. Next, the third set of sentences is input into a language model to generate a first summary sentence related to the first category. Further, a truth value determination is made as to whether or not there is an implicative relationship between the first summary sentence and the third set of sentences by a fact-checking unit. If it is determined to be false, the third set of sentences is input into the language model to regenerate the first summary sentence.

Prior Art Documents

Patent Documents

[0003]

Patent Document 1

Summary of the Invention

Problems to be Solved by the Invention

[0004] In the method of Patent Document 1, if there is no implicative relationship between the summary sentence generated by the language model and the sentence input into the language model, the summary sentence is regenerated. However, the summary sentence output from the language model by regeneration does not always have an implicative relationship with the sentence input into the language model, and a summary of good quality is not always obtained.

[0005] One objective included in this disclosure is to provide a technology that enables the generation of a summary of good quality.

Means for Solving the Problem

[0006] A summary generation system according to one aspect included in this disclosure includes a summary text processing unit that generates a summary text obtained by summarizing the original text, and an evaluation processing unit that determines whether the quality of the summary text satisfies a predetermined quality condition, and if the quality does not satisfy the quality condition, creates supplementary information for improving the quality, generates a summary text using the supplementary information, and determines the quality of the generated summary text.

Advantages of the Invention

[0007] According to one aspect included in this disclosure, it becomes possible to generate a summary of good quality.

Brief Description of the Drawings

[0008]

Figure 1

Figure 2

Figure 3

Figure 4

Figure 5

Figure 6

Figure 7

Figure 8

Figure 9

Figure 10

Figure 11

Figure 12

Figure 13

Embodiments for Carrying Out the Invention

[0009] Hereinafter, embodiments of the present invention will be described with reference to the drawings.

[0010] FIG. 1 is a conceptual diagram for explaining the outline of text summarization processing. Text summarization processing is a process of generating a summary text by summarizing the original text. The original text is not particularly limited, and for example, it may be a text with a large amount of information contained therein, such as a document created as part of a work product accompanying information system development.

[0011] The text summarization processing is executed in the following flow using a text summary generation system 10, a natural language generation artificial intelligence system 20, and a database 30. The text summary generation system 10 is a computer system that provides a user with a function for generating a summary. The natural language generation artificial intelligence system 20 is a computer system that generates a natural language text using natural language generation artificial intelligence. Hereinafter, artificial intelligence may be abbreviated as AI. AI is the initials of Artificial Intelligence. The natural language generation AI is, for example, a trained machine learning model. The database 30 is a group of data stored in a storage device in a readable manner. Various data before, during, and after the text summarization processing are stored in the database 30. Hereinafter, the database may be abbreviated as DB.

[0012] (1) The administrator 91 makes prior settings for the text summarization generation system 10. The prior settings include the setting of the sufficiency threshold and the frequent occurrence determination threshold. The administrator 91 is the person who manages and operates the text summarization generation system 10. The text summarization generation system 10 is managed by the administrator 91 and provides functions to the user 92 and is used by the user 92.

[0013] (2) The user 92 makes a summarization request to the text summarization generation system 10. This summarization request includes the original text to be summarized, the required summary length, and the specification of the essential keywords. The sufficiency threshold is a threshold used to determine the quality of the summary text. Details of the sufficiency threshold will be described later. The required summary length is the specification of the length of the summary text, and is specified by the number of characters as an example. The essential keywords are words and phrases such as words specified as the keywords to be included in the summary text.

[0014] (3) The text summarization generation system 10 makes a summarization request to the natural language generation AI system 20. This summarization request includes instruction information (not shown) and the original text. The instruction information is the text for instructing the natural language generation AI system 20 to generate a summary, which is a so-called prompt.

[0015] (4) The natural language generation AI system 20 sends the summary text as the processing result to the text summarization generation system 10. Also, the processing result is stored in the database 30 in a readable manner.

[0016] (5) The text summarization generation system 10 determines whether the quality of the summary text sent from the natural language generation AI system 20 meets the predetermined quality conditions. Details of the quality conditions and the determination method will be described later. If the quality of the summary text does not meet the quality conditions, the text summarization generation system 10 makes a summary request again to the natural language generation AI system 20 together with supplementary information for improving the quality of the summary text.

[0017] (6) When the natural language generation AI system 20 that has received a request for a summary sends the summary text as a processing result to the text summarization generation system 10. Also, the processing result is stored in the database 30 in a readable manner. (5) to (6) are repeated until the quality of the summary text meets the quality criteria.

[0018] (7) When the quality of the summary text meets the quality criteria, the text summarization generation system 10 provides the user 92 with the summary text having a sufficiently high accuracy.

[0019] Figure 2 is a sequence diagram of the text summarization process.

[0020] In the above (2), when the user 92 inputs a request to the text summarization generation system 10 together with the original text, the required summary length, and the essential keywords, the text summarization generation system 10 extracts frequently occurring keywords from the original text. Frequently occurring keywords are words that appear with a high frequency within the original text. Since words that appear with a high frequency in the original text are presumed to be words related to important information in the original text, they are regarded as words that should also appear in the summary text. The frequently occurring keywords and their extraction method will be described later.

[0021] After that, in the above (3), the text summarization generation system 10 sends the creation of its summary together with the original text to the natural language generation AI system 20, and in the above (4), receives the summary text as a processing result from the natural language generation AI system 20.

[0022] Upon receiving the summary text of the processing result, the text summarization generation system 10 decomposes the original text and the summary text into individual sentences respectively, calculates the distance between each sentence of the original text and each sentence of the summary text, and extracts the maximum value (the maximum similarity value). Also, the text summarization generation system 10 checks the sufficiency threshold value used for comparison with the maximum similarity value. Furthermore, the text summarization generation system 10 performs a search for essential keywords and a search for frequently occurring keywords on the summary text. Additionally, the text summarization generation system 10 checks the length of the summary text. Then, using that information, the text summarization generation system 10 determines whether the quality of the summary text meets a predetermined quality condition. The determination includes a determination based on the maximum similarity value, a determination based on essential keywords, a determination based on frequently occurring keywords, and a determination based on the length of the summary text. The details of each determination method will be described later.

[0023] When the quality of the summary text does not meet the quality condition, the text summarization generation system 10 creates supplementary information for improving the quality, and in the above (5), sends a summary request again to the natural language generation AI system 20 together with the supplementary information, and in the above (6), receives the summary text as the processing result.

[0024] The text summarization generation system 10 repeatedly determines whether the quality of the summary text meets the quality condition and re-requests the summary to the natural language generation AI system 20. When the quality of the summary text meets the quality condition, in the above (7), it provides the highly accurate summary text to the user 92.

[0025] Figure 3 is a block diagram of the text summarization generation system.

[0026] The text summarization generation system 10 has a storage unit 101 and a processing unit 107. The text summarization generation system 10 is a computer equipped with a processing device (CPU: Central Processing Unit) and a storage device as hardware. Each part of the processing unit 107 is realized by the processing device reading and executing a software program defined and stored in the storage device.

[0027] The storage unit 101 stores a document management DB 102, a frequently occurring keyword management DB 103, an essential keyword management DB 104, a user setting management DB 105, and a summary article management DB 106.

[0028] The processing unit 107 includes a user interface unit 108, a sentence construction processing unit 109, an information management unit 112, a generation AI interface unit 116, a summary article processing unit 119, and an evaluation processing unit 122.

[0029] The sentence construction processing unit 109 includes a morphological analysis unit 110 and a sentence construction unit 111.

[0030] The information management unit 112 includes a document management unit 113, a frequently occurring keyword calculation unit 114, and a user setting information management unit 115.

[0031] The generation AI interface unit 116 includes an instruction information creation unit 117 and a generation AI call unit 118.

[0032] The summary article processing unit 119 includes a summary article generation unit 120 and a summary article management unit 121.

[0033] The evaluation processing unit 122 includes a sentence distance calculation unit 123, a user setting information comparison unit 124, a keyword comparison unit 125, and a supplementary information generation unit 126.

[0034] As a pre - setting process, setting values of parameters used for article summary processing are pre - recorded in the user setting management DB 105. By the administrator 91 setting the setting values of the parameters, the user 92 can use the article summary generation system 10. As a summary generation process, in response to the request of the user 92, the article summary generation system 10 generates a summary article and provides the summary article to the user 92. At that time, the user 92 is provided with a GUI (Graphical User Interface) for using the article summary generation system 10 by the user interface unit 108.

[0035] <Pre-setting process> The user setting information management unit 115 performs pre-setting processing according to the operations by the administrator 91.

[0036] Figure 4 is a flowchart of the pre-setting process. In step S101, the user setting information management unit 115 registers the sufficiency threshold value in the user setting management DB 105. In step S102, the user setting information management unit 115 registers the frequent occurrence determination threshold value in the user setting management DB 105.

[0037] Figure 5 is a diagram showing an example of the data in the user setting management DB.

[0038] Corresponding to the management ID, an article name, an administrator, an abstract length, a sufficiency threshold value, and a frequent occurrence determination threshold value are set. The "management ID" is assigned for each original article. The "article name" is the name of the original article. The "administrator" is the identification information of the administrator of the original article. The "abstract length" is information specifying the length of the abstract article. The "abstract length" may be set in advance by the administrator 91, or may be set by the user 92 when generating the abstract article. The "sufficiency threshold value" is a threshold value used for comparison with the maximum similarity value of each sentence of the original article. The "frequent occurrence determination threshold value" is a threshold value used for extracting frequent keywords.

[0039] <Abstract generation process> Figure 6 is a flowchart of the abstract generation process.

[0040] In the abstract generation process, first, various information is acquired from the user 92 by the user interface unit 108. The document management unit 113 determines, in step S201, whether the information acquired from the user 92 is sufficient. If the information is insufficient, in step S202, the document management unit 113 notifies the user 92 of the insufficient information via the GUI by the user interface unit 108 and ends the process. If the information is sufficient, in step S203, the original text, essential keywords, required abstract length, and user data are registered in the database 30. The original text is registered in the document management DB 102. The essential keywords are registered in the essential keyword management DB 104. The required abstract length, threshold, etc. are registered in the user setting management DB 105.

[0041] FIG. 7 is a diagram showing an example of the data in the essential keyword management DB.

[0042] The essential keyword management DB 104 stores the data of the essential keywords.

[0043] Corresponding to the management ID assigned to each essential keyword, the article name, keyword, appearance count, and usage are recorded. The "article name" is the name of the original text. The "keyword" is the text data (string) of the essential keyword. The "appearance count" is information indicating the number of times the frequent keyword appears in the abstract text. The "usage" is information indicating whether the frequent keyword is used in the abstract text. Until it is determined whether the essential keyword is used in the abstract text, the "appearance count" is set to 0, and the "usage" column is left blank.

[0044] Subsequently, in step S204, frequently occurring keywords are extracted. At this time, first, the morphological analysis unit 110 performs morphological analysis on the original text. As a result, the original text is segmented into morphemes. The frequently occurring keyword calculation unit 114 extracts predetermined types of phrases from the morphemes obtained by the morphological analysis unit 110 and counts the number of occurrences of each phrase in the original text. Then, the frequently occurring keyword calculation unit 114 extracts phrases whose number of occurrences is equal to or greater than the frequent occurrence determination threshold as frequently occurring keywords. There is no particular limitation on what types of phrases are targeted as frequently occurring keywords. For example, proper nouns, compound nouns, nouns, etc. may be targeted.

[0045] For example, when the target of the frequently occurring keyword is a compound noun, if the original text is "··· The business process is as follows. (1) Member registration: It is a prerequisite that the user registers as a member. (2) Event creation: The operation staff creates an event (an event application page is generated). After obtaining approval from the management staff, it is published on the web. ···" and the frequent occurrence determination threshold is 1, the compound nouns "business process", "member registration", "user", "event creation", "event application", "operation staff", and "management staff" are extracted as frequently occurring keywords.

[0046] The extracted frequently occurring keywords are registered in the frequently occurring keyword management DB 103.

[0047] FIG. 8 is a diagram showing an example of the data in the frequently occurring keyword management DB.

[0048] The frequently occurring keyword management DB 103 stores data on frequently occurring keywords.

[0049] In association with the management ID assigned to each frequently occurring keyword, the article name, keyword, number of occurrences, and usage are recorded. The "article name" is the name of the original article. The "keyword" is the text data (character string) of the frequently occurring keyword. The "number of occurrences" is information indicating the number of times the frequently occurring keyword appears in the original article. The "usage" is information indicating whether the frequently occurring keyword is used in the summary article. Until it is determined whether the frequently occurring keyword is used in the summary article, the "usage" column remains blank.

[0050] Subsequently, in step S205, a summary article is generated. For the generation of the summary article, the natural language generation AI in the natural language generation AI system 20 is utilized. The summary article generation unit 120 requests the instruction information creation unit 117 to generate a summary article from the original article by specifying the summary length. The received instruction information creation unit 117 creates instruction information for generating a summary article of the specified summary length and requests the natural language generation AI to generate a summary article via the generation AI call unit 118. The summary article generated by the natural language generation AI is sent to the summary article generation unit 120 via the generation AI call unit 118 and the instruction information creation unit 117.

[0051] In step S206, the sentence construction unit 111 divides the original article into sentences. For example, the sentence construction unit 111 can divide the original article into sentences by separating it at full stops. The original article divided into sentences and the user's data are stored in the document management DB 102.

[0052] Figure 9 is a diagram showing an example of the data in the document management DB.

[0053] The document management unit 113 assigns a management ID to each sentence split from the original sentence by the sentence construction processing unit 109 and records it in the document management DB 102. In the document management DB 102, corresponding to the management ID for each sentence, the sentence name, user, sentence, maximum similarity value, and determination information of the sentence are recorded. The "sentence name" is the name of the original sentence. The "user" is the identification information of the user of the original sentence. The "sentence" is the text data (character string) of the sentence. The "maximum similarity value" is an index value for determining the quality of the summary sentence. Until the value is calculated, the column of the "maximum similarity value" is left blank. The "determination" is the result of comparing the maximum similarity value of the sentence with a predetermined threshold (sufficiency threshold). If the maximum similarity value is equal to or greater than the sufficiency threshold, the "determination" is sufficient, and if the maximum similarity value is less than the sufficiency threshold, the "determination" is insufficient. In the example of FIG. 9, further, an example where the "determination" of ignoring is made is shown. The determination of ignoring indicates that the sentence has been ignored in the generation of the summary in the original sentence. For example, a threshold smaller than the sufficiency threshold may be set, and if the maximum similarity value is less than or equal to the threshold, it may be determined as ignored. If there is a sentence designated as a part to be ignored when generating a summary in the original sentence, the maximum similarity value of that sentence will be a sufficiently small value. By the "determination" of that sentence being ignored, it can be confirmed that the sentence has been ignored as specified. Also, when the essential keyword or frequently occurring keyword included in the sentence of the original sentence does not appear in the summary sentence, it may be determined that the sentence has been ignored in the generation of the summary. Also, before the maximum similarity value is calculated or before the maximum similarity value is compared with the sufficiency threshold, the column of the "determination" is left blank.

[0054] Returning to FIG. 6, in step S207, the sentence construction unit 111 splits the summary sentence into sentences. For example, the sentence construction unit 111 can split the summary sentence into sentences by delimiting it with a period. The data of the summary sentence split into sentences is stored in the summary sentence management DB 106 by the summary sentence management unit 121.

[0055] FIG. 10 is a diagram showing an example of the data in the summary sentence management DB.

[0056] In the abstract text management DB 106, the text name, sentence, number of trials, and sentence length are recorded in association with the management ID. The "management ID" is assigned to each sentence of the abstract text. The "text name" is the name of the abstract text. The same name as the text name of the original text may be assigned to the text name of the abstract text. The "sentence" is the text data (string) of the sentence. The "number of trials" is information indicating how many times the abstract text was obtained by trying to generate it using the natural language generation AI. The "sentence length" is information indicating the length (number of characters) of the sentence.

[0057] Returning to FIG. 6, in step S208, the inter-sentence distance calculation unit 123 calculates the distance between sentences as a string of sentences for all combinations of the sentences of the original text and the sentences of the abstract text, and determines the similarity based on the distance. If the distance between sentences is close, it means that those sentences are similar to each other, so the similarity between sentences can be represented by the distance. Further, the inter-sentence distance calculation unit 123 calculates the maximum similarity value for each sentence of the original text based on the similarity between each sentence of the original text and each sentence of the abstract text. Note that since the distance or similarity between strings in this embodiment is an index used for comparison, it may be determined relatively, and the method of determining the absolute value is not particularly limited. Therefore, either a normalized value or a non-normalized value may be used. Also, either value that is in an inverse relationship with each other may be used.

[0058] The method for calculating the similarity will be described.

[0059] The method of quantifying the similarity between strings is not particularly limited, but in this embodiment, the Jaro distance or the Jaro-Winkler distance will be used. These are methods of calculating the distance from the number of matching characters within the specified number of characters and the presence or absence of replacements, and expressing the similarity from the number of matching characters. It is characterized by the point of measuring the similarity between sentences by focusing on one character, and the point of fixing the word detection range and making the matching rate variable. Since it is possible not to strictly consider the coincidence of the positions of characters and the similarity can be easily controlled, the Jaro distance and the Jaro-Winkler distance are suitable for evaluating the quality of the summary sentence. However, you may calculate the similarity by other methods such as the Levenshtein distance, Gestalt pattern matching, the optimal transport distance between word vectors, and the rotation distance between word vectors.

[0060] The similarity (distance) between the original text sentence and the summary sentence of this embodiment is calculated as follows, with the sentence of the original text as the first string and the sentence of the summary text as the second string. The distance Φ between the first string and the second string = W 1 ·(c / d)+W 2 ·(c / r)+W t ·((c-t) / c) Here, W 1 is the weight applied to the first string. W 2 is the weight applied to the second string. W t is the weight applied to the replacement. d is the length (number of characters) of the first string. r is the length (number of characters) of the second string. c is the number of characters that match (hereinafter also referred to as "matching characters") within the interval of (max(d,r) / 2 - 1) characters between the first string and the second string. t is half of the number of characters that need to be replaced to match the permutation of the matching characters arranged in the order of appearance in the first string and the permutation of the matching characters arranged in the order of appearance in the second string.

[0061] Figure 11 is a flowchart for explaining the flow of calculating the similarity.

[0062] In step S301, the inter-sentence distance calculation unit 123 defines a first character string and a second character string. In the example of FIG. 11, the first character string is "あいうえおかきくけこ" (aiueookikukeko), and the second character string is "あえいおうおさしすせそ" (aeiouosa shisuso).

[0063] In step S302, the inter-sentence distance calculation unit 123 defines weights W 1 , W 2 , W t . The weights W 1 , W 2 , W t are not particularly limited. In the example of FIG. 11, W 1 = 5, W 2 = 10, W t = 1 are defined. By defining W 2 as a value larger than W 1 , the weight is strengthened in the summary sentence compared to the original sentence. By defining W t as a value smaller than W 2 and W 1 , the weight of the rearrangement of the word order is weakened.

[0064] In step S303, the inter-sentence distance calculation unit 123 calculates the string length (d) of the first character string and the string length (r) of the second character string. In the example of FIG. 11, d = 10 and r = 11.

[0065] In step S304, the inter-sentence distance calculation unit 123 defines an interval. In the example of FIG. 11, since d = 10 and r = 11, the interval (max(d, r) / 2 - 1) = 4.5.

[0066] In step S305, the inter-sentence distance calculation unit 123 extracts the overlapping characters between the first character string and the second character string. In the example of FIG. 11, "あいうえお" (aiueo) is extracted from the first character string, and "あえいおうお" (aeiouo) is extracted from the second character string.

[0067] In step S306, the inter-sentence distance calculation unit 123 calculates the number of characters t to be replaced. In the example of FIG. 11, since there are 4 replacements, t = 2.

[0068] In step S307, the inter-sentence distance calculation unit 123 calculates the number of matching characters c. In the example of FIG. 11, since 6 characters match, c = 6.

[0069] In step S308, the inter-sentence distance calculation unit 123 calculates the distance Φ. In the example of FIG. 11, the distance Φ = W 1 ·(c / d)+W 2 ·(c / r)+W t ·((c - t) / c)=9.12.

[0070] Based on the similarity calculated as above, the maximum similarity value is calculated. The method for calculating the maximum similarity value is as follows. For each first string, the inter-sentence distance calculation unit 123 identifies the maximum value among the similarities of the first string with each of the second strings as the maximum similarity value for the first string.

[0071] Returning to FIG. 6, in step S209, the keyword comparison unit 125 searches the summary text using the essential keywords registered in the essential keyword management DB 104 as keys. Further, in step S210, the keyword comparison unit 125 searches the summary text using the frequently occurring keywords registered in the frequently occurring keyword management DB 103 as keys. Further, in step S211, the user setting information comparison unit 124 checks the length of the summary text.

[0072] Then, in step S212, the inter-sentence distance calculation unit 123 determines whether the maximum similarity value for any sentence (first string) of the original text is less than the sufficiency threshold value registered in the user setting management DB 105.

[0073] FIG. 12 and FIG. 13 are diagrams for explaining the determination of the quality of the summary text based on the maximum similarity value. In FIGS. 12 and 13, graphs are shown with each sentence of the original text on the horizontal axis and the maximum similarity value on the vertical axis.

[0074] Looking at the graph in FIG. 12, the maximum similarity value is below the sufficiency threshold in any one of the sentences. Therefore, in this case, it is determined that the quality of the summary text does not meet the quality conditions. Looking at the graph in FIG. 13, the maximum similarity value is above the sufficiency threshold in any of the sentences. Therefore, in this case, it is not determined that the quality of the summary text does not meet the quality conditions.

[0075] If the maximum similarity value for any one of the sentences is less than the sufficiency threshold, the sentence distance calculation unit 123 determines that the quality of the summary text does not meet the quality conditions and proceeds to step S216.

[0076] If the maximum similarity value for any of the sentences is greater than or equal to the sufficiency threshold, at step S213, the keyword comparison unit 125 determines whether the essential keywords are present in the summary text. If any of the essential keywords are not present in the summary text, the keyword comparison unit 125 determines that the quality of the summary text does not meet the quality conditions and proceeds to step S216.

[0077] If all of the essential keywords are present in the summary text, at step S214, the keyword comparison unit 125 determines whether the frequently occurring keywords are present in the summary text. If any of the frequently occurring keywords are not present in the summary text, the keyword comparison unit 125 determines that the quality of the summary text does not meet the quality conditions and proceeds to step S216.

[0078] If all of the frequently occurring keywords are present in the summary text, at step S215, the user setting information comparison unit 124 checks whether the length of the summary text is less than or equal to the required summary length registered in the user setting management DB 105.

[0079] If the length of the abstract text is not less than the required abstract length, the user setting information comparison unit 124 determines that the quality of the abstract text does not meet the quality conditions and proceeds to step S216.

[0080] In step S216, the supplementary information generation unit 126 creates supplementary information with content that improves the quality index that did not meet the quality conditions.

[0081] At this time, if it is determined that the quality of the abstract text does not meet the quality conditions because the maximum similarity value for any sentence in the original text was less than the sufficiency threshold in the determination of step S212, the supplementary information generation unit 126 points out the possibility of lack of information included in the sentence where the maximum similarity value was less than the sufficiency threshold, and / or creates supplementary information instructing to include that information. Thereby, when the abstract text is generated again, the lack of that information can be suppressed.

[0082] Also, if it is determined that the quality of the abstract text does not meet the quality conditions because the essential keyword was not included in the abstract text in the determination of step S213, the supplementary information generation unit 126 points out the possibility of lack of that essential keyword, and / or creates supplementary information instructing to include the essential keyword. Thereby, when the abstract text is generated again, the lack of information related to that essential keyword can be suppressed.

[0083] Also, if it is determined that the quality of the abstract text does not meet the quality conditions because the frequently occurring keyword was not included in the abstract text in the determination of step S214, the supplementary information generation unit 126 points out the possibility of lack of that frequently occurring keyword, and / or creates supplementary information instructing to include the frequently occurring keyword. Thereby, when the abstract text is generated again, the lack of information related to that frequently occurring keyword can be suppressed.

[0084] Also, if it is determined in the determination of step S215 that the length of the summary text exceeds the required summary length and thus the quality of the summary text does not meet the quality conditions, the supplementary information generation unit 126 creates supplementary information that points out that the length of the summary text exceeds the required summary length and / or instructs to make the length of the summary text not exceed the required summary length. Thereby, when generating the summary text again, it is possible to suppress the length of the summary text from exceeding the required summary length.

[0085] In addition, in this embodiment, an example of setting an upper limit value of the required summary length for the length of the summary text is shown, but other configurations are also possible. For example, a required summary length range may be set by an upper limit value and a lower limit value, and the summary text may be generated so as to fall within that range. In that case, if the length of the summary text exceeds the upper limit value, supplementary information that points out that the length of the summary text exceeds the upper limit value and / or instructs to make the length of the summary text not exceed the upper limit value is created, and if the length of the summary text is below the lower limit value, supplementary information that points out that the length of the summary text is below the lower limit value and / or instructs to make the length of the summary text not less than the lower limit value may be created.

[0086] After step S216, the process returns to step S205, and the instruction information creation unit 117 creates instruction information using the supplementary information and causes the natural language generation AI to generate a summary text.

[0087] The embodiment described above is an exemplification for explaining the present invention, and is not intended to limit the scope of the present invention only to those embodiments. Those skilled in the art can implement the present invention in various other modes without departing from the scope of the present invention.

[0088] The above embodiment includes the following matters. However, the matters included in the above embodiment are not limited to those shown below.

[0089] (Matter 1) The summary generation system includes a summary text processing unit that generates a summary text by summarizing the original text, and an evaluation processing unit that determines whether the quality of the summary text meets a predetermined quality condition. If the quality does not meet the quality condition, supplementary information for improving the quality is created, a summary text is generated using the supplementary information, and the quality of the generated summary text is determined. Since the summary text is generated by the supplementary information so as to improve the quality if the quality of the summary text does not meet the predetermined quality condition, it becomes possible to generate a summary text with good quality.

[0090] (Item 2) In the summary generation system according to Item 1, the evaluation processing unit sets each individual sentence included in the original text as a first character string, and each individual sentence included in the summary text as a second character string. For each first character string, the maximum value among the similarities of the first character string with each of the second character strings is specified as the maximum similarity value for the first character string, and based on the maximum similarity value for the first character string, it is determined whether the quality of the summary text meets the quality condition. Thereby, the quality of the summary text can be appropriately and easily evaluated by the similarity between character strings.

[0091] (Item 3) In the summary generation system according to Item 2, the similarity between the first character string and the second character string is the weight applied to the first character string as W 1 and the weight applied to the second character string as W 2 and the weight applied to replacement as W t and the length of the first character string as d, the length of the second character string as r, the number of matching characters c that are characters that match within the interval of (max(d,r) / 2 - 1) characters between the first character string and the second character string, and 1 / 2 of the number of characters to be replaced in order to match the permutation of the matching characters arranged in the order in which they appear in the first character string and the permutation of the matching characters arranged in the order in which they appear in the second character string as t, Φ = W 1 ·(c / d)+W 2 ·(c / r)+W t· It is determined based on the distance Φ between the first string and the second string, which is calculated by the formula ((c - t) / c).

[0092] (Item 4) In the summary generation system described in Item 3, W 2 >W 1 That is. By using the similarity calculated with emphasis on the summary text rather than the original text for the purpose of generating an appropriate summary text, an appropriate evaluation of the quality of the summary text becomes possible.

[0093] (Item 5) In the summary generation system described in Item 4, W 2 >W 1 >W t That is. By considering that the word order can be swapped even if the summary is appropriate and using the similarity calculated with allowance for the swapped word order, an appropriate evaluation of the quality of the summary text becomes possible.

[0094] (Item 6) In the summary generation system described in Item 2, if the maximum similarity value for any first string is less than a predetermined sufficiency threshold, the evaluation processing unit determines that the quality of the summary text does not meet the quality condition. ⇒ By determining that the summary text does not meet the condition if the maximum similarity value for any first string is less than a predetermined sufficiency threshold, it is possible to suppress the omission of matters included in the original text in the summary text.

[0095] (Item 7) In the summary generation system described in Item 6, the evaluation processing unit points out the possibility of omission of information included in the first string for which the maximum similarity value is less than the sufficiency threshold, and / or creates supplementary information instructing to include the information included in the first string. Thereby, when the maximum similarity value for any first string is less than a predetermined sufficiency threshold, it is possible to regenerate the summary text so that the information of that first string is not omitted.

[0096] (Item 8) In the summary generation system described in Item 6, if the given essential phrase is not included in the summary text, the evaluation processing unit further determines that the quality of the summary text does not meet the quality conditions. By determining that the summary text does not meet the conditions if the essential phrase is not included in the summary text, it is possible to prevent important information contained in the original text from being missing in the summary text.

[0097] (Item 9) In the summary generation system described in Item 8, if the essential phrase is not included in the summary text, the evaluation processing unit points out the possibility of omission of the essential phrase and / or creates supplementary information for instructing to include the essential phrase. When the essential phrase is not included in the summary text, it is possible to regenerate the summary text so that information related to the essential phrase is not missing.

[0098] (Item 10) The summary generation system described in Item 6 further includes an information management unit that extracts, as frequently occurring phrases, phrases that appear in the original text more than a predetermined frequent occurrence determination threshold, and if the frequently occurring phrase is not included in the summary text, the evaluation processing unit further determines that the quality of the summary text does not meet the quality conditions. By determining that the summary text does not meet the conditions if the frequently occurring phrase is not included in the summary text, it is possible to prevent information that frequently appears in the original text from being missing in the summary text.

[0099] (Item 11) In the summary generation system described in Item 10, if the frequently occurring phrase is not included in the summary text, the evaluation processing unit creates supplementary information for instructing to include the frequently occurring phrase. When the frequently occurring phrase is not included in the summary text, it is possible to regenerate the summary text so that information related to the frequently occurring phrase is not missing.

[0100] (Item 12) In the summary generation system described in Item 6, if the length of the summary text is outside the predetermined summary length range, the evaluation processing unit further determines that the quality of the summary text does not meet the quality conditions. By determining that the summary text does not meet the conditions if the length of the summary text is outside the summary length range, it is possible to prevent the length of the summary text from deviating from an appropriate range.

[0101] (Item 13) In the summary generation system described in Item 12, if the length of the summary text is outside the summary length range, the evaluation processing unit creates supplementary information indicating that the length of the summary text is outside the summary length range and / or instructing to make the length of the summary text within the summary length range. When the length of the summary text is outside the summary length range, the summary text can be regenerated so as to have a length within the summary length range.

Explanation of Signs

[0102] 10…Article summary generation system, 20…Natural language generation artificial intelligence system, 30…Database, 91…Administrator, 92…User, 101…Memory unit, 102…Article management DB, 102…Document management DB, 103…Frequently appearing keyword management DB, 104…Essential keyword management DB, 105…User setting management DB, 106…Summary article management DB, 107…Processing unit, 108…User interface unit, 109…Sentence construction processing unit, 110…Morphological analysis unit, 111…Sentence construction unit, 112…Information management unit, 113…Document management unit, 114…Frequently appearing keyword calculation unit, 115…User setting information management unit, 116…Generated AI interface unit, 117…Instruction information creation unit, 118…Generated AI call unit, 119…Summary article processing unit, 120…Summary article generation unit, 121…Summary article management unit, 122…Evaluation processing unit, 123…Inter-sentence distance calculation unit, 124…User setting information comparison unit, 125…Keyword comparison unit, 126…Supplementary information generation unit

Claims

1. An abstract text processing unit that generates an abstract text summarizing the original text, determines whether the quality of the abstract text meets a predetermined quality condition. If the quality does not meet the quality condition, supplementary information for improving the quality is created, an abstract text is generated using the supplementary information, and an evaluation processing unit that determines the quality of the generated abstract text. An abstract generation system having the above.

2. The evaluation processing unit sets each individual sentence included in the original text as a first character string, sets each individual sentence included in the abstract text as a second character string, and for each first character string, determines the maximum value among the similarities with each of the second character strings of the first character string as the maximum similarity value for the first character string, and based on the maximum similarity value for the first character string, determines whether the quality of the abstract text meets the quality condition. The abstract generation system according to Claim 1.

3. The similarity between the first character string and the second character string is Let the weight applied to the first string be W 1 and Let the weight applied to the second string be W 2 and Let the weight to be replaced be W t and Let the length of the first character string be d, Let the length of the second character string be r, Let the number of matching characters, which are characters that match within the interval of (max(d, r) / 2 - 1) characters between the first character string and the second character string, be c, Let 1 / 2 of the number of characters to be replaced to match the permutation of the matching characters arranged in the order in which they appear in the first character string and the permutation of the matching characters arranged in the order in which they appear in the second character string be t. Φ = W 1 · (c / d) + W 2 · (c / r) + W t Determined based on the distance Φ between the first character string and the second character string, calculated by the formula · ((c - t) / c) The abstract generation system according to Claim 2.

4. W 2 > W 1 is true, The abstract generation system according to Claim 3.

5. W 2 > W 1 > W t is, The abstract generation system according to Claim 4.

6. If the maximum similarity value for any first character string is less than a predetermined sufficiency threshold, the evaluation processing unit determines that the quality of the abstract text does not meet the quality condition. The abstract generation system according to Claim 2.

7. The evaluation processing unit points out the possibility of information loss included in the first character string for which the maximum similarity value is less than the sufficiency threshold, and / or creates supplementary information that instructs to include the information included in the first character string. The abstract generation system according to Claim 6.

8. The evaluation processing unit further determines that the quality of the abstract text does not meet the quality condition if a given essential phrase is not included in the abstract text. The abstract generation system according to Claim 6.

9. If the evaluation processing unit determines that the essential phrase is not included in the summary sentence, it creates supplementary information that points out the possibility of omission of the essential phrase and / or gives an instruction to include the essential phrase. The summary generation system according to claim 8.

10. The system further includes an information management unit that extracts, as frequently occurring phrases, phrases that appear in the original text more than a predetermined frequent-occurrence determination threshold value. If the evaluation processing unit determines that the frequently occurring phrase is not included in the summary sentence, it determines that the quality of the summary sentence does not meet the quality condition. The summary generation system according to claim 6.

11. If the evaluation processing unit determines that the frequently occurring phrase is not included in the summary sentence, it creates supplementary information that gives an instruction to include the frequently occurring phrase. The summary generation system according to claim 10.

12. If the evaluation processing unit determines that the length of the summary sentence is outside a predetermined summary length range, it determines that the quality of the summary sentence does not meet the quality condition. The summary generation system according to claim 6.

13. If the evaluation processing unit determines that the length of the summary sentence is outside the summary length range, it points out that the length of the summary sentence is outside the summary length range and / or creates supplementary information that gives an instruction to bring the length of the summary sentence within the summary length range. The summary generation system according to claim 12.

14. Generate a summary sentence that summarizes the original text. Determine whether the quality of the summary sentence meets a predetermined quality condition. If the quality does not meet the quality condition, create supplementary information for improving the quality, generate a summary sentence using the supplementary information, and determine the quality of the generated summary sentence. A summary generation method for a computer to execute.

15. Generate a summary sentence that summarizes the original text. Determine whether the quality of the summary sentence meets a predetermined quality condition. If the quality does not meet the quality condition, create supplementary information for improving the quality, generate a summary sentence using the supplementary information, and determine the quality of the generated summary sentence. A summary generation program for causing a computer to execute.

Citation Information

Patent Citations

  • Method and system for summarizing document using hyperscale language model

    JP2023053867A

Cited By

  • Character string processing device and character string processing method

    WO2026120671A1