Method, apparatus, electronic device and readable medium for generating abstract of text
By filtering text sentences with high correlation as candidate text sentences in the generation of long text summary, the problems of poor abstract consistency and deviation from the meaning of text expression are solved, and the consistency of text summary and the matching degree of original text are improved.
Patent Information
- Application Number
- CN202111277688.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-10-29
- Publication Date
- 2025-07-15
- Estimated Expiration
- 2041-10-29
AI Technical Summary
The prior art has problems of poor summary coherence and deviation from the meaning of text expression in the generation of long text summary.
By obtaining the correlation scores between each original text sentence of the preset text and other text sentences, filter out candidate text sentences with high correlation, and generate text summary based on these candidate text sentences to ensure the consistency of the summary and the matching degree with the original text.
It effectively ensures the consistency of text abstracts and the accuracy of the original text expression, ensuring that the generated abstracts can effectively express the meaning of the original text.
Smart Images

Figure CN114090731B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of text processing, and in particular to a method for generating a summary of a text, a device for generating a summary of a text, an electronic device, and a computer-readable medium. Background Art
[0002] With the explosive growth of text information, people can come into contact with a vast amount of text information every day, such as news, meeting records, blogs, chats, reports, papers, microblogs, etc. Among them, for summary work such as long target text summaries, such as work meeting records and summaries, most of them are completed manually, which undoubtedly consumes a great deal of the time and manpower of workers. Therefore, it has become increasingly important to extract important content from a large amount of text information, and the technology of automatic target text summarization, which enables users to obtain information more quickly and accurately, has thus emerged. Extracting important content from a large amount of text information has become an urgent need for us, and automatic target text summarization provides an efficient solution. Automatic target text summarization technology effectively compresses and refines document information, helps users retrieve the relevant information they need from a vast amount of information, avoids the problem of excessive redundant and one-sided information that may be generated by retrieving through a search engine, and effectively solves the problem of information overload. However, in the process of generating a summary of a long text using related technologies, on the one hand, in order to ensure content integrity, there are problems such as a large amount and complexity of the summary content and weak generalization. On the other hand, in order to make the summary concise enough, important information is easily ignored, resulting in discontinuous content and the inability to ensure the central idea of the text, leading to deviation from the true meaning expressed by the text. Summary of the Invention
[0003] Embodiments of the present invention provide a method, a device, an electronic device, and a computer-readable storage medium for generating a summary of a text, so as to solve or partially solve the problem that in the process of generating a summary of a text in related technologies, there are problems such as poor coherence of the summary and easy deviation from the meaning expressed by the text.
[0004] Embodiments of the present invention disclose a method for generating a summary of a text, including:
[0005] Obtaining a preset text, where the preset text includes a plurality of original text sentences;
[0006] Determining a relevance score between each of the original text sentences and other text sentences;
[0007] Extracting candidate text sentences from each of the original text sentences according to the relevance score;
[0008] Generating a target text summary corresponding to the preset text according to the candidate text sentences.
[0009] Optionally, determining the correlation scores between each of the original text sentences and other text sentences includes:
[0010] Inputting the original text sentence into a sentence correlation model to obtain the correlation scores between the original text sentence and other text sentences in the preset text.
[0011] Optionally, extracting candidate text sentences from each of the original text sentences according to the correlation scores includes:
[0012] Using the respective correlation scores of the original text sentence to generate a sentence score for the original text sentence;
[0013] Taking the original text sentences in the preset text whose sentence scores are greater than or equal to a preset score threshold as candidate text sentences of the preset text.
[0014] Optionally, generating a target text summary corresponding to the preset text according to the candidate text sentences includes:
[0015] Determining a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding respective correlation scores;
[0016] Using the starting text sentence and the at least one associated text sentence to generate a target text summary corresponding to the preset text.
[0017] Optionally, determining a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding respective correlation scores includes:
[0018] Taking the candidate text sentence with the highest sentence score in the preset text as the starting text sentence;
[0019] Taking the candidate text sentences in the preset text that are after the starting text sentence as target text sentences;
[0020] Determining at least one associated text sentence according to the correlation score of the starting text sentence, the correlation scores between each of the target text sentences and the target text sentences.
[0021] Optionally, determining at least one associated text sentence according to the correlation score of the starting text sentence, the correlation scores between each of the target text sentences and the target text sentences includes:
[0022] Taking the candidate text sentence with the highest correlation score corresponding to the starting text sentence as the associated text sentence associated with the starting text sentence;
[0023] Determine whether there is a candidate text sentence after the associated text sentence in the preset text;
[0024] If there is a candidate text sentence after the associated text sentence in the preset text, then use the candidate text sentence with the highest relevance score corresponding to the associated text sentence as the new associated text sentence, and return to the step of determining whether there is a candidate text sentence after the associated text sentence;
[0025] When all candidate text sentences in the target text sentence are traversed, obtain at least one associated text sentence.
[0026] Optionally, after determining whether there is a candidate text sentence after the associated text sentence in the preset text, the method further includes:
[0027] If there is no candidate text sentence after the associated text sentence in the preset text, stop traversing the target text sentence and obtain at least one associated text sentence.
[0028] Optionally, before obtaining at least one associated text sentence when all candidate text sentences in the target text sentence are traversed, the method further includes:
[0029] Obtain the current text summary composed of the starting text sentence and the associated text sentence, and obtain the first text length of the current text summary;
[0030] Use the preset text threshold and the text length of the preset text to determine the second text length;
[0031] If the text length is greater than or equal to the second text length, stop traversing the target text sentence, and use the current text summary as the target text summary of the preset text.
[0032] An embodiment of the present invention also discloses a text summary generation device, including:
[0033] A preset text acquisition module, configured to acquire a preset text, where the preset text includes a plurality of original text sentences;
[0034] A relevance determination module, configured to determine the relevance scores between each of the original text sentences and other text sentences;
[0035] A candidate text sentence determination module, configured to extract candidate text sentences from each of the original text sentences according to the relevance scores;
[0036] A text summary generation module, configured to generate a target text summary corresponding to the preset text according to the candidate text sentences.
[0037] Optionally, the relevance determination module is specifically configured to:
[0038] Input the original text sentence into the sentence relevance model to obtain the relevance scores between the original text sentence and other text sentences in the preset text.
[0039] Optionally, the candidate text sentence determination module includes:
[0040] A sentence score generation sub-module, configured to generate a sentence score of the original text sentence by using the respective relevance scores of the original text sentence;
[0041] A candidate text sentence determination sub-module, configured to use other text sentences whose relevance scores corresponding to the original text sentence are greater than or equal to a preset score threshold as candidate text sentences corresponding to the original text sentence.
[0042] Optionally, the text summary generation module includes:
[0043] A text sentence determination sub-module, configured to determine a starting text sentence and at least one associated text sentence according to the sentence scores of each candidate text sentence and the respective relevance scores;
[0044] A text summary generation module, configured to generate a target text summary corresponding to the preset text by using the starting text sentence and the at least one associated text sentence.
[0045] Optionally, the text sentence determination sub-module includes:
[0046] A starting text sentence determination unit, configured to use the candidate text sentence with the largest sentence score in the preset text as the starting text sentence;
[0047] A target text sentence determination unit, configured to use the candidate text sentences in the preset text that are after the starting text sentence as target text sentences;
[0048] An associated text sentence determination unit, configured to determine at least one associated text sentence according to the relevance score of the starting text sentence, the relevance scores of each target text sentence and the target text sentence.
[0049] Optionally, the associated text sentence determination unit is specifically configured to:
[0050] Use the candidate text sentence with the highest relevance score corresponding to the starting text sentence as the associated text sentence associated with the starting text sentence;
[0051] Determine whether there are candidate text sentences in the preset text that are after the associated text sentence;
[0052] If there is a candidate text sentence after the associated text sentence in the preset text, then use the candidate text sentence with the highest relevance score corresponding to the associated text sentence as the new associated text sentence, and return the step of determining whether there is a candidate text sentence after the associated text sentence;
[0053] When all candidate text sentences in the target text sentence have been traversed, at least one associated text sentence is obtained.
[0054] Optionally, the associated text sentence determining unit is further specifically configured to:
[0055] If there is no candidate text sentence after the associated text sentence in the preset text, stop traversing the target text sentence and obtain at least one associated text sentence.
[0056] Optionally, the text sentence determining sub-module further includes:
[0057] A first text length determining unit, configured to obtain a current text summary formed by the starting text sentence and the associated text sentence, and obtain a first text length of the current text summary;
[0058] A second text length determining unit, configured to determine a second text length by using a preset text threshold and the text length of the preset text;
[0059] A text summary generating unit, configured to, if the text length is greater than or equal to the second text length, stop traversing the target text sentence and use the current text summary as the target text summary of the preset text.
[0060] An embodiment of the present invention also discloses an electronic device, including a processor, a communication interface, a memory, and a communication bus, wherein the processor, the communication interface, and the memory communicate with each other through the communication bus;
[0061] The memory is used to store a computer program;
[0062] The processor is configured to, when executing the program stored in the memory, implement the method as described in the embodiment of the present invention.
[0063] An embodiment of the present invention also discloses one or more computer-readable media, on which instructions are stored, and when executed by one or more processors, cause the processors to execute the method as described in the embodiment of the present invention.
[0064] The embodiments of the present invention have the following advantages:
[0065] In an embodiment of the present invention, the preset text may be a long text with a text length greater than or equal to a certain threshold. Then, in the process of generating a text summary for the long text, the original text sentences of the preset text can be obtained first. Next, the relevance score between each original text sentence and other text sentences in the preset text is determined, and candidate text sentences corresponding to each original text sentence are screened according to the relevance score of the original text sentence. Then, a text summary of the preset text can be generated based on the candidate text sentences. Thus, when generating a summary for a long text, by the relevance between each sentence in the text, sentences with high relevance are screened out as the summary of the text, effectively ensuring the coherence of the text summary, and the sentences are extracted based on the original text, so that the generated summary can effectively express the meaning of the original text, ensuring the matching degree between the summary and the original text. BRIEF DESCRIPTION OF THE DRAWINGS
[0066] Figure 1 is a flowchart of the steps of a method for generating a text summary provided in an embodiment of the present invention;
[0067] Figure 2 is a block diagram of the structure of a device for generating a text summary provided in an embodiment of the present invention;
[0068] Figure 3 is a block diagram of an electronic device provided in an embodiment of the present invention;
[0069] Figure 4 is a schematic diagram of a computer-readable medium provided in an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0070] To make the above objects, features, and advantages of the present invention more obvious and understandable, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0071] As an example, automatic text summarization can effectively compress and refine document information, helping users retrieve the relevant information they need from a vast amount of information, avoiding the problem that users may retrieve too much redundant and one-sided information through a search engine, or reducing the problem that users need to read a large amount of document information, and effectively solving the problem of information overload.
[0072] Among them, in the related art, the method for generating a summary of a long text is mainly to segment the input original text and then use an end-to-end model to obtain the generated text summary. In this process, the summaries generated for each segment are directly concatenated together to obtain the text summary of the long text. However, this method is too dependent on the quality of the segmentation technology. When the long text is segmented, the relevance between each paragraph is poor, so the relevance between paragraphs cannot be guaranteed, and further the relevance between sentences in the generated text summary will be poor, seriously affecting the coherence and semantic expression of the text summary.
[0073] In this regard, one of the core inventive points of the embodiments of the present invention lies in, for a preset text, by obtaining each original text sentence of the preset text, then determining the relevance score between each original text sentence and other text sentences in the preset text, and based on the relevance scores of the original text sentences, screening out candidate text sentences that can be used to form a text summary from all the original text sentences, and then generating a text summary of the preset text according to the candidate text sentences. Thus, when generating a summary of a long text, by the relevance between each sentence in the text, high-relevance sentences are screened out as the summary of the text, effectively ensuring the coherence of the text summary, and the sentences are extracted based on the original text, enabling the generated summary to effectively express the meaning of the original text and ensuring the matching degree between the summary and the original text.
[0074] To enable those skilled in the art to better understand the technical solutions of the embodiments of the present invention, some technical terms involved in the embodiments of the present invention are explained and described below:
[0075] Original text sentence, which can be a sentence in the preset text. For example, if the original text includes 10 sentences, each sentence can be an original text sentence.
[0076] Other text sentences. In the preset text, for a certain original text sentence, all other sentences in the text can be the other sentences of this original text sentence. For example, if the original text includes original text sentence a, original text sentence b, and original text sentence c, etc., then for original text sentence a, original text sentence b and original text sentence c can be the other text sentences of original text sentence a; for original text sentence b, original text sentence a and original text sentence c can be the other text sentences of original text sentence a; similarly, for original text sentence c as well, which will not be elaborated here.
[0077] Candidate text sentence. In the process of generating a text summary, in order to ensure the relevance between each sentence in the text summary, high-relevance original text sentences can be screened out from the preset text as the sentences for generating the text summary. For example, if the original text sentences include original text sentence a, original text sentence b, and original text sentence c, etc., and among them, the high-relevance original text sentences include original text sentence a and original text sentence c, then original text sentence a and original text sentence c can be used as candidate text sentences for the subsequent sentence materials for generating the text summary.
[0078] Starting text sentence, which can be a sentence screened out from the candidate text sentences as the starting sentence in the text summary. By determining the starting text sentence, irrelevant sentences in the text summary can be effectively reduced, and while ensuring the matching degree between the text summary and the original text, the text length of the text summary can be reduced.
[0079] After determining the starting text sentence, the candidate text sentences after the starting text sentence in the preset text can be used as the target text sentence, so as to screen out at least one associated text sentence from the target text sentences according to the relevance between sentences.
[0080] An associated text sentence is a sentence that can be screened out from the target text sentences according to the relevance between sentences and is used to form a text summary.
[0081] Specifically, referring to Figure 1 , it shows the step flow chart of a method for generating a summary of a text provided in an embodiment of the present invention, which may specifically include the following steps:
[0082] Step 101, obtain a preset text, where the preset text includes a number of original text sentences;
[0083] In practice, for the preset text, it may include a text with a text length greater than or equal to a certain word count threshold. For example, the word count threshold may be 1024, and the preset text may be a long text with a length greater than or equal to 1024 words, etc. It should be noted that in the embodiments of the present invention, an example is given with the preset text being a long text. It can be understood that the present invention is not limited thereto.
[0084] Step 102, determine the relevance scores between each of the original text sentences and other text sentences;
[0085] For each original text sentence in the preset text, in order to ensure the coherence of the generated text summary, the relevance between each original text sentence and other text sentences in the preset text can be obtained first, and then the text summary corresponding to the preset text can be determined according to the relevance.
[0086] In a specific implementation, each original text sentence can be input into a sentence relevance model to obtain the relevance score between each original text sentence and other text sentences in the preset text. Among them, the sentence relevance model can be a recursive neural network model, which can use the question-answer pair as a positive sample, and at the same time use a similarity model to construct several negative samples for each question, and then use the positive sample and the negative sample as the training data of the model for model training. Thus, in the process of generating a summary of the text, the relevance score between each sentence and other sentences in the text can be obtained through the sentence relevance model.
[0087] In one example, assuming that the preset text includes statements such as original text sentence a, original text sentence b, original text sentence c, original text sentence d, and original text sentence e, etc., each original text sentence can be input into the sentence relevance model to obtain the relevance scores with other text sentences, including the relevance score A1 between original text sentence a and original text sentence b, the relevance score A2 with original text sentence c, the relevance score A3 with original text sentence d, and the relevance score A4 with original text sentence e; the relevance score B1 between original text sentence b and original text sentence a, the relevance score B2 with original text sentence c, the relevance score B3 with original text sentence d, and the relevance score B4 with original text sentence e; similarly, each relevance score C1, C2, C3, and C4 of original text sentence c, each relevance score D1, D2, D3, and D4 of original text sentence d, and each relevance score E1, E2, E3, and E4 of original text sentence e can be obtained.
[0088] Step 103, extract candidate text sentences from each of the original text sentences according to the relevance scores;
[0089] Regarding the relevance between an original text sentence and other text sentences in the preset text, the degree of relevance can be judged by the level of the relevance score. When the relevance score is higher, the relevance between the sentences is higher; when the relevance score is lower, the relevance between the sentences is lower. For this, the relevance scores between the original text sentence and other original text sentences in the preset text can be used to generate the sentence score of the original text sentence first, and then the original text sentences in the preset text with sentence scores greater than or equal to the preset score threshold can be used as candidate text sentences of the preset text, so as to screen out candidate sentences with satisfied relevance in the preset text, so as to generate a text summary according to the sentences with high relevance and ensure the coherence of the text summary. Optionally, the sentence score can be the average value of all relevance scores in the same original text sentence.
[0090] In one example, the respective relevance scores A1, A2, A3, and A4 of the original text sentence a, the average value A'; the respective relevance scores B1, B2, B3, and B4 of the original text sentence b, the average value B'; the respective relevance scores C1, C2, C3, and C4 of the original text sentence c, the average value C'; the respective relevance scores D1, D2, D3, and D4 of the original text sentence d, the average value D'; the respective relevance scores E1, E2, E3, and E4 of the original text sentence e, the average value E', etc., and the preset score threshold can be 0.9. Then, when the average value is less than 0.9, it indicates that the relevance between the original text sentence and the corresponding sentence is relatively low. To ensure the coherence of the text summary, the corresponding sentence is not used as a candidate text sentence. Assume that the average values A', C', D', and E' of the original text sentences a, c, d, and e are all greater than or equal to 0.9. Then, the original text sentences a, c, d, and e can be used as candidate text sentences, and the original text sentence b can be discarded. Thus, after obtaining the relevance between each original text sentence and other text sentences in the preset text, the sentences with high relevance can be screened out as candidate text sentences according to the level of relevance, so as to generate a text summary based on the sentences with high relevance and ensure the coherence of the text summary.
[0091] Step 104, generate a target text summary corresponding to the preset text according to the candidate text sentences.
[0092] In an embodiment of the present invention, the starting text sentence and at least one associated text sentence can be determined according to the sentence scores of each candidate text sentence and the corresponding respective relevance scores. Then, the starting text sentence and the at least one associated text sentence are used to generate a target text summary corresponding to the preset text.
[0093] In the process of generating the text summary, the candidate text sentence with the highest sentence score can be used as the starting text sentence of the text summary. For example, if A' is greater than B', C', D', and E', then the original text sentence a can be used as the starting text sentence of the text summary of the preset text, and starting from this original text sentence a, subsequent associated text sentences are screened out.
[0094] After the starting text sentence is determined from the preset text, the candidate text sentences located after the starting text sentence in the preset text can be used as target text sentences. Then, according to the relevance score of the starting text sentence and the relevance scores between the target text sentences and the target text sentences, at least one associated text sentence is determined. The target text sentence can be a sentence located after the starting text sentence in the preset text. For example, if the starting text sentence is the original text sentence a, the target text sentences can include the original text sentences c, d, and e; if the starting text sentence is the original text sentence c, the target text sentences can include d, e, etc. That is, when determining the associated text sentence, only the sentences behind are determined.
[0095] In a specific implementation, after determining the starting text sentence, the candidate text sentence with the highest relevance score corresponding to the starting text sentence can be used as the associated text sentence associated with the starting text sentence. Then, it is determined whether there is a candidate text sentence after the associated text sentence in the preset text. If there is a candidate text sentence after the associated text sentence in the preset text, the candidate text sentence with the highest relevance score corresponding to the associated text sentence is used as the new associated text sentence, and the step of determining whether there is a candidate text sentence after the associated text sentence is returned. When all candidate text sentences in the target text sentence are traversed, at least one associated text sentence is obtained; if there is no candidate text sentence after the associated text sentence in the preset text, the traversal of the target text sentence is stopped, and at least one associated text sentence is obtained. After determining the associated text sentence, the starting text sentence and at least one associated text sentence can be combined to obtain the text summary of the preset text. Thus, when generating a summary of a long text, through the relevance between each sentence in the text, sentences with high relevance are selected as the summary of the text, effectively ensuring the coherence of the text summary, and the sentences are extracted based on the original text, enabling the generated summary to effectively express the meaning of the original text and ensuring the matching degree between the summary and the original text.
[0096] In addition, during the process of determining the associated text sentence, the current text summary composed of the starting text sentence and the associated text sentence can be obtained, and the first text length of the current text summary can be obtained. Then, the preset text threshold and the text length of the preset text are used to determine the second text length. If the text length is greater than or equal to the second text length, the traversal of the target text sentence is stopped, and the current text summary is used as the target text summary of the preset text. Thus, when the text length of the determined text summary reaches the corresponding condition, the text summary is output, effectively ensuring the length of the text summary, avoiding the text length of the text summary from being too long, and improving the conciseness of the text summary.
[0097] Among them, the preset text threshold can be a percentage. Through this percentage and the text length of the preset text, it can be used to limit the text length of the text summary. For example, the preset text threshold can be 10%. After obtaining the text length of the preset text, multiplying it by 10% can obtain the upper limit of the text length of the text summary. During the process of determining the associated text sentence, if the text length of the current text summary composed of the starting text sentence and the associated text sentence has reached 10% of the text length of the preset text, the screening of the associated text sentence is stopped, and the current text summary is directly output as the final summary to ensure the text length of the text summary and improve the conciseness of the text summary.
[0098] In the above example, after candidate text sentences a, c, d, and e are filtered from the preset text according to the relevance scores, the generation of the text summary can be performed based on the sentence scores of each candidate text sentence and the corresponding relevance scores. Assuming that the starting text sentence is candidate text sentence a, the target text sentences can include candidate text sentences c, d, and e. Then, according to the relevance score of candidate text sentence a, the associated text sentence associated with the starting text sentence can be determined as candidate text sentence c (which can be determined according to the high and low of the relevance scores. When there are the same situations, one can be selected). Then, it can be determined whether there are candidate text sentences after candidate text sentence c in the preset text. There are still candidate text sentences d and e after candidate text sentence c. Then, according to the relevance scores between candidate text sentences d, e and candidate text sentence c, the next associated text sentence can be selected. If it is candidate text sentence e, at this time candidate text sentence e is already the end of the sentence, that is, it indicates that the traversal of the target text sentences is completed. Then, it is determined that the associated text sentences include candidate text sentence c and candidate text sentence e. Then, the text summary of the preset text is combined according to the arrangement order of the sentences in the preset text, and the text summary of the preset text is generated (that is, candidate text sentence a + candidate text sentence c + candidate text sentence e). Thus, when generating a summary of a long text, through the relevance between each sentence in the text, the sentences with high relevance are filtered out as the summary of the text, effectively ensuring the coherence of the text summary, and the sentences are extracted based on the original text, so that the generated summary can effectively express the meaning of the original text, ensuring the matching degree between the summary and the original text.
[0099] It should be noted that the embodiments of the present invention include but are not limited to the above examples. It can be understood that under the guidance of the idea of the embodiments of the present invention, those skilled in the art can also set according to actual needs, and the present invention does not limit this.
[0100] In the embodiments of the present invention, the preset text can be a long text with a text length greater than or equal to a certain threshold. Then, in the process of generating the text summary of the long text, the original text sentences of the preset text can be obtained first, and then the relevance scores between each original text sentence and other text sentences in the preset text are determined. According to the relevance scores of the original text sentences, the candidate text sentences corresponding to each original text sentence are filtered out. Then, the text summary of the preset text can be generated according to the candidate text sentences. Thus, when generating a summary of a long text, through the relevance between each sentence in the text, the sentences with high relevance are filtered out as the summary of the text, effectively ensuring the coherence of the text summary, and the sentences are extracted based on the original text, so that the generated summary can effectively express the meaning of the original text, ensuring the matching degree between the summary and the original text.
[0101] To enable those skilled in the art to better understand the technical solutions of the embodiments of the present invention, an exemplary illustration is given below through an example:
[0102] Assume that the target article is a long text. According to the sentence order, it can be split into 5 sentences, including a, b, c, d, and e. Then, each sentence can be input into the sentence relevance model to obtain the relevance scores between each sentence and other sentences, as follows:
[0103] For sentence a, the relevance scores include A1, A2, A3, A4, and the average value is A';
[0104] For sentence b, the relevance scores include B1, B2, B3, B4, and the average value is B';
[0105] For sentence c, the relevance scores include C1, C2, C3, C4, and the average value is C';
[0106] For sentence d, the relevance scores include D1, D2, D3, D4, and the average value is D';
[0107] For sentence e, the relevance scores include E1, E2, E3, E4, and the average value is E'.
[0108] Next, according to the relevance scores of each sentence, the candidate sentences of each sentence can be screened out. Specifically, if the score threshold is set to 0.9, the sentences with relevance scores greater than or equal to 0.9 can be used as candidate sentences. For example, the average values of sentences a, c, d, and e are greater than or equal to 0.9, so these sentences are used as candidate sentences, and sentence b is filtered out.
[0109] In the process of generating a text summary, the average scores of each sentence can be compared to determine the starting sentence. For example, if the average value A' of sentence a is the highest, then sentence a can be used as the starting sentence, and sentences c, d, and e can be used as target sentences to screen out relevant sentences associated with the starting sentence from the target sentences. Among them, the candidate sentences of sentence a include sentences c and e, then the one with a high correlation score can be selected from sentences c and e as the associated sentence of sentence a, assumed to be sentence c. At this time, taking sentence c as a node, the next sentence is determined from the candidate sentences of sentence c. The candidate sentences of sentence c include sentences d and e, then the sentence with a higher correlation score can be used as the next sentence. For example, if the correlation score corresponding to sentence e is higher than the correlation score corresponding to sentence d, then sentence e is selected as the associated sentence of sentence c, and taking sentence e as a new node, continue to traverse. Since in the above example, sentence e is the end of the sentence, the traversal stops, and sentences a, c, and e are combined to generate the text summary of the target article. Thus, when generating a text summary, through the correlation between each statement in the text, sentences with high correlation are screened out as the summary of the text, effectively ensuring the coherence of the text summary, and the sentences are extracted based on the original text, so that the generated summary can effectively express the meaning of the original text, ensuring the matching degree between the summary and the original text.
[0110] It should be noted that for the method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of the present invention are not limited by the described action sequence, because according to the embodiments of the present invention, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions involved are not necessarily essential for the embodiments of the present invention.
[0111] Referring to Figure 2 , a structural block diagram of a text summary generation device provided in an embodiment of the present invention is shown, which may specifically include the following modules:
[0112] A preset text acquisition module 201, configured to acquire a preset text, where the preset text includes a plurality of original text sentences;
[0113] A correlation determination module 202, configured to determine the correlation scores between each of the original text sentences and other text sentences;
[0114] A candidate text sentence determination module 203, configured to extract candidate text sentences from each of the original text sentences according to the correlation scores;
[0115] A text summary generation module 204, configured to generate a target text summary corresponding to the preset text according to the candidate text sentences.
[0116] In an alternative embodiment, the correlation determination module 202 is specifically configured to:
[0117] Input the original text sentence into a sentence correlation model to obtain the correlation scores between the original text sentence and other text sentences in the preset text.
[0118] In an alternative embodiment, the candidate text sentence determination module 203 includes:
[0119] A sentence score generation sub-module, configured to generate a sentence score for the original text sentence by using the respective correlation scores of the original text sentence;
[0120] A candidate text sentence determination sub-module, configured to use other text sentences whose correlation scores corresponding to the original text sentence are greater than or equal to a preset score threshold as the candidate text sentences corresponding to the original text sentence.
[0121] In an alternative embodiment, the text summary generation module 204 includes:
[0122] A text sentence determination sub-module, configured to determine a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding respective correlation scores;
[0123] A text summary generation module 204, configured to generate a target text summary corresponding to the preset text by using the starting text sentence and the at least one associated text sentence.
[0124] In an alternative embodiment, the text sentence determination sub-module includes:
[0125] A starting text sentence determination unit, configured to use the candidate text sentence with the highest sentence score in the preset text as the starting text sentence;
[0126] A target text sentence determination unit, configured to use the candidate text sentences in the preset text that are after the starting text sentence as the target text sentences;
[0127] An associated text sentence determination unit, configured to determine at least one associated text sentence according to the correlation score of the starting text sentence, and the correlation scores between each of the target text sentences and the target text sentences.
[0128] In an alternative embodiment, the associated text sentence determination unit is specifically configured to:
[0129] Use the candidate text sentence with the highest correlation score corresponding to the starting text sentence as the associated text sentence associated with the starting text sentence;
[0130] Determine whether there is a candidate text sentence after the associated text sentence in the preset text;
[0131] If there is a candidate text sentence after the associated text sentence in the preset text, then use the candidate text sentence with the highest relevance score corresponding to the associated text sentence as the new associated text sentence, and return to the step of determining whether there is a candidate text sentence after the associated text sentence;
[0132] When all candidate text sentences in the target text sentence are traversed, at least one associated text sentence is obtained.
[0133] In an alternative embodiment, the associated text sentence determination unit is further specifically configured to:
[0134] If there is no candidate text sentence after the associated text sentence in the preset text, then stop traversing the target text sentence and obtain at least one associated text sentence.
[0135] In an alternative embodiment, the text sentence determination sub-module further includes:
[0136] A first text length determination unit, configured to obtain a current text summary formed by the starting text sentence and the associated text sentence, and obtain a first text length of the current text summary;
[0137] A second text length determination unit, configured to determine a second text length by using a preset text threshold and the text length of the preset text;
[0138] A text summary generation unit, configured to, if the text length is greater than or equal to the second text length, then stop traversing the target text sentence, and use the current text summary as the target text summary of the preset text.
[0139] For the device embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and for the related parts, refer to the partial description of the method embodiment.
[0140] In addition, an embodiment of the present invention further provides an electronic device, as Figure 3 shown, including a processor 301, a communication interface 302, a memory 303, and a communication bus 304, wherein the processor 301, the communication interface 302, and the memory 303 communicate with each other through the communication bus 304,
[0141] The memory 303 is used to store a computer program;
[0142] The processor 301 is configured to, when executing the program stored in the memory 303, implement the following steps:
[0143] Obtain a preset text, where the preset text includes a number of original text sentences;
[0144] Determine the correlation scores between each of the original text sentences and other text sentences;
[0145] Extract candidate text sentences from each of the original text sentences according to the correlation scores;
[0146] Generate a target text summary corresponding to the preset text according to the candidate text sentences.
[0147] In an alternative embodiment, the determining the correlation scores between each of the original text sentences and other text sentences includes:
[0148] Input the original text sentence into a sentence correlation model to obtain the correlation scores between the original text sentence and other text sentences in the preset text.
[0149] In an alternative embodiment, the extracting candidate text sentences from each of the original text sentences according to the correlation scores includes:
[0150] Use the respective correlation scores of the original text sentence to generate a sentence score for the original text sentence;
[0151] Take the original text sentences in the preset text whose sentence scores are greater than or equal to a preset score threshold as candidate text sentences of the preset text.
[0152] In an alternative embodiment, the generating a target text summary corresponding to the preset text according to the candidate text sentences includes:
[0153] Determine a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding respective correlation scores;
[0154] Use the starting text sentence and the at least one associated text sentence to generate a target text summary corresponding to the preset text.
[0155] In an alternative embodiment, the determining a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding respective correlation scores includes:
[0156] Take the candidate text sentence with the largest sentence score in the preset text as the starting text sentence;
[0157] Take the candidate text sentences in the preset text that are after the starting text sentence as target text sentences;
[0158] Determine at least one associated text sentence according to the relevance score of the starting text sentence and the relevance scores of each of the target text sentences with the target text sentence.
[0159] In an alternative embodiment, the determining at least one associated text sentence according to the relevance score of the starting text sentence and the relevance scores of each of the target text sentences with the target text sentence includes:
[0160] Use the candidate text sentence with the highest relevance score corresponding to the starting text sentence as the associated text sentence associated with the starting text sentence;
[0161] Determine whether there is a candidate text sentence after the associated text sentence in the preset text;
[0162] If there is a candidate text sentence after the associated text sentence in the preset text, use the candidate text sentence with the highest relevance score corresponding to the associated text sentence as the new associated text sentence, and return to the step of determining whether there is a candidate text sentence after the associated text sentence;
[0163] When all candidate text sentences in the target text sentence are traversed, obtain at least one associated text sentence.
[0164] In an alternative embodiment, after determining whether there is a candidate text sentence after the associated text sentence in the preset text, the method further includes:
[0165] If there is no candidate text sentence after the associated text sentence in the preset text, stop traversing the target text sentence and obtain at least one associated text sentence.
[0166] In an alternative embodiment, before obtaining at least one associated text sentence when all candidate text sentences in the target text sentence are traversed, the method further includes:
[0167] Obtain the current text summary composed of the starting text sentence and the associated text sentence, and obtain the first text length of the current text summary;
[0168] Use a preset text threshold and the text length of the preset text to determine a second text length;
[0169] If the text length is greater than or equal to the second text length, stop traversing the target text sentence and use the current text summary as the target text summary of the preset text.
[0170] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus, an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For the sake of convenience of representation, only a thick line is used in the figure, but it does not mean that there is only one bus or one type of bus.
[0171] The communication interface is used for communication between the above terminal and other devices.
[0172] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.
[0173] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.
[0174] As Figure 4 shown, in another embodiment provided by the present invention, a computer-readable storage medium 401 is further provided. Instructions are stored in the computer-readable storage medium. When it runs on a computer, it enables the computer to execute the text summary generation method described in the above embodiment.
[0175] In another embodiment provided by the present invention, a computer program product containing instructions is further provided. When it runs on a computer, it enables the computer to execute the text summary generation method described in the above embodiment.
[0176] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present invention are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that can be accessed by a computer, or a data storage device such as a server or data center that includes one or more integrated available media. The available medium can be a magnetic medium (such as a floppy disk, hard disk, magnetic tape), an optical medium (such as a DVD), or a semiconductor medium (such as a solid state disk (SSD)).
[0177] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.
[0178] Each embodiment in this specification is described in a related manner. For the same or similar parts among the embodiments, reference can be made to each other. Each embodiment focuses on the differences from other embodiments. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and reference can be made to the corresponding part of the method embodiment for the related content.
[0179] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principle of the present invention are all included in the protection scope of the present invention.
Claims
1. A method for generating an abstract of a text, characterized in that, Including: Obtain a preset text, where the preset text includes several original text sentences; Determine the relevance scores between each of the original text sentences and other text sentences; Extract candidate text sentences from each of the original text sentences according to the relevance scores; Generate a target text summary corresponding to the preset text according to the candidate text sentences, including: Determine a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding relevance scores, including: taking the candidate text sentence with the largest sentence score in the preset text as the starting text sentence; taking the candidate text sentences located after the starting text sentence in the preset text as target text sentences; taking the target text sentence with the highest relevance score corresponding to the starting text sentence as the associated text sentence associated with the starting text sentence; judging whether there are candidate text sentences located after the associated text sentence in the preset text; if there are candidate text sentences after the associated text sentence in the preset text, taking the candidate text sentence with the highest relevance score corresponding to the associated text sentence as the new associated text sentence, and returning to the step of judging whether there are candidate text sentences located after the associated text sentence; when all candidate text sentences in the target text sentences are traversed, obtain at least one associated text sentence; Generate a target text summary corresponding to the preset text by using the starting text sentence and the at least one associated text sentence.
2. The method according to claim 1, wherein The determining the relevance scores between each of the original text sentences and other text sentences includes: Input the original text sentence into a sentence relevance model to obtain the relevance score between the original text sentence and other text sentences in the preset text.
3. The method according to claim 1, characterized in that, The extracting candidate text sentences from each of the original text sentences according to the relevance scores includes: Generate the sentence score of the original text sentence by using each relevance score of the original text sentence; Take the original text sentences in the preset text whose sentence scores are greater than or equal to a preset score threshold as the candidate text sentences of the preset text.
4. The method according to claim 1, characterized in that After judging whether there are candidate text sentences located after the associated text sentence in the preset text, the method further includes: If there are no candidate text sentences located after the associated text sentence in the preset text, stop traversing the target text sentences and obtain at least one associated text sentence.
5. The method according to claim 1, wherein Before the step of obtaining at least one associated text sentence when all candidate text sentences in the target text sentences are traversed, the method further includes: Obtain the current text summary composed of the starting text sentence and the associated text sentence, and obtain the first text length of the current text summary; Determine a second text length by using a preset text threshold and the text length of the preset text; If the text length is greater than or equal to the second text length, stop traversing the target text sentences and take the current text summary as the target text summary of the preset text.
6. An apparatus for generating an abstract of a text, characterized in that Including: A preset text acquisition module, configured to acquire a preset text, where the preset text includes several original text sentences; A relevance determination module for determining the relevance scores between each of the original text sentences and other text sentences; A candidate text sentence determination module for extracting candidate text sentences from each of the original text sentences according to the relevance scores; A text summary generation module for generating a target text summary corresponding to the preset text according to the candidate text sentences, including: A text sentence determination sub-module for determining a starting text sentence and at least one associated text sentence according to the sentence scores of each of the candidate text sentences and the corresponding relevance scores, including: a starting text sentence determination unit for using the candidate text sentence with the largest sentence score in the preset text as the starting text sentence; a target text sentence determination unit for using the candidate text sentences located after the starting text sentence in the preset text as target text sentences; an associated text sentence determination unit for using the target text sentence with the highest relevance score corresponding to the starting text sentence as the associated text sentence associated with the starting text sentence; determining whether there are candidate text sentences in the preset text that are located after the associated text sentence; if there are candidate text sentences after the associated text sentence in the preset text, using the candidate text sentence with the highest relevance score corresponding to the associated text sentence as the new associated text sentence, and returning to the step of determining whether there are candidate text sentences after the associated text sentence; when all candidate text sentences in the target text sentences are traversed, obtaining at least one associated text sentence A text summary generation module for generating a target text summary corresponding to the preset text by using the starting text sentence and the at least one associated text sentence.
7. The device according to claim 6, characterized in that, The candidate text sentence determination module includes: A sentence score generation sub-module for generating the sentence score of the original text sentence by using the relevance scores of the original text sentence; A candidate text sentence determination sub-module for using the other text sentences whose relevance scores corresponding to the original text sentence are greater than or equal to a preset score threshold as the candidate text sentences corresponding to the original text sentence.
8. An electronic device, characterized in that, It includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete communication with each other through the communication bus; The memory is used for storing a computer program; When the processor is used to execute the program stored on the memory, it realizes the method according to any one of claims 1-5.
9. One or more computer-readable media, on which instructions are stored, and when executed by one or more processors, cause the processors to execute the method according to any one of claims 1-5.
Citation Information
Patent Citations
Automatic text summarization method, device and electronic device
CN109101489A