A method, apparatus, electronic device, and storage medium for generating a text summary

By preprocessing the target text and sorting the eigenvector weights, combining the BERT and T5-PEGASUS models, a more coherent and accurate text summary is generated, solving the problem of poor sentence coherence in the prior art.

CN114138936BActive Publication Date: 2025-05-30PERFECT WORLD HLDG GRP CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202111456066.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2021-12-01
Publication Date
2025-05-30
Estimated Expiration
2041-12-01

AI Technical Summary

Technical Problem

In the prior art, the text summary algorithm based on Word2vec+TextRank has poor sentence coherence, and the deep neural network-based model lacks coding ability for long text, resulting in the generated summary being inaccurate and coherent enough.

Method used

By preprocessing the target text, generating sentence feature vectors, and sorting them based on the weight of the sentence feature vectors, selecting sentences that meet the preset length range to generate text summary, and using BERT and T5-PEGASUS models to generate and polish text summary.

Benefits of technology

On the premise of ensuring the completeness of the topic content, a more smooth and easy-to-read abstract was generated, solving the problem of poor sentence coherence in the prior art and improving the accuracy of the abstract.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114138936B_ABST
    Figure CN114138936B_ABST
Patent Text Reader

Abstract

The present application provides a method, device, electronic device and storage medium for generating a text summary, wherein the method comprises: preprocessing an input target text to obtain a plurality of first sentences; generating a sentence feature vector corresponding to the first sentence when the plurality of first sentences meet a preset condition, and determining a weight corresponding to the sentence feature vector; sorting the plurality of first sentences according to the weight, and selecting a plurality of second sentences that meet a preset length range from the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted from size according to the weight; and selecting a preset number of third sentences that are ranked top based on the plurality of second sentences to generate a text summary of the target text. The above scheme solves the problem that the sentences generated by the text summary generation method have poor coherence, resulting in the generated summary being inaccurate.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of data processing, and particularly to a method, apparatus, electronic device, and storage medium for generating a text summary. Background Art

[0002] With the rapid development of Internet information, the text information on the network has shown an explosive growth. In order to enable users to quickly and efficiently obtain the information they are interested in, it is necessary to compress long text information into short texts with concise content, which can help users save a large amount of time costs.

[0003] Currently, there are two ways to extract text summaries:

[0004] Way 1) Implement news text summary extraction based on word2vec + TextRank, that is, extractive text summary. Implement text summary extraction based on a graph, such as the TextRank algorithm. Take sentences as vertices, and the similarity between sentences as the weight of the edge. Determine the key sentences according to the weight scores of the vertices.

[0005] Way 2) Implement generative news text summary based on a deep neural network model, that is, generative text summary. Use a neural network model to train a summary generation model, take (summary, source text) data pairs as training data, and then implement text summary generation.

[0006] However, for the Word2vc + Textrank algorithm in the above Way 1, the text representation model word2vec has a simple network structure and weak feature extraction ability. For the Textrank algorithm, it only depends on sentence similarity, has a large amount of calculation, and this algorithm extracts key sentences from the source text, so it will lead to problems such as poor sentence coherence, difficult to control the number of words, and unclear main idea of the target sentence. That is to say, the quality of its summary depends on the original text. For the above Way 2, implementing text summary generation based on a neural network, currently the maximum encoding length of existing deep learning models is only 1024. For long texts, it is impossible to extract effective features from overly long sentences. Therefore, the semantic vectors generated by encoding will lose a large amount of information, resulting in inaccurate generated summaries. Summary of the Invention

[0007] The purpose of the embodiments of the present application is to provide a method, apparatus, electronic device, and storage medium for generating a text summary, which solves the problem that the sentences generated by the existing methods for generating text summaries have poor coherence, resulting in inaccurate generated summaries. The specific technical solutions are as follows:

[0008] In the first aspect of the implementation of this application, a method for generating a text summary is first provided, including: preprocessing the input target text to obtain a plurality of first sentences; generating a sentence feature vector corresponding to the first sentence and determining the weight corresponding to the sentence feature vector when the plurality of first sentences meet a preset condition; sorting the plurality of first sentences according to the weight and selecting a plurality of second sentences that meet a preset length range from the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted according to the weight from large to small; and generating a text summary of the target text based on a preset number of third sentences with higher rankings selected from the plurality of second sentences.

[0009] Optionally, the preprocessing the input target text to obtain a plurality of first sentences includes: splitting the target text into sentences to obtain a plurality of fourth sentences; performing word segmentation on the plurality of fourth sentences; and removing the first target words and target characters in the plurality of fourth sentences after word segmentation to obtain the plurality of first sentences.

[0010] Optionally, the generating a sentence feature vector corresponding to the first sentence includes: performing word segmentation processing on the first sentence to obtain a plurality of word segments corresponding to the first sentence; and performing word vectorization on the plurality of word segments and averaging the sum of the word vectors obtained after word vectorization to obtain the sentence feature vector.

[0011] Optionally, the sentence feature vector includes a title feature vector and a non-title feature vector corresponding to the title in the target text; the determining the weight corresponding to the sentence feature vector includes: determining the similarity between the title feature vector and the non-title feature vector and determining a first weight based on the similarity; determining the position of the corresponding sentence in the target text based on the non-title feature vector and determining a second weight based on the position; determining a third weight based on whether the second target word is included in the sentence corresponding to the non-title feature vector; determining a fourth weight based on the coverage rate of the third target word in the sentence corresponding to the non-title feature vector; and determining the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight.

[0012] Optionally, determining the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight includes: calculating a first product result of the first weight and a first coefficient; calculating a second product result of the second weight and a second coefficient; calculating a third product result of the third weight and a third coefficient; calculating a fourth product result of the fourth weight and a fourth coefficient; and determining the sum of the first product result, the second product result, the third product result, and the fourth product result as the weight corresponding to the sentence feature vector; wherein the sum of the first coefficient, the second coefficient, the third coefficient, and the fourth coefficient is 1.

[0013] Optionally, generating a text summary of the target text based on a preset number of third sentences selected from among a plurality of second sentences and having a relatively high ranking includes: selecting a fifth sentence that meets a preset length from among the preset number of third sentences; performing redundancy processing on the selected fifth sentence; and generating a text summary of the target text based on the fifth sentence after redundancy processing.

[0014] Optionally, the method further includes: when the plurality of first sentences do not meet the preset condition, generating a text summary of the target text based on the plurality of first sentences that do not meet the preset condition.

[0015] Optionally, the preset condition refers to the number of sentences being less than or equal to a first preset threshold, or the preset condition refers to the total length of the text being less than or equal to a second preset threshold.

[0016] In a second aspect of the implementation of the present application, there is also provided a device for generating a text summary, including: a first processing module configured to perform preprocessing on an input target text to obtain a plurality of first sentences; a second processing module configured to generate a sentence feature vector corresponding to the first sentence and determine the weight corresponding to the sentence feature vector when the plurality of first sentences meet a preset condition; a third processing module configured to sort the plurality of first sentences according to the weight and select a plurality of second sentences that meet a preset length range from among the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted in descending order of weight; and a first generation module configured to generate a text summary of the target text based on a preset number of third sentences selected from among the plurality of second sentences and having a relatively high ranking.

[0017] Optionally, the first processing module includes: a sentence splitting unit configured to split the target text into a plurality of fourth sentences; a first word segmentation unit configured to perform word segmentation on the plurality of fourth sentences; and a removal unit configured to remove a first target word and target characters from the plurality of fourth sentences after word segmentation to obtain the plurality of first sentences.

[0018] Optionally, the second processing module includes: a second word segmentation unit for performing word segmentation on the first sentence to obtain a plurality of word segments corresponding to the first sentence; and a first processing unit for performing word vectorization on the plurality of word segments, and averaging the sum of the word vectors obtained after word vectorization to obtain the sentence feature vector.

[0019] Optionally, the sentence feature vector includes a title feature vector and a non-title feature vector corresponding to the title in the target text; the second processing module includes: a first determination unit for determining the similarity between the title feature vector and the non-title feature vector, and determining a first weight based on the similarity; a second determination unit for determining the position of the corresponding sentence in the target text based on the non-title feature vector, and determining a second weight based on the position; a third determination unit for determining a third weight based on whether the second target word is included in the sentence corresponding to the non-title feature vector; a fourth determination unit for determining a fourth weight based on the coverage rate of the third target word in the sentence corresponding to the non-title feature vector; and a fifth determination unit for determining the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight.

[0020] Optionally, the fifth determination unit includes: a first calculation subunit for calculating a first product result of the first weight and a first coefficient; a second calculation subunit for calculating a second product result of the second weight and a second coefficient; a third calculation subunit for calculating a third product result of the third weight and a third coefficient; a fourth calculation subunit for calculating a fourth product result of the fourth weight and a fourth coefficient; and a first determination subunit for determining the sum of the first product result, the second product result, the third product result, and the fourth product result as the weight corresponding to the sentence feature vector; wherein the sum of the first coefficient, the second coefficient, the third coefficient, and the fourth coefficient is 1.

[0021] Optionally, the first generation module includes: a selection unit for selecting a fifth sentence that meets a preset length from a preset number of the third sentences; a second processing unit for performing redundancy processing on the selected fifth sentence; and a generation unit for generating a text summary of the target text based on the fifth sentence after redundancy processing.

[0022] Optionally, the apparatus further includes: a second generation module for generating a text summary of the target text based on the plurality of first sentences that do not meet the preset condition when the plurality of first sentences do not meet the preset condition.

[0023] Optionally, the preset condition refers to that the number of sentences is less than or equal to a first preset threshold, or the preset condition refers to that the total length of the text is less than or equal to a second preset threshold.

[0024] In the third aspect of the implementation of the present application, there is also provided an electronic device, including a processor, a communication interface, a memory and a communication bus, wherein the processor, the communication interface and the memory communicate with each other through the communication bus; the memory is used to store computer programs; the processor is used to implement the method steps described in the first aspect when executing the program stored in the memory.

[0025] In a fourth aspect of the implementation of the present application, a computer-readable storage medium is further provided, wherein the computer-readable storage medium stores instructions, which, when executed on a computer, enable the computer to execute the method described in the first aspect.

[0026] In an embodiment of the present application, for a target text to be generated as a text summary, it can be preprocessed to obtain the corresponding multiple first sentences, and then the multiple first sentences are generated into corresponding multiple sentence feature vectors, and the weights corresponding to the sentence feature vectors are determined, and then sorted based on the weights, and multiple second sentences are selected from the sorted multiple first sentences, and a preset number of third sentences with top rankings are further selected from the multiple second sentences to generate a text summary of the target text. That is to say, in an embodiment of the present application, corresponding sentence feature vectors are generated for sentences in the target text, and sentences that meet a preset length range are selected according to the weights of the sentence feature vectors, thereby realizing the extraction of important sentences in the target text, so that a more fluent and readable summary text is generated while ensuring the completeness of the target subject content, thereby solving the problem that the sentences generated by the text summary generation method in the prior art have poor coherence, resulting in the generated summary being inaccurate. BRIEF DESCRIPTION OF THE DRAWINGS

[0027] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, the drawings required for use in the embodiments or the description of the prior art are briefly introduced below.

[0028] Figure 1 A schematic diagram of a method for generating a Chinese text summary according to an embodiment of the present application;

[0029] Figure 2 A flowchart of a method for implementing end-to-end extraction of long news information based on a joint model of an improved graph model and deep learning according to an embodiment of the present application;

[0030] Figure 3 It is a structural schematic diagram of a device for generating a Chinese text summary according to an embodiment of the present application;

[0031] Figure 4is a schematic structural diagram of a device for generating a Chinese text abstract according to another embodiment of the present application;

[0032] Figure 5 Schematic diagram of the structure of an electronic device according to an embodiment of the present application. DETAILED DESCRIPTION

[0033] In order to make the purpose, technical solution and advantages of the embodiments of the present application clearer, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application.

[0034] In the following description, the suffixes such as "module" and "unit" used to represent elements are only used to facilitate the description of the present application, and they themselves have no specific meaning. Therefore, "module" and "component" can be used interchangeably.

[0035] The following will describe the technical solution in the embodiment of the present application in conjunction with the accompanying drawings in the embodiment of the present application. The embodiment of the present application provides a method for generating a text summary, such as Figure 1 As shown, the method comprises the following steps:

[0036] Step 102, preprocessing the input target text to obtain a plurality of first sentences;

[0037] In the example of the embodiment of the present application, the target text may be a news release, paper, document, etc. with a relatively large amount of text content. In addition, the language of the text content in the target text may be Chinese, English, or other types of languages.

[0038] Step 104, when the plurality of first sentences satisfy the preset condition, generating a sentence feature vector corresponding to the first sentence, and determining a weight corresponding to the sentence feature vector;

[0039] In the embodiment of the present application, the weight is used to characterize the importance of the first sentence corresponding to the sentence feature vector in the target text, that is, the larger the weight value is, the more suitable it is for generating a text summary of the target text.

[0040] Step 106, sorting the plurality of first sentences according to the weights, and selecting a plurality of second sentences that meet a preset length range from the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted in size according to the weights; and

[0041] Step 108 : Generate a text summary of the target text based on selecting a preset number of third sentences ranked top among the plurality of second sentences.

[0042] Through the above steps 102 to 108, for the target text for which a text summary is to be generated, it is possible to preprocess it to obtain a plurality of corresponding first sentences, and then generate corresponding sentence feature vectors for the plurality of first sentences, and determine the weights corresponding to the sentence feature vectors. Then, sort based on the weights, and select a plurality of second sentences from the sorted plurality of first sentences. Further, select a preset number of third sentences with a higher ranking from the plurality of second sentences to generate the text summary of the target text. That is to say, in the embodiments of the present application, sentence feature vectors corresponding to the sentences in the target text are generated, and sentences that meet the preset length range are selected according to the weights of the sentence feature vectors, realizing the extraction of important sentences in the target text, so that a more smooth and readable summary text is generated on the premise of ensuring the completeness of the target theme content, thus solving the problem that the sentences generated by the existing text summary generation method have poor coherence and the generated summary is not accurate enough.

[0043] In an alternative implementation manner of the embodiments of the present application, for the method of preprocessing the input target text in step 102 above to obtain a plurality of first sentences, it may further include:

[0044] Step 11, perform sentence splitting on the target text to obtain a plurality of fourth sentences;

[0045] Step 12, perform word segmentation on the plurality of fourth sentences; and

[0046] Step 13, remove the first target words and target characters in the plurality of fourth sentences after word segmentation to obtain a plurality of first sentences.

[0047] For the above steps 11 to 13, in a specific example, the target text is segmented into multiple fourth sentences using punctuation marks, such as [。?!?!;;.] as delimiters. If the target text is in Chinese, a word segmentation tool can be used to segment each sentence of the text, and then the first target word in the segmented sentence, such as a stop word, can be removed. When dealing with it specifically, the Chinese stop word list can be referred to, and the target characters in the segmented sentence, such as punctuation marks, emoticons (such as [◆|◇|●|■|▼|▲|★|▍]), etc., can be removed. If the target text is in English, each sentence can be segmented using spaces, and the first target word in the sentence, such as a stop word, can be removed. When dealing with it specifically, the English stop word list can be referred to, and the target characters in the segmented sentence, such as punctuation marks, emoticons (such as [◆|◇|●|■|▼|▲|★|▍]), etc., can be removed. It should be noted that stop words refer to some words that do not contain much information and have no meaning for the entire text, especially some meaningless high-frequency words. The existence of these words greatly affects the weight in text summary extraction. The stop word list can be preset or supplemented through actual processing. After removing the first target word and target characters, the processed word segmentation can be stored in a list.

[0048] By segmenting and word-segmenting the target text, and removing the first target word and target characters, the text content in the target text that is unimportant for abstract generation can be filtered, improving the accuracy of subsequent text abstract generation.

[0049] In another optional implementation manner of the embodiment of the present application, for the method of generating a sentence feature vector corresponding to the first sentence involved in the above step 104, it may further include:

[0050] Step 21, performing word segmentation processing on the first sentence to obtain multiple word segments corresponding to the first sentence; and

[0051] Step 22, performing word vectorization on the multiple word segments, and averaging the multiple word vectors obtained after word vectorization to obtain a sentence feature vector.

[0052] In the example of the embodiment of the present application, word vectorization of sentence tokenization can be performed based on BERT (Bidirectional Encoder Representation from Transformers, a deep bidirectional pre-trained transformer). Essentially, BERT is based on a bidirectional Transformer encoder, and the Transformer performs more prominently in capturing the bidirectional relationships of statements. This model incorporates more syntactic and semantic information in the BERT word vectors, can model the text to express more text semantic features, can also better solve the problem of polysemy, and at the same time, BERT is trained with words as units, which can overcome the out-of-vocabulary problem faced by Word2Vec to a certain extent. Finally, it can improve the accuracy of the weights of the TextRank algorithm and improve the quality of the abstract.

[0053] In another alternative implementation manner of the embodiment of the present application, for the sentence feature vectors involved in the embodiment of the present application, including the title feature vector and the non-title feature vector corresponding to the title in the target text. Based on this, for the method of determining the weight corresponding to the sentence feature vector involved in the above step 104, it can further include:

[0054] Step 31, determine the similarity between the title feature vector and the non-title feature vector, and determine the first weight based on the similarity;

[0055] Among them, the title feature vector refers to the feature vector corresponding to the title in the target text, and the non-title feature vector refers to the feature vectors corresponding to each sentence after sentence splitting and tokenization in the body of the target text. Taking the target text as a current affairs news as an example, since the title of the real-time news usually reflects the core content of a real-time news, the higher the similarity between the sentence in the real-time news and the title, the more it indicates that the sentence can represent the main idea of the news content, and the more important it is, then a higher weight should be given to this sentence. In a specific example, the weight calculation formula can be:

[0056]

[0057] Among them, cos(s topic , s j ) represents the cosine similarity between sentence s j and title s tocpic .

[0058] Step 32, determine the position of the corresponding sentence in the target text based on the non-title feature vector, and determine the second weight based on the position;

[0059] Among them, the importance of sentences in different positions in the target text is relatively different. Generally speaking, the importance of sentences at the beginning and end of a paragraph is relatively high. Taking the target text as a news text as an example, a news text usually makes a general statement about the news in the first few sentences at the beginning, and may make a summary-like generalization in the last one or two sentences. Therefore, increasing the weights of the first sentence, the sentences in the front, and the summary sentences can make the position information be noticed. Usually, the probability of selecting the first sentence of the text as the abstract is 0.85, and the probability of selecting the last sentence as the abstract is 0.07. The specific weight values can be set accordingly according to the actual situation, and the above is only an example. In the embodiments of the present application, the weight of a sentence can be adjusted according to the position of the sentence. The greater the weight increase is given to the earlier sentences in the target text, and the smaller the weight increase is given to the later sentences in the last paragraph of the text. The position weight calculation formula is as follows:

[0060]

[0061] Among them, s is the number of sentences selected at the beginning of the article, num is the total number of sentences in the article, r is the number of sentences selected at the end of the article, e 1 and e 2 are the weight adjustment thresholds for the first few sentences and the last few sentences respectively. In the embodiments of the present application, the values of e 1 and e 2 can be set according to the characteristics of long news texts. For example, e 1 = 0.3, e 2 = 0.2.

[0062] Step 33, determine the third weight based on whether the sentence corresponding to the non-title feature vector includes the second target word;

[0063] Among them, the second target word in the embodiments of the present application refers to the relatively important words in the target text, such as summary words, key keywords, etc. In a specific example, taking the target text as a news article as an example, in a news article, there will be summary words such as "in short", "ultimately", "generally speaking", "footing", "key point", "In short", "finally", etc. These special words (i.e., the second target words) usually summarize the main content of the news or article. Therefore, the sentences including such special words are more important, and higher weights can be set for such sentences. The sentence special word weight calculation formula is expressed as:

[0064]

[0065] Among them, s pw is the sentence in the target text.

[0066] Step 34: Determine the fourth weight based on the coverage rate of the third target word in the sentence corresponding to the non-title feature vector.

[0067] In this regard, in a specific example, taking the target text as a news article, the main content of the news article can generally be reflected by some keywords. The more keywords appear in a sentence, the higher the importance of the sentence. First, a keyword list can be obtained in the preprocessing stage, and the keyword coverage rate of each sentence in the article can be calculated using this keyword list. If more keywords appear in a sentence, its keyword coverage rate is higher, that is, a higher weight is assigned to this sentence. The calculation of the keyword coverage rate weight is as follows:

[0068] W k,j = K wc (S j )

[0069] K wc (S j ) = len(keywords(s j )) / len(s j ), 0 ≤ K wc (S j ) ≤ 1

[0070] Among them, K wc (S j ) represents the coverage rate of keywords in sentence S j , len(keywords(sj)) represents the number of keywords included in this sentence, and len(s j ) represents the total number of words in this sentence after word segmentation; it should be noted that the list of sentences after word segmentation has removed irrelevant words such as stop words and punctuation marks.

[0071] Step 35: Determine the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight.

[0072] In the above steps 31 to 35, since the similarity between the sentence and the title, the position of the sentence, whether the sentence includes special words, and the keyword coverage rate in the sentence can all represent the importance of the sentence in the target text to different degrees, therefore, based on the similarity between the sentence and the title, the position of the sentence, whether the sentence includes special words, and the keyword coverage rate in the sentence to determine the weight of the sentence can comprehensively and intuitively express the importance of the sentence in the target text, so that subsequent sorting of sentences based on the weight and selecting sentences from the sorting results to generate a summary can be more accurate.

[0073] In an alternative implementation manner of the embodiment of the present application, for the method of determining the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight involved in step 35 above, it may further include:

[0074] Step 41, calculate the first product result of the first weight and the first coefficient;

[0075] Step 42, calculate the second product result of the second weight and the second coefficient;

[0076] Step 43, calculate the third product result of the third weight and the third coefficient;

[0077] Step 44, calculate the fourth product result of the fourth weight and the fourth coefficient; and

[0078] Step 45, determine the sum of the first product result, the second product result, the third product result, and the fourth product result as the weight corresponding to the sentence feature vector, where the sum of the first coefficient, the second coefficient, the third coefficient, and the fourth coefficient is 1.

[0079] For steps 41 to 45 above, the importance of different features for abstract extraction is different. Therefore, weight coefficients are introduced for each part of the features to balance the proportion of the weight influence factors of each part. The importance of each feature can be dynamically adjusted according to the characteristics of different long texts. For example, the target text in the embodiment of the present application is a long news text, and its weights can be downgraded according to W topic,j >W pos,j >W k,j >W s,j >W spw,j to construct the final sentence weight calculation formula as follows:

[0080] W final =λ s W s,j +λ pos W pos,j +λ topic W topic,j +λ spw W spw,j +λ k W k,j

[0081] λ s +λ pos +λ topic +λ spw +λ k =1

[0082] Among them, λ is the weight coefficient of the weight influence factors of each part (including the above-mentioned first coefficient, second coefficient, third coefficient, and fourth coefficient), and its value ranges from 0 to 1. According to the analysis of specific experimental data, λ is assigned; in addition, W spw,j is the weight corresponding to the sentence position, W spw,j is the weight indicating whether the sentence includes special words, W k,j is the weight corresponding to the word coverage rate, W s,j is the weight corresponding to the similarity between sentences, W final is the final weight value. The above λ is only an example, and the specific value of λ can be set accordingly according to different target texts. By setting different weight coefficients according to different situations, the weight of the sentence can be obtained more accurately, and the importance of the sentence in the target text can be better reflected.

[0083] In another optional implementation manner of the embodiment of the present application, for the method of generating the text summary of the target text by selecting a preset number of third sentences with a relatively high ranking from the multiple second sentences involved in the above step 108, it may further include:

[0084] Step 51, select a fifth sentence that meets the preset length from the preset number of third sentences;

[0085] Step 52, perform redundancy processing on the selected fifth sentence; and

[0086] Step 53, generate the text summary of the target text based on the fifth sentence after redundancy processing.

[0087] For the above steps 51 to 53, since the weights are determined according to the similarity between texts, the finally generated summary sentences may be very similar, resulting in redundant summary information. Therefore, in the embodiment of the present application, in order to further improve the accuracy of the generated summary, redundancy processing can be performed through the MMR (Maximal Marginal Relevance) algorithm. Through the MMR algorithm, not only the relevance between the sentence and the key information of the current text set can be considered, but also the difference between sentences can be taken into account, so that the selected summary sentences before and after show the difference between information while ensuring the relevance between the sentence and the central content of the text set, thereby controlling the redundancy of the output result. Therefore, when extracting a single sentence as the summary, no redundancy processing is performed. When extracting multiple sentences, the sorting result is combined with the MMR algorithm for redundancy processing to obtain a summary containing more key information.

[0088] The MMR algorithm controls the redundancy of the text summary by introducing a penalty factor a, which can take 0.75 in the specific example of this application. It introduces the difference between sentences into the calculation, re-scores the sentences, and finally extracts the final summary according to the scores. The formula is as follows:

[0089] Score(MMR)=a*score(sen i )+(1-a)*sim(sen i ,sen i -1)

[0090] Among them, sen i is the sorted sentence. The first sorted sentence does not need to be calculated and is used as the first sentence of the extracted summary. Taking the first sentence as a reference, the similarity penalty is performed with the subsequent sentences in turn. When the similarity between sentences is large, they are selected in the order in which these sentences appear in the candidate abstract sentences, and the remaining sentences are regarded as redundant sentences and removed from the candidate abstract sentence group. Finally, the text in the front of the sorting is extracted as the summary, and the extracted summary is returned in the original text order.

[0091] In addition, in the embodiments of this application, the method of the embodiments of this application may further include:

[0092] Step 110, when multiple first sentences do not meet the preset conditions, generate a text summary of the target text based on the multiple first sentences that do not meet the preset conditions; where the preset condition means that the number of sentences is less than or equal to the first preset threshold, or the preset condition means that the total length of the text is less than or equal to the second preset threshold.

[0093] That is to say, the preset condition means whether the number of sentences is less than the preset threshold, or whether the total text length is less than the preset threshold. For example, whether the number of sentences is less than 2, or whether the total text of multiple sentences is less than 6; of course, the above 2 and 6 are only examples and can be adjusted and set according to the actual situation.

[0094] The following is an example of this application in combination with the specific implementation manners of this application. The specific implementation manner provides a method for realizing end-to-end extraction of long news information based on a joint model of an improved graph model and deep learning, as Figure 2 shown. The steps of this method include:

[0095] Step 201, input the text;

[0096] Step 202, determine the language of the data text. When the judgment result indicates that the text language is English, execute Step 210. When the judgment result indicates that the text language is Chinese, execute Step 203;

[0097] Step 203, Chinese data preprocessing;

[0098] Among them, the Chinese data preprocessing includes: sentence splitting, word segmentation, stop word removal, and removal of special symbols such as emoticons;

[0099] Step 204, determine whether the number of sentences is less than or equal to 2, or the total length of the processed text is less than or equal to 6; if the judgment result is yes, execute Step 205, and if the judgment result is no, execute Step 206;

[0100] Step 205, return the source text, and then execute Step 208;

[0101] Step 206, return the list of all preprocessed segmented text;

[0102] Step 207, BERT+STLLP_TextRank is used to implement text extraction;

[0103] Among them, when constructing the TextRank network graph, the nodes are changed from text sentences to sentence vectors generated by the pre-trained model BERT. BERT is essentially a bidirectional Transformer encoder, and Transformer performs more prominently in capturing the bidirectional relationships of statements. This model incorporates more syntactic and semantic information into the BERT word vectors, can model the text to express more text semantic features, can also better solve the problem of polysemy, and at the same time BERT is trained with characters as units, which can overcome the out-of-vocabulary problem faced by Word2Vec to a certain extent. Finally, it can improve the accuracy of the TextRank algorithm weights and improve the summary quality.

[0104] Step 208, the T5-PEGASUS model is used to train the text summary generation model;

[0105] Among them, the T5-PEGASUS model is based on mT5 as the basic architecture and initial weights, and uses PEGASUS-style pseudo-summary pre-training on Chinese corpora, that is, draws on the idea of PEGASUS to construct the corpus, and finally has good text generation performance, especially excellent few-shot learning ability, which can effectively solve problems such as insufficient training sets and high annotation costs. In addition, T5 is a text-to-text transfer Transformer model, which uses a standard encoder-decoder model. The pre-training of T5 includes both supervised and unsupervised parts, and the training objective is similar to that of BERT, except that it is changed to the Seq2Seq version. For the supervised part, common NLP supervised task data is collected and uniformly converted into seq2seq tasks for training.

[0106] PEGASUS is a pre-trained model customized for abstracts. It can be used as a general generative pre-training task. PEGASUS is a standard Transformer (prefix encoder-decoder predictor) with both an encoder and a decoder. The pre-training objectives include GSG (Gap Sentences Generation) and MLM (Masked Language Model). For the original three sentences, one sentence is entirely masked by [MASK1] and used as the target text to be generated. The other two sentences have some tokens randomly masked by [MASK2] and are used as inputs.

[0107] For the specific implementation of T5-PEGASUS, assume a document has n sentences. We select approximately n / 4 sentences (which can be non-consecutive) from them, such that the text formed by these n / 4 sentences and the text formed by the remaining 3n / 4 sentences have the longest possible common subsequence. Then we regard the text formed by the 3n / 4 sentences as the original text and the text formed by the n / 4 sentences as the abstract, thus forming a pseudo-abstract data pair of "(original text, abstract)". We can then use these data pairs to train the Seq2Seq model. If there are no duplicate sentences in the document, the sentences in the original text and the abstract will have no intersection. Therefore, the generation task is not a simple copy of the original text. That is, after polishing the previously extracted text abstract, a better reading effect can be achieved.

[0108] Step 209: Output the generated Chinese text abstract.

[0109] Step 210: Preprocess the English data.

[0110] Among them, the preprocessing of the English data includes: sentence splitting, word tokenization, stop word removal, and removal of special symbols such as emojis.

[0111] Step 211: Determine whether the number of sentences is less than or equal to 2, or the total length of the processed text is less than or equal to 6. If the judgment result is yes, execute Step 212; if the judgment result is no, execute Step 213.

[0112] Step 212: Return the source text and then execute Step 207.

[0113] Step 213: Train the text abstract generation model using the English T5 model.

[0114] Among them, the principle of generating Chinese abstracts with the T5 pre-training model is the same as above. During the feature extraction process, the T5 model needs to set the length of abstract generation (such as 200) for training. However, when calling the model for abstract generation, the length of the generated abstract is determined according to the proportion of the article. If the abstract length specified by the model training is smaller than the actual length of the abstract that needs to be generated, for example, the abstract length that needs to be generated according to the proportion of the article is 300-400, and the shortest abstract length that needs to be generated is 300, but the model can only generate up to 200, then irrelevant characters will appear for completion, which will cause repeated characters and irrelevant characters to appear at the end of the generated English abstract. Therefore, in the model feature extraction stage, it is necessary to process each annotated training set, that is, add iconic ending words such as "End of Token" after the target abstract, so that in the subsequent model prediction, the irrelevant text after the "End of Token" terminator will be truncated, which can effectively avoid the above problems.

[0115] Step 214: output the generated English text summary.

[0116] It can be seen that through the above steps 201 to 214, firstly, based on the BERT+STLP-TextRank summary extraction algorithm model, the long news text summary extraction is preliminarily realized, which can effectively solve the problems that the subsequent model cannot encode too long text and the summary generation theme is incomplete due to the simple use of the model. In addition, for Chinese, the extracted summary is polished based on the T5-PEGASUS pre-trained model to realize the generation of Chinese news text summary. For English, the extracted summary is polished based on the T5 pre-trained model to realize the generation of English news text summary, so that a more fluent and easy-to-read summary text is generated while ensuring the completeness of the subject content.

[0117] The present application embodiment provides a device for generating a text summary, such as Figure 3 As shown, including:

[0118] A first processing module 32, configured to pre-process the input target text to obtain a plurality of first sentences;

[0119] A second processing module 34 is used to generate a sentence feature vector corresponding to the first sentence and determine a weight corresponding to the sentence feature vector when the plurality of first sentences meet a preset condition;

[0120] A third processing module 36 is used to sort the plurality of first sentences according to the weights, and select a plurality of second sentences that meet a preset length range from the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted in size according to the weights; and

[0121] A first generation module 38, configured to generate a text summary of the target text based on a preset number of third sentences selected from multiple second sentences and ranked higher.

[0122] Optionally, the first processing module 32 in the embodiments of the present application may further include: a sentence splitting unit, configured to split the target text into multiple fourth sentences; a first word segmentation unit, configured to perform word segmentation on the multiple fourth sentences; and a removal unit, configured to remove a first target word and target characters in the multiple word-segmented fourth sentences to obtain multiple first sentences.

[0123] Optionally, the second processing module 34 in the embodiments of the present application may further include: a second word segmentation unit, configured to perform word segmentation on the first sentence to obtain multiple word segments corresponding to the first sentence; and a first processing unit, configured to perform word vectorization on the multiple word segments, and calculate the average after adding the multiple word vectors obtained by word vectorization to obtain a sentence feature vector.

[0124] Optionally, the sentence feature vector in the embodiments of the present application includes a title feature vector and a non-title feature vector corresponding to the title in the target text; based on this, the second processing module 34 may further include: a first determination unit, configured to determine the similarity between the title feature vector and the non-title feature vector, and determine a first weight based on the similarity; a second determination unit, configured to determine the position of the corresponding sentence in the target text based on the non-title feature vector, and determine a second weight based on the position; a third determination unit, configured to determine a third weight based on whether the second target word is included in the sentence corresponding to the non-title feature vector; a fourth determination unit, configured to determine a fourth weight based on the coverage rate of the third target word in the sentence corresponding to the non-title feature vector; and a fifth determination unit, configured to determine the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight.

[0125] Optionally, the fifth determination unit in the embodiments of the present application may further include: a first calculation subunit, configured to calculate a first product result of the first weight and a first coefficient; a second calculation subunit, configured to calculate a second product result of the second weight and a second coefficient; a third calculation subunit, configured to calculate a third product result of the third weight and a third coefficient; a fourth calculation subunit, configured to calculate a fourth product result of the fourth weight and a fourth coefficient; and a first determination subunit, configured to determine the sum of the first product result, the second product result, the third product result, and the fourth product result as the weight corresponding to the sentence feature vector; where the sum of the first coefficient, the second coefficient, the third coefficient, and the fourth coefficient is 1.

[0126] Optionally, the first generation module 38 in the embodiments of the present application may further include: a selection unit, configured to select a fifth sentence that meets a preset length from a preset number of third sentences; a second processing unit, configured to perform redundancy processing on the selected fifth sentence; and a generation unit, configured to generate a text summary of the target text based on the fifth sentence after the redundancy processing.

[0127] As Figure 4 shown, the device in the embodiments of the present application may further include:

[0128] A second generation module 42, configured to generate a text summary of the target text based on a plurality of first sentences that do not meet a preset condition when the plurality of first sentences do not meet the preset condition.

[0129] Optionally, the preset condition in the embodiments of the present application refers to that the number of sentences is less than or equal to a first preset threshold, or the preset condition refers to that the total length of the text is less than or equal to a second preset threshold.

[0130] The embodiments of the present application further provide an electronic device, as Figure 5 shown, including a processor 501, a communication interface 502, a memory 503, and a communication bus 504, where the processor 501, the communication interface 502, and the memory 503 communicate with each other through the communication bus 504.

[0131] The memory 503 is used to store a computer program.

[0132] When the processor 501 is configured to execute the program stored on the memory 503, it implements the Figure 1 method steps, and the function it plays is the same as the Figure 1 method steps.

[0133] The communication bus mentioned in the above terminal may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. This communication bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 5 only a thick line is shown in

[0134] but it does not mean that there is only one bus or one type of bus.

[0135] The memory may include a Random Access Memory (RAM), or may also include a non-volatile memory, such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0136] The aforementioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0137] In another embodiment provided by the present application, there is also provided a computer-readable storage medium, in which instructions are stored, and when it runs on a computer, it causes the computer to execute the method for generating a text summary described in any one of the above embodiments.

[0138] In another embodiment provided by the present application, there is also provided a computer program product containing instructions, and when it runs on a computer, it causes the computer to execute the method for generating a text summary described in any one of the above embodiments.

[0139] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer may be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions may be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions may be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.) means. The computer-readable storage medium may be any available medium that can be accessed by a computer or a data storage device such as a server or data center that includes one or more integrated available media. The available medium may be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid state disk (SSD)).

[0140] It should be noted that, in this document, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including", or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article, or device that includes a series of elements includes not only those elements but also other elements that are not explicitly listed, or also includes elements that are inherent to such process, method, article, or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article, or device that includes the element.

[0141] Each embodiment in this specification is described in a related manner. The same or similar parts among the embodiments can be referred to each other, and the differences between each embodiment and other embodiments are emphasized. In particular, for the system embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the description of the method embodiment.

[0142] The above are only the preferred embodiments of the present application and are not intended to limit the protection scope of the present application. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application are all included in the protection scope of the present application.

Claims

1. A method for generating text summaries, It is characterized in that include: Preprocessing the input target text to obtain multiple first sentences; In the case where the plurality of first sentences meet a preset condition, a sentence feature vector corresponding to the first sentence is generated, and a weight corresponding to the sentence feature vector is determined; wherein the preset condition refers to that the number of sentences is less than or equal to a first preset threshold, or the preset condition refers to that the total length of the text is less than or equal to a second preset threshold; Sorting the plurality of first sentences according to the weights, and selecting a plurality of second sentences that meet a preset length range from the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted in size according to the weights; and Generating a text summary of the target text based on selecting a preset number of third sentences ranked top from the plurality of second sentences; The generating of the sentence feature vector corresponding to the first sentence includes: performing word segmentation processing on the first sentence to obtain a plurality of word segments corresponding to the first sentence; and performing word vectorization on the plurality of word segments, and adding and averaging the plurality of word vectors obtained after the word vectorization to obtain the sentence feature vector; Among them, the sentence feature vector includes a title feature vector and a non-title feature vector corresponding to the title in the target text; the determining of the weight corresponding to the sentence feature vector includes: determining the similarity between the title feature vector and the non-title feature vector, and determining a first weight based on the similarity; determining the position of the corresponding sentence in the target text based on the non-title feature vector, and determining a second weight based on the position; determining a third weight based on whether the sentence corresponding to the non-title feature vector includes a second target word; determining a fourth weight based on the coverage rate of the third target word in the sentence corresponding to the non-title feature vector; and determining the weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight and the fourth weight.

2. The method according to claim 1, It is characterized in that The input target text is preprocessed to obtain a plurality of first sentences, including: Sentence the target text to obtain a plurality of fourth sentences; performing word segmentation on the plurality of fourth sentences; and The first target words and target characters in the plurality of fourth sentences after word segmentation are removed to obtain the plurality of first sentences.

3. The method according to claim 1, It is characterized in that Determining a weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight includes: Calculate a first product result of the first weight and a first coefficient; Calculate a second product of the second weight and the second coefficient; Calculating a third product of the third weight and the third coefficient; Calculating a fourth product of the fourth weight and the fourth coefficient; and The sum of the first product result, the second product result, the third product result and the fourth product result is determined as the weight corresponding to the sentence feature vector; wherein the sum of the first coefficient, the second coefficient, the third coefficient and the fourth coefficient is 1.

4. The method according to claim 1, It is characterized in that The step of generating a text summary of the target text based on selecting a preset number of third sentences ranked top from the plurality of second sentences comprises: Selecting a fifth sentence satisfying a preset length from a preset number of the third sentences; Performing redundancy processing on the selected fifth sentence; and Based on the fifth sentence after redundancy processing, a text summary of the target text is generated.

5. The method according to claim 1, It is characterized in that The method further comprises: In a case where the plurality of first sentences do not satisfy the preset condition, a text summary of the target text is generated based on the plurality of first sentences that do not satisfy the preset condition.

6. A device for generating a text summary, It is characterized in that include: A first processing module, used for preprocessing the input target text to obtain a plurality of first sentences; A second processing module is used to generate a sentence feature vector corresponding to the first sentence and determine a weight corresponding to the sentence feature vector when the plurality of first sentences meet a preset condition; wherein the preset condition refers to that the number of sentences is less than or equal to a first preset threshold, or the preset condition refers to that the total length of the text is less than or equal to a second preset threshold; A third processing module is used to sort the plurality of first sentences according to the weights, and select a plurality of second sentences that meet a preset length range from the sorted plurality of first sentences; wherein the sorted plurality of first sentences are sorted in size according to the weights; and A first generating module, configured to generate a text summary of the target text based on selecting a preset number of third sentences ranked top from the plurality of second sentences; The second processing module includes: a second word segmentation unit, which is used to perform word segmentation processing on the first sentence to obtain multiple word segments corresponding to the first sentence; and a first processing unit, which is used to perform word vectorization on the multiple word segments, and add and average the multiple word vectors obtained after the word vectorization to obtain a sentence feature vector; Wherein, the sentence feature vector includes a title feature vector corresponding to the title and a non-title feature vector in the target text, and the second processing module includes: a first determination unit configured to determine a similarity between the title feature vector and the non-title feature vector, and determine a first weight based on the similarity; a second determination unit configured to determine a position of a corresponding sentence in the target text based on the non-title feature vector, and determine a second weight based on the position; a third determination unit configured to determine a third weight based on whether a second target word is included in the sentence corresponding to the non-title feature vector; a fourth determination unit configured to determine a fourth weight based on a coverage rate of a third target word in the sentence corresponding to the non-title feature vector; and a fifth determination unit configured to determine a weight corresponding to the sentence feature vector based on the first weight, the second weight, the third weight, and the fourth weight.

7. An electronic device, characterized in that it includes a processor, a communication interface, a memory, and a communication bus. Among them, the processor, the communication interface, and the memory complete mutual communication through the communication bus; the memory is used for storing a computer program; the processor, when executing the program stored on the memory, implements the method for generating a text summary according to any one of claims 1-5.

8. A computer-readable storage medium, on which a computer program is stored, characterized in that when the program is executed by a processor, it implements the method for generating a text summary according to any one of claims 1-5.

Citation Information

Patent Citations

  • Abstraction generation method and device, terminal equipment and storage medium

    CN110837556A

  • Multi-language multi-document abstract extraction method based on weighted TextRank

    CN112948543A