A text summarization generation method and computer device based on discrimination of abstraction levels

Through the abstraction discrimination model, the abstraction degree is identified, combined with abstract extraction and generation model, the problem of low abstract accuracy and efficiency in the existing methods is solved, and efficient and accurate text abstract generation is achieved.

CN114996443BActive Publication Date: 2025-07-08BEIJING ZHONGKE ZHIJIA TECH CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210590856.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-05-27
Publication Date
2025-07-08
Estimated Expiration
2042-05-27

AI Technical Summary

Technical Problem

The existing text summary generation methods lack targeting, resulting in low accuracy and efficiency of generated abstractions, and the inability to effectively process texts of different degrees of abstraction.

Method used

The abstract degree discrimination model is introduced, and the abstract degree of text is identified through the pre-trained model, which is divided into three levels: weak abstraction, medium abstraction, and strong abstraction. Different abstract generation methods are used respectively, combined with abstract extraction and generation models, and abstract generation is generated using multi-classification and multi-channel methods.

Benefits of technology

It improves the accuracy and efficiency of abstract generation, solves the problem of unlogged words and inadequate summary, makes full use of text characteristics, and reduces resource occupation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN114996443B_ABST
    Figure CN114996443B_ABST
Patent Text Reader

Abstract

The present invention relates to a method for generating text abstracts based on discrimination of abstraction levels and a computer device, belonging to the technical field of natural language processing. The method for generating text abstracts of the present invention includes the following steps: discriminating the abstraction level of the text to be abstracted through a pre-trained abstraction level discrimination model to obtain the abstraction level label of the text to be abstracted; predicting the extraction probability of each sentence in the text to be abstracted through a pre-trained abstract extraction model based on the clause sequence of the text to be abstracted; and performing abstract extraction according to the abstraction level label of the text to be abstracted and the extraction probability of each sentence in the text to be abstracted to obtain the abstract of the text to be abstracted. It solves the problem that the existing text abstract generation methods lack pertinence to the writing characteristics and abstraction levels of texts, and generate abstracts without discrimination, resulting in low accuracy and efficiency of the generated abstracts.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly to a method for generating a text summary based on discrimination of the degree of abstraction and a computer device. Background Art

[0002] In recent years, data mainly in the form of text has flooded people's lives through microblogs, news, web pages, etc. In order to quickly obtain the important content of the text in a short time, text summary technology has emerged as the times require.

[0003] Text summary technology belongs to the field of Natural Language Processing (NLP). It can, through a series of algorithms or training models, achieve the purpose of extracting important information in the text and automatically generating a summary. Text summary technology greatly saves the time people spend reading text and greatly improves reading efficiency, and has a wide range of applications in the fields of finance, justice, military industry, Internet, etc.

[0004] The existing text summary technologies can be divided into extractive, generative, and extractive-generative hybrid types. The basic principle of extractive text summarization is to select key sentences from the original text to form a summary. Since the integrity of the sentences is retained, this method has a low error rate in grammar and syntax, ensuring the accuracy and readability of the extraction results, but it also leads to the disadvantages of discontinuous summary content, redundancy, and insufficient refinement and summary; the representative algorithm is the TextRank method based on graph ranking. Generative summaries can achieve the automatic generation of sentences, do not completely inherit the original corpus, and allow new words or phrases to be included in the summary, with high flexibility. Currently, generative summaries generally adopt the sequence-to-sequence (seq2seq) model based on deep learning. The results of generative summaries have better performance than extractive models in a large number of cases, but they also introduce problems such as out-of-vocabulary words and repeated decoding. Hybrid text summarization first locks some key sentences through the extractive method, and then uses these key sentences as the input of the generative model to finally obtain the generated summary. This method inherits both the advantages and disadvantages of the extractive and generative models, is more suitable for long text summaries, but requires a large amount of high-quality manually annotated corpus, and has certain requirements for resources and computing power. Summary of the Invention

[0005] In view of the above analysis, the present invention aims to provide a method for generating a text summary based on discrimination of the degree of abstraction and a computer device; it solves the problem that the existing text summary generation methods have no pertinence to the writing characteristics and degree of abstraction of the text, and generate summaries without discrimination, resulting in low accuracy and efficiency of the generated summaries.

[0006] The object of the present invention is mainly achieved through the following technical solutions:

[0007] On the one hand, the present invention provides a method for generating a text summary based on the discrimination of the degree of abstraction, including the following steps:

[0008] Discriminate the degree of abstraction level of the text to be summarized through a pre-trained degree of abstraction discrimination model to obtain the degree of abstraction label of the text to be summarized;

[0009] Based on the sentence sequence of the text to be summarized, predict the extraction probability of each sentence in the text to be summarized through a pre-trained summary extraction model;

[0010] According to the degree of abstraction label of the text to be summarized and the extraction probability of each sentence in the text to be summarized, perform summary extraction to obtain the summary of the text to be summarized.

[0011] Furthermore, the training of the degree of abstraction discrimination model includes:

[0012] Input a training sample set with text degree of abstraction annotation labels, and predict the probability distribution of the degree of abstraction labels of the samples in the training sample set;

[0013] Through iterative update of the loss between the probability distribution of the degree of abstraction labels and the degree of abstraction annotation labels of the corresponding text, train to obtain the degree of abstraction discrimination model.

[0014] Furthermore, the degree of abstraction discrimination model includes an encoder, a convolutional layer, a pooling layer, and a linear layer;

[0015] The encoder uses the Bert pre-trained model; according to the token sequence of the input text, through word embedding and context encoding, obtain the encoder hidden vector with the context representation of the input text; the input text includes the samples in the training sample set or the text to be summarized;

[0016] The convolutional layer is used to learn a set of word vectors with the length of the vocabulary according to the encoder hidden vector;

[0017] The pooling layer is used to compress the data of the word vector set of the convolutional layer in the word dimension to obtain the document-level vector representation;

[0018] The linear layer is used to reduce the dimension of the document-level vector representation;

[0019] The output after dimensionality reduction is converted through an activation function to obtain the probability distribution of the degree of abstraction labels; the label with the highest probability in the probability distribution is set as the degree of abstraction label of the current text.

[0020] Furthermore, the training sample set is obtained through data preprocessing;

[0021] The training sample set includes Chinese texts and the original abstracts corresponding to the texts.

[0022] The data preprocessing includes data denoising, data annotation, and vocabulary building. The data annotation includes: annotating the abstraction level label of the Chinese text and the extraction or non-extraction label for each sentence in the Chinese text. The constructed vocabulary is used as the vocabulary of the word segmenter for the text summarization system.

[0023] Furthermore, determine whether the original abstract in the training sample set is a common subsequence of the corresponding Chinese text, and annotate the abstraction level label for the Chinese text: If the original abstract is a common subsequence of its corresponding text, determine that the abstraction level is low, and annotate the Chinese text as L0; if the abstract is a non-common subsequence of its corresponding text, determine that the abstraction level is medium, and annotate it as L1; if the abstract is not a common subsequence of its corresponding text, determine that the abstraction level is high, and annotate it as L2.

[0024] By traversing all combinations between clauses in the Chinese text, calculate the similarity between the combination and the original abstract of the corresponding text, and label the sentences included in the combination with the highest similarity as extraction labels, and the remaining sentences as non-extraction labels.

[0025] Furthermore, the abstract extraction model includes an encoder, a convolutional layer, a pooling layer, and a linear layer.

[0026] The encoder is used to obtain a set of word vectors through word embedding and context encoding based on the clause sequence of the text to be summarized. The set of word vectors is compressed by the pooling layer to obtain a set of sentence vectors. The set of sentence vectors is learned by the convolutional layer, and the hidden vector of the convolutional layer is output. The hidden vector of the convolutional layer is reduced in dimension and transformed by an activation function through the linear layer to obtain the extraction probability of each sentence in the text to be summarized.

[0027] Furthermore, according to the abstraction level label of the text to be summarized, a corresponding extraction probability threshold is preset; select the sentences that meet the extraction probability threshold as the extraction summary.

[0028] For the extraction summary corresponding to the text labeled as L0, directly use it as the final summary of the original text.

[0029] Input the extraction summaries corresponding to the texts labeled as L1 and L2 into the abstract optimization model for optimization to obtain the final summary.

[0030] Furthermore, the abstract optimization model includes an encoder and a decoder.

[0031] The encoder is used to receive the extraction summary of the Chinese text and obtain the encoder hidden vector through word embedding and position encoding.

[0032] The decoder is used to predict the probability distribution of words in the target summary text according to the encoder hidden vector by using the self-attention mechanism;

[0033] Based on the word with the highest probability at each time step in the probability distribution, the final summary is obtained.

[0034] Furthermore, the summary optimization model predicts the probability distribution of words by the following formula:

[0035]

[0036]

[0037]

[0038] where P(w) is the probability distribution of the predicted word at time step t;

[0039] p gen is the probability of generating a new word, (1 - p gen ) is the probability of copying a word in the input sequence. When the predicted abstraction level label L ab = L1, p gen = 0, and the model only selects words from the original extracted summary as the summary; when the predicted abstraction level label L ab = L2, p gen is obtained by the above formula, P vocab(w) is the probability distribution of words in the vocabulary predicted by T5; is a vector of the vocabulary length, where the value at the word position in the input sequence is the attention distribution of each word and the values at the remaining positions are 0; is the context vector of the encoder at time step t, a t is the attention distribution of the input word at time step t; s t is the decoder state at time step t, x t is the decoder input at time step t; w x , b ptr , V', V, b, b' are learnable parameters.

[0040] On the other hand, a computer device for text summary generation is also disclosed, which is characterized by including at least one processor and at least one memory communicatively connected to the processor;

[0041] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the foregoing text summary generation method.

[0042] Advantages of this technical solution:

[0043] 1. The present invention extracts text features through a pre-trained model, predicts the types of text summaries, divides the text to be summarized into three types, and adopts different summary extraction and generation methods according to different types. That is, a classifier is added on the basis of the summary extraction model and the summary optimization generation model, and text summary generation is carried out through a multi-classification and multi-channel method, which improves the efficiency and accuracy of summary generation.

[0044] 2. The present invention uses the vocabulary of the input text and the complete vocabulary for summary generation respectively according to the abstraction level of the text to be summarized, makes full use of the characteristics of the text itself, effectively reduces resource occupation, and improves the efficiency of text summary generation.

[0045] 3. The text summary generation method of the present invention identifies the writing characteristics of the text, locates the text abstraction level, and performs summary extraction and optimization generation according to the writing characteristics of the text. It makes full use of the advantages of the summary extraction model and the optimization generation model, and at the same time solves the problems of out-of-vocabulary words and incomplete summarization in the summary, improving the efficiency and accuracy of summary generation.

[0046] Other features and advantages of the present invention will be described in the following specification, and some of them will become obvious from the specification, or be understood by implementing the present invention. The objectives and other advantages of the present invention can be realized and obtained through the structures specifically pointed out in the written specification, claims, and drawings. Brief Description of the Drawings

[0047] The drawings are only for the purpose of showing specific embodiments, and are not considered as limiting the present invention. Throughout the drawings, the same reference signs denote the same components.

[0048] Figure 1 It is a flowchart of the text summary generation method according to an embodiment of the present invention;

[0049] Figure 2 It is a schematic structural diagram of the summary abstraction level discrimination model according to an embodiment of the present invention;

[0050] Figure 3 It is a schematic structural diagram of the summary extraction model according to an embodiment of the present invention;

[0051] Figure 4 It is a schematic diagram of summary extraction according to the abstraction level label and sentence extraction probability according to an embodiment of the present invention;

[0052] Figure 5 It is a schematic structural diagram of the summary optimization model according to an embodiment of the present invention;

[0053] Figure 6 It is a schematic diagram of the text summary generation method according to an embodiment of the present invention. Detailed Description of the Embodiments

[0054] The preferred embodiments of the present invention will be specifically described below in conjunction with the accompanying drawings. The accompanying drawings form a part of this application and are used together with the embodiments of the present invention to explain the principles of the present invention, rather than to limit the scope of the present invention.

[0055] A method for generating a text summary based on the discrimination of the degree of abstraction in this embodiment is as Figure 1 shown and includes the following steps:

[0056] Discriminate the degree of abstraction level of the text to be summarized through a pre-trained degree of abstraction discrimination model to obtain the degree of abstraction label of the text to be summarized;

[0057] Based on the sentence sequence of the text to be summarized, predict the extraction probability of each sentence in the text to be summarized through a pre-trained summary extraction model;

[0058] Perform summary extraction according to the degree of abstraction label of the text to be summarized and the extraction probability of each sentence in the text to be summarized to obtain the summary of the text to be summarized.

[0059] In practical applications, according to the writing characteristics of the text itself, the text can be divided into three levels according to the degree of abstraction: 1. Texts with weak abstraction, that is, the key sentences and central sentences in the text are clear, and the summary can be directly extracted from the original text; 2. Texts with medium abstraction, that is, there are sentences expressing the central idea in the text, but the summary needs to be further refined; 3. Texts with strong abstraction, that is, there are no central sentences or main idea sentences in the text, and the summary needs to be generated after understanding and summarizing the original text.

[0060] In the prior art, the automatic summary methods using a single extraction method or generation method, and the methods using a combination of extraction and generation, do not take into account the situations of the three degrees of abstraction and cannot be well applicable to the above three types of texts at the same time, resulting in low accuracy and efficiency of the summaries generated by the existing methods.

[0061] The method for generating a text summary based on the discrimination of the degree of abstraction disclosed in the present invention introduces a degree of abstraction discrimination model. Before generating the summary, it first identifies the writing characteristics of the text, locates the degree of abstraction of the text, and then uses different methods to generate the text summary according to the degree of abstraction. It solves the problems of out-of-vocabulary words and incomplete summarization in the summaries generated by the existing methods, and improves the efficiency and accuracy of summary generation.

[0062] As a specific embodiment, the method for generating a text summary of the present invention can be implemented through the following steps:

[0063] Step S1: Discriminate the degree of abstraction level of the text to be summarized through a pre-trained degree of abstraction discrimination model to obtain the degree of abstraction label of the text to be summarized;

[0064] Among them, the training of the abstract degree discrimination model includes:

[0065] Input a training sample set with text abstraction level labels, and predict the probability distribution of abstraction level labels of samples in the training sample set;

[0066] The abstract degree discrimination model is trained by iteratively updating the loss of the probability distribution of the abstract degree labels and the abstract degree annotation labels of the corresponding texts.

[0067] The training sample set used is obtained through data preprocessing; the training sample set includes Chinese text and the original summary corresponding to the text, and the Chinese text is annotated with text abstraction degree annotation labels and sentence extraction labels;

[0068] Specifically, first obtain a Chinese text summary dataset, wherein the Chinese text summary dataset includes Chinese text and the original summary corresponding to the text; exemplarily, the Chinese text summary dataset may include: short text datasets, such as LCSTS, THUCNews, etc.; long text datasets, such as nlpcc2017 summary data, Shen Ce Cup 2018 summary data, etc.

[0069] The samples in the Chinese text summary dataset are divided into a training sample set, a test sample set, and a validation sample set in proportion;

[0070] The Chinese text summary data set is preprocessed, including data denoising, data labeling and constructing a vocabulary; wherein, it is necessary to perform data denoising on all texts in the Chinese text summary data set, and only perform data labeling on the texts in the training sample set; data labeling includes: labeling the Chinese texts in the training sample set with abstraction levels, and labeling each sentence in the Chinese text with extraction or non-extraction labels; the constructed vocabulary is used as the word segmenter vocabulary of the model, that is, the text summarization method of this embodiment uses the same vocabulary constructed in the data preprocessing process as the word segmenter vocabulary when training with training samples and in actual applications after the training is completed.

[0071] Preferably, the following methods can be used for data denoising: delete data with missing content in the data set, delete special symbols and extra spaces in the data set, convert all full-width characters to half-width characters, convert all traditional Chinese characters to simplified Chinese characters, and convert all uppercase English letters to lowercase. The training sample set after data denoising is denoted as D.

[0072] This embodiment constructs a vocabulary in the following way: all the denoised texts in the data set are segmented, and all the words obtained after the segmentation constitute the vocabulary of the data set; the vocabulary of the existing Bert pre-training model and the vocabulary of the T5 pre-training model are merged into the vocabulary of the data set to expand the vocabulary of the data set to obtain the vocabulary of this embodiment, which is denoted as V Large .

[0073] Further, the data annotation includes:

[0074] Determine whether the abstract in the training sample set is a common subsequence of its corresponding Chinese text, and label the Chinese text with an abstraction level label: If the abstract is a common subsequence of its corresponding text, determine that the abstraction level is low, and label the Chinese text as L0; if the abstract is a non-common subsequence of its corresponding text, determine that the abstraction level is medium and label it as L1; if the abstract is not a common subsequence of its corresponding text, determine that the abstraction level is high and label it as L2;

[0075] In addition, by traversing all combinations between clauses in the Chinese text, calculate the similarity between the combination and the corresponding text abstract, and label the sentences included in the combination with the highest similarity as extraction labels, and the remaining sentences as non-extraction labels.

[0076] Preferably, the data annotation is performed in the following manner: First, truncate the overly long text, and only retain the first 3000 characters for texts with more than 3000 characters; then label the training samples, which are divided into two parts:

[0077] 1. For each text in the training sample, label 0, 1, 2, representing the labels of weak, medium, and strong abstraction levels, denoted as L0, L1, and L2 respectively.

[0078] Specifically, for any text, determine the common subsequence between the abstract and the original text. If the abstract is a continuous common subsequence of the original text, it indicates that the text has weak abstraction, the text abstraction level is at level 0, and the abstraction level label is L0; if the abstract is a non-continuous common subsequence of the original text, it indicates that the text has medium abstraction, the abstraction level is at level 1, and the abstraction level label is L1; if the abstract is not a subsequence of the original text, that is, some or all of the words in the abstract do not appear in the original text, it indicates that the text has strong abstraction, the abstraction level is at level 2, and the abstraction level label is L2.

[0079] Preferably, the abstraction level labels in this embodiment use the one-hot form as the input for model training, that is, L0 = [1, 0, 0], L1 = [0, 1, 0], L2 = [0, 0, 1]; as shown in Table 1.

[0080] Table 1 Text Abstraction Level Judgment Criteria and Examples

[0081]

[0082] 2. For each sentence in the training sample, label 1 or 0, representing the extraction label or the non-extraction label.

[0083] Specifically, for any text, first, the sentences are segmented according to punctuation marks, and the jieba toolkit is used for word segmentation. Traverse all combinations between sentences (including single sentences); calculate the similarity between each combination and the abstract, label the sentences included in the combination with the highest similarity as 1, that is, extract the label, and label the remaining sentences as 0, that is, do not draw the label; form a list of whether the sentences are extracted or not according to the original order of appearance, and use L sents to represent.

[0084] Among them, the calculation of similarity uses Rouge1:

[0085]

[0086] Among them, the target text is the combination between sentences.

[0087] Furthermore, the abstraction degree discrimination model includes an encoder, a convolutional layer, a pooling layer, and a linear layer; as Figure 2 shown, among them,

[0088] The encoder uses the Bert pre-trained model; according to the tokenized sequence of the input text, through word embedding and context encoding, an encoder hidden vector with the context representation of the input text is obtained; the input text includes samples in the training sample set or the text to be abstracted;

[0089] The convolutional layer is used to learn a set of word vectors with the length of the vocabulary according to the encoder hidden vector;

[0090] The pooling layer is used to compress the data of the word vector set of the convolutional layer in the word dimension to obtain a document-level vector representation;

[0091] The linear layer is used to reduce the dimension of the document-level vector representation;

[0092] The output after dimension reduction is transformed through an activation function to obtain an abstraction degree label probability distribution; the label with the highest probability in the probability distribution is set as the abstraction degree label of the current text.

[0093] Preferably, the abstraction degree discrimination model is trained by the following method:

[0094] The encoder uses the Bert pre-trained model. First, load the pre-trained weights of the Bert model, and set the Bert tokenizer vocabulary as the pre-constructed vocabulary V Large . The pre-trained weights are fixed and not updated during the encoding phase.

[0095] Take the original tokenized sequence of the training sample from the dataset D:

[0096] X = [x1, x2…x n , n is the number of words, as the input, and through word embedding and context encoding, a set of 768-dimensional word vectors is obtained

[0097] The convolutional layer is a 7-layer convolutional neural network. H te is input into the 7-layer convolutional neural network for learning. The convolutional kernel size of each layer of the convolutional neural network is 3, and the dilation rates are (1, 2, 4, 8, 4, 2, 1) respectively. Zero-padding is performed on each layer to ensure that the input and output dimensions are consistent. The seventh layer of the convolutional neural network outputs a set of 768-dimensional word vectors.

[0098] H tc After passing through the global average pooling layer to compress the data of the word dimension, a 768-dimensional document-level vector representation H is obtained. * , H * is reduced to three dimensions through a linear layer, and the sigmoid activation function is used:

[0099]

[0100] After conversion by the activation function, the probability distribution P of the abstraction level labels predicted by the model is obtained. ab ;

[0101] The probability distribution P of the predicted abstraction level labels ab is compared with the labeled abstraction level label L of the corresponding text (0,1,2) to calculate the cross-entropy loss:

[0102] Loss1 = CrossEntropy(P ab , L (0,1,2) );

[0103] Backpropagate the loss and update the weights of the convolutional layer and the linear layer by minimizing Loss1.

[0104] After iterative updates until the model converges and Loss1 no longer decreases, the abstraction level discrimination model training is completed.

[0105] Use the trained abstraction level discrimination model to predict the probability distribution P of the abstraction level labels of all texts in the dataset D ab , and change the maximum value in each P corresponding to each text ab to 1 and the rest to 0 to obtain the abstraction level labels L of all texts in D predicted by the model. ab .

[0106] Step S2: Based on the sentence sequence of the text to be summarized, use a pre-trained summary extraction model to predict the extraction probability of each sentence in the text to be summarized;

[0107] Specifically, the summary extraction model includes an encoder, a convolutional layer, a pooling layer, and a linear layer; as Figure 3As shown in the figure, an encoder is used to obtain a set of word vectors according to the clause sequence of the text to be summarized through word embedding and context encoding; the set of word vectors is compressed by a pooling layer to obtain a set of sentence vectors; the set of sentence vectors is learned by a convolutional layer, and the hidden vector of the convolutional layer is output; the hidden vector of the convolutional layer is reduced in dimension by a linear layer and transformed by an activation function to obtain the extraction probability of each sentence in the text to be summarized.

[0108] Preferably, the abstract extraction model is trained by the following method:

[0109] The structure of the abstract extraction model is similar to that of the abstract degree discrimination model, and the encoder uses the Bert pre-trained model; first, load the pre-trained weights of the Bert model, and set the Bert tokenizer vocabulary to the pre-constructed vocabulary V Large . The pre-trained weights are fixed and not updated during the encoding phase.

[0110] Take the original text clause sequence S = [s1, s2... s m (m is the number of sentences) from the training sample set D as the input, and obtain a set of 768-dimensional word vectors H through word embedding and context encoding * ;

[0111]

[0112] where represents the 768-dimensional word vector of the j-th word in the m-th sentence of S;

[0113] Compress the vectors of each sentence word dimension in H * to one dimension through global average pooling to obtain a set of 768-dimensional sentence vectors

[0114] H se Learn through a 7-layer convolutional neural network. The convolutional kernel size of each layer of the convolutional neural network is 3, and the dilation rates are (1, 2, 4, 8, 4, 2, 1). Zero padding is performed on each layer to ensure that the input and output dimensions are the same, and a set of 768-dimensional sentence vectors is obtained

[0115] H sc Reduce the 768 dimensions to 1 dimension through a linear layer, and use the sigmoid activation function to control the output range between 0 and 1:

[0116] Obtain the sentence extraction probability where respectively represent the extraction probabilities of sentences s1, s2... s m of,

[0117] The currently predicted sentence extraction probability Pext The extraction tag list L corresponding to the sentence annotation sents Calculate the cross-entropy loss:

[0118] Loss2 = CrossEntropy(L sents , P ext );

[0119] Backpropagate the loss and update the weights of the convolutional layer and linear layer by minimizing Loss2 to update the model weights.

[0120] Through iterative updates until the model converges and Loss2 no longer decreases, the extraction of the model is completed.

[0121] Use the trained summary extraction model to predict the sentence extraction probability P of all texts in D ext , and in subsequent processing, based on the abstraction level label L of the text ab and the extraction probability P ext perform summary extraction.

[0122] Step S3, according to the abstraction level label of the text to be summarized and the extraction probability of each sentence in the text to be summarized, perform summary extraction to obtain the summary of the text to be summarized

[0123] Specifically, according to the abstraction level label of the text to be summarized, preset the corresponding extraction probability threshold; select the sentences that meet the extraction probability threshold as the extraction summary;

[0124] For the extraction summary corresponding to the text labeled L0, directly use it as the final summary of the original text;

[0125] Input the extraction summaries corresponding to the texts labeled L1 and L2 into the summary optimization model for optimization to obtain the final summary.

[0126] Preferably, the process of summary extraction is as Figure 4 shown.

[0127] Set the following extraction probability thresholds according to the abstraction level label:

[0128]

[0129] According to the above extraction probability threshold p t and the abstraction level label, select the sentences corresponding to the values with a probability greater than p t in the sentence extraction probability as the extraction summary, denoted as X ext .

[0130] Such as Figure 4As shown in the figure, the extraction summary corresponding to the text with the abstraction level label L0 is directly used as the final summary of the original text; the extraction summaries corresponding to the texts with the abstraction level labels L1 and L2 are input into the summary optimization model for optimization to obtain the final summaries of the corresponding texts;

[0131] Specifically, the summary optimization model includes an encoder and a decoder, as Figure 5 shown:

[0132] The encoder is used to receive the extraction summary of the Chinese text and obtain the encoder hidden vector through word embedding and position encoding;

[0133] The decoder is used to predict the probability distribution of the words of the target summary text according to the encoder hidden vector by using the self-attention mechanism;

[0134] Based on the word with the highest probability at each time step in the probability distribution, the final summary is obtained.

[0135] Preferably, the summary optimizer is trained by the following method:

[0136] The encoder and decoder of the summary optimizer use the T5 pre-trained model and are trained in the way of pre-training + fine-tuning. First, load the T5 pre-trained weights, which can be updated during the training phase; set the T5 tokenizer vocabulary to the pre-constructed vocabulary V Large .

[0137] The encoder of T5 receives the text extraction summary X ext , and obtains the word vector set through word embedding and position encoding,

[0138] During decoding, adjust the probability distribution of predicting words at time step t according to the following formula:

[0139]

[0140]

[0141]

[0142] where P(w) is the probability distribution of predicting words at time step t; p gen is the probability of generating a new word; (1 - p gen ) is the probability of copying a word in the input sequence; P vocab(w) is the probability distribution of words in the vocabulary predicted by T5; is a vector of the length of the vocabulary, where the value at the position of the word in the input sequence is the attention distribution of each word and the values at the remaining positions are 0; is the context vector of the encoder at time step t, at is the attention distribution of the input word at time step t; s t is the decoder state at time step t, x t is the decoder input at time step t; w x , b ptr , V', V, b, b' are learnable parameters;

[0143] When the predicted abstraction level label L ab = L1, p gen = 0, and the model only selects words from the original extracted summary as the summary; when the predicted abstraction level label L ab = L2, p gen is obtained through the above formula; that is, the model will comprehensively consider the probability of copying words from the original extracted summary and the probability of generating new words for the final summary generation.

[0144] It should be noted that for text types with abstraction levels L1 and L2, the final summary is obtained by using the extraction plus generation method; according to the text type, for text with abstraction level L1, the final summary is directly generated using the vocabulary corresponding to the input text; for text with abstraction level L2, the final summary is generated using the complete pre - constructed vocabulary, which makes full use of the characteristics of the text itself, effectively reduces resource occupancy, and improves the efficiency and accuracy of summary generation.

[0145] Preferably, define the loss at each time step t and the total loss:

[0146]

[0147]

[0148] is the target word predicted at time step t, that is, the word in the original text summary at time step t, is the probability of the predicted target word, and T is the total number of time steps of the decoder.

[0149] Back - propagate the loss, and update T5 and T w w x , b ptr , V', V, b, b' weights.

[0150] After iterative update until the model converges and Loss T no longer decreases, the summary optimizer training is completed.

[0151] According to the abstract degree tags of the text to be summarized, the present invention uses the vocabulary of the input text and the complete vocabulary respectively to generate summaries, making full use of the characteristics of the text itself. For texts with weak abstraction, the extraction method is used, and for texts with general or strong abstraction, the copy mechanism is used to optimize the vocabulary, effectively reducing resource occupancy and improving the efficiency of text summary generation.

[0152] Another embodiment of the present invention also discloses a computer device for a text summary generation method, including at least one processor and at least one memory communicatively connected to the processor;

[0153] The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the foregoing text summary generation method.

[0154] Based on the abstract extraction model and the abstract optimization generation model, the present invention adds an abstract degree discrimination model and generates text summaries through a multi-classification and multi-channel method. The present invention extracts text features through a pre-trained model, predicts the text summary type, and classifies the text into two types: suitable for extraction (text with an abstract degree tag of L0) and suitable for extraction plus generation (text with abstract degree tags of L1 and L2). For the text type suitable for extraction plus generation, according to the text type, two summary generation methods using the input vocabulary and the complete vocabulary are adopted, making full use of the characteristics of the text itself, effectively reducing resource occupancy, and improving the summary generation efficiency and accuracy.

[0155] The foregoing text summary generation system is respectively tested and verified through the test set and the verification set in the dataset, and the test and verification results both show that the summary generated by the text summary generation system of the present invention has high accuracy and strong readability.

[0156] In summary, the present invention discloses a text summary generation method based on abstract degree discrimination, as Figure 6 shown. First, determine the text to be summarized, preprocess the text to be summarized, and remove invalid characters, spaces, etc.; add an abstract degree discrimination model on the basis of the extraction model and the generation model, and predict the probability distribution P of the abstract degree tag of the text to be summarized through the abstract degree discrimination model ab , change the maximum value in P ab to 1 and the other values to 0 to obtain the predicted abstract degree tag L of the text to be summarized ab , and divide the abstract degree into three levels: weak abstraction (level 0), medium abstraction (level 1), and strong abstraction (level 2). Different methods are used to generate text summaries according to the abstract degree of the text to be summarized. Specifically, first predict the sentence extraction probability of the text to be summarized, that is, P ext ; according to the abstract degree tag L of the text to be summarized aband the sentence extraction probability P of the text to be summarized ext Select several sentences from the text to be summarized as the extracted summary; determine whether the extracted summary needs to be optimized by the summary optimization model; for the extracted summary that does not need to be optimized, directly output it as the final summary; for the extracted summary that needs to be optimized, according to L ab Set the word generation probability p gen , and output the final summary after optimization by the summary optimization model; that is, use the extraction method for texts with weak abstractness, and use the copy mechanism to optimize the vocabulary for texts with general or strong abstractness. The text summary generation method of the present invention fully utilizes the advantages of the extraction model and the generation model by identifying the text writing characteristics and positioning the text abstraction level, and at the same time solves the problems of out-of-vocabulary words and incomplete summarization in the summary, improving the summary generation efficiency and accuracy.

[0157] Those skilled in the art can understand that all or part of the processes of implementing the methods in the above embodiments can be completed by instructing relevant hardware through a computer program, and the program can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium is a disk, an optical disk, a read-only memory or a random access memory, etc.

[0158] The above is only a preferred specific embodiment of the present invention, but the protection scope of the present invention is not limited thereto. Any changes or substitutions that can be easily thought of by those skilled in the art within the technical scope disclosed by the present invention should be covered by the protection scope of the present invention.

Claims

1. A text summarization generation method based on discrimination of abstraction level, characterized in that Including the following steps: Determine the abstraction level of the text to be summarized through a pre-trained abstraction level discrimination model, and obtain the abstraction level label of the text to be summarized; The training of the abstraction level discrimination model includes: inputting a training sample set with text abstraction level annotation labels, and predicting the probability distribution of the abstraction level labels of the samples in the training sample set; through loss iteration update of the probability distribution of the abstraction level labels and the abstraction level annotation labels of the corresponding text, training to obtain the abstraction level discrimination model; the training sample set is obtained through data preprocessing; the training sample set includes Chinese texts and the original summaries corresponding to the texts; the data preprocessing includes data denoising, data annotation, and building a vocabulary; the data annotation includes: annotating the abstraction level label of the Chinese text and the extraction or non-extraction label of each sentence in the Chinese text; the built vocabulary is used as the tokenizer vocabulary of the model; judge whether the original summary in the training sample set is a common subsequence of the Chinese text corresponding to it, and annotate the abstraction level of the Chinese text: if the original summary is a continuous common subsequence of its corresponding text, judge that the abstraction level is low, and annotate the Chinese text as L0; if the summary is a non-continuous common subsequence of its corresponding text, judge that the abstraction level is medium, and annotate it as L1; if the summary is not a common subsequence of its corresponding text, judge that the abstraction level is high, and annotate it as L2; by traversing all combinations between clauses in the Chinese text, calculate the similarity between the combination and the original summary of the corresponding text, and label the sentences included in the combination with the highest similarity as extraction labels, and the remaining sentences as non-extraction labels; Based on the clause sequence of the text to be summarized, predict the extraction probability of each sentence in the text to be summarized through a pre-trained summary extraction model; Perform summary extraction according to the abstraction level label of the text to be summarized and the extraction probability of each sentence in the text to be summarized, and obtain the summary of the text to be summarized.

2. The method for generating a text summary according to claim 1, wherein The abstraction level discrimination model includes an encoder, a convolutional layer, a pooling layer, and a linear layer; The encoder uses the Bert pre-trained model; according to the token sequence of the input text, through word embedding and context encoding, obtain the encoder hidden vector with the context representation of the input text; the input text includes samples in the training sample set or the text to be summarized; The convolutional layer is used to learn a set of word vectors with the length of the vocabulary according to the encoder hidden vector; The pooling layer is used to compress the data of the word vector set of the convolutional layer in the word dimension to obtain a document-level vector representation; The linear layer is used to reduce the dimension of the document-level vector representation; The output after dimension reduction is converted through an activation function to obtain the probability distribution of the abstraction level labels; the label with the highest probability in the probability distribution is set as the abstraction level label of the current text.

3. The method for generating a text abstract according to claim 1, characterized in that The summary extraction model includes an encoder, a convolutional layer, a pooling layer, and a linear layer; The encoder is used to obtain a set of word vectors through word embedding and context encoding according to the sequence of clauses of the text to be summarized; the set of word vectors is compressed by a vector through a pooling layer to obtain a set of sentence vectors; the set of sentence vectors is learned through a convolutional layer, and a convolutional layer hidden vector is output; the convolutional layer hidden vector is reduced in dimension through a linear layer and transformed by an activation function to obtain the extraction probability of each sentence in the text to be summarized.

4. The method for generating a text summary according to claim 3, wherein, According to the abstraction degree label of the text to be summarized, a corresponding extraction probability threshold is preset; the sentences that meet the extraction probability threshold are selected as the extracted summary. The extraction summary corresponding to the text labeled L0 is directly used as the final summary of the original text. The extraction summaries corresponding to the texts labeled L1 and L2 are input into a summary optimization model for optimization to obtain the final summary.

5. The method for generating a text summary according to claim 4, characterized in that, The summary optimization model includes an encoder and a decoder. The encoder is used to receive the extraction summary of the Chinese text and obtain an encoder hidden vector through word embedding and position encoding. The decoder is used to predict the probability distribution of the words of the target summary text according to the encoder hidden vector by using a self-attention mechanism. Based on the word with the highest probability at each time step in the probability distribution, the final summary is obtained.

6. The method for generating a text summary according to claim 5, wherein, The summary optimization model predicts the probability distribution of words by the following formula: where P(w) is the probability distribution of the predicted word at time step t. p gen is the probability of generating a new word, and (1 - p gen ) is the probability of copying a word in the input sequence. When the predicted abstraction level label L ab = L1, p gen = 0, and the model only selects words from the original extractive summary as the summary; when the predicted abstraction level label L ab = L2, p gen is obtained through the above formula, and P vocab(w) is the probability distribution of the vocabulary words predicted by T5; is a vector of the vocabulary length, where the value at the word position in the input sequence is the attention distribution of each word and the values at the remaining positions are 0; is the context vector of the encoder at time step t, and a t is the attention distribution of the input word at time step t; s t is the decoder state at time step t, and x t is the decoder input at time step t; V′, V, b, b′ are learnable parameters.

7. A computer device for text summary generation, characterized in that, It includes at least one processor and at least one memory communicatively connected to the processor. The memory stores instructions executable by the processor, and the instructions are used to be executed by the processor to implement the text summary generation method according to any one of claims 1-6.

Citation Information

Patent Citations

  • Abstract generation method and device

    CN109635103A

  • Text abstract generation method based on advanced semantics

    CN109992775A