A text summary generation method, device and storage medium
By performing word segmentation processing and training model loss function analysis on the text data set, the abstract model is generated using the encoder, decoder and attention module, the problem of word order disorder in automatic abstract of Chinese text is solved, and the readability and accuracy of the abstract is improved.
Patent Information
- Application Number
- CN202210332378.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-30
- Publication Date
- 2025-09-05
- Estimated Expiration
- 2042-03-30
AI Technical Summary
In the existing automatic summary technology of Chinese text, the word order problem of generated text results in poor readability of the summary, which is difficult to effectively solve in the existing technology.
By performing word segmentation on the text data set, the training model is constructed and the loss function analysis is performed. The abstract model is generated using the encoder, decoder and attention modules, and the word order of the abstract is optimized for the generation of the summary by combining the cross entropy loss function and the ROUGE algorithm.
It improves the word order problem in generating abstracts, improves the readability and word order of abstracts, and improves the accuracy of generating abstracts.
Smart Images

Figure CN114662483B_ABST
Abstract
Description
Technical Field
[0001] The present invention mainly relates to the field of language processing technology, and specifically relates to a text summary generation method, device and storage medium. Background Art
[0002] Existing automatic summarization technology for Chinese text still has many shortcomings that need to be improved and addressed. Improving the readability of generated text is a crucial issue. There are many reasons for this readability issue, but addressing word order is crucial. When generating summaries, words are reordered, inevitably leading to word order issues in the generated summaries. Summary of the Invention
[0003] The technical problem to be solved by the present invention is to provide a text summary generation method, device and storage medium in response to the shortcomings of the existing technology.
[0004] The present invention solves the above technical problems with the following technical solution: A method for generating a text summary comprises the following steps:
[0005] Importing a text data set and performing word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information;
[0006] Constructing a training model, and training the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information;
[0007] performing a loss function analysis on the training model according to the plurality of original prediction sequences and the plurality of summary information, and obtaining a summary generation model according to the analysis result;
[0008] Each of the original text information and the summary information corresponding to each of the original text information are respectively input into the summary generation model for prediction analysis to obtain an evaluation score corresponding to each of the original text information, and all the evaluation scores are used as the results of text summary generation.
[0009] Another technical solution of the present invention to solve the above technical problem is as follows: a text summary generation device, comprising:
[0010] A word segmentation processing module is used to import a text data set and perform word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information;
[0011] A model training module is used to construct a training model, and train the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information;
[0012] a loss function analysis module, configured to perform a loss function analysis on the training model based on the plurality of original prediction sequences and the plurality of summary information, and obtain a summary generation model based on the analysis results;
[0013] The summary generation result acquisition module is used to input each of the original text information and the summary information corresponding to each of the original text information into the summary generation model for prediction analysis, obtain the evaluation score corresponding to each of the original text information, and use all the evaluation scores as the results of text summary generation.
[0014] Another technical solution of the present invention to solve the above technical problem is as follows: A text summary generation device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the text summary generation method described above is implemented.
[0015] Another technical solution of the present invention to solve the above technical problem is as follows: a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the text summary generation method as described above is implemented.
[0016] The beneficial effects of the present invention are as follows: a plurality of original text information and summary information are obtained by word segmentation processing of a text data set, an original prediction sequence is obtained by training a training model with each original text information respectively, a loss function of the training model is analyzed based on the plurality of original prediction sequences and the plurality of summary information, a summary generation model is obtained based on the analysis results, and each original text information and summary information are respectively input into the summary generation model for prediction analysis to obtain the result of text summary generation, so that the word order problem in the generated summary is improved, the word order of the text is smoother, and thus the readability of the summary is improved. BRIEF DESCRIPTION OF THE DRAWINGS
[0017] Figure 1 A flowchart of a text summary generation method provided by an embodiment of the present invention;
[0018] Figure 2 This is a module block diagram of a text summary generation device provided by an embodiment of the present invention. DETAILED DESCRIPTION
[0019] The principles and features of the present invention are described below with reference to the accompanying drawings. The examples given are only used to explain the present invention and are not used to limit the scope of the present invention.
[0020] Figure 1 A flowchart of a text summary generation method provided by an embodiment of the present invention.
[0021] like Figure 1 As shown, a text summary generation method includes the following steps:
[0022] Importing a text data set and performing word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information;
[0023] Constructing a training model, and training the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information;
[0024] performing a loss function analysis on the training model according to the plurality of original prediction sequences and the plurality of summary information, and obtaining a summary generation model according to the analysis result;
[0025] Each of the original text information and the summary information corresponding to each of the original text information are respectively input into the summary generation model for prediction analysis to obtain an evaluation score corresponding to each of the original text information, and all the evaluation scores are used as the results of text summary generation.
[0026] Preferably, the text dataset may be an LCSTS dataset.
[0027] In the above embodiment, multiple original text information and summary information are obtained by word segmentation processing of the text data set, and the original prediction sequence is obtained by training the training model with each original text information. The loss function of the training model is analyzed based on the multiple original prediction sequences and multiple summary information. The summary generation model is obtained based on the analysis results, and each original text information and summary information are input into the summary generation model for prediction analysis to obtain the result of text summary generation, so that the word order problem in the generated summary is improved, the word order of the text is smoother, and the readability of the summary is improved.
[0028] Optionally, as an embodiment of the present invention, the process of performing word segmentation processing on the text dataset to obtain multiple original text information and summary information corresponding to each of the original text information includes:
[0029] The text dataset is segmented using a Python tool to obtain a plurality of original text information and summary information corresponding to each of the original text information.
[0030] It should be understood that the LCSTS dataset (ie, the text dataset) is segmented using jieba of the python function package (ie, the python tool), words in a sentence are separated by spaces, and a dictionary file is created.
[0031] It should be understood that the main function of the python function package (i.e. the python tool) jieba is to perform Chinese word segmentation, which can perform simple word segmentation, parallel word segmentation, and command line word segmentation. Of course, its functions are not limited to this. It currently also supports keyword extraction, part-of-speech tagging, word position query, etc.
[0032] Specifically, the input text dataset (i.e., the text dataset) is preprocessed, including segmenting the LCSTS dataset (i.e., the text dataset) and establishing a word list, dividing the original text (i.e., the original information) and the summary information in the dataset (i.e., the text dataset) and storing them in two different documents respectively, and establishing a pause word word list.
[0033] It should be understood that the data set (i.e., the text data set) contains the original information and the summary information. The original information and the summary information in each data set (i.e., the text data set) are separated, the original information is written into the src file, and the summary information is written into the tgt file.
[0034] In the above embodiment, Python tools are used to perform word segmentation processing on a text dataset to obtain a plurality of original text information and summary information, providing accurate data support for subsequent processing, thereby improving the word order problem in generating the summary.
[0035] Optionally, as an embodiment of the present invention, the training model includes an encoder, a decoder and an attention module;
[0036] The process of constructing a training model and training the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information includes:
[0037] Inputting each of the original text messages into the encoder for encoding analysis, obtaining an encoding vector corresponding to each of the original text messages, an initial context vector corresponding to each of the original text messages, a plurality of target encoder hidden states corresponding to each of the original text messages, and a plurality of target context vectors corresponding to each of the original text messages;
[0038] Inputting each of the encoding vectors and a plurality of target encoder hidden states corresponding to each of the original text messages into the decoder for decoding analysis, thereby obtaining an initialized decoder hidden state corresponding to each of the original text messages and a plurality of target decoder hidden states corresponding to each of the original text messages;
[0039] The multiple target decoder hidden states corresponding to each of the original text messages, the encoding vectors corresponding to each of the original text messages, the initialized decoder hidden states corresponding to each of the original text messages, the initial context vectors corresponding to each of the original text messages, and the multiple target context vectors corresponding to each of the original text messages are respectively input into the attention module for word prediction analysis to obtain the original prediction sequence corresponding to each of the original text messages.
[0040] Preferably, the encoder may be a bidirectional LSTM long short-term memory artificial neural network, and the decoder may be a unidirectional LSTM long short-term memory artificial neural network.
[0041] It should be understood that the bidirectional LSTM long short-term memory artificial neural network, or LSTM, is a special type of RNN that can learn long-term dependency information in longer texts; the bidirectional LSTM is composed of two LSTMs superimposed on each other, the first LSTM starts inputting from the beginning of the sentence, and the other starts inputting from the last word of the sentence, and finally the two results are integrated and processed.
[0042] It should be understood that the unidirectional LSTM (Long Short-Term Memory) artificial neural network can only predict the output at the next moment based on the time series information at the previous moment. However, in some problems, the output at the current moment is not only related to the previous state, but may also be related to the future state. For example, predicting the missing word in a sentence requires not only considering the previous context but also the following content, truly making a context-based judgment.
[0043] It should be understood that the attention module, also known as the attention mechanism, helps focus on the most relevant information in the source sequence. The attention mechanism assigns high weights to the parts that are important to the result to retain the key information. The attention module is added between the decoder and the output layer. It is responsible for selecting task-related information from the encoder output sequence and passing it, along with the final state of the decoder output sequence, as features for text generation to the output layer.
[0044] It should be understood that a neural network model (i.e., the training model) is established based on the self-attention mechanism of seq2seq.
[0045] Specifically, a simple Encoder-Decoder framework cannot effectively focus on the input target, which means that models like seq2seq cannot achieve their maximum effectiveness when used alone. The encoder (i.e., the encoder) encodes the input (i.e., the original information) into context information (i.e., the initial context vector or the target context vector). During decoding, each output will be decoded using this context information (i.e., the initial context vector or the target context vector) without distinction. What the attention model does is to encode the encoder into different context information according to each time step of the sequence. During decoding, each different context information is combined for decoding output, resulting in a more accurate result.
[0046] In the above embodiment, the original prediction sequence is obtained by training the training model with each original text information, and the input end information can be obtained dynamically and on demand, so that the word order problem in the generated summary is improved, the word order of the text is smoother, and the readability of the summary is improved.
[0047] Optionally, as an embodiment of the present invention, each of the original text information includes a plurality of sequentially arranged word vectors, and the process of inputting each of the original text information into the encoder for encoding analysis to obtain an encoding vector corresponding to each of the original text information, an initial context vector corresponding to each of the original text information, a plurality of target encoder hidden states corresponding to each of the original text information, and a plurality of target context vectors corresponding to each of the original text information includes:
[0048] S211: Inputting the first word vector of each original message into the encoder for a first encoding process, thereby obtaining an initial encoder hidden state corresponding to each original message and an initial context vector corresponding to each original message;
[0049] S212: inputting the next word vector in each of the original text information and the initial encoder hidden state into the encoder in the order of arrangement of the word vectors to perform a second encoding process, thereby obtaining a target encoder hidden state corresponding to each of the word vectors and a target context vector corresponding to each of the word vectors;
[0050] S213: Use the target encoder hidden state as the next initial encoder hidden state, and return to step S212 until all word vectors are input into the encoder, thereby obtaining multiple target encoder hidden states corresponding to each of the original text information and multiple target context vectors corresponding to each of the original text information, and use the last target encoder hidden state in each of the original text information as the encoding vector corresponding to each of the original text information.
[0051] It should be understood that the bidirectional LSTM encoder (i.e., the encoder) sequentially receives the text in the source text (i.e., the original text information) of the dataset, and the encoder sequentially reads the entire input sequence (i.e., the original text information). At each time step, a word (i.e., the word vector) is sent to the encoder. At each time step, the encoder outputs a hidden state (i.e., the initial encoder hidden state or the target encoder hidden state) and context information (i.e., the initial context vector or the target context vector). The hidden state of the encoder at the final time step is used as the encoding information of the input sentence (i.e., the encoding vector).
[0052] In the above embodiment, each original text information is input into the encoder for encoding analysis to obtain an encoding vector, an initial context vector, multiple target encoder hidden states and multiple target context vectors. The input end information can be obtained dynamically and on demand, so that the word order problem in the generated summary is improved, the word order of the text is smoother, and the readability of the summary is improved.
[0053] Optionally, as an embodiment of the present invention, the process of inputting each of the encoding vectors and the multiple target encoder hidden states corresponding to each of the original text information into the decoder for decoding analysis to obtain the initialized decoder hidden state corresponding to each of the original text information and the multiple target decoder hidden states corresponding to each of the original text information includes:
[0054] S221: Initializing the decoder according to each encoding vector to obtain a hidden state of the initialized decoder corresponding to each original text information;
[0055] S222: Inputting each of the encoding vectors and the initialized decoder hidden state corresponding to each of the original text information into the decoder for a first decoding, obtaining an initial decoder hidden state corresponding to each of the original text information, and using the initial decoder hidden state as the target decoder hidden state corresponding to the word vector of the second word in the encoding vector;
[0056] S223: inputting each of the encoding vectors, the target encoder hidden state corresponding to the next word vector, and the initial decoder hidden state into the decoder in the order of arrangement of the word vectors, and performing a second decoding to obtain the target decoder hidden state corresponding to each of the word vectors;
[0057] S224: Use the target decoder hidden state as the next initial decoder hidden state and return to step S223 until all word vectors are input into the decoder, thereby obtaining multiple target decoder hidden states corresponding to each of the original text information.
[0058] It should be understood that a unidirectional LSTM is used as the decoder, which uses the encoded information of the input sentence and the output of the previous time step (i.e., the initial decoder hidden state) and the hidden state (i.e., the target encoder hidden state) as input in each time step, and outputs the hidden state at the current moment (i.e., the target decoder hidden state).
[0059] In the above embodiment, each encoding vector and multiple target encoder hidden states are respectively input into the decoder for decoding analysis to obtain an initialized decoder hidden state and multiple target decoder hidden states. This allows for dynamic and on-demand acquisition of input end information, thereby improving the readability of the summary.
[0060] Optionally, as an embodiment of the present invention, the process of inputting a plurality of target decoder hidden states corresponding to each of the original textual information, an encoding vector corresponding to each of the original textual information, an initialized decoder hidden state corresponding to each of the original textual information, an initial context vector corresponding to each of the original textual information, and a plurality of target context vectors corresponding to each of the original textual information into the attention module for word prediction analysis to obtain an original prediction sequence corresponding to each of the original textual information includes:
[0061] Inputting each of the initialized decoder hidden states and the initial context vector corresponding to each of the original text information into the attention module for a first word prediction, thereby obtaining a first predicted context vector corresponding to each of the original text information;
[0062] Inputting each of the target decoder hidden states and the target context vectors corresponding to each of the word vectors into the attention module in sequence according to the arrangement order of the word vectors to perform a second word prediction, thereby obtaining a second predicted context vector corresponding to each of the word vectors;
[0063] Performing a first splicing on each of the first predicted context vectors and the encoding vector corresponding to each of the original text information to obtain a first splicing vector corresponding to each of the original text information;
[0064] Performing a second splicing on each of the second predicted context vectors and the encoding vector corresponding to each of the original text information to obtain a second splicing vector corresponding to each of the word vectors;
[0065] Inputting each of the first concatenated vectors into the decoder for a third decoding to obtain a first predicted word corresponding to each of the original text information;
[0066] Inputting each of the second concatenated vectors into the decoder in sequence according to the arrangement order of the word vectors for fourth decoding to obtain a second predicted word corresponding to each of the word vectors;
[0067] The first predicted words corresponding to each of the original text information and the plurality of second predicted words corresponding to each of the original text information are combined respectively to obtain an original prediction sequence corresponding to each of the original text information.
[0068] It should be understood that the hidden state of the decoder (i.e., the initialized decoder hidden state or the target decoder hidden state) and the context information of the encoder (i.e., the initial context vector or the target context vector) are input into the attention module, and the output result is used as the context information (i.e., the first predicted context vector or the second predicted context vector), which is spliced with the encoded information of the encoder and input into the decoder. The decoder reads the entire target sequence word by word based on the output context information, predicts the same sequence offset at each time step, and predicts the next word based on the previous word.
[0069] In the above embodiment, multiple target decoder hidden states, encoding vectors, initialized decoder hidden states, initial context vectors and multiple target context vectors are respectively input into the attention module for word prediction analysis to obtain the original prediction sequence, so that the word order problem in the generated summary is improved, and the word order of the text is smoother, thereby improving the readability of the summary and the accuracy of the results.
[0070] Optionally, as an embodiment of the present invention, each summary information includes multiple summary word vectors, and the process of performing loss function analysis on the training model based on the multiple original prediction sequences and the multiple summary information, and obtaining the summary generation model based on the analysis results includes:
[0071] The loss function corresponding to each of the original prediction sequences and the summary information corresponding to each of the original text information is calculated by the first formula to obtain the loss function corresponding to each of the original text information. The first formula is:
[0072]
[0073] Among them, P α(i) (Y i )→P k (Y i ),
[0074] Among them, Y i is the i-th summary word vector in the summary information, L is the loss function, P α(i) (Y i ) is the probability distribution of the i-th summary word vector at the mapping position, i∈{1,...,n}, P1,...,P m are all the predicted words in the original prediction sequence, k∈{1,...,m}, k is the position after mapping on the summary word vector {1,...,m}, α(i) is the position of the i-th summary word vector in the summary information on the prediction sequence after mapping through the alignment function α;
[0075] The training model is updated according to the Adam gradient descent algorithm and all loss functions to obtain a summary generation model.
[0076] It should be understood that the alignment cross entropy loss function is added, the prediction sequence (i.e., the original prediction sequence) and the corresponding target sequence (i.e., the summary information) are input to calculate the cross entropy loss, and the back propagation calculation updates the model parameters to reduce the cross entropy loss function value (i.e., the loss function).
[0077] It should be understood that the loss function is used as the loss function training model (i.e., the training model), and the parameters of the neural network (i.e., the training model) are updated by the Adam gradient descent algorithm in python.
[0078] It should be understood that the Adam gradient descent algorithm is a stochastic optimization method with adaptive momentum, which can calculate the adaptive learning rate of each parameter and is often used as an optimizer algorithm in deep learning.
[0079] Specifically, the alignment cross entropy loss function is introduced and the operation is as follows:
[0080] Let Y={Y1,...,Y n} is the target sequence containing n tokens (i.e., the summary information), Pre is the model prediction sequence containing m tokens (i.e., the original prediction sequence), P1,...,P mFor the probability distribution of tokens in Pre, define an alignment function α to map the target position to the predicted position, that is, α:{1,...,n}→{1,...,m}.
[0081] The loss function is defined as:
[0082]
[0083] Wherein, α(i) represents the corresponding position of the i-th token of the target sequence (i.e., the summary information) on Pre (i.e., the original prediction sequence) after being mapped by the alignment function α. α(i) (Y i ) represents the probability distribution of the target sequence (ie, the summary information) at the mapping position.
[0084] Let j∈{1,...,n} be a position on {1,...,n}, its position after mapping on {1,...,m} is k, and k∈{1,...,m} then:
[0085] P α(j) (Y j )→P k (Y j ),
[0086] The loss function is used to calculate the alignment cross entropy of Y (ie, the summary information) and Pre (ie, the original prediction sequence).
[0087] In the above embodiment, a loss function analysis is performed on the training model based on multiple original prediction sequences and multiple summary information, and a summary generation model is obtained based on the analysis results. A backpropagation operation is used to minimize the loss function, so that the word order of the generated text is smoother and has better readability.
[0088] Optionally, as an embodiment of the present invention, the process of inputting each of the original text information and the summary information corresponding to each of the original text information into the summary generation model for prediction analysis to obtain an evaluation score corresponding to each of the original text information includes:
[0089] Inputting each of the original text information and the summary information corresponding to each of the original text information into the summary generation model to predict a target prediction sequence, thereby obtaining a target prediction sequence corresponding to each of the original text information;
[0090] The target prediction sequence and the summary information corresponding to each of the original text information are scored respectively using the ROUGE algorithm to obtain an evaluation score corresponding to each of the original text information.
[0091] It should be understood that the ROUGE algorithm, namely the Python function package ROUGE, is a commonly used machine translation and article summary evaluation indicator, which is mainly calculated based on the recall rate.
[0092] It should be understood that ROUGE is used to compare the generated summary (ie, the target prediction sequence) with the corresponding summary of the source text (ie, the summary information) to evaluate the effect of the generated summary (ie, the evaluation score).
[0093] It should be understood that, according to the input text (ie, the original text information), the corresponding summary content (ie, the target prediction sequence) is generated.
[0094] In the above embodiment, each original text information and summary information are respectively input into the summary generation model to perform target prediction sequence prediction to obtain a target prediction sequence, and the ROUGE algorithm is used to score the target prediction sequence and summary information respectively to obtain an evaluation score, which can well evaluate the effect of generating the summary and more intuitively know the result of the summary generation.
[0095] Optionally, as another embodiment of the present invention, the present invention mainly utilizes a public Chinese data set, adjusts the cross-entropy function, and calculates the cross-entropy loss based on the alignment between the label sequence (i.e., the summary information) and the prediction sequence (i.e., the target prediction sequence), so that the word order of the generated text is smoother and has better readability.
[0096] Alternatively, as another embodiment of the present invention, the method of the present invention is as follows:
[0097] First, the Chinese text of the data set (i.e., the text data set) is preprocessed with jieba for word segmentation; the text after word segmentation (i.e., the multiple original text information) is input into the encoder and encoded into a word vector, and the input end information is dynamically and on demand obtained through the attention mechanism (i.e., the attention module); a cross-entropy function is added to the decoding end (i.e., the decoder), and the loss of the relative position between the decoded text (i.e., the original prediction sequence) and the target text (i.e., the summary information) is calculated, and the backpropagation operation is used to minimize the loss function to complete the model training. Finally, the summary text (i.e., the target prediction sequence) is obtained through the output model (i.e., the summary generation model).
[0098] Optionally, as another embodiment of the present invention, the present invention adds an aligned cross entropy function to the summary generation model, and provides more accurate training for the summary model by ignoring absolute position and focusing on relative order and lexical matching, so that the word order problem in the generated summary is improved, thereby improving the readability of the summary.
[0099] Figure 2 This is a module block diagram of a text summary generation device provided by an embodiment of the present invention.
[0100] Alternatively, as another embodiment of the present invention, Figure 2 As shown, a text summary generation device includes:
[0101] A word segmentation processing module is used to import a text data set and perform word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information;
[0102] A model training module is used to construct a training model, and train the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information;
[0103] a loss function analysis module, configured to perform a loss function analysis on the training model based on the plurality of original prediction sequences and the plurality of summary information, and obtain a summary generation model based on the analysis results;
[0104] The summary generation result acquisition module is used to input each of the original text information and the summary information corresponding to each of the original text information into the summary generation model for prediction analysis, obtain the evaluation score corresponding to each of the original text information, and use all the evaluation scores as the results of text summary generation.
[0105] Alternatively, another embodiment of the present invention provides a text summary generation device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the above-described text summary generation method is implemented. The device may be a computer or other device.
[0106] Optionally, another embodiment of the present invention provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the above-mentioned text summary generation method is implemented.
[0107] It should be noted that, in this document, relational terms such as first and second, etc., are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the terms "comprises," "comprising," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that includes a list of elements includes not only those elements but also other elements not explicitly listed, or elements inherent to such process, method, article, or apparatus.
[0108] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described devices and units can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.
[0109] In the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of units is merely a logical functional division. In actual implementation, other division methods may be used, such as combining or integrating multiple units or components into another system, or ignoring or not implementing certain features.
[0110] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected based on actual needs to achieve the objectives of the embodiments of the present invention.
[0111] In addition, the functional units in the various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0112] If the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention is essentially or the part that contributes to the prior art, or all or part of the technical solution can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the various embodiments of the present invention. The aforementioned storage medium includes: U disk, mobile hard disk, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), disk or optical disk, and other media that can store program code.
[0113] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.
Claims
1. A text summary generation method, characterized in that: The steps include: Importing a text data set and performing word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information; Constructing a training model, and training the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information; performing a loss function analysis on the training model according to the plurality of original prediction sequences and the plurality of summary information, and obtaining a summary generation model according to the analysis result; Inputting each of the original text information and the summary information corresponding to each of the original text information into the summary generation model for prediction analysis, obtaining an evaluation score corresponding to each of the original text information, and using all the evaluation scores as the result of text summary generation; Each summary information includes a plurality of summary word vectors, and the process of performing loss function analysis on the training model according to the plurality of original prediction sequences and the plurality of summary information, and obtaining a summary generation model according to the analysis results includes: The loss function corresponding to each of the original prediction sequences and the summary information corresponding to each of the original text information is calculated by the first formula to obtain the loss function corresponding to each of the original text information. The first formula is: Among them, P α(i) (Y i )→P k (Y i ), Among them, Y i is the i-th summary word vector in the summary information, L is the loss function, P α(i) (Y i ) is the probability distribution of the i-th summary word vector at the mapping position, i∈{1,...,n}, P1,...,P m are all the predicted words in the original prediction sequence, k∈{1,...,m}, k is the position after mapping on the summary word vector {1,...,m}, α(i) is the position of the i-th summary word vector in the summary information on the prediction sequence after mapping through the alignment function α; The training model is updated according to the Adam gradient descent algorithm and all loss functions to obtain a summary generation model.
2. The text summary generation method according to claim 1, characterized in that The process of performing word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information includes: The text dataset is segmented using a Python tool to obtain a plurality of original text information and summary information corresponding to each of the original text information.
3. The text summary generation method according to claim 1, characterized in that The training model includes an encoder, a decoder, and an attention module; The process of constructing a training model and training the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information includes: Inputting each of the original text messages into the encoder for encoding analysis, obtaining an encoding vector corresponding to each of the original text messages, an initial context vector corresponding to each of the original text messages, a plurality of target encoder hidden states corresponding to each of the original text messages, and a plurality of target context vectors corresponding to each of the original text messages; Inputting each of the encoding vectors and a plurality of target encoder hidden states corresponding to each of the original text messages into the decoder for decoding analysis, thereby obtaining an initialized decoder hidden state corresponding to each of the original text messages and a plurality of target decoder hidden states corresponding to each of the original text messages; The multiple target decoder hidden states corresponding to each of the original text messages, the encoding vectors corresponding to each of the original text messages, the initialized decoder hidden states corresponding to each of the original text messages, the initial context vectors corresponding to each of the original text messages, and the multiple target context vectors corresponding to each of the original text messages are respectively input into the attention module for word prediction analysis to obtain the original prediction sequence corresponding to each of the original text messages.
4. The text summary generation method according to claim 3, characterized in that Each of the original text information includes a plurality of sequentially arranged word vectors. The process of inputting each of the original text information into the encoder for encoding analysis to obtain an encoding vector corresponding to each of the original text information, an initial context vector corresponding to each of the original text information, a plurality of target encoder hidden states corresponding to each of the original text information, and a plurality of target context vectors corresponding to each of the original text information includes: S211: Inputting the first word vector of each original message into the encoder for a first encoding process, thereby obtaining an initial encoder hidden state corresponding to each original message and an initial context vector corresponding to each original message; S212: inputting the next word vector in each of the original text information and the initial encoder hidden state into the encoder in the order of arrangement of the word vectors to perform a second encoding process, thereby obtaining a target encoder hidden state corresponding to each of the word vectors and a target context vector corresponding to each of the word vectors; S213: Use the target encoder hidden state as the next initial encoder hidden state, and return to step S212 until all word vectors are input into the encoder, thereby obtaining multiple target encoder hidden states corresponding to each of the original text information and multiple target context vectors corresponding to each of the original text information, and use the last target encoder hidden state in each of the original text information as the encoding vector corresponding to each of the original text information.
5. The text summary generation method according to claim 4, characterized in that The process of inputting each of the encoding vectors and the multiple target encoder hidden states corresponding to each of the original text information into the decoder for decoding analysis to obtain the initialized decoder hidden state corresponding to each of the original text information and the multiple target decoder hidden states corresponding to each of the original text information includes: S221: Initializing the decoder according to each encoding vector to obtain a hidden state of the initialized decoder corresponding to each original text information; S222: Inputting each of the encoding vectors and the initialized decoder hidden state corresponding to each of the original text information into the decoder for a first decoding, obtaining an initial decoder hidden state corresponding to each of the original text information, and using the initial decoder hidden state as the target decoder hidden state corresponding to the word vector of the second word in the encoding vector; S223: inputting each of the encoding vectors, the target encoder hidden state corresponding to the next word vector, and the initial decoder hidden state into the decoder in the order of arrangement of the word vectors, and performing a second decoding to obtain the target decoder hidden state corresponding to each of the word vectors; S224: Use the target decoder hidden state as the next initial decoder hidden state and return to step S223 until all word vectors are input into the decoder, thereby obtaining multiple target decoder hidden states corresponding to each of the original text information.
6. The text summary generation method according to claim 5, characterized in that The process of inputting a plurality of target decoder hidden states corresponding to each of the original text information, an encoding vector corresponding to each of the original text information, an initialized decoder hidden state corresponding to each of the original text information, an initial context vector corresponding to each of the original text information, and a plurality of target context vectors corresponding to each of the original text information into the attention module for word prediction analysis to obtain an original prediction sequence corresponding to each of the original text information includes: Inputting each of the initialized decoder hidden states and the initial context vector corresponding to each of the original text information into the attention module for a first word prediction, thereby obtaining a first predicted context vector corresponding to each of the original text information; Inputting each of the target decoder hidden states and the target context vectors corresponding to each of the word vectors into the attention module in sequence according to the arrangement order of the word vectors to perform a second word prediction, thereby obtaining a second predicted context vector corresponding to each of the word vectors; Performing a first splicing on each of the first predicted context vectors and the encoding vector corresponding to each of the original text information to obtain a first splicing vector corresponding to each of the original text information; Performing a second splicing on each of the second predicted context vectors and the encoding vector corresponding to each of the original text information to obtain a second splicing vector corresponding to each of the word vectors; Inputting each of the first concatenated vectors into the decoder for a third decoding to obtain a first predicted word corresponding to each of the original text information; Inputting each of the second concatenated vectors into the decoder in sequence according to the arrangement order of the word vectors for fourth decoding to obtain a second predicted word corresponding to each of the word vectors; The first predicted words corresponding to each of the original text information and the plurality of second predicted words corresponding to each of the original text information are combined respectively to obtain an original prediction sequence corresponding to each of the original text information.
7. The text summary generation method according to claim 3, characterized in that: The process of inputting each of the original text information and the summary information corresponding to each of the original text information into the summary generation model for prediction analysis to obtain the evaluation score corresponding to each of the original text information includes: Inputting each of the original text information and the summary information corresponding to each of the original text information into the summary generation model to predict a target prediction sequence, thereby obtaining a target prediction sequence corresponding to each of the original text information; The target prediction sequence and the summary information corresponding to each of the original text information are scored respectively using the ROUGE algorithm to obtain an evaluation score corresponding to each of the original text information.
8. A text summary generation device, characterized in that: include: A word segmentation processing module is used to import a text data set and perform word segmentation processing on the text data set to obtain a plurality of original text information and summary information corresponding to each of the original text information; A model training module is used to construct a training model, and train the training model according to each of the original text information to obtain an original prediction sequence corresponding to each of the original text information; a loss function analysis module, configured to perform a loss function analysis on the training model based on the plurality of original prediction sequences and the plurality of summary information, and obtain a summary generation model based on the analysis results; a summary generation result obtaining module, configured to input each of the original text information and the summary information corresponding to each of the original text information into the summary generation model for prediction analysis, obtain evaluation scores corresponding to each of the original text information, and use all the evaluation scores as the results of text summary generation; Each summary information includes a plurality of summary word vectors, and the loss function analysis module is specifically used to: The loss function corresponding to each of the original prediction sequences and the summary information corresponding to each of the original text information is calculated by the first formula to obtain the loss function corresponding to each of the original text information. The first formula is: Among them, P α(i) (Y i )→P k (Y i ), Among them, Y i is the i-th summary word vector in the summary information, L is the loss function, P α(i) (Y i ) is the probability distribution of the i-th summary word vector at the mapping position, i∈{1,...,n}, P1,...,P m are all the predicted words in the original prediction sequence, k∈{1,...,m}, k is the position after mapping on the summary word vector {1,...,m}, α(i) is the position of the i-th summary word vector in the summary information on the prediction sequence after mapping through the alignment function α; The training model is updated according to the Adam gradient descent algorithm and all loss functions to obtain a summary generation model.
9. A computer-readable storage medium storing a computer program, characterized in that: When the computer program is executed by a processor, the text summary generation method according to any one of claims 1 to 7 is implemented.
Citation Information
Patent Citations
Text abstract generation method and device, computer equipment and storage medium
CN109657051A
Text abstract generation method and device, equipment and storage medium
CN113268586A