Abstract generation method and device, electronic equipment and medium
By constructing the target comparison learning summary group and adjusting the parameters of the summary model, the problem of abstracts being unrelated to the source text or duplicate content in the generative summary model is solved, and the accuracy and diversity of summary generation are improved.
Patent Information
- Application Number
- CN202311629377.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-11-30
- Publication Date
- 2025-06-03
AI Technical Summary
The existing generative summary model has exposure errors when generating text summary, resulting in the generated summary being unrelated to the source text or the content being duplicated, resulting in poor results and low accuracy.
By constructing a target comparison learning summary group, the target summary prediction probability and quality evaluation values of candidate summary candidate summary are calculated, and the summary model parameters are adjusted to improve the accuracy of summary generation.
The problem of high probability of generated prediction summary and poor actual evaluation is effectively avoided, and the accuracy and diversity of summary generation is improved.
Smart Images

Figure CN120086361A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the technical field of natural language processing, and in particular, to a method, apparatus, electronic device, and medium for generating abstracts. Background Art
[0002] With the advent of the big data era, information data has grown explosively, generating hundreds of millions of data and information every day, and inevitably facing the problem of information overload. Based on this, the text summarization generation technology has emerged. The text summarization generation technology refers to the technology of highly summarizing the original information through a concise piece of text. Currently, text summarization generation mainly includes: extractive summarization and generative summarization. Extractive summarization is to use relevant algorithms to extract ready-made sentences from the source text as summary sentences, and generative summarization is based on natural language processing (Natural Language Processing, NLP) technology. According to the content of the source text, the generative summary model automatically generates a natural language description instead of extracting sentences from the original text.
[0003] The current generative summary models mainly use neural network models with an encoder-decoder structure (such as the transformer model, seq2seq model) as the basic model, and then train the model through maximum likelihood estimation (Maximum Likelihood Estimate, MLE). Given a reference summary sequence, the prediction probability of generating the reference summary is maximized. However, the generative summary model trained by it has an exposure error, resulting in the generated summary being irrelevant to the source text or the content in the generated summary being repeated when generating text summaries through the trained generative summary model, resulting in poor text summary generation effect and low accuracy. Summary of the Invention
[0004] Embodiments of this application provide a method, apparatus, electronic device, and storage system for generating abstracts, aiming to solve the problems of poor text summary generation effect and low accuracy in the prior art.
[0005] In a first aspect, this application provides a method for generating an abstract, and the method for generating an abstract includes:
[0006] Obtain the original text to be processed;
[0007] Preprocess the original text to be processed to obtain the target original text;
[0008] Perform abstract generation processing on the target original text through a target abstract model to obtain the target text abstract of the original text to be processed.
[0009] As a feasible embodiment of the present application, the process of generating a target text summary of the target original text through the target summary model to obtain the target text summary of the to-be-processed original text includes:
[0010] Processing the target original text through the target summary model to obtain a candidate summary of the target original text;
[0011] Obtaining a reference summary corresponding to the candidate summary, constructing a target contrastive learning summary group based on the reference summary and the candidate summary, and calculating the target summary prediction probability and the target quality evaluation value of each contrastive learning summary in the target contrastive learning summary group;
[0012] Processing the candidate summary according to the target summary prediction probability and the target quality evaluation value to determine the target text summary of the to-be-processed original text.
[0013] As a feasible embodiment of the present application, the process of generating a target text summary of the target original text through the target summary model to obtain the target text summary of the to-be-processed original text includes:
[0014] Obtaining a predicted word sequence of the target original text, querying a summary word list through the target summary model, and obtaining initial candidate words whose association degree with the target predicted word in the predicted word sequence is higher than a preset probability value;
[0015] Determining a target candidate word of the target predicted word from the initial candidate words according to the demand evaluation value of the initial candidate words; the demand evaluation value is an evaluation parameter reflecting the relevance between the initial candidate words and the predicted word sequence;
[0016] Updating the target candidate word to the predicted word sequence, and generating a candidate summary whose summary prediction probability is greater than or equal to a preset summary threshold based on the updated predicted word sequence, where the summary prediction probability is the product of the word element prediction probabilities of the predicted words in the candidate summary.
[0017] As a feasible embodiment of the present application, before determining the target candidate word of the target predicted word from the initial candidate words according to the demand evaluation value of the initial candidate words, it further includes:
[0018] For each of the initial candidate words, obtaining the text similarity between the initial candidate word and each predicted word in the predicted word sequence;
[0019] Obtaining an average similarity value according to the text similarity between the initial candidate word and the predicted word, and setting the average similarity value as the word element difference value of the initial candidate word.
[0020] Obtain the demand evaluation value of the initial candidate word according to the lemma difference value and the lemma prediction probability of the initial candidate word.
[0021] As a feasible embodiment of the present application, the preprocessing of the to-be-processed original text to obtain the target original text includes:
[0022] Perform word segmentation on the to-be-processed original text according to the trained target word segmentation model to obtain the segmented original text;
[0023] Filter non-text symbols in the segmented original text to obtain the target original text.
[0024] As a feasible embodiment of the present application, the target summary model is obtained based on the following steps:
[0025] Process the first training text through a pre-trained initial summary model to obtain training candidate summaries, and construct a training comparison summary group according to the training reference summary of the second training text and each of the training candidate summaries;
[0026] Calculate the text similarity between each training comparison summary in the training comparison summary group and the reference summary of the first training text, and determine the quality evaluation value of the training comparison summary according to the text similarity;
[0027] Calculate a first comparison loss value according to the training summary prediction probability and the training quality evaluation value of the training comparison summary in the training comparison summary group; the training quality evaluation value is used to reflect the matching degree between the training comparison summary in the training comparison summary group and the reference summary of the first training text;
[0028] Adjust the model parameters of the initial summary model based on the first comparison loss value to obtain the target summary model.
[0029] As a feasible embodiment of the present application, the calculating the first comparison loss value according to the training summary prediction probability and the training quality evaluation value of the training comparison summary in the training comparison summary group includes:
[0030] For each training comparison summary group, sort the training reference summary of the second training text and each of the training candidate summaries according to the training quality evaluation value to obtain a first comparison learning sequence;
[0031] Perform a difference calculation on the training comparison summary group according to the quality evaluation value of the training comparison summary in the training comparison summary group, the maximum quality evaluation value and the minimum quality evaluation value in the first comparison learning sequence to obtain the sequence difference value of the training comparison summary group;
[0032] Obtain the summary prediction probability of the contrastive learning summary under the condition of the first training text based on the model parameters of the initial summary model and the probability density function of the model parameters;
[0033] Calculate the first contrast loss value according to the summary difference value of the training contrast summary group and the summary prediction probability of the training contrast summary in the training contrast summary group.
[0034] In a second aspect, an embodiment of the present application provides a summary generation device, which includes:
[0035] A text acquisition module, configured to acquire an original text to be processed;
[0036] A preprocessing module, configured to preprocess the original text to be processed to obtain a target original text;
[0037] A summary generation module, configured to perform summary generation processing on the target original text through a target summary model to obtain a target text summary of the original text to be processed.
[0038] In a third aspect, an embodiment of the present application further provides an electronic device, including a processor and a memory, where the memory stores multiple instructions; the processor loads the instructions from the memory to execute the summary generation method described in any one of the above embodiments, or the steps of the summary generation method described above.
[0039] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, which stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the summary generation method described in any one of the above embodiments, or the steps of the summary generation method described above.
[0040] In the embodiment of the present application, by acquiring the original text to be processed; preprocessing the original text to be processed to obtain a target original text; and performing summary generation processing on the target original text through a target summary model to obtain a target text summary of the original text to be processed. It is realized to perform summary generation processing on the original text through a target summary model based on contrastive learning, avoiding the problem that the prediction summary probability of the generated prediction summary is high while the prediction summary is poor in actual evaluation, thereby improving the accuracy of summary generation. Description of the Drawings
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present disclosure, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the following drawings are only some embodiments of the present disclosure. For those skilled in the art, without creative efforts, other drawings can be obtained based on these drawings.
[0042] Figure 1 A flowchart of the abstract generation method provided by an embodiment of the present application;
[0043] Figure 2 A structural diagram of a preset neural network model provided by an embodiment of the present application;
[0044] Figure 3 A flowchart of an embodiment of the training of the target abstract model in the abstract generation method provided by an embodiment of the present application;
[0045] Figure 4 A schematic diagram of calculating the abstract prediction probability provided by an embodiment of the present application;
[0046] Figure 5 A schematic diagram of calculating the quality evaluation value provided by an embodiment of the present application;
[0047] Figure 6 A flowchart of an embodiment of calculating the first loss value in the abstract generation method provided by an embodiment of the present application;
[0048] Figure 7 A structural diagram of an embodiment of the abstract generation device provided by an embodiment of the present application;
[0049] Figure 8 A structural diagram of an embodiment of an electronic device provided by an embodiment of the present application; Detailed implementation manners
[0050] Next, the technical solutions in the embodiments of the present application will be clearly and completely described in conjunction with the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments of the present application, all other embodiments obtained by those skilled in the art without creative efforts fall within the protection scope of the present application.
[0051] In the description of the present application, the terms "first" and "second" are only used for descriptive purposes and cannot be understood as indicating or implying relative importance or implicitly specifying the quantity of the indicated technical features. Thus, the features defined with "first" and "second" may explicitly or implicitly include one or more of the described features. In the description of the present application, "a plurality" means two or more unless otherwise specifically defined.
[0052] In the description of the present application, the term "for example" is used to mean "serving as an example, illustration, or explanation". Any embodiment described as "for example" in the present application is not necessarily to be construed as more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to implement and use the present invention. In the following description, details are set forth for purposes of explanation. It should be understood that those of ordinary skill in the art can recognize that the present invention can be implemented without these specific details. In other instances, well-known structures and processes are not elaborated in detail to avoid obscuring the description of the present invention with unnecessary details. Therefore, the present invention is not intended to be limited to the embodiments shown, but rather to be in line with the broadest scope consistent with the principles and features disclosed in the present application.
[0053] An embodiment of the present application provides a method, apparatus, electronic device, and computer-readable storage medium for abstract model training and abstract generation.
[0054] Specifically, the embodiment of the present application will be described from the perspective of the abstract generation method. The abstract generation method can be specifically integrated in an electronic device, that is, the abstract generation method provided by the embodiment of the present application can be executed by the electronic device. Optionally, the electronic device can be a terminal device. The terminal device can be a mobile phone, a tablet computer, a smart Bluetooth device, a laptop computer, a game console, or a personal computer (PC), etc. Optionally, the electronic device can be a server, which can be an independent server or a server network or server cluster composed of servers, including but not limited to a computer, a network host, a single network server, multiple network servers, or a cloud server composed of multiple servers. Among them, the cloud server is composed of a large number of computers or network servers based on cloud computing. It can be seen that the abstract generation method provided by the embodiment of the present application can run on a local terminal device or a server.
[0055] The following will be described in detail with reference to the accompanying drawings. In this embodiment, the execution entity is taken as an electronic device as an example. It should be noted that the description order of the following embodiments does not limit the preferred order of the embodiments. Although the logical order is shown in the flowchart, in some cases, the order shown or described can be executed in a different order from that shown in the drawings.
[0056] Figure 1 is a schematic flowchart of the abstract generation method provided by the embodiment of the present application. As Figure 1 shown, the abstract generation method at least includes the following steps:
[0057] S101, obtain the original text to be processed;
[0058] The abstract generation method in this embodiment is applied to an electronic device. The type and number of the electronic devices are not specifically limited. That is, the electronic device can be one or more intelligent terminals or servers. In a specific embodiment, the electronic device can be an intelligent computer with an abstract generation function.
[0059] Specifically, during the operation of the electronic device, it responds to an abstract generation request triggered for the original text to be processed. Among them, the abstract generation request is an operation instruction that drives the electronic device to generate an abstract for the specified original text to be processed. Optionally, the triggering method of the abstract generation request is not specifically limited here. That is, the abstract generation request can be actively triggered by the user. For example, the user clicks the abstract generation control through a mouse, keyboard, or touch screen to actively trigger the abstract generation request carrying the original text to be processed. Optionally, the abstract generation request can also be automatically triggered by the electronic device. For example, the electronic device has pre-set an abstract generation process and actively triggers the abstract generation request when detecting specific text.
[0060] Specifically, the electronic device responds to the abstract generation request and obtains the original text to be processed corresponding to the abstract generation request. Among them, the original text to be processed is the original text data to be subjected to abstract extraction, including specific text and symbols.
[0061] S102. Preprocess the original text to be processed to obtain a target original text;
[0062] Specifically, before performing the abstract generation process on the original text to be processed, in order to improve the accuracy and efficiency of abstract generation, the electronic device also preprocesses the original text to be processed to obtain the target original text corresponding to the original text to be processed.
[0063] Specifically, the electronic device has pre-set a trained target word segmentation model for text word segmentation, and after obtaining the original text to be processed, uses the trained target word segmentation model to perform word segmentation on the original text to be processed to obtain a segmented original text. Among them, the trained target word segmentation model can be any one or more of a trained hidden Markov model, a machine learning model, and a deep learning model.
[0064] Specifically, the electronic device inputs the original text to be processed into the trained target word segmentation model and processes the original text to be processed through the word segmentation strategy in the target word segmentation model to obtain a segmented original text. Among them, the word segmentation strategy can be a dictionary word segmentation strategy, a probability word segmentation strategy, etc. The segmented original text is a set of segmented text obtained by splitting the continuous original text to be processed into words with specific meanings.
[0065] Specifically, after the electronic device performs word segmentation on the original text to be processed and obtains the segmented original text, there may be non-text symbols such as special characters in the segmented original text, which may affect the accuracy of subsequent summary generation. Therefore, the electronic device also identifies the non-text symbols in the segmented original text and filters the non-text symbols in the segmented original text to obtain the target original text. The target original text is a text set obtained by performing preprocessing such as word segmentation and special symbol filtering on the original text to be processed.
[0066] S103. Perform summary generation processing on the target original text through the target summary model to obtain the target text summary of the original text to be processed.
[0067] Specifically, after the electronic device obtains the preprocessed target original text, it inputs the target original text into the target summary model, performs summary generation processing on the target original text through the target summary model, obtains the target text summary corresponding to the original text to be processed, and outputs the target text summary.
[0068] Among them, the target summary model is obtained by adjusting the model parameters of the pre-trained initial summary model based on the first contrast loss value. The first contrast loss value is calculated according to the summary prediction probability and quality evaluation value of the contrast learning summary in the contrast learning summary group. The contrast learning summary group is constructed based on the candidate summary obtained by processing the first training text through the pre-trained initial summary model and the reference summary of the second training text. The quality evaluation value is used to reflect the matching degree between the contrast learning summary and the reference summary of the first training text in the contrast learning summary group.
[0069] Among them, the above initial summary model is a summary model obtained by pre-training a preset neural network model based on an encoder-decoder structure (for example: transformer model, Seq2Seq model). For example, as Figure 2 shown Figure 2 is a schematic structural diagram of the preset neural network model provided by the embodiment of the present application. The terminal inputs the training text into the encoder of the preset neural network model and the reference summary of the training text into the decoder of the preset neural network model, and adjusts the model parameters of the preset neural network model through maximum likelihood estimation to obtain the pre-trained initial summary model, so that the initial summary model can process the original text and obtain the predicted summary of the original text.
[0070] Specifically, after the electronic device inputs the target original text into the target summary model, it processes the target original text through the target summary model to obtain the candidate summary of the target original text. That is, the terminal obtains the predicted word sequence of the target original text and queries the summary word table through the target summary model to obtain the initial candidate words whose association degree with the target predicted word in the predicted word sequence is higher than the preset probability value.
[0071] Specifically, after obtaining the initial candidate words, the demand evaluation values corresponding to the initial candidate words are further obtained, and the target candidate words of the target prediction word are determined according to the demand evaluation values corresponding to the initial candidate words, where the demand evaluation value is an evaluation parameter reflecting the relevance between the initial candidate word and the prediction word sequence.
[0072] Among them, the calculation method of the demand evaluation value corresponding to the initial candidate word is as follows:
[0073]
[0074] Among them, S t (v) is the demand evaluation value corresponding to the initial candidate word, α is a preset hyperparameter, p θ (v|x <t ) is the word element prediction probability of the initial candidate word v, x <t is the prediction word sequence, is the word element difference of the initial candidate word v, h v is the feature vector of the initial candidate word v, is the feature vector of the j-th prediction word x j in the prediction word sequence.
[0075] That is, the terminal obtains the text similarity between each initial candidate word and each prediction word in the prediction word series through the target summary model, and obtains the average similarity according to the text similarity between the initial candidate word and the prediction word, sets the average similarity as the word element difference value of the initial candidate word, and calculates the demand evaluation value of the initial candidate word according to the word source difference value and the word element prediction probability of the initial candidate word.
[0076] Specifically, after obtaining the target candidate word, update the target candidate word to the prediction word sequence, and generate candidate summaries with a summary prediction probability greater than or equal to the prediction summary threshold based on the updated prediction word sequence. Among them, the summary prediction probability is the product of the word element prediction probabilities of the prediction words in the candidate summary.
[0077] The electronic device obtains each target candidate word based on the target summary model. After updating all the target candidate words to the prediction word sequence until the summary prediction probability is greater than or equal to the preset summary threshold, the prediction word sequence with a summary prediction probability greater than or equal to the preset summary threshold is set as the target prediction word sequence, and the text of the target prediction word sequence in the arrangement order is the candidate summary.
[0078] Among them, the summary prediction probability P a corresponding to the prediction word sequence is:
[0079] P a = p θ (x 2 |x 1)×…×p θ (v|x <t );
[0080] That is, the summary prediction probability P corresponding to the predicted word sequence a is the product of the token prediction probabilities of the predicted words in the predicted word sequence.
[0081] Specifically, the electronic device also obtains the reference summary corresponding to the candidate summary, constructs a target contrastive learning group based on the reference summary and the candidate summary, calculates the target summary prediction probability and the target quality evaluation value of each contrastive learning summary in the target contrastive learning group, and processes the candidate summary according to the target summary prediction probability and the target quality evaluation value to determine the target summary of the original text to be processed. Among them, the reference summary can be a reference text in the same batch as the text to be processed, or a summary text extracted from a reference text generated by the text to be processed through other encoding methods.
[0082] Specifically, the electronic device sets the obtained candidate summary and reference summary as contrastive learning summaries, forms a contrastive summary set, and selects any two contrastive learning summaries from the contrastive summary set to form a target contrastive learning summary group to complete the construction of the target contrastive learning summary group.
[0083] Specifically, after constructing each target contrastive learning summary group, a target text summary is also generated according to the summary prediction probability and the target quality evaluation value in the target contrastive learning summary group.
[0084] Among them, the target quality evaluation value reflects the text similarity between the contrastive learning summary in the target contrastive learning summary group and the reference summary of the target original text.
[0085] For each target contrastive learning summary group, calculate the text similarity between the contrastive learning summary in the target contrastive learning summary group and the reference summary of the target original text; determine the target quality evaluation value of the contrastive learning summary according to the text similarity.
[0086] Specifically, according to the ROUGE value (i.e., text similarity) of each contrastive learning summary and the reference summary S of the target original text * , determine the corresponding target quality evaluation value according to the ROUGE value, that is, the target quality evaluation value η i of the contrastive learning summary S i is:
[0087] η i =ROUGE(S i ,S * ).
[0088] In one example, the ROUGE value, i.e., the text similarity value, can be directly used as the target quality evaluation value, or a variant of the ROUGE value (e.g., multiplied by a weight or added with a base value) can be used as the target quality evaluation value, which is not specifically limited in the embodiments of the present application.
[0089] Specifically, after the target summary model determines the target summary prediction probability and the target quality evaluation value of each contrastive learning summary, the contrastive learning summary with a target quality evaluation value greater than a preset quality evaluation threshold and a target summary prediction probability greater than a preset prediction probability threshold is used as the target summary text, and the target text summary of the to-be-processed original text is generated based on each target summary text.
[0090] In this embodiment, the electronic device obtains the to-be-processed original text; preprocesses the to-be-processed original text to obtain the target original text; and performs summary generation processing on the target original text through the target summary model to obtain the target text summary of the to-be-processed original text. It realizes the summary generation processing of the original text through the target summary model based on contrastive learning, avoids the problem that the prediction summary probability of the generated prediction summary is high while the prediction summary is poor in actual evaluation, thereby improving the accuracy of summary generation.
[0091] As Figure 3 shown, Figure 3 is a schematic flowchart of an embodiment of the training of the target summary model in the summary generation method provided by the embodiments of the present application. Specifically, in this embodiment, the summary generation method further includes steps 201 to 204:
[0092] 201. Process the first training text through a pre-trained initial summary model to obtain training candidate summaries, and construct a training contrast summary group according to the training reference summary of the second training text and each of the training candidate summaries;
[0093] 202. Calculate the text similarity between each training contrast summary in the training contrast summary group and the reference summary of the first training text, and determine the quality evaluation value of the training contrast summary according to the text similarity;
[0094] 203. Calculate a first contrast loss value according to the training summary prediction probability and the training quality evaluation value of the training contrast summaries in the training contrast summary group;
[0095] 204. Adjust the model parameters of the initial summary model based on the first contrast loss value to obtain the target summary model.
[0096] Based on the above embodiments, in this embodiment, the electronic device pre-sets an initial summary model and performs summary generation training on the initial summary model to obtain a target summary model. Among them, the initial summary model is a summary model obtained by pre-training a preset neural network model based on the encoder-decoder structure (for example: transformer model, Seq2Seq model),
[0097] In the embodiments of the present application, the first training text is input into the pre-trained initial summary model to obtain candidate summaries of the first training text and the summary prediction probabilities of each candidate summary. For example, as Figure 4 shown Figure 4 is a schematic diagram for calculating the summary prediction probability provided by the embodiments of the present application. Taking the initial summary model as the Seq2Seq model as an example, the first training text is input into the pre-trained initial summary model to obtain candidate summary 1, candidate summary 2,..., candidate summary i,..., candidate summary n of the first training text and the summary prediction probability P a .
[0098] In the embodiments of the present application, when obtaining the predicted word sequence of the first training text through the initial summary model, the next predicted word is predicted through the current predicted word sequence. Therefore, at the initial stage of the predicted word sequence, there may only be an initial predicted word, which can be a preset given word or a word from the original text, and no specific limitation is made in the embodiments of the present application. Among them, the predicted word sequence includes at least one predicted word, and the predicted words are arranged in sequence according to the prediction order. The above predicted words are the predicted words used to form the candidate summary.
[0099] It can be understood that each preset word in the summary vocabulary may be the next predicted word of the target predicted word. Therefore, the word element prediction probability P l of each preset word in the summary vocabulary corresponding to the target predicted word is obtained through the initial summary model. The word element preset probability is set as the association degree between the target predicted word and the preset word. The preset words in the summary vocabulary with a word element prediction probability greater than or equal to the preset word element threshold P l ' are used as the initial candidate words of the target predicted word. Among them, multiple preset words are pre-set in the above summary vocabulary, and the above target predicted word is the last predicted word in the predicted word sequence arranged in order.
[0100] For example, if there are 100 preset words in the summary vocabulary and the preset word element threshold P l ' is 0.5, then the preset words with a word element prediction probability greater than or equal to 0.5 are selected from the summary vocabulary as the initial candidate words of the predicted word.
[0101] In one example, a demand threshold, i.e., a preset demand threshold, can be set in advance. The initial candidate words whose training demand evaluation values are greater than or equal to the preset demand threshold are set as the target candidate words of the target prediction word, that is, to determine the target candidate words of the target prediction word from the initial candidate words according to the training demand evaluation values of the initial candidate words. The above training demand evaluation value is used to reflect the relevance between the initial candidate word and the prediction word sequence.
[0102] In one example, the number of candidate words, i.e., the preset number of candidate words m, can be set in advance. The demand evaluation sequence is obtained by sorting according to the magnitudes of the training demand evaluation values of the initial candidate words, and m initial candidate words are obtained from the demand evaluation sequence according to the rule of descending training demand evaluation values as the target candidate words of the target prediction word, that is, to determine the target candidate words of the target prediction word from the initial candidate words according to the training demand evaluation values of the initial candidate words.
[0103] It can be understood that the above preset demand threshold or the preset number of candidate words can be set according to the actual situation and is not specifically limited in the embodiments of the present application.
[0104] In the embodiments of the present application, first calculate the text similarity between each initial candidate word and each prediction word in the preset word sequence, then calculate the similarity average value according to the text similarity, and use the similarity average value as the word element difference value of the initial candidate word.
[0105] Specifically, the training demand evaluation value S t (v) of the initial candidate word v is shown in the following formula:
[0106]
[0107] Among them, α is a preset hyperparameter, p θ (v|x <t ) is the word element prediction probability of the initial candidate word v, x <t is the prediction word sequence, is the word element difference of the initial candidate word v, h v is the feature vector of the initial candidate word v, is the feature vector of the j-th prediction word x j in the prediction word sequence.
[0108] In the embodiments of the present application, taking the word element difference of the initial candidate word as the penalty term of the training demand evaluation value of the initial candidate word can avoid the occurrence of repeated prediction words, thereby further improving the accuracy of the predicted summary generated by the target summary model for the original text.
[0109] Update the target candidate word as the next predicted word of the target predicted word to the predicted word sequence to obtain the updated predicted word sequence, and calculate the summary prediction probability P corresponding to the updated predicted word sequence a , if the summary prediction probability P of the updated predicted word sequence a is less than the predicted summary threshold P a ′, then continue to obtain the next predicted word according to the above steps until the summary prediction probability is greater than or equal to the preset summary threshold, and set the predicted word sequence with the summary prediction probability greater than or equal to the preset summary threshold as the target predicted word sequence. The text of the target predicted word sequence in the order of arrangement is the candidate summary. Among them, the summary prediction probability is the product of the token prediction probabilities of the predicted words in the preset word sequence
[0110] Among them, the summary prediction probability P corresponding to the predicted word sequence a is
[0111] P a = p θ (x 2 |x 1 ) × … × p θ (v|x <t );
[0112] That is, the summary prediction probability P corresponding to the predicted word sequence a is the product of the token prediction probabilities of the predicted words in the predicted word sequence
[0113] In the embodiments of the present invention, the sampling space (i.e., the predicted word sequence) of the first training text is constructed by the token prediction probability of the preset word in the summary word table and the demand evaluation value, and stops until the product of the token prediction probabilities of all the predicted words in the sampling space (i.e., the summary prediction probability corresponding to the sampling space) is greater than or equal to the preset summary threshold. That is, after giving the initial candidate word, the prediction of the next predicted word is carried out. Taking the predicted word sequence x <t as an example, the x t for predicting the target predicted word is that the sampling space V (P) is
[0114] V (P) = argmin|S| s.t. ∑ v∈S (p θ (v|x <t )) ≥ P′ a .
[0115] As can be seen from the above embodiments, if the number of target candidate words of the target prediction word is greater than or equal to 1, and if there are m target candidate words for the target prediction word, then m prediction word sequences are obtained after updating. When continuing with the next prediction word, there are also m target candidate words, and at this time, there will be m×m prediction word sequences. Therefore, the embodiments of the present application can obtain multiple candidate summaries and the summary prediction probabilities of each candidate summary.
[0116] In the embodiments of the present application, the initial candidate words of the target prediction word are obtained from the summary word list according to the word element prediction probabilities of each preset word, and then the target candidate words of the target prediction word are obtained from the initial candidate words according to the demand evaluation values of the initial candidate words, realizing the prediction of the next continuous prediction word. It can ensure the coherence of the semantics before and after during the summary generation process, prevent model degradation, and enhance the diversity of the generated candidate summaries to improve the accuracy of the target summary model in predicting the summary. On the other hand, since there are many prediction words in the summary word list, obtaining the initial candidate words first and then obtaining the target candidate words from the initial candidate words can save computing resources and improve the training efficiency of the summary model.
[0117] In the embodiments of the present application, a second training text of the same batch as the first training text is obtained, and the reference summary of the second training text and the candidate summary of the first training text are set as the training contrast learning summaries to form a training contrast summary sample set S^. Any two training contrast learning summaries in the training contrast summary sample set are used as a training contrast learning summary group to complete the construction of the training contrast learning summary group.
[0118] For each training contrast learning summary group, calculate the text similarity between the training contrast learning summary in the training contrast learning summary group and the reference summary of the first training text; determine the training quality evaluation value of the training contrast learning summary according to the text similarity. The training quality evaluation value is used to reflect the text similarity between the contrast learning summary in the contrast learning summary group and the reference summary of the first training text.
[0119] Specifically, according to the ROUGE function, the ROUGE value (i.e., the text similarity) of each training contrast learning summary and the reference summary S * of the first training text is calculated, and the corresponding quality evaluation value is determined according to the ROUGE value. As Figure 5 shown, Figure 5 is a schematic diagram of calculating the quality evaluation value provided by the embodiments of the present application. That is, the quality evaluation value η i of the contrast learning summary S i is:
[0120] η i = ROUGE(S i , S * ).
[0121] In one example, the ROUGE value, i.e., the text similarity value, can be directly used as the quality evaluation value, or a variant of the ROUGE value (e.g., multiplied by a weight or added with a base value) can be used as the quality evaluation value, which is not specifically limited in the embodiments of the present application.
[0122] In some alternative embodiments, as Figure 6 shown, Figure 6 FIG. is a schematic flowchart of an embodiment for calculating a first loss value in the abstract generation method provided by the embodiments of the present application. In step S203: according to the training summary prediction probability and the training quality evaluation value of the training summary in the training comparison summary group, calculate a first comparison loss value, which can be implemented by at least the following steps:
[0123] S301. For each of the training comparison summary groups, sort the training reference summary of the second training text and each of the training candidate summaries according to the training quality evaluation value to obtain a first comparison learning sequence.
[0124] S302. According to the quality evaluation value of the training comparison summary in the training comparison summary group, the maximum quality evaluation value and the minimum quality evaluation value in the first comparison learning sequence, perform a difference calculation on the training comparison summary group to obtain a sequence difference value of the training comparison summary group;
[0125] S303. Calculate the first comparison loss value according to the summary difference value of the training comparison summary group and the summary prediction probability of the training comparison summary in the training comparison summary group.
[0126] In the embodiments of the present application, the difference between the training quality evaluation values of the training comparison learning summaries in the training comparison learning summary group can be directly used as the summary difference value of the training comparison learning summary group.
[0127] Optionally, the reference summary of the second training text and each candidate summary can also be sorted according to the training quality evaluation value to obtain a first comparison learning sequence; according to the training quality evaluation value of the training comparison learning summary in the training comparison learning summary group, the maximum quality evaluation value and the minimum quality evaluation value in the first comparison learning sequence, perform a difference calculation on the comparison learning summary group to obtain a summary difference value of the comparison learning summary group.
[0128] In the embodiments of the present application, the training comparison learning summaries in the comparison summary set S^ are sorted in descending order according to the quality evaluation value of the training comparison learning summary to obtain a first comparison learning sequence:
[0129] {S 1 ,…,S i ,…,S n};
[0130] Specifically, the summary difference value of the training contrast learning summary group (S i , S j ) is:
[0131] where η i > η j That is, i < j.
[0132] In the embodiment of the present application, after sorting, the maximum quality evaluation value and the minimum quality evaluation value in the contrast summary set S^ are obtained. Based on the maximum quality evaluation value and the minimum quality evaluation value in the first contrast learning sequence, the difference calculation of the contrast learning summary group can refer to the differences between different summaries, and can further improve the accuracy of the target summary model generated summary.
[0133] Specifically, the electronic device also calculates the training summary prediction probability of each training contrast learning summary in each training contrast learning summary group. Among them, the calculation method of the training summary prediction probability is the same as the calculation method of the summary prediction probability, and the calculation method of the training summary prediction probability is as follows:
[0134] where, if the training contrast learning summary group is (S i , S j ), the training summary prediction probability f(s i ) of the training contrast learning summary S i is:
[0135]
[0136] The training summary prediction probability f(S j ) of the training contrast learning summary S j is:
[0137]
[0138] where D is the first training text.
[0139] Specifically, the electronic device also calculates the first contrast loss value according to the summary difference value of the training contrast learning summary group and the training summary prediction probability of the training contrast learning summary in the training contrast learning summary group.
[0140] Specifically, the first contrast loss function L ctr is as follows:
[0141] L ctr = -∑ i ∑ j>i max(0, f(S j ) - f(S i ) + Δ ij λ);
[0142] Among them, λ is a hyperparameter of the preset control intensity.
[0143] In the embodiment of the present application, the model parameters of the initial summary model are adjusted based on the first contrast loss value until the adjusted initial summary meets the preset conditions, and the target summary model is obtained. The preset conditions mentioned here can be that the first contrast loss value is less than the preset loss threshold and / or the number of model training iterations is greater than or equal to the preset iteration threshold.
[0144] The summary generation method provided by the embodiment of the present application processes the first training text through a pre-trained initial summary model to obtain a candidate summary of the first training text, then constructs a contrast learning summary group with the reference summary of the second training text and the candidate summary of the first training text, and calculates the first contrast loss value according to the summary prediction probability and quality evaluation value of the contrast learning summaries in the contrast learning summary group, so as to adjust the model parameters of the initial summary model based on the first contrast loss value to obtain the target summary model. In the above embodiment, the reference summary of the second training text and the candidate summary of the first training text are used as contrast learning summaries, and the contrast loss is calculated according to the summary prediction probability and quality evaluation value of the contrast learning summaries in the contrast learning summary group to adjust the model parameters of the initial summary model, which precise the differences between the contrast samples, enables the initial summary model to learn better the ability to evaluate the generated candidate summaries during the training process, makes the prediction summary probability of the prediction summary of the original text generated by the target summary model more in line with the actual situation, and avoids the problem that the prediction summary probability of the generated prediction summary is high while the prediction summary is poor in the actual evaluation, thereby improving the accuracy of the summary generated by the target summary model and improving the user experience.
[0145] In some optional embodiments, before obtaining the candidate summary by processing the first training text through the pre-trained initial summary model, the summary generation method provided by the embodiment of the present application further includes:
[0146] Concatenate and segment the third training text and the reference summary of the third training text to obtain a second contrast learning sequence;
[0147] Calculate the contrast loss according to any two training tokens in the second contrast learning sequence to obtain a second contrast loss value;
[0148] Adjust the model parameters of the preset neural network model based on the second contrast loss value to obtain a pre-trained initial summary model.
[0149] The above preset neural network model is a neural network model with an encoder-decoder structure. For details, refer to the above text and will not be elaborated here.
[0150] The obtained third training text and its reference abstract can be concatenated first and then tokenized to obtain multiple training tokens, and the second contrast learning sequence X = {x 1 , x 2 , …, x n} is obtained by sorting according to the positions of the training tokens in the original text; the contrast loss is calculated for any two training tokens in the second contrast learning sequence to obtain a second contrast loss value for adjusting the model parameters in the preset neural network model. The second contrast loss value L CL is as follows:
[0151]
[0152] where ρ ∈ [-1, 1] is a predefined boundary parameter that can adjust the intensity of the contrast loss; is the feature vector of the i-th training token x i , is the feature vector of the j-th training token x j ; is the similarity between the training token x i and the training token x j , which can be text similarity; |x| is the sequence length.
[0153] It can be understood that in the initial stage of training the summary model, in addition to calculating the second contrast loss value, it is also necessary to calculate the maximum likelihood value of the preset neural network model through maximum likelihood estimation, which is specifically described as follows:
[0154] The third training text D is input into the encoder of the preset neural network model, and the reference summary S * of the third training text is input into the decoder of the preset neural network model for maximum likelihood estimation to obtain the first maximum likelihood estimation value L MLE of the preset neural network model:
[0155]
[0156] where θ is the model parameter of the preset neural network model g, is the probability density function of the model parameter, represents a partial reference summary sequence is a predefined start symbol, is the label smoothing probability function, which is specifically shown in the following formula:
[0157]
[0158] where N is the dictionary size and β is the smoothing probability parameter.
[0159] In the embodiments of the present application, the sum of the second contrast loss value and the first maximum likelihood estimate value is used as the total pre-training loss function L in the pre-training stage. The model parameters of the preset neural network model are adjusted through the total pre-training loss function until the preset neural network model meets the preset conditions, and an initial summary model after pre-training is obtained.
[0160] Since the problem of anisotropic distribution will occur during pre-training through maximum likelihood estimation, that is, the similarity between the representations of each token in the model is relatively high, resulting in a relatively concentrated representation space of the model and poor ability to distinguish the differences between tokens, thus affecting the accuracy of the target summary model in generating summaries. In the present application, a second contrast learning sequence is constructed through the third training text and its reference summary, and contrast learning is performed through the training tokens in the second contrast learning sequence. The model will increase the distance between the representations of different tokens, encourage the model to distinguish and learn isotropic token representations, reduce the similarity between tokens, so as to obtain an isotropic identification space, and further improve the accuracy of the target summary model in generating summaries.
[0161] In some alternative embodiments, to coordinate the summary prediction probability and summary quality of the target summary model, the following constraint conditions are added:
[0162]
[0163] Among them, represents the probability of generating a reference summary, represents the probability of generating a non-reference summary, and M() represents the quality evaluation function.
[0164] In some alternative embodiments, adjusting the model parameters of the initial summary model based on the first contrast loss value to obtain the target summary model may further include:
[0165] Based on the first training text and its reference summary, obtaining a third contrast loss value; and based on the first training text and its reference summary, obtaining a second maximum likelihood estimate value; according to the first contrast loss value L ctr 、the third contrast loss value L′ CL 、and the second maximum likelihood estimate value L′ MLE calculate the fine-tuning total loss function L mul , and adjust the model parameters of the initial summary model according to the fine-tuning total loss function L mul to obtain the target summary model.
[0166] Among them, the total loss function L mul is:
[0167] L mul =L′ MLE +L′ CL +γL ctr ;
[0168] Among them, γ is a preset weight value.
[0169] It can be understood that the third contrast loss value L′ CL Referring to the technical solution in the second contrast loss function calculated in the above-mentioned embodiments, the second maximum likelihood estimate value L′ MLE Referring to the technical solution for calculating the first maximum likelihood estimate value in the above-mentioned embodiments, it will not be elaborated here.
[0170] In the embodiments of the present application, the contrast loss and the cross-entropy loss can be complementary. Therefore, the contrast loss is defined at the sequence level, and the cross-entropy loss is defined at the character level to ensure that the model distributes balanced probability mass over the entire sequence, further improving the accuracy of the target neural network model.
[0171] To better implement the abstract generation method provided in the embodiments of the present application, based on the abstract generation method provided in the embodiments of the present application, the embodiments of the present application further provide an abstract model training device. The abstract model training device is applied to an electronic device, such as Figure 7 shown Figure 7 is a schematic structural diagram of an embodiment of the abstract generation device provided in the embodiments of the present application. The abstract generation device 400 provided in the embodiments of the present application includes:
[0172] A text acquisition module 410, configured to acquire the original text to be processed;
[0173] A preprocessing module 420, configured to preprocess the original text to be processed to obtain a target original text;
[0174] An abstract generation module 430, configured to perform abstract generation processing on the target original text through a target abstract model to obtain a target text abstract of the original text to be processed.
[0175] In some alternative embodiments, the abstract generation device also processes the target original text through the target abstract model to obtain candidate abstracts of the target original text;
[0176] Obtain a reference abstract corresponding to the candidate abstract, construct a target contrast learning abstract group according to the reference abstract and the candidate abstract, and calculate the target abstract prediction probability and the target quality evaluation value of each contrast learning abstract in the target contrast learning abstract group;
[0177] Process the candidate abstract according to the target abstract prediction probability and the target quality evaluation value to determine the target text abstract of the original text to be processed.
[0178] In some alternative embodiments, the abstract generation device performs abstract generation processing on the target original text through a target abstract model to obtain a target text abstract of the original text to be processed, including:
[0179] Obtain a predicted word sequence of the target original text, and through the target abstract model, query an abstract word list to obtain initial candidate words whose association degree with a target predicted word in the predicted word sequence is higher than a preset probability value;
[0180] Determine a target candidate word of the target predicted word from the initial candidate words according to the demand evaluation value of the initial candidate words; the demand evaluation value is an evaluation parameter reflecting the relevance between the initial candidate words and the predicted word sequence;
[0181] Update the target candidate word to the predicted word sequence, and generate a candidate abstract whose abstract prediction probability is greater than or equal to a preset abstract threshold based on the updated predicted word sequence, where the abstract prediction probability is the product of the word element prediction probabilities of the predicted words in the candidate abstract.
[0182] In some alternative embodiments, before the abstract generation device determines a target candidate word of the target predicted word from the initial candidate words according to the demand evaluation value of the initial candidate words, it further includes:
[0183] For each of the initial candidate words, obtain the text similarity between the initial candidate word and each predicted word in the predicted word sequence;
[0184] Obtain an average similarity value according to the text similarity between the initial candidate word and the predicted word, and set the average similarity value as the word element difference value of the initial candidate word;
[0185] Obtain the demand evaluation value of the initial candidate word according to the word element difference value and the word element prediction probability of the initial candidate word.
[0186] In some alternative embodiments, the abstract generation device preprocesses the original text to be processed to obtain a target original text, including:
[0187] Perform word segmentation processing on the original text to be processed according to a trained target word segmentation model to obtain a word-segmented original text;
[0188] Filter non-text symbols in the word-segmented original text to obtain a target original text.
[0189] In some alternative embodiments, the abstract generation device processes a first training text through a pre-trained initial abstract model to obtain training candidate abstracts, and constructs a training comparison abstract group according to the training reference abstract of the second training text and each of the training candidate abstracts;
[0190] Calculate the text similarity between each training comparison summary in the training comparison summary group and the reference summary of the first training text, and determine the quality evaluation value of the training comparison summary according to the text similarity;
[0191] Calculate a first contrast loss value according to the training summary prediction probability and the training quality evaluation value of the training comparison summary in the training comparison summary group; the training quality evaluation value reflects the matching degree between the training comparison summary in the training comparison summary group and the reference summary of the first training text;
[0192] Adjust the model parameters of the initial summary model based on the first contrast loss value to obtain a target summary model.
[0193] In some alternative embodiments, the summary generation device calculates a first contrast loss value according to the training summary prediction probability and the training quality evaluation value of the training comparison summary in the training comparison summary group, including:
[0194] For each training comparison summary group, sort the training reference summary of the second training text and each training candidate summary according to the training quality evaluation value to obtain a first contrast learning sequence;
[0195] Perform a difference calculation on the training comparison summary group according to the quality evaluation value of the training comparison summary in the training comparison summary group, the maximum quality evaluation value and the minimum quality evaluation value in the first contrast learning sequence to obtain the sequence difference value of the training comparison summary group;
[0196] Based on the model parameters of the initial summary model and the probability density function of the model parameters, obtain the training summary prediction probability of the training comparison summary under the condition of the first training text;
[0197] Calculate the first contrast loss value according to the summary difference value of the training comparison summary group and the summary prediction probability of the training comparison summary in the training comparison summary group.
[0198] Correspondingly, an embodiment of the present application further provides an electronic device, which can be a terminal, and the terminal can be a terminal device such as a smart phone, a tablet computer, a notebook computer, a touch screen, a game console, a personal computer (PC, Personal Computer), a personal digital assistant (Personal Digital Assistant, PDA), etc. Alternatively, the electronic device can be a server.
[0199] As Figure 8 shown, Figure 8It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. The electronic device 500 includes a processor 501 having one or more processing cores, a memory 502 having one or more computer-readable storage media, and a computer program stored on the memory 502 and executable on the processor. Among them, the processor 501 is electrically connected to the memory 502. Those skilled in the art can understand that the structure of the electronic device shown in the figure does not constitute a limitation on the electronic device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.
[0200] The processor 501 is the control center of the electronic device 500, connecting various parts of the entire electronic device 500 through various interfaces and lines. By running or loading software programs and / or units stored in the memory 502, and calling data stored in the memory 502, it executes various functions of the electronic device 500 and processes data, thereby monitoring the electronic device 500 as a whole. The processor 501 can be a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), etc., and can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present disclosure.
[0201] In the embodiments of the present disclosure, the processor 501 in the electronic device will load the instructions corresponding to the processes of one or more application programs into the memory 502 according to the following steps, and the processor 501 will run the application programs stored in the memory 502 to implement various functions, such as:
[0202] Processing the first training text through a pre-trained initial summary model to obtain candidate summaries;
[0203] Constructing a contrastive learning summary group according to the reference summary of the second training text and each of the candidate summaries;
[0204] Calculating a first contrastive loss value according to the summary prediction probability and quality evaluation value of the contrastive learning summaries in the contrastive learning summary group; the quality evaluation value is used to reflect the matching degree between the contrastive learning summaries in the contrastive learning summary group and the reference summary of the first training text;
[0205] Adjusting the model parameters of the initial summary model based on the first contrastive loss value to obtain a target summary model.
[0206] Another example:
[0207] Obtaining the original text to be processed;
[0208] Processing the original text to be processed through a preset target summary model to obtain a text summary of the original text to be processed;
[0209] The target abstract model is trained by the abstract generation method described in any of the above embodiments.
[0210] Optionally, as Figure 8 shown, the electronic device 500 further includes: a touch display screen 503, a radio frequency circuit 504, an audio circuit 505, an input unit 506, and a power supply 507. Among them, the processor 501 is electrically connected to the touch display screen 503, the radio frequency circuit 504, the audio circuit 505, the input unit 506, and the power supply 507 respectively. Those skilled in the art can understand that Figure 8 the structure of the electronic device shown in
[0211] does not constitute a limitation on the electronic device, and may include more or fewer components than shown, or combine some components, or have different component arrangements.
[0212] The radio frequency circuit 504 can be used to receive and transmit radio frequency signals to establish wireless communication with a network device or other electronic devices, and to receive and transmit signals between the network device or other electronic devices.
[0213] The audio circuit 505 can be used to provide an audio interface between the user and the electronic device through a speaker and a microphone. The audio circuit 505 can transmit the electrical signal converted from the received audio data to the speaker, which converts it into a sound signal for output; on the other hand, the microphone converts the collected sound signal into an electrical signal, which is received by the audio circuit 505 and then converted into audio data. After the audio data is output to the processor 501 for processing, it is transmitted through the radio frequency circuit 504 to, for example, another electronic device, or the audio data is output to the memory 502 for further processing. The audio circuit 505 may also include an earphone jack to provide communication between a peripheral earphone and the electronic device.
[0214] The input unit 506 can be used to receive input digital, character information or user characteristic information (such as fingerprint, iris, facial information, etc.), and to generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function controls.
[0215] The power supply 507 is used to supply power to each component of the electronic device 500. Optionally, the power supply 507 can be logically connected to the processor 501 through a power management system, so as to implement functions such as management of charging, discharging, and power consumption management through the power management system. The power supply 507 may also include any components such as one or more DC or AC power supplies, a recharge system, a power failure detection circuit, a power converter or inverter, and a power status indicator.
[0216] Although Figure 8 not shown in the figure, the electronic device may also include a camera, a sensor, a Wi-Fi module, a Bluetooth module, etc., which will not be elaborated here.
[0217] In the above embodiments, the descriptions of the various embodiments have their own emphases. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0218] Those of ordinary skill in the art can understand that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions. The instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.
[0219] To this end, embodiments of the present disclosure provide a computer-readable storage medium storing multiple computer programs that can be loaded by a processor to execute any one of the abstract generation methods provided by the embodiments of the present disclosure. The computer programs can perform the steps of the following abstract generation method:
[0220] Obtain the original text to be processed;
[0221] Preprocess the original text to be processed to obtain the target original text;
[0222] Perform abstract generation processing on the target original text through a target abstract model to obtain the target text abstract of the original text to be processed.
[0223] For the specific implementation of each of the above operations, reference can be made to the previous embodiments, which will not be elaborated here.
[0224] Among them, the computer-readable storage medium may include: read-only memory (ROM, Read Only Memory), random access memory (RAM, Random Access Memory), magnetic disk or optical disc, etc.
[0225] Since the computer programs stored in the computer-readable storage medium can execute any one of the abstract generation methods provided by the embodiments of the present disclosure, the beneficial effects achievable by any one of the abstract generation methods provided by the embodiments of the present disclosure can be realized. For details, reference can be made to the previous embodiments, which will not be elaborated here.
[0226] In the above embodiments of the abstract generation method, abstract model training device, abstract generation device, computer-readable storage medium, and electronic device, the descriptions of each embodiment have their own focuses. For parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments. Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes and beneficial effects of the above-described speaker tracking device, computer-readable storage medium, electronic device, and their corresponding units can refer to the description of the speaker tracking method in the above embodiments, which will not be elaborated here specifically.
[0227] The above has introduced in detail a method, device, electronic device, computer-readable storage medium for abstract model training and abstract generation provided by the embodiments of the present disclosure. Specific examples are used in this article to elaborate on the principles and implementation manners of the present disclosure. The descriptions of the above embodiments are only used to help understand the method and its core idea of the present disclosure; at the same time, for those skilled in the art, based on the idea of the present disclosure, there will be changes in the specific implementation manners and application scopes. In summary, the content of this specification should not be construed as a limitation to the present disclosure.
Claims
1. A method for generating an abstract, characterized in that, the method for generating an abstract includes: obtaining the original text to be processed; performing preprocessing on the original text to be processed to obtain a target original text; performing abstract generation processing on the target original text through a target abstract model to obtain the target text abstract of the original text to be processed.
2. The method for generating an abstract according to claim 1, characterized in that, the performing abstract generation processing on the target original text through a target abstract model to obtain the target text abstract of the original text to be processed includes: processing the target original text through the target abstract model to obtain candidate abstracts of the target original text; obtaining the reference abstract corresponding to the candidate abstracts, constructing a target contrastive learning abstract group according to the reference abstract and the candidate abstracts, and calculating the target abstract prediction probability and the target quality evaluation value of each contrastive learning abstract in the target contrastive learning abstract group; processing the candidate abstracts according to the target abstract prediction probability and the target quality evaluation value to determine the target text abstract of the original text to be processed.
3. The method for generating an abstract according to claim 2, characterized in that, the performing abstract generation processing on the target original text through a target abstract model to obtain the target text abstract of the original text to be processed includes: obtaining the predicted word sequence of the target original text, querying an abstract word table through the target abstract model, and obtaining initial candidate words whose association degree with the target predicted word in the predicted word sequence is higher than a preset probability value; determining the target candidate word of the target predicted word from the initial candidate words according to the demand evaluation value of the initial candidate words; the demand evaluation value is an evaluation parameter reflecting the relevance between the initial candidate words and the predicted word sequence; updating the target candidate words to the predicted word sequence, and generating candidate abstracts whose abstract prediction probability is greater than or equal to a preset abstract threshold based on the updated predicted word sequence, where the abstract prediction probability is the product of the word element prediction probabilities of the predicted words in the candidate abstracts.
4. The method for generating an abstract according to claim 3, characterized in that, before determining the target candidate word of the target predicted word from the initial candidate words according to the demand evaluation value of the initial candidate words, it further includes: for each of the initial candidate words, obtaining the text similarity between the initial candidate word and each predicted word in the predicted word sequence; obtaining an average similarity value according to the text similarity between the initial candidate word and the predicted word, and setting the average similarity value as the word element difference value of the initial candidate word; obtaining the demand evaluation value of the initial candidate word according to the word element difference value and the word element prediction probability of the initial candidate word.
5. The method for generating an abstract according to any one of claims 1-4, characterized in that, the performing preprocessing on the original text to be processed to obtain a target original text includes: performing word segmentation processing on the original text to be processed according to a trained target word segmentation model to obtain a word-segmented original text; filtering non-text symbols in the word-segmented original text to obtain a target original text.
6. The abstract generation method according to claim 1, wherein, the target abstract model is obtained based on the following steps: Processing the first training text through a pre-trained initial abstract model to obtain training candidate abstracts, and constructing a training comparison abstract group according to the training reference abstract of the second training text and each of the training candidate abstracts; Calculating the text similarity between each training comparison abstract in the training comparison abstract group and the reference abstract of the first training text, and determining the quality evaluation value of the training comparison abstract according to the text similarity; Calculating a first comparison loss value according to the training abstract prediction probability and the training quality evaluation value of the training comparison abstracts in the training comparison abstract group; the training quality evaluation value reflects the matching degree between the training comparison abstracts in the training comparison abstract group and the reference abstract of the first training text; Adjusting the model parameters of the initial abstract model based on the first comparison loss value to obtain the target abstract model.
7. The abstract generation method according to claim 6, wherein, the calculating the first comparison loss value according to the training abstract prediction probability and the training quality evaluation value of the training comparison abstracts in the training comparison abstract group includes: For each training comparison abstract group, sorting the training reference abstract of the second training text and each of the training candidate abstracts according to the training quality evaluation value to obtain a first comparison learning sequence; Performing a difference calculation on the training comparison abstract group according to the quality evaluation value of the training comparison abstracts in the training comparison abstract group, the maximum quality evaluation value and the minimum quality evaluation value in the first comparison learning sequence to obtain the sequence difference value of the training comparison abstract group; Obtaining the training abstract prediction probability of the training comparison abstract under the condition of the first training text based on the model parameters of the initial abstract model and the probability density function of the model parameters; Calculating the first comparison loss value according to the abstract difference value of the training comparison abstract group and the abstract prediction probability of the training comparison abstracts in the training comparison abstract group.
8. An abstract generation device, wherein, the abstract generation device includes: A text acquisition module configured to acquire the original text to be processed; A preprocessing module configured to preprocess the original text to be processed to obtain a target original text; An abstract generation module configured to perform abstract generation processing on the target original text through a target abstract model to obtain the target text abstract of the original text to be processed.
9. An electronic device, wherein, it includes a processor and a memory, and the memory stores multiple instructions; the processor loads the instructions from the memory to execute the steps of the abstract generation method according to any one of claim 1 or the abstract generation method according to claims 1-7.
10. A computer-readable storage medium, wherein, the computer-readable storage medium stores multiple instructions, and the instructions are suitable for being loaded by a processor to execute the steps of the abstract generation method according to any one of claim 1 or the abstract generation method according to claims 1-7.