Generative Summary Error Correction Method for Fact Consistency
By introducing search modules and error correction modules into the generative text summary, the three-classification task and feature-level masking technology are used to solve the problem of fact inconsistency in the generative text summary, efficient and accurate error correction effect is achieved, and the fact consistency of the abstract text is significantly improved.
Patent Information
- Application Number
- CN202210474905.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-04-29
- Publication Date
- 2025-06-27
- Estimated Expiration
- 2042-04-29
AI Technical Summary
The existing generative text summary method is prone to produce summary content that is inconsistent with the original text facts, and it is difficult for the existing technology to effectively modify errors with finer granularity and higher logical levels.
A generative summary error correction method is proposed. By establishing a search module and an error correction module, the search module retrieves evidence information from the source text in the form of a three-class task. The error correction module is constructed by an encoder, a mask and a decoder, and performs feature-level masking and decoding to generate a fact-consistent error correction text.
This method can efficiently locate evidence information from the source text, accurately modify the fine-grained factual inconsistency errors in the text to be corrected, and significantly improve the factual consistency of the content of the summary text.
Smart Images

Figure CN115358215B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing. Specifically, it relates to a generative summary correction method for factual consistency. Background Art
[0002] In the process of the rapid development of the Internet, a huge amount of data content with various forms has been generated. For various types of text data, the Automatic Text Summarization technology can extract the core content of the given corpus and describe the main idea of the original text with a relatively short summary text. This technology is beneficial to compressing the storage space of text data, reducing the storage cost of text data, helping to improve the retrieval efficiency of text data, and has important practical significance and application value for further realizing information integration.
[0003] The existing automatic text summarization technology is divided into extractive text summarization method and generative text summarization method: The extractive text summarization method directly selects important paragraphs, sentences and words in the original text as the summary content, with poor logicality and coherence between sentences and information redundancy; The generative text summarization method constructs the summary content with novel words through a process similar to the human thinking mode, by understanding, paraphrasing and condensing the original text, and its sentence content is more coherent.
[0004] However, in order to retain the abstract ability during training, the generative model based on neural network weakens the restriction on the generated sentences. Therefore, the generative text summarization method based on the generative model is prone to generate summary content inconsistent with the facts in the original text, resulting in factual inconsistency. The existing technology generally reduces factual consistency errors by optimizing the model structure or training method. However, for the summary generation model that has been deployed as a service or has been built, improving factual consistency according to the above methods requires additional human and material costs to reconstruct and train the model. Or it can effectively correct the factual inconsistency errors in the summary content, but it is only limited to the entity and character levels and cannot correct errors with finer granularity and higher logical levels (such as semantic errors, logical errors, grammar errors, etc.).
[0005] Based on the above technical requirements, there is an urgent need for a text processing method to modify and correct the summary text generated by the generative model and inconsistent with the source text. Summary of the Invention
[0006] An embodiment of the present application provides a generative summary error correction method for factual consistency, so as to at least solve the problem of correcting summary texts generated by generative models in related technologies that have inconsistent content with the source text.
[0007] In an embodiment of the present application, a generative summary error correction method for factual consistency is proposed, characterized in that the method includes:
[0008] Establish a retrieval module, which is a pre-trained model fine-tuned through a three-classification task of a data set;
[0009] Establish an error correction module, which is configured to be constructed by an encoder, a masker, and a decoder;
[0010] After preprocessing the summary to be corrected and the corresponding source text, input them in pairs into the retrieval module. The retrieval module, in the form of a three-classification task, retrieves evidence information from the corresponding source text to match each sentence content of the summary to be corrected, so as to obtain a matching set;
[0011] The error correction module maps the matching set to a feature vector E through the encoder, masks the feature vector E in the feature space through the masker to obtain a masked feature E', and the decoder decodes the masked feature E' to generate an error correction text;
[0012] Evaluate the error correction text from two aspects of content and factuality. After calculating the benefit of the error correction text, generate a record and add it to the sample pool;
[0013] Select the N records with the largest sample benefits from the sample pool to calculate the policy gradient loss, guide the parameter learning of the error correction module, and then empty the sample pool;
[0014] Use the learned retrieval module and error correction module to correct the summary to be corrected in the target data.
[0015] In one embodiment, the retrieval module retrieves evidence information from the source text in the form of a three-classification task of "summary sentence - source text sentence" pairs.
[0016] In one embodiment, the feature vector E = [n, m, sq_len, e_dim], where the summary text to be corrected Sum = [c1, c2,..., cn], there are n sentences in total, that is, n represents the number of summary sentences to be corrected), and the corresponding source text Doc = [e1, e2,..., em]) has m sentences in total, that is, m represents the number of corresponding source text sentences), e_dim represents the hidden layer dimension size of the vector representation, and sq_len represents the length of the token sequence after word segmentation for each combination in the input data;
[0017] Through slicing operation, extract the hidden layer representations corresponding to the [CLS] identifiers at the head of each combination in the feature vector E to obtain E cls =[n, m, e_dim]; then, E cls The input is classified into three categories of n*m, and the classification results include support S, neutral N, and contradiction C;
[0018] Take each clause of the abstract text in the target data as the target sentence in turn, and count the classification results of its corresponding combination. If the classification results of all combinations where the target sentence is located are S, then this sentence does not need to be corrected. If there is N or C in the classification results of all combinations where the target sentence is located, then this sentence needs to be corrected.
[0019] In one embodiment, for a target sentence containing a combination with a classification result of N, it is necessary to calculate the similarity between the target sentence and the corresponding source text clause in this type of combination, and take the K source text clauses with the largest similarity as evidence information; finally, form the matching set with all the collected evidence information and the corresponding target sentence, and the first part of each group of matches is the evidence information, and the second part is the abstract sentence to be corrected.
[0020] In one embodiment, the masker performs feature-level masking on the feature vector E in the feature space through a soft mask of the hybrid attention mechanism.
[0021] In one embodiment, when decoding, the decoder generates two types of corrected texts respectively by using two word lookup strategies of "probability selection" and "Softmax selection" based on the original vocabulary distribution of the top layer output.
[0022] In one embodiment, the dataset uses the FEVER dataset.
[0023] In one embodiment, the encoder uses the encoder of the SimCSE-RoBERTa model, and the masker uses the masker based on the Transformer Encoder model.
[0024] Through the embodiments of the present application, the posterior error correction method for the factual consistency of abstract texts can efficiently locate evidence information from the source text and accurately correct fine-grained factual inconsistency errors in the text to be corrected. This method first constructs a retrieval module based on a fine-tuned pre-trained model, and efficiently and accurately retrieves evidence information from the source text to match the abstract text to be corrected in a three-classification manner, making the most of the prior knowledge learned by the pre-trained model during the pre-training process to improve the matching accuracy. Secondly, after mapping the input into a feature representation, the error correction module constructed by this method masks the feature representation at the feature level, finely masking out data noise, so as to ensure that the decoder can decode the corrected text content consistent with the facts of the original text, greatly improving the accuracy of error correction. Compared with traditional posterior correction methods, the present invention can correct more types of errors and greatly improve the factual consistency of the abstract text content. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0026] Figure 1 is a schematic flowchart of an error correction method according to an embodiment of the present application;
[0027] Figure 2 is a flowchart of an alternative error correction method according to an embodiment of the present application;
[0028] Figure 3 is a schematic diagram of a retrieval module according to an embodiment of the present application;
[0029] Figure 4 is a schematic diagram of an error correction module according to an embodiment of the present application;
[0030] Figure 5 is a schematic diagram of the training of an alternative error correction module according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0031] The present application will be described in detail below with reference to the drawings and in conjunction with the embodiments. It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects and do not necessarily need to be used to describe a specific order or sequence.
[0033] Such as Figure 1As shown, in view of the problem that traditional posterior correction methods cannot finely identify and modify inconsistent content in the summary text with the source text, and the training dataset is inconsistent with the data distribution that the model needs to correct during actual application, a generative summary error correction method for factual consistency is proposed. The retrieval module is used to retrieve matching evidence information from the source text for the summary text to be corrected to obtain a matching set. Then, after the error correction module maps the matching set to the feature space, fine-grained masking is performed at the feature level to mask noise errors, and then the decoder decodes the error-corrected text that is factually consistent.
[0034] Specifically, as Figure 1 shown, a generative summary error correction method for factual consistency is proposed, and the method includes:
[0035] S202, establish a retrieval module, where the retrieval module is a pre-trained model fine-tuned through a three-classification task of the dataset;
[0036] S204, establish an error correction module, where the error correction module is configured to be constructed by an encoder, a masker, and a decoder;
[0037] S206, after preprocessing the summary to be corrected and the corresponding source text, input them in pairs into the retrieval module. The retrieval module, in the form of a three-classification task, retrieves evidence information from the corresponding source text to match each sentence content of the summary to be corrected, so as to obtain a matching set;
[0038] S208, the error correction module maps the matching set to a feature vector E through the encoder, masks the feature vector E in the feature space through the masker to obtain a masked feature E', and the decoder decodes the masked feature E' to generate an error-corrected text;
[0039] S210, evaluate the error-corrected text from two aspects of content and factuality, calculate the benefit of the error-corrected text, and generate a record to be added to the sample pool;
[0040] S212, select the N records with the largest sample benefit from the sample pool to calculate the policy gradient loss to guide the parameter learning of the error correction module, and then clear the sample pool;
[0041] S214, use the learned retrieval module and error correction module to correct the summary to be corrected in the target data.
[0042] It should be noted that the retrieval module retrieves evidence information from the source text in the form of a three-classification task of "abstract statement - source text statement" pairs, which can ensure the accuracy and efficiency of retrieval. The text error correction reward is calculated from two aspects: content and factuality, and a record is generated and added to the sample pool. Then, the re-play mechanism is used to select the optimal samples for gradient calculation and parameter learning. While ensuring that the error correction effect of the error correction module gets better with each iteration, it avoids dependence on true value labels.
[0043] In one embodiment, the above method further includes:
[0044] S1. The feature vector E = [n, m, sq_len, e_dim], where the summary text to be corrected Sum = [c1, c2,..., cn], with a total of n sentences, that is, n represents the number of summary sentences to be corrected), and the corresponding source text Doc = [e1, e2,..., em] has a total of m sentences, that is, m represents the number of corresponding source text sentences), e_dim represents the size of the hidden layer of the vector representation, and sq_len represents the length of the token sequence after word segmentation for each combination in the input data;
[0045] S2. Through slicing operations, the hidden layer representation corresponding to the [CLS] identifier at the head of each combination in the feature vector E is taken out to obtain E cls = [n, m, e_dim]; then, E cls is input for n*m three-classification, and the classification results include support S, neutral N, and contradiction C;
[0046] S3. Each clause of the summary text in the target data is used as the target sentence in turn, and the classification results of its corresponding combination are counted. If the classification results of all combinations where the target sentence is located are S, then this sentence does not need error correction. If there is an N or C in the classification results of all combinations where the target sentence is located, then this sentence needs error correction.
[0047] S4. For the target sentence containing a combination with a classification result of N, it is necessary to calculate the similarity between the target sentence and the corresponding source text clause in this type of combination, and take the K source text clauses with the largest similarity as evidence information; finally, all the collected evidence information and the corresponding target sentences form the matching set, and in each group of matches, the first part is the evidence information and the second part is the summary sentence to be corrected.
[0048] It should be noted that the masker performs feature-level masking on the feature vector E in the feature space through the soft mask of the hybrid attention mechanism, and performs feature-level masking on the input vector features in the feature space, which can mask noise more finely, correct more types of errors, and achieve better error correction effects. When decoding, the decoder generates two types of error-corrected texts respectively by adopting two word-searching strategies of "probability selection" and "Softmax selection" based on the original vocabulary distribution output at the top layer. Through the self-criticism policy gradient, it ensures that the benefit "baseline" of the error-corrected texts generated by the error correction module is continuously improved.
[0049] As Figures 2 to 5 shown, the following specific example is used for expansion and explanation:
[0050] Step 10, first, import the SimCSE-RoBERTa model and load the pre-trained weights as the main structure of the retrieval module. Secondly, introduce a three-classifier to form the retrieval module shown in the dashed box in Figure 3 . Then, use the FEVER dataset (this dataset contains approximately 185k instances, each instance consists of a target statement, evidence information from Wikipedia, and a corresponding label, and the label indicates whether the relationship between the target statement and the evidence information is consistent or insufficient to prove) to fine-tune the parameters of the retrieval module according to the three-class loss function shown in formula (1) through a three-classification task (given a pair of sentences, predict whether the relationship between the second sentence and the first sentence is entailment, contradiction, or neutral). Finally, the parameter configuration of the retrieval module is adjusted from θ1 to θ1'.
[0051]
[0052] where hi is the hidden layer representation of the i-th text pair output by the SimCSE-RoBERTa model, is the three-dimensional true value label vector of the i-th text pair (for example, indicates that the relationship of this text pair belongs to the second category, that is, contradiction), and M represents the number of samples randomly sampled from the dataset in one iteration process (that is, batch_size = M).
[0053] Step 20, construct an error correction module with an encoder (part A), a masker (part B), and a decoder (part C), and its structure is as Figure 4As shown by the black dashed box in the figure. The encoder is composed of a pre-trained SimCSE-RoBERTa model, and the masker and decoder are respectively composed of L1 and L2 layer Transformer Encoder structures. In particular, the masker introduces a hybrid attention mechanism on the basis of the multi-head self-attention layer of the original Transformer Encoder to form a multi-head attention mask layer, aiming to control the feature representation mask and the attention direction through soft mask and hard mask respectively; the decoder structure is similar to the masker, and it only uses the hard mask in the controlled multi-head attention layer to control the attention direction and does not use the soft mask.
[0054] Step 30, preprocess the text to be corrected summary Sum = [c1, c2,..., c n and the corresponding source text Doc = [e1, e2,..., e m . Here, c and e are respectively a clause in the text to be corrected summary and the source text, and the text to be corrected summary and its corresponding source text have n and m clauses respectively. First, combine each clause in Sum with each clause in Doc, generating a total of m×n combinations. Secondly, for each combination, add the [CLS] identifier at the starting position of the combination to distinguish each combination, and add the [SEP] identifier between the two clauses within the combination to distinguish the two clauses. Then, use the BPE (Byte-Pair Encoding) text encoding method for the two clauses within the combination to divide each word in the sentence into sub-word units unified in the vocabulary (this unit is called a token). Finally, input the processed data into the retrieval module for further processing. The processed data is as Figure 3 shown in the bottom square. In particular, in the figure, for the sake of simplified representation, c and e are actually token sequences after word segmentation, which need to be distinguished from the concept of c and e representing clauses described above.
[0055] Step 40, as Figure 3As shown in the figure, the retrieval module obtained in step 10 is used to match the evidence information for the abstract to be corrected and the corresponding source text in a three-classification manner. First, the SimCSE-RoBERTa encoder in the retrieval module is used to encode the data output in step 30 into a vector representation E = [n, m, sq_len, e_dim], where e_dim represents the size of the hidden layer dimension of the vector representation, and sq_len represents the length of the token sequence after word segmentation for each combination in the input data (for the convenience of model processing, it is set to a fixed length here, padded if insufficient and truncated if exceeding the length); secondly, through slicing operations, the hidden layer representation corresponding to the [CLS] identifier at the head of each combination in the vector representation E is taken out to obtain E cls = [n, m, e_dim]; then, E cls is input into a three-classifier for classification, and the classification results include S (support), N (neutral), and C (contradiction); then, each clause of the abstract text is used as the target statement in turn, and the classification results of its corresponding combination are counted. If the classification results of all combinations where the target statement is located are S, then this sentence does not need to be corrected. If there is N or C in the classification results of all combinations where the target statement is located, then this sentence needs to be corrected (the clauses of the source text in the combination are the evidence information to be used when correcting); further, for the target statement containing a combination with a classification result of N, it is necessary to calculate the similarity between the target statement and the corresponding source text clause in this type of combination, and take the K source text clauses with the largest similarity as the evidence information; finally, all the collected evidence information and the corresponding target statements form a matching set CE_Map, which is in the form of {<[e2, e5], c2>, <[e1], c5>, ……, <[e4, e m , e7], c n >}, where the first part of each group of matches is the evidence information (in square brackets), and the second part is the abstract sentence to be corrected.
[0056] Step 50, as shown in part A of Figure 4 the figure, the encoder of the error correction module uses the pre-trained SimCSE-RoBERTa model to map the matching set obtained in step 40 into a feature representation E = SimCSE-RoBERTa(CE_Map) in the feature space. In particular, CE_Map also needs to be preprocessed before encoding, including BPE word segmentation, adding the [CLS] identifier at the head, adding the [SEP] identifier between the evidence information and the statement to be modified in each group of matches, and adding the [EOS] identifier at the end of each group of matches.
[0057] Step 60, the masker of the error correction module is as shown in part B of Figure 4 the figure. First, the error correction module generates a hard mask matrix M for the preprocessed matching set obtained in step 50 hardTo control the attention direction of the Multi-Head Attention Mask layer in the control masker and the Controlled Multi-Head Attention layer in the decoder. The element in the i-th row and j-th column of the hard mask matrix represents the attention relationship between the i-th token and the j-th token in each group of matches. As shown in formula (2), when the element is 0, it indicates that attention interaction is allowed between the two tokens, and when the element is negative infinity, it indicates that attention interaction is not allowed between the two tokens.
[0058]
[0059] Secondly, the Multi-Head Attention Mask layer in the Transformer Encoder structure of each layer in the masker is processed according to formulas (3) to (7) through the hybrid attention mechanism (assuming this is the l-th layer), and feature-level masking is performed on the input vector features in the feature space.
[0060]
[0061]
[0062]
[0063] O l =A l (V+padding(M soft )) (6)
[0064] In formula (3), H l-1 represents the hidden layer representation output by the (l-1)-th layer Transformer Encoder structure. In particular, H 0 =E; and are the query weight matrix, key weight matrix, and value weight matrix to be learned respectively; Q, K, and V are the Query matrix, Key matrix, and Value matrix respectively. In formula (4), A l is the attention weight matrix of the Multi-Head Attention Mask layer in the l-th layer Transformer Encoder structure, which is controlled by the hard mask matrix M hard ; d k is the hidden layer dimension size, used to scale the weights so that they are distributed within a reasonable range. In formula (5), M soft is the soft mask matrix, which is used to mask the vector feature representation in the feature space; A adis the Adapter matrix, which aims to scale the soft mask matrix to the same dimension as the summary statement sequence to be corrected through dimensionality reduction, and is obtained by slicing the attention weight matrix A l ; is the Soft-Mask weight matrix to be learned. In formula (6), after padding zeros to Msoft (i.e., the padding function), it is directly added to the Value matrix obtained from formula (3) in the feature space, and then multiplied by the attention weight matrix to obtain the output O of the multi-head attention mask layer l . At this time, the hidden layer representation of the summary statement to be corrected in this output is masked, while the hidden layer representation of the evidence information is not masked
[0065] Finally, the l-th layer Transformer Encoder structure processes the output O of the multi-head attention mask layer through two layer normalizations with residual connections and a position-wise feed forward network l to obtain the hidden layer feature representation H input to the next layer of the masker as processed by formulas (7) and (8) l .
[0066]
[0067]
[0068] where LN(·) represents the layer normalization operation and PFF(·) represents the position-wise feed forward network
[0069] Step 70, as shown in part C of Figure 4 , the decoder is similar to the masker and is composed of L2 layer Transformer Encoder structures, but it does not contain the soft mask operation. Specifically, first, the Controlled Multi-Head Attention layer in the l-th layer decoder processes the hidden layer feature representation output from the previous layer according to formulas (3), (4), and (9)
[0070] O l = A l V (9)
[0071] Second, the output O of the Controlled Multi-Head Attention layer is still processed according to formulas (7) and (8) through two layer normalizations with residual connections and a position-wise feed forward network to obtain the hidden layer feature representation H input to the next layer of the decoder l l .
[0072] Specifically, in the decoder, H0 The feature representation H output by the top layer of the masker (i.e., the L1 layer) L1 .
[0073] Next, the hidden layer feature representation H output by the top layer of the decoder (i.e., the L2 layer) L2 is extended to the same dimension as the vocabulary size to obtain the original vocabulary distribution Obs_dis (Observation distribution).
[0074] Finally, as shown in the upper left corner Figure 5 , word selection and sentence generation operations are performed on Obs_dis according to the vocabulary according to the two strategies of "selecting words and looking up words by probability" and "Softmax looking up words" respectively, and two corrected texts are obtained. The final corresponding vocabulary distributions of the two are the "vocabulary distribution based on probability selection" Act_dis and the "vocabulary distribution based on greedy selection" Greedy_dis. In particular, the "selecting words and looking up words by probability" strategy means that "if there is a vocabulary distribution [0.7, 0.2, 0.1], then even if the probability of a certain word being selected is low (e.g., 0.1), there is still a 0.1 probability of being selected"; on the contrary, when the "Softmax looking up words" strategy is adopted, the word with the highest probability is fixedly selected by the Softmax function (e.g., 0.7).
[0075] Step 80, as shown in the upper right corner Figure 5 , the benefits of the two corrected texts generated in step 70 are comprehensively evaluated in terms of content and factuality. The factual benefit is calculated by formula (10), the content benefit is calculated by formula (11), and the total benefit is calculated by formula (12).
[0076] reward fact = αFactCC(ModSum)+(1 - α)QEGA(ModSum) (10)
[0077] reward content = β1R1 + β2R2+(1 - β1 - β2)RL (11)
[0078] reward = γ·reward content +(1 - γ)·reward fact (12)
[0079] Among them, ModSum is the error-corrected text; FactCC and QAGS are fact consistency evaluation models proposed by Kryscinski et al. and Wang et al. respectively. Here, we calculate the fact consistency of the error-corrected text with the source text as a reference; R1, R2, and RL respectively refer to the Rouge-1, Rouge-2, and Rouge-L metrics of the error-corrected text and the source text; α, β1, β1, and γ are all balance factors ranging from 0 to 1. After experiments, when α = 0.52, β1 = 0.43, β1 = 0.2, and γ = 0.46, the error correction module achieves the optimal result.
[0080] Next, generate a record in the form of record i =<Obs_dis i , Act_dis i , reward i > and add it to the sample pool for subsequent processing. In particular, during the model training process, the training samples of one batch will be used 10 times in one iteration, and a total of 10×batch_size records will be generated in the sample pool.
[0081] Step 90, as Figure 5 shown in the lower right corner, use the re-play mechanism to select the N records with the highest rewards from the sample pool and calculate the self-criticism policy gradient loss through formulas (13) and (14) to train the parameters in the error correction module, and then empty the sample pool. In each iteration, continuously repeat steps 30 to 90 until the gradient loss converges.
[0082]
[0083]
[0084] Among them, reward(·) is the reward function shown in formulas (10) to (12). Formula (14) trains the parameters of the error correction module by minimizing the difference between the "vocabulary distribution selected based on probability" Act_dis and the "vocabulary distribution selected based on greed" Greedy_dis, explores the sentence correction space by the "probability-based word selection and lookup" strategy, and drives the upward shift of the baseline where the "Softmax word lookup" strategy is located.
[0085] Step 100, use the fine-tuned retrieval module and error correction module to correct the summary text to be corrected in the real data according to steps 30 to 70, and generate the corrected summary text content.
[0086] The above are only the preferred embodiments of the present application. It should be noted that for those of ordinary skill in the art, without departing from the principle of the present application, several improvements and refinements can be made, and these improvements and refinements should also be regarded as the protection scope of the present application.
Claims
1. A generative summary text error correction method for factual consistency, characterized in that, The method includes: Establishing a retrieval module, which is a pre-trained model fine-tuned through a three-classification task of a dataset; Establishing an error correction module, which is configured to be constructed by an encoder, a masker, and a decoder; After preprocessing the summary to be error-corrected and the corresponding source text, inputting them in pairs into the retrieval module. The retrieval module retrieves evidence information from the corresponding source text to match each sentence content of the summary to be error-corrected in the form of a three-classification task, so as to obtain a matching set; The error correction module maps the matching set into a feature vector through the encoder E , masks the feature vector in the feature space through the masker E to obtain a masked feature E’ , and the decoder decodes the masked feature E’ to generate error-corrected text; Evaluating the error-corrected text from two aspects of content and factuality, calculating the benefit of the error-corrected text, and generating a record to be added to the sample pool; Select the records with the largest sample returns from the sample pool N to calculate the policy gradient loss and guide the parameter learning of the error correction module. After that, clear the sample pool; Using the learned retrieval module and error correction module to correct the summary to be error-corrected in the target data; among them, The retrieval module retrieves evidence information from the source text in the form of a three-classification task of "summary sentence - source text sentence" pairs; The feature vector E = n , m , sq_len , e_dim , where the summary text to be corrected Sum = [c1, c2, …, cn], and there are n sentences, that is n indicating the number of sentences in the summary text to be corrected and the corresponding source text Doc = [e1, e2, …, em], and there are m sentences, that is m indicating the number of sentences in the corresponding source text, e_dim indicating the size of the hidden layer dimension represented by the vector, sq_len indicating the length of the token sequence after word segmentation for each combination in the input data; Through the slicing operation, the feature vector is extracted E The corresponding hidden layer representation of the [CLS] identifier of each combined header in is obtained E cls =[ n, m, e_dim ]; Next, E cls Input n*m Three categories, the classification results include support S, neutral N and conflict C; Taking each clause of the summary text in the target data as the target sentence in turn, and counting the classification results of the combinations where it is located. If the classification results of all combinations where the target sentence is located are S, then this sentence does not need to be error-corrected. If there is N or C in the classification results of all combinations where the target sentence is located, then this sentence needs to be error-corrected.
2. The method according to claim 1, wherein For the target sentence containing the combination with the classification result of N, it is necessary to calculate the similarity between the target sentence and the corresponding source text clause in this type of combination, and take the K source text clause with the largest similarity as the evidence information; finally, form the matching set with all the collected evidence information and the corresponding target sentence. In each group of matches, the first part is the evidence information and the second part is the abstract sentence to be corrected.
3. The method according to claim 1, characterized in that, The masker performs feature-level masking on the feature vector in the feature space through the soft mask of the hybrid attention mechanism E for feature-level masking.
4. The method according to claim 1, wherein When decoding, the decoder generates two types of error-corrected texts respectively by using two word retrieval strategies of "probability selection" and "Softmax selection" based on the original vocabulary distribution output at the top layer.
5. The method according to claim 1, characterized in that, The dataset uses the FEVER dataset.
6. The method according to claim 1, characterized in that, The encoder uses the encoder of the SimCSE-RoBERTa model, and the masker uses the masker based on the Transformer Encoder model.
Citation Information
Patent Citations
Chinese spelling error correction method based on multi-task learning
CN114065738A
Text processing method and device
CN114357122A