A Long Text Information Stance Detection Method Based on the RoBERTa Model
Through the long text information position detection method based on the RoBERTa model, combined with text cutting, hierarchical attention mechanism and key sentence marking technology, the problems of text length limitation and noise impact in long text information position detection are solved, and higher model accuracy and global information fusion capabilities are achieved.
Patent Information
- Application Number
- CN202210717351.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-06-23
- Publication Date
- 2025-07-01
- Estimated Expiration
- 2042-06-23
AI Technical Summary
The prior art faces the problems of text length limitation, remote attention deficiency and noise influence in the detection of long text information, resulting in the model's deviation and accuracy reduction in global information fusion.
The long text information position detection method based on the RoBERTa model is adopted to break through text length limitations through text cutting technology, and a hierarchical attention mechanism and key sentence marking layer are designed, combining BiLSTM and CRF modules to carry out global information fusion and key evidence sentence marking.
Effectively utilize long text information to improve the model's attention to global information, reduce local information loss, improve model accuracy, and reduce the interference of long text noise on model prediction.
Smart Images

Figure CN115203406B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the fields of deep learning and natural language processing, and particularly to a long text information stance detection method based on the RoBERTa model. Background Art
[0002] The rapid development of mobile Internet and social media has created a good environment for information dissemination. On the one hand, the production and dissemination thresholds of news have been greatly reduced, making the dissemination channels diversified and the content diversified, and people can obtain richer information; on the other hand, the rise of self-media has led to uneven news quality, and various eye-catching false news, reverse news, rumors and other information have spread wantonly, to a certain extent, impacting the influence of official and traditional media. Evidence-based news fact-checking aims to distinguish true and false news and narrow the distance between people and news facts. As a sub-task of news fact-checking, stance detection often involves two data sets: long texts as evidence, and short texts to be checked as statements.
[0003] Stance detection is essentially a text classification problem. Currently, the mainstream methods to solve this problem include:
[0004] (1) Traditional machine learning methods based on models such as Support Vector Machine (SVM), Naive Bayes, etc.;
[0005] (2) Deep learning methods based on models such as Recurrent Neural Network (RNN), Text Convolutional Neural Network (TextCNN), etc.; and
[0006] (3) Deep learning methods based on large-scale pre-trained models such as BERT (Bidirectional Encoder Representations from Transformers).
[0007] Since the stance detection model needs to have the ability to capture semantic-level information, compared with methods (1) and (2), using deep pre-trained models such as BERT can achieve better results in the short text field. However, the length of texts in practical applications is often long, exceeding the maximum text length supported by the BERT model (512 words). As an improved version of the BERT model, the RoBERTa (A Robustly Optimized BERT) model still makes the same limitation on text length. Therefore, the stance detection task in news fact-checking mainly faces the following three challenges:
[0008] (1) When preprocessing the input long text, the truncation method is mostly used, including head truncation, tail truncation, middle truncation, etc., which easily leads to the loss of text information;
[0009] (2) Since the BERT model is based on a multi-layer attention mechanism, the increase in text length leads to insufficient long-range attention, causing the model to deviate in the fusion of global information;
[0010] (3) Documents with too long text contain a large amount of noise, and insufficient long-range attention can cause short-range noise around the key sentence to be assigned a higher weight, ultimately affecting the model accuracy.
[0011] Therefore, designing a stance detection method that can effectively utilize long text information is a technical problem that urgently needs to be solved in the fact-checking task. Summary of the Invention
[0012] The technical problem solved by the present invention is how to make full use of the global information of the evidence document for stance detection of long text information, including (1) breaking through the text length limitation of the RoBERTa model; (2) designing an effective global information fusion method; and (3) designing an effective key evidence sentence marking method.
[0013] The specific technical solution adopted by the present invention is as follows:
[0014] A long text information stance detection method based on the RoBERTa model, which inputs the evidence document and the statement sentence to be detected into a pre-trained stance detection model to predict the truth or falsehood of the statement;
[0015] Among them, the stance detection model consists of an encoder layer, a hierarchical attention mechanism layer, a key sentence marking layer, and a classification layer. The hierarchical attention mechanism layer includes a word-level attention mechanism and a sentence-level attention mechanism. First, the concatenated evidence document and the claim sentence are tokenized and transformed to obtain word indices. The index sequence composed of all word indices is segmented with overlap to obtain a series of index segments, which are input into the encoder layer. In the encoder layer, the RoBERTa model is used to encode each index segment to obtain word vectors. After removing the overlapping parts of the word vectors, they are re-concatenated to obtain the word vector sequence of each sentence in the evidence document and the claim sentence. Then, in the hierarchical attention mechanism layer, all word vectors in the word vector sequence of the claim sentence are averaged and fused to obtain the claim sentence vector, which is used as the query of the hierarchical attention mechanism. The word vector sequences of each clause in the evidence document are weighted and fused by the word-level attention mechanism to obtain clause vectors. The clause vectors of each clause in the evidence document are weighted and fused by the sentence-level attention mechanism to obtain the evidence document vector. At the same time, in the key sentence marking layer, the clause vectors of each clause in the evidence document are marked as key sentences or non-key sentences. The clause vectors of all key sentences are weighted and fused to obtain the weighted average vector of key sentences. Finally, the concatenated weighted average vector of key sentences, the proof document vector, and the claim sentence vector are input into the classification layer to output the classification result of the truth or falsehood of the claim.
[0016] Preferably, the key sentence marking layer of the stance detection model consists of a BiLSTM and a CRF module. The clause vectors of all clauses in the evidence document are input into the BiLSTM module as a sequence, and the scores of each clause belonging to the key sentence or non-key sentence category are output. These scores are input into the CRF module to output the key sentence marking of the clause sequence in the evidence document.
[0017] Preferably, when the key sentences of the evidence document are not marked, based on the semi-supervised learning method of Self-training, the BiLSTM and CRF modules in the key sentence marking layer of the stance detection model are iteratively trained. The key sentence marking in each round of training is predicted and output by the stance detection model obtained in the previous round of training, and only the claim sentence is used as the key sentence in the first round of training.
[0018] Preferably, before the index sequence is input into the RoBERTa model for text encoding, it needs to be segmented into equal lengths to obtain a series of index segments. The length of each segment is the maximum input length supported by the RoBERTa model (i.e., 512 words), and any two adjacent index segments have overlapping parts.
[0019] Preferably, the total loss used in the training of the stance detection model is the weighted sum of the key sentence marking layer loss and the classification layer loss. The loss function used in the key sentence marking layer is the negative log-likelihood loss, and the loss function used in the classification layer is the cross-entropy loss function.
[0020] Preferably, in the word-level attention mechanism, only the word vectors of the evidence document are used as Query, Value, and Key for self-attention fusion to obtain the self-attention sentence vector. At the same time, all word vectors in the word vector sequence of the claim sentence are averaged and fused to obtain the claim sentence vector. The claim sentence vector is used as Query, and the word vector sequences of each clause in the evidence document are used as Key and Value for external attention fusion to obtain the external attention sentence vector.
[0021] Preferably, in the sentence-level attention mechanism, the claim sentence vector, the self-attention sentence vector, and the external attention sentence vector are used as Query, Key, and Value respectively for attention fusion to obtain the evidence document vector.
[0022] Preferably, the RoBERTa model is the RoBERTa-large model, and each word vector, sentence vector, and document vector adopts the default output dimension of the word vector of RoBERTa-large, which is 1024 dimensions.
[0023] Preferably, the classification layer is composed of a fully connected layer and a Softmax layer.
[0024] Preferably, the Softmax layer outputs a multi-classification result representing the true or false degree of the claim.
[0025] Compared with the prior art, the beneficial effects of the present invention are as follows:
[0026] The present invention introduces the RoBERTa model based on text segmentation in the long text information stance detection task, introduces the BiLSTM and CRF modules for marking key evidence sentences, and at the same time introduces the semi-supervised learning method based on Self-training for training the BiLSTM and CRF modules. Compared with the original traditional stance detection technology, the innovative process relying on text segmentation in the present invention solves the problem of the text length limitation of the RoBERTa model, enables the RoBERTa model to pay more attention to global information, and avoids the loss of local information caused by length limitation. Moreover, the key sentence marking relying on the BiLSTM and CRF modules improves the interpretability of the model and at the same time suppresses the interference of long text noise on the final prediction of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] Figure 1 It is a schematic diagram of the constituent modules of the stance detection model;
[0028] Figure 2 Schematic diagram of the network framework of the stance detection model;
[0029] Figure 3 Schematic diagram of the semi-supervised learning method based on Self-training;
[0030] Figure 4 General flowchart of the stance detection method in the embodiment. Specific implementation manners
[0031] To make the above objects, features and advantages of the present invention more obvious and understandable, the following will describe in detail the specific implementation manners of the present invention in conjunction with the accompanying drawings, and elaborate on specific details to facilitate a full understanding of the present invention. However, the present invention can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the connotation of the present invention. Therefore, the present invention is not limited by the specific embodiments disclosed below. Without conflict, the technical features in various embodiments of the present invention can be combined accordingly.
[0032] In a preferred embodiment of the present invention, a method for detecting the stance of long text information based on the RoBERTa model is provided. The evidence document and the claim sentence to be detected are input into a pre-trained stance detection model to predict the truth or falsehood of the claim.
[0033] Among them, the stance detection model is a pre-constructed and trained network model. As Figure 1 shown, the stance detection model consists of an encoder layer, a hierarchical attention mechanism layer, a key sentence marking layer and a classification layer. The hierarchical attention mechanism layer includes a word-level attention mechanism layer and a sentence-level attention mechanism layer. After the evidence document and the claim sentence to be detected are input into the stance detection model, the specific process inside the model is as shown in the following 1) to 4):
[0034] 1) First, after the evidence document and the claim sentence are concatenated, they are tokenized and converted into word indices. The index sequence composed of all word indices is divided by overlapping to form a series of index segments, and then input into the encoder layer. The RoBERTa model encodes each index segment respectively to obtain word vectors, and then the overlapping parts of the obtained word vectors are removed and re-concatenated to obtain the word vector sequence of each sentence in the evidence document and the claim sentence.
[0035] It should be noted that the RoBERTa model is an improved version of the BERT model. Its specific model structure and principle belong to the prior art. For details, please refer to https: / / arxiv.org / abs / 1907.11692, and no further specific introduction will be made here.
[0036] The evidence document of the present invention is a long text, which contains a series of clauses; the statement sentence is a short text, generally containing only one sentence. After the evidence document and the statement sentence are concatenated, it is still a long text. The encoder layer is used to encode the evidence document and the statement sentence to obtain long text word vectors. As a preferred implementation manner of an embodiment of the present invention, before performing text encoding by the RoBERTa model, the text words need to be first converted into inputtable indices. Before the index sequence is input into the RoBERTa model for text encoding, it needs to be equally segmented into a series of index segments (the total number of segments is denoted as N segments), and the length of each segment is the maximum input length supported by RoBERTa, that is, 512, and any two adjacent index segments have an overlapping part.
[0037] As a preferred implementation manner of an embodiment of the present invention, the specific method of the above step 1) is as follows: Concatenate and segment the evidence document and the statement sentence, and then convert them into word indices that can be input into the model. Selectively segment the index sequence composed of all word indices, and the segmentation principle is: the length of each index segment is 512, and two adjacent index segments have an overlapping part. For example: For an index sequence with a length less than 512, no segmentation is performed; for an index sequence with a length between 512 and 1024, it is segmented into two index segments with a length of 512 each that contain an overlapping part; for an index sequence with a length between 1024 and 1536, it is segmented into 3 index segments with a length of 512 each; and when segmenting, any two adjacent index segments contain a common overlapping part. Sequentially send the segmented input index segments into the RoBERTa model for encoding to obtain n segments of word vector encodings. Remove the overlapping parts of the n segments of word vector encodings and concatenate them to obtain all the word vector encodings w of the long text = [w1, w2, w3, …, w i ,w i+1 ,…,w j , where j is the total length of the evidence and the statement sentence, i is the text length of the statement sentence, and i - j + 1 is the length of the evidence document. All the word vector encodings w of the long text can be divided according to each sentence in the evidence document and the statement sentence, and the word vectors of each sentence form a word vector sequence. Therefore, w can be expressed as w = [w 1,1 , w 1,2 , …, w 2,1 ,w 2,2 , …, w n,m , where w n,m represents the m-th word vector in the n-th sentence.
[0038] As a preferred implementation manner of an embodiment of the present invention, the above RoBERTa model can adopt the RoBERTa-large version. Therefore, the word vectors, sentence vectors, and document vectors in the present invention all adopt the default output dimension of the word vectors of RoBERTa-large, that is, 1024 dimensions.
[0039] 2) Then, all the word vectors in the declared word vector sequence are averaged and fused to obtain the declared sentence vector, which is used as the query of the hierarchical attention mechanism. The word vector sequences of each clause in the evidence document are weighted and fused by the word-level attention mechanism to obtain the sentence vectors of each clause in the evidence document. The sentence vectors of each clause in the evidence document are then weighted and fused by the sentence-level attention mechanism to obtain the document vector.
[0040] It should be noted that the hierarchical attention mechanism layer includes attention mechanisms at the word and sentence levels, which are used to obtain the document vector encoding of long texts. Among them, when performing attention fusion at the two levels, the declared sentence vector c used is directly averaged from all the word vectors of the declaration output in the encoder layer. In the hierarchical attention mechanism, first, the word vectors of the evidence document are fed into the word-level attention layer to obtain the attention-weighted sentence vector. The sentence vectors of all clauses in the evidence document are denoted as e=(e1, e1, e3, …, es), where s is the total number of sentences in the evidence document. Preferably, the clauses can be divided according to punctuation marks. After obtaining the sentence vectors of all clauses in the evidence document, each sentence vector is then fed into the sentence-level attention layer to obtain the attention-weighted and fused evidence document vector.
[0041] The specific principle of the attention mechanism belongs to the prior art and includes two categories: self-attention and external attention. As a preferred implementation manner of the embodiment of the present invention, self-attention and external attention are introduced into the word-level attention mechanism. Specifically, only the word vectors of the evidence document are used as Query, Value, and Key for self-attention fusion to obtain the self-attention sentence vector. At the same time, all the word vectors in the declared word vector sequence are averaged and fused to obtain the declared sentence vector c. The declared sentence vector is used as Query, and the word vector sequences of each clause in the evidence document are used as Key and Value for external attention fusion to obtain the external attention sentence vector. Further, in the sentence-level attention mechanism, the declared sentence vector c, the self-attention sentence vector, and the external attention sentence vector are used as Query, Key, and Value respectively for attention fusion to obtain the document vector d.
[0042] 3) At the same time, the sentence vectors of each clause in the evidence document are respectively marked as key sentences or non-key sentences after passing through the key sentence marking layer, and the sentence vectors of all key sentences are weighted and fused to obtain the key sentence weighted average vector s.
[0043] As a preferred implementation manner of the embodiment of the present invention, the above key sentence tagging layer is jointly composed of a BiLSTM module and a CRF module, and is used to tag each sentence vector of the evidence document. The sentence vectors of all clauses in the evidence document are used as a sequence to be input into the BiLSTM module, and scores of two categories, namely key sentences and non-key sentences, of each sentence are output. Then, the scores are input into the CRF module to output the key sentence tagging of the clause sequence in the evidence document. The key sentence tagging process in the CRF module considers all sentence sequences of the evidence document, and the category with the highest score among all combinations of the sentence sequences is used as the tagging of all sentence sequences of the evidence document.
[0044] As a preferred implementation manner of the embodiment of the present invention, when weighted fusion is performed on the sentence vectors of all key sentences, the sentence vectors in the key sentence set I can pass through a common attention layer, so as to be fused into a weighted average vector of the key sentences.
[0045] 4) Finally, the weighted average vector s of the key sentences, the document vector d, and the sentence vector c of the claim are concatenated and then input into the classification layer to output the final classification result of the truth or falsehood of the claim.
[0046] As a preferred implementation manner of the embodiment of the present invention, the above classification layer is composed of a fully connected layer plus a Softmax layer. The concatenated vector formed by concatenating the weighted average vector of the key sentences, the document vector, and the sentence vector of the claim is input into the fully connected layer for a fully connected operation, and then the multi-classification probability of the truth or falsehood of the claim is output through the Softmax layer. According to the multi-classification probability, the true or false label of the claim can be obtained. It should be noted that the Softmax layer can output the binary classification probability that the claim is true and the claim is false, or can output the multi-classification result representing different degrees of truth or falsehood of the claim. For example, the classification result can be set to six types: extremely false (pants-fire), false, barely true, half true, mostly true, true.
[0047] Thus, the overall framework of the stance detection model in the embodiment of the present invention is as Figure 2 shown.
[0048] It should be noted that before the above-mentioned stance detection model is actually applied, it needs to be pre-trained using a training dataset. And it should be particularly noted that if the evidence documents in the training data do not contain key sentence tags, the Self-training method needs to be adopted to iteratively train the BiLSTM and CRF modules in the key sentence tagging layer to automatically generate key sentences. That is to say, during the iterative training process of the entire stance detection model, the BiLSTM and CRF modules in the key sentence tagging layer are iteratively trained based on the semi-supervised learning method of Self-training. The key sentence tags for each round of training are predicted and output by the stance detection model obtained from the previous round of training, and only declarative sentences are used as key sentences in the first round of training. The specific Self-training is as Figure 3 shown. The set of all sentences in the evidence document is denoted as D, the set of key sentences is denoted as I, and the set of unlabeled sentences is denoted as N. The initial evidence document only contains one declarative key sentence, denoted as 1, and the rest are non-key sentences, denoted as 0. Before the start of the second round of training, use the stance detection model trained in the first round to predict all sentences in N, and select the clause with the highest probability of the output value being 1 and greater than the threshold α, and add it to the set of key sentences I. For each round, the result of predicting the clauses by the stance detection model obtained from the previous round of training is used as the true value of the key sentences for the next round of training; the iteration ends until the model reaches the expected effect or the probability that all unlabeled sentences are predicted as key sentences is less than the threshold.
[0049] It should be noted that the Self-training of the BiLSTM and CRF modules in the key sentence tagging layer is not independent, but is iteratively trained along with the entire stance detection model. During one round of iterative training of the entire stance detection model, not only the network parameters of the BiLSTM and CRF modules in the key sentence tagging layer need to be updated, but also the network parameters of other learnable parts in modules such as the RoBERTa model and the fully connected network need to be updated. However, since the key sentences in the evidence document are not labeled, the set of key sentences needs to be updated round by round through Self-training during the iteration process. Therefore, as a preferred implementation manner of the embodiment of the present invention, the total loss adopted for training the above-mentioned stance detection model is the weighted sum of the loss of the key sentence tagging layer and the loss of the classification layer, where the loss function adopted by the key sentence tagging layer is the negative log-likelihood loss, and the loss function adopted by the classification layer is the cross-entropy loss function. The training of the entire stance detection model ends when the maximum number of iterative rounds is reached or the evidence key sentences do not increase and the classification accuracy of the validation set does not rise, and the finally obtained stance detection model can be used for actual applications.
[0050] Next, the long text information stance detection method based on the RoBERTa model shown in the above embodiment will be applied to a specific example to demonstrate its specific implementation and technical effects.
[0051] Embodiment
[0052] This embodiment includes two datasets: the evidence set E and the claim set C. The evidence set is document-level text with a length between 200 and 1500 words. The claim set is labeled sentence-level short text, which is sourced from the speeches of political figures or user posts on social media such as Facebook, with a length within 30 words. Both the evidence set and the claim set are from the Politifact website. Politifact is a website that fact-checks international political news manually. Reviewers collect evidence and give judgment results, which are used as the labels of the claim set, including six types: Pants on Fire, False, Barely True, Half True, Mostly True, True.
[0053] It should be noted that the evidence and claims in the original dataset are not one-to-one matched and need to be matched in the data preprocessing stage.
[0054] The overall process of implementing long text information stance detection in this embodiment is as Figure 4 shown as follows:
[0055] 1. Data preprocessing stage
[0056] This invention involves two datasets: the evidence set E' of long text and the claim set C' of short text. Among them, the claim set C' is labeled data.
[0057] Step 1, perform text preprocessing on the evidence set E' and the claim set C', removing information irrelevant to stance detection such as website addresses, to obtain the preprocessed evidence set E and claim set C.
[0058] Step 2, for each short text claim c_i in the claim set C, use the BM25 algorithm to search the evidence set E to match the most relevant piece of evidence e_i. c_i and e_i together form a training data, and all the training data samples form the training dataset T.
[0059] 2. Model construction
[0060] Construct a stance detection model. It includes an encoder layer, a hierarchical attention mechanism layer, a key sentence marking layer, and a classification layer. Among them, the hierarchical attention mechanism layer includes a word-level attention layer and a sentence-level attention layer.
[0061] The specific construction process is as follows:
[0062] Step 1, RoBERTa is an optimized version of the BERT model for obtaining word vectors. The RoBERTa-large version is adopted, which can encode words into 1024-dimensional word vectors. The maximum length of the document is set to 1536 words. The Batch_sizes is set to 16, and each batch of documents is encoded into a tensor of [16, 1536, 1024] dimensions through the RoBERTa layer.
[0063] Step 2, the hierarchical attention mechanism layer includes: the self-attention mechanism at the word level and the self-attention mechanism at the sentence level, which are used to obtain the attention-weighted sentence representation and document representation respectively. The word-level attention formula is as follows: s i = ∑ j α i,j w i,i
[0064] where: the value of α i,j is obtained by the linear layer, and w i,j is the j-th word in the i-th sentence. S i represents the sentence vector of the i-th sentence.
[0065] The sentence-level attention is implemented as: h i = Σ j β i,j w i,j , d = ∑ i λ i h i
[0066] where β i,j is proportional to exp(S(c, w i,j )) and λ i is proportional to exp(S′(c, h i )); S and S′ are bilinear scoring functions, c is the statement sentence vector representation, and d is the document vector representation.
[0067] Step 3, select the BiLSTM and CRF modules for marking key sentences. BiLSTM consists of two LSTMs: one receives forward input and the other receives backward input. BiLSTM improves the way the algorithm utilizes context and effectively increases the amount of information available to the network. CRF is a discriminative module. Let y be the tagging sequence and x be the sentence vector sequence, and find the most likely tagging sequence:
[0068] P(y|x) = exp(Score(x, y)) / Σexp(Score(x, y′))
[0069] where: Score represents the function for calculating the score;
[0070] Step 4: Select the fully connected layer and the Softmax function as the classifier. Concatenate the weighted average vector s of the key sentences, the document vector d, and the sentence vector c of the claim, and then input the concatenated vector into the classifier to output the classification probability of the concatenated vector.
[0071] The finally obtained stance detection model is as Figure 2 shown. The data processing flow between each module is as described above and will not be elaborated here.
[0072] 3. Model Training
[0073] Step 1: Batch the training dataset T according to a fixed batch size, with a total of N*.
[0074] Step 2: Perform the following operations on all the training data within a batch: Split them into several sentences according to punctuation marks, obtain the start and end positions of each clause, and do not save the actual data formed by the splitting, but only record the starting position of each sentence.
[0075] Step 3: Divide all the sentence sets of the evidence document into the key sentence set I and the unlabeled set N. Initially, I only contains the claim sentences.
[0076] Step 4: Set the total loss of the stance detection model as the weighted sum of the key sentence marking layer loss and the classification layer loss. The loss function adopted by the key sentence marking layer is the negative log-likelihood loss function, and the loss function adopted by the classification layer is the cross-entropy loss function. Input the training data into the stance detection model for training. At the end of each round of training, use the obtained model to predict the unlabeled sentence set N, select the sentence with the highest predicted probability of being a key sentence and greater than the threshold α, delete it from the unlabeled set N, and add it to the key sentence set I.
[0077] Step 5: Use the key sentence set I as the pseudo-label for a new round of sentence marking task to train the stance detection model.
[0078] Step 6: Continuously iterate and train the stance detection model until the maximum number of iteration rounds is reached, or no new evidence sentences appear, or the classification prediction accuracy of the validation set does not increase, and then end the training to obtain the final stance detection model.
[0079] 4. Applying the Model for Stance Prediction
[0080] Use the evidence documents and claim sentences of the test set and the evidence set as the input to the trained stance detection model. Finally, the Softmax layer predicts the probability that the claim belongs to the true or false category, thereby realizing text stance detection.
[0081] In an example of applying this model for stance detection, the key sentence markings obtained by the BiLSTM and CRF modules are as follows:
[0082] Table 1
[0083] Sentence order 0 1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 Marker 1 1 0 0 0 1 1 1 0 1 0 1 1 1 1 1
[0084] The final prediction result is: 2, that is, barely-true. The prediction result is consistent with the true value. It can be seen that when this method conducts stance detection, it can give accurate results while marking the corresponding key evidence sentences, reducing the noise impact of long documents and providing interpretability for model inference. The innovative process relying on text segmentation solves the problem of the text length limitation of the RoBERTa model, enabling the RoBERTa model to pay more attention to global information and avoiding the loss of local information caused by length limitation.
[0085] The above-described embodiments are only a preferred solution of the present invention, but they are not intended to limit the present invention. Those of ordinary skill in the relevant technical fields can still make various changes and modifications without departing from the spirit and scope of the present invention. Therefore, all technical solutions obtained by means of equivalent replacement or equivalent transformation fall within the protection scope of the present invention.
Claims
1. A long text information stance detection method based on the RoBERTa model, characterized in that: Input the evidence document and the claim sentence to be detected into a pre-trained stance detection model to predict the truth or falsehood of the claim; Among them, the stance detection model consists of an encoder layer, a hierarchical attention mechanism layer, a key sentence marking layer, and a classification layer. The hierarchical attention mechanism layer includes a word-level attention mechanism and a sentence-level attention mechanism. First, tokenize and transform the concatenated evidence document and claim sentence to obtain word indices. Perform overlapping segmentation on the index sequence composed of all word indices to obtain a series of index segments, and input them into the encoder layer. In the encoder layer, use the RoBERTa model to encode each index segment to obtain word vectors. After removing the overlapping parts of the word vectors and re-concatenating them, obtain the word vector sequence of each sentence in the evidence document and the claim sentence. Then, in the hierarchical attention mechanism layer, average and fuse all the word vectors in the word vector sequence of the claim sentence to obtain the claim sentence vector, which is used as the query of the hierarchical attention mechanism. Perform word-level attention mechanism weighted fusion on the word vector sequence of each clause in the evidence document to obtain the sentence vector. Perform sentence-level attention mechanism weighted fusion on the sentence vectors of each clause in the evidence document to obtain the evidence document vector. At the same time, in the key sentence marking layer, mark the sentence vector of each clause in the evidence document as a key sentence or a non-key sentence. Perform weighted fusion on the sentence vectors of all key sentences to obtain the key sentence weighted average vector. Finally, input the concatenated key sentence weighted average vector, the proof document vector, and the claim sentence vector into the classification layer to output the classification result of the truth or falsehood of the claim.
2. The long text information stance detection method based on the RoBERTa model according to claim 1, characterized in that The key sentence marking layer of the stance detection model consists of a BiLSTM and a CRF module. Take the sentence vectors of all clauses in the evidence document as a sequence and input them into the BiLSTM module to output the scores of each clause belonging to the key sentence or non-key sentence category. Input the scores into the CRF module to output the key sentence marking of the clause sequence in the evidence document.
3. The long text information stance detection method based on the RoBERTa model according to claim 2, characterized in that, When the key sentences of the evidence document are not marked, based on the semi-supervised learning method of Self-training, iteratively train the BiLSTM and CRF modules in the key sentence marking layer of the stance detection model; The key sentence marking for each round of training is predicted and output by the stance detection model obtained from the previous round of training, and only the claim sentence is used as the key sentence in the first round of training.
4. The long text information stance detection method based on the RoBERTa model according to claim 1, wherein, Before the index sequence is input into the RoBERTa model for text encoding, it needs to be segmented into equal lengths to obtain a series of index segments. The length of each segment is the maximum input length supported by the RoBERTa model, that is, 512 words, and any two adjacent index segments have an overlapping part.
5. The long text information stance detection method based on the RoBERTa model according to claim 1, characterized in that The total loss used for training the stance detection model is the weighted sum of the key sentence marking layer loss and the classification layer loss. The loss function used in the key sentence marking layer is the negative log-likelihood loss, and the loss function used in the classification layer is the cross-entropy loss function.
6. The long text information stance detection method based on the RoBERTa model according to claim 1, wherein, In the word-level attention mechanism, only the word vectors of the evidence document are used as Query, Value, and Key for self-attention fusion to obtain a self-attention sentence vector. At the same time, all word vectors in the word vector sequence of the claim sentence are averaged and fused to obtain a claim sentence vector. Using the claim sentence vector as Query, and the word vector sequences of each clause in the evidence document as Key and Value, external attention fusion is performed to obtain an external attention sentence vector.
7. The long text information stance detection method based on the RoBERTa model according to claim 6, wherein, In the sentence-level attention mechanism, the claim sentence vector, the self-attention sentence vector, and the external attention sentence vector are used as Query, Key, and Value respectively for attention fusion to obtain an evidence document vector.
8. The long text information stance detection method based on the RoBERTa model according to claim 1, wherein, The RoBERTa model is the RoBERTa-large model, and each word vector, sentence vector, and document vector adopts the default output dimension of the word vector of RoBERTa-large, which is 1024 dimensions.
9. The long text information stance detection method based on the RoBERTa model according to claim 1, wherein The classification layer consists of a fully connected layer and a Softmax layer.
10. The long text information stance detection method based on the RoBERTa model according to claim 9, wherein The Softmax layer outputs a multi-classification result representing the true or false degree of the claim.
Citation Information
Patent Citations
Public opinion detection method, device and equipment for news event, and storage medium
CN113282754A
Media false news detection method and system based on long text feature extraction optimization
CN113704473A