Three-classification hybrid transfer learning method and system for information epidemic situation unreal information discrimination
Through the three-class hybrid transfer learning method, combined with pre-training and fine-tuning models, a variety of deep learning models are used to process the "information epidemic" data, solving the problems of low accuracy and overfitting of false information in the existing technology, and achieving more efficient and accurate information screening.
Patent Information
- Application Number
- CN202510144247.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-10
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-02-10
AI Technical Summary
The existing technology fails to fully utilize the characteristics of health information and the professional knowledge of health workers in the "information epidemic", resulting in low accuracy in identifying false information, and complex models are prone to overfitting when handling such tasks.
The three-class mixed transfer learning method is adopted, and the combination of pre-training models and fine-tuning models is used to pre-process and feature extraction of "information epidemic" related data through the combination of pre-training models and fine-tuning models. The subdivided data is "undetermined", "false" or "true".
It improves the accuracy and practical application effect of false information screening, reduces the phenomenon of overfitting, and enhances the generalization ability and flexibility of the model.
Smart Images

Figure CN120067415A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of information management, and in particular to a three-class hybrid transfer learning method and system for identifying false information in the context of "infodemic". Background Art
[0002] The term "Infodemic" is composed of the root words of "Information" and "Epidemic", referring to the phenomenon of excessive spread of a large amount of information in the context of an infectious disease epidemic. During an infectious disease epidemic, the public's demand for health information surges, making health information an important part of the "Infodemic".
[0003] (Epidemic), and it refers to the phenomenon of excessive spread of a large amount of information in the context of an infectious disease epidemic. During an infectious disease epidemic, the public's demand for health information surges, making health information an important part of the "Infodemic".
[0004] Xu Nuo et al. proposed a health rumor detection method based on I-BERT-BiLSTM. By extracting the abstract of document-level long-sequence text and inputting it into a deep neural network with a multi-layer attention mechanism as the framework for feature extraction, and finally inputting it into BiLSTM for rumor classification. The experimental results show that the model achieved a high accuracy on a self-built Chinese health rumor dataset containing 2000 health rumor data and 2000 health non-rumor data.
[0005] Wang Xiaoyan et al. mined 18 features of WeChat health rumors in terms of text rhetoric, posting motivation, publishing norms, and editing layout. Using a deep neural network model to obtain pre-classification results, which were used as deep semantic feature representations to expand the artificially constructed feature set, and combined with a machine learning model to complete the detection task of health rumors. The experimental results verified the effectiveness and innovation of this method in the detection of WeChat health rumors. The dataset used included 1092 real articles and 1015 rumor articles.
[0006] Hussain et al. proposed a cascaded group multi-head attention model for COVID-19 fake news detection. The model combines deep convolution and cascaded multi-head attention mechanisms to capture information in local and global contexts through multi-head attention, achieving a comprehensive understanding of fake news. On a Urdu Twitter dataset containing accurate tweets (942) and false tweets (2676), the classification effect of this model exceeded that of the current state-of-the-art models.
[0007] Xia et al. proposed a fake news detection framework based on complex adaptive system theory and information dissemination theory. This framework extracts and integrates abnormal knowledge of COVID-19 fake news, constructs a knowledge base, and applies it to a hybrid model based on CNN-BiLSTM-AM for fake news detection. The application of this model on an English dataset (including 5,600 true news and 5,100 fake news) significantly improves the performance of various evaluation metrics.
[0008] The above research mainly focuses on classifying "infodemic" into true and false categories, and does not fully utilize the characteristics of health information in "infodemic". The professional knowledge of health workers in the field of health information can provide effective support for the identification of false information in "infodemic". Specifically, the false information in "infodemic" can be further divided into three categories: the first category is information that health workers cannot independently judge as true or false; the second category is information that health workers can independently judge as false; the third category is information that health workers can independently judge as true.
[0009] Luo et al. proposed a three-class detection method based on a deep learning model for the complexity of distinguishing true and false information in the COVID-19 "infodemic". This study combines the professional knowledge of health workers in the field of health information and classifies "infodemic" into three categories: true, false, and uncertain. It uses fastText, three RNN-based models, two CNN-based models, and two Transformer-based models to conduct classification experiments on Chinese and English data. The research results show that due to the limited data volume of the three-class "infodemic" task, models with pre-training or simpler architectures perform better in terms of performance than complex models, and complex models are prone to overfitting when dealing with such problems.
[0010] Less research focuses on using the characteristics of health information in "infodemic" and the professional knowledge of health workers in the field of health information to provide effective support for identifying false information in "infodemic". In addition, the three-class "infodemic" task is limited by insufficient data volume, it is difficult to improve the model accuracy, and complex models are more likely to overfit when dealing with such tasks.
[0011] Classifying "infodemic" into true and false categories fails to fully utilize the characteristics of health information in "infodemic" and also fails to give full play to the professional knowledge of health workers in the field of health information, resulting in a relatively low actual accuracy rate for identifying false information in "infodemic" and limiting its actual application effect.
[0012] The three-class "infodemic" task that integrates the characteristics of health information and the professional knowledge of health workers in the field of health information is limited by insufficient data volume, it is difficult to improve the model accuracy, and complex models are more likely to overfit when dealing with such tasks. Summary of the Invention
[0013] The technical problem to be solved by the present invention is to provide a three-class hybrid transfer learning method and system for identifying false information in "infodemic", which can more accurately identify false information in "infodemic" and improve the practical application effect of the method.
[0014] To solve the above technical problem, the technical solution of the present invention is as follows:
[0015] In a first aspect, a three-class hybrid transfer learning system for identifying false information in "infodemic", the system includes: a pre-trained model and a fine-tuning model;
[0016] Using TF-IDF technology to preprocess data related to "infodemic", and extracting key feature words therefrom;
[0017] Through the BERT model, generating word embedding representations for the extracted "infodemic" key words and conventional false information;
[0018] Fusing the word embedding of the "infodemic" key words with the BERT word embedding and its text output of the conventional false information to obtain the fused data; using the fused data and performing binary classification training according to the conventional "false" or "true" labels to construct a pre-trained model;
[0019] In the fine-tuning model, the "undetermined" category is used to classify records that cannot be clearly determined as "false" or "true"; using the text output that fuses the "infodemic" key words and conventional false information generated by the BERT layer of the pre-trained model;
[0020] Combining the BERT model, the TextCNN model, and the fastText model to process the input features of the BERT model, training the processed features using the fine-tuning model, and dividing the data related to "infodemic" into three categories: "undetermined", "false", or "true".
[0021] Further, using TF-IDF technology to preprocess data related to "infodemic", and extracting key feature words therefrom, including:
[0022] Using TF-IDF to preprocess "infodemic" records, and extracting the top 10% key words, where "infodemic" is represented as D = {d 1 , d 2 , …, d n}, and each d i represents an "infodemic" record;
[0023] For each record d iPerform word segmentation. For Chinese texts, use the Jieba word segmentation tool; for English texts, use regular expressions for word segmentation; use TF-IDF to process the "infodemic" records after word segmentation, and calculate the TF-IDF score of each term t in document d i as follows:
[0024] TF-IDF(t, d i ) = TF(t, d i ) × IDF(t);
[0025] where, N represents the total number of records, |{d: t ∈ d}| represents the number of records containing term t. After obtaining the TF-IDF scores of each term in all "infodemic" records D, select the top 10% of the terms with the highest TF-IDF values as the "infodemic" keywords, and the keywords provide the most relevant information about "infodemic".
[0026] Furthermore, through the BERT model, generate word embedding representations for the extracted "infodemic" keywords and regular misinformation, including:
[0027] The set of "infodemic" keywords is represented as X Healthwords = {t 1 , t 2 , …, t k}, where k represents the number of "infodemic" keywords, and each keyword t i is processed by the BERT model to generate the corresponding word embedding representation e i as follows:
[0028] E BERT-Healthwords = BERT embedding (X Healthwords ) = [e 1 , e 2 , …, e k ;
[0029] where, BERT embedding represents the operation of generating word embeddings using the BERT model, and E BERT-Healthwords represents the result of generating word embeddings using the BERT model.
[0030] X Gen,j represents the token sequence of the j-th instance of regular misinformation. Process X Gen,j using the BERT model to generate the word embedding representation of X Gen,j as follows;
[0031] E BERT-Gen,j = BERT embedding (XGen,j )
[0032] Among them, BERT embedding represents the operation of generating word embeddings using the BERT model, and E BERT-Gen,j represents the result of generating word embeddings using the BERT model.
[0033] Fuse the word embeddings of conventional misinformation and "infodemic" keywords to create a unified representation. Concatenate the word embedding E of conventional misinformation BERT-Gen,j with the word embedding e of each "infodemic" keyword i , and flatten the result into a single vector. The calculation formula is as follows:
[0034]
[0035] Among them, concat represents the concatenation operation, represents the concatenation method, Flatten represents the flattening operation, and F j represents the fusion of the word embeddings of conventional misinformation and "infodemic" keywords.
[0036] Furthermore, use the fused data and perform binary classification training according to the conventional "false" or "true" labels to construct a pre-trained model, including:
[0037] Process X using the BERT model Gen,j , generate the text output of X Gen,j . This text output includes an embedding vector with context information and a pooled output. Obtain the pooled output of X Gen,j . The calculation formula is as follows:
[0038] pooled_output j = BERT pooled (X Gen,j )
[0039] Among them, BERT pooled represents the operation of generating the pooled output using the BERT model, and pooled_output j represents the result of generating the pooled output using the BERT model.
[0040] Combine pooled_output j and F j , process the fused feature vector through a fully connected layer, and perform binary classification prediction on the conventional misinformation integrated with "infodemic" keywords. The calculation formula is as follows:
[0041]
[0042] Among them, fc represents the fully connected layer, and σ is the sigmoid activation function. It is the probability of predicting to belong to the "true" category in the binary classification task.
[0043] Furthermore, in the fine-tuning model, the "undetermined" category is used to classify records that cannot be clearly determined as "false" or "true"; the text output that combines the keywords of "infodemic" and conventional misinformation generated by the BERT layer of the pre-trained model includes:
[0044] Access the text output of conventional misinformation generated by the BERT layer in the pre-trained model to obtain a batch of embedding vectors and pooled outputs of the "infodemic" text containing context information. The calculation formula is as follows:
[0045]
[0046] Among them, X Health represents the token sequence of the input batch of "infodemic" instances, and BERT pooled represents the operation of generating a pooled output using the BERT model. pooled_output represents the result of generating a pooled output using the BERT model. P represents that the result of generating a pooled output by the BERT model is a matrix, and the dimension of this matrix is is the set of real numbers, B is the batch size, L is the sequence length, and BERT encoded represents the operation of generating contextually informed embedding vectors using the BERT model. contextualized_token_embedding represents the result of generating contextually informed embedding vectors using the BERT model. H represents that the result of generating contextually informed embedding vectors by the BERT model is a matrix, and the dimension of this matrix is H is the hidden layer dimension;
[0047] The pooled output is processed by the dropout layer, and the contextually informed embedding vectors are processed by the TextCNN pooling layer and the fastText pooling layer. The calculation formula is as follows:
[0048] P drop = Dropout(P);
[0049] C 1 = AdaptiveMaxPool1D(Conv1D(H, k 1 ));
[0050] C 2 = AdaptiveMaxPool1D(Conv1D(H, k 2 ));
[0051] F mean = AdaptiveAvgPool1D(H);
[0052] F max = AdaptiveMaxPool1D(H);
[0053] where P drop represents the result after applying dropout regularization to the pooling output; C 1 and C 2 are the outputs of the TextCNN pooling layer, which performs convolution operations with kernel sizes of k 1 and k 2 respectively on the embedding vectors containing context information, performs adaptive max pooling, Conv1D represents a one-dimensional convolution operation, and AdaptiveMaxPool1D represents an adaptive max pooling operation; F mean and F max are the outputs of the fastText pooling layer, which performs adaptive average pooling and adaptive max pooling respectively on the embedding vectors containing context information, and AdaptiveAvgPool1D represents an adaptive average pooling operation.
[0054] Furthermore, by combining the BERT model, the TextCNN model, and the fastText model, the input features of the BERT model are processed, and the processed features are trained using a fine-tuning model. The data related to "infodemic" is classified into three categories: "undetermined", "false", or "true", including:
[0055] Concatenate the processing results of the dropout layer, the TextCNN pooling layer, and the fastText pooling layer to obtain the final feature representation, and its calculation formula is as follows:
[0056] F = [P drop , C 1 , C 2 , F mean , F max ;
[0057] where F represents the concatenated feature vector;
[0058] Pass the concatenated feature vector to the fully connected layer to perform a three-class prediction on the "infodemic" records, and its calculation formula is as follows:
[0059]
[0060] where φ represents the softmax activation function, represents the predicted probability distribution of each category in the three-class task.
[0061] In a second aspect, a three-class hybrid transfer learning method for identifying misinformation in "infodemic" includes:
[0062] Preprocess the data related to "infodemic" using TF-IDF technology to extract key feature words from it;
[0063] Generate word embedding representations for the extracted "infodemic" keywords and conventional misinformation through a BERT model;
[0064] Fuse the word embeddings of the "infodemic" keywords with the BERT word embeddings of the conventional misinformation and their text outputs to obtain the fused data; Use the fused data and perform binary classification training according to the conventional "false" or "true" labels to construct a pre-trained model;
[0065] In the fine-tuning model, use the "undetermined" category to classify records that cannot be clearly determined as "false" or "true"; Use the text output that fuses the "infodemic" keywords and conventional misinformation generated by the BERT layer of the pre-trained model;
[0066] Combine the BERT model, TextCNN model, and fastText model to process the input features of the BERT model, use the fine-tuning model to train the processed features, and divide the data related to "infodemic" into three categories: "undetermined", "false", or "true".
[0067] In a third aspect, a computing device includes:
[0068] One or more processors;
[0069] A storage device for storing one or more programs, which when executed by the one or more processors cause the one or more processors to implement the described method.
[0070] In a fourth aspect, a computer-readable storage medium stores a program that, when executed by a processor, implements the described method.
[0071] The above solutions of the present invention have at least the following beneficial effects:
[0072] 1. Utilize the professional knowledge of health workers in the field of health information, label "infodemic" into three categories, and develop a fine-tuning model by combining BERT, TextCNN, and fastText to classify "infodemic" into undetermined, false, or true categories, increasing the practical application effect of the method.
[0073] 2. Effectively utilize the conventional misinformation marked as false or true, combined with the keyword "infodemic", to design a pre-trained model. By incorporating a large amount of conventional misinformation with the keyword "infodemic", improve the accuracy and generalization ability of the classification method, and alleviate the overfitting phenomenon of the model in such tasks. Description of the Drawings
[0074] Figure 1 It is a schematic diagram of a three-class hybrid transfer learning system for identifying misinformation related to "infodemic" provided by an embodiment of the present invention.
[0075] Figure 2 It is a schematic diagram of the confusion matrix of the three-class task of identifying misinformation related to "infodemic" in Chinese of the present invention.
[0076] Figure 3 It is a schematic diagram of the confusion matrix of the three-class task of identifying misinformation related to "infodemic" in English of the present invention. Detailed Embodiment
[0077] Hereinafter, exemplary embodiments of the present disclosure will be described in more detail with reference to the drawings. Although the exemplary embodiments of the present disclosure are shown in the drawings, it should be understood that the present disclosure can be implemented in various forms and should not be limited by the embodiments set forth herein. On the contrary, these embodiments are provided so that the present disclosure can be more thoroughly understood and the scope of the present disclosure can be fully conveyed to those skilled in the art.
[0078] As Figure 1 shown, an embodiment of the present invention proposes a three-class hybrid transfer learning system for identifying misinformation related to "infodemic", including a pre-trained model and a fine-tuning model;
[0079] Use the TF-IDF technology to preprocess the data related to "infodemic" and extract the key feature words therefrom;
[0080] Through the BERT model, generate word embedding representations for the extracted "infodemic" keywords and conventional misinformation;
[0081] Fuse the word embedding of the "infodemic" keyword with the BERT word embedding and its text output of the conventional misinformation to obtain the fused data; use the fused data and perform binary classification training according to the conventional "false" or "true" labels to construct a pre-trained model;
[0082] In the fine-tuning model, the "undetermined" category is used to classify the records that cannot be clearly determined as "false" or "true"; utilize the text output that fuses the "infodemic" keyword and the conventional misinformation generated by the BERT layer of the pre-trained model;
[0083] Combined with the BERT model, TextCNN model, and fastText model, the input features of the BERT model are processed, and the processed features are trained using a fine-tuned model. The data related to "infodemic" is subdivided into three categories: "undetermined", "false", or "true".
[0084] In the embodiments of the present invention, by combining the BERT model, TextCNN model, and fastText model, the system can comprehensively utilize the advantages of multiple deep learning models to capture text features from different perspectives, thereby improving the accuracy of classifying data related to "infodemic". This method of multi-model fusion helps to reduce the limitations of a single model and makes the classification results more reliable. Using the TF-IDF technology to preprocess the data can effectively extract key feature words, which helps the system to more accurately understand and analyze text data. At the same time, the word embedding representation generated by the BERT model can capture the context information of words, further enhancing the system's ability to process complex text data. By combining the pre-trained model and the fine-tuned model, the system can not only utilize the general knowledge learned by the pre-trained model but also be optimized for specific tasks through the fine-tuned model. This method of transfer learning enables the system to better adapt to different data and task requirements and improves the flexibility of the system. Traditional binary classification models can often only classify data into two categories: "true" or "false", but in practical applications, some information may not be clearly determined as true or false. By introducing the "undetermined" category and the professional knowledge of health workers in the field of health information, the system can more accurately reflect the real situation of "infodemic" data, avoid misjudging uncertain information, and improve the efficiency of identifying false information. In the context of a public health emergency, false information about "infodemic" spreads extremely fast on social media and the Internet. Therefore, it is crucial to promptly and accurately identify false information about "infodemic". Through the fast training and inference capabilities of the deep learning model, the system can quickly classify a large amount of "infodemic" data, improving the efficiency of identifying false information.
[0085] Infodemic dataset: Chinese and English datasets are used. These datasets contain social media text records annotated by health workers, and all records are labeled as undetermined, false, or true. The Chinese dataset contains 1,055 records, of which 435 are labeled as undetermined, 281 are labeled as false, and 339 are labeled as true. The data sources include manually verified Weibo posts, the WeChat mini-program "Verification", and other authoritative sources. The English dataset contains 2,490 records, with 830 records in each category. The data sources include public fact-checking websites, Twitter API, and other social media platforms.
[0086] Conventional misinformation datasets: The Chinese dataset is sourced from the Zhiyuan Institute of Computing - Internet False News Detection Challenge. After removing low-quality entries, this dataset contains 38,455 records, of which 19,170 are labeled as false and 19,284 are labeled as true. The English dataset is from Kaggle Fake News Detection Datasets, and each record consists of a title and a body. Since the bodies in the true category are usually longer and more formal, while those in the false category are shorter, only the title part is selected. Many titles in the false category have the first letter of each word capitalized, while the true category does not follow this rule. To ensure consistency, all capital letters in the titles of both categories are converted to lowercase. Also, some titles in the false category end with bracketed tags (such as (VIDEO), (IMAGE), or (TWEET)), and these tags are removed. After removing low-quality entries, this dataset contains 42,986 records, of which 21,569 are labeled as false and 21,417 are labeled as true.
[0087] In a preferred embodiment of the present invention, the TF-IDF technique is used to preprocess the data related to "infodemic", and key feature words are extracted therefrom, including:
[0088] Using TF-IDF to preprocess the "infodemic" records, and extracting the top 10% of the keywords therefrom, where "infodemic" is represented as D = {d 1 , d 2 , …, d n}, and each d i represents an "infodemic" record;
[0089] Performing word segmentation on each record d i . For Chinese text, the Jieba word segmentation tool is used; for English text, regular expressions are used for word segmentation; using TF-IDF to process the segmented "infodemic" records, and calculating the TF-IDF score of each term t in the document d i , and its calculation formula is as follows:
[0090] TF-IDF(t, d i ) = TF(t, d i ) × IDF(t);
[0091] Where, N represents the total number of records, |{d: t ∈ d}| represents the number of records containing the term t. After obtaining the TF-IDF scores of each term in all "infodemic" records D, the top 10% of the terms with the highest TF-IDF values are selected as the "infodemic" keywords, and the keywords provide the most relevant information about "infodemic".
[0092] In the embodiment of the present invention, by extracting the top 10% of keywords, the dimensionality of the data can be significantly reduced, which helps the subsequent training and inference of the model, reduces the computational complexity, and improves the processing efficiency. The TF-IDF technique can assign higher weights to those terms that appear frequently in a certain "infodemic" record but are not common in other records. In this way, the extracted keywords can better reflect the uniqueness and importance of each record. By extracting the keywords most relevant to "infodemic", a large amount of redundant and irrelevant information can be removed, making the data input into the subsequent pre-trained model more refined and effective. By using the Jieba tokenization tool to process Chinese text and using regular expressions to process English text, the system can adapt to the tokenization needs of different languages. This multi-language processing ability makes the system more flexible and versatile, and can process "infodemic" data from different language environments.
[0093] In a preferred embodiment of the present invention, through the BERT model, word embedding representations are generated for the extracted "infodemic" keywords and conventional misinformation, including:
[0094] The set of "infodemic" keywords is represented as X Healthwords ={t 1 , t 2 , …, t k}, where k represents the number of "infodemic" keywords, and each keyword t i is processed by the BERT model to generate the corresponding word embedding representation e i , and its calculation formula is as follows:
[0095] E BERT-Healthwords =BERT embedding (X Healthwords ) = [e 1 , e 2 , …, e k ;
[0096] Among them, BERT embedding represents the operation of generating word embeddings using the BERT model, and E BERT-Healthwords represents the result of generating word embeddings using the BERT model.
[0097] X Gen,j represents the token sequence of the j-th instance of conventional misinformation. Processing X Gen,j using the BERT model generates the word embedding representation of X Gen,j , and its calculation formula is as follows;
[0098] E BERT-Gen,j =BERT embedding (X Gen,j );
[0099] Among them, BERT embedding represents the operation of generating word embeddings using the BERT model, and E BERT-Gen,j represents the result of generating word embeddings using the BERT model.
[0100] Fuse the word embeddings of conventional misinformation and the keywords of "infodemic", create a unified representation, and concatenate the word embeddings E of conventional misinformation BERT-Gen,j with the word embeddings e of each "infodemic" keyword i and flatten the result into a single vector. The calculation formula is as follows:
[0101]
[0102] Among them, concat represents the concatenation operation, represents the concatenation method, Flatten represents the flattening operation, and F j represents the fusion of the word embeddings of conventional misinformation and the keywords of "infodemic".
[0103] In the embodiment of the present invention, the BERT model is a pre-trained deep bidirectional model, which can generate word embedding representations rich in context information. By processing the "infodemic" keywords and conventional misinformation through BERT, more accurate and rich semantic representations can be obtained, which helps the subsequent classification tasks. By concatenating the word embeddings of conventional misinformation with the word embeddings of "infodemic" keywords, the system can simultaneously capture the overall semantics of misinformation and the specific keyword information related to "infodemic". This fusion method enhances the integrity and pertinence of the feature representation, which helps the subsequent fine-tuning model to more accurately identify misinformation related to "infodemic". By using the fused word embedding representation as the input feature, more comprehensive information can be provided to the subsequent classifier, thereby improving the accuracy of classification. This representation method fully considers the relevance between the "infodemic" keywords and the recognition of conventional misinformation, which helps the model to make more accurate judgments. Although multiple word embedding representations are fused, by flattening them into a single vector, effective dimensionality reduction of the features is actually performed. This helps to reduce the computational complexity of the subsequent classifier and improve the processing efficiency.
[0104] In a preferred embodiment of the present invention, the fused data is used, and binary classification training is performed according to the conventional "false" or "true" labels to construct a pre-trained model, including:
[0105] Use the BERT model to process X Gen,j , generate the text output of X Gen,j , the text output includes the embedding vector containing context information and the pooled output, and obtain X Gen,jThe pooled output has the following calculation formula:
[0106] pooled_output j =BERT pooled (X Gen,j );
[0107] Among them, BERT pooled represents the operation of generating the pooled output using the BERT model, and pooled_output j represents the result of generating the pooled output using the BERT model.
[0108] Combine pooled_output j and F j , and process the fused feature vectors through a fully connected layer to perform binary classification prediction on the conventional misinformation incorporating the keyword "infodemic". The calculation formula is as follows:
[0109]
[0110] Among them, fc represents the fully connected layer, σ is the sigmoid activation function, is the probability of predicting to belong to the "true" category in the binary classification task.
[0111] In the embodiments of the present invention, the pre-trained model is trained with a large amount of data and can learn general language representations and patterns. This enables the model to have better generalization ability when processing new data, and can make relatively accurate judgments even in the face of unseen information. Through pre-training, the model has learned rich language knowledge and patterns. Therefore, in the subsequent fine-tuning process for specific tasks, it can converge faster and reduce the training time. For the data in the specific field of "infodemic", by incorporating relevant keywords, the model can better capture the specific features and patterns in this field. This enables the model to have higher sensitivity and accuracy when processing information related to "infodemic". By using the pre-trained model, the design of the subsequent task model can be simplified. The pre-trained model provides powerful feature extraction capabilities, enabling the subsequent model to only focus on the design of the classification layer for specific tasks, reducing the complexity of model design. The pre-trained model is usually trained on a large-scale dataset, which makes the model have better robustness to various language phenomena and noisy data.
[0112] In a preferred embodiment of the present invention, in the fine-tuning model, the "undetermined" category is used to classify records that cannot be clearly determined as "false" or "true"; the text output that combines the keyword "infodemic" and conventional misinformation is generated using the BERT layer of the pre-trained model, including:
[0113] Access the regular misinformation text output generated by the BERT layer in the pre-trained model, and obtain a batch of embedding vectors and pooled outputs of the "infodemic" text containing context information. The calculation formulas are as follows:
[0114]
[0115] Among them, X Health represents the token sequence of the "infodemic" instance input batch, BERT pooled represents the operation of generating a pooled output using the BERT model, pooled_output represents the result of generating a pooled output using the BERT model, P represents that the result of generating a pooled output by the BERT model is a matrix, and the dimension of this matrix is is the set of real numbers, B is the batch size, L is the sequence length, BERT encoded represents the operation of generating a contextualized embedding vector using the BERT model, contextualized_token_embedding represents the result of generating a contextualized embedding vector using the BERT model, H represents that the result of generating a contextualized embedding vector by the BERT model is a matrix, and the dimension of this matrix is H is the hidden layer dimension;
[0116] The pooled output is processed by the dropout layer, and the contextualized embedding vector is processed by the TextCNN pooling layer and the fastText pooling layer. The calculation formulas are as follows:
[0117] P drop = Dropout(P);
[0118] C 1 = AdaptiveMaxPool1D(Conv1D(H, k 1 ));
[0119] C 2 = AdaptiveMaxPool1D(Conv1D(H, k 2 ));
[0120] F mean = AdaptiveAvgPool1D(H);
[0121] F max = AdaptiveMaxPool1D(H);
[0122] Among them, P drop represents the result after applying dropout regularization to the pooled output; C 1 and C2 is the output of the pooling layer of TextCNN. The pooling layer performs convolution operations with kernel sizes of k 1 and k 2 on the embedding vectors containing context information, and performs adaptive max pooling. Conv1D represents a one-dimensional convolution operation, and AdaptiveMaxPool1D represents an adaptive max pooling operation; F mean and F max are the outputs of the pooling layer of fastText. The pooling layer performs adaptive average pooling and adaptive max pooling on the embedding vectors containing context information respectively. AdaptiveAvgPool1D represents an adaptive average pooling operation.
[0123] In the embodiments of the present invention, by introducing the "undetermined" category, the model can more carefully process those records that cannot be clearly determined as "false" or "true". This three-classification (false, true, undetermined) method helps to reduce misjudgments and improve the accuracy of information classification. The text output generated by the BERT model, including the embedding vectors containing context information and the pooling output, provides richer and deeper semantic information for the model. This helps the model to make more robust judgments when processing complex or ambiguous information. By combining the pooling layers of TextCNN and fastText, the model can extract text features from multiple perspectives. TextCNN efficiently extracts local patterns and significant features from word embeddings, and is very suitable for identifying keyword groups and local dependencies. FastText captures global sequence-level features and provides a lightweight and effective solution for representing the overall trend and main features in limited data scenarios. This multi-level feature extraction method improves the model's ability to understand text information. Applying dropout regularization to the pooling output helps to reduce the risk of overfitting in the training process of the model. Dropout can randomly discard some network connections, so that the model does not overly rely on certain specific connections during training, thereby enhancing the generalization ability of the model.
[0124] In a preferred embodiment of the present invention, by combining the BERT model, the TextCNN model, and the fastText model, the input features of the BERT model are processed, and the processed features are trained using a fine-tuning model. The data related to "infodemic" is divided into three categories: "undetermined", "false", or "true", including:
[0125] Concatenate the processing results of the dropout layer, the TextCNN pooling layer, and the fastText pooling layer to obtain the final feature representation, and its calculation formula is as follows:
[0126] F = [P drop , C 1 , C2 , F mean , F max ;
[0127] Among them, F represents the concatenated feature vector;
[0128] The concatenated feature vector is passed to the fully connected layer to perform a three-class prediction on the "infodemic" record, and its calculation formula is as follows:
[0129]
[0130] Among them, φ represents the softmax activation function, represents the prediction probability distribution of each category in the three-class task.
[0131] In the embodiments of the present invention, by concatenating the processing results of the dropout layer, the TextCNN pooling layer, and the fastText pooling layer, the model can fuse the feature information from different levels. This feature fusion method enhances the model's representation ability for text data, enabling the model to capture the semantic and structural information in the text more comprehensively. TextCNN efficiently extracts local patterns and significant features from word embeddings, while fastText captures global sequence-level features. Combining the outputs of these models with the pooling output of BERT processed by dropout regularization enables the fine-tuned model to understand the text from multiple perspectives, improving the processing ability for complex text data. Introducing the "undetermined" category and fine-tuning based on the pre-trained model enables the model to better adapt to small data sets in the three-class task, not only improving the model performance but also effectively saving computational resources and time. The entire process realizes end-to-end training. From the original text input to the final three-class output, all components can be optimized in a unified framework. This end-to-end training method helps the model better adapt to the task requirements, improving the training efficiency and model performance.
[0132] Table 1. Test results of the three-class task for identifying false information of "infodemic" in Chinese
[0133] <![CDATA[P macro > <![CDATA[R macro > <![CDATA[F1 macro > ACC FastText 0.7652 0.7673 0.7655 0.7806 TextRNN 0.7119 0.7174 0.7109 0.7173 TextRNN_Att 0.7417 0.7491 0.7441 0.7553 TextRCNN 0.7229 0.7208 0.7204 0.7342 TextCNN 0.7683 0.7696 0.7689 0.7806 DPCNN 0.6993 0.7002 0.6972 0.7173 Transformer 0.6595 0.6708 0.6618 0.6709 BERT 0.8120 0.8136 0.8075 0.8101 Hybrid 0.8435 0.8521 0.8471 0.8523
[0134] Table 2. Test results of the three-class task for identifying false information of "infodemic" in English
[0135]
[0136]
[0137] Table 3. Test results of each category of the three-class task for identifying false information of "infodemic" in Chinese
[0138]
[0139] Note: 0, 1, and 2 respectively represent the "infodemic" data marked as "uncertain", "false", and "true".
[0140] Table 4. Test results of each category in the three-class classification task for identifying false information of "infodemic" in English
[0141]
[0142] Note: 0, 1, and 2 respectively represent the "infodemic" data marked as "uncertain", "false", and "true".
[0143] An embodiment of the present invention also provides a three-class hybrid transfer learning method for identifying false information of "infodemic", including:
[0144] Using the TF-IDF technology to preprocess the data related to "infodemic" and extract key feature words from it;
[0145] Through the BERT model, generate word embedding representations for the extracted "infodemic" keywords and conventional false information;
[0146] Fuse the word embedding of the "infodemic" keywords with the BERT word embedding and its text output of the conventional false information to obtain the fused data; use the fused data and perform binary classification training according to the conventional "false" or "true" labels to construct a pre-trained model;
[0147] In the fine-tuning model, use the "undetermined" category to classify records that cannot be clearly determined as "false" or "true"; utilize the text output of the fused "infodemic" keywords and conventional false information generated by the BERT layer of the pre-trained model;
[0148] Combine the BERT model, TextCNN model, and fastText model to process the input features of the BERT model, use the fine-tuning model to train the processed features, and subdivide the data related to "infodemic" into three categories: "undetermined", "false", or "true".
[0149] It should be noted that this system corresponds to the above method, and all implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0150] An embodiment of the present invention also provides a computing device, including: a processor and a memory storing a computer program. When the computer program is run by the processor, it executes the method as described above. All implementation methods in the above method embodiments are applicable to this embodiment and can achieve the same technical effects.
[0151] An embodiment of the present invention also provides a computer-readable storage medium storing instructions, which, when run on a computer, cause the computer to execute the method described above. All implementation manners in the above method embodiments are applicable to this embodiment and can also achieve the same technical effects.
Claims
1. A three-classification hybrid transfer learning system for identifying false information about "information epidemics", characterized by: The system includes a pre-trained model and a fine-tuned model; Use TF-IDF technology to pre-process data related to "information epidemic" and extract key feature words from it; Through the BERT model, word embedding representations are generated for the extracted "information epidemic" keywords and common false information; The word embedding of the keyword "information epidemic" is fused with the BERT word embedding of conventional false information and its text output to obtain fused data; using the fused data and based on the conventional "false" or "true" labels, binary classification training is performed to build a pre-trained model; In the fine-tuning model, the "Undetermined" category is used to classify records that cannot be clearly determined as "false" or "true". The text output generated by the BERT layer of the pre-trained model combines the "infodemic" keywords and general false information; The BERT model, TextCNN model and fastText model are combined to process the input features of the BERT model, and the processed features are trained using the fine-tuning model to subdivide the data related to the "information epidemic" into three categories: "uncertain", "false" or "true".
2. According to claim 1, a three-classification hybrid transfer learning system for identifying false information about "information epidemics" is characterized by: TF-IDF technology is used to pre-process the data related to "information epidemic" and extract key feature words, including: TF-IDF is used to pre-process the "information epidemic" records and extract the top 10% keywords from them, where "information epidemic" is represented by D = {d1, d2, ..., d n }, each d i Represents an "information epidemic" record; For each record d i For word segmentation, Jieba word segmentation tool is used for Chinese text; regular expressions are used for English text; TF-IDF is used to process the "information epidemic" records after word segmentation, and the number of each term t in document d is calculated. i The TF-IDF score in is calculated as follows: TF-IDF(t,d i )=TF(t,d i )×IDF(t); in, represents the total number of records, |{d:t∈d}| represents the number of records containing term t. After obtaining the TF-IDF score of each term in all "information epidemic" records D, the 10% terms with the highest TF-IDF values are selected as "information epidemic" keywords. The keywords provide the most relevant information about "information epidemic".
3. According to claim 2, a three-classification hybrid transfer learning system for identifying false information about "information epidemics" is characterized by: Through the BERT model, word embedding representations are generated for the extracted "information epidemic" keywords and common false information, including: The keyword set of "information epidemic" is represented by X Healthwords ={t1,t2,…,t k }, where k represents the number of information epidemic keywords, and each keyword t i After being processed by the BERT model, the corresponding word embedding representation e is generated i , and its calculation formula is as follows: AND BERT-Healthwords =BERT embedding (X Healthwords )=[e1,e2,…,e k ]; Among them, BERT embedding Indicates the operation of generating word embedding using the BERT model, E BERT-Healthwords Indicates the result of generating word embedding using the BERT model; X Gen,j The tag sequence representing the jth instance of conventional false information, processed by the BERT model X Gen,j , generating X Gen,j The word embedding representation is calculated as follows: E BERT-Gen,j =BERT embedding (X Gen,j ); Among them, BERT embedding Indicates the operation of generating word embedding using the BERT model, E BERT-Gen,j Indicates the result of generating word embedding using the BERT model; The word embeddings of conventional false information and "information epidemic" keywords are combined to create a unified representation, and the word embeddings of conventional false information are E BERT-Gen,j The word embedding of each "infodemic" keyword i Concatenate them and flatten the result into a single vector, which is calculated as follows: Among them, concat represents the concatenation operation. Indicates the splicing method, Flatten indicates the flattening operation, and F j Represents word embedding that fuses conventional misinformation and “infodemic” keywords.
4. According to claim 3, a three-classification hybrid transfer learning system for identifying false information about "information epidemic" is characterized by: Use the fused data and perform binary classification training based on the conventional "fake" or "real" labels to build a pre-trained model, including: Use BERT model to process X Gen,j , generating X Gen,j The text output includes the embedding vector containing context information and the pooling output, and obtains X Gen,j The pooling output is calculated as follows: pooled_output j =BERT pooled (X Gen,j ); Among them, BERT pooled Indicates the operation of generating pooled output using the BERT model, pooled_output j Indicates the result of generating pooled output using the BERT model; pooled_output j and F j Combined with the fused feature vector processed by the fully connected layer, a binary classification prediction is performed on the conventional false information that incorporates the keyword "information epidemic". The calculation formula is as follows: Among them, fc represents the fully connected layer, σ is the sigmoid activation function, The probability of predicting the "true" category in a binary classification task.
5. According to claim 4, a three-classification hybrid transfer learning system for identifying false information about "information epidemic" is characterized by: In the fine-tuning model, the "Undetermined" category is used to classify records that cannot be clearly determined as "false" or "true". The text output generated by the BERT layer of the pre-trained model combines the "information epidemic" keywords and general false information, including: Access the regular false information text output generated by the BERT layer in the pre-trained model, and obtain a batch of "information epidemic" texts containing contextual information embedding vectors and pooling outputs. The calculation formula is as follows: Among them, X Health A sequence of tokens representing an input batch of "information epidemic" instances, BERT pooled Indicates the operation of generating pooled output using the BERT model, pooled_output indicates the result of generating pooled output using the BERT model, and P indicates that the result of generating pooled output by the BERT model is a matrix, and the dimension of the matrix is is a set of real numbers, B is the batch size, L is the sequence length, BERT encoded Indicates the operation of using the BERT model to generate an embedding vector containing contextual information. contextualized_token_embedding indicates the result of using the BERT model to generate an embedding vector containing contextual information. H indicates that the result of the BERT model generating an embedding vector containing contextual information is a matrix. The dimension of the matrix is H is the hidden layer dimension; The pooled output is processed by the dropout layer, and the embedding vector containing contextual information is processed by the TextCNN pooling layer and the fastText pooling layer. The calculation formula is as follows: P drop =Dropout(P); C1=AdaptiveMaxPool1D(Conv1D(H,k1)); C2=AdaptiveMaxPool1D(Conv1D(H,k2)); F mean =AdaptiveAvgPool1D(H); F max =AdaptiveMaxPool1D(H); Among them, P drop represents the result of applying dropout regularization to the pooled output; C1 and C2 are the outputs of the TextCNN pooling layer. The pooling layer performs adaptive maximum pooling on the embedded vector containing context information using convolution operations with kernel sizes of k1 and k2, respectively. Conv1D represents a one-dimensional convolution operation, and AdaptiveMaxPool1D represents an adaptive maximum pooling operation; F mean and F max It is the output of the fastText pooling layer. The pooling layer performs adaptive average pooling and adaptive maximum pooling on the embedded vector containing context information. AdaptiveAvgPool1D represents the adaptive average pooling operation.
6. According to claim 5, a three-classification hybrid transfer learning system for identifying false information about "information epidemics" is characterized by: Combining the BERT model, the TextCNN model, and the fastText model, the input features of the BERT model are processed, and the processed features are trained using the fine-tuning model to subdivide the data related to the "information epidemic" into three categories: "uncertain", "false", or "true", including: Concatenate the processing results of the dropout layer, TextCNN pooling layer, and fastText pooling layer to obtain the final feature representation. The calculation formula is as follows: F=[P drop ,C1,C2,F mean ,F max ]; Among them, F represents the concatenated feature vector; The concatenated feature vector is passed to the fully connected layer to perform three-classification prediction on the "information epidemic" record. The calculation formula is as follows: Among them, φ represents the softmax activation function, Represents the predicted probability distribution of each category in the three-category classification task.
7. A three-classification hybrid transfer learning method for identifying false information about "information epidemic", characterized by: Applicable to the system according to any one of claims 1 to 6, comprising: Use TF-IDF technology to pre-process data related to "information epidemic" and extract key feature words from it; Through the BERT model, word embedding representations are generated for the extracted "information epidemic" keywords and common false information; The word embedding of the keyword "information epidemic" is fused with the BERT word embedding of conventional false information and its text output to obtain fused data; using the fused data and based on the conventional "false" or "true" labels, binary classification training is performed to build a pre-trained model; In the fine-tuning model, the "Undetermined" category is used to classify records that cannot be clearly determined as "false" or "true". The text output generated by the BERT layer of the pre-trained model combines the "infodemic" keywords and general false information; The BERT model, TextCNN model and fastText model are combined to process the input features of the BERT model, and the processed features are trained using the fine-tuning model to subdivide the data related to the "information epidemic" into three categories: "uncertain", "false" or "true".
8. A computing device, characterized in that include: one or more processors; A storage device for storing one or more programs, when the one or more programs are executed by the one or more processors, the one or more processors implement the method according to claim 7.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a program, which implements the method according to claim 7 when executed by a processor.
Citation Information
Patent Citations
Bad information identification method based on TextCNN-Bert fusion model algorithm
CN116796740A
Social media rumor detection method based on active learning iteration
CN117992651A
False information identification method based on knowledge enhancement and semantic difference
CN118013040A
Few-sample supervision Chinese fact checking system enhanced by rumor detection data
CN118133151A
Moss block
KR102489713B1