Tibetan rumor model training method and device and Tibetan rumor model detection method and device
Through the method of combining CINO pre-training model and convolutional neural network with multiple self-attention mechanism, the problems of complex grammar and scarcity of resources in Tibetan rumor detection are solved, and efficient and accurate rumor detection is achieved.
Patent Information
- Application Number
- CN202510399298.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-01
- Publication Date
- 2025-07-18
AI Technical Summary
The existing rumor detection methods are difficult to effectively adapt to the language characteristics of Tibetan language, especially in a low-resource environment, which cannot effectively capture the grammatical and semantic characteristics of Tibetan language, and it is difficult to deal with complex morphological changes and unique grammatical structures.
The CINO pre-trained model is used to encode Tibetan texts, generate initial representation vectors, extract local semantic features using convolutional neural networks, and process global context information through a multi-head self-attention mechanism, and adaptive weighting fusion is performed in combination with learnable weight parameters, and finally rumor detection and classification is performed through the full connection layer.
It effectively solves the modeling difficulties of Tibetan language under low resource conditions, improves the accuracy and robustness of Tibetan rumor detection, can adapt to complex and changeable practical application scenarios, and achieves efficient detection of rumors.
Smart Images

Figure CN120336525A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of natural language processing, and particularly to a Tibetan rumor detection training method, detection method and device based on a hybrid convolutional-self attention model. The present invention processes Tibetan texts through deep learning technology to achieve automatic detection of rumor information. Background Art
[0002] With the rapid development of the Internet, social media and online platforms have become the main channels for information dissemination. However, the spread of false information, especially rumors, on the Internet may have serious negative impacts on society, economy, culture, etc. As an important ethnic minority language, Tibetan faces unique challenges in the field of rumor detection. These challenges include the lack of language resources, complex morphological changes and unique grammatical structures, etc., and traditional rumor detection methods are difficult to effectively adapt to the language characteristics of Tibetan. At present, natural language processing technologies based on deep learning are widely applied to rumor detection tasks. Among them, the convolutional neural network (CNN) can effectively extract local features, the recurrent neural network (RNN) or long short-term memory network (LSTM) is good at capturing sequence dependencies, and the Transformer architecture performs excellently in modeling global dependencies through the multi-head self-attention mechanism. However, the performance of these methods is limited in low-resource language environments.
[0003] In view of this, the present invention is particularly proposed. Summary of the Invention
[0004] The technical problem to be solved by the present invention is to overcome the deficiencies of the prior art and provide a Tibetan rumor model training method, detection method and device, which solve the problems raised in the above background art.
[0005] The basic idea of the technical solution adopted by the present invention to solve the above technical problems is as follows:
[0006] A Tibetan rumor detection method includes the following steps:
[0007] S1: Use the CINO pre-trained model to encode the Tibetan text to generate an initial representation vector of the Tibetan text;
[0008] S2: Use a convolutional neural network to extract features from the initial representation vector to obtain local semantic features;
[0009] S3: Use the multi-head self-attention mechanism to process the local semantic features to obtain global context information;
[0010] S4: Perform adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation;
[0011] S5: Based on the fused feature representation, perform rumor detection and classification through a fully connected layer, and output the detection result.
[0012] Optionally, the steps of using the CINO pre-trained model to encode the Tibetan text and generate the initial representation vector of the Tibetan text are as follows:
[0013] S101: Input the sequence x = {x1, x2,..., x n}, where n represents the sequence length, and each is a vector d of dimension. The goal is to predict a label y ∈ {0, 1}, where y = 0 represents a rumor and y = 1 represents a non-rumor;
[0014] S102: For the preprocessed Tibetan text, use the CINO pre-trained model for encoding. Through model calculation, generate the preliminary representation vector h (0) of the text sequence. Its expression: h (0) = CINO(x), where h (0) is the preliminary semantic representation of the text sequence. In addition, encode the semantic and syntactic features crucial for downstream tasks. Finally, use the generated initial representation vector h (0) as the input for subsequent feature extraction and classification.
[0015] Optionally, the steps of using a convolutional neural network to extract features from the initial representation vector and obtain local semantic features are as follows:
[0016] S201: Design and initialize convolutional kernels of multiple sizes to capture semantic features at the affix, phrase, and sentence levels respectively. Among them, the sizes of the convolutional kernels are [3, 5, 7]. Then, input the initial representation vector into the convolutional neural network to perform one-dimensional convolutional operations to capture n-gram features of different granularities respectively;
[0017] S202: Each convolutional kernel slides the window to extract features and generates a local feature map h (1) . Then, generate a query matrix Q, a key matrix K, and a value matrix V from the local feature map h (1) . Their expressions are Q = h (1) *W q , K = h (1) *W k , V = h (1) *W v , where is a learnable weight matrix, k is the output feature dimension of the convolutional kernel, used to project the initial representation vector h (0) into the feature space, and * represents the 1D convolutional operation;
[0018] S203: Calculate the positional correlation weights within a local range using the query matrix Q and the key matrix K, and its expression is as follows: where d k is the dimension of the key matrix, which is used to scale the correlation scores to prevent the gradients from being unstable due to overly large numerical values. QK T is the dot product of the query matrix Q and the key matrix K, generating an n×n correlation matrix;
[0019] S204: Input the feature z generated by local attention into the feed-forward network (FFN), and further optimize the local feature representation through two-layer linear transformation and activation function. Its expression is: FFN(z) = ReLU(zW1 + b1)W2 + b2, where z represents the embedding involved locally and serves as the final local feature representation, W1 and W2 are weight matrices used for the first and second layer linear transformations respectively, b1 and b2 are the corresponding bias vectors, and z = Attention(Q, K, V);
[0020] S205: Apply ReLU activation to the feature FFN(z) generated by the feed-forward network to enhance the non-linear expression ability, perform normalization on the activated feature, and finally, input the normalized feature matrix into the max-pooling layer to extract the significant features in each local region, thereby obtaining the local semantic feature h (l) .
[0021] Optionally, the steps of using the multi-head self-attention mechanism to process the local semantic feature to obtain global context information are as follows:
[0022] Based on the input local semantic feature, generate the query matrix Q, key matrix K, and value matrix V of the multi-head self-attention mechanism. The specific calculation formula for each attention head head i is: head i = Attention(Q i , K i , V i ), where Q i , K i , V are the i-th heads of the query, key, and value matrices respectively, and d k is the scaling factor used to prevent the attention scores from being overly large;
[0023] Concatenate the outputs of all attention heads and perform a linear transformation through the weight matrix W o to generate the representation h (global) of the global context information. The combined formula is: MultiHead(Q, K, V) = concat(had1,..., head H )W O ·, where HHH is the number of attention heads and W ois a learnable weight matrix;
[0024] Use a feed-forward network (FFN) to represent the global feature h generated by the self-attention mechanism (global) For further optimization, the feedforward network consists of two layers of linear transformation and nonlinear activation. The specific formula is: Z1 = ReLU (h (l) W1+b1), Z2=Z1W2 + b2, where W1, W2 are the weight matrices of the feedforward network, b1, b2 are bias terms, and X is the self-attention output feature.
[0025] Optionally, the steps of performing adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation are:
[0026] S401: weighted fusion of local semantic features and optimized global context information Z2 according to learnable weight parameters, calculated as: h (fusion) =α·h (locla) +β·Z2, where represents the fused feature representation, α controls the weight of local features, and β controls the weight of global features. And satisfy: α+β=1;
[0027] S402: The fused feature h (fusion) As the input of the subsequent classification task, the core semantics of local semantic features and global context information are preserved.
[0028] Optionally, based on the fused feature representation, rumor detection and classification are performed through a fully connected layer, and the steps of outputting the detection result are:
[0029] S501: Receive the fused features and use the formula For classification, σ is the sigmoid activation function, w fc is the weight matrix of the final fully connected layer, b fc is the bias term;
[0030] The predicted probability Compare with the preset threshold τ and output the detection result: If It is predicted as non-rumor, where τ is the threshold of the classification task.
[0031] A training method for a Tibetan rumor detection model comprises the following steps:
[0032] Load the Rumor_Ti dataset, divide the dataset into training set, validation set and test set, and at the same time preprocess the data, including noise reduction, word segmentation and sample verification, to ensure the data quality and the grammatical structure of the language;
[0033] Initialize the model parameters, including the weight matrices of the convolutional layer, Transformer encoder layer and feed-forward network. The convolutional layer sets the number of filters to 128 and the convolutional kernel size to [3, 5, 7]. The Transformer encoder layer is configured with 8 attention heads and the model dimension is 256.
[0034] Input the preprocessed data into the model in batches, and generate the initial embedding features h (0) for each batch, where the feature dimension is B×n×d, where B is the batch size, h is the sequence length, and d is the embedding feature dimension;
[0035] Use the binary cross-entropy (BCE) loss function to optimize the classification model, and its expression is: where N is the number of samples, y i is the true label of the sample, is the predicted probability of the sample;
[0036] Use the Adam optimizer to update the model parameters, with an initial learning rate of 1e-3, combined with a weight decay of 1e-5 to avoid overfitting. Calculate the gradient through the backpropagation algorithm and gradually reduce the value of the loss function;
[0037] Adopt the cosine annealing strategy to dynamically adjust the learning rate, so that the learning rate gradually decreases during the training process, thereby accelerating convergence and improving the model stability, and its expression is:
[0038] where η t is the learning rate of the current iteration, T is the total number of training epochs, η min and η max are the minimum and maximum values of the learning rate respectively;
[0039] After each training epoch, use the validation set to evaluate the model performance, and save the model weights with the lowest validation set loss or the highest accuracy;
[0040] Stop training when the validation set loss does not decrease for 5 consecutive epochs to ensure that the model does not overfit or waste computing resources.
[0041] A Tibetan rumor detection device, including
[0042] Text encoding module: Use the CINO pre-trained model to encode the Tibetan text and generate the initial representation vector of the Tibetan text;
[0043] Local feature extraction module: Use a convolutional neural network to extract features from the initial representation vector to obtain local semantic features;
[0044] Global feature modeling module: Use the multi-head self-attention mechanism to process the local semantic features to obtain global context information;
[0045] Feature fusion module: Perform adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation;
[0046] Classification module: Based on the fused feature representation, perform rumor detection classification through a fully connected layer and output the detection result.
[0047] After adopting the above technical solutions, the present invention has the following beneficial effects compared with the prior art. Of course, any product implementing the present invention does not necessarily need to achieve all the advantages described below at the same time:
[0048] 1. The Tibetan rumor detection method of the present invention achieves significant beneficial effects through the following techniques: First, use the CINO pre-trained model to encode Tibetan texts to generate initial representation vectors that can capture the unique grammar and semantic features of Tibetan, effectively solving the problem of modeling difficulties for Tibetan as a low-resource language under data-scarce conditions. Second, extract local semantic features through a convolutional neural network, and use convolutional kernels to capture n-gram patterns of different granularities in the text, enhancing the ability to recognize phrases and local patterns. At the same time, the multi-head self-attention mechanism further processes the local features, models long-distance dependencies, and generates global context information, solving the modeling difficulties of complex Tibetan syntax and long-distance dependencies.
[0049] 2. The present invention realizes the adaptive weighted fusion of local features and global context through learnable weight parameters, dynamically adjusts the feature weight ratio, and enables the model to have both fine-grained feature expression and overall semantic understanding capabilities. Finally, the fused features are classified through a fully connected layer to achieve efficient detection of scientific, social, and other types of rumors. The detection results have high accuracy and strong robustness, and can adapt to complex and changing actual application scenarios.
[0050] The following further describes the specific implementation manners of the present invention in detail with reference to the accompanying drawings. Description of the Drawings
[0051] Figure 1 is the workflow diagram based on the CINO model;
[0052] Figure 2 is the schematic diagram of the T-LAM architecture;
[0053] Figure 3 is the schematic diagram of model performance comparison;
[0054] Figure 4 It is a flow chart for Tibetan rumor detection. Specific implementation manners
[0055] Now, the present invention will be further described in detail with reference to the accompanying drawings.
[0056] Please refer to Figures 1-4 As shown, in this embodiment, a Tibetan rumor detection method is provided, including the following steps:
[0057] S1: Use the CINO pre-trained model to encode the Tibetan text to generate an initial representation vector of the Tibetan text;
[0058] S2: Adopt a convolutional neural network to extract features from the initial representation vector to obtain local semantic features;
[0059] S3: Use the multi-head self-attention mechanism to process the local semantic features to obtain global context information;
[0060] S4: Perform adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation;
[0061] S5: Based on the fused feature representation, perform rumor detection classification through a fully connected layer and output the detection result.
[0062] In this embodiment, the step of using the CINO pre-trained model to encode the Tibetan text to generate an initial representation vector of the Tibetan text is as follows:
[0063] S101: Input the sequence x = {x1, x2,..., x n}, where n represents the sequence length, and each is a vector d of dimension. The goal is to predict a label y ∈ {0, 1}, where y = 0 represents a rumor and y = 1 represents non-rumor;
[0064] S102: For the preprocessed Tibetan text, use the CINO pre-trained model for encoding. Through model calculation, generate a preliminary representation vector h (0) of the text sequence. Its expression: h (0) = CINO(x), where h (0) is the preliminary semantic representation of the text sequence, including the embedding representation of each token unit, retaining local and global context information. In addition, encode semantic and syntactic features crucial for downstream tasks. Finally, use the generated initial representation vector h (0) as the input for subsequent feature extraction and classification.
[0065] In this embodiment, the step of using a convolutional neural network to extract features from the initial representation vector to obtain local semantic features is as follows:
[0066] S201: Design and initialize convolutional kernels of multiple sizes to capture semantic features at the affix, phrase, and sentence levels respectively. Among them, the sizes of the convolutional kernels are [3, 5, 7]. Then, input the initial representation vector into the convolutional neural network for one-dimensional convolutional operations to capture n-gram features of different granularities respectively. Among them, small-sized convolutional kernels extract short-distance features (such as word-level relationships), and large-sized convolutional kernels capture longer-range phrase-level and sentence-level features. These operations can retain local context information and provide a basis for further attention mechanism modeling;
[0067] S202: Each convolutional kernel slides the window to extract features and generates a local feature map h (1) , and then, generate a query matrix Q, a key matrix K, and a value matrix V from the local feature map h (1) . Their expressions are Q = h (1) *W q , K = h (1) *Wk, V = h (1) *W v , where is a learnable weight matrix, k is the output feature dimension of the convolutional kernel, and is used to project the initial representation vector h (0) into the feature space;
[0068] S203: Use the query matrix Q and the key matrix K to calculate the position correlation weights within the local range. Its expression is: where d k is the dimension of the key matrix, which is used to scale the correlation scores to prevent the gradient from being unstable due to overly large values. QK T is the dot product of the query matrix Q and the key matrix K, generating an n×n correlation matrix;
[0069] S204: Input the feature z generated by local attention into a feed-forward network (FFN), and further optimize the local feature representation through two-layer linear transformation and activation function. Its expression is: FFN(z) = ReLU(zW1 + b1)W2 + b2, where z is the feature generated by local attention and serves as the final local feature representation, W1, W2 are weight matrices used for the first and second layer linear transformations respectively, b1, b2 are the corresponding bias vectors, and z = Attention(Q, K, V);
[0070] S205: ReLU activation is performed on the features FFN(z) generated by the feed-forward network to enhance the non-linear expression ability. The activated features are then normalized. Finally, the normalized feature matrix is input into the max-pooling layer to extract the significant features in each local region, and the local semantic features h can be obtained. (l) , The significant features include but are not limited to local context patterns: such as the semantic relationship between consecutive words; key semantic features: such as the importance and semantic information of words; boundary information: such as the feature representation of clause or sentence boundaries.
[0071] In this embodiment, the steps of using the multi-head self-attention mechanism to process the local semantic features and obtain the global context information are as follows: The steps of using the multi-head self-attention mechanism to process the local semantic features and obtain the global context information are as follows:
[0072] Based on the input local semantic features, query matrix Q, key matrix K, and value matrix V of the multi-head self-attention mechanism are generated. For each attention head head i , the specific calculation formula is: head i = Attention(Q i , K i , V i ), where Q i , K i , V are the i-th heads of the query, key, and value matrices respectively, and d k is the scaling factor;
[0073] Concatenate the outputs of all attention heads and perform a linear transformation through the weight matrix W o to generate the representation h of the global context information (global) , and the combined formula is: MultiHead(Q, K, V) = concat(had1,..., head H )W O ·, where HHH is the number of attention heads, and W o is a learnable weight matrix;
[0074] Use a feed-forward network (FFN) to further optimize the global feature representation h generated by the self-attention mechanism. The feed-forward network consists of two layers of linear transformation and non-linear activation, and the specific formula is: Z1 = ReLU(h (global) W1 + b1), Z2 = Z1W2 + b2, where W1, W2 are the weight matrices of the feed-forward network, b1, b2 are the bias terms, and X is the self-attention output feature. (l)
[0075] In this embodiment, the steps of adaptively weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain the fused feature representation are as follows:
[0076] S401: Weightedly fuse the local semantic features and the optimized global context information Z2 according to the learnable weight parameters. The calculation formula is: h (fusion) = α·h (local) + β·Z2, where represents the fused feature representation, α controls the weight of the local features, β controls the weight of the global features, and satisfy: α + β = 1;
[0077] S402: Use the fused feature h (fusion) as the input for the subsequent classification task, retaining the core semantics of the local semantic features and the global context information.
[0078] In this embodiment, based on the fused feature representation, the steps of performing rumor detection and classification through a fully connected layer and outputting the detection result are as follows:
[0079] S501: Receive the fused feature and classify it through the formula , where σ is the sigmoid activation function, w fc is the weight matrix of the final fully connected layer, and b fc is the bias term;
[0080] Compare the predicted probability with the preset threshold τ and output the detection result: If then predict as a rumor; if then predict as non-rumor, where τ is the threshold for the classification task.
[0081] A training method for a Tibetan rumor detection model includes the following steps:
[0082] Load the Rumor_Ti dataset, divide the dataset into a training set, a validation set, and a test set, and at the same time preprocess the data, including noise reduction, word segmentation, and sample verification, to ensure the data quality and the grammatical structure of the language;
[0083] Initialize the model parameters, including the weight matrices of the convolutional layer, the Transformer encoder layer, and the feed-forward network. The convolutional layer sets the number of filters to 128 and the convolutional kernel size to [3, 5, 7]. The Transformer encoder layer is configured with 8 attention heads and the model dimension is 256.
[0084] Input the preprocessed data into the model in batches, and generate initial embedding features h (0) for each batch. The feature dimension is B × n × d, where B is the batch size, h is the sequence length, and d is the embedding feature dimension;
[0085] Optimize the classification model using the Binary Cross Entropy (BCE) loss function, whose expression is: where N is the number of samples, y i is the true label of the sample, is the predicted probability of the sample;
[0086] Update the model parameters using the Adam optimizer with an initial learning rate of 1e-3 and combine it with a weight decay of 1e-5 to avoid overfitting. Calculate the gradient through the backpropagation algorithm and gradually reduce the value of the loss function;
[0087] Adopt the cosine annealing strategy to dynamically adjust the learning rate, so that the learning rate gradually decreases during training, thereby accelerating convergence and improving the model stability. Its expression is:
[0088] where η t is the learning rate of the current iteration, T is the total number of training epochs, η min and η max are the minimum and maximum values of the learning rate respectively;
[0089] After each training epoch, evaluate the model performance using the validation set and save the model weights with the lowest validation set loss or the highest accuracy;
[0090] Stop training when the validation set loss has not decreased for 5 consecutive epochs to ensure that the model does not overfit or waste computing resources.
[0091] A Tibetan rumor detection device, including
[0092] Text encoding module: Use the CINO pre-trained model to encode the Tibetan text and generate the initial representation vector of the Tibetan text;
[0093] Local feature extraction module: Adopt a convolutional neural network to extract features from the initial representation vector to obtain local semantic features;
[0094] Global feature modeling module: Use the multi-head self-attention mechanism to process the local semantic features to obtain global context information;
[0095] Feature fusion module: Perform adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation;
[0096] Classification module: Based on the fused feature representation, perform rumor detection classification through a fully connected layer and output the detection result.
[0097] 4. Experimental settings and results
[0098] 4.1 Baseline model
[0099] To evaluate the performance of Transformer-based models with the Token-Level Attention Mechanism (T-LAM) in Tibetan rumor detection, it was compared with several baseline models commonly used in low-resource language processing. These models represent various architectures designed to capture local and global dependencies in text, providing a solid basis for evaluating the effectiveness of T-LAM. Table 1 summarizes the selected models and their configurations.
[0100] Overview of Baseline Models:
[0101] 1. CNN: A standard Convolutional Neural Network (CNN) that extracts local n-gram features through convolution operations. This model is used as a baseline for evaluating local dependency capture.
[0102] Key Configuration: Three convolutional layers with filter sizes [3, 5, 7], [3, 5, 7], [3, 5, 7], and 128 filters per layer.
[0103] 2. BiLSTM: A Bidirectional Long Short-Term Memory (BiLSTM) network that captures sequential dependencies in both directions. It is typically effective for sequential text tasks, making it a relevant comparison for T-LAM.
[0104] Key Configuration: Two bidirectional LSTM layers with a hidden size of 128.
[0105] 3. Transformer: A vanilla Transformer architecture that models long-term dependencies through multi-head self-attention, serving as a base comparison for T-LAM.
[0106] Key Configuration: Four encoder layers, eight attention heads, and a model dimension of 256.
[0107] 3. RoBERTa-Ti: A pre-trained RoBERTa model fine-tuned on Tibetan text, aiming to leverage large-scale language modeling for low-resource languages.
[0108] Key Configuration: Twelve Transformer layers with a hidden size of 768.
[0109] 4. CINO: A multilingual pre-trained model tailored for minority languages including Tibetan, providing a strong baseline for resource-scarce language tasks.
[0110] Key Configuration: Twenty-four Transformer layers trained on a diverse multilingual corpus.
[0111] 5. T-LAM: The proposed architecture integrates causal convolutions into the self-attention mechanism to capture local and long-term dependencies. T-LAM enhances locality modeling through its convolutional self-attention mechanism, making it particularly suitable for the language challenges of Tibetan.
[0112] Key configurations: Four Transformer encoder layers with 8 attention heads, a model dimension of 256, and three convolutional layers using filter sizes [3, 5, 7], [3, 5, 7], [3, 5, 7] and 128 filters.
[0113]
[0114] Table 1: Summary of baseline models and configurations
[0115] Rationale for T-LAM
[0116] The baseline models were chosen to highlight specific features that T-LAM integrates or improves:
[0117] Local feature extraction: While CNNs and BiLSTMs are good at capturing local patterns, they struggle to effectively model global dependencies. T-LAM overcomes this limitation by incorporating causal convolutions into the self-attention mechanism.
[0118] Global context modeling: Models like RoBERTa-Ti and CINO are effective at capturing long-range dependencies but are less adept at modeling fine-grained local features, especially in morphologically complex languages like Tibetan. T-LAM addresses this gap through its convolutional self-attention mechanism.
[0119] Hybrid approach: T-LAM leverages the strengths of CNNs and Transformers. Its convolutional self-attention mechanism uses causal convolutions to generate queries and keys, enhancing sensitivity to local context while retaining the ability to model global dependencies. The evaluation framework highlights T-LAM's contributions relative to traditional and state-of-the-art models, emphasizing its ability to handle the language and computational challenges of Tibetan rumor detection. As Figure 1 shown, the design of T-LAM enables the efficient integration of short-range and long-range dependencies, making it a robust solution for low-resource language tasks. Table 1 summarizes the configurations of the baseline models. This structured comparison allows for a rigorous evaluation of T-LAM, demonstrating its advantages in local feature modeling and global context understanding for the complex task of Tibetan rumor detection.
[0120] 4.2 Experimental details
[0121] 4.2.1 Dataset and preprocessing
[0122] The Rumor_Ti dataset was created specifically for this study and consists of 4,477 labeled Tibetan samples, carefully curated to reflect real-world rumor scenarios. The dataset is balanced, with 2,366 non-rumor samples (labeled 1) and 2,111 rumor samples (labeled 0). To further enhance the model's ability to detect different types of misinformation, the rumor samples are divided into three different types: scientific rumors, social rumors, and other rumors. This classification enriches the dataset and helps the model handle different rumor contexts more effectively.
[0123] Rumor categories:
[0124] Scientific rumors: These rumors include false claims related to health, technology, and scientific progress, such as misleading medical information or misunderstood research findings. Keywords such as "medicine", "technology", and "research" are used to identify these samples.
[0125] Social rumors: These rumors involve misinformation about social, cultural, or interpersonal issues, often causing social tension or misunderstanding. Keywords such as "festival", "community", and "heritage" are used for classification.
[0126] Other rumors: This category covers misinformation that does not fall into the scientific or social categories, such as conspiracy theories or general false information. It provides a broad test bed for the model's generalization capabilities.
[0127] Dataset partitioning: To ensure balanced evaluation, the dataset is divided into a training set (3134 samples), a validation set (671 samples), and a test set (672 samples), maintaining the ratio of rumor and non-rumor data. This setting ensures fair comparison during training and evaluation.
[0128] Each subset maintains the proportional distribution of rumor and non-rumor samples, ensuring consistent data representation across the training, validation, and testing phases.
[0129] Preprocessing pipeline:
[0130] The preprocessing pipeline is designed to handle the unique language properties of Tibetan text while ensuring high data quality. The steps include:
[0131] Noise reduction: Removing irrelevant or noisy data, such as special characters or malformed text, while preserving meaningful language patterns.
[0132] Tokenization: Using a custom Tibetan tokenizer optimized for the CINO pre-trained architecture to split the text into meaningful units while preserving the grammar structure of the language.
[0133] Sample Validation: Remove invalid samples (e.g., incomplete or incorrectly formatted sentences) to ensure high linguistic and contextual quality in the dataset.
[0134] Utility of Dataset Classification:
[0135] Categorizing rumor samples into science, society, and other types provides necessary details for training. By allowing the model to focus on class-specific features, this approach enhances its ability to capture the nuances of misinformation in different contexts:
[0136] Science Rumors: These utilize keyword-based patterns and structured semantics, which are highly consistent with the global dependency modeling of Transformer-based architectures.
[0137] Social Rumors: These require a more context-sensitive approach to capture cultural nuances, which benefits from the local perception mechanism of T-LAM.
[0138] Other Rumors: This category tests the adaptability and generalization ability of the model by including a variety of and often ambiguous misinformation patterns.
[0139] Data Preparation for T-LAM: The preprocessing pipeline is tailored to support the hybrid convolutional self-attention mechanism of T-LAM. By integrating causal convolutions into the query and key generation process, T-LAM can capture local and global context information. This enables the model to address the linguistic and contextual challenges inherent in Tibetan text, providing a solid foundation for model training and evaluation. This comprehensive approach ensures that T-LAM can handle the complexities of Tibetan rumor detection well, using a dataset that reflects real-world challenges while supporting advanced modeling techniques.
[0140] 4.3 Comparative Experiments
[0141] To evaluate the effectiveness and robustness of the T-LAM model, extensive comparative experiments were conducted with several state-of-the-art baseline models. The evaluation used five standard performance metrics: accuracy, precision, recall, F1-score, and the area under the receiver operating characteristic curve (AUC-ROC). The baseline models selected for comparison included traditional deep learning architectures (CNN and BiLSTM), Transformer-based models (Transformer), and recent specialized models (RoBERTa-Ti and CINO). Table 3 shows a comprehensive comparison of the experimental results of all models in the test set.
[0142] Performance of T-LAM
[0143] Performance of T-LAM
[0144] The T-LAM model demonstrates competitive performance with an AUC-ROC score as high as 93.5%, showcasing its strong ability to distinguish rumors from non-rumors in Tibetan texts. Although the accuracy of 86.4% is slightly lower than that of the best-performing Bert-LAM (89.6%), T-LAM achieves a balanced accuracy (84.2%) and recall (87.1%), with an F1 score of 85.6%, indicating its robustness in handling false positives and false negatives. These results highlight the potential of the model for reliable real-world applications.
[0145] 2. Benefits of the Hybrid Architecture
[0146] T-LAM combines the advantages of convolutional layers and Transformer encoders, effectively capturing local and global features in Tibetan texts:
[0147] Convolutional layers: These layers are adept at capturing local language patterns such as n-grams and suffix-based morphology, which are crucial for the complex grammar structures of Tibetan. This feature improves the model's accuracy in identifying subtle features vital for rumor classification.
[0148] Transformer encoders: By capturing long-range dependencies, Transformer encoders ensure that the model comprehensively understands the broader context within a sentence, which is essential for processing Tibetan texts.
[0149] 3. Comparison with Baseline Models
[0150] CNN: The accuracy reaches 80.8% and the AUC-ROC is 89.0%. Although it performs well in capturing local features, its inability to capture global dependencies limits its overall performance.
[0151] BiLSTM: Demonstrates the lowest accuracy (74.3%) and AUC-ROC (82.2%) among all models. Despite its sequential modeling capabilities, BiLSTM still struggles with the complex morphology and resource-poor nature of Tibetan.
[0152] Transformer: The independent Transformer has a precision of 81.3% and an AUC-ROC of 89.1%, with moderate performance. However, its lack of convolutional layers reduces its ability to effectively capture local features.
[0153] RoBERTa-Ti and CINO: Both models performed competitively, with accuracy scores of 85.6% and 84.9% respectively, and AUC-ROC values of 90.2% and 91.1% respectively. Their pre-training on large-scale multilingual datasets enhanced generalization. However, T-LAM outperformed them in terms of recall and AUC-ROC, indicating its superior effectiveness in capturing true positives.
[0154] Bert-LAM: As the model with the best performance in terms of accuracy (89.6%) and AUC-ROC (95.1%), Bert-LAM demonstrated excellent performance. It utilized advanced pre-training techniques and specialized features, giving it an edge over other models.
[0155] As Figure 3 shown, T-LAM had an AUC-ROC score of 93.5%, demonstrating its effectiveness in distinguishing rumors from non-rumors. This performance is particularly notable considering the challenges associated with resource-scarce languages such as Tibetan. The model's architecture combines convolutional layers for local feature extraction and Transformer encoders for capturing global dependencies, enabling it to effectively handle the subtle features of Tibetan text.
[0156] In the context of rumor detection, accuracy and recall are key metrics as they reflect the model's ability to accurately identify rumors while minimizing false positives (unnecessary alerts) and false negatives (missed rumors). T-LAM achieved an accuracy of 84.2% and a recall of 87.1%, indicating that its hybrid structure effectively addressed the linguistic complexity of Tibetan. The model balanced the trade-off between precision and recall, enhancing its reliability in accurately detecting Tibetan rumors without significantly affecting either metric.
[0157]
[0158]
[0159] Table 2: Comparison results of T-LAM and baseline models on the test set
[0160] 2. Benefits of the hybrid architecture
[0161] T-LAM combines the advantages of convolutional layers and Transformer encoders to capture local and global features in Tibetan text, thereby improving its overall performance:
[0162] Convolutional Layers: The convolutional layers in T-LAM effectively capture local language features such as n-grams and suffixes, which are crucial for handling the complex grammar structures and suffix-based morphology of Tibetan. These layers help the model identify key text patterns that contribute to the accurate classification of rumors.
[0163] Transformer Encoder: The Transformer encoder captures long-range dependencies across the entire text, enabling T-LAM to understand the broader context of each sequence. This feature ensures that the model can integrate information from distant words, leading to a comprehensive understanding of Tibetan text, which is essential for accurate rumor detection.
[0164] 3. Comparison with Baseline Models
[0165] CNN: With an accuracy of 80.8% and an AUC-ROC of 89.0%, it shows the effectiveness of capturing local features but lacks the ability to integrate context information in longer sequences.
[0166] BiLSTM: The model has the lowest performance, with an accuracy of 74.3% and an AUC-ROC of 82.2%. Although BiLSTM can capture sequential dependencies, it struggles with the complex grammar structures of Tibetan and has difficulty generalizing in such resource-constrained environments.
[0167] Transformer: The standalone Transformer model achieves moderate performance with an accuracy of 81.3% and an AUC-ROC of 89.1%. While it captures long-range dependencies, the lack of convolutional layers limits its ability to effectively extract local context features.
[0168] RoBERTa-Ti and CINO: Both pre-trained models achieved competitive results, with accuracies of 85.6% and 84.9% respectively, and AUC-ROC values higher than 90%. These models benefit from large-scale pre-training on various multilingual datasets, enhancing their generalization ability. However, T-LAM outperforms them in terms of recall and AUC-ROC, indicating its higher efficiency in capturing true positives and minimizing false negatives.
[0169] 4. Effectiveness in Low-Resource Language Contexts
[0170] T-LAM's AUC-ROC score of 93.5% demonstrates its effectiveness in distinguishing rumors from non-rumors. This performance is particularly notable considering the challenges associated with low-resource languages such as Tibetan. The model's architecture combines convolutional layers for local feature extraction and a Transformer encoder for capturing global dependencies, enabling it to effectively handle the subtle features of Tibetan text.
[0171] In the context of rumor detection, accuracy and recall are key metrics as they reflect the model's ability to accurately identify rumors while minimizing false positives (unnecessary alerts) and false negatives (missed rumors). T-LAM achieved an accuracy of 84.2% and a recall of 87.1%, indicating that its hybrid structure effectively addressed the linguistic complexity of Tibetan. The model balanced the trade-off between precision and recall, enhancing its reliability in accurately detecting rumors without significantly affecting either metric.
[0172] 4.4 Ablation Experiments
[0173] To rigorously evaluate T-LAM's architectural decisions, a comprehensive ablation study was conducted following a predefined experimental protocol. This systematic investigation examined the contributions of three key components: (1) the convolutional feature extractor, (2) Transformer-based context modeling, and (3) CINO pre-trained embeddings. By conducting controlled experiments, the impact of each component on the model's ability to capture local linguistic features and global context dependencies necessary for Tibetan rumor detection was measured.
[0174] 4.4.1 Experimental Setup
[0175] Four variants of the model were constructed for the ablation study:
[0176] 1. The full T-LAM architecture (baseline)
[0177] 2. T-LAM without convolutional layers (Conv)
[0178] 3. T-LAM without the Transformer encoder (Trans)
[0179] 4. T-LAM without CINO pre-trained embeddings (CINO)
[0180] All experiments maintained consistent hyperparameters across variants to ensure fair comparison: a batch size of 16, a learning rate of 2×10 -5 , and a maximum sequence length of 512 tokens. Each variant was trained for 50 epochs using the Adam optimizer with a weight decay of 0.01.
[0181]
[0182] Table 3: Results of the ablation experiments
[0183] 4.4.2 Quantitative Analysis:
[0184] The impact of each component on T-LAM's performance was analyzed based on the results in Table 3.
[0185] Convolutional Feature Extractor: Removing the convolutional layers led to a moderate performance degradation. The accuracy decreased by 2.6 percentage points (from 84.2% to 81.6%), the F1-score decreased by 3.1 percentage points (from 85.6% to 82.5%). Additionally, the AUC-ROC decreased by 2.5 percentage points (from 93.5% to 91.0%). These results indicate that the convolutional layers significantly enhance the model's ability to capture local n-gram features and morphological patterns unique to Tibetan text, although they are not essential for the overall architecture.
[0186] Transformer Encoder: Removing the Transformer encoder resulted in a slight performance degradation. The precision decreased by 0.6 percentage points (from 84.2% to 83.6%), while the F1-score remained unchanged. The AUC-ROC decreased by 1.1 percentage points (from 93.5% to 92.4%). These findings suggest that while the Transformer encoder is crucial for modeling long-range dependencies, the convolutional layers partially compensate for its absence by focusing on local features.
[0187] CINO Pre-trained Embeddings: Excluding the CINO embeddings led to the most significant performance degradation across all metrics. The accuracy decreased by 14.0 percentage points (from 86.4% to 72.4%), while the precision decreased by 12.2 percentage points (from 84.2% to 72.0%). The recall and F1-score decreased by 13.1 and 12.6 percentage points respectively, and the AUC-ROC decreased by 13.5 percentage points (from 93.5% to 80.0%). These results emphasize the crucial importance of language-specific pre-trained embeddings for effectively capturing the nuances of the Tibetan language in resource-constrained settings.
[0188] 4.4.3 Component Impact Analysis:
[0189] The ablation study revealed several insights into the architectural advantages and dependencies of T-LAM:
[0190] 1. Convolutional Feature Extractor: The convolutional layers enhance the model's proficiency in capturing local features and sequential patterns, which are crucial for managing complex suffixes and syntactic structures in Tibetan text. Excluding these layers led to a slight but noticeable decrease in all performance metrics, highlighting their value in enriching the model's representation of local language features.
[0191] 2. Transformer Encoder: The Transformer encoder plays a crucial role in capturing long-range dependencies. Removing it reduces the model's ability to handle extended context information, thus affecting the recall and AUC-ROC. However, the convolutional layers, by focusing on local features, somewhat mitigate the absence of the encoder, resulting in only a moderate performance reduction.
[0192] 3. CINO Pre-trained Embeddings: Excluding CINO pre-trained embeddings led to the most significant decline in all metrics, highlighting the importance of language-specific embeddings for achieving high accuracy, precision, recall, and AUC-ROC in resource-constrained environments. These embeddings provided a fundamental understanding of the Tibetan language nuances, which were crucial for accurately processing Tibetan texts in the rumor detection task.
[0193] 4.4.4 Impact:
[0194] The ablation study revealed key insights into the architectural advantages and dependencies of the T-LAM model:
[0195] The hierarchical feature extraction strategy integrated convolutional layers with the Transformer-based attention mechanism, effectively leveraging the complementary advantages of both components.
[0196] After removing specific components (excluding CINO embeddings), the performance degradation was relatively small, indicating a certain degree of redundancy in the architecture, thus enhancing its robustness.
[0197] The crucial role of pre-trained embeddings highlighted the importance of language-specific pre-training for resource-constrained languages, especially for morphologically complex languages like Tibetan.
[0198] Table 3 emphasized the different contributions of each T-LAM component. While convolutional layers enhanced the model's ability to capture local features, the Transformer encoder captured extended dependencies. However, CINO embeddings were the foundation for achieving robust and accurate performance, especially in resource-constrained environments.
[0199] 4.5 Stability Experiments:
[0200] To evaluate the stability and robustness of the T-LAM model, we conducted a series of experiments in five independent runs, each initialized with a different random seed. The main purpose of these experiments was to assess the consistency of the model's performance under different initialization conditions, thus ensuring its reliability in different training scenarios.
[0201] Result Analysis
[0202] Table 4 shows the average performance metrics of T-LAM across five experimental runs, along with their respective standard deviations. The metrics evaluated included Accuracy, Precision, Recall, F1Score, and AUC-ROC.
[0203] The standard deviations observed in all metrics are very low, indicating that T-LAM can consistently provide high performance regardless of the specific initialization. For example, the standard deviations of accuracy (±1.3), precision (±2.4), recall (±2.2), F1-score (±1.4), and AUC-ROC (±0.9) suggest a high level of stability in the model's training and evaluation processes.
[0204] Implications for practical applications
[0205] The demonstrated performance consistency is particularly important for practical applications because changes in training data or environmental conditions are inevitable in real-world scenarios. The results show that T-LAM can maintain stable and reliable performance even when the initialization changes. This robustness makes T-LAM a reliable model for Tibetan rumor detection tasks, capable of providing consistent and reliable outputs despite small fluctuations in the training scenarios. This stability is especially valuable in resource-constrained environments, as reliable model behavior is crucial for effective deployment and operational reliability.
[0206]
[0207] Table 4: Results of the stability experiment
[0208] Result analysis:
[0209] Table 4 shows the average performance metrics of T-LAM over five runs, along with their respective standard deviations. These metrics include Accuracy, Precision, Recall, F1Score, and AUC-ROC.
[0210] The low standard deviations observed in all metrics indicate that T-LAM can consistently achieve high performance regardless of the specific initialization. Specifically, the standard deviations of accuracy (±1.3), precision (±2.4), recall (±2.2), F1-score (±1.4), and AUC-ROC (±0.9) reflect a high level of robustness in the model's training and evaluation processes.
[0211] Implications for practical applications:
[0212] This performance consistency is crucial for practical applications because changes in training data or environmental conditions are common. The results show that T-LAM can maintain stable and reliable performance even when the training initialization changes. Therefore, T-LAM can be considered a reliable model for Tibetan rumor detection tasks, capable of providing strong performance even with small fluctuations in the training scenarios. This stability is particularly valuable in resource-constrained environments, as reliable and consistent model behavior is essential for effective deployment.
[0213] This invention introduces T-LAM, a hybrid model that combines a Convolutional Neural Network (CNN) with a Transformer-based self-attention mechanism to address the specific challenges of Tibetan rumor detection. By leveraging the CNN and Transformer components, T-LAM effectively captures local semantic features and long-range dependencies, which are crucial for handling the complex syntactic structures and resource-poor nature of the Tibetan language. Integrating CINO embeddings designed specifically for minority languages further enhances the performance of T-LAM, achieving an accuracy of 86.4%, outperforming existing models such as CNN, BiLSTM, Transformer, RoBERTa-Ti, and CINO. Future research may focus on integrating other pre-trained models, such as monolingual Tibetan or cross-lingual embeddings, to improve semantic understanding and accuracy. Expanding the Rumor_Ti dataset to include a wider range of rumor types and dialects will also enhance the robustness of the model. Transfer learning from related languages can improve the performance of Tibetan and other low-resource languages. Exploring the applicability of Ti-TranCNN to other resource-poor languages may further expand its impact. Improving the interpretability of Ti-TranCNN is also crucial for providing clearer insights into key language features, enhancing the transparency and trust in model predictions.
[0214] In summary, T-LAM represents a step forward in Tibetan rumor detection, providing a solid foundation for further research aimed at improving the detection capabilities of low-resource languages. This work has the potential to positively impact the Tibetan community by addressing the urgent need for accurate rumor detection and providing a scalable model that can be adapted to other minority languages facing similar NLP challenges.
[0215] This paper presents T-LAM, a novel hybrid architecture that combines a Convolutional Neural Network (CNN) and a Transformer-based self-attention mechanism to tackle the unique challenges of Tibetan rumor detection. By using the CNN to capture localized semantic features and the Transformer to model long-term dependencies, T-LAM addresses the unique syntactic complexity and morphological richness of the Tibetan language, which is characterized by its non-canonical structures and extensive use of suffixes. The incorporation of CINO embeddings designed explicitly for minority languages further enhances the effectiveness of the model by providing a rich representation of language features, including morphological, syntactic, and semantic. T-LAM achieves an accuracy of 89.6%, significantly outperforming existing models such as CNN, BiLSTM, Transformer, RoBERTa-Ti, and CINO, highlighting its ability to overcome the limitations of current methods in low-resource NLP tasks.
[0216] Future research directions
[0217] Based on the progress made in this study, several promising future research directions have been identified: Integrating other pre-trained models: Incorporating monolingual Tibetan embeddings or cross-lingual embeddings can deepen the model's semantic understanding, especially for capturing the nuances unique to Tibetan grammar and word usage. Pre-trained models, such as those tailored for morphologically rich or resource-poor languages, can enhance the T-LAM's context understanding, thereby improving accuracy and robustness. Dataset expansion: The Rumor_Ti dataset is the fundamental resource for this study, but expanding it to include a more diverse range of rumor types and dialectal variations will enhance the model's ability to generalize across different language contexts. For example, incorporating rumors from different Tibetan dialects or cultural domains will ensure the model's effective performance in various real-world scenarios. Transfer learning: Leveraging knowledge transfer from related languages, such as Dzongkha or other Sino-Tibetan languages, can improve the model's ability to handle low-resource languages with linguistic similarities. This approach can serve as a cost-effective strategy to enhance performance without requiring a large labeled dataset for each target language. Broader applicability: Exploring the adaptability of T-LAM to other low-resource languages with similar language challenges, such as complex syntactic structures or rich morphological systems, may expand its impact. By demonstrating its effectiveness beyond Tibetan, T-LAM can become a benchmark model for addressing resource-poor NLP challenges in a wider range of languages. Interpretability: Enhancing the interpretability of T-LAM is crucial for providing actionable insights into the linguistic features influencing the model's predictions. Techniques such as attention visualization or feature importance analysis can shed light on how the model identifies rumors, thereby improving the transparency and trustworthiness of its outputs. This is particularly important in sensitive applications where understanding the reasons behind predictions is essential for user trust and ethical deployment. In summary, T-LAM represents a significant advancement in the field of Tibetan rumor detection, offering a powerful solution to the challenges posed by resource-poor languages. This work lays a solid foundation for further research aimed at improving the detection capabilities of Tibetan and other minority languages. By addressing the pressing need for accurate rumor detection, T-LAM has the potential to have a meaningful impact on Tibetan-speaking communities. Additionally, its scalability and adaptability make it a promising tool for addressing similar NLP challenges in other minority languages.
[0218] The present invention is not limited to the above embodiments. Anyone should be aware that structural changes made under the inspiration of the present invention, as long as they have the same or similar technical solutions as the present invention, fall within the protection scope of the present invention. The technologies, shapes, and structures not described in detail in the present invention are all well-known technologies.
Claims
1. A Tibetan rumor detection method, characterized in that, It includes the following steps: S1: Use the CINO pre-trained model to encode the Tibetan text to generate an initial representation vector of the Tibetan text; S2: Adopt a convolutional neural network to extract features from the initial representation vector to obtain local semantic features; S3: Use the multi-head self-attention mechanism to process the local semantic features to obtain global context information; S4: Perform adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation; S5: Based on the fused feature representation, perform rumor detection classification through a fully connected layer and output the detection result.
2. The Tibetan rumor detection method according to claim 1, wherein The step of using the CINO pre-trained model to encode the Tibetan text to generate an initial representation vector of the Tibetan text is: S101: Input sequence x = {x1, x2,..., x n}, where n represents the sequence length, and each is a dimensional vector d. The goal is to predict a label y ∈ {0, 1}, where y = 0 represents rumor and y = 1 represents non-rumor; S102: For the preprocessed Tibetan text, use the CINO pre-trained model for encoding. Through model calculation, generate the initial representation vector h of the text sequence (0) , and its expression: h (0) = CINO(x), where h (0) is the initial semantic representation of the text sequence. In addition, encode the semantic and syntactic features crucial for downstream tasks. Finally, use the generated initial representation vector h (0) as the input for subsequent feature extraction and classification.
3. A Tibetan rumor detection method according to claim 1, characterized in that, The step of adopting a convolutional neural network to extract features from the initial representation vector to obtain local semantic features is: S201: Design and initialize convolutional kernels of multiple sizes to capture semantic features at the affix, phrase, and sentence levels respectively. Among them, the sizes of the convolutional kernels are [3, 5, 7]. Then, input the initial representation vector into the convolutional neural network for one-dimensional convolutional operations to capture n-gram features of different granularities respectively; S202: Each convolutional kernel's sliding window extracts features to generate a local feature map h (1) , then, from the local feature map h (1) query matrix Q, key matrix K, and value matrix V are generated, and their expressions are Q = h (1) * W q , K = h (1) * W k , V = h (1) * W v , where is a learnable weight matrix, k is the output feature dimension of the convolutional kernel, used to project the initial representation vector h (0) into the feature space, and * represents a 1D convolution operation; S203: Calculate the positional correlation weights within a local range using the query matrix Q and the key matrix K, and its expression is: where d k is the dimension of the key matrix, which is used to scale the correlation score to prevent the gradient from being unstable due to overly large values. QK T is the dot product of the query matrix Q and the key matrix K, generating an n×n correlation matrix; S204: Input the feature z generated by local attention into the feed-forward network (FFN), and further optimize the local feature expression through two-layer linear transformation and activation function. Its expression is: FFN(z) = ReLU(zW1 + b1)W2 + b2, where z represents the embedding of local participation and serves as the final local feature representation, W1 and W2 are weight matrices used for the first and second layer linear transformations respectively, b1 and b2 are the corresponding bias vectors, and z = Attention(Q, K, V); S205: Apply ReLU activation to the features FFN(z) generated by the feed-forward network to enhance the non-linear expression ability. Then, normalize the activated features. Finally, input the normalized feature matrix into the max-pooling layer to extract the significant features in each local region, and the local semantic feature h can be obtained. (l) 。 4. A Tibetan rumor detection method according to claim 1, characterized in that The step of using the multi-head self-attention mechanism to process the local semantic features to obtain global context information is: Generate the query matrix Q, key matrix K, and value matrix V of the multi-head self-attention mechanism based on the input local semantic features, for each attention head head i The specific calculation formula is: head i = Attention(Q i , K i , V i ) where Q i , K i , V are the i-th heads of the query, key, and value matrices respectively, and d k is the scaling factor; Concatenate the outputs of all attention heads and perform a linear transformation through the weight matrix W o to generate a representation h of the global context information (global) , and the combined formula is: MultiHead(Q, K, V) = concat(head1,..., head H )W O . where HHH is the number of attention heads and W o is a learnable weight matrix; The global feature representation h generated by the self-attention mechanism is further optimized using a feed-forward network (FFN). The feed-forward network consists of two layers of linear transformation and non-linear activation, and the specific formula is: Z1 = ReLU(h (global) W1 + b1), Z2 = Z1W2 + b2., where W1 and W2 are the weight matrices of the feed-forward network, b1 and b2 are the bias terms, and X is the self-attention output feature. (l) 5. A Tibetan rumor detection method according to claim 1, characterized in that, The step of performing adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation is: S401: Fuse the local semantic features and the optimized global context information Z2 with learnable weight parameters. The calculation formula is: h (fusion) = α · h (local) + β · Z2, where represents the fused feature representation, α controls the weight of the local features, β controls the weight of the global features, and satisfy: α + β = 1; S402: Use the fused feature h (fusion) as the input for the subsequent classification task, retaining the core semantics of local semantic features and global context information.
6. A Tibetan rumor detection method according to claim 1, characterized in that, The step of performing rumor detection classification through a fully connected layer based on the fused feature representation and outputting the detection result is: S501: Receive the fused features and classify them through the formula , where σ is the sigmoid activation function, w fc is the weight matrix of the final fully connected layer, and b fc is the bias term; Compare the predicted probability with a preset threshold τ and output the detection result: If then it is predicted as a rumor; if then it is predicted as non-rumor, where τ is the threshold for the classification task.
7. A training method for a Tibetan rumor detection model, characterized in that, It includes the following steps: Load the Rumor_Ti dataset, divide the dataset into a training set, a validation set, and a test set, and at the same time preprocess the data, including noise reduction, word segmentation, and sample verification, to ensure the data quality and the grammar structure of the language; Initialize the model parameters, including the weight matrices of the convolutional layer, the Transformer encoder layer, and the feed-forward network. Set the number of filters in the convolutional layer to 128, the convolutional kernel size to [3, 5, 7], configure 8 attention heads in the Transformer encoder layer, and the model dimension to 256; Input the preprocessed data into the model batch by batch, and generate the initial embedding feature h for each batch (0) , where the feature dimension is B×n×d, where B is the batch size, n is the sequence length, and d is the embedding feature dimension; Optimize the classification model using the binary cross - entropy (BCE) loss function, and its expression is: where N is the number of samples, y i is the true label of the sample, is the predicted probability of the sample; Use the Adam optimizer to update the model parameters, with an initial learning rate of 1e-3, combined with a weight decay of 1e-5 to avoid overfitting, calculate the gradient through the backpropagation algorithm, and gradually reduce the loss function value; Adopt the cosine annealing strategy to dynamically adjust the learning rate, so that the learning rate gradually decreases during the training process, thereby accelerating convergence and improving the model stability. Its expression is: Among them, η t is the learning rate of the current iteration, T is the total number of training rounds, η min and η max are the minimum and maximum values of the learning rate respectively; After each training epoch, use the validation set to evaluate the model performance and save the model weights with the lowest validation set loss or the highest accuracy; Stop training when the validation set loss does not decrease for 5 consecutive epochs to ensure that the model does not overfit or waste computing resources.
8. A Tibetan rumor detection device, characterized in that It includes Text encoding module: Use the CINO pre-trained model to encode the Tibetan text to generate an initial representation vector of the Tibetan text; Local feature extraction module: Use a convolutional neural network to extract features from the initial representation vector to obtain local semantic features; Global feature modeling module: Use the multi-head self-attention mechanism to process the local semantic features to obtain global context information; Feature fusion module: Perform adaptive weighted fusion of the local semantic features and the global context information with learnable weight parameters to obtain a fused feature representation; Classification module: Based on the fused feature representation, perform rumor detection classification through a fully connected layer and output the detection result.
Citation Information
Cited By
Tibetan language three-dialect parallel corpus data set generation method based on multi-dialect text-to-speech model
CN120708597A
A method for generating parallel corpus datasets of the three major Tibetan dialects based on a multi-dialect text-to-speech model
CN120708597B
Multi-weather robust target detection method and device for electric power inspection and storage medium
CN120807897A
Electric power engineering compliance file automatic generation system based on NLP big language model
CN120874854A
Method and device for improving text sequence single classification anomaly detection capability
CN121658655A