Multi-feature fusion rumor detection method, system and device based on knowledge distillation

Through knowledge distillation technology and hierarchical gated interactive fusion network, the dynamic coupling relationship between multiple features is explicitly modeled, which solves the problems of single feature representation and model complexity in rumor detection and achieves efficient and accurate rumor detection.

CN120763753APending Publication Date: 2025-10-10CHONGQING UNIVERSITY OF SCIENCE AND TECHNOLOGY +1
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510939041.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-08
Publication Date
2025-10-10

AI Technical Summary

Technical Problem

Existing rumor detection methods have limitations in single feature representation, inefficient multi-feature fusion strategies, lack of dynamic contextual association, and the contradiction between model complexity and practicality, resulting in low recognition accuracy, high false detection rate, and poor robustness, especially in complex rumor scenarios, which limits their practicality.

Method used

A multi-feature fusion method based on knowledge distillation is adopted. Through hierarchical gated interactive fusion and lightweight model migration technology, the dynamic coupling relationship between multimodal features is explicitly modeled. The knowledge migration and parameter compression of the teacher model and student network are combined to improve the detection accuracy and efficiency.

Benefits of technology

It significantly improves the accuracy and generalization ability of rumor detection, solves the problems of feature information dilution and model complexity in traditional methods, and realizes a lightweight real-time detection solution.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120763753A_ABST
    Figure CN120763753A_ABST
Patent Text Reader

Abstract

The invention provides a multi-feature fusion rumor detection method, system and device based on knowledge distillation, and mainly solves the problems that an existing model is high in calculation overhead, insufficient in feature fusion and insufficient in emotion utilization. The method comprises the steps of firstly obtaining multi-dimensional data such as social media original texts and comments; extracting deep semantic representation by using a pre-training model, and analyzing comment emotion features in combination with a hybrid neural network; then, features such as semantics, emotions, emoticons and populations are input into a hierarchical gating interactive fusion network (GIFN), and weights are dynamically adjusted to achieve effective fusion of multi-granularity features; in order to reduce complexity, a knowledge distillation framework is designed: a deep GIFN is used as a teacher network to generate a soft label, and a lightweight student network (LSTM) is guided to perform training. According to the trained student model, the parameter quantity is remarkably reduced, meanwhile, good detection performance is kept, the student model can be conveniently deployed in an actual content auditing system or edge equipment, and social content rumors can be efficiently recognized and judged.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of natural language processing and social media information detection, and provides a rumor detection method, system and device based on knowledge distillation and multi-feature fusion, specifically involving semantic mining of social media texts, dynamic fusion of multiple features, lightweight model inference technology, and model compression and classification performance optimization under the knowledge distillation framework. Background Art

[0002] In recent years, social media platforms have become the primary channel for rumor propagation. Their rapid spread and evolution pose a serious threat to social order and public perception. Therefore, efficiently identifying and detecting rumor information has become a critical issue that needs to be addressed. Compared to detection methods that rely on a single semantic feature, modeling the use of multiple features (such as text semantics, sentiment, communication structure, and user behavior) can more comprehensively capture the inherent patterns and propagation characteristics of rumors. Current research is increasingly focusing on developing detection methods that integrate multiple features, but existing technologies still have significant shortcomings.

[0003] Most existing rumor detection methods have the following main problems:

[0004] (1) Limitations of single feature representation: Traditional methods usually perform detection based on the surface semantics or keyword statistical features of the text, ignoring the impact of deep features such as sentiment polarity, propagation path, and user credibility on rumor discrimination, resulting in insufficient generalization ability of the model for complex rumor scenarios.

[0005] (2) Inefficient multi-feature fusion strategy: Existing methods often use shallow fusion methods such as direct splicing or weighted addition to fuse multiple features, which fail to explicitly model the dynamic dependencies between features (such as the synergy between sentiment tendency and communication structure), resulting in redundant noise interference and weakening the information utilization of key features.

[0006] (3) Lack of dynamic contextual association: Most fusion methods only adopt static early fusion or late fusion frameworks, without considering the temporal correlation between features during the rumor propagation process (such as the association between the evolution of user behavior and semantic diffusion), making it difficult to capture the dynamic laws of rumor propagation.

[0007] (4) The contradiction between model complexity and practicality: Although the multi-feature model based on deep learning has strong representation capabilities, it has the problems of large number of parameters and slow inference speed. It is difficult to deploy in real-time detection scenarios, which limits its practical application value.

[0008] These issues have led to bottlenecks in existing rumor detection technologies, such as low recognition accuracy, high false positive rates, and poor robustness. This severely limits the practicality of existing methods, especially when faced with highly concealed rumors with complex propagation paths. Therefore, there is an urgent need to design a rumor detection solution that combines deep multi-feature correlation mining with lightweight inference capabilities to improve detection accuracy and efficiency in complex scenarios. Summary of the Invention

[0009] To address the above-mentioned problems, the present invention aims to propose a multi-feature fusion rumor detection method, system, and device based on knowledge distillation. By combining hierarchical gated interactive fusion with lightweight model migration technology, it is possible to explicitly model the dynamic coupling relationship between multimodal features such as semantics, emotions, emoticons, and internet buzzwords, thereby breaking through the performance bottleneck of traditional single-path feature extraction and shallow fusion methods. At the same time, by introducing a knowledge distillation framework, while retaining the deep feature association pattern, knowledge transfer and parameter compression between the teacher model and the student network are achieved, alleviating the contradiction between detection accuracy and computational efficiency, and ultimately improving the accuracy, generalization ability, and practical deployment applicability of rumor detection.

[0010] In order to solve the above technical problems, the present invention adopts the following technical solutions:

[0011] A multi-feature fusion rumor detection method based on knowledge distillation, the method comprising:

[0012] (1) Obtain the data to be tested from social media, including original Weibo text, related comment data, Emoji emoticons, and Internet buzzwords;

[0013] (2) The pre-trained BERT model is used to extract semantic features from the original text and comments of Weibo, generating context-aware deep semantic representations. At the same time, based on the BiLSTM-CNN hybrid network structure, the sentiment polarity of the text is modeled and the sentiment semantic feature vector is output.

[0014] (3) The semantic features and sentiment features extracted in step 2 are input into the hierarchical gated interactive fusion network (GIFN) along with the emoji embedding vector and the internet buzzwords vocabulary vector. The cascaded gated units are used to dynamically adjust the weight ratios of different features, construct a nonlinear mapping relationship between feature dimensions, and generate a comprehensive representation that integrates multi-granularity features.

[0015] (4) Construct a teacher model based on GIFN and a student model based on BiLSTM network, and jointly train them through knowledge distillation technology. The soft label probability distribution, hidden layer feature response and attention weight matrix output by the teacher model serve as supervision signals to guide the student model to learn feature association patterns;

[0016] (5) Deploy the trained student model to the inference device to perform feature fusion, classification decision and output rumor detection results on the input social media data.

[0017] Preferably, according to the present invention, the specific steps of acquiring and preprocessing data in step 1 are:

[0018] (1) Obtain cross-platform social media data through generalized data collection tools;

[0019] (2) Noise filtering of comment texts, including stripping out independent emoticons and preset internet buzzwords, and performing semantic completion and context-based expansion mapping on low-quality texts;

[0020] (3) Based on the interaction time series information between user comments and original texts, we construct time window synchronized associated data pairs and annotate them with sentiment tendency labels to form a training set.

[0021] According to the preferred embodiment of the present invention, feature extraction is performed on different data in step 2, and the implementation process is as follows:

[0022] The pre-trained BERT model's multiple attention heads are used to concurrently extract contextual semantic features from the original Weibo text and associated comments. A multi-layer Transformer encoder is used to capture long-range dependencies in the text. Dynamic feature dimensionality reduction and hierarchical aggregation are then performed on the output layer embedding vector to generate a dimensionally aligned deep semantic vector.

[0023] Simultaneously, fine-grained sentiment analysis of the text is performed using a BiLSTM-CNN hybrid network:

[0024] The input text is mapped into a distributed representation through the word embedding layer, input into a bidirectional LSTM network to model forward and backward semantic associations, and outputs the hidden state vector of each position. A CNN sub-network with multi-scale convolution kernels is then used to extract the key sentiment features of the local text fragment and concatenate them with the BiLSTM hidden state vector across channels. A multi-head self-attention mechanism is then used to weight the importance of the concatenated features to generate a semantic sentiment vector that integrates long-term dependencies and local sentiment polarity.

[0025] Furthermore, a symbol-text joint embedding method is used for Emoji emoticons, and the implicit semantics of emoticons are modeled with polysemy based on the pre-trained language model; for Internet buzzwords, domain-adapted word vector fine-tuning technology is used, combined with attention weights to dynamically adjust their influence coefficient in emotional characteristics.

[0026] According to the preferred embodiment of the present invention, the specific implementation process of step 3 is as follows:

[0027] The emoji expression symbol is converted into a high-dimensional dense vector through a preset embedding layer, a pre-trained BERT model is used to extract a time sequence context representation of a network popular language, and L2 norm normalization is respectively performed;

[0028] Secondly, two-level gating structures are arranged in a hierarchical gated interactive fusion network (GIFN):

[0029] A primary gating unit: semantic features and emotional features are spliced, a different feature correlation score is calculated through a learnable parameter matrix, and a sentiment-enhanced semantic vector is generated;

[0030] A senior gating unit: the primary gating output is cross-channel spliced with an emoji vector and a popular language vector, a bidirectional dynamic weight distribution module with a Tanh activation is adopted, and a feature weight is calculated through the following formula:

[0031] α i =Softmax(W g T ·[f sem ;f emoji ;f slang ])

[0032] Wherein, W g is a trainable parameter matrix, and alpha i represents the contribution coefficient of each feature in the decision space;

[0033] Finally, a cross-entropy constraint is applied to the features between channels based on a sliding window mechanism, so that different features establish an inter-class discrimination boundary in the hidden space, and finally a fusion feature vector with a dimension of 768 is output for knowledge distillation.

[0034] According to the present application, the specific implementation process of step 4 is:

[0035] The teacher model is defined as a multi-layer attention Transformer architecture containing a hierarchical gated interactive fusion network (GIFN), and the student model is a lightweight network composed of a bidirectional long short-term memory network (BiLSTM) and a full connection layer;

[0036] And the following multi-task supervision objectives are used to realize knowledge distillation:

[0037] Soft label distillation loss: the Kullback-Leibler divergence is used to measure the difference between the soft label probability distribution P T output by the teacher model containing a temperature parameter tau and the probability P S output by the student model;

[0038] Hidden layer response matching loss: the cosine similarity is used to constrain the last hidden state hT With the student model BiLSTM output feature h S The distribution of aligns;

[0039] Attention weight alignment loss: By minimizing the teacher model self-attention matrix A T With the student model temporal attention matrix A S Frobenius norm distance, migration association pattern;

[0040] The teacher model is trained earlier than the student model, and the teacher model parameters are frozen during the joint training phase. The student model parameters are optimized only by the stochastic gradient descent algorithm, where the total loss function is

[0041]

[0042] in is the cross entropy loss of the student model on the true label, λ1, λ2, λ3 are the weight coefficients of adaptive adjustment;

[0043] The temperature parameter τ is dynamically adjusted during the distillation process. In the initial stage, τ = 5 is set to smooth the probability distribution, and it linearly decays to τ = 1 as the number of training rounds increases to enhance the discrimination of difficult samples.

[0044] According to the preferred embodiment of the present invention, the specific implementation process of step 5 is as follows:

[0045] Real-time preprocessing and standardized input are performed on the social media data stream to be detected. The deep semantic features of the text, emotional semantic features, Emoji expression vector representation and Internet buzzword context vector are simultaneously extracted through parallel computing paths. Hierarchical feature weights are calibrated and dynamically spliced ​​through the BiLSTM student network. The cascade neural network is used to perform nonlinear mapping on the fused heterogeneous features. The softmax activation function is used to output multi-classification probability distribution, and the final rumor category is determined according to the preset confidence threshold.

[0046] A multi-feature fusion rumor detection system based on knowledge distillation, used to implement the multi-feature fusion rumor detection method based on knowledge distillation described in any one of claims 1-6, characterized in that the system includes: a data acquisition and preprocessing module, a feature extraction module, a hierarchical gated fusion network module (GIFN), and a knowledge distillation and reasoning module.

[0047] The data acquisition and preprocessing module is used to: collect original Weibo texts, comment sequences, Emoji emoticons and Internet buzzwords in social media in real time, and perform denoising and cleaning, normalized encoding and input format unification operations; the feature extraction module is used to: integrate pre-trained BERT to extract deep semantic features of text, BiLSTM-CNN hybrid network to capture emotional semantic features, and simultaneously map Emoji symbols and Internet buzzwords into low-dimensional vector representations through the word embedding layer; the hierarchical gated fusion network module (GIFN) is used to: dynamically weight feature weights through gating units, and use nonlinear activation functions and attention mechanisms to achieve cross-dimensional deep feature interaction; the knowledge distillation and reasoning module is used to: integrate the knowledge distillation framework of the GIFN teacher model and the BiLSTM student model, perform soft label probability supervision and parameter compression training, and deploy a lightweight student model to achieve real-time feature fusion classification.

[0048] A multi-feature fusion rumor detection device based on knowledge distillation, wherein the device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. At the same time, when the processor executes the computer program, it implements the steps of the multi-feature fusion rumor detection method or system based on knowledge distillation described in any one of the above technical solutions.

[0049] Beneficial effects of the present invention:

[0050] 1. The present invention proposes a rumor detection method based on multi-feature fusion, which constructs event representation from four dimensions: semantic features, emotional features, emoticon features, and Internet buzzword features. It breaks through the limitation of traditional feature extraction methods that rely on single text semantic analysis, and significantly improves the information coverage of rumor discrimination.

[0051] 2. This paper designs a hierarchical gated interactive fusion network (GIFN), which dynamically adjusts the weight distribution of multi-dimensional features through cascaded gating units and constructs nonlinear coupling relationships across dimensions. Compared with existing shallow weighted fusion methods, this method can explicitly model the dynamic coordination mechanism between multiple features, solving the problem of feature information dilution caused by static weight settings in traditional fusion methods.

[0052] 3. This invention innovatively integrates deep feature fusion and lightweight transfer learning techniques: a teacher model based on GIFN is constructed to model deep feature associations, and knowledge distillation is used to transfer the fusion model to a lightweight student network. This dual knowledge transfer architecture preserves the interpretability of deep feature crosstalk while effectively reducing the number of parameters through a lightweight model structure. Compared to traditional staged approaches that combine feature fusion and model compression, this significantly improves inference speed while maintaining detection accuracy, providing a more feasible solution for real-time rumor detection in resource-constrained scenarios. BRIEF DESCRIPTION OF THE DRAWINGS

[0053] Figure 1 It is a schematic diagram of the process provided by the present invention;

[0054] Figure 2 It is a schematic diagram of the overall structure of the system provided by the present invention;

[0055] Figure 3 It is a schematic diagram of the GIFN structure provided by the present invention. DETAILED DESCRIPTION

[0056] The technical solutions in the embodiments of the present invention are described clearly and completely below. Obviously, the described embodiments are only some preliminary embodiments of the present invention, rather than all embodiments.

[0057] It should be understood that the step numbers used herein are only for the convenience of description and are not intended to limit the order in which the steps are executed.

[0058] It should be understood that the terms used in the present specification are only for the purpose of describing specific embodiments and are not intended to limit the present invention. As used in the present specification and the appended claims, the singular forms "a", "an" and "the" are intended to include the plural forms as well, unless the context clearly indicates otherwise.

[0059] The term "and / or" refers to and includes any and all possible combinations of one or more of the associated listed items.

[0060] Example 1

[0061] A multi-feature fusion rumor detection method based on knowledge distillation, such as Figure 1 As shown, the following steps are included:

[0062] (1) Obtain the data to be tested from social media, including original Weibo text, related comment data, Emoji emoticons, and Internet buzzwords;

[0063] (2) The pre-trained BERT model is used to extract semantic features from the original text and comments of Weibo, generating context-aware deep semantic representations. At the same time, based on the BiLSTM-CNN hybrid network structure, the sentiment polarity of the text is modeled and the sentiment semantic feature vector is output.

[0064] (3) The semantic features and sentiment features extracted in step 2 are input into the hierarchical gated interactive fusion network (GIFN) along with the emoji embedding vector and the internet buzzwords vocabulary vector. The cascaded gated units are used to dynamically adjust the weight ratios of different features, construct a nonlinear mapping relationship between feature dimensions, and generate a comprehensive representation that integrates multi-granularity features.

[0065] (4) Construct a teacher model with GIFN as the core and a student model based on the BiLSTM network, and jointly train the two through knowledge distillation technology, wherein the soft label probability distribution, hidden layer feature response and attention weight matrix output by the teacher model are used as a supervision signal to guide the student model to learn the feature association mode;

[0066] (5) Deploy the trained student model to the inference device to perform feature fusion, classification decision and output rumor detection results on the input social media data.

[0067] Embodiment 2

[0068] The rumor detection method based on knowledge distillation according to Embodiment 1 is different in that:

[0069] In step (1), the data acquisition and preprocessing means: acquiring cross-platform social media data through a general data acquisition tool; filtering noise from the comment text, including stripping independent emoticons and preset popular network phrase segments, and performing semantic completion and context association-based expansion mapping on low-quality text; according to the interactive timing information of user comments and the original text, constructing a time window synchronized associated data pair, and labeling the sentiment tendency label to form a training set.

[0070] Embodiment 3

[0071] The rumor detection method based on knowledge distillation according to Embodiment 1 or 2 is different in that:

[0072] The specific implementation process of feature extraction on different data in step (2) is as follows:

[0073] The multiple attention heads of the pre-trained BERT model are used to extract the context semantic features of the microblog original text and the associated comments in parallel, the long-range dependency relationship in the text is captured through the multi-layer Transformer encoder, and the output layer embedding vector is dynamically reduced in dimension and aggregated in level to generate a deep semantic vector with aligned dimensions;

[0074] Synchronously, the text is analyzed in detail through the BiLSTM-CNN hybrid network:

[0075] The input text is mapped to a distributed representation through the word embedding layer, and the forward and backward semantic associations are modeled through the bidirectional LSTM network, and the hidden state vector at each position is output; then the CNN subnetwork with multiple scale convolution kernels is used to extract the key sentiment features of the local text segment, and the hidden state vector of the BiLSTM is cross-channel spliced; then the multi-head self-attention mechanism is used to weight the importance of the spliced features, and a semantic sentiment vector that integrates long-term dependency and local sentiment polarity is generated;

[0076] Furthermore, a symbol-text joint embedding method is used for Emoji emoticons, and the implicit semantics of emoticons are modeled with polysemy based on the pre-trained language model; for Internet buzzwords, domain-adapted word vector fine-tuning technology is used, combined with attention weights to dynamically adjust their influence coefficient in emotional characteristics.

[0077] Example 4

[0078] The difference between the multi-feature fusion rumor detection method based on knowledge distillation described in Example 1, 2, or 3 is that:

[0079] The specific implementation process of step (3) is as follows:

[0080] The Emoji emoticons are converted into high-dimensional dense vectors through a preset embedding layer, and the temporal context representation of Internet buzzwords is extracted using a pre-trained BERT model, and L2 norm normalization is performed on each.

[0081] Secondly, a two-level gating structure is set in the hierarchical gated interactive fusion network (GIFN):

[0082] Primary gating unit: concatenates semantic features with sentiment features, calculates the correlation scores of different features through a learnable parameter matrix, and generates sentiment-enhanced semantic vectors;

[0083] Advanced Gating Unit: Cross-channel concatenation of the primary gating output with the Emoji vector and buzzword vector. A bidirectional dynamic weight allocation module with Tanh activation is used to calculate the feature weight using the following formula:

[0084]

[0085] Where W g is the trainable parameter matrix, α i Represents the contribution coefficient of each feature in the decision space;

[0086] Finally, a cross-entropy constraint is imposed on inter-channel features based on the sliding window mechanism, forcing different features to establish inter-class discrimination boundaries in the latent space, and finally outputting a fused feature vector with a dimension of 768 for knowledge distillation.

[0087] Example 5

[0088] The difference between the multi-feature fusion rumor detection method based on knowledge distillation described in Example 1, 2, 3 or 4 is that:

[0089] The specific implementation process of step (4) is as follows:

[0090] The teacher model is defined as a multi-layer attention Transformer architecture containing a hierarchical gated interactive fusion network (GIFN), and the student model is a lightweight network consisting of a bidirectional long short-term memory network (BiLSTM) and a fully connected layer;

[0091] Knowledge distillation is then achieved through the following multi-task supervision objectives:

[0092] Soft label distillation loss: Kullback-Leibler divergence is used to measure the soft label probability distribution P output by the teacher model with temperature parameter τ T And the student model output probability P S differences;

[0093] Hidden layer response matching loss: Use cosine similarity to constrain the teacher model GIFN's last hidden state h T With the student model BiLSTM output feature h S The distribution of aligns;

[0094] Attention weight alignment loss: By minimizing the teacher model self-attention matrix A T With the student model temporal attention matrix A S Frobenius norm distance, migration association pattern;

[0095] The teacher model is trained earlier than the student model, and the teacher model parameters are frozen during the joint training phase. The student model parameters are optimized only by the stochastic gradient descent algorithm, where the total loss function is

[0096]

[0097] in is the cross entropy loss of the student model on the true label, λ1, λ2, λ3 are the weight coefficients of adaptive adjustment;

[0098] The temperature parameter τ is dynamically adjusted during the distillation process. In the initial stage, τ = 5 is set to smooth the probability distribution, and it linearly decays to τ = 1 as the number of training rounds increases to enhance the discrimination of difficult samples.

[0099] Example 6

[0100] The difference between the multi-feature fusion rumor detection method based on knowledge distillation described in Example 1, 2, 3, 4 or 5 is that:

[0101] The specific implementation process of step (5) is as follows:

[0102] Real-time preprocessing and standardized input are performed on the social media data stream to be detected. The deep semantic features of the text, emotional semantic features, Emoji expression vector representation and Internet buzzword context vector are simultaneously extracted through parallel computing paths. Hierarchical feature weights are calibrated and dynamically spliced ​​through the BiLSTM student network. The cascade neural network is used to perform nonlinear mapping on the fused heterogeneous features. The softmax activation function is used to output multi-classification probability distribution, and the final rumor category is determined according to the preset confidence threshold.

[0103] Example 7

[0104] A multi-feature fusion rumor detection system based on knowledge distillation includes: a data acquisition and preprocessing module, a feature extraction module, a hierarchical gated fusion network module (GIFN), and a knowledge distillation and reasoning module.

[0105] The data acquisition and preprocessing module is used to: collect original Weibo texts, comment sequences, Emoji emoticons and Internet buzzwords in social media in real time, and perform denoising and cleaning, normalized encoding and input format unification operations; the feature extraction module is used to: integrate pre-trained BERT to extract deep semantic features of text, BiLSTM-CNN hybrid network to capture emotional semantic features, and simultaneously map Emoji symbols and Internet buzzwords into low-dimensional vector representations through the word embedding layer; the hierarchical gated fusion network module (GIFN) is used to: dynamically weight feature weights through gating units, and use nonlinear activation functions and attention mechanisms to achieve cross-dimensional deep feature interaction; the knowledge distillation and reasoning module is used to: integrate the knowledge distillation framework of the GIFN teacher model and the BiLSTM student model, perform soft label probability supervision and parameter compression training, and deploy a lightweight student model to achieve real-time feature fusion classification.

[0106] In this embodiment, the specific implementation of the data acquisition and preprocessing module is as follows:

[0107] (1) Obtain cross-platform social media data through generalized data collection tools;

[0108] (2) Noise filtering of comment texts, including stripping out independent emoticons and preset internet buzzwords, and performing semantic completion and context-based expansion mapping on low-quality texts;

[0109] (3) Based on the interaction time series information between user comments and original texts, we construct time window synchronized associated data pairs and annotate them with sentiment tendency labels to form a training set.

[0110] In this embodiment, the feature extraction module is specifically implemented as follows:

[0111] The pre-trained BERT model is used to extract the context semantic features of the microblog original text and the associated comments in parallel through multiple attention heads, capture the long-range dependencies in the text through a multi-layer Transformer encoder, and perform dynamic feature dimension reduction and hierarchical aggregation on the output layer embedding vectors to generate a deep semantic vector with aligned dimensions.

[0112] Simultaneously, fine-grained sentiment analysis is performed on the text through a BiLSTM-CNN hybrid network:

[0113] The input text is mapped to a distributed representation through a word embedding layer, input into a bidirectional LSTM network to model forward and backward semantic associations, and output hidden state vectors at each position. Then, a CNN subnetwork with multiple scale convolution kernels is used to extract key sentiment features of local text segments, and the hidden state vectors of BiLSTM are cross-channel spliced. Then, a multi-head self-attention mechanism is used to weight the importance of the spliced features, generating semantic sentiment vectors that integrate long-term dependencies and local sentiment polarity.

[0114] Further, a symbol-text joint embedding method is used for Emoji emoticons, and the implicit semantics of the emoticons are modeled based on a pre-trained language model. Network slang is fine-tuned using domain-adapted word vectors, and the influence coefficient of attention weight is dynamically adjusted.

[0115] In this embodiment, the specific implementation of the hierarchical gated fusion network module is as follows:

[0116] The Emoji emoticons are converted into high-dimensional dense vectors through a pre-set embedding layer, and the temporal context representation of network slang is extracted using a pre-trained BERT model and normalized by L2 norm respectively.

[0117] Secondly, a two-level gating structure is set in the hierarchical gated interaction fusion network (GIFN):

[0118] Primary gating unit: concatenate semantic features and sentiment features, calculate different feature correlation scores through a learnable parameter matrix, and generate sentiment-enhanced semantic vectors.

[0119] Advanced gating unit: cross-channel splicing of primary gating output, Emoji vector, and popular language vector, using a bidirectional dynamic weight distribution module with Tanh activation, calculating feature weights by the following formula:

[0120] α i =Softmax(W g T ·[f sem ;f emoji ;f slang ])

[0121] Where W g is the trainable parameter matrix, α i Represents the contribution coefficient of each feature in the decision space;

[0122] Finally, a cross-entropy constraint is imposed on inter-channel features based on the sliding window mechanism, forcing different features to establish inter-class discrimination boundaries in the latent space, and finally outputting a fused feature vector with a dimension of 768 for knowledge distillation.

[0123] In this embodiment, the specific implementation of the knowledge distillation and reasoning module is as follows:

[0124] The teacher model is defined as a multi-layer attention Transformer architecture containing a hierarchical gated interactive fusion network (GIFN), and the student model is a lightweight network consisting of a bidirectional long short-term memory network (BiLSTM) and a fully connected layer;

[0125] Knowledge distillation is then achieved through the following multi-task supervision objectives:

[0126] Soft label distillation loss: Kullback-Leibler divergence is used to measure the soft label probability distribution P output by the teacher model with temperature parameter τ T And the student model output probability P S differences;

[0127] Hidden layer response matching loss: Use cosine similarity to constrain the teacher model GIFN's last hidden state h T With the student model BiLSTM output feature h S The distribution of aligns;

[0128] Attention weight alignment loss: By minimizing the teacher model self-attention matrix A T With the student model temporal attention matrix A S Frobenius norm distance, migration association pattern;

[0129] The teacher model is trained earlier than the student model, and the teacher model parameters are frozen during the joint training phase. The student model parameters are optimized only by the stochastic gradient descent algorithm, where the total loss function is:

[0130]

[0131] in is the cross entropy loss of the student model on the true label, λ1, λ2, λ3 are the weight coefficients of adaptive adjustment;

[0132] The temperature parameter τ is dynamically adjusted during the distillation process. In the initial stage, τ = 5 is set to smooth the probability distribution, and it linearly decays to τ = 1 as the number of training rounds increases to enhance the discrimination of difficult samples.

[0133] Real-time preprocessing and standardized input are performed on the social media data stream to be detected. The deep semantic features of the text, emotional semantic features, Emoji expression vector representation and Internet buzzword context vector are simultaneously extracted through parallel computing paths. Hierarchical feature weights are calibrated and dynamically spliced ​​through the BiLSTM student network. The cascade neural network is used to perform nonlinear mapping on the fused heterogeneous features. The softmax activation function is used to output multi-classification probability distribution, and the final rumor category is determined according to the preset confidence threshold.

[0134] Example 8

[0135] A multi-feature fusion rumor detection device based on knowledge distillation, wherein the device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. At the same time, when the processor executes the computer program, it implements the steps of the multi-feature fusion rumor detection method based on knowledge distillation described in any one of Examples 1-7.

Claims

1. A multi-feature fusion rumor detection method based on knowledge distillation, characterized in that: The steps include: (1) Obtain the data to be tested from social media, including original Weibo text, related comment data, Emoji emoticons, and Internet buzzwords; (2) The pre-trained BERT model is used to extract semantic features from the original text and comments of Weibo, generating context-aware deep semantic representations. At the same time, based on the BiLSTM-CNN hybrid network structure, the sentiment polarity of the text is modeled and the sentiment semantic feature vector is output. (3) The semantic features and sentiment features extracted in step 2 are input into the hierarchical gated interactive fusion network (GIFN) along with the emoji embedding vector and the internet buzzwords vocabulary vector. The cascaded gated units are used to dynamically adjust the weight ratios of different features, construct a nonlinear mapping relationship between feature dimensions, and generate a comprehensive representation that integrates multi-granularity features. (4) Construct a teacher model based on GIFN and a student model based on BiLSTM network, and jointly train them through knowledge distillation technology. The soft label probability distribution, hidden layer feature response and attention weight matrix output by the teacher model serve as supervision signals to guide the student model to learn feature association patterns; (5) Deploy the trained student model to the inference device to perform feature fusion, classification decision and output rumor detection results on the input social media data.

2. The rumor detection method according to claim 1, characterized in that: The step 1 comprises: (1) Obtain cross-platform social media data through generalized data collection tools; (2) Noise filtering of comment texts, including stripping out independent emoticons and preset internet buzzwords, and performing semantic completion and context-based expansion mapping on low-quality texts; (3) Based on the interaction time series information between user comments and original texts, we construct time window synchronized associated data pairs and annotate them with sentiment tendency labels to form a training set.

3. The multi-feature fusion rumor detection method based on knowledge distillation according to claim 1 is characterized in that: The step 2 includes: The pre-trained BERT model's multiple attention heads are used to concurrently extract contextual semantic features from the original Weibo text and associated comments. A multi-layer Transformer encoder is used to capture long-range dependencies in the text. Dynamic feature dimensionality reduction and hierarchical aggregation are then performed on the output layer embedding vector to generate a dimensionally aligned deep semantic vector. Simultaneously, fine-grained sentiment analysis of the text is performed using a BiLSTM-CNN hybrid network: The input text is mapped into a distributed representation through the word embedding layer, input into a bidirectional LSTM network to model forward and backward semantic associations, and outputs the hidden state vector of each position. A CNN sub-network with multi-scale convolution kernels is then used to extract the key sentiment features of the local text fragment and concatenate them with the BiLSTM hidden state vector across channels. A multi-head self-attention mechanism is then used to weight the importance of the concatenated features to generate a semantic sentiment vector that integrates long-term dependencies and local sentiment polarity. Furthermore, a symbol-text joint embedding method is used for Emoji emoticons, and the implicit semantics of emoticons are modeled with polysemy based on the pre-trained language model; for Internet buzzwords, domain-adapted word vector fine-tuning technology is used, combined with attention weights to dynamically adjust their influence coefficient in emotional characteristics.

4. The multi-feature fusion rumor detection method based on knowledge distillation according to claim 1 is characterized in that: The step 3 includes: The Emoji emoticons are converted into high-dimensional dense vectors through a preset embedding layer, and the temporal context representation of Internet buzzwords is extracted using a pre-trained BERT model, and L2 norm normalization is performed on each. Secondly, a two-level gating structure is set in the hierarchical gated interactive fusion network (GIFN): Primary gating unit: concatenates semantic features with sentiment features, calculates the correlation scores of different features through a learnable parameter matrix, and generates sentiment-enhanced semantic vectors; Advanced Gating Unit: Cross-channel concatenation of the primary gating output with the Emoji vector and buzzword vector. A bidirectional dynamic weight allocation module with Tanh activation is used to calculate the feature weight using the following formula: α i =Softmax(W g T ·[f sem ;f emoji ;f slang ]) Where W g is the trainable parameter matrix, α i Represents the contribution coefficient of each feature in the decision space; Finally, a cross-entropy constraint is imposed on inter-channel features based on the sliding window mechanism, forcing different features to establish inter-class discrimination boundaries in the latent space, and finally outputting a fused feature vector with a dimension of 768 for knowledge distillation.

5. The multi-feature fusion rumor detection method based on knowledge distillation according to claim 1 is characterized in that: The step 4 comprises: The teacher model is defined as a multi-layer attention Transformer architecture containing a hierarchical gated interactive fusion network (GIFN), and the student model is a lightweight network consisting of a bidirectional long short-term memory network (BiLSTM) and a fully connected layer; Knowledge distillation is then achieved through the following multi-task supervision objectives: Soft label distillation loss: Kullback-Leibler divergence is used to measure the soft label probability distribution P output by the teacher model with temperature parameter τ T And the student model output probability P S differences; Hidden layer response matching loss: Use cosine similarity to constrain the teacher model GIFN's last hidden state h T With the student model BiLSTM output feature h S The distribution of aligns; Attention weight alignment loss: By minimizing the teacher model self-attention matrix A T With the student model temporal attention matrix A S Frobenius norm distance, migration association pattern; The teacher model is trained earlier than the student model, and the teacher model parameters are frozen during the joint training phase. The student model parameters are optimized only by the stochastic gradient descent algorithm, where the total loss function is in is the cross entropy loss of the student model on the true label, λ1, λ2, λ3 are the weight coefficients of adaptive adjustment; The temperature parameter τ is dynamically adjusted during the distillation process. In the initial stage, τ = 5 is set to smooth the probability distribution, and it linearly decays to τ = 1 as the number of training rounds increases to enhance the discrimination of difficult samples.

6. The multi-feature fusion rumor detection method based on knowledge distillation according to claim 1 is characterized in that: The step 5 comprises: Real-time preprocessing and standardized input are performed on the social media data stream to be detected. The deep semantic features of the text, emotional semantic features, Emoji expression vector representation and Internet buzzword context vector are simultaneously extracted through parallel computing paths. Hierarchical feature weights are calibrated and dynamically spliced ​​through the BiLSTM student network. The cascade neural network is used to perform nonlinear mapping on the fused heterogeneous features. The softmax activation function is used to output multi-classification probability distribution, and the final rumor category is determined according to the preset confidence threshold.

7. A multi-feature fusion rumor detection system based on knowledge distillation, used to implement the multi-feature fusion rumor detection method based on knowledge distillation according to any one of claims 1 to 6, characterized in that: The system includes: a data acquisition and preprocessing module, a feature extraction module, a hierarchical gated fusion network module (GIFN), and a knowledge distillation and reasoning module; The data acquisition and preprocessing module is used to: collect original Weibo texts, comment sequences, Emoji emoticons and Internet buzzwords in social media in real time, and perform denoising and cleaning, normalized encoding and input format unification operations; the feature extraction module is used to: integrate pre-trained BERT to extract deep semantic features of text, BiLSTM-CNN hybrid network to capture emotional semantic features, and simultaneously map Emoji symbols and Internet buzzwords into low-dimensional vector representations through the word embedding layer; the hierarchical gated fusion network module (GIFN) is used to: dynamically weight feature weights through gating units, and use nonlinear activation functions and attention mechanisms to achieve cross-dimensional deep feature interaction; the knowledge distillation and reasoning module is used to: integrate the knowledge distillation framework of the GIFN teacher model and the BiLSTM student model, perform soft label probability supervision and parameter compression training, and deploy a lightweight student model to achieve real-time feature fusion classification.

8. A multi-feature fusion rumor detection device based on knowledge distillation, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the computer program, it implements the steps of the multi-feature fusion rumor detection method or system based on knowledge distillation as described in any one of claims 1 to 7, specifically covering data acquisition and preprocessing, multi-feature extraction and fusion, cross-module interaction based on hierarchical gating, and lightweight reasoning process driven by knowledge distillation.

Citation Information

Cited By

  • Model training method, risk behavior identification method, system, device and medium

    CN120910569A

  • Efficient interaction method and system for connecting large model and multi-source data

    CN120996012A

  • Lightweight multi-modal representation learning method based on multilayer attention mechanism

    CN121051701A

  • Engineering vehicle motor fault diagnosis method and system based on TAEB model

    CN121350504A

  • An engineering vehicle motor fault diagnosis method and system based on a TAEB model

    CN121350504B