Semantic guidance-based short video false news adaptive detection method

Through the adaptive detection method of short video fake news based on semantic guidance, multimodal features are extracted and analyzed, and the problems of difficulty in classification of true and false short videos and imbalance in classification in the existing technology are solved, and efficient detection and classification of fake news is achieved.

CN120107849APending Publication Date: 2025-06-06GUANGXI UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510145815.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-10
Publication Date
2025-06-06

AI Technical Summary

Technical Problem

The prior art is difficult to correctly classify the content of real and fake short videos, and rely too much on comments and user interactions after short videos are published, making it difficult to achieve efficient detection during the review stage and lack of classification balance.

Method used

A short video fake news adaptive detection method based on semantic guidance is proposed. Multimodal features are extracted through feature extraction and encoding modules, and the semantic verification module performs differential semantic analysis, and the distinctive semantic coding module improves feature differences, and the boundary analysis and calculation of classification boundaries are analyzed and calculated through boundary analysis and aggregation module to realize the true and false classification of short videos.

Benefits of technology

This method can effectively capture the differences in multimodal representations, balance the classification of fake news short videos, improve classification accuracy and balance, and solve the problem of classification imbalance and dependence on post-reviews.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120107849A_ABST
    Figure CN120107849A_ABST
Patent Text Reader

Abstract

According to the short video false news adaptive detection method based on semantic guidance, modeling is carried out on multiple modal characteristics of vision, text, audio and emotion in a video to obtain original representations of different modalities, and difference semantic analysis is carried out on the original representations of different modalities based on a text modal condition through a semantic verification module; a distinguishable semantic coding module uses difference semantics to enhance distinguishable features of original modal semantics, multi-modal feature differences of different types of videos are improved, a boundary analysis module analyzes classification boundaries of different modals according to the obtained multi-modal distinguishable features, learnable vectors are used as the classification boundaries of all the modals, and the classification boundaries of all the modals are analyzed. And according to the multi-modal classification boundary, the boundary aggregation module calculates multi-level feature aggregation for the multi-modal boundary to obtain true and false classification labels of the video. The invention not only analyzes the multi-mode representation difference of true and false short videos, but also provides a classification method of false news short videos with balanced categories.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of fake news short video authentication, and specifically relates to a semantically guided short video fake news adaptive detection method. Background Art

[0002] Short video fake news detection is a task to detect fake news in short videos. The main challenge of the task is how to correctly distinguish between real videos and fake news videos when the real and fake contents are similar. Early work was mainly based on manual feature methods. These methods extracted speech, emotion and video title vectors for support vector machine model learning decision boundaries. Since the manually extracted features are relatively fixed, these methods can only handle short video fake news detection in limited fields. Subsequent work combined more information on the basis of deep neural networks, such as the relationship between video comments and videos, but still did not solve the problem of classification imbalance caused by the high degree of content similarity between real and fake short videos. Later work used large language models to mine implicit opinions in short videos and used diffusion methods to simulate the spread of opinions. However, it was too dependent on comments and user interactions after the short videos were released, making it difficult for the platform to achieve efficient detection in the review stage. At the same time, the models of such work also need to be improved in terms of classification balance. Therefore, it is necessary to study a short video fake news detection method that can maintain classification balance. Summary of the invention

[0003] In order to solve the problems existing in the above-mentioned prior art, the present invention proposes a semantic-guided adaptive detection method for short video fake news, with the aim of analyzing the multimodal representation differences between true and false short videos and providing a classification method for fake news short videos with balanced categories, so as to solve the problems of unbalanced classification due to the high degree of similarity in content between true and false short videos and over-reliance on comments and user interactions after the release of short videos, making it difficult for the platform to achieve efficient detection in the review stage.

[0004] In order to achieve the above object, the specific scheme of the present invention is as follows:

[0005] A semantically guided short video fake news adaptive detection method comprises the following steps:

[0006] S1, extracts multimodal feature modeling of visual features, text features, audio features and emotional features in the video through feature extraction and encoding modules to obtain original representations of different modalities;

[0007] S2, performs differential semantic analysis on the original representations of different modalities based on the text modality through the semantic verification module;

[0008] S3, improves the multimodal feature differences of different categories of videos through a distinguishable semantic coding module, and uses differential semantic enhancement coding to obtain multimodal distinguishable features;

[0009] S4, analyzes the classification boundaries of different modalities based on the obtained multi-modal distinguishable features through the boundary analysis module, uses the learnable vector as the classification boundary of each modality, inputs it into the multi-head attention mechanism, and obtains the boundary features between the real video and the fake news video in each modality;

[0010] S5, the boundary aggregation module calculates multi-level feature aggregation for the classification boundaries of each modality to obtain true or false classification labels for the video.

[0011] Furthermore, the steps of the difference semantic analysis in step S2 are as follows:

[0012] S21, using the attention mechanism to calculate the conditional features between text features and other modal features, and normalize the text features and conditional features;

[0013] S22, performing a dot product operation on the normalized text features and conditional features;

[0014] S23, scale and transform the dot product result using the temperature scaling factor to obtain semantic verification scores of different modalities.

[0015] Furthermore, the calculation formula of the conditional feature in step S21 is as follows:

[0016]

[0017] In the formula, f v Represents other modal features; f pos represents the position vector; f t Represents text features; d k represents the scaling factor; f c represents conditional features; T represents the transposed function;

[0018] The formula for obtaining the semantic verification scores of different modalities in step S23 is as follows:

[0019]

[0020] Where F represents a linear function; α and σ represent scaling factors; and S represents the semantic verification scores of different modalities.

[0021] Furthermore, the steps of improving the multimodal feature differences of videos of different categories in step S3 are as follows:

[0022] S31, taking text features and other modal features as query vectors and key vectors, and taking other modal features as value vectors, inputs them into the attention mechanism to calculate other modal context features based on the text;

[0023] S32, adding the text-based other modal context features and other modal features calculated in step S31 as a new query vector and a key vector, and inputting the other modal features as a value vector into the multi-head attention mechanism, further calculating the discriminative other modal features, and adding the discriminative other modal features to the original other modal features to obtain a multimodal feature representation with implied context;

[0024] S33, summing the text features and other text-based modal context features calculated in step S31, and multiplying the sum result by the verification score to generate a final distinguishable feature representation.

[0025] Furthermore, the calculation formula of the other modal context features based on text in step S31 is as follows:

[0026]

[0027] In the formula, f v Represents other modal features; f t Represents text features; d k represents the scaling factor; t l represents other modal context features; T represents the transposition function;

[0028] The calculation formula of the multimodal feature representation with implicit context in step S32 is as follows:

[0029]

[0030] In the formula, f cont Represents other modal features that are discriminative; represents the multimodal feature representation containing context; S represents the semantic verification score of different modalities.

[0031] Furthermore, the calculation formula for obtaining the boundary features between the real video and the fake news video in each modality described in step S4 is as follows:

[0032]

[0033] In the formula, q v represents the actual calculation; represents the zero vector; represents the key vector and value vector; b v Representing the boundary features between real videos and fake news videos in each modality;

[0034] Furthermore, the formula for obtaining the true or false classification label for the video in step S5 is as follows:

[0035] b′ l =ReLU(LN(Conv([b v1 ,b v2 , b v3 ,b v4 ]))),

[0036] b″ l =Conv((b′ l ) T ),

[0037] f f =FFN((b″ l ) T ),

[0038] In the formula, ReLU represents the activation function; LN represents the regularization layer; Conv represents the convolutional layer; FFN represents the feedforward neural network; b v1 、b v2 、b v3 、b v4 Respectively represent the boundary features of text, video, audio, and emotion; f f Represents the classification label of the final output.

[0039] Advantages of the present invention

[0040] The present invention is a short video fake news adaptive detection method based on semantic guidance. It is a neural network algorithm based on the Transformer model as the skeleton. It combines the differences in the representation of semantic features of different modalities, enhances the algorithm to capture the differences in multimodal representations, and uses the adaptive adjustment advantages of learnable vectors to fit the classification boundaries between complex multimodal features, so as to classify fake news short videos in a balanced manner. In view of the problems of modal information heterogeneity and classification imbalance, the present invention uses semantic guidance to emphasize the key differences between modalities, enhances the model's ability to understand the content of short videos, and designs a distinguishable semantic mining module and a boundary learning module, which can conditionally learn and extract modal features consistent with the emotions of the story context, thereby effectively detecting fake news in short videos. BRIEF DESCRIPTION OF THE DRAWINGS

[0041] Figure 1 This is a flow chart of the semantically guided short video fake news adaptive detection method of the present invention. DETAILED DESCRIPTION

[0042] The present invention is further explained and illustrated below in conjunction with the accompanying drawings and specific embodiments. It should be noted that this specific embodiment is not intended to limit the scope of rights of the present invention.

[0043] like Figure 1 As shown, this specific embodiment provides a semantic-guided short video fake news adaptive detection method, which includes the following steps:

[0044] S1, extracts multimodal feature modeling of visual features, text features, audio features and emotional features in the video through feature extraction and encoding modules to obtain original representations of different modalities;

[0045] The visual feature refers to sampling the video to obtain frames, and then encoding them using the CLIP model, which is denoted as V = [v 1 ,v 2 ,v 3 ,...,v i ].

[0046] The text feature refers to the subtitles and titles in the video, which are concatenated and encoded using the BERT model, denoted as T = [t 1 ,t 2 ,t 3 ,...,t i ].

[0047] The audio feature refers to the audio content in the video, which is encoded using the Wav2Vec model and is denoted as A=[a 1 ,a 2 ,a 3 ,...,a i ].

[0048] The emotional feature refers to the use of a pre-trained audio emotion model to classify the audio features to obtain the corresponding emotional coding representation, denoted as E = [e 1 ,e 2 ,e 3 ,...,e i ].

[0049] S2, through the semantic verification module, the original representations of different modalities are subjected to differential semantic analysis based on the text modality. The specific steps are as follows:

[0050] S21, adding the position vector and other modal features as the query vector, using the text features as the key vector and the value vector, and using the attention mechanism to calculate the conditional features between the text features and other modal features, and normalizing the text features and the conditional features;

[0051] The calculation formula of the conditional characteristics is as follows:

[0052]

[0053] In the formula, f v Represents other modal features; f pos represents the position vector; f t Represents text features; d k represents the scaling factor; f c represents conditional features; T represents the transposed function;

[0054] S22, performing a dot product operation on the normalized text features and conditional features;

[0055] S23, scale and transform the dot product result using the temperature scaling factor to obtain semantic verification scores of different modalities.

[0056] The formula for obtaining the semantic verification scores of different modalities is as follows:

[0057]

[0058] Where F represents a linear function; α and σ represent scaling factors; and S represents the semantic verification scores of different modalities.

[0059] S3, improves the multimodal feature differences of different categories of videos through a distinguishable semantic coding module, and uses differential semantic enhancement coding to obtain multimodal distinguishable features;

[0060] The steps of improving the multimodal feature differences of different categories of videos in step S3 are as follows:

[0061] S31, taking the text features and other modal features as the query vector and key vector, and taking the other modal features as the value vector, inputting them into the attention mechanism, and calculating the other modal context features based on the text; the calculation formula of the other modal context features based on the text is as follows:

[0062]

[0063] In the formula, f v Represents other modal features; f t Represents text features; d k represents the scaling factor; t l represents other modal context features; T represents the transposition function;

[0064] S32, adding the text-based other modal context features and other modal features calculated in step S31 as a new query vector and a key vector, and inputting the other modal features as a value vector into the multi-head attention mechanism, further calculating the discriminative other modal features, and adding the discriminative other modal features to the original other modal features to obtain a multimodal feature representation with implied context;

[0065] The calculation formula of the multimodal feature representation with implicit context in step S32 is as follows:

[0066]

[0067] In the formula, f cont Represents other modal features that are discriminative; represents the multimodal feature representation containing context; S represents the semantic verification score of different modalities.

[0068] S33, summing the text features and other text-based modal context features calculated in step S31, and multiplying the sum result by the verification score to generate a final distinguishable feature representation.

[0069] S4, through the boundary analysis module, the classification boundaries of different modes are analyzed according to the obtained multi-modal distinguishable features, and the learnable vectors are used as the classification boundaries of each mode. The initial values ​​are randomly sampled from the uniform distribution to initialize the two query vectors, q v For actual calculations, the other is a zero vector Q v and Add them together to get a new query vector, and then add the query vector to the multimodal discrimination feature As key vectors and value vectors, they are input into the multi-head attention mechanism to obtain the boundary features between real videos and fake news videos in each modality. The calculation formula for obtaining the boundary features between real videos and fake news videos in each modality is as follows:

[0070]

[0071] In the formula, q v represents the actual calculation; represents the zero vector; represents the key vector and value vector; b v represents the boundary features between real videos and fake news videos in each modality; f v· Indicates other modal features.

[0072] S5, the boundary aggregation module calculates multi-level feature aggregation for the classification boundaries of each modality to obtain true or false classification labels for the video.

[0073] Specifically, the query vectors and boundary features of the four modalities are first cascaded to generate a feature vector bl containing the information of the four modalities, and then the feature vectors are aligned and fused at multiple levels. Finally, the dimensionally aligned feature vector is cascaded with the learnable query vector and input into the feedforward neural network for calculation, thereby generating the final classification feature:

[0074] Channel-level alignment: Use a one-dimensional convolutional layer to perform a convolution operation on bl, thereby aligning the boundary features of the four modalities at the channel level. Dimension-level alignment: Use the transpose operation to swap the first and second dimensions of bl, and then use a one-dimensional convolutional layer to perform a convolution operation again:

[0075] b′ l =ReLU(LN(Conv([b v1 ,b v2 ,b v3 ,b v4 ]))),

[0076] b″ l =Conv((b′ l ) T ),

[0077] f f =FFN((b″ l ) T ),

[0078] In the formula, ReLU represents the activation function; LN represents the regularization layer; Conv represents the convolutional layer; FFN represents the feedforward neural network; b v1 、b v2 、b v3 、b v4 Respectively represent the boundary features of text, video, audio, and emotion; f f Represents the classification label of the final output.

[0079] The following is an evaluation of the story ending generation method based on the pre-trained model in this embodiment on the FakeSV and FakeTT corpora. The FakeSV public dataset contains 3624 samples. The dataset is divided into a training set, a validation set, and a test set, with 2558 samples, 542 samples, and 542 samples, respectively. The FakeTT public dataset contains 1911 samples, and the dataset is divided into a training set, a validation set, and a test set, with 1313 samples, 299 samples, and 299 samples, respectively.

[0080] In the evaluation phase, the method of this embodiment and the method of the prior art were compared and evaluated in the following three parts:

[0081] (1) Accuracy (Acc) is the most intuitive and commonly used classification performance evaluation indicator. It represents the ratio of the number of samples correctly classified by the classifier to the total number of samples. The larger the value, the better.

[0082] (2) F macro average F1 score, which is the harmonic mean of precision and recall. The larger the value, the better.

[0083] (3) Matthews Correlation Coefficient (MCC). The Matthews Correlation Coefficient is a comprehensive evaluation indicator that takes into account true positives, true negatives, false positives, and false negatives. Its value ranges from -1 to 1. An MCC of 1 indicates a perfect prediction, an MCC of 0 indicates a random prediction, and an MCC of -1 indicates a completely opposite prediction. The larger the value, the better.

[0084] During the training of the FakeSV and FakeTT corpus, a learning rate of 1e-5 and the Adam optimizer were used for 30 epochs of training. The number of heads of the attention mechanism was set to 5, and the harmonic coefficients were set to 1 and 0.5. Compared with the prior art methods, this embodiment has been effectively improved in terms of classification accuracy and balance.

[0085] Finally, the comparison of the method of this embodiment and other prior art methods in terms of evaluation indicators is shown in Table 1:

[0086]

[0087]

[0088] From the results of the above specific embodiments, it can be seen that compared with the prior art methods, the method of this specific embodiment can alleviate the classification imbalance problem that is prone to occur, can maintain similar classification scores for positive and negative samples, and can improve classification accuracy and balance.

[0089] The above implementation cases are only preferred embodiments of the present invention and are not intended to limit the present invention. For those skilled in the art, various changes and modifications can be made based on the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention belong to the protection scope of the present invention.

Claims

1. A semantically guided short video fake news adaptive detection method, characterized in that: The steps include: S1, extracts multimodal feature modeling of visual features, text features, audio features and emotional features in the video through feature extraction and encoding modules to obtain original representations of different modalities; S2, performs differential semantic analysis on the original representations of different modalities based on the text modality through the semantic verification module; S3, improves the multimodal feature differences of different categories of videos through a distinguishable semantic coding module, and uses differential semantic enhancement coding to obtain multimodal distinguishable features; S4, analyzes the classification boundaries of different modalities based on the obtained multi-modal distinguishable features through the boundary analysis module, uses the learnable vector as the classification boundary of each modality, inputs it into the multi-head attention mechanism, and obtains the boundary features between the real video and the fake news video in each modality; S5, the boundary aggregation module calculates multi-level feature aggregation for the classification boundaries of each modality to obtain true or false classification labels for the video.

2. According to claim 1, a semantically guided short video fake news adaptive detection method is characterized in that: The steps of differential semantic analysis in step S2 are as follows: S21, using the attention mechanism to calculate the conditional features between text features and other modal features, and normalize the text features and conditional features; S22, performing a dot product operation on the normalized text features and conditional features; S23, scale and transform the dot product result using the temperature scaling factor to obtain semantic verification scores of different modalities.

3. The method according to claim 2, characterized in that The calculation formula of the conditional feature in step S21 is as follows: In the formula, f v Indicates other modal features; f pos represents the position vector; f t Represents text features; d k represents the scaling factor; f c represents conditional features; T represents the transposed function; The formula for obtaining the semantic verification scores of different modalities in step S23 is as follows: Where F represents a linear function; α and σ represent scaling factors; and S represents the semantic verification scores of different modalities.

4. The method according to claim 1, characterized in that: The steps of improving the multimodal feature differences of different categories of videos in step S3 are as follows: S31, taking text features and other modal features as query vectors and key vectors, and taking other modal features as value vectors, inputs them into the attention mechanism to calculate other modal context features based on the text; S32, adding the text-based other modal context features and other modal features calculated in step S31 as a new query vector and a key vector, and inputting the other modal features as a value vector into the multi-head attention mechanism, further calculating the discriminative other modal features, and adding the discriminative other modal features to the original other modal features to obtain a multimodal feature representation with implied context; S33, summing the text features and other text-based modal context features calculated in step S31, and multiplying the sum result by the verification score to generate a final distinguishable feature representation.

5. The method according to claim 4, characterized in that The calculation formula of the other modal context features based on text in step S31 is as follows: In the formula, f v Indicates other modal features; f t Represents text features; d k represents the scaling factor; t l represents other modal context features; T represents the transposition function; The calculation formula of the multimodal feature representation with implicit context in step S32 is as follows: In the formula, f cont Indicates other modal features that are discriminative; represents the multimodal feature representation containing context; S represents the semantic verification score of different modalities.

6. The method according to claim 1, characterized in that The calculation formula for obtaining the boundary features between the real video and the fake news video in each modality described in step S4 is as follows: In the formula, q v represents the actual calculation; represents the zero vector; represents the key vector and value vector; b v Representing the boundary features between real videos and fake news videos in each modality; f v· Indicates other modal features.

7. The method according to claim 1, characterized in that The formula for obtaining the true or false classification label for the video in step S5 is as follows: b′ l =ReLU(LN(Conv([b v1 ,b v2 ,b v3 ,b ,4 ]))), b″ l =Conv((b′ l ) T ), f f =FFN((b″ l ) T ), In the formula, ReLU represents the activation function; LN represents the regularization layer; Conv represents the convolutional layer; FFN represents the feedforward neural network; b v1 , b v2 , b v3 , b v4 Respectively represent the boundary features of text, video, audio, and emotion; f f Represents the classification label of the final output.