A false news video detection method based on information enhancement and guided denoising

By extracting label features using a large language model and performing semantic enhancement and multimodal denoising, the problem of difficulty in extracting key information and noise interference in fake news video detection is solved, achieving more efficient fake news detection.

CN121280973BActive Publication Date: 2026-05-22INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
INST OF AUTOMATION CHINESE ACAD OF SCI
Filing Date
2025-12-04
Publication Date
2026-05-22

AI Technical Summary

Technical Problem

Existing technologies are insufficient to effectively detect fake news videos, especially in multimodal videos where key information is difficult to extract and noise interference is severe, resulting in low detection accuracy.

Method used

A fake news video detection process is constructed by using a method based on information enhancement and guided denoising. This process utilizes a large language model to extract label features, and combines semantic enhancement and multimodal feature processing with a label-guided denoising mechanism. The process includes label feature extraction, semantic enhancement, multimodal feature fusion, and denoising.

Benefits of technology

It improves the accuracy and efficiency of fake news video detection, achieves more reliable verification of news authenticity, and can more accurately identify fake news.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121280973B_ABST
    Figure CN121280973B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of video detection, and discloses a false news video detection method based on information enhancement and guided denoising, which comprises the following steps: obtaining a news video to be detected, and extracting label features of the news video to be detected by using a large language model; based on the label features, multi-modal features of the news video to be detected are extracted, and the label features are subjected to semantic enhancement to obtain label enhanced features; label guided denoising is carried out based on the label enhanced features and the multi-modal features, multi-modal denoised features are obtained, and whether the news video to be detected is a false news video is determined according to the multi-modal denoised features. The application extracts key information of a multi-modal video by means of the semantic understanding ability of the large language model, extracts multi-modal features, subjects the label features to semantic enhancement, combines a label guided denoising mechanism, constructs a complete and effective false news video detection process, improves the detection precision and efficiency of the false news video, and realizes more reliable news authenticity verification.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of video detection technology, specifically to a method for detecting fake news videos based on information enhancement and guided denoising. Background Technology

[0002] In recent years, short video platforms have rapidly gained popularity, gradually becoming a major channel for news dissemination and accelerating the spread of video news (including fake video news). Studies show that 39% of adults under the age of 30 obtain news information through short video platforms, and nearly half of users consider it their primary source of news information. With the increasing prevalence of short video platforms, fake video news, through emotional manipulation and one-sided information presentation, induces audiences to believe and spread false content, and its negative impact on social stability is becoming increasingly prominent, urgently requiring effective detection technologies to address the issue.

[0003] Unlike traditional text and image news, short video news integrates multimodal information such as visual, audio, and text, significantly increasing the complexity of fake news detection. Some researchers have proposed enhancing news feature expression through graph aggregation and using image-text consistency for fact-checking; others have analyzed the fake news creation process from the perspective of material selection and editing features. However, existing methods still face two major challenges: First, the ambiguity of the content makes it difficult to extract key information. The unstructured nature of short videos makes core information (such as visual keyframes and text descriptions) subject to interference and subjective expression, making it difficult for traditional methods to generate structured semantic representations. Second, multimodal redundant noise interferes with the effective location of information. Irrelevant scenes and redundant discussions obscure core facts, and existing cross-modal alignment methods lack noise filtering mechanisms, making it difficult for models to focus on key content. Summary of the Invention

[0004] In view of this, the present invention provides a method for detecting fake news videos based on information enhancement and guided denoising, in order to solve the problems of difficulty in extracting key information and low detection accuracy in multimodal videos.

[0005] In a first aspect, the present invention provides a method for detecting fake news videos based on information enhancement and guided denoising, the method comprising:

[0006] Acquire news videos to be detected and extract label features of the news videos to be detected using a large language model;

[0007] Based on the label features, multimodal features of the news video to be detected are extracted, and the label features are semantically enhanced to obtain the label-enhanced features;

[0008] Label-guided denoising is performed based on label enhancement features and multimodal features to obtain multimodal denoising features, and the detection of whether the news video to be detected is a fake news video is determined based on the multimodal denoising features.

[0009] The present invention provides a fake news video detection method based on information enhancement and guided denoising. It uses the semantic understanding capability of a large language model to extract key information from multimodal videos, extracts multimodal features and enhances the semantics of the tag features to enrich the feature dimensions and strengthen the semantic expression of the tags. Combined with a tag-guided denoising mechanism, it constructs a complete and effective fake news video detection process, improves the detection accuracy and efficiency of fake news videos, and achieves more reliable verification of the authenticity of news.

[0010] In one optional implementation, the tag features include: content authenticity tags and visual insight tags. Tag features are extracted from the news video to be detected using a large language model, including:

[0011] Extract text and keyframes from the news video to be detected;

[0012] Based on the text, a large language model is used to extract content authenticity tags, and based on the video keyframes, a large language model is used to extract visual insight tags.

[0013] The fake news video detection method based on information enhancement and guided denoising provided by this invention uses a large language model to extract content authenticity tags and visual insight tags, and structurally represents news content from two levels: global scene and local details, thereby enhancing the reliability and completeness of key information and effectively mitigating the interference caused by content ambiguity.

[0014] In one optional implementation, the multimodal features include: text features, visual features, and audio features. Extracting the multimodal features of the news video to be detected includes:

[0015] Identify textual, visual, and audio information in news videos to be detected;

[0016] Natural language processing models are used to extract text features from text information, convolutional neural network models are used to extract visual features from visual information, and Mel spectrogram and audio neural network models are used to extract audio features from audio information.

[0017] The method for detecting fake news videos based on information enhancement and guided denoising provided by this invention fully explores the value of different modal information by identifying multimodal information, and extracts multimodal features from multimodal information using different models, thereby achieving multi-dimensional and comprehensive feature extraction of news videos. By relying on the adaptive model, it accurately captures the unique features of each modality, making the detection based on richer and more accurate features, and improving the accuracy and reliability of detection.

[0018] In one optional implementation, semantic enhancement is performed on the label features to obtain label-enhanced features, including:

[0019] Obtain multiple related videos of the news video to be detected, and extract the tag features of each related video;

[0020] By utilizing the tag features of the news video to be detected and the tag features of each related video, a tag semantic similarity matrix is ​​constructed to quantify the semantic association strength between tag pairs.

[0021] A dynamic nearest neighbor sampling strategy is used to select the most similar positive samples of the same category and the most similar negative samples of different categories for each label.

[0022] A three-layer fully connected encoder is used to map the labels to a low-dimensional semantic space, and a loss function is used for optimization training to obtain label-enhanced features.

[0023] The present invention provides a fake news video detection method based on information enhancement and guided denoising. By enhancing the semantics of tag features and enriching semantic references with relevant videos, and by optimizing precise sampling and encoding, the method strengthens the accuracy and discriminativeness of tag semantic expression, enabling tag features to more accurately reflect news semantics and improving the accuracy of feature matching and judgment in subsequent fake news detection.

[0024] In one optional implementation, label-guided denoising is performed based on label enhancement features and multimodal features to obtain multimodal denoising features, including:

[0025] The weights of each modal feature are set using binary masks based on prior knowledge;

[0026] Multimodal features are used as query vectors and label features are used as key-value pairs. Multi-head attention is then performed to obtain multimodal denoising features.

[0027] In one optional implementation, multimodal features are used as query vectors and label features are used as key-value pairs. Multi-head attention computation is performed to obtain multimodal denoising features, including:

[0028] Multimodal features are obtained by processing label features;

[0029] The multimodal label features are subjected to mean pooling, and the weights of each modal label feature are calculated using a two-layer gating network.

[0030] Multimodal denoising features are calculated based on the modal label features and their corresponding weights.

[0031] The fake news video detection method based on information enhancement and guided denoising provided by this invention accurately filters out irrelevant feature interference by leveraging prior knowledge, and uses an attention mechanism to deeply associate multimodal features with labels. Then, through mean pooling and gating networks to refine modal weights, the denoised multimodal features are made to better fit the semantics of the labels, providing cleaner and more accurate feature support for fake news detection and improving the reliability and accuracy of detection results.

[0032] In one optional implementation, determining whether a news video to be detected is a fake news video based on multimodal denoising features includes:

[0033] The multimodal denoising features are input into the encoder to obtain the multimodal fusion features;

[0034] Based on multimodal fusion features, a classifier is used to determine whether the news video to be detected is a fake news video or a real news video.

[0035] The present invention provides a method for detecting fake news videos based on information enhancement and guided denoising. It utilizes an encoder to integrate multimodal information, eliminate modal differences and enhance complementarity, so that the fused features accurately reflect the essence of the news. The classifier makes judgments based on the fused features, and with the help of the true and false patterns learned by the model, it can efficiently and accurately distinguish between fake news videos and real news videos.

[0036] Secondly, the present invention provides a fake news video detection device based on information enhancement and guided denoising, the device comprising:

[0037] The tag extraction module is used to acquire news videos to be detected and to extract tag features of the news videos to be detected using a large language model.

[0038] The information enhancement module is used to extract multimodal features of the news video to be detected based on the label features, and to perform semantic enhancement on the label features to obtain label-enhanced features;

[0039] The guided denoising module is used to perform label-guided denoising based on label enhancement features and multimodal features to obtain multimodal denoising features, and to determine whether the news video to be detected is a fake news video based on the multimodal denoising features.

[0040] Thirdly, the present invention provides a computer device, comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the computer instructions to perform the method described in the first aspect or any corresponding embodiment thereof.

[0041] Fourthly, the present invention provides a computer-readable storage medium storing computer instructions for causing a computer to perform the method described in the first aspect or any corresponding embodiment thereof. Attached Figure Description

[0042] To more clearly illustrate the specific embodiments of the present invention or the technical solutions in the prior art, the drawings used in the description of the specific embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0043] Figure 1 This is a flowchart illustrating a method for detecting fake news videos based on information enhancement and guided denoising according to an embodiment of the present invention.

[0044] Figure 2 This is a flowchart illustrating another method for detecting fake news videos based on information enhancement and guided denoising according to an embodiment of the present invention;

[0045] Figure 3 This is a schematic diagram of a key video frame in a fake news video detection method based on information enhancement and guided denoising according to an embodiment of the present invention;

[0046] Figure 4 This is a schematic diagram illustrating the principle of tag contrast enhancement in the fake news video detection method based on information enhancement and guided denoising according to an embodiment of the present invention;

[0047] Figure 5 This is a schematic diagram of the complete data processing process in the fake news video detection method based on information enhancement and guided denoising according to an embodiment of the present invention;

[0048] Figure 6 This is a structural block diagram of a fake news video detection device based on information enhancement and guided denoising according to an embodiment of the present invention;

[0049] Figure 7 This is a schematic diagram of the hardware structure of a computer device according to an embodiment of the present invention. Detailed Implementation

[0050] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0051] This invention provides a method for detecting fake news videos based on information enhancement and guided denoising. By constructing a complete and effective fake news video detection process, it aims to effectively extract key information from multimodal videos and improve detection accuracy.

[0052] According to an embodiment of the present invention, a method for detecting fake news videos based on information enhancement and guided denoising is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.

[0053] This embodiment provides a method for detecting fake news videos based on information augmentation and guided denoising, which can be used in the aforementioned computer system. Figure 1 This is a flowchart of a fake news video detection method based on information enhancement and guided denoising according to an embodiment of the present invention, such as... Figure 1 As shown, the process includes the following steps:

[0054] Step S101: Obtain the news video to be detected and extract the label features of the news video to be detected using a large language model.

[0055] Specifically, the news videos to be detected consist of multiple consecutive video frames. A large-scale language model is used to extract multi-level labels from the visual keyframes and text descriptions of the video news, including visual insight labels and content authenticity labels. Visual insight labels are used to capture the global scene semantics of the video, while content authenticity labels focus on the credibility assessment of key local details.

[0056] Step S102: Based on the label features, extract the multimodal features of the news video to be detected, and perform semantic enhancement on the label features to obtain the label-enhanced features.

[0057] Specifically, the news videos to be detected contain multiple modalities, including visual, auditory, and textual information. Feature encoding is performed on the visual, audio, and textual modalities, and semantic consistency is enhanced by constraining the label features. By constructing positive and negative sample pairs, it is ensured that similar labels from different modalities remain tightly clustered in the feature space.

[0058] Step S103: Perform label-guided denoising based on label enhancement features and multimodal features to obtain multimodal denoising features, and determine whether the news video to be detected is a fake news video based on the multimodal denoising features.

[0059] Specifically, single-label guided denoising is performed on multimodal features. A gated multi-head attention mechanism is designed based on visual insight labels and content authenticity labels. Visual insight labels guide global topic feature filtering, while content authenticity labels drive the purification of local detail features, achieving layered denoising.

[0060] Dynamic weight fusion is performed on the global and local denoised multimodal features. This fusion strategy automatically adjusts the feature weights based on the credibility between modalities, achieving adaptive feature fusion. The fused and denoised multimodal features are then input into a classifier to detect fake news videos. Finally, a multilayer perceptron classifier is used to output the detection probability of fake news.

[0061] The fake news video detection method based on information enhancement and guided denoising provided in this embodiment extracts key information from multimodal videos by leveraging the semantic understanding capabilities of a large language model. By extracting multimodal features and semantically enhancing the label features, the feature dimensions are enriched and the semantic expression of the labels is strengthened. Combined with a label-guided denoising mechanism, a complete and effective fake news video detection process is constructed, improving the detection accuracy and efficiency of fake news videos and achieving more reliable verification of news authenticity.

[0062] This embodiment provides a method for detecting fake news videos based on information augmentation and guided denoising, which can be used in the aforementioned computer system. Figure 2 This is a flowchart of a fake news video detection method based on information enhancement and guided denoising according to an embodiment of the present invention, such as... Figure 2 As shown, the process includes the following steps:

[0063] Step S201: Obtain the news video to be detected and extract the label features of the news video to be detected using a large language model.

[0064] Specifically, the tag features include: content authenticity tags and visual insight tags, and step S201 above includes:

[0065] Step S2011: Extract the text and video keyframes of the news video to be detected.

[0066] Specifically, the text in the news video to be detected refers to the text information appearing in the video, and the keyframes are keyframes selected from the news video to be detected. To extract explicit structured information from complex news videos, a multi-level information enhancement module was designed. Visual insight tags (VI-Tag) capture overall background information, while content authenticity tags (CV-Tag) identify detailed features, achieving an organic integration of global and local information. To overcome the limitations of single-perspective content, VI-Tag and CV-Tag are designed as a complementary and collaborative mechanism to jointly enhance the information representation capabilities of fake news detection.

[0067] Visual insight tags need to be extracted from visual information, while content authenticity tags need to be extracted from text. Preprocessing is required for the news video to be detected, which involves extracting the text and keyframes of the video. This process is a mature existing technology and will not be elaborated here.

[0068] Step S2012: Extract content authenticity tags from the text using a large language model, and extract visual insight tags from the video keyframes using a large language model.

[0069] Specifically, VI-Tags are used to identify the themes and major events of news videos at a global level. Cue words from a large language model are leveraged to parse keyframes of the news video, thereby summarizing the overall scene. For example, ... Figure 3 As shown, the input video keyframes are accompanied by the prompt: "You are good at understanding video images and captions. These {images} are keyframes of the video. Please summarize the key information of this video news in one sentence and use it as a tag." The output of the visual insight tag would be: "This video describes the debate surrounding the authenticity of images of strange fish." This is just an example and is not a limitation. While CV-Tags can provide detailed information, they often lack overall context in complex scenes, easily leading to one-sided or misleading interpretations. VI-Tags, on the other hand, provide a more stable analytical foundation, helping to comprehensively understand the main themes and events of the video.

[0070] At the local level, CV-Tags extract key details and assess the credibility of specific claims, utilizing large-scale language models to analyze news text and uncover detailed factual clues. For example, given the input text: "Headline: A meeting was recently held; User information: User verified," and the prompt: "You excel at detailed news understanding and fake news detection. Please assess and briefly analyze the credibility of {content} and provide the output in the following format: The news mentions..., consistent with reality or scientific knowledge (inconsistent / consistent)," the output of the content authenticity tag would be "The news mentions a meeting or interaction, which is consistent with reality." This is just an example and not a limitation. Compared to complex visual data, text content is more suitable for extracting local tags because the detailed information it provides is easier to directly verify and fact-check. Unlike VI-Tags, which provide the overall theme of the news, CV-Tags emphasize key details and assist in fake news detection by assessing their authenticity.

[0071] VI-Tag and CV-Tag complement each other through their respective informational advantages: VI-Tag extracts the overall visual context (e.g., scene type, subject behavior) from keyframes of the video, providing a macro framework for content understanding; while CV-Tag focuses on the fine-grained semantics of news text (e.g., data description, names of people), verifying micro-facts through language models. Their complementarity is reflected in the following: when the global scene identified by VI-Tag has a semantic relationship with the text details detected by CV-Tag, the system enhances the credibility of that information; conversely, if CV-Tag finds a contradiction between the text and the scene determined by VI-Tag, it triggers a detection warning. This complementary mechanism does not rely on complex feature fusion, but rather significantly improves the ability to identify tampered videos (real footage with false subtitles) or misleading editing through cross-validation of the spatial context of the video modality and the logical details of the text modality.

[0072] The fake news video detection method based on information enhancement and guided denoising provided in this embodiment uses a large language model to extract content authenticity tags and visual insight tags, and structurally represents news content from two levels: global scene and local details, thereby enhancing the reliability and completeness of key information and effectively mitigating the interference caused by content ambiguity.

[0073] Step S202: Based on the label features, extract the multimodal features of the news video to be detected, and perform semantic enhancement on the label features to obtain the label-enhanced features.

[0074] Specifically, the multimodal features include: text features, visual features, and audio features. In step S202 above, the multimodal features of the news video to be detected are extracted, including:

[0075] Step S2021: Identify the text information, visual information, and audio information in the news video to be detected.

[0076] Specifically, the news video to be detected consists of multiple consecutive frames, each containing textual, visual, and audio information. For example, given a news video containing... A dataset of news items (one frame of video). Each news item N consists of three modalities: visual information, auditory information, and textual information, and can be represented as follows: Where N represents a single news item, V represents visual information, A represents audio information, and T represents text information.

[0077] Step S2022: Extract text features from text information using a natural language processing model, extract visual features from visual information using a convolutional neural network model, and extract audio features from audio information using a Mel spectrogram and an audio neural network model.

[0078] Specifically, different feature extraction methods are used for different modalities. For the text modality, specific optimizations are made for different language features: For the Chinese dataset FakeSV, a Chinese whole-word masking pre-trained model (Robustly Optimized BERT Approach-Whole Word Masking, RoBERTa-wwm) is used. This model can effectively handle the characteristics of Chinese word segmentation and outputs word-level feature vectors with a dimension of 768. ,in, (This represents the text feature vector). For the English dataset FakeTT, the standard BERT-based-cased model is used, whose twelve-layer Transformer architecture can fully capture English grammatical structures. For visual information, feature extraction is performed using the ImageNet pre-trained VGG19 model. ,in, The video frame features are represented by a convolutional network that extracts 4096-dimensional frame-level features, where m is the number of sampled frames (default setting is 16 frames). For audio information, the original waveform is first converted into a 128-dimensional Mel spectrogram, and then input into the VGGish network to extract 128-dimensional temporal features. , This represents the audio temporal features, where k is the time step. All features must undergo mean-variance normalization before input.

[0079] The fake news video detection method based on information enhancement and guided denoising provided in this embodiment fully explores the value of different modal information by identifying multimodal information, and extracts multimodal features from multimodal information using different models. This achieves multi-dimensional and comprehensive feature extraction of news videos, and accurately captures the unique features of each modality by relying on the adaptive model. This makes the detection based on richer and more accurate features, thereby improving the accuracy and reliability of the detection.

[0080] In step S202 above, semantic enhancement is performed on the label features to obtain enhanced label features, including:

[0081] Step S2023: Obtain multiple related videos of the news video to be detected, and extract the tag features of each related video.

[0082] Specifically, in addition to obtaining feature representations of multimodal information, semantic enhancement is performed on the generated label features through a contrastive learning mechanism to improve the reliability and discriminativeness of the label representations. For example... Figure 4The diagram illustrates the principle of the contrastive learning mechanism. This process requires multiple related videos to the news video being detected, with each video serving as an element to facilitate the construction of a tag semantic similarity matrix. Therefore, the tag features extracted from each related video are visual insight tags (VI-Tag) and content authenticity tags (CV-Tag).

[0083] Step S2024: Using the tag features of the news video to be detected and the tag features of each related video, construct a tag semantic similarity matrix to quantify the semantic association strength between tag pairs.

[0084] Specifically, based on observations of the high correlation between the authenticity of content on the same topic during the news dissemination process, a semantic similarity matrix of tags is constructed. , where each element The label features of the corresponding video are located in the similarity matrix, which precisely quantifies the semantic association strength between the label pairs. x and y represent the row and column of the similarity matrix, respectively.

[0085] Step S2025: Use the dynamic nearest neighbor sampling strategy to select the most similar positive samples in the same category and the most similar negative samples in different categories for each label.

[0086] Specifically, a dynamic k-nearest neighbor sampling strategy is adopted, and each anchor point (CC) is labeled. Select the most similar ones in the same category Positive samples Most similar to different categories negative samples ,in, Indicates positive sample label, This represents the label of a negative sample. For example, such as... Figure 4 As shown, if the anchor label is "A chemical plant exploded, thick smoke...", then its corresponding positive sample label can be "The chemical plant exploded...in thick smoke", and its corresponding negative sample label can be "Thousands of firefighters participated in the rescue operation...". This is just an example, but it is not limited to this.

[0087] In step S2026, a three-layer fully connected encoder is used to map the labels to a low-dimensional semantic space, and a loss function is used for optimization training to obtain label-enhanced features.

[0088] Specifically, a three-layer fully connected encoder maps the labels to a low-dimensional semantic space. The three-layer fully connected encoder is represented as follows:

[0089]

[0090] in, Let represent the initial features of the i-th label. The process utilizes an improved multi-sample triplet loss function for optimization training.

[0091]

[0092] in, This represents the dynamically adjusted boundary threshold (initial value 0.5). and This represents the sample weights calculated based on the attention mechanism. For positive sample attention weights, For negative sample attention weights.

[0093] The fake news video detection method based on information enhancement and guided denoising provided in this embodiment enhances the semantics of tag features, enriches semantic references with relevant videos, and strengthens the accuracy and discriminativeness of tag semantic expression through precise sampling and encoding optimization. This allows tag features to more accurately reflect news semantics and improves the accuracy of feature matching and judgment in subsequent fake news detection.

[0094] Step S203: Perform label-guided denoising based on label enhancement features and multimodal features to obtain multimodal denoising features, and determine whether the news video to be detected is a fake news video based on the multimodal denoising features.

[0095] Specifically, in step S203 above, label-guided denoising is performed based on label enhancement features and multimodal features to obtain multimodal denoising features, including:

[0096] Step S2031: Set the weights of each modal feature based on the binary mask of prior knowledge.

[0097] Specifically, for each original modal feature ,in, Indicates the original text features, Representing original visual features, To represent the original audio features, a gated attention mechanism was designed to achieve label-guided feature enhancement. A binary mask based on prior knowledge was introduced during attention calculation. This limits the attention weight of irrelevant features, allowing the model to focus on feature regions that are highly relevant to the semantics of the labels.

[0098] Step S2032: Using multimodal features as query vectors and label features as key-value pairs, perform multi-head attention calculation to obtain multimodal denoising features.

[0099] Specifically, such as Figure 5 The diagram illustrates the complete process of processing news videos in this embodiment, where multimodal features are obtained through feature extraction. and label-enhanced features and In the label-guided denoising stage, multimodal features are used as query vectors and label enhancement features are used as key-value pairs. Through multi-head attention computation (head=8), the label-aware feature representation is obtained.

[0100]

[0101] in, Representation layer normalization, For linear projection layers, express The query represents the word vector currently being "queried"; The key represents the "identity label" of each word, used to align with or score the query; the value represents the content information of each word, which is the final weighted aggregation object.

[0102] In some optional implementations, step S2032 above includes:

[0103] Step a1: Process multimodal features based on label features to obtain multimodal label features.

[0104] Specifically, such as Figure 5 As shown, after processing by the gating attention module, multimodal label features are obtained. , , , , , The internal processing of the gating attention module is a mature existing technology and will not be described in detail here.

[0105] Step a2: Perform mean pooling on the multimodal label features respectively, and use a two-layer gating network to calculate the weights of each modal label feature.

[0106] Specifically, a dynamic gating fusion network is designed to integrate the purification results guided by different tags. The multimodal tag features processed by CV-Tag and VI-Tag are respectively subjected to mean pooling, and the weights are calculated using a two-layer gating network.

[0107]

[0108] in, The weight vector representing the output of the gating network is obtained by performing a non-linear transformation (ReLU) on the input feature vector, followed by two fully connected layers and a Softmax function. It is used to control the fusion ratio of features guided by different labels. The input features are typically the concatenated or spliced ​​CV-Tag and VI-Tag features, or the mean of both. and These represent the weight matrix and bias of the first fully connected layer, respectively. and These represent the weight matrix and bias of the second fully connected layer, respectively.

[0109] Step a3: Calculate the multimodal denoising features based on the modal label features and their corresponding weights.

[0110] Specifically, after feature fusion, the final denoised multimodal denoising feature representation is as follows:

[0111]

[0112] in, , This represents the two attention weights at the output of the gating network, used for weighted fusion. and Two modal features; This indicates the modal characteristics after purification guided by the Visual Insight Tag (VI-Tag). This indicates the modal characteristics after being cleaned up by Content Authenticity Tags (CV-Tag); This represents the Hadamard product, which is an element-wise multiplication.

[0113] The fake news video detection method based on information enhancement and guided denoising provided in this embodiment accurately filters out irrelevant feature interference by leveraging prior knowledge, and uses an attention mechanism to deeply associate multimodal features with labels. Then, through mean pooling and gating networks to refine modal weights, the denoised multimodal features are made to better fit the semantics of the labels, providing cleaner and more accurate feature support for fake news detection and improving the reliability and accuracy of detection results.

[0114] Specifically, step S203 above, which determines whether the news video to be detected is a fake news video based on multimodal denoising features, includes:

[0115] Step S2033: Input the multimodal denoising features into the encoder to obtain the multimodal fusion features.

[0116] Specifically, the denoised multimodal features are input into the Transformer encoder to obtain multimodal fused features, and finally, mean pooling is used to obtain the joint representation:

[0117]

[0118] in, Indicates multimodal fusion features, Denoising features representing text information Denoising features representing visual information This represents the noise reduction features of audio information.

[0119] Step S2034: Based on multimodal fusion features, use a classifier to determine whether the news video to be detected is a fake news video or a real news video.

[0120] Specifically, based on multimodal fusion features, the detection probability of fake news is output by a classifier to determine whether the news video to be detected is a fake news video or a real news video. The classifier can be a binary classifier, and the classification process is a mature existing technology, which will not be elaborated here.

[0121] The fake news video detection method based on information enhancement and guided denoising provided in this embodiment utilizes an encoder to integrate multimodal information, eliminate modal differences and enhance complementarity, so that the fused features accurately reflect the essence of the news. The classifier makes judgments based on the fused features, and with the help of the true and false patterns learned by the model, it can efficiently and accurately distinguish between fake news videos and real news videos.

[0122] Specifically, to verify the effectiveness of the fake news video detection method based on information augmentation and guided denoising provided in this embodiment, a comprehensive experimental evaluation was conducted on two publicly available benchmark datasets, FakeSV and FakeTT. The FakeSV dataset contains 12,843 Chinese short video news samples, of which 38.7% are fake news, while the FakeTT dataset contains 9,572 English short video samples, of which 41.2% are fake news. Both datasets provide complete visual frame sequences, audio streams, and text information. The experiments divided the datasets into training, validation, and test sets in a 7:2:1 ratio to ensure a balanced data distribution.

[0123] Table 1 compares the proposed method with other methods on the FakeSV and FakeTT datasets, using two evaluation metrics: accuracy and F1 score (Macro F1, the harmonic mean of precision and recall, used to comprehensively evaluate model performance).

[0124] Table 1 Model Performance Comparison Table

[0125]

[0126] In the experimental evaluation, the method proposed in this embodiment was systematically compared with other representative benchmark models, which can be divided into two main categories based on their methodological characteristics: multimodal methods and methods based on large language models. Regarding multimodal methods, six advanced detection models were selected as benchmarks. HCFC-Hou classifies by combining linguistic, acoustic, and user engagement features; HCFC-Medina utilizes the TF-IDF features of the top 100 comments and video title verification features for saliency; FANVM uses adversarial networks to model topic distribution and detect stance inconsistency; TikTec proposes a visual-audio collaborative attention fusion module to capture cross-modal associations; SVFEND uses a cross-modal Transformer and self-attention mechanism to achieve multimodal interaction; and FakingRecipe analyzes material selection preferences from an emotional semantic perspective and considers spatiotemporal editing features. Regarding methods based on large language models, two state-of-the-art LLM benchmarks were selected. GPT-4 uses a zero-shot cue method to detect news headlines and transcribed text; GPT-4V is its enhanced version, adding processing capabilities for visual input. These benchmark models cover different technical approaches, from traditional feature engineering to cutting-edge large language models, providing a multi-dimensional comparative reference for comprehensively evaluating the superiority of the method of this invention.

[0127] As can be seen from the results in Table 1, the present invention significantly outperforms other methods on both datasets.

[0128] This embodiment also provides a fake news video detection device based on information enhancement and guided denoising. This device is used to implement the above embodiments and preferred embodiments, and details already described will not be repeated. As used below, the term "module" can be a combination of software and / or hardware that performs a predetermined function. Although the device described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.

[0129] This embodiment provides a fake news video detection device based on information enhancement and guided denoising, such as... Figure 6 As shown, it includes:

[0130] The tag extraction module 601 is used to acquire the news video to be detected and extract the tag features of the news video to be detected using a large language model.

[0131] The information enhancement module 602 is used to extract multimodal features of the news video to be detected based on the label features, and to perform semantic enhancement on the label features to obtain label-enhanced features.

[0132] The guided denoising module 603 is used to perform label-guided denoising based on label enhancement features and multimodal features to obtain multimodal denoising features, and to determine whether the news video to be detected is a fake news video based on the multimodal denoising features.

[0133] In some alternative implementations, the tag extraction module 601 includes:

[0134] The information extraction unit is used to extract the text and video keyframes of the news video to be detected.

[0135] The tag extraction unit is used to extract content authenticity tags from text using a large language model, and to extract visual insight tags from video keyframes using a large language model.

[0136] In some alternative implementations, the information enhancement module 602 includes:

[0137] The information recognition unit is used to identify textual, visual, and audio information in the news video to be detected.

[0138] The multimodal feature extraction unit is used to extract text features from text information using a natural language processing model, extract visual features from visual information using a convolutional neural network model, and extract audio features from audio information using a Mel spectrogram and an audio neural network model.

[0139] The related video tag feature extraction unit is used to obtain multiple related videos of the news video to be detected and extract the tag features of each related video.

[0140] The similarity matrix construction unit is used to construct a tag semantic similarity matrix by utilizing the tag features of the news video to be detected and the tag features of each related video, so as to quantify the semantic association strength between tag pairs.

[0141] The sample selection unit is used to select the most similar positive samples in the same category and the most similar negative samples in different categories for each label using a dynamic nearest neighbor sampling strategy.

[0142] The feature mapping unit is used to map labels to a low-dimensional semantic space using a three-layer fully connected encoder, and to optimize training using a loss function to obtain label-enhanced features.

[0143] In some alternative implementations, the guided noise reduction module 603 includes:

[0144] The weight setting unit is used to set the weights of each modal feature based on the binary mask of prior knowledge.

[0145] The multi-head attention computation unit is used to perform multi-head attention computation on multimodal features as query vectors and label features as key-value pairs to obtain multimodal denoising features.

[0146] The feature fusion unit is used to input multimodal denoising features into the encoder to obtain multimodal fused features.

[0147] The virtual / real judgment unit is used to determine whether the news video to be detected is a fake news video or a real news video based on multimodal fusion features and a classifier.

[0148] Further functional descriptions of the above modules and units are the same as those in the corresponding embodiments described above, and will not be repeated here.

[0149] The fake news video detection device based on information enhancement and guided noise reduction in this embodiment is presented in the form of a functional unit. Here, a unit refers to an ASIC (Application Specific Integrated Circuit) circuit, a processor and memory that execute one or more software or fixed programs, and / or other devices that can provide the above functions.

[0150] This invention also provides a computer device having the above-described features. Figure 6 The device shown is a fake news video detection device based on information enhancement and guided denoising.

[0151] Please see Figure 7 , Figure 7 This is a schematic diagram of the structure of a computer device provided in an optional embodiment of the present invention, such as... Figure 7 As shown, the computer device includes one or more processors 10, memory 20, and interfaces for connecting the components, including high-speed interfaces and low-speed interfaces. The components communicate with each other via different buses and can be mounted on a common motherboard or otherwise installed as needed. The processors can process instructions executed within the computer device, including instructions stored in or on memory to display graphical information of a GUI on external input / output devices (such as display devices coupled to the interfaces). In some alternative implementations, multiple processors and / or multiple buses can be used with multiple memories and multiple memory modules, if desired. Similarly, multiple computer devices can be connected, each providing some of the necessary operations (e.g., as a server array, a group of blade servers, or a multiprocessor system). Figure 7 Take a processor 10 as an example.

[0152] Processor 10 may be a central processing unit, a network processor, or a combination thereof. Processor 10 may further include a hardware chip. The hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The programmable logic device may be a complex programmable logic device (CAMP), a field-programmable gate array (FPGA), a general-purpose array logic (GDA), or any combination thereof.

[0153] The memory 20 stores instructions executable by at least one processor 10 to cause the at least one processor 10 to perform the method shown in the above embodiments.

[0154] The memory 20 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created based on the use of the computer device. Furthermore, the memory 20 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some alternative embodiments, the memory 20 may optionally include memory remotely located relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0155] The memory 20 may include volatile memory, such as random access memory; the memory may also include non-volatile memory, such as flash memory, hard disk or solid-state drive; the memory 20 may also include a combination of the above types of memory.

[0156] The computer device also includes a communication interface 30 for communicating with other devices or communication networks.

[0157] This invention also provides a computer-readable storage medium. The methods described above according to embodiments of the invention can be implemented in hardware or firmware, or implemented as computer code that can be recorded on a storage medium, or implemented as computer code downloaded via a network and originally stored on a remote storage medium or a non-transitory machine-readable storage medium and then stored on a local storage medium. Thus, the methods described herein can be processed by software stored on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. The storage medium can be a magnetic disk, optical disk, read-only memory, random access memory, flash memory, hard disk, or solid-state drive, etc.; further, the storage medium can also include combinations of the above types of memory. It is understood that computers, processors, microprocessor controllers, or programmable hardware include storage components capable of storing or receiving software or computer code, which, when accessed and executed by the computer, processor, or hardware, implements the methods shown in the above embodiments.

[0158] Although embodiments of the invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention, and such modifications and variations all fall within the scope defined by the appended claims.

Claims

1. A method for detecting fake news videos based on information enhancement and guided denoising, characterized in that, The method includes: The news video to be detected is acquired, and the tag features of the news video to be detected are extracted using a large language model; Based on the tag features, multimodal features of the news video to be detected are extracted, and the tag features are semantically enhanced to obtain tag-enhanced features; Label-guided denoising is performed based on label enhancement features and multimodal features to obtain multimodal denoising features. Then, based on these multimodal denoising features, it is determined whether the news video to be detected is a fake news video, including: The weights of each modal feature are set using binary masks based on prior knowledge; Using multimodal features as query vectors and label features as key-value pairs, multi-head attention computation is performed to obtain multimodal denoising features, including: processing multimodal features based on label features to obtain multimodal label features; performing mean pooling on the multimodal label features respectively, and calculating the weights of each modal label feature using a two-layer gating network; and calculating the multimodal denoising features based on each modal label feature and its corresponding weight.

2. The method according to claim 1, characterized in that, The tag features include: content authenticity tags and visual insight tags. The tag features of the news video to be detected are extracted using a large language model, including: Extract the text and video keyframes of the news video to be detected; Based on the text, a large language model is used to extract content authenticity tags, and based on the video keyframes, a large language model is used to extract visual insight tags.

3. The method according to claim 1, characterized in that, The multimodal features include: text features, visual features, and audio features. Extracting the multimodal features of the news video to be detected includes: Identify textual, visual, and audio information in the news video to be detected; Text features are extracted from the text information using a natural language processing model, visual features are extracted from the visual information using a convolutional neural network model, and audio features are extracted from the audio information using a Mel spectrogram and an audio neural network model.

4. The method according to claim 1, characterized in that, The label features are semantically enhanced to obtain label-enhanced features, including: Obtain multiple related videos of the news video to be detected, and extract the tag features of each related video; By utilizing the tag features of the news video to be detected and the tag features of each related video, a tag semantic similarity matrix is ​​constructed to quantify the semantic association strength between tag pairs. A dynamic nearest neighbor sampling strategy is used to select the most similar positive samples of the same category and the most similar negative samples of different categories for each label. A three-layer fully connected encoder is used to map the labels to a low-dimensional semantic space, and a loss function is used for optimization training to obtain label-enhanced features.

5. The method according to claim 1, characterized in that, Determining whether the news video to be detected is a fake news video based on the multimodal denoising features includes: The multimodal denoising features are input into the encoder to obtain multimodal fusion features; Based on the multimodal fusion features, a classifier is used to determine whether the news video to be detected is a fake news video or a real news video.

6. A fake news video detection device based on information enhancement and guided denoising, characterized in that, The device includes: The tag extraction module is used to acquire news videos to be detected and to extract tag features of the news videos to be detected using a large language model. The information enhancement module is used to extract multimodal features of the news video to be detected based on the tag features, and to perform semantic enhancement on the tag features to obtain tag-enhanced features; A guided denoising module is used to perform label-guided denoising based on label enhancement features and multimodal features to obtain multimodal denoising features, and to determine whether the news video to be detected is a fake news video based on the multimodal denoising features, including: The weight setting unit is used to set the weights of each modal feature based on a binary mask with prior knowledge. The multi-head attention computation unit is used to perform multi-head attention computation on multimodal features as query vectors and label features as key-value pairs to obtain multimodal denoising features. This includes: processing multimodal features based on label features to obtain multimodal label features; performing mean pooling on the multimodal label features respectively; calculating the weights of each modal label feature using a two-layer gating network; and calculating the multimodal denoising features based on each modal label feature and its corresponding weight.

7. A computer device, characterized in that, include: A memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to perform the method of any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions for causing a computer to perform the method of any one of claims 1 to 5.