False News Video Detection Method Based on Knowledge-Enhanced Graph Attention Network

By introducing knowledge-enhanced graph attention networks into video fake news detection, combined with multimodal feature fusion and prediction, the problem that the existing technology cannot fully capture video fake features is solved, and higher detection accuracy and early recognition capabilities are achieved.

CN119723429BActive Publication Date: 2025-06-13HEFEI HUALIHUI INTELLECTUAL PROPERTY OPERATION CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510230611.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-28
Publication Date
2025-06-13
Estimated Expiration
2045-02-28

AI Technical Summary

Technical Problem

The prior art cannot fully capture false features in video content in video fake news detection, especially in processing visual, voice and timing information of video.

Method used

A fake news video detection method based on knowledge-enhanced graph attention network is proposed, subtitles are extracted through optical character recognition technology, text semantic features are extracted by pre-trained language model, speech information is extracted by video editing toolkit, speech emotion features are extracted by pre-trained speech model, visual features are extracted by sliding window model, and fusion and prediction are carried out through multi-layer perception machines.

Benefits of technology

By integrating multimodal features, the accuracy and robustness of fake news video detection are improved, and the semantic understanding ability of scene graphs and the ability to early recognition of fake news are effectively improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119723429B_ABST
    Figure CN119723429B_ABST
Patent Text Reader

Abstract

The present invention proposes a method for detecting fake news videos based on a knowledge-enhanced graph attention network. The method includes: extracting semantic features of text based on complete text information through a pre-trained language model; extracting emotional features of speech using a pre-trained speech model; extracting visual features from key frames using a pre-trained sliding window model; obtaining the global comment feature of the video by using a pre-trained stance detection model in combination with the stance feature of the video and the weight value of the comment; obtaining the social feature of the user using the social feature of the video publisher; obtaining the temporal scene graph feature through a scene graph sequence; fusing multi-modal features to obtain a multi-modal feature representation; and obtaining a prediction result based on the multi-modal feature representation. The present invention provides a comprehensive and efficient fake news video detection solution, significantly improving the detection accuracy and practical application value.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of computer vision and natural language processing, and particularly relates to a method for detecting fake news videos based on a knowledge-enhanced graph attention network. Background Art

[0002] With the rapid development of the Internet and social media, the problem of the spread of fake news has become increasingly serious, especially for fake news in video content. As a powerful communication medium, the visual and auditory effects of videos make fake news more deceptive. Traditional methods for detecting fake news mainly rely on text analysis, usually focusing on text information such as the title, article content, or comments of the news. However, this method cannot fully utilize the visual, speech, and temporal information in the video. Therefore, there are many limitations in the existing technology for detecting fake news in videos, and it is impossible to comprehensively capture the fake features in the video content.

[0003] Currently, the technologies for detecting fake news in videos can be divided into methods based on visual information, methods based on speech / subtitle information, and methods based on text comment analysis. However, these methods are usually single-modal and cannot comprehensively analyze various types of information in the video. For example, visual content may not be able to independently determine whether a video is fake news, and speech or subtitle analysis is also difficult to fully reflect the overall semantic information of the video. In addition, the video itself has temporal characteristics, and the change and development of information are dynamic. Many existing methods have not effectively processed these temporal information, resulting in deficiencies in the early detection and change capture of fake news. Summary of the Invention

[0004] In view of the above situation, the main purpose of the present invention is to propose a method for detecting fake news videos based on a knowledge-enhanced graph attention network to solve the above technical problems.

[0005] The present invention proposes a method for detecting fake news videos based on a knowledge-enhanced graph attention network, and the method includes the following steps:

[0006] Step 1: Use optical character recognition technology (OCR) to extract the subtitles in the video, and then splice the subtitles with the video title to obtain complete text information;

[0007] Based on the complete text information, extract the semantic features of the text through a pre-trained language model (BERT);

[0008] Step 2: Use an open-source toolkit for video editing (MoviePy) to extract the speech information from the video, and then use a pre-trained speech model (HuBERT) to extract the emotional features of the speech;

[0009] Step 3: Extract key frames in the video by calculating the differences between adjacent frames, and then use the pre-trained sliding window model (Swin Transformer) to extract the key frames to obtain visual features;

[0010] Step 4: Use the pre-trained stance detection model (StanceBERTa) to extract the stance features of video comments, and then combine the stance features of the video and the weight values of the comments to obtain the global comment features of the video;

[0011] Step 6: Extract the social features of the video publisher, and then use the social features of the video publisher to obtain the social features of the user;

[0012] Step 6: Use an unbiased scene graph generation method to convert the key frame sequence of the video into a scene graph sequence, and extract the scene graph sequence through a scene graph attention network to obtain temporal scene graph features;

[0013] Step 7: Fuse the semantic features of the text, the emotional features of the speech, the visual features, the global comment features of the video, the social features of the user, and the temporal scene graph features to obtain a global multi-modal feature representation;

[0014] Step 8: Input the global multi-modal feature representation into a multi-layer perceptron to obtain a prediction result.

[0015] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0016] 1. By fusing text, speech, vision, comment, and social features, the present invention improves the accuracy and robustness of fake news video detection;

[0017] 2. By introducing an external knowledge graph for knowledge enhancement, the present invention effectively improves the semantic understanding ability of the scene graph;

[0018] 3. By using a dynamic scene graph attention network to capture the temporal features in the video, the present invention improves the ability to identify fake news at an early stage.

[0019] Additional aspects and advantages of the present invention will be given in part in the following description, become apparent in part from the following description, or be learned through the embodiments of the present invention. Brief Description of the Drawings

[0020] Figure 1 It is a flowchart of the fake news video detection method based on a knowledge-enhanced graph attention network proposed by the present invention;

[0021] Figure 2 It is a model architecture diagram of the fake news video detection method based on a knowledge-enhanced graph attention network proposed by the present invention. Detailed Embodiments

[0022] Embodiments of the present invention will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present invention and should not be construed as limiting the present invention.

[0023] Referring to the following description and the accompanying drawings, these and other aspects of the embodiments of the present invention will become clear. In these descriptions and drawings, some specific embodiments of the embodiments of the present invention are specifically disclosed to represent some ways of implementing the principles of the embodiments of the present invention, but it should be understood that the scope of the embodiments of the present invention is not limited by this.

[0024] Please refer to Figure 1 , an embodiment of the present invention proposes a method for detecting fake news videos based on a knowledge-enhanced graph attention network. The method includes the following steps:

[0025] Step 1: Use optical character recognition technology to extract the subtitles in the video, and then splice the subtitles with the video title to obtain complete text information;

[0026] Based on the complete text information, extract through a pre-trained language model to obtain the semantic features of the text;

[0027] In step 1, based on the complete text information, extract through a pre-trained language model to obtain the semantic features of the text. The relational formula for the corresponding process is:

[0028] ;

[0029] Wherein, represents the semantic features of the text, represents the complete text information, represents the extraction process through the pre-trained language model.

[0030] Step 2: Use an open-source toolkit for video editing to extract the speech information from the video, and then use a pre-trained speech model to extract the speech information to obtain the emotional features of the speech;

[0031] In step 2, use an open-source toolkit for video editing to extract the speech information from the video, and then use a pre-trained speech model to extract the speech information to obtain the emotional features of the speech. The specific steps are as follows:

[0032] Perform noise reduction processing on the speech information to obtain the noise-reduced speech information. The relational formula for the corresponding process is:

[0033] ;

[0034] Among them, represents the denoised speech information, indicating being processed by the denoising function, represents the speech information;

[0035] The denoised speech information is normalized to obtain the normalized speech information, and the relational formula existing in the corresponding process is:

[0036] ;

[0037] Among them, represents the normalized speech information, indicating being processed by the normalization function;

[0038] The normalized speech information is input into the pre-trained speech model to obtain the emotional features of the speech, and the relational formula existing in the corresponding process is:

[0039] ;

[0040] Among them, represents the emotional features of the speech, indicating being processed by the pre-trained speech model.

[0041] Step 3: Extract the key frames in the video by calculating the differences between adjacent frames, and then use the pre-trained sliding window model to extract the key frames to obtain the visual features;

[0042] In Step 3, the key frames in the video are extracted by calculating the differences between adjacent frames, and then the key frames are extracted using the pre-trained sliding window model to obtain the visual features. The specific steps are as follows:

[0043] Perform frame difference calculation on the video frame sequence to obtain the difference values between adjacent frames. The relational formula existing in the corresponding process is:

[0044] ;

[0045] Among them, represents the th frame and the th frame of the video, represents the th frame of the video;

[0046] According to the difference values, filter the key frames through the set difference value threshold to obtain the key frame sequence;

[0047] Input the key-frame sequence into the pre-trained sliding window model to obtain the visual features of the key frames. The relational expression in the corresponding process is as follows:

[0048] ;

[0049] where, represents the visual feature of the -th key frame, represents being processed by the pre-trained sliding window model, represents the -th key frame;

[0050] Collect the visual features of all key frames to obtain the visual feature set of key frames;

[0051] Process the visual feature set of key frames through the self-attention mechanism to obtain visual features. The relational expression in the corresponding process is as follows:

[0052] ;

[0053] where, represents the visual feature, represents being processed by the self-attention mechanism, represents the visual feature set of key frames.

[0054] Step 4: Use the pre-trained stance detection model to extract the stance features of video comments, and then combine the stance features of the video and the weight values of the comments to obtain the global comment features of the video;

[0055] In Step 4, use the pre-trained stance detection model to extract the stance features of video comments, and then combine the stance features of the video and the weight values of the comments to obtain the global comment features of the video. The specific steps are as follows:

[0056] Collect the comment data set related to the video. Each comment in the comment data set consists of the text content of the comment and the number of likes;

[0057] Use the pre-trained stance detection model to encode the text content of each comment in the comment to obtain the stance features. The relational expression in the corresponding process is as follows:

[0058] ;

[0059] where, represents the stance feature, represents being processed by the pre-trained stance detection model, represents the -th comment text content;

[0060] Calculate the weight value of a comment based on the number of likes for each comment. The relational expression for the corresponding process is as follows:

[0061] ;

[0062] Among them, represents the weight value of the th comment, represents the number of likes for the th comment, represents the total number of comments on the video, represents the number of likes for the th comment;

[0063] Combine the weight values of the comments to perform a weighted sum of all stance features to obtain the global comment feature of the video. The relational expression for the corresponding process is as follows:

[0064] ;

[0065] Among them, represents the global comment feature of the video.

[0066] Step 5: Extract the social features of the video publisher, and then use the social features of the video publisher to obtain the social features of the user;

[0067] In Step 5, extract the social features of the video publisher, and then use the social features of the video publisher to obtain the social features of the user. The specific steps are as follows:

[0068] Obtain the relevant attributes of the publisher from the metadata of the video, and then construct the user initial feature vector according to the relevant attributes of the publisher. The relational expression for the corresponding process is as follows:

[0069] ;

[0070] Among them, represents the user initial feature vector, represents the number of fans, represents the number of friends, represents the number of videos published, represents the interaction characteristics of the publisher;

[0071] Perform normalization processing on the user initial feature vector to obtain the social features of the user. The relational expression for the corresponding process is as follows:

[0072] ;

[0073] Among them, represents the social features of the user.

[0074] Step 6: Use an unbiased scene graph generation method to convert the key frame sequence of the video into a scene graph sequence, and extract the temporal scene graph features from the scene graph sequence through a scene graph attention network;

[0075] In Step 6, use an unbiased scene graph generation method to convert the key frame sequence of the video into a scene graph sequence, and extract the temporal scene graph features from the scene graph sequence through a scene graph attention network. The specific steps are as follows:

[0076] Apply a pre-trained object detection model to the key frame sequence to detect visual objects in each key frame. Among them, each visual object includes a class label and a bounding box, and the relational expression in the corresponding process is:

[0077] ;

[0078] Among them, represents the visual object in the th key frame, represents the detection process of the pre-trained object detection model;

[0079] Model the spatial and semantic relationships between visual objects to generate a preliminary relationship graph between visual objects;

[0080] Further optimize the preliminary relationship graph between visual objects to generate the scene graph corresponding to the key frame;

[0081] Arrange the scene graphs corresponding to all key frames in chronological order to form a scene graph sequence;

[0082] Based on the nodes and relationships in the scene graph sequence, obtain a query keyword set. The relational expression in the corresponding process is:

[0083] ;

[0084] Among them, represents the query keyword set, all represent query keywords, represents a node, represents a relationship;

[0085] Use the query keyword set to retrieve knowledge triples in an external knowledge graph;

[0086] Filter the knowledge triples and retain the knowledge related to the objects and relationships in the scene graph;

[0087] Add the knowledge related to the objects and relationships in the scene graph as supplementary nodes to the scene graph sequence to obtain a knowledge-augmented scene graph sequence;

[0088] Construct a scene graph attention network using the nodes in the scene graph sequence. The relational expression for the construction process is as follows:

[0089] ;

[0090] Among them, represents the scene graph attention network, represents the node the aggregated and updated feature, represents the number of multi-head attentions, represents the concatenation operation, represents the node the set of adjacent nodes of, represents at the th attention type, the adjacent node to the node the attention weight of, , and all represent parameter matrices, represents the feature representation of the node , represents the feature representation of the node , represents the feature representation of the edge , represents after the normalization operation, represents after being processed by the activation function, represents the transpose of the attention weight;

[0091] Based on the structure of the recurrent neural network, use the output of the scene graph attention network at the previous moment as the input of the scene graph attention network at the next moment to construct a dynamic scene graph attention network;

[0092] Input the scene graph sequence after knowledge augmentation into the dynamic scene graph attention network to obtain the temporal scene graph features. The relational expression for the corresponding process is as follows:

[0093] ;

[0094] Among them, represents the temporal scene graph features, represents being processed by the self-attention mechanism, represents being processed by the dynamic scene graph attention network, all represent the scene graphs after knowledge augmentation.

[0095] Step 7: Integrate the semantic features of the text, the emotional features of the speech, the visual features, the global comment features of the video, the social features of the user, and the temporal scene graph features to obtain the global multi-modal feature representation;

[0096] In step 7, the semantic features of the text, the emotional features of the speech, the visual features, the global comment features of the video, the social features of the user, and the temporal scene graph features are fused to obtain a global multi-modal feature representation. The specific steps are as follows:

[0097] Use a fully connected neural network to align the semantic features of the text, the emotional features of the speech, the visual features, the global comment features of the video, the social features of the user, and the temporal scene graph features to the same feature space, and obtain the aligned feature representation. The relational expression for the corresponding process is:

[0098] ;

[0099] Among them, represents the aligned feature representation, represents being processed by a fully connected neural network; represents the semantic features of the text , the emotional features of the speech , the visual features , the global comment features of the video , the social features of the user and the temporal scene graph features any one of;

[0100] Based on the aligned feature representation, process it through a self-attention mechanism to obtain the modal weights. The relational expression for the corresponding process is:

[0101] ;

[0102] Among them, represents the modal weights, represents being processed by an evaluation function of modal feature importance;

[0103] According to the modal weights, perform a weighted sum on the aligned feature representation to obtain the global multi-modal feature representation. The relational expression for the corresponding process is:

[0104] ;

[0105] Among them, represents the global multi-modal feature representation.

[0106] Specifically, in this step, the evaluation function of modal feature importance is implemented through self-attention.

[0107] Step 8, input the global multi-modal feature representation into a multi-layer perceptron to obtain a prediction result;

[0108] In step 8, input the global multi-modal feature representation into a multi-layer perceptron to obtain a prediction result. The relational expression for the corresponding process is:

[0109] ;

[0110] Among them, represents the prediction result, represents being processed by a multi-layer perceptron, represents the parameters of the linear layer, represents the bias term.

[0111] It should be understood that each part of the present invention can be implemented by hardware, software, firmware, or a combination thereof. In the above embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented by hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0112] In the description of this specification, the descriptions referring to terms such as "one embodiment", "some embodiments", "example", "specific example", or "some examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present invention. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in a suitable manner in any one or more embodiments or examples.

[0113] The above-described embodiments merely represent several implementation manners of the present invention, and their descriptions are relatively specific and detailed, but should not be construed as limiting the scope of the patent of the present invention. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present invention, several modifications and improvements can still be made, and these all belong to the protection scope of the present invention. Therefore, the protection scope of the patent of the present invention should be subject to the appended claims.

Claims

1. A fake news video detection method based on knowledge-enhanced graph attention network, characterized in that: The method comprises the following steps: Step 1: Use optical character recognition technology to extract subtitles from the video, and then splice the subtitles with the video title to obtain complete text information; Based on the complete text information, the semantic features of the text are extracted through the pre-trained language model; Step 2: Use an open source toolkit for video editing to extract voice information from the video, and then use a pre-trained voice model to extract the voice information to obtain the emotional features of the voice; Step 3: extract key frames in the video by calculating the difference between adjacent frames, and then use the pre-trained sliding window model to extract the key frames to obtain visual features; Step 4: Use the pre-trained stance detection model to extract the stance features of the video comments, and then combine the stance features of the video and the weight of the comments to obtain the global comment features of the video; Step 5: extract the social features of the video publisher, and then use the social features of the video publisher to obtain the social features of the user; Step 6: Use an unbiased scene graph generation method to convert the key frame sequence of the video into a scene graph sequence, extract the scene graph sequence through a scene graph attention network, and obtain the temporal scene graph features; Step 7: Fuse the semantic features of the text, the emotional features of the speech, the visual features, the global comment features of the video, the social features of the user, and the temporal scene graph features to obtain a global multimodal feature representation; Step 8: Input the global multimodal feature representation into the multi-layer perceptron to obtain the prediction result.

2. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 1 is characterized in that: In step 1, based on the complete text information, the semantic features of the text are extracted through the pre-trained language model, and the corresponding relationship in the process is: ; in, Represents the semantic features of the text, Represents complete text information. Indicates that it has been extracted by a pre-trained language model.

3. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 2 is characterized in that: In step 2, an open source toolkit for video editing is used to extract voice information from the video, and then a pre-trained voice model is used to extract the voice information to obtain the emotional features of the voice. The specific steps are as follows: The voice information is subjected to noise reduction processing to obtain the noise-reduced voice information. The relationship between the corresponding process is: ; in, Indicates the voice information after noise reduction. It means that it has been processed by the noise reduction function. Indicates voice information; The denoised speech information is normalized to obtain the normalized speech information. The corresponding relationship is: ; in, represents the normalized speech information, Indicates that it has been processed by the normalization function; The normalized speech information is input into the pre-trained speech model to obtain the emotional characteristics of the speech. The corresponding relationship in the process is: ; in, Indicates the emotional characteristics of speech. Represents pre-trained speech model processing.

4. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 3 is characterized in that: In step 3, the key frames in the video are extracted by calculating the difference between adjacent frames, and then the key frames are extracted using a pre-trained sliding window model to obtain visual features. The specific steps are as follows: The frame difference calculation is performed on the video frame sequence to obtain the difference value between adjacent frames. The relationship between the corresponding process is: ; in, Indicates Frame and The difference between frames, Indicates that after being processed by the inter-frame difference calculation function, Indicates the video frame, Indicates the video frame; According to the difference value, the key frames are screened by a set difference value threshold to obtain a key frame sequence; The key frame sequence is input into the pre-trained sliding window model to obtain the visual features of the key frames. The relationship between the corresponding process is: ; in, Indicates The visual features of key frames, represents the pre-trained sliding window model processing, Indicates keyframes; Collect the visual features of all key frames to obtain a visual feature set of the key frames; The visual feature set of the key frame is processed through the self-attention mechanism to obtain the visual feature. The relationship between the corresponding process is: ; in, Represents visual features, It means that it has been processed by the self-attention mechanism. A collection of visual features representing a keyframe.

5. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 4 is characterized in that: In step 4, the stance features of the video comments are extracted using the pre-trained stance detection model, and then the global comment features of the video are obtained by combining the stance features of the video and the weight values ​​of the comments. The specific steps are as follows: Collect a comment dataset related to the video. Each comment in the comment dataset consists of the text content of the comment and the number of likes. The pre-trained stance detection model is used to encode the text content of each comment to obtain stance features. The corresponding relationship is: ; in, Indicates the characteristics of the position, It means processed by the pre-trained stance detection model. Indicates The text content of the comment; The weight of the comment is calculated based on the number of likes for each comment. The corresponding relationship is: ; in, Indicates The weight of the comments, Indicates Number of likes for comments, Indicates the total number of comments on the video. Indicates Number of likes for a comment; Combine the weight values ​​of the comments to perform weighted summation of all stance features to obtain the global comment features of the video. The corresponding relationship is: ; in, Represents the global comment features of the video.

6. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 5 is characterized in that: In step 5, the social features of the video publisher are extracted, and then the social features of the user are obtained by using the social features of the video publisher. The specific steps are as follows: The publisher's related attributes are obtained from the video's metadata, and then the user's initial feature vector is constructed based on the publisher's related attributes. The corresponding process has the following relationship: ; in, represents the user's initial feature vector, Indicates the number of fans, Indicates the number of friends, Indicates the number of published videos. Represents the interactive characteristics of the publisher; The user's initial feature vector is normalized to obtain the user's social features. The corresponding relationship is: ; in, Represents the social characteristics of the user.

7. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 6 is characterized in that: In step 6, the key frame sequence of the video is converted into a scene graph sequence using an unbiased scene graph generation method, and the scene graph sequence is extracted through a scene graph attention network to obtain a temporal scene graph feature. The specific steps are: as follows: Apply the pre-trained object detection model to the keyframe sequence to detect the visual objects in each keyframe, where each visual object includes a category label and a bounding box. The relationship between the corresponding process is: ; in, Indicates visual objects in keyframes, Represents the detection process of the pre-trained object detection model; Modeling the spatial and semantic relationships between visual objects to generate a preliminary relationship graph between visual objects; Further optimize the preliminary relationship graph between visual objects to generate a scene graph corresponding to the key frames; Arrange the scene graphs corresponding to all key frames in chronological order to form a scene graph sequence; Based on the nodes and relationships in the scene graph sequence, a set of query keywords is obtained, and the relationship between the corresponding process is: ; in, Represents a set of query keywords. All represent query keywords. Represents a node, To express a relationship; Use the query keyword set in the external knowledge graph to retrieve knowledge triples; Filter the knowledge triples to retain the knowledge related to the objects and relations in the scene graph; The knowledge related to the objects and relations in the scene graph is added to the scene graph sequence as a supplementary node to obtain the scene graph sequence after knowledge expansion; The scene graph attention network is constructed using the nodes in the scene graph sequence. The relationship of the construction process is: ; in, represents the scene graph attention network, Representation Node Aggregate the updated features, represents the number of multi-head attention, Represents a splicing operation, Representation Node The set of adjacent nodes of Indicated in Adjacent nodes in attention type For Node The attention weight, , and They all represent parameter matrices, Representation Node The characteristic representation of Representation Node The characteristic representation of Represents edge The characteristic representation of It means that after normalization operation, It means that it has been processed by the activation function. represents the transpose of the attention weights; Based on the structure of recurrent neural network, the output of the scene graph attention network at the previous moment is used as the input of the scene graph attention network at the next moment to construct a dynamic scene graph attention network; The scene graph sequence after knowledge expansion is input into the dynamic scene graph attention network to obtain the temporal scene graph features. The relationship between the corresponding process is: ; in, Represents the temporal scene graph features, It means that it is processed by the self-attention mechanism. It means that it has been processed by the dynamic scene graph attention network. Both represent scene graphs after knowledge expansion.

8. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 7 is characterized in that: In step 7, the semantic features of the text, the emotional features of the speech, the visual features, the global comment features of the video, the social features of the user and the temporal scene graph features are fused to obtain a global multimodal feature representation. The specific steps are as follows: A fully connected neural network is used to align the semantic features of text, emotional features of speech, visual features, global comment features of video, social features of users, and temporal scene graph features into the same feature space to obtain the aligned feature representation. The relationship between the corresponding processes is: ; in, represents the aligned feature representation, Indicates processing by a fully connected neural network; Represents the semantic features of text , emotional characteristics of speech , visual features , global comment features of the video , user's social characteristics and temporal scene graph features Any of the following: Based on the aligned feature representation, the modal weight is obtained through the self-attention mechanism. The corresponding relationship is: ; in, represents the modal weight, It means that it has been processed by evaluating the importance function of modal features; According to the modal weights, the aligned feature representations are weighted and summed to obtain the global multimodal feature representation. The corresponding process has the following relationship: ; in, Represents global multimodal feature representation.

9. The method for detecting fake news videos based on knowledge-enhanced graph attention network according to claim 8, characterized in that: In step 8, the global multimodal feature representation is input into the multilayer perceptron to obtain the prediction result. The relationship between the corresponding process is: ; in, Represents the prediction result, It means processed by multi-layer perceptron. represents the linear layer parameters, Represents the bias term.

Citation Information

Patent Citations

  • Multi-platform collaborative new media content monitoring management system based on big data

    CN113177164A

  • Video processing method and device, and computer readable storage medium

    CN117061816A