Fake news video detection method and device based on creative process perspective guidance

By analyzing audio and text emotional features, visual and text semantic features, and combining cross-modal Transformer models and multi-layer perceptrons, the creation process of fake news videos can be identified, solving the misleading problem of fake news video detection in existing technologies and achieving more efficient detection performance.

CN118506235BActive Publication Date: 2025-10-03INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202410608989.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-05-16
Publication Date
2025-10-03
Estimated Expiration
2044-05-16

AI Technical Summary

Technical Problem

Existing fake news video detection methods are easily misled on short video platforms and find it difficult to effectively distinguish between true and fake news videos, especially since editing behaviors are difficult to identify due to the popularity of video editing tools and the openness of the platforms.

Method used

By extracting audio and text emotional features, visual and text semantic features, and fusing them with a cross-modal Transformer model, we analyze the spatial and temporal editing behavior characteristics of videos and use a multi-layer perceptron to detect fake news.

Benefits of technology

It significantly improves the performance of fake news video detection, can accurately identify fake news videos without relying on external evidence, and enhances the performance of existing detection models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN118506235B_ABST
    Figure CN118506235B_ABST
Patent Text Reader

Abstract

The present invention proposes a method and device for detecting fake news videos guided by the perspective of the creative process, through feature modeling inspired by material selection behavior. The material selection behavior of video creation is analyzed from the two aspects of emotion and semantics, and the news authenticity prediction result inspired by the material selection behavior is obtained. Feature modeling inspired by material editing behavior. The material editing behavior of video creation is analyzed from the two aspects of spatial editing behavior and temporal editing behavior, and the news video authenticity prediction result inspired by the material editing behavior is obtained. A dual-perspective fake news video detection method guided by the creative process is used to obtain the news video authenticity prediction result from the two perspectives of comprehensive material selection and material editing. The present invention no longer only analyzes the shallow pattern of the content presented in the fake news video, but analyzes and speculates the creation process of the fake news video.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the fields of computer application technology and natural language processing technology, and in particular to a method, device, electronic device and storage medium for detecting fake news videos guided by a creative process perspective. Background Art

[0002] In recent years, the popularity of short video platforms like Douyin and Kuaishou has led to a gradual shift in traditional news presentation formats towards short videos. However, the prevalence of video news has also led to the proliferation of fake news videos, posing a threat to the internet content ecosystem. Fake news video detection has become an emerging and urgent need in multimodal fake news detection. Fake news videos typically consist of a headline and a video. The fake news video detection task is often formulated as a binary classification problem, where the goal is to classify an input news video as "true" or "fake."

[0003] Existing fake news video detection methods mostly follow the same approach used for text or image news, focusing on uncovering detection clues by analyzing the authenticity of multimodal content (e.g., detecting deepfakes) and modeling cross-modal correlations in the presented content. However, the characteristics of fake news videos on short video platforms differ, posing challenges to traditional detection methods. First, the widespread availability of video editing tools has lowered the barrier to entry for video editing, making editing commonplace. The presence or absence of editing is no longer an effective criterion for distinguishing true from false news videos. Second, the open nature of short video platforms makes it possible for historically authentic news videos to be downloaded, re-edited, and uploaded for use in fake news production. These characteristics further blur the line between true and false news, making traditional detection methods susceptible to misleading and incorrect judgments. Summary of the Invention

[0004] The purpose of this invention is to effectively and automatically detect fake news videos. In response to the problem that existing technologies have limited authenticity perception analysis of news short videos, a fake news video detection method guided by the creative process perspective is proposed.

[0005] Specifically, the present invention is as follows Figure 5 As shown in the figure, a fake news video detection method based on the creative process perspective is proposed, which includes:

[0006] Emotional feature extraction step 1: obtaining a video to be detected as fake news, extracting audio emotional features of the video and text emotional features of the video title and subtitle, and fusing the audio emotional features and text emotional features to obtain a multimodal emotional feature;

[0007] Semantic feature extraction step 2: Interval sampling of the video's key frames, extracting the visual semantic features of the key frames, and extracting the textual semantic features of the subtitles; using a cross-modal Transformer model to interact with the visual semantic features and the textual semantic features, obtaining visually enhanced text features and text-enhanced visual features, and splicing the two into the Transformer model for fusion to obtain multimodal semantic features;

[0008] In step 3 of spatial feature extraction, rich text visual frames in the video are selected based on the size of the text region in the video image; the text region is framed from the rich text visual frame to obtain a text frame, and the text frame is encoded to obtain a prompt feature; the prompt feature and the image feature of the rich text visual frame are input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text region to obtain a spatial editing behavior feature;

[0009] Temporal feature extraction step 4: extracting the text temporal sequence of the video based on the number, content, and time interval of the text segments in the video; extracting the visual temporal sequence of the video based on the number, key frames, and time interval of the visual segments in the video; extracting the features of each segment in the text temporal sequence and the visual temporal sequence respectively to obtain text segment features and visual segment features, adding position coding and duration coding to the text segment features and the visual segment features respectively, and fusing them to obtain temporal editing behavior features;

[0010] In feature fusion detection step 5, the multimodal emotional feature and the multimodal semantic feature are spliced ​​and input into a multi-layer perceptron to obtain a selection behavior feature; the spatial editing behavior feature and the temporal editing behavior feature are spliced ​​and input into a multi-layer perceptron to obtain an editing behavior feature; the selection behavior feature and the editing behavior feature are fused and input into a binary classification model to obtain a fake news video detection result for the video.

[0011] The fake news video detection method based on the creative process perspective is shown, wherein the emotional feature extraction step includes:

[0012] In terms of emotion, an audio pre-training model and a text pre-training model are used to extract the audio emotion features and the text emotion features respectively. The audio emotion features and the text emotion features are concatenated and sent to the Transformer model for fusion to obtain the multimodal emotion features.

[0013] The method for detecting fake news videos based on the creative process perspective shown in FIG, wherein the spatial feature extraction step includes:

[0014] A text detector is used to locate the text area from the rich-text visual frame to obtain a plurality of text boxes, which are then encoded using a prompt encoder to obtain the prompt feature. An image classification model is used to encode the rich-text visual frame to obtain the image feature, which is then input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a guide feature, which is then downsampled to obtain the spatial editing behavior feature.

[0015] The fake news video detection method based on the creative process perspective guidance is shown, wherein the temporal feature extraction step includes:

[0016] The text time sequence , n is the number of text segments, For the text fragment content, is the time interval corresponding to the fragment;

[0017] The visual timing sequence , m is the number of visual segments, is the key frame of the visual segment, is the time interval corresponding to the fragment; , are the start frame index and end frame index of the segment;

[0018] The visual time sequence and the text time sequence are input into the hierarchical temporal structure extractor as the material time sequence. For the text time sequence, the text segment features are obtained by splicing multiple text segments in the same time period and inputting them into the encoder. For the visual time sequence, the visual segment features are obtained by fusing multiple frames corresponding to the same time period using self-attention.

[0019] The position code The encoding strategy is:

[0020]

[0021] in , i is the temporal relative position of the fragment, The position encoding vector of the fragment dimension, k range value ;

[0022] The duration code The encoding includes relative duration encoding and absolute duration encoding; the relative duration and absolute duration are:

[0023]

[0024] Where fps is the frame rate of the video, and vframes is the total number of frames of the video;

[0025] Duration encoding of the i-th segment Expressed as:

[0026]

[0027] Add position encoding and duration coding Then, we get the fragment features ;

[0028] Multiple segments are interacted using the self-attention mechanism to obtain the temporal structure features of a specific modality:

[0029]

[0030] The hierarchical temporal structure extractor is used to process the text temporal sequence input and the visual temporal sequence input respectively to obtain the text temporal editing behavior features. and visual time editing behavior characteristics ; Using Transformer fusion and Get the editing behavior characteristics of this time .

[0031] The present invention also proposes a fake news video detection device B based on the creative process perspective guidance, such as Figure 6 shown, including:

[0032] The emotional feature extraction module M1 obtains the video to be detected as fake news, extracts the audio emotional features of the video and the text emotional features of the video title and subtitle, and fuses the audio emotional features and the text emotional features to obtain a multimodal emotional feature;

[0033] Semantic feature extraction module M2 samples the key frames of the video at intervals, extracts the visual semantic features of the key frames, and extracts the textual semantic features of the subtitles. It uses a cross-modal Transformer model to interact with the visual semantic features and the textual semantic features to obtain visually enhanced text features and text-enhanced visual features. The two are then concatenated and fed into the Transformer model for fusion to obtain multimodal semantic features.

[0034] The spatial feature extraction module M3 selects rich text visual frames in the video based on the size of the text area in the video image; selects the text area from the rich text visual frame to obtain a text box, and encodes the text box to obtain a prompt feature; the prompt feature and the image feature of the rich text visual frame are input into the bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain the spatial editing behavior feature;

[0035] The temporal feature extraction module M4 extracts the text temporal sequence of the video based on the number, content, and time interval of the text segments in the video; extracts the visual temporal sequence of the video based on the number, key frames, and time interval of the visual segments in the video; extracts the features of each segment in the text temporal sequence and the visual temporal sequence respectively to obtain text segment features and visual segment features, adds position coding and duration coding to the text segment features and the visual segment features respectively, and then fuses them to obtain temporal editing behavior features;

[0036] The feature fusion detection module M5 splices the multimodal emotional feature and the multimodal semantic feature into a multi-layer perceptron to obtain a selection behavior feature; splices the spatial editing behavior feature and the temporal editing behavior feature into a multi-layer perceptron to obtain an editing behavior feature; and fuses the selection behavior feature and the editing behavior feature into a binary classification model to obtain a false news video detection result for the video.

[0037] The fake news video detection device based on the creative process perspective guidance shown in the figure, wherein the emotional feature extraction module includes:

[0038] In terms of emotion, an audio pre-training model and a text pre-training model are used to extract the audio emotion features and the text emotion features respectively. The audio emotion features and the text emotion features are concatenated and sent to the Transformer model for fusion to obtain the multimodal emotion features.

[0039] The fake news video detection device based on the creative process perspective guidance shown in the figure, wherein the spatial feature extraction module includes:

[0040] A text detector is used to locate the text area from the rich-text visual frame to obtain a plurality of text boxes, which are then encoded using a prompt encoder to obtain the prompt feature. An image classification model is used to encode the rich-text visual frame to obtain the image feature, which is then input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a guide feature, which is then downsampled to obtain the spatial editing behavior feature.

[0041] The fake news video detection device based on the creative process perspective guidance shown in the figure, wherein the time feature extraction module includes:

[0042] The text time sequence , n is the number of text segments, For the text fragment content, is the time interval corresponding to the fragment;

[0043] The visual timing sequence , m is the number of visual segments, is the key frame of the visual segment, is the time interval corresponding to the fragment; , are the start frame index and end frame index of the segment;

[0044] The visual time sequence and the text time sequence are input into the hierarchical temporal structure extractor as the material time sequence. For the text time sequence, the text segment features are obtained by splicing multiple text segments in the same time period and inputting them into the encoder. For the visual time sequence, the visual segment features are obtained by fusing multiple frames corresponding to the same time period using self-attention.

[0045] The position code The encoding strategy is:

[0046]

[0047] in , i is the temporal relative position of the fragment, The first position encoding vector of the segment dimension, k range value ;

[0048] The duration code The encoding includes relative duration encoding and absolute duration encoding; the relative duration and absolute duration are:

[0049]

[0050] Where fps is the frame rate of the video, and vframes is the total number of frames of the video;

[0051] Duration encoding of the i-th segment Expressed as:

[0052]

[0053] Add position encoding and duration coding Then, we get the fragment features ;

[0054] Multiple segments are interacted using the self-attention mechanism to obtain the temporal structure features of a specific modality:

[0055]

[0056] The hierarchical temporal structure extractor is used to process the text temporal sequence input and the visual temporal sequence input respectively to obtain the text temporal editing behavior features. and visual time editing behavior characteristics ; Using Transformer fusion and Get the editing behavior characteristics of this time .

[0057] The present invention also proposes an electronic device, comprising the aforementioned fake news video detection device based on creation process perspective guidance.

[0058] The electronic device is connected to an information display device, which is used to display the fake news video detection results using display parameters and attributes set by the user or through an artificial intelligence model.

[0059] The present invention also proposes a storage medium for storing a computer program for executing the fake news video detection method based on the creative process perspective guidance.

[0060] In view of the deficiencies in the prior art, the present invention proposes a

[0061] It can be seen from the above scheme that the advantages of the present invention are:

[0062] Without introducing external evidence, user comments and feedback and other information, the present invention achieves significant performance improvement compared with the existing technology. In addition, the feature modeling method inspired by material editing behavior in the present invention can be combined with existing detection models to help the existing detection models achieve performance improvement. BRIEF DESCRIPTION OF THE DRAWINGS

[0063] Figure 1 Schematic diagram of the false news video detection method of the present invention

[0064] Figure 1A Detailed diagram of a method for detecting fake news videos corresponding to an embodiment of the method of the present invention;

[0065] Figure 2 Flowchart for extracting material and selecting behavioral features for the present invention;

[0066] Figure 2A Detailed flow chart of the process of extracting material and selecting behavioral features for the present invention

[0067] Figure 3 This is a flow chart of the present invention for extracting the characteristics of material editing behavior;

[0068] Figure 3A A detailed flow chart of the present invention for extracting the characteristics of material editing behavior;

[0069] Figure 4 Bidirectional attention module and downsampling network diagram;

[0070] Figure 5 Flow chart of the method of the present invention;

[0071] Figure 6 This is a module diagram of the device of the present invention;

[0072] Figure 7 This is a schematic structural diagram of a first electronic device of the present invention;

[0073] Figure 8 This is a schematic diagram of the application environment structure of the first electronic device of the present invention;

[0074] Figure 9 This is a schematic structural diagram of a second electronic device according to the present invention.

[0075] Reference numerals:

[0076] 11- News Video;

[0077] 12-Material selection behavior characteristics;

[0078] 13-Material editing behavior characteristics;

[0079] 14- Fusion features;

[0080] 15-Classification results;

[0081] 100-Fake News Video Detection Implementation Framework;

[0082] 21-Audio;

[0083] 22-Title subtitle;

[0084] 23-Keyframe;

[0085] 24-emotional branch;

[0086] 25-semantic branch;

[0087] 26-Multilayer Perceptron;

[0088] 200- Extract material and select behavioral feature framework;

[0089] 31-Space editing behavior;

[0090] 32-Time editing behavior;

[0091] 33-Multilayer Perceptron;

[0092] 300- Extracting material editing behavior feature framework;

[0093] 41-bidirectional attention module;

[0094] 42-downsampling network;

[0095] 1-Emotional feature extraction step;

[0096] 2-Semantic feature extraction step;

[0097] 3-Spatial feature extraction step;

[0098] 4-temporal feature extraction step;

[0099] 5-Feature fusion detection step;

[0100] M1-emotional feature extraction module;

[0101] M2-semantic feature extraction module;

[0102] M3-spatial feature extraction module;

[0103] M4-temporal feature extraction module;

[0104] M5-feature fusion detection module;

[0105] A-First electronic device;

[0106] B-Fake news video detection device based on creative process perspective guidance;

[0107] C-data acquisition equipment;

[0108] D-information display device;

[0109] 1000- second electronic device;

[0110] Ⅰ-computing unit;

[0111] II-ROM;

[0112] III-RAM;

[0113] IV-bus;

[0114] V-interface;

[0115] VI-input unit;

[0116] VII-output unit;

[0117] VIII-Storage medium;

[0118] IX-Communication unit. DETAILED DESCRIPTION

[0119] It should be noted that the processor described in the present invention is the control center of an electronic device and can be a single processor or a collective term for multiple processing elements. For example, it can be one or more central processing units (CPUs), application-specific integrated circuits (ASICs), or one or more integrated circuits configured to implement embodiments of the present invention, such as one or more digital signal processors (DSPs) or one or more field programmable gate arrays (FPGAs).

[0120] Optionally, the processor can perform various functions of the electronic device by running or executing a software program stored in the memory, and calling data stored in the memory.

[0121] In a specific implementation, as an example, the processor may include one or more CPUs. Each of these processors may be a single-core processor (single-CPU) or a multi-core processor (multi-CPU). The processor herein may refer to one or more devices, circuits, and / or processing cores for processing data (e.g., computer program instructions). Electronic devices may include servers, desktop computers, laptop computers, smartphones, tablet computers, embedded computers, etc., where the embedded computers include vehicles and robots, etc.

[0122] The memory is used to store the software program for executing the solution of the present invention, and the execution is controlled by the processor. The specific implementation method can refer to the above method embodiment and will not be repeated here.

[0123] It should be noted that the structure of the electronic device shown in the drawings of the present invention does not constitute a limitation thereto, and the actual knowledge structure recognition device may include more or fewer components than shown in the drawings, or a combination of certain components, or a different arrangement of components.

[0124] The above embodiments can be implemented in whole or in part via software, hardware (e.g., circuits), firmware, or any other combination thereof. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product comprises one or more computer instructions or computer programs. When loaded or executed on a computer, the processes or functions described in accordance with the embodiments of the present invention are fully or partially performed. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired means (e.g., infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium accessible by a computer or a data storage device such as a server or data center that contains a collection of one or more available media. The available medium can be magnetic media (e.g., floppy disks, hard disks, magnetic tapes), optical media (e.g., DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0125] It should also be understood that the term "and / or" in this document simply describes an association between related objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " in this document generally indicates an "or" relationship between the related objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0126] In this disclosure, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, "at least one of a, b, or c" can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0127] It should also be understood that in various embodiments of the present invention, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0128] In the several embodiments provided by the present invention, it should be understood that the disclosed devices, apparatuses, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of the device or unit, which can be electrical, mechanical or other forms.

[0129] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0130] In addition, each functional unit in each embodiment of the present invention may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0131] If the functions are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention. The aforementioned storage media include various media that can store program code, such as USB flash drives, mobile hard drives, read-only memories (ROM), random access memories (RAM), magnetic disks, or optical disks.

[0132] To improve the performance of existing video fake news detection methods, this paper proposes a fake news video detection method guided by the creative process perspective. This method goes beyond simply analyzing the superficial patterns of fake news video content and instead analyzes and speculates on the creation process of fake news videos. Because fake news creators on short video platforms often lack firsthand, authentic news materials and professional news production skills, and often deliberately create fake news for personal gain, the creation processes of real and fake news videos differ, resulting in different characteristics in the final news videos. These unique traces can serve as important clues for fake news detection.

[0133] To illustrate the above-mentioned features and effects of the present invention more clearly and easily, the following embodiments are specifically described below with reference to the accompanying drawings. This specification discloses one or more embodiments incorporating the features of the present invention. The disclosed embodiments are for illustrative purposes only. The scope of protection of the present invention is not limited to the disclosed embodiments; the present invention is defined by the appended claims.

[0134] The goal of this invention is to input a video sample data (containing only the video and the title text) and output a true / false binary classification judgment of the video. The overall framework of the fake news video detection method guided by the creative process perspective is as follows Figure 1 、 Figure 1A 、 Figure 2 、 Figure 2A 、 Figure 3 and Figure 3A The following is a detailed introduction to each module of this method.

[0135] 1. Feature Modeling Inspired by Material Selection Behavior:

[0136] Model the creator's material selection behavior from both emotional and semantic perspectives.

[0137] In terms of emotion, the audio and text content of the video are selected as the analysis objects, and the audio emotion features are extracted respectively using the audio pre-training model (such as HuBERT) and the text pre-training model (such as XML-RoBERTA) fine-tuned on the emotion dataset. and text sentiment features The emotional features from the two modalities are concatenated and fed into the Transformer model for fusion to obtain multimodal emotional features. .

[0138] In terms of semantics, text and visual content are selected as analysis objects. The visual content input is multiple frames of images sampled at equal intervals. A pre-trained model (such as CLIP) is used to extract text semantic features at the character / frame level. and visual semantic features , using a cross-modal Transformer model based on a collaborative attention mechanism to interact text and visual semantic features to obtain visually enhanced text features and text-enhanced visual features , the text semantic features and visual features after mean pooling are spliced ​​and sent to the Transformer model based on the self-attention mechanism for fusion to obtain multimodal semantic features .

[0139] The fake news prediction results inspired by material selection behavior are obtained by splicing multimodal sentiment features and semantic features into a multi-layer perceptron:

[0140]

[0141] 2. Feature Modeling Inspired by Material Editing Behavior:

[0142] This method models the creator's material selection behavior from both spatial and temporal aspects.

[0143] In terms of space, this paper mainly models the visual features of superimposed text. First, according to the size of the text area, the rich text visual frame (the frame with the largest text area) is selected as the input of this part. . Using text detector from Locate the text area and get several text boxes , use the prompt encoder to encode these text boxes to obtain prompt features , this encoding can guide the model to focus on the text area of ​​the image. On the other hand, first use the pre-trained model (such as ViT) to Encode the image features of the rich text visual frame ,Will and Input the bidirectional attention module to interact and use guide Focus the text area and get .Will The final feature obtained after downsampling is used as the spatial editing behavior feature The specific design of the bidirectional attention module and downsampling network is as follows: Figure 4 As shown in the figure, self-attention aggregates information on the prompt representation to obtain prompt features that reflect the spatial location information of the text region on the image. The prompt-image attention network calculates the attention of the text region prompt features to the image representation to obtain the interactive text region prompt representation vector. The multi-layer perceptron updates the text region prompt representation to obtain the updated text region prompt representation. The image-prompt attention network calculates the attention of the image representation to the text region prompt representation to obtain the image representation guided by the text region prompt.

[0144] The image representation guided by the text area prompt is passed through a downsampling network (the network structure is convolutional network, layer normalization, activation layer, convolution layer, activation layer, as shown in the figure) to perform feature dimensionality reduction to obtain the final spatial editing behavior features.

[0145] In terms of time, this method mainly models the splicing behavior of the time sequence. The input of this part includes: text time sequence (n is the number of text segments, For text snippets, is the time interval of the corresponding segment) and the visual timing sequence (m is the number of visual segments, is the key frame of the visual segment, is the time interval of the corresponding segment). The format is , which are the start frame index and end frame index of the segment. The video is cut into multiple segments based on the visual picture. The visual segment is a single segment of the visually cut segment.

[0146] This part designs a hierarchical temporal structure extractor that can be applied to multiple modalities. The module takes the temporal sequence of materials of different modalities as input, first performs intra-segment fusion on each material segment, and obtains the content features of the segment. Specifically, for a text temporal sequence, the content features of the text segments are obtained by splicing multiple text segments in the same time period into the encoder. (i represents the i-th text or visual segment). For visual temporal sequences, the content features of the visual segment are obtained by fusing the corresponding multiple frames in the same time period using self-attention. (i represents the i-th clip). In order to model the subtle effects of exposure duration and temporal position of different material clips, the content features of the clips are On this basis, position coding is added and duration coding Positional encoding The encoding strategy is

[0147]

[0148] in , i is the temporal relative position of the fragment, The position encoding vector of the fragment dimension, k range value The position encoding is an array of dim_PE numbers, where each number represents a dimension. The encoding consists of two parts: relative duration encoding and absolute duration encoding. The relative duration and absolute duration are calculated as:

[0149]

[0150] Where fps is the video frame rate and vframes is the total number of video frames. The present invention uses equal frequency bucketing to encode the duration, discretizing the continuous duration into different buckets and encoding the buckets as feature vectors. When encoding a specific duration, the duration data is first assigned to the corresponding bucket, and then the encoding features of the corresponding bucket are used as the encoding vector of the duration. Duration encoding of the i-th segment Expressed as:

[0151]

[0152] Group refers to finding the bucket (duration interval) corresponding to a video's duration data. Emb refers to encoding, which converts the duration information (i.e., the duration interval) into a feature vector.

[0153] Add position encoding and duration coding After that, the fragment feature is expressed as:

[0154]

[0155] By interacting multiple segments using the self-attention mechanism, we can obtain the temporal structure characteristics of a specific modality:

[0156]

[0157] SA refers to Self Attention Network, and MEAN refers to Mean Pooling.

[0158] The hierarchical temporal structure extractor is used to process the text sequence input and the visual sequence input respectively to obtain the text temporal editing behavior characteristics. and visual time editing behavior characteristics . Using Transformer fusion and Get the final multimodal time editing behavior characteristics The fake news prediction results inspired by material editing are obtained by splicing multimodal spatial editing behavior features and temporal editing behavior features into a multi-layer perceptron:

[0159]

[0160] 3. Creation Process-Guided Dual-Perspective Fake News Video Detection

[0161] The creative process-inspired dual-view fake news video detection method uses a post-fusion approach to integrate the fake news prediction results inspired by material selection and the fake news prediction results inspired by material editing to form the final fake news prediction judgment:

[0162]

[0163] in is the fusion function.

[0164] During model training, this method updates the model parameters through multi-branch weighted cross entropy loss:

[0165]

[0166]

[0167] Among them, Y is the actual label of the video, To predict the label, and are weight hyperparameters, and is the cross entropy of the prediction result of a single branch, is the cross entropy of the dual-branch comprehensive prediction results.

[0168] The following is a system embodiment corresponding to the above method embodiment. This embodiment can be implemented in conjunction with the above embodiment. The relevant technical details mentioned in the above embodiment are still valid in this embodiment and will not be repeated here to reduce repetition. Accordingly, the relevant technical details mentioned in this embodiment can also be applied to the above embodiment.

[0169] The present invention also proposes a fake news video detection device B based on the creative process perspective guidance, such as Figure 6 shown, including:

[0170] The emotional feature extraction module M1 obtains the video to be detected as fake news, extracts the audio emotional features of the video and the text emotional features of the video title and subtitle, and fuses the audio emotional features and the text emotional features to obtain a multimodal emotional feature;

[0171] Semantic feature extraction module M2 samples the key frames of the video at intervals, extracts the visual semantic features of the key frames, and extracts the textual semantic features of the subtitles. It uses a cross-modal Transformer model to interact with the visual semantic features and the textual semantic features to obtain visually enhanced text features and text-enhanced visual features. The two are then concatenated and fed into the Transformer model for fusion to obtain multimodal semantic features.

[0172] The spatial feature extraction module M3 selects rich text visual frames in the video based on the size of the text area in the video image; selects the text area from the rich text visual frame to obtain a text box, and encodes the text box to obtain a prompt feature; the prompt feature and the image feature of the rich text visual frame are input into the bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain the spatial editing behavior feature;

[0173] The temporal feature extraction module M4 extracts the text temporal sequence of the video based on the number, content, and time interval of the text segments in the video; extracts the visual temporal sequence of the video based on the number, key frames, and time interval of the visual segments in the video; extracts the features of each segment in the text temporal sequence and the visual temporal sequence respectively to obtain text segment features and visual segment features, adds position coding and duration coding to the text segment features and the visual segment features respectively, and then fuses them to obtain temporal editing behavior features;

[0174] The feature fusion detection module M5 splices the multimodal emotional feature and the multimodal semantic feature into a multi-layer perceptron to obtain a selection behavior feature; splices the spatial editing behavior feature and the temporal editing behavior feature into a multi-layer perceptron to obtain an editing behavior feature; and fuses the selection behavior feature and the editing behavior feature into a binary classification model to obtain a false news video detection result for the video.

[0175] The fake news video detection device based on the creative process perspective guidance shown in the figure, wherein the emotional feature extraction module includes:

[0176] In terms of emotion, an audio pre-training model and a text pre-training model are used to extract the audio emotion features and the text emotion features respectively. The audio emotion features and the text emotion features are concatenated and sent to the Transformer model for fusion to obtain the multimodal emotion features.

[0177] The fake news video detection device based on the creative process perspective guidance shown in the figure, wherein the spatial feature extraction module includes:

[0178] A text detector is used to locate the text area from the rich-text visual frame to obtain a plurality of text boxes, which are then encoded using a prompt encoder to obtain the prompt feature. An image classification model is used to encode the rich-text visual frame to obtain the image feature, which is then input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a guide feature, which is then downsampled to obtain the spatial editing behavior feature.

[0179] The fake news video detection device based on the creative process perspective guidance shown in the figure, wherein the time feature extraction module includes:

[0180] The text time sequence , n is the number of text segments, For the text fragment content, is the time interval corresponding to the fragment;

[0181] The visual timing sequence , m is the number of visual segments, is the key frame of the visual segment, is the time interval corresponding to the fragment; , are the start frame index and end frame index of the segment;

[0182] The visual time sequence and the text time sequence are input into the hierarchical temporal structure extractor as the material time sequence. For the text time sequence, the text segment features are obtained by splicing multiple text segments in the same time period and inputting them into the encoder. For the visual time sequence, the visual segment features are obtained by fusing multiple frames corresponding to the same time period using self-attention.

[0183] The position code The encoding strategy is:

[0184]

[0185] in , i is the temporal relative position of the fragment, The position encoding vector of the fragment dimension, k range value ;

[0186] The duration code The encoding includes relative duration encoding and absolute duration encoding; the relative duration and absolute duration are:

[0187]

[0188] Where fps is the frame rate of the video, and vframes is the total number of frames of the video;

[0189] Duration encoding of the i-th segment Expressed as:

[0190]

[0191] Add position encoding and duration coding Then, we get the fragment features ;

[0192] Multiple segments are interacted using the self-attention mechanism to obtain the temporal structure features of a specific modality:

[0193]

[0194] The hierarchical temporal structure extractor is used to process the text temporal sequence input and the visual temporal sequence input respectively to obtain the text temporal editing behavior features. and visual time editing behavior characteristics ; Using Transformer fusion and Get the editing behavior characteristics of this time .

[0195] like Figure 7 As shown, the present invention further proposes, in another embodiment, a first electronic device A including the aforementioned fake news video detection device B based on creation process perspective guidance.

[0196] like Figure 8 As shown, the first electronic device A can also be connected to the data acquisition device C and the information display device D through a wired or wireless information transmission scheme. The data acquisition device C is used to collect and obtain videos to be identified and classified, such as the news video described in the embodiment of the present invention, and the information display device D is used to display the video classification results obtained by the analysis of the present invention.

[0197] The information display device D can organize and process the data output by the first electronic device A based on the information display mechanism to improve the readability of the data output by the first electronic device A. The information display mechanism can be manually preset, for example, the data output by the first electronic device A is visually displayed, which can be based on the display parameters and / or attributes set by the user. The display parameters can be, for example, the display data range, and the display attributes can be, for example, the display font, color, whether to scroll, etc. The user is presented with the key information specified by the user, and the user can understand this information more promptly without having to access the secondary page or scroll the page, saving the user's operation. Or the information display mechanism can be an artificial intelligence AI display model, which can learn the user's key information based on the user's previous usage habits, such as viewing time, number of clicks, number of edits, etc., and then automatically present the user with rich and necessary key information.

[0198] In another embodiment, the present invention further proposes a storage medium VIII for storing a computer program for executing the method for detecting fake news videos based on the creative process perspective. It should be understood that the storage medium in the embodiments of the present invention may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM), which serves as an external cache. By way of example and not limitation, many forms of random access memory (RAM) are available, such as static RAM (SRAM), dynamic random access memory (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDR SDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), and direct RAM bus RAM (DR RAM).

[0199] Figure 9 A schematic block diagram of a second electronic device 1000 that can be used to implement an embodiment of the present invention is shown. The second electronic device 1000 electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The second electronic device 1000 can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or required herein. The second electronic device 1000 may be the same as or different from the first electronic device A.

[0200] Second electronic device 1000 includes a computing unit I, which can perform various appropriate actions and processes based on a computer program stored in a read-only memory (ROM) II or loaded from a storage medium VIII into a random access memory (RAM) III. RAM III may also store various programs and data required for the operation of device 1000. Computing unit I, ROM II, and RAM III are interconnected via a bus IV. An input / output (I / O) interface V is also connected to bus IV.

[0201] Multiple components in the second electronic device 1000 are connected to the I / O interface V, including: an input unit VI, such as a keyboard and mouse; an output unit VII, such as various types of displays and speakers; a storage medium VIII, such as a magnetic disk and optical disk; and a communication unit IX, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit IX allows the second electronic device 1000 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0202] Computing unit I can be various general-purpose and / or specialized processing components with processing and computing capabilities. Some examples of computing unit I include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. Computing unit I performs the various methods and processes described above, such as method steps 1-5. For example, in some embodiments, the method can be implemented as a computer software program tangibly embodied on a machine-readable medium, such as storage medium VIII. In some embodiments, part or all of the computer program can be loaded and / or installed onto device 1000 via ROM II and / or communication unit IX. When the computer program is loaded into RAM III and executed by computing unit I, one or more steps of the method described above can be performed. Alternatively, in other embodiments, computing unit I can be configured to perform the method by any other suitable means (e.g., via firmware).

[0203] In summary, the present invention uses feature modeling inspired by material selection behavior. The material selection behavior of video creation is analyzed from the perspectives of emotion and semantics, and different modal combinations are targeted to extract features from specific angles. A hierarchical fusion mechanism is adopted to first fuse the multimodal features under a single perspective, and then fuse the multi-perspective features. The news authenticity prediction results inspired by the material selection behavior are obtained. The present invention uses feature modeling inspired by material editing behavior. The material editing behavior of video creation is analyzed from the perspectives of spatial editing behavior and temporal editing behavior, including visual feature modeling of superimposed text layers and multimodal hierarchical temporal structure extraction. The news video authenticity prediction results inspired by material editing behavior are obtained. The present invention uses a dual-perspective fake news video detection method guided by the creative process. The analysis results of the two creative processes of material selection and material editing are integrated by post-fusion, and the cross-entropy weighted sum of the single-branch prediction results and the dual-branch comprehensive prediction results is used as the loss function to adjust the overall training optimization of the model. The news video authenticity prediction results from the two perspectives of comprehensive material selection and material editing are obtained.

[0204] Although the embodiments of the present invention have been disclosed above, they are not limited to the applications listed in the description and implementation methods. They can be fully applied to various fields suitable for the present invention. For those familiar with the art, additional modifications can be easily implemented. Therefore, without departing from the general concept defined by the claims and the scope of equivalents, the present invention is not limited to the specific details and illustrations shown and described herein.

Claims

1. A fake news video detection method based on the creative process perspective, characterized by: include: The emotional feature extraction step obtains a video to be detected as fake news, extracts audio emotional features of the video and text emotional features of the video title and subtitle, and fuses the audio emotional features and the text emotional features to obtain a multimodal emotional feature; The semantic feature extraction step involves sampling the key frames of the video at intervals, extracting the visual semantic features of the key frames, and extracting the textual semantic features of the subtitles. The cross-modal Transformer model is used to interact with the visual semantic features and the textual semantic features to obtain visually enhanced text features and text-enhanced visual features. The two are then concatenated and fed into the Transformer model for fusion to obtain multimodal semantic features. The spatial feature extraction step selects the rich text visual frames in the video based on the size of the text area in the video; Select a text area from the rich text visual frame to obtain a text box, and encode the text box to obtain a prompt feature; The prompt feature and the image feature of the rich text visual frame are input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a spatial editing behavior feature; The temporal feature extraction step extracts the text temporal sequence of the video based on the number, content and time interval of the text segments in the video; extracts the visual temporal sequence of the video based on the number, key frames and time interval of the visual segments in the video; extracts the features of each segment in the text temporal sequence and the visual temporal sequence respectively to obtain text segment features and visual segment features, adds position coding and duration coding to the text segment features and the visual segment features respectively, and fuses them to obtain temporal editing behavior features; In the feature fusion detection step, the multimodal emotional feature and the multimodal semantic feature are spliced ​​and input into a multilayer perceptron to obtain a selection behavior feature; The spatial editing behavior feature and the temporal editing behavior feature are spliced ​​and input into a multi-layer perceptron to obtain the editing behavior feature; The selection behavior features and editing behavior features are fused and input into the binary classification model to obtain the fake news video detection result of the video.

2. The fake news video detection method based on creative process perspective guidance according to claim 1 is characterized in that: The emotion feature extraction step includes: In terms of emotion, an audio pre-training model and a text pre-training model are used to extract the audio emotion features and the text emotion features respectively. The audio emotion features and the text emotion features are concatenated and sent to the Transformer model for fusion to obtain the multimodal emotion features.

3. The fake news video detection method based on creative process perspective guidance according to claim 1 is characterized in that: The spatial feature extraction step includes: A text detector is used to locate the text area from the rich-text visual frame to obtain a plurality of text boxes, which are then encoded using a prompt encoder to obtain the prompt feature. An image classification model is used to encode the rich-text visual frame to obtain the image feature, which is then input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a guide feature, which is then downsampled to obtain the spatial editing behavior feature.

4. The method for detecting fake news videos based on the creative process perspective guidance according to claim 1, characterized in that: The temporal feature extraction step includes: The text time sequence , n is the number of text segments, For the text fragment content, is the time interval corresponding to the fragment; The visual timing sequence , m is the number of visual segments, is the key frame of the visual segment, is the time interval corresponding to the fragment; , are the start frame index and end frame index of the segment; The visual time sequence and the text time sequence are input into the hierarchical temporal structure extractor as the material time sequence. For the text time sequence, the text segment features are obtained by splicing multiple text segments in the same time period and inputting them into the encoder. For the visual time sequence, the visual segment features are obtained by fusing multiple frames corresponding to the same time period using self-attention. The position code The encoding strategy is: in , i is the temporal relative position of the fragment, The position encoding vector of the fragment dimension, k range value ; The duration code The encoding includes relative duration encoding and absolute duration encoding; the relative duration and absolute duration are: Where fps is the frame rate of the video, and vframes is the total number of frames of the video; Duration encoding of the i-th segment Expressed as: Add position encoding and duration coding Then, we get the fragment features ; Multiple segments are interacted using the self-attention mechanism to obtain the temporal structure features of a specific modality: The hierarchical temporal structure extractor is used to process the text temporal sequence input and the visual temporal sequence input respectively to obtain the text temporal editing behavior features. and visual time editing behavior characteristics ; Using Transformer fusion and Get the editing behavior characteristics of this time .

5. A fake news video detection device based on the creative process perspective guidance, characterized in that: include: The emotional feature extraction module obtains the video to be detected as fake news, extracts the audio emotional features of the video and the text emotional features of the video title and subtitle, and fuses the audio emotional features and the text emotional features to obtain a multimodal emotional feature; The semantic feature extraction module samples the key frames of the video at intervals, extracts the visual semantic features of the key frames, and extracts the textual semantic features of the subtitles. The cross-modal Transformer model is used to interact with the visual semantic features and the textual semantic features to obtain visually enhanced text features and text-enhanced visual features. The two are then concatenated and fed into the Transformer model for fusion to obtain multimodal semantic features. The spatial feature extraction module selects the rich text visual frames in the video based on the size of the text area in the video; Select a text area from the rich text visual frame to obtain a text box, and encode the text box to obtain a prompt feature; The prompt feature and the image feature of the rich text visual frame are input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a spatial editing behavior feature; The temporal feature extraction module extracts the text temporal sequence of the video based on the number, content and time interval of the text segments in the video; extracts the visual temporal sequence of the video based on the number, key frames and time interval of the visual segments in the video; extracts the features of each segment in the text temporal sequence and the visual temporal sequence respectively to obtain text segment features and visual segment features, adds position coding and duration coding to the text segment features and the visual segment features respectively, and then fuses them to obtain temporal editing behavior features; The feature fusion detection module splices the multimodal emotional features and the multimodal semantic features and inputs them into a multi-layer perceptron to obtain the selection behavior features; splicing the spatial editing behavior feature and the temporal editing behavior feature and inputting them into a multi-layer perceptron to obtain the editing behavior feature; The selection behavior features and editing behavior features are fused and input into the binary classification model to obtain the fake news video detection result of the video.

6. The fake news video detection device based on creative process perspective guidance according to claim 5, characterized in that: The emotion feature extraction module includes: In terms of emotion, an audio pre-training model and a text pre-training model are used to extract the audio emotion features and the text emotion features respectively. The audio emotion features and the text emotion features are concatenated and sent to the Transformer model for fusion to obtain the multimodal emotion features.

7. The fake news video detection device based on creative process perspective guidance according to claim 5 is characterized in that: The spatial feature extraction module includes: A text detector is used to locate the text area from the rich-text visual frame to obtain a plurality of text boxes, which are then encoded using a prompt encoder to obtain the prompt feature. An image classification model is used to encode the rich-text visual frame to obtain the image feature, which is then input into a bidirectional attention module for interaction, and the prompt feature is used to guide the image feature to focus on the text area to obtain a guide feature, which is then downsampled to obtain the spatial editing behavior feature.

8. The fake news video detection device based on creative process perspective guidance according to claim 5, characterized in that: The temporal feature extraction module includes: The text time sequence , n is the number of text segments, For the text fragment content, is the time interval corresponding to the fragment; The visual timing sequence , m is the number of visual segments, is the key frame of the visual segment, is the time interval corresponding to the fragment; , are the start frame index and end frame index of the segment; The visual time sequence and the text time sequence are input into the hierarchical temporal structure extractor as the material time sequence. For the text time sequence, the text segment features are obtained by splicing multiple text segments in the same time period and inputting them into the encoder. For the visual time sequence, the visual segment features are obtained by fusing multiple frames corresponding to the same time period using self-attention. The position code The encoding strategy is: in , i is the temporal relative position of the fragment, The first position encoding vector of the segment dimension, k range value ; The duration code The encoding includes relative duration encoding and absolute duration encoding; the relative duration and absolute duration are: Where fps is the frame rate of the video, and vframes is the total number of frames of the video; Duration encoding of the i-th segment Expressed as: Add position encoding and duration coding Then, we get the fragment features ; Multiple segments are interacted using the self-attention mechanism to obtain the temporal structure features of a specific modality: The hierarchical temporal structure extractor is used to process the text temporal sequence input and the visual temporal sequence input respectively to obtain the text temporal editing behavior features. and visual time editing behavior characteristics ; Using Transformer fusion and Get the editing behavior characteristics of this time .

9. An electronic device, characterized in that: A false news video detection device based on creation process perspective guidance as described in any one of claims 5-8.

10. The electronic device according to claim 9, wherein The electronic device is connected to an information display device, which is used to display the fake news video detection results using display parameters and attributes set by the user or through an artificial intelligence model.

11. A storage medium for storing a computer program for executing the fake news video detection method based on creative process perspective guidance as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Cross-sample false news video detection method and system

    CN116863366A

  • False news detection method based on news transmission process and related device

    CN117194806A