A method and system for detecting a generative deepfake video
By constructing a multi-dimensional interbrain synchronization feature representation and a multi-view interbrain synchronization feature fusion network, the robustness and accuracy problems of existing deepfake video detection methods are solved, and automated and standardized video authenticity detection is achieved.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- ANHUI UNIV
- Filing Date
- 2026-02-05
- Publication Date
- 2026-05-01
AI Technical Summary
Existing deepfake video detection methods rely on volatile algorithmic traces, have poor robustness, and are unstable based on human subjective judgment, making them difficult to apply on a large scale. Existing neural detection models have failed to effectively integrate the spatial topological structure and heterogeneous information of multidimensional EEG signals.
A multi-dimensional interbrain synchronization feature representation was constructed. EEG signals were collected through a two-person hyperscanning experiment. A multi-view interbrain synchronization feature fusion network was used, which combined a graph topology branch guided by prior knowledge and a multi-view feature encoding branch. The interbrain correlation matrix ISC, phase-locked value matrix PLV, and coherence matrix COH were fused. The attention mechanism was used for feature fusion to predict the authenticity of video content.
It improves the robustness and accuracy of deepfake video detection, realizes automated and standardized detection, overcomes the limitations of existing methods, and significantly enhances the generalization ability and accuracy of detection.
Smart Images

Figure CN121682133B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of electroencephalogram (EEG) signal processing technology, and in particular to a generative deepfake video detection method and system. Background Technology
[0002] With the advent of the era of deep learning and big data, advanced generative artificial intelligence is experiencing explosive growth. In particular, the application of diffusion models has driven the evolution of AI-generated content (AIGC) videos, from simple digital splicing and face-swapping to the creation of a large amount of text-to-video (T2V) and image-to-video (I2V) content. Therefore, developing robust detection techniques to combat these emerging deepfake videos has become a key and urgent challenge in the field of artificial intelligence.
[0003] Current mainstream research on deepfake video detection primarily relies on computer vision and deep learning models, such as convolutional neural networks (CNNs) and transformers. The core principle of these methods is to train models to identify subtle digital artifacts or statistical inconsistencies left behind by the generation algorithm during the forgery process, such as differences in the frequency domain or facial reconstruction errors. However, these computer vision-based methods have fundamental limitations. They heavily depend on specific algorithmic artifacts, leading to an endless "arms race" between the detector and the generator. Once the generation technology is updated, the detection model quickly becomes ineffective, exhibiting poor robustness. Furthermore, these deep learning models are essentially "black boxes," and their opaque decision-making processes make their detection logic difficult for humans to understand, significantly limiting the system's reliability. Finally, most existing research focuses excessively on specific forgery types such as "face swapping," while being insufficiently prepared for forgeries generated by I2V models that involve full-scene manipulation.
[0004] The second type of existing technology is experience-based detection based on subjective human judgment. This method does not rely on algorithms, but rather on human observation, experience, and intuition to discover flaws in fake content. The shortcomings of this type of subjective experience-based detection are also obvious. Its judgment results are extremely subjective and unstable, greatly influenced by personal experience, attention, and cognitive biases. This time-consuming and labor-intensive method is also completely unsuitable for the automated, large-scale screening of massive amounts of videos on social media platforms.
[0005] Neuroscience research has found that even if an individual cannot accurately distinguish between real and fake videos at a conscious level, their brain can still keenly detect subtle inconsistencies in deepfake content. In contrast, facial expressions and speech in real videos are highly compatible with the cognitive system developed by humans over long-term evolution. Deepfake videos, even if visually highly realistic, may present (e.g., facial muscle movements, speech waveforms, or physical dynamics) that conflict with the neural representations stored in the brain about the real world (or specific individuals), thus generating measurable, specific electroencephalogram (EEG) signals. While individual EEG studies have confirmed different neural responses to deepfake videos, there is still a significant gap in exploring shared sensory processing. A more advanced technological approach, "hyperelectroencephalography" (EEG), is beginning to be used in this field. This technology focuses not only on the individual but also quantifies the state of neural synchronization between brains by "simultaneously recording" the brain activity of two or more individuals watching the same video (i.e., binary EEG). This neural coupling has been shown to occur in various collaborative tasks, directly measuring the "collective dynamics" among individuals in "shared perceptual experiences." While a recent study validated the feasibility of using interbrain correlation analysis to identify differences in processing real and artificial stimuli, these efforts are technically limited. For example, they are often "limited to a single synchronization indicator" or primarily rely on PLV features. Most importantly, existing neural detection models often neglect the inherent spatial topology of EEG signals and struggle to effectively integrate synchronization indicators with different attributes, such as "interbrain" and "intrabrain" information. Therefore, there is an urgent need to provide a novel generative deepfake video detection method and system to address these issues. Summary of the Invention
[0006] The technical problem to be solved by the present invention is to provide a generative deepfake video detection method and system. By constructing a comprehensive, multi-dimensional brain-to-brain synchronization feature representation and using a specially designed deep learning architecture, the method can simultaneously process "brain-to-brain synchronization mode" and "intra-brain temporal-spectral dynamics" to achieve detection of deepfake videos.
[0007] To solve the above-mentioned technical problems, one technical solution adopted by the present invention is to provide a generative deepfake video detection method, comprising the following steps:
[0008] S1: Design a dual-person hyperscanning experimental paradigm under video stimulation and collect dual-person EEG signals to construct a dual-person EEG database under real and deepfake video stimulation.
[0009] S2: Perform data preprocessing on the acquired raw EEG signals from both individuals;
[0010] S3: Extract the multidimensional neural synchronization matrix of the two-person EEG from the preprocessed data, including the interbrain correlation matrix ISC, the phase-locked value matrix PLV, and the coherence matrix COH;
[0011] S4: Construct a multi-view brain-to-brain synchronous feature fusion network, which includes a graph topology branch guided by prior knowledge and a multi-view feature encoding branch. Each synchronization matrix is input into the multi-view feature encoding branch. A brain network spatial topology graph is constructed based on electrode prior knowledge. The brain network spatial topology graph is input into the graph topology branch guided by prior knowledge. The features output by the two branches are fused by an attention mechanism to predict the authenticity of the video content.
[0012] In a preferred embodiment of the present invention, in step S1, the experimental paradigm adopts a multi-factor within-subjects design, the core variable of which is the authenticity dimension based on real videos and AI-generated videos. At the same time, the video stimuli cover biological attributes of living and non-living things as well as two kinds of interference factors: positive, negative and neutral emotional valence. All video stimuli are presented to each pair of participants in a random order, and each trial follows a watch-judge-rest structure.
[0013] In a preferred embodiment of the present invention, step S2, the process of preprocessing the acquired raw dual-person EEG signals, includes:
[0014] S201: Downsampling of EEG signals from two individuals;
[0015] S202: Apply a 0.1-60 Hz bandpass filter to retain the frequency range of interest, and then apply a 48-52 Hz notch filter to remove power frequency noise;
[0016] S203: The filtered data is subjected to whole-brain average rereference to suppress common-mode noise and improve signal purity;
[0017] S204: Independent component analysis algorithm is used to decompose the EEG signals of two people into independent components for the purpose of separating and removing artifacts;
[0018] S205: Time-frequency decomposition is performed using wavelet transform into five frequency bands, including Delta, Theta, Alpha, Beta, and Gamma, to extract power and phase information that varies with time.
[0019] In a preferred embodiment of the present invention, in step S3, the three indices of interbrain correlation (ISC), interbrain coherence (Coh), and phase-locked value (PLV) of the two-person EEG signals are calculated to generate corresponding feature matrices for the five frequency bands of the original two-person EEG signal decomposition.
[0020] In a preferred embodiment of the present invention, in step S4, the multi-view interbrain synchronous feature fusion network includes a prior knowledge-guided graph topology branch, a multi-view feature encoding branch, and a feature fusion module.
[0021] Furthermore, the prior knowledge-guided graph topology branching is used to deeply mine the spatial topological features of interbrain synchronization, specifically including the following implementation steps:
[0022] S401: Construct a graph structure G=(V, E) containing N nodes, where the node set V corresponds to the N electrode channels of the EEG acquisition device, and E is the set of connected edges after filtering in the inter-subject neural synchronization graph matrix. For the i-th node in the graph structure, directly extract the i-th row or column vector corresponding to the synchronization matrix as the initial feature of that node. This feature physically characterizes the synchronous connectivity pattern between a specific brain region of subject A and the whole brain region of subject B;
[0023] S402: Obtain the 3D physical coordinates of each electrode in the standard MNI brain space, generate coordinate embeddings through nonlinear mapping to provide absolute position information, calculate random walk probabilities based on the connection topology between electrodes to provide relative structural information of nodes in the graph, and combine the two prior codes with the initial features. Merge to generate the node input for layer 0. ;
[0024] S403: Set up an improved graph attention layer of L layers to iteratively update node features, where L > 1. In each layer, the allocation of attention weights is explicitly guided by introducing an electrode prior adjacency matrix. This is repeated L times, so that node features can perform multiple rounds of information interaction and integration on the whole brain topology.
[0025] S404: The obtained set of node features containing deep spatial topological semantics is integrated into a whole-brain topological feature vector through a global aggregation operation. .
[0026] Furthermore, the multi-view feature encoding branch inputs the interbrain correlation matrix ISC, phase-locked value matrix PLV, and coherence matrix COH into three independent improved residual network branches. Each branch independently learns the deep nonlinear features of the corresponding index and outputs feature vectors of interbrain coherence features from three different perspectives. Phase-locked loop characteristics Correlation characteristics between brains The three feature vectors are processed by a feature fusion module based on a multi-view attention mechanism to output a fused feature. The final global feature vector is constructed by feature concatenation. , global feature vector The input is fed into the classifier module, which outputs the predicted probability that the video is "real" or "fake".
[0027] Furthermore, the feature fusion module integrates the interbrain coherence features of feature vectors from three different perspectives. Phase-locked loop characteristics Correlation characteristics between brains The input tensor is concatenated to form the input tensor. The weights of each perspective are calculated using the Query-Key-Value attention paradigm, and then weighted and fused to output a multi-view brain synchronous fusion feature vector. .
[0028] Furthermore, the steps for calculating the weights of each perspective and then weighting and fusing them include:
[0029] First, let's define the three input variables. , , After being concatenated, they form an input tensor with a dimension of 768. Then, a fully connected layer maps the input tensor to a query Q, a key K, and a value V:
[0030]
[0031] in, A learnable linear mapping weight matrix is used to project the input features onto a high-dimensional semantic space; subsequently, a normalized attention weight map A is generated through dot product and activation function.
[0032]
[0033] in, It is the scaling factor in the attention mechanism. The dimension of the feature vector;
[0034] The calculated weight map A is used to weight the value vector V, and the multi-view brain-brain synchronous fusion feature vector is output. :
[0035]
[0036] in, This indicates element-wise multiplication or matrix multiplication.
[0037] To solve the above-mentioned technical problems, another technical solution adopted by the present invention is: to provide a generative deepfake video detection system, comprising:
[0038] The dual-person EEG signal acquisition module is used to design dual-person hyperscanning experimental paradigms under video stimulation and acquire dual-person EEG signals to construct dual-person EEG databases under real and deepfake video stimulation.
[0039] A dual-person EEG signal preprocessing module is used to preprocess the raw dual-person EEG signals acquired by the dual-person EEG signal acquisition module.
[0040] The multidimensional neural synchronization matrix extraction module is used to extract the multidimensional neural synchronization matrix of the two-person EEG from the data preprocessed by the two-person EEG signal preprocessing module, including the interbrain correlation matrix ISC, the phase-locked value matrix PLV, and the coherence matrix COH.
[0041] The feature processing and video forgery result output module is used to construct a multi-view brain synchronous feature fusion network, which includes a graph topology branch guided by prior knowledge and a multi-view feature encoding branch. Each synchronization matrix is input into the multi-view feature encoding branch. A brain network spatial topology graph is constructed based on electrode prior knowledge. The brain network spatial topology graph is input into the graph topology branch guided by prior knowledge. The features output by the two branches are fused by an attention mechanism to predict the authenticity of the video content.
[0042] The beneficial effects of this invention are:
[0043] (1) This invention solves the technical defect of existing computer vision (CV)-based detection methods that rely too much on transient digital artifacts. This invention does not rely on volatile algorithm traces, but is based on stable human neurocognitive mechanisms (i.e., brain-brain neural synchronization). Since there is a fundamental conflict between the forged content and the real-world representation stored in the brain, this method can detect it from a physiological level, thereby improving the robustness and generalization of the detection.
[0044] (2) This invention solves the technical defects of existing methods based on human subjective judgment, which are "extremely subjective and unstable" and "time-consuming and labor-intensive". This invention provides a repeatable and standardized detection scheme by collecting objective and quantifiable binary EEG biological signals and using deep learning models for automated analysis, which completely overcomes the problems of unreliability, low efficiency and inability to be applied on a large scale by human judgment;
[0045] (3) To address the technical problems of existing binary EEG detection methods being "limited to a single synchronization index" and failing to construct multi-dimensional feature vectors, this invention constructs a comprehensive, multi-dimensional representation of inter-brain synchronization features. This representation systematically integrates multiple indicators such as inter-brain correlation (ISC), phase-locked value (PLV), and coherence (COH) to more comprehensively and deeply characterize the inter-brain synchronization change patterns induced by fake videos. Most importantly, to address the technical problems of existing neural detection models often ignoring the inherent spatial topology of EEG signals and the difficulty in effectively integrating synchronization indicators with different attributes (such as "inter-brain" and "intra-brain" information), this invention proposes a multi-view inter-brain synchronization feature fusion network (Hyper-FusionNet). This architecture introduces graph neural networks (GNNs) to capture the spatial dependencies between electrodes and utilizes attention mechanisms to "co-process and fuse" multiple heterogeneous neural features, thereby solving the deficiency of existing simple models in effectively decoding complex brain network dynamics, and ultimately significantly improving the accuracy of deepfake video detection. Attached Figure Description
[0046] Figure 1 This is a flowchart of a generative deepfake video detection method according to the present invention;
[0047] Figure 2 This is a model architecture diagram of the generative deepfake video detection method.
[0048] Figure 3 This is a diagram of the experimental acquisition device for dual-person electroencephalogram (EEG) signals of the present invention;
[0049] Figure 4 This is a schematic diagram of the dual-person hyperscanning experimental paradigm described in this invention;
[0050] Figure 5 This is a flowchart of the preprocessing of the EEG signals from the two individuals;
[0051] Figure 6 This is a diagram showing the results of the brain coherence network analysis of real and fake video stimuli.
[0052] Figure 7 This is a statistical graph showing the results of PLV (Programmable Logic Detection) synchronization of brainwaves under the stimulation of real and fake videos.
[0053] Figure 8 This is a topographical distribution map of the differences in EEG phase synchronization under the stimulation of real and fake videos;
[0054] Figure 9 This is a graph showing the statistical results of the interbrain correlation analysis (ISC).
[0055] Figure 10 This is a structural block diagram of the generative deepfake video detection system. Detailed Implementation
[0056] The preferred embodiments of the present invention will now be described in detail with reference to the accompanying drawings, so that the advantages and features of the present invention can be more easily understood by those skilled in the art, thereby providing a clearer and more explicit definition of the scope of protection of the present invention.
[0057] Please see Figure 1 and Figure 2 The embodiments of the present invention include:
[0058] A generative deepfake video detection method includes the following steps:
[0059] S1: Design a dual-person hyperscanning experimental paradigm under video stimulation and collect dual-person EEG signals to construct a dual-person EEG database under real and deepfake video stimulation.
[0060] The core of this system lies in achieving "electroencephalography (EEG) scanning," which involves simultaneously recording the brain activity of two individuals while they are watching a video. For example... Figure 1 As shown, the system comprises two parallel, independent EEG acquisition systems, such as two "Boruien Brain" wireless dry electrode EEG acquisition systems. Each participant wears an EEG cap equipped with 32-channel dry electrodes, arranged according to the international 10-20 system standard to provide full scalp coverage. During the experiment, the two participants sit comfortably side-by-side and jointly watch video stimulation material presented on a single "monitor." In this example, video playback and the experimental workflow are precisely managed by a stimulation computer using the dedicated experimental presentation software "E-Prime 3." To achieve "precise time alignment" between the two independent EEG data streams—crucial for subsequent interbrain synchronization analysis—the system employs the "Laboratory Flow Lambda" (LSL) protocol. The stimulation computer sends synchronized event markers to the router via the LSL protocol while presenting the video or making judgments. The two EEG acquisition systems also wirelessly transmit their respective acquired EEG data streams to the router in real time via Wi-Fi. Finally, the router aggregates the two EEG data streams and the LSL event marker stream, transmits them together to the EEG data acquisition computer, and is ultimately responsible for analyzing the acquired synchronized data and detecting video forgery.
[0061] The experimental paradigm of this invention employs a multi-factor within-subjects design, such as... Figure 2As shown. To systematically separate the effects of realism, content category, and emotional valence on neural activity, this invention employs a 2 (video realism: real vs. AI-generated) × 2 (content category: living vs. non-living) × 3 (emotional valence: positive, negative, neutral) factor design. To achieve this design, this invention constructs a stimulus material library containing, for example, 36 video clips. The real video materials are selected from publicly available datasets; for example, the "living" real videos (12 clips) are selected from the DeepFakeDetection (DFD) dataset used in Facebook's 2020 DeepFake Detection Competition; the "non-living" real videos (9 clips) are selected from platforms such as YouTube. The corresponding deepfake video materials are created using the "Seedance 1.0" platform generation model. In this invention, each AI-generated video is a visually and semantically coherent copy of its corresponding real video, with careful matching between the two in terms of scene composition, emotional tone, narrative structure, resolution, and size. For example, an AI version of a real video depicting a person smiling will also depict that person smiling. This one-to-one matching strategy is crucial, as it ensures that "realism" is the primary variable distinguishing the two sets of videos, thus guaranteeing that any subsequent observed neural differences can be attributed to the "synthetic nature" of the content, rather than other visual or narrative confounding factors.
[0062] Upon arrival at the lab, participants first underwent a 3-minute resting-state EEG recording, which served as the neural baseline for subsequent analysis. The main experimental task then commenced. This task consisted of a series of trials, with all video stimuli presented to each pair of participants in a randomized order to prevent any learning or sequence effects. Each trial followed a precise and consistent "View-Judge-Rest" structure. In each round, participants first watched a 10-second video clip. Immediately after the video ended, a prompt appeared on the screen, instructing participants to make a judgment (either "real" or "AI-generated") by pressing a button. After the judgment, the screen went black, initiating a 6-second rest period. This rest period allowed participants' cognitive states to return to baseline before the next trial began. The entire experimental procedure lasted approximately 30 minutes. Throughout the experiment, participants were required to remain silent, minimize physical movement, and avoid eye contact, physical contact, or any other interaction with each other.
[0063] S2: Perform data preprocessing on the acquired raw EEG signals from both individuals; combined with... Figure 5 The preprocessing procedure is as follows:
[0064] First, the EEG signals from both individuals were downsampled to a lower sampling rate of 250 Hz to reduce the computational load.
[0065] The data then undergoes a series of filtering processes: first, a 0.1-60 Hz bandpass filter is applied to preserve the frequency range of interest, and then a 48-52 Hz notch filter is applied to specifically remove power frequency noise.
[0066] The filtered data is then subjected to whole-brain average rereference to suppress common-mode noise and improve signal purity.
[0067] To remove artifacts, the system employs Independent Component Analysis (ICA) algorithm, which is used to identify and remove noise components caused by non-brain activities such as eye movements, electromyography, or electrocardiography.
[0068] Finally, to perform frequency domain specificity analysis, wavelet transform was applied to decompose the time-frequency data into five classic frequency bands (i.e., , , , and The power and phase information that changes over time are extracted.
[0069] S3: Extract the multidimensional neural synchronization matrix of the two-person EEG from the preprocessed data, including the interbrain correlation matrix ISC, the phase-locked value matrix PLV, and the coherence matrix COH;
[0070] To overcome the limitations of existing methods that rely solely on a "single synchronization index," this invention constructs a comprehensive, multi-dimensional representation of inter-brain synchronization features. This representation systematically integrates multiple complementary synchronization indices, primarily including coherence (Coh), phase locking value (PLV), and inter-subject correlation (ISC). These three dual-person EEG multidimensional neural synchronization indices are described in detail below:
[0071] (1) Coherence is a linear correlation index measured in the frequency domain that considers both phase and amplitude consistency between two signals. This invention can robustly estimate coherence using, for example, Welch's method. The calculation formula is as follows:
[0072]
[0073] in, It is a specific frequency; It is a signal and The square magnitude of the average cross-power spectral density between them; and and These are signals and At this frequency The average power spectral density on.
[0074] (2) Phase-locked value (PLV) is used to quantify the consistency of the phase difference between two signals at a specific frequency. A key characteristic of this metric is that it is independent of the signal amplitude. Its calculation formula is as follows:
[0075]
[0076] in, This represents the total number of time points; and Two participants are in the passage. Time point and frequency The instantaneous phase on; It is the imaginary unit.
[0077] (3) Interbrain correlation (ISC) identifies shared neural processing by finding spatial patterns that maximize the temporal correlation between EEG signals from two participants. This invention employs a framework based on correlation component analysis (CCA). This process first calculates the cross-covariance matrix. and subject covariance matrix :
[0078]
[0079]
[0080] in, and These represent the EEG data of the first and second participants, respectively. After shrinking and regularizing the within-participant covariance matrix to enhance robustness, we obtain... The system then finds the spatial components that maximize the generalized eigenvalues by solving the following generalized eigenvalue problem:
[0081]
[0082] in, and It is a spatial filter matrix; The corresponding feature value represents the intensity of shared neural activity captured by this spatial pattern, i.e., the ISC value.
[0083] like Figure 6As shown in the figure, this visually illustrates the comparison of brain coherence (Coh) network topology connectivity at the Beta and Gamma high-frequency bands. The colors of the connecting lines in the figure map coherence intensity (Colorbar ranges from 0.2 to 0.43), with blue representing low coherence and red representing high coherence.
[0084] Comparative analysis shows that when watching real videos (Real), Figure 5 Left column: The connections between the two brains are relatively sparse and predominantly cool-toned (blue-green, lower values), indicating that excessive synchronization coupling did not form between brains when processing natural, realistic video streams. When watching fake videos (… Figure 5 (Right column): The situation has been significantly reversed, with both the Beta and Gamma bands exhibiting a widespread state of "hypersynchrony". Numerous dense, strong red connections have emerged between the two hemispheres (connection strength approaching the upper limit of 0.43).
[0085] This visualization strongly demonstrates that although subjects may not be able to consciously distinguish fake videos, deepfake content (possibly due to subtle unnatural artifacts or spatiotemporal inconsistencies) unconsciously triggers a higher intensity of shared neural responses and cognitive load between the two brains, resulting in a significant enhancement of high-frequency coherence. It's possible that subtle physical violations exist in the fake videos, constituting a persistent "prediction error." This error forces the brain into a state of "active scrutiny," requiring the mobilization of more high-frequency cognitive resources to correct the error and integrate information. It is this shared "cognitive effort" that leads to the abnormally high level of synchronous coupling between brains.
[0086] like Figure 7 As shown in the figure, this graph presents the statistical analysis results of phase-locked value synchronization (PLV) between different videos. Figure 7 (a) shows the statistical results of phase synchronization comparison between real and fake videos in neutral videos; (b) shows the statistical results of phase synchronization comparison between real and fake videos in living videos; (c) shows the statistical results of phase synchronization comparison between real and fake videos in inanimate videos; and (d) shows the statistical results of scent synchronization comparison between inanimate and living videos. The PLV analysis reveals a strong "content-dependent" effect. For example, when watching living video content (such as...) Figure 7 In (b) of the above, real videos elicited significantly higher PLVs compared to fake videos, exhibiting a pro-realism characteristic; however, when watching inanimate video content (such as... Figure 7 In (d) of the above, the situation reversed; the AI-generated fake videos induced significantly higher PLVs across all five frequency bands, indicating stronger synchronization of faked interbrain connections. Figure 8As shown, the figure presents the topographic distribution of the differences in phase-locking value (PLV) induced by real video and deepfake video on the scalp at five EEG frequency bands, which also reveals key visualization results of the spatial characteristics of neural activity when the brain processes real and fake visual content.
[0087] like Figure 9 As shown in the figure, this graph presents the results of statistical analysis of interbrain correlations (ISC). Figure 9 (a) shows the results of the interbrain ISC comparison between real and fake videos in neutral video contexts; (b) shows the results of the interbrain ISC comparison between real and fake videos in inanimate video contexts. ISC analysis further confirms this content dependence. When viewing "neutral" content (such as… Figure 9 In (a) of the study, real video consistently elicited significantly higher ISCs across all five frequency bands than fake video, suggesting that natural stimuli drive more reliable shared neural processing. However, similar to the results for PLV, when watching “inanimate” content (such as…),… Figure 9 In (b) of the above, the pattern is reversed, and the AI-generated fake videos induce significantly higher ISC in the Delta and Gamma bands.
[0088] Specifically, this invention calculates the three indices (ISC, PLV, COH) mentioned above, generating corresponding feature matrices for each of the five frequency bands. Unlike existing technologies that simply stack these matrices into a high-dimensional tensor, this invention treats ISC, PLV, and COH as three different "views," preserving their independent matrix forms (e.g., 32×32×32 matrices) so that they can be input into different branches of the neural network for specialized feature encoding, avoiding feature confusion caused by early forced stacking.
[0089] S4: Construct a multi-view brain-to-brain synchronous feature fusion network, which includes a graph topology branch guided by prior knowledge and a multi-view feature encoding branch. Each synchronization matrix is input into the multi-view feature encoding branch. A brain network spatial topology graph is constructed based on electrode prior knowledge. The brain network spatial topology graph is input into the graph topology branch guided by prior knowledge. The features output by the two branches are fused by an attention mechanism to predict the authenticity of the video content.
[0090] The multi-view interbrain synchronous feature fusion network includes a prior knowledge-guided graph topology branch, a multi-view feature encoding branch, and a feature fusion module. The following details each branch and module:
[0091] (1) Graph topology branching guided by prior knowledge:
[0092] Electroencephalogram (EEG) signals not only exhibit numerical synchronization but also possess clear anatomical spatial relationships. For example... Figure 9 As shown in (c), this invention constructs a graph neural network branch based on electrode location, and deeply mines the spatial topological features of interbrain synchronization by stacking L layers of prior-guided graph attention modules. The specific implementation steps are as follows:
[0093] (1.1) Construction and Feature Initialization of Graph Structure: First, a graph structure G=(V, E) containing N nodes is constructed, where the node set V corresponds to the N electrode channels of the EEG acquisition device (N=32 in this embodiment), and E is the set of connected edges after filtering in the inter-subject neural synchronization graph matrix. Feature Definition: In order to capture the binary inter-brain interaction features, this invention example selects the inter-brain synchronization matrix (such as the ISC matrix) calculated by the preceding module as the input of the graph nodes. Specifically, for the i-th node in the graph, the vector of the i-th row (or column) corresponding to the synchronization matrix is directly extracted as the initial feature of that node. This feature physically characterizes the synchronous connectivity pattern between a specific brain region of subject A and the whole brain region of subject B.
[0094] (1.2) Injection of Multidimensional Position Priors: To enable the neural network to perceive the physical distribution and structural role of electrodes, this invention injects two types of prior knowledge into the initial features: Physical coordinate prior: Obtaining the three-dimensional physical coordinates of each electrode in the standard MNI brain space, generating coordinate embeddings through nonlinear mapping to provide absolute position information; Structural topology prior: Calculating the random walk positional encoding (RWPE) based on the connection topology between electrodes to provide relative structural information of nodes in the graph. The above two prior encodings are integrated with the initial features. Merge to generate the node input for layer 0. .
[0095] (1.3) Prior-guided graph attention computation: This is the core computational step of this branch. The model sets up L layers (L>1) of improved graph Transformer layers to iteratively update node features. In each layer, the allocation of attention weights is explicitly guided by introducing an electrode prior adjacency matrix. Specifically, for the i-th node in the l-th layer, its feature update process includes the following mathematical operations:
[0096] (1.3.1) Linear Projection: First, the input features are mapped to the query, key, and value spaces:
[0097]
[0098] in, Let represent the input feature vector of the i-th electrode node in the L-th improved graph Transformer layer; specifically, This represents the initial input feature vector of the i-th electrode node. It is formed by fusing the i-th row vector in the synchronization matrix with the position prior code (PE). The position prior code (PE) includes an absolute position embedding generated based on the electrode MNI coordinates (x, y, z) and a random walk position code (RWPE) calculated based on the inter-electrode connection topology, which is used to provide relative structural information of the node. These are the query weight matrix, key weight matrix, and value weight matrix, which are randomly initialized and assigned initial values in the graph attention layer, and are continuously learned and optimized during training. Let i represent the query vector of node i, the key vector of neighbor node j, and the information value vector of neighbor node j, respectively.
[0099] (1.3.2) Prior-guided attention scoring: Calculate the attention coefficient between node i and its neighbor node j. This section introduces the core innovation—prior bias. This value is based on a pre-calculated physical Euclidean distance between electrodes i and j on the scalp surface, and the value is inversely proportional to the physical distance:
[0100]
[0101] in, As a sparse normalization function, compared to the traditional Softmax, it can eliminate noisy connections and retain the most significant spatial topological features of brain networks. It is the scaling factor in the attention mechanism. The dimension of the feature vector is used to balance the numerical range of the dot product operation. When aggregating information, not only data correlation is considered (…). ), and must also comply with anatomical distance constraints ( ).
[0102] (1.3.3) Feature Aggregation and Update: Based on the above coefficients, the neighborhood information is weighted and aggregated to obtain the output of this layer:
[0103]
[0104] in, This represents the updated node features after the l-th layer graph attention operation, obtained by weighted summation of neighboring node information. : represents the set of neighboring nodes of node i, where v is the set of all nodes in the graph. This process is repeated L times, allowing node features to undergo multiple rounds of information interaction and integration across the whole-brain topology.
[0105] (1.4) Whole-brain topological feature generation: After processing by the L-layer prior guided graph attention layer, the model obtains a set of node features containing deep spatial topological semantics. Finally, through global aggregation operations (such as linear mapping and concatenation), the features of N nodes are integrated into a whole-brain topological feature vector. The vector is then fed into the feature fusion module, where it is fused with the output of the multi-view feature encoding branch to jointly participate in the detection of fake videos.
[0106] (2) Multi-view feature encoding branch:
[0107] To address the three heterogeneous synchronization metrics—inter-brain correlation (ISC), phase-locked value (PLV), and coherence (COH)—this invention designs three parallel feature extraction sub-networks. For example... Figure 9 As shown in (d), each subnetwork employs an improved residual network (e.g., the improved version). As a backbone network, it addresses the low resolution of the interbrain synchronization matrix. Due to the characteristics of [the network], the improved residual network removes the first layer. Large convolutional kernels and pooling layers, replaced with The small convolution kernel was used, and the downsampling step size of subsequent residual blocks was adjusted.
[0108] The SC, PLV, and COH matrices are respectively input into these three independent improved matrices. Within each branch, each branch independently learns the deep nonlinear features of the corresponding index, and the output feature vectors are denoted as follows: This "independent encoding" strategy effectively preserves the unique statistical distribution characteristics of each synchronization index, avoiding feature interference caused by direct stacking in traditional methods.
[0109] (3) Feature fusion module:
[0110] To fully integrate interbrain synchronization features from different frequency and phase domains, as well as the topological spatial features extracted in the previous stage, this invention designs a decision module that includes "multi-view attention fusion" and "dual-stream feature stitching." The specific implementation steps are as follows:
[0111] (3.1) The attention-based multi-view feature fusion pre-sequence multi-view feature encoding branch (ResNet branch) outputs feature vectors from three different perspectives: interbrain coherence features Phase-locked loop characteristics Correlation characteristics between brains To adaptively integrate these features, this invention introduces a multi-view attention mechanism, as shown in module (b) of the figure. This mechanism aims to dynamically calculate the importance weights of different synchronization metrics. The specific calculation process is as follows:
[0112] (3.1.1) Feature stacking: Stack or concatenate the feature vectors from the three perspectives to form the input tensor. .
[0113] (3.1.2) Attention Weight Calculation: The weights of each perspective are calculated using the Query-Key-Value attention paradigm. First, the three input variables... , , After being spliced (C), a tensor with dimension (768) is formed. Then, a fully connected layer maps the input tensor to a query Q, a key K, and a value V:
[0114]
[0115] in, The weight matrix is a learnable linear mapping used to project input features into a high-dimensional semantic space. Before model training begins, the weight matrix is randomly initialized and then iteratively optimized during training using the backpropagation algorithm.
[0116] Subsequently, a normalized attention weight map A is generated using dot product and activation function:
[0117]
[0118] in, It is the scaling factor in the attention mechanism. In this embodiment, the dimension of the feature vector is introduced. The scaling factor is used to prevent the dot product result from becoming too large, which could lead to gradient vanishing and thus maintain training stability. Note this. With (1.3.2) The values are the same.
[0119] The Sigmoid activation function is used here to capture the non-linear interactions between features from different perspectives.
[0120] (3.1.3) Weighted fusion: The value vector V is weighted using the calculated weight map A, and the fused multi-view brain-brain synchronous fusion feature vector is output. :
[0121]
[0122] in, This indicates element-wise multiplication or matrix multiplication, depending on the specific implementation. This step ensures that the model can automatically focus on the most discriminative synchronization metrics based on the characteristics of the current sample (e.g., automatically assigning higher weights to PLV in certain forgery scenarios).
[0123] (3.2) Global concatenation of dual-stream features: This model contains two parallel information streams: data-driven stream and the fused features obtained in the above steps. This mainly characterizes the numerical intensity, texture, and statistical distribution of interbrain synchronization; prior driving flow: features output by the topological branch in step (1.4). This primarily characterizes the anatomical spatial distribution and topological structure of brain synchronization. To combine these two complementary types of information, this invention employs a feature concatenation operation to construct the final global feature vector:
[0124]
[0125] This design enables the model to possess both the ability of deep learning to perceive weak signals and the ability to spatially interpret prior knowledge from neuroscience.
[0126] (3.3) Forgery detection and classification: Finally, the global feature vector is... The input is fed into the classifier module. The classifier consists of a multilayer perceptron (MLP), which typically includes fully connected layers, non-linear activation functions (such as ReLU), and dropout layers (to prevent overfitting).
[0127]
[0128] The output layer uses the Softmax or Sigmoid function to output the predicted probability that the video is "real" or "fake".
[0129] Based on rigorous subject cross-validation (15-fold LOPO-CV), this invention has achieved a significant breakthrough in detection performance:
[0130] (1) Significantly improved accuracy: As shown in Table 1, the method of this invention (Ours - Hyper-FusionNet) achieved an average accuracy of 89.23% and an F1 score of 89.04%, which is significantly better than the traditional SVM (75.78%) and CNN (80.16%) methods, and also better than the standard ResNet-18 baseline (82.43%).
[0131] (2) The necessity of multi-branch and graph networks: Comparative experiments show that if the graph neural network branches are removed (Ours has no GNN branches), the accuracy rate is 87.05%; while adding graph branches improves it to 89.23%. This directly proves that introducing EEG spatial topology information is crucial for identifying subtle neural differences induced by deepfakes.
[0132] Table 1: Performance of different models in deepfake video classification under various combinations of EEG synchronization features.
[0133]
[0134] (3) Strong robustness: The present invention reports a 95% confidence interval [84.54, 93.92], indicating that the model maintains a high degree of consistent detection performance among different subjects and has a strong generalization ability.
[0135] In summary, this invention, through a specially designed Hyper-FusionNet, achieves deep decoding of brain synchronization signals in the "time-frequency-space-multimodal" dimensions, providing a high-precision and robust neural computing solution for deepfake video detection.
[0136] See Figure 10 This invention also provides a generative deepfake video detection system, including a dual-person EEG signal acquisition module, a dual-person EEG signal preprocessing module, a multi-dimensional neural synchronous feature extraction module, and a feature processing and video forgery result output module.
[0137] The dual-person EEG signal acquisition module is used to design a dual-person hyperscanning experimental paradigm under video stimulation and acquire dual-person EEG signals to construct a dual-person EEG database under real and deepfake video stimulation.
[0138] The dual-person EEG signal preprocessing module is used to preprocess the raw dual-person EEG signals acquired by the dual-person EEG signal acquisition module.
[0139] The multidimensional neural synchronization matrix extraction module is used to extract the multidimensional neural synchronization matrix of the two-person EEG from the data preprocessed by the two-person EEG signal preprocessing module, including the interbrain correlation matrix ISC, the phase-locked value matrix PLV, and the coherence matrix COH.
[0140] The feature processing and video forgery result output module is used to construct a multi-view brain synchronous feature fusion network, which includes a graph topology branch guided by prior knowledge and a multi-view feature encoding branch. Each synchronization matrix is input into the multi-view feature encoding branch. A brain network spatial topology graph is constructed based on electrode prior knowledge. The brain network spatial topology graph is input into the graph topology branch guided by prior knowledge. The features output by the two branches are fused by an attention mechanism to predict the authenticity of the video content.
[0141] This example provides a generative deepfake video detection system that can execute the generative deepfake video detection method provided by this invention. It can perform any combination of the steps of the method example and has the corresponding functions and beneficial effects of the method.
[0142] The above description is merely an embodiment of the present invention and does not limit the patent scope of the present invention. Any equivalent structural or procedural transformations made based on the content of the present invention's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of the present invention.
Claims
1. A generative deepfake video detection method, characterized in that, Includes the following steps: S1: Design a dual-person hyperscanning experimental paradigm under video stimulation and collect dual-person EEG signals to construct a dual-person EEG database under real and deepfake video stimulation. S2: Perform data preprocessing on the acquired raw EEG signals from both individuals; S3: Extract the multidimensional neural synchronization matrix of the two-person EEG from the preprocessed data, including the interbrain correlation matrix ISC, the phase-locked value matrix PLV, and the coherence matrix COH; S4: Construct a multi-view interbrain synchronization feature fusion network, which includes a prior knowledge-guided graph topology branch and a multi-view feature encoding branch. Each synchronization matrix is input into the multi-view feature encoding branch. A brain network spatial topology graph is constructed based on electrode prior knowledge. This brain network spatial topology graph is input into the prior knowledge-guided graph topology branch. The features output from the prior knowledge-guided graph topology branch and the multi-view feature encoding branch are fused and concatenated to predict the authenticity of the video content. The prior knowledge-guided graph topology branch is used for deep mining of the spatial topological features of interbrain synchronization, specifically including the following implementation steps: S401: Construct a graph structure G=(V, E) containing N nodes, where the node set V corresponds to the N electrode channels of the EEG acquisition device, and E is the set of connected edges after filtering in the inter-subject neural synchronization graph matrix. For the i-th node in the graph structure, directly extract the i-th row or column vector corresponding to the synchronization matrix as the initial feature of that node. This initial feature It physically characterizes the synchronous connectivity patterns between a specific brain region of subject A and the whole brain region of subject B; S402: Obtain the 3D physical coordinates of each electrode in the standard MNI brain space, generate coordinate embeddings through nonlinear mapping to provide absolute position information, calculate random walk probabilities based on the connection topology between electrodes to provide relative structural information of nodes in the graph, and combine the two prior codes with the initial features. Merge to generate the node input for layer 0. ; S403: Set up an improved graph attention layer of L layers to iteratively update node features, where L > 1. In each layer, the allocation of attention weights is explicitly guided by introducing an electrode prior adjacency matrix. This is repeated L times, so that node features can perform multiple rounds of information interaction and integration on the whole brain topology. S404: The obtained set of node features containing deep spatial topological semantics is integrated into a whole-brain topological feature vector through a global aggregation operation. .
2. The generative deepfake video detection method according to claim 1, characterized in that, In step S1, the experimental paradigm adopts a multi-factor within-subjects design, with the core variable being the authenticity dimension based on real videos and AI-generated videos. Meanwhile, the video stimuli cover biological attributes of living and non-living things, as well as two interfering factors: positive, negative, and neutral emotional valence. All video stimuli are presented to each pair of participants in a random order, and each trial follows a watch-judge-rest structure.
3. The generative deepfake video detection method according to claim 1, characterized in that, In step S2, the process of preprocessing the acquired raw EEG signals from both individuals includes: S201: Downsampling of EEG signals from two individuals; S202: Apply a 0.1-60 Hz bandpass filter to retain the frequency range of interest, and then apply a 48-52 Hz notch filter to remove power frequency noise; S203: The filtered data is subjected to whole-brain average rereference to suppress common-mode noise and improve signal purity; S204: Independent component analysis algorithm is used to decompose the EEG signals of two people into independent components for the purpose of separating and removing artifacts; S205: Time-frequency decomposition is performed using wavelet transform into five frequency bands, including Delta, Theta, Alpha, Beta, and Gamma, to extract power and phase information that varies with time.
4. The generative deepfake video detection method according to claim 1, characterized in that, In step S3, the three indices of interbrain correlation (ISC), interbrain coherence (Coh), and phase-locked value (PLV) of the two-person EEG signals are calculated to generate corresponding feature matrices for the five frequency bands of the original two-person EEG signal decomposition.
5. The generative deepfake video detection method according to claim 1, characterized in that, The multi-view feature encoding branch inputs the interbrain correlation matrix ISC, phase-locked value matrix PLV, and coherence matrix COH into three independent improved residual network branches. Each branch independently learns the deep nonlinear features of the corresponding index and outputs feature vectors of interbrain coherence features from three different perspectives. Phase-locked loop characteristics Correlation characteristics between brains The three feature vectors are processed by a feature fusion module based on a multi-view attention mechanism to output a fused feature. The final global feature vector is constructed by feature concatenation. , global feature vector The input is fed into the classifier module, which outputs the predicted probability of whether the video is real or fake.
6. The generative deepfake video detection method according to claim 1, characterized in that, In step S4, the multi-view interbrain synchronous feature fusion network further includes a feature fusion module.
7. The generative deepfake video detection method according to claim 6, characterized in that, The feature fusion module combines feature vectors from three different perspectives with brain coherence features. Phase-locked loop characteristics Correlation characteristics between brains The input tensor is concatenated to form the input tensor. The weights of each perspective are calculated using the Query-Key-Value attention paradigm, and then weighted and fused to output a multi-view brain synchronous fusion feature vector. .
8. The generative deepfake video detection method according to claim 7, characterized in that, The steps for calculating the weights of each viewpoint and performing weighted fusion include: First, let's define the three input variables. , , After being concatenated, they form an input tensor with a dimension of 768. Then, a fully connected layer maps the input tensor to a query Q, a key K, and a value V: ; in, A learnable linear mapping weight matrix is used to project the input features onto a high-dimensional semantic space; subsequently, a normalized attention weight map A is generated through dot product and activation function. ; in, It is the scaling factor in the attention mechanism. The dimension of the feature vector; The calculated weight map A is used to weight the value vector V, and the multi-view brain-brain synchronous fusion feature vector is output. : ; in, This indicates element-wise multiplication or matrix multiplication.
9. A generative deepfake video detection system, characterized in that, include: The dual-person EEG signal acquisition module is used to design dual-person hyperscanning experimental paradigms under video stimulation and acquire dual-person EEG signals to construct dual-person EEG databases under real and deepfake video stimulation. A dual-person EEG signal preprocessing module is used to preprocess the raw dual-person EEG signals acquired by the dual-person EEG signal acquisition module. The multidimensional neural synchronization matrix extraction module is used to extract the multidimensional neural synchronization matrix of the two-person EEG from the data preprocessed by the two-person EEG signal preprocessing module, including the interbrain correlation matrix ISC, the phase-locked value matrix PLV, and the coherence matrix COH. The feature processing and video forgery result output module is used to construct a multi-view brain-to-brain synchronization feature fusion network. It includes a prior knowledge-guided graph topology branch and a multi-view feature encoding branch. Each synchronization matrix is input into the multi-view feature encoding branch. A brain network spatial topology graph is constructed based on electrode prior knowledge. This graph is then input into the prior knowledge-guided graph topology branch. The features output from the prior knowledge-guided graph topology branch and the multi-view feature encoding branch are fused and concatenated to predict the authenticity of the video content. The prior knowledge-guided graph topology branch is used for in-depth mining of the spatial topological features of brain-to-brain synchronization, specifically including the following implementation steps: S401: Construct a graph structure G=(V, E) containing N nodes, where the node set V corresponds to the N electrode channels of the EEG acquisition device, and E is the set of connected edges after filtering in the inter-subject neural synchronization graph matrix. For the i-th node in the graph structure, directly extract the i-th row or column vector corresponding to the synchronization matrix as the initial feature of that node. This initial feature It physically characterizes the synchronous connectivity patterns between a specific brain region of subject A and the whole brain region of subject B; S402: Obtain the 3D physical coordinates of each electrode in the standard MNI brain space, generate coordinate embeddings through nonlinear mapping to provide absolute position information, calculate random walk probabilities based on the connection topology between electrodes to provide relative structural information of nodes in the graph, and combine the two prior codes with the initial features. Merge to generate the node input for layer 0. ; S403: Set up an improved graph attention layer of L layers to iteratively update node features, where L > 1. In each layer, the allocation of attention weights is explicitly guided by introducing an electrode prior adjacency matrix. This is repeated L times, so that node features can perform multiple rounds of information interaction and integration on the whole brain topology. S404: The obtained set of node features containing deep spatial topological semantics is integrated into a whole-brain topological feature vector through a global aggregation operation. .
Citation Information
Patent Citations
Learning attention state evaluation method based on multi-dimensional feature fusion network
CN117173758A
Time-frequency space electroencephalogram emotion recognition method based on three-dimensional space position embedding
CN121101563A