A method and system for reliable detection of generated video content
By employing a three-path feature extraction and a multi-scale consistency fluctuation kernel mechanism, combined with pseudo-label self-learning, the problem of insufficient detection complexity for multimodal forgery attacks in existing technologies is solved, achieving high-precision, anti-interference, and adaptive detection of deepfake videos.
Patent Information
- Application Number
- CN202511152682.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-18
- Publication Date
- 2025-10-31
- Estimated Expiration
- 2045-08-18
AI Technical Summary
Existing technologies are insufficient in detecting multimodal joint forgery attacks, especially local tampering and time-series frame replacement, and lack continuous adaptive mechanisms in open environments, making it difficult to effectively identify the credibility of deepfake videos.
We employ a three-path modeling approach (spatial appearance, motion, and semantics) to extract features. We combine cross-segment bidirectional residual evolution and a multi-scale consistent fluctuation kernel mechanism. We construct a credibility scoring mechanism by quantifying residual stability through Gaussian perturbation and kernel matrix analysis, and introduce a pseudo-label self-learning mechanism to optimize the model.
It significantly improves the detection accuracy and anti-interference ability of deepfake videos, enhances the ability to identify complex forgeries, and ensures long-term effectiveness and adaptability in open environments.
Smart Images

Figure CN120726538B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of multimodal content security detection technology, specifically to a method and system for reliable detection of video generated content. Background Technology
[0002] With the rapid development of deepfake technology, video forgery methods based on generative adversarial networks and diffusion models have made continuous breakthroughs in realism and generation efficiency, leading to the large-scale dissemination of forged videos on social media and public information platforms, which seriously threatens the credibility of digital content, social trust systems and national security.
[0003] Chinese invention patent CN112016500B discloses a method and system for identifying abnormal group behavior based on multi-scale temporal information fusion. The scheme includes acquiring a video frame sequence containing only normal samples and extracting optical flow maps between adjacent frames; sampling the same video segment at different sampling rates to obtain video sequences at different time scales; acquiring skipped frame video frame sequences from the aforementioned video frame sequences and extracting optical flow maps from the skipped frames to obtain video frame sequences and optical flow map sequences at different time scales; training a deep convolutional network using normal frames, skipped frame video sequences, and optical flow map sequences as input; using the trained deep convolutional network to perform abnormal behavior detection; the detection network evaluates the obtained predicted images to obtain anomaly scores; and judging whether a video frame is an abnormal frame based on the score, thus completing the anomaly detection of the video frame.
[0004] Existing technologies mostly focus on single-dimensional feature analysis (such as spatial artifacts or motion trajectories), which have insufficient complexity when dealing with multimodal joint forgery attacks (such as local tampering and temporal frame skipping replacement). Further research is needed on continuous adaptive mechanisms in open environments to improve long-term generalization capabilities. Summary of the Invention
[0005] The purpose of this invention is to address the problems existing in the background technology by proposing a method and system for reliable detection of video generated content.
[0006] The technical solution of this invention: a method for reliable detection of video generated content, comprising the following specific implementation steps:
[0007] S1. Divide the input samples into non-overlapping segments and construct three parallel modeling paths. Use a convolutional encoder to extract static features through the spatial appearance path, an optical flow encoder to extract dynamic features through the motion path, and a semantic encoder to extract semantic features through the semantic path. After aligning and fusing the features of the three paths, a unified dynamic representation sequence is constructed through standardization and position encoding.
[0008] S2. Extract temporal evolution differences through cross-segment bidirectional residual evolution modeling, inject Gaussian perturbation to quantize space and motion feature stability, and calculate residual stability index based on residual differences before and after perturbation to evaluate vulnerable areas of the video.
[0009] S3. The average perturbation residual within the sliding window is calculated by setting a multi-scale time window. For each scale, a kernel function based on Euclidean distance is defined to quantify the similarity between windows and construct a kernel matrix. The kernel symmetry deviation index is extracted and then weighted and fused with the residual stability index to generate a consistent energy field.
[0010] S4. By summarizing the multi-scale consistent energy field, perturbation response and residual stability index, a credibility tuple is constructed and normalized. A credibility score is generated by dynamic weighted fusion based on index variance. The optimal judgment threshold is adaptively learned to distinguish the authenticity of the video. A secondary perturbation analysis is performed on the boundary blurred samples to fuse the score.
[0011] S5. By filtering fuzzy samples near the threshold, high-confidence pseudo-labels are generated based on the initial score and the secondary perturbation score. New forgery patterns are identified by clustering the perturbation response and the consistency energy field characteristics, and representation centers are constructed. The model parameters are fine-tuned by combining the binary cross-entropy loss and the KL divergence distribution alignment loss.
[0012] Preferably, the three parallel modeling paths include:
[0013] Path A, Spatial Appearance Path: For each video segment, a lightweight convolutional encoder (MobileNet) is applied frame-by-frame to extract static visual features of spatial appearance, which are then stacked into d values for each frame. A 3D feature sequence tensor ;
[0014] Path B, Motion Path: For each video segment, calculate the inter-frame optical flow sequence, use the PWCNet optical flow estimator to obtain the inter-frame optical flow field, and use a motion encoder for convolutional pooling to generate features with dimension d. M motion feature sequence ;
[0015] Pathway C, Semantic Evolution Pathway: For each video frame, a pre-trained visual language model is applied to extract semantic representation vectors. These vectors are stacked into a semantic sequence and then used to model the temporal evolution through Bi-GRU and self-attention mechanisms to generate dynamic representation vectors. .
[0016] Preferably, the calculation process for the difference in residuals before and after the disturbance is as follows:
[0017] Compressed representations are obtained by performing time-averaged pooling on the dynamic features of the segments. The forward and reverse residual vectors of adjacent segments are calculated and concatenated into a residual tensor.
[0018] Gaussian perturbations are injected into a subset of the spatial and motion features of each segment to generate perturbation features. The perturbation features are then concatenated with the unperturbed semantic features to form a new representation. Subsequently, residual modeling is repeated to calculate the residual tensor after perturbation. The difference between the residuals before and after perturbation is measured based on Euclidean distance to obtain the perturbation response measure for each segment pair.
[0019] Preferably, the residual stability index is: ;
[0020] RSI represents the residual stability index; This represents a constant to prevent division by zero; N represents the number of time slices, satisfying: ; This represents the floor function; T represents the number of video frames. The length of any video segment; This represents the residual tensor formed by concatenating the forward and reverse residuals to create a fragment pair; This represents the Euclidean distance between the residual tensors before and after the perturbation of the i-th segment pair, i.e., the perturbation response metric.
[0021] Preferably, the kernel matrix construction process is as follows:
[0022] Define a finite set of time scales: ;
[0023] Where S is a set of multi-scale time windows, containing L sliding window scales of different lengths; This represents the length of the l-th time window in the scale set;
[0024] For each scale Apply a sliding window across the entire residual sequence and calculate the average perturbation residual within each window: ;
[0025] in, Representing scale Below, the average tensor of the perturbation residuals within the sliding window starting from the j-th time position;
[0026] For any two scale windows and Define a consistency kernel:
[0027] ;
[0028] in, Representing scale The kernel similarity values of the perturbation residuals for the j-th and k-th time windows; Represents the natural exponential function; Representing scale The bandwidth parameter of the kernel function;
[0029] For each time scale, an exponential kernel function based on Euclidean distance is defined to calculate the similarity value of perturbation residuals between windows of the same scale, and a kernel matrix is constructed: ;
[0030] in, Representing scale The perturbation residual kernel similarity matrix is shown below, with rows and columns corresponding to different time window indices.
[0031] Preferably, the process for generating a uniform energy field is as follows:
[0032] Construct a symmetry deviation index: ;
[0033] in, Representing scale Nuclear symmetry deviation;
[0034] The KSD at several scales is weighted and fused with the residual stability index RSI to obtain a comprehensive consistency score: ;
[0035] in, The weights for multi-scale consistent energy field fusion satisfy the following conditions: CEF stands for Consistency Energy Field, which is the final overall credibility score of the video.
[0036] Preferably, the process of adaptively learning the optimal threshold to distinguish the authenticity of videos is as follows:
[0037] By summarizing the multi-scale consistent energy field (CEF) and perturbation response metric of each video segment A local confidence tuple and a global index set are constructed using the residual stability index RSI, and nonlinear normalization is performed based on the statistical mean of the training set.
[0038] ;
[0039] in, This represents the original global value of the j-th indicator, where j=1 is CEF, j=2 is D, and j=3 is RSI. Let represent the statistical mean of the training set for the j-th indicator; This represents the scaling factor; This represents the j-th index value after normalization;
[0040] Define the final credibility score S: ;
[0041] Weights are automatically generated based on indicator stability. : ;
[0042] Where S represents the final credibility fusion score; Indicates the normalization index The fusion weights for the final score S; This represents the variance of the j-th normalized index at the fragment level;
[0043] By defining an expected misjudgment risk function that integrates the false positive rate and the false negative rate, the optimal decision boundary threshold that minimizes risk loss is statistically learned on a dynamic sliding window, and the video authenticity classification decision is made based on this threshold.
[0044] Preferably, the feedback execution process for performing secondary perturbation analysis and fusion scoring on samples with ambiguous boundaries is as follows: when the absolute difference between the final credible score and the adaptive judgment threshold is less than a set value, the feedback process is triggered, the perturbation response and multi-scale consistency analysis are re-executed to generate a secondary score, and the first score and the secondary score are weighted and fused to generate an enhanced result and update historical indicators to optimize threshold learning.
[0045] Preferably, the implementation process of fine-tuning the model parameters by combining the binary cross-entropy loss and the KL divergence distribution alignment loss is as follows:
[0046] Take the video V to be processed, obtain its final reliable score S, if it satisfies If the video is considered a potentially high-value ambiguous sample, then the following fusion-based pseudo-label generation function is used:
[0047] ;
[0048] in, This indicates increased tolerance for false label detection; This represents the automatically generated pseudo-label value. If it can be confirmed as a fake, it is 1; if it is real, it is 0; if it is uncertain, it is "uncertain" and will not be used in training.
[0049] Construct a disturbance response memory cache unit: Using similarity clustering algorithm to analyze disturbance response and consistency Cluster analysis is performed to identify new forgery style groups with significant structural differences, and a local representation center vector C is constructed for each forgery style. k ;
[0050] Where M is the perturbation response memory cache unit, containing the residual tensor, perturbation response, consistency kernel feature, and pseudo-label of all automatically labeled samples; C k Let be the cluster center vector of the k-th class of fake style in the perturbation response space;
[0051] Select samples with high reliability of pseudo-labels ( ), enter fine-tuning set D fine And perform lightweight fine-tuning using the following loss function:
[0052] ;
[0053] in, This represents the overall loss function for fine-tuning the training. Represents the weighting coefficients of the KL divergence loss term; This indicates that the current model is in the input fragment F i The probability distribution of the predicted output; This represents the binary cross-entropy loss; This indicates that the current detection model detects video segment F. i Output forgery credibility score; This represents the KL divergence.
[0054] The technical solution of the present invention: a video-generated content credibility detection system, which is used to perform the above-mentioned video-generated content credibility detection method, comprising:
[0055] Memory;
[0056] processor;
[0057] A computer program stored in the memory and capable of running on the processor;
[0058] The processor executes a computer program to implement the aforementioned method for reliable detection of video generated content.
[0059] Compared with the prior art, the above-mentioned technical solution of the present invention has the following beneficial technical effects:
[0060] This invention designs a method and system for reliable detection of video generated content. The overall technical solution achieves substantial breakthroughs in detection accuracy, anti-interference capability, and continuous evolution. Specifically: First, based on independent modeling and unified fusion of spatial appearance, motion, and semantic pathways, it comprehensively captures static texture anomalies, dynamic trajectory fragmentation, and semantic evolution contradictions in deepfake videos, effectively solving the limitations of single-dimensional detection. Second, a cross-segment bidirectional residual evolution mechanism combined with adversarial perturbation sensitivity quantization accurately identifies the vulnerability of temporal fragmentation and forged regions, significantly enhancing the ability to identify covert forgeries such as frame skipping and local tampering. Third, a multi-scale consistency fluctuation kernel mechanism effectively resists complex interference and improves model generalization by quantifying the structural stability of the video at multiple time scales. Finally, a pseudo-label-guided self-learning mechanism dynamically updates the forgery style memory and optimizes the decision boundary, continuously adapting to new forgery techniques and ensuring the long-term effectiveness of the system in open environments. Attached Figure Description
[0061] Figure 1This is a flowchart of a video generation content credibility detection method proposed in this invention. Detailed Implementation
[0062] Example 1, as Figure 1 As shown, the present invention proposes a method for reliable detection of video generated content, which includes the following specific implementation steps:
[0063] S1. The input video is decomposed, modeled, and uniformly represented from multi-dimensional information paths to construct a dynamic representation tensor that can be used for subsequent forgery detection. The structured representation of video semantic behavior is achieved through temporal segmentation, three-path dynamic modeling, and multi-dimensional feature alignment and fusion. The specific implementation process is as follows:
[0064] S11, Given the input video Divide it into N non-overlapping segments, each segment having a length of This yields a set of fragments: ;
[0065] Where V represents the original input video sequence, containing T frames, each frame being an RGB image with a size of H×M; T represents the number of video frames; H and W represent the height and width of the video frame, respectively; and N represents the number of time segments, satisfying: ; This represents the floor function; This represents the i-th time segment, which is composed of Δt frames continuously extracted from V.
[0066] S12. To capture the dynamic evolution of video content across different dimensions, three parallel modeling pathways are constructed: spatial appearance pathway, motion pathway, and semantic pathway. Each pathway independently extracts temporal features in its own dimension and aligns them along the timeline. Specifically:
[0067] (1) Path A, spatial appearance path, which models the static changes of video frames at the visual pixel level and is suitable for detecting static camouflage features such as local texture fusion, edge forgery, and face camouflage. Specifically:
[0068] For each segment Input Lightweight Convolutional Encoder For each frame, features are extracted and stacked into a tensor:
[0069] ;
[0070] in, This represents the spatial appearance feature sequence extracted from the i-th segment, with one d per frame. A dimensional vector; This represents the spatial feature vector of the j-th frame in the i-th segment; This represents the spatial encoder function, which extracts the static visual features of each frame. In this embodiment, the lightweight convolutional network MobileNet is used.
[0071] (2) Path B, motion path, which captures the inter-frame motion trend in the video and is used to identify dynamic forgeries such as unreasonable motion, region drift, and frame skipping. Specifically:
[0072] Calculate the inter-frame optical flow sequence for each video segment:
[0073] ;
[0074] ;
[0075] in, Represents the sequence of all inter-frame optical flow in the i-th segment, with a length of . ; This represents the optical flow field between the j-th frame and the (j+1)-th frame in the i-th segment, and the two-dimensional vector describes the direction and amplitude of pixel motion. The optical flow estimator (PWCNet is used in this embodiment) extracts the two-dimensional motion vector field between frames; Represents two frames of images and Calculate optical flow between them;
[0076] Using motion encoders Convolutional pooling encoding of the optical flow sequence:
[0077] ;
[0078] in, This represents an optical flow encoder that performs convolutional pooling on an optical flow sequence and outputs fixed-length motion features. d represents the sequence of motion features extracted from the i-th segment; M Characteristic dimensions representing the motor pathway;
[0079] (3) Pathway C, semantic evolution path, models the continuity and evolutionary stability of frame-level semantic concepts, specifically:
[0080] For each frame, a visual language model (such as CLIP, BLIP) is applied to obtain a semantic vector:
[0081] ;
[0082] Obtain the frame-level semantic sequence: ;
[0083] Input Time Modeler (This embodiment uses Bi-GRU+Self-Attention) to dynamically model the semantic sequence and generate evolution vectors: ;
[0084] in, The semantic encoder, in this embodiment, is set as the visual encoding part in a pre-trained graph-text model (such as CLIP, BLIP), which extracts the semantic representation vector of each frame; This represents the semantic vector of the j-th frame in the i-th segment; Represents the semantic sequence of the i-th segment (the semantic vectors of each frame are stacked in time). This represents the dynamic semantic representation of the i-th segment; This represents a semantic modeler that captures the temporal evolution trend of semantics; d S This represents the semantic vector dimension of each frame;
[0085] It should be noted that CLIP (Contrastive Language-Image Pretraining) aligns image and text features through contrastive learning, is trained using 400 million network image-text pairs, supports zero-shot image classification (such as matching an image to any text category), and is widely used in image-text retrieval, content moderation, and multimodal research. BLIP (Bootstrapping Language-Image Pretraining) focuses on generation and understanding, combining a Captioner (generating descriptions), a Filter (cleaning noisy data), and a Retriever (image-text matching). It employs model bootstrapping techniques to optimize training data quality and excels in image caption generation (such as generating natural language titles for images) and visual question answering (VQA) tasks.
[0086] S13. Perform unified alignment and normalization fusion to construct the final unified representation tensor for subsequent analysis, specifically:
[0087] Align the three-channel outputs to a consistent number of frames in the time dimension, and unify the dimension through channel transformation:
[0088] ; ; ;
[0089] in, This represents the spatially aligned and compressed feature sequence, with all segments normalized to a consistent frame number. With dimension ; Indicates motion path alignment output; Indicates semantic path alignment output; This indicates that the sequence length of different pathways is uniformly determined by interpolation / sampling in the time dimension.
[0090] The final tensor is constructed by fusion. : ;
[0091] in, The i-th segment represents the final fused feature, which includes three types of representations: spatial, motion, and semantic. This indicates a tensor splicing operation (along the channel dimension). Indicates the uniform frame number after alignment; This represents the uniform dimension to which each path is compressed;
[0092] For the final tensor By performing channel normalization (LayerNorm) and incorporating learnable positional encoding, the model's ability to perceive temporal location is improved. Ultimately, all segments are combined to form a dynamic representation sequence of the entire video. ;
[0093] Where F represents the dynamic representation sequence of the entire video, which is used by the subsequent detection model.
[0094] S2. By constructing a cross-segment bidirectional residual evolution mechanism, introducing adversarial perturbation sensitivity assessment, and defining credibility indicators, the mid-to-high-level structural behavior analysis of the credibility of video-generated content is achieved, thereby identifying potential forgery signals such as evolutionary anomalies, motion trajectory fragmentation, and perturbation vulnerability in deepfakes. The specific implementation process is as follows:
[0095] S21. Construct a cross-segment bidirectional residual evolution modeling mechanism, introduce inverse residual modeling, observe temporal fragmentation through the symmetry of forward and inverse evolution, extract the temporal evolution differences between segments, and perform symmetry analysis on their forward and inverse residuals to capture anomalous temporal fragmentation, specifically:
[0096] For each dynamic feature fragment First, average pooling is performed in the time dimension to obtain the compressed representation of the segment: ;
[0097] in, Representing fragments Average feature representation over the time dimension; This represents the fused feature vector of the i-th segment at time step t, with dimension 1. ;
[0098] For adjacent segments and Construct a bidirectional residual vector:
[0099] Positive residual vector (representing the natural evolutionary trend): ;
[0100] Inverse residual vector (indicating the plausibility of the inversion): ;
[0101] Combined representation of the residual information of the i-th segment pair: ;
[0102] in, Represents the positive residual vector of the i-th pair of adjacent segments, describing the natural evolution of the video from i to i+1; This represents the inverse residual vector of the i-th pair of segments, used to determine the rationality of the inversion; This represents the residual tensor formed by concatenating the forward and reverse residuals to create a fragment pair;
[0103] S22. Perform local perturbation injection and residual response mapping. By injecting perturbation into the test segment and measuring the degree of change in the response to the perturbation, the stability of the segment is quantified, and it is determined whether it may be a forged region. Specifically:
[0104] For each fused fragment feature Inject Gaussian distributed perturbations (perturbing only the first two paths, namely spatial and motion paths): Construct the perturbation features, i.e., construct the perturbation vector: ;
[0105] in, Represented as fragment The generated Gaussian perturbation noise; express The identity matrix controls the covariance structure to be diagonal; The hyperparameter representing the control of disturbance intensity is set manually; This represents the generated perturbed fused representation fragment; This represents a subset of the fused features of the spatial appearance and motion path of the i-th segment in time for each frame. Source: Previous It is static image content, later It is inter-frame optical flow or dynamic motion information; This represents a subset of semantic pathway features for each frame in segment i, without any perturbation. The concatenation operation of features along the channel dimension (i.e., concatenating two tensors along the last dimension).
[0106] All segments after perturbation Repeat step S21 to calculate the corresponding residual tensor after perturbation:
[0107] ;
[0108] in, This means that the residual model is re-modeled for the perturbed features to obtain the perturbed residual, which is the i-th residual pair after the perturbation; This represents the positive residual vector of the i-th pair of adjacent segments after the perturbation; This represents the inverse residual vector of the i-th pair of adjacent segments after the perturbation;
[0109] Calculate the difference between the residual vectors before and after the perturbation to obtain the perturbation response metric for each segment pair:
[0110] ;
[0111] in, This represents the Euclidean distance between the residual tensors before and after the perturbation of the i-th segment pair, i.e., the perturbation response metric; This represents the L2 norm, or Euclidean distance, which measures the change in residuals caused by a disturbance.
[0112] S23. Calculate and determine the reliability of the residual stability index (RSI), converting the disturbance response into an interpretable numerical index as the basis for the video's reliability. Specifically:
[0113] ;
[0114] Wherein, RSI represents the residual stability index, with a range of [0,1]; This represents a constant that prevents division by zero, thus enhancing numerical stability.
[0115] S3. A dynamic consistency fluctuation kernel mechanism is proposed, which jointly models the perturbation response, residual evolution, and inter-scale variation structure to quantify the consistency characteristics of video across multiple time scales. The specific implementation process is as follows:
[0116] S31. Construct a multi-scale time segment and generated residual fusion view, and reconstruct the perturbation residual tensor according to different time scales, specifically:
[0117] Define a finite set of time scales:
[0118] ;
[0119] Where S is a set of multi-scale time windows, containing L sliding window scales of different lengths; This represents the length of the l-th time window in the scale set;
[0120] For each scale Apply a sliding window across the entire residual sequence and calculate the average perturbation residual within each window: ;
[0121] in, Representing scale Below, the average tensor of the perturbation residuals within the sliding window starting from the j-th time position, ;
[0122] S32. After obtaining the perturbation residual sequences at different scales, a set of residual consistency kernel functions is defined for each scale, and the patterns of local structural changes in the video are captured in a nonlinear manner, specifically as follows:
[0123] For any two scale windows and Their consistency kernel is defined as:
[0124] ;
[0125] Where j and k represent the positions of the j-th and k-th time windows, respectively; Representing scale The kernel similarity values of the perturbation residuals for the j-th and k-th time windows; Represents the natural exponential function; Representing scale The bandwidth parameter (standard deviation) of the kernel function. ;
[0126] For all windows<j,k> Construct a complete scale kernel matrix: ;
[0127] in, Representing scale The perturbation residual kernel similarity matrix is shown below, with rows and columns corresponding to different time window indices;
[0128] S33. Construct two types of quantification mechanisms to extract structured indicators from the kernel matrix, specifically:
[0129] Construct a symmetry deviation index: ;
[0130] in, Representing scale Kernel symmetry deviation.
[0131] The KSD at multiple scales is weighted and fused with the residual stability index RSI to obtain a comprehensive consistency score: ;
[0132] in, This represents the weighting of the multi-scale consistent energy field fusion, used to weight the contributions at each scale, satisfying... CEF stands for Consistency Energy Field, which is the final overall score for video credibility.
[0133] S4. Construct a multi-dimensional fusion credibility scoring mechanism, combining multi-scale consistent energy field (CEF) with perturbation response metric and residual stability index. Through designing a dynamic adaptive fusion model and credibility boundary inference algorithm, accurately determine whether a video is forged. The specific implementation process is as follows:
[0134] S41. Using the Consistent Energy Field (CEF) from step S3 as the core reliability indicator, simultaneously collect the perturbation response measurement from step S2. Using the residual stability index RSI, we summarize the three dimensions of the index for each video segment i to construct its local reliability tuple. For the entire video, construct a global credibility indicator set I: ;
[0135] To eliminate dimensional differences, all indicators are uniformly subjected to nonlinear normalization:
[0136] ;
[0137] in, Represents the credibility tuple of the i-th video segment, containing the multidimensional evaluation metrics of that segment; The multi-scale consistency energy field index reflects the degree of structural alignment of the segment in multi-scale residual analysis. This represents a perturbation response metric, quantifying the degree of change in the dynamic performance of the segment after countering perturbation injection; represents the residual stability index, which measures the stability of the perturbation residual distribution over time; I represents the set of global credibility indices for the entire video segment, representing the global levels of multi-scale consistency, perturbation responsibility, and stability, respectively. This represents the expected value of the index for all segments i, i.e., the average value at the video level; This represents the original global value of the j-th indicator (j=1 is CEF, j=2 is D, j=3 is RSI); Let represent the statistical mean of the training set for the j-th indicator; This represents the scaling factor, used to control the sensitivity of the normalization function; This represents the j-th index value after normalization, which serves as the basis for the fusion score.
[0138] S42. Design a fusion model based on dynamic weight adjustment of a multi-dimensional index space to generate the final credibility score S, specifically:
[0139] Define the final credibility score S: ;
[0140] Weights are automatically generated based on indicator stability. : ;
[0141] Where S represents the final credibility fusion score, which ranges between (0,1), and the larger the value, the more likely it is to be a real video; Indicates the normalization index The fusion weight for the final score S depends on its cross-segment variance, i.e. its credibility in video judgment; This represents the variance of the j-th normalization index at the segment level, reflecting its volatility in the time series dimension. The smaller the value, the more stable the index, and the more important it is in the judgment.
[0142] S43. Construct an adaptive confidence boundary inference mechanism to dynamically adjust the confidence judgment threshold T, and introduce a confidence region judgment function to determine whether a sample falls into the boundary ambiguity zone, thereby achieving a more reliable recognition strategy, specifically:
[0143] Define the expected misjudgment risk function Combining the false positive rate (FPR) and the false negative rate (FNR), the optimal confidence boundary T is derived: ;
[0144] Statistical learning through a dynamic sliding window: ;
[0145] Use the calculated threshold To determine the credibility of the test samples:
[0146] ;
[0147] in, This indicates the false positive rate, which is the proportion of genuine videos that are mistakenly identified as fake videos. This indicates the false negative rate, which is the proportion of fake videos that are mistakenly identified as real videos. Represents the risk loss function; represents the risk balance coefficient, which controls the relative importance of FPR and FNR in the loss function; T represents the decision boundary for classifying videos as real or fake. The optimal value representing the decision boundary for classifying a video as real or fake is learned by an adaptive algorithm;
[0148] S44. A confidence assessment and self-feedback update mechanism is added to handle uncertain samples near the decision boundary, specifically:
[0149] If the final score is too close to the judgment boundary, that is: If the decision is not made with sufficient confidence, it is considered that the model does not have confidence in its judgment of the sample, and then the feedback optimization process begins.
[0150] in, This represents the confidence boundary tolerance constant;
[0151] The video undergoes a second perturbation process, following steps S2 and S3, to generate a second round of perturbation response and consistency analysis. This is then combined with the initial score for reweighted fusion. And update historical statistical indicators to optimize future threshold learning;
[0152] in, This represents the score result after confidence enhancement, which integrates the initial judgment score S and the secondary perturbation analysis score. get; This represents the confidence score calculated after secondary perturbation of the fuzzy boundary samples in the feedback process;
[0153] Therefore, by using dynamic perturbation comparison to enhance the discrimination ability of uncertain samples, simulating the "secondary judgment process of the human eye", the model's performance on fuzzy boundaries is greatly improved.
[0154] S5. Construct a self-learning mechanism that combines pseudo-label enhancement guidance, dynamic forgery memory module updates, and positive / negative sample boundary correction feedback to achieve continuous adaptation to new forged samples. The specific implementation process is as follows:
[0155] S51. Discover and utilize video samples (with scores close to the threshold) where the model is not yet clear as potential learning samples, and automatically generate pseudo-labels for soft training, specifically:
[0156] Take the video V to be processed and obtain its final credibility score S;
[0157] If satisfied If so, the video is considered a potentially high-value fuzzy sample;
[0158] Use the following blended pseudo-tag generation function:
[0159] ;
[0160] in, This indicates an increased tolerance for false labeling, used in both the initial score (S) and the secondary score. Based on this, a more stringent confidence range is set to improve the reliability of fake tags; This represents the automatically generated pseudo-label value. If it can be confirmed as a fake, it is 1; if it is real, it is 0; if it is uncertain, it is "uncertain" and will not be used in training.
[0161] S52. Identify the perturbation behavior patterns of unknown types of forged videos and dynamically enhance the detection system, specifically:
[0162] Construct a disturbance response memory cache unit:
[0163] ;
[0164] The disturbance response is analyzed using a similarity clustering algorithm (DBSCAN in this embodiment). and consistency Cluster analysis was performed to identify new forgery style groups with significant structural differences;
[0165] Construct the local representation center vector C for each type of forgery style k This is used to update the discriminator's attention mechanism or to forge the distribution prior.
[0166] Where M is the perturbation response memory cache unit, a set containing the residual tensors, perturbation responses, consistency kernel features, and pseudo-labels of all automatically labeled samples; C k Let be the cluster center vector of the k-th forgery style in the perturbation response space, used to model new forgery types;
[0167] S53. Use the trusted boundary samples (automatically generated pseudo-labels) and the identified new forgery style features to fine-tune the original discrimination model, so that it maintains robustness and generalization in new scenarios, specifically:
[0168] Select samples with high reliability of pseudo-labels ( ), enter fine-tuning set D fine ;
[0169] Lightweight fine-tuning can be performed using the following loss function:
[0170] ;
[0171] in, The overall loss function for fine-tuning training consists of two parts: cross-entropy loss and distribution alignment loss. This represents the weighting coefficient of the KL divergence loss term, used to adjust the weight balance between cross-entropy loss and forgery alignment. This indicates that the current model is in the input fragment F i The probability distribution of the predicted output is used to compare the degree of matching between the real pseudo-label and the model output; This represents the binary cross-entropy loss, used to measure the accuracy of the detector's judgment of false labels; This indicates that the current detection model detects video segment F. i Output forgery credibility score; This represents the KL divergence.
[0172] Example 2: A video-generated content credibility detection system proposed in this invention is used to execute a video-generated content credibility detection method proposed in Example 1, comprising:
[0173] Memory;
[0174] processor;
[0175] A computer program stored in the memory and capable of running on the processor;
[0176] The processor executes a computer program to implement a video generation content credibility detection method as described in Embodiment 1 above.
[0177] The embodiments of the present invention have been described in detail above with reference to the accompanying drawings. However, the present invention is not limited thereto. Various changes can be made within the scope of knowledge possessed by those skilled in the art without departing from the spirit of the present invention.
Claims
1. A method for detecting the credibility of video-generated content, characterized in that, The specific implementation steps include the following: S1. Divide the input samples into non-overlapping segments and construct three parallel modeling paths. Use a convolutional encoder to extract static features through the spatial appearance path, an optical flow encoder to extract dynamic features through the motion path, and a semantic encoder to extract semantic features through the semantic path. After aligning and fusing the features of the three paths, a unified dynamic representation sequence is constructed through standardization and position encoding. The three parallel modeling pathways include: Path A, Spatial Appearance Path: For each video segment, a lightweight convolutional encoder (MobileNet) is applied frame-by-frame to extract static visual features of spatial appearance, which are then stacked into d values for each frame. A 3D feature sequence tensor ; Path B, Motion Path: For each video segment, calculate the inter-frame optical flow sequence, use the PWCNet optical flow estimator to obtain the inter-frame optical flow field, and use a motion encoder for convolutional pooling to generate features with dimension d. M motion feature sequence ; Pathway C, Semantic Evolution Pathway: For each video frame, a pre-trained visual language model is applied to extract semantic representation vectors. These vectors are stacked into a semantic sequence and then used to model the temporal evolution through Bi-GRU and self-attention mechanisms to generate dynamic representation vectors. ; S2. Extract temporal evolution differences through cross-segment bidirectional residual evolution modeling, inject Gaussian perturbation to quantize space and motion feature stability, and calculate residual stability index based on residual differences before and after perturbation to evaluate vulnerable areas of the video. S3. The average perturbation residual within the sliding window is calculated by setting a multi-scale time window. For each scale, a kernel function based on Euclidean distance is defined to quantify the similarity between windows and construct a kernel matrix. The kernel symmetry deviation index is extracted and then weighted and fused with the residual stability index to generate a consistent energy field. The kernel matrix construction process is as follows: Define a finite set of time scales: ; Where S is a set of multi-scale time windows, containing L sliding window scales of different lengths; This represents the length of the l-th time window in the scale set; For each scale Apply a sliding window across the entire residual sequence and calculate the average perturbation residual within each window: ; in, Representing scale Below, the average tensor of the perturbation residuals within the sliding window starting from the j-th time position; For any two scale windows and Define a consistency kernel: ; in, Representing scale The kernel similarity values of the perturbation residuals for the j-th and k-th time windows; Represents the natural exponential function; Representing scale The bandwidth parameter of the kernel function; For each time scale, an exponential kernel function based on Euclidean distance is defined to calculate the similarity value of perturbation residuals between windows of the same scale, and a kernel matrix is constructed: ; in, Representing scale The perturbation residual kernel similarity matrix is shown below, with rows and columns corresponding to different time window indices; S4. By summarizing the multi-scale consistent energy field, perturbation response and residual stability index, a credibility tuple is constructed and normalized. A credibility score is generated by dynamic weighted fusion based on index variance. The optimal judgment threshold is adaptively learned to distinguish the authenticity of the video. A secondary perturbation analysis is performed on the boundary blurred samples to fuse the score. S5. By filtering fuzzy samples near the threshold, high-confidence pseudo-labels are generated based on the initial score and the secondary perturbation score. New forgery patterns are identified by clustering the perturbation response and the consistency energy field features, and representation centers are constructed. The model parameters are fine-tuned by combining the binary cross-entropy loss and the KL divergence distribution alignment loss.
2. The video generation content credibility detection method according to claim 1, characterized in that, The calculation process for the difference in residuals before and after the disturbance is as follows: Compressed representations are obtained by performing time-averaged pooling on the dynamic features of the segments. The forward and reverse residual vectors of adjacent segments are calculated and concatenated into a residual tensor. Gaussian perturbations are injected into a subset of the spatial and motion features of each segment to generate perturbation features. The perturbation features are then concatenated with the unperturbed semantic features to form a new representation. Subsequently, residual modeling is repeated to calculate the residual tensor after perturbation. The difference between the residuals before and after perturbation is measured based on Euclidean distance to obtain the perturbation response measure for each segment pair.
3. The video generation content credibility detection method according to claim 2, characterized in that, The residual stability index is: ; RSI represents the residual stability index; This represents a constant to prevent division by zero; N represents the number of time slices, satisfying: ; This represents the floor function; T represents the number of video frames. The length of any video segment; This represents the residual tensor formed by concatenating the forward and reverse residuals to create a fragment pair; This represents the Euclidean distance between the residual tensors before and after the perturbation of the i-th segment pair, i.e., the perturbation response metric.
4. The video generation content credibility detection method according to claim 3, characterized in that, The process of generating a uniform energy field is as follows: Construct a symmetry deviation index: ; in, Representing scale Nuclear symmetry deviation; The KSD at several scales is weighted and fused with the residual stability index RSI to obtain a comprehensive consistency score: ; in, The weights for multi-scale consistent energy field fusion satisfy the following conditions: CEF stands for Consistency Energy Field, which is the final overall credibility score of the video.
5. The video generation content credibility detection method according to claim 4, characterized in that, The adaptive learning-based optimal threshold for distinguishing video authenticity is as follows: By summarizing the multi-scale consistent energy field (CEF) and perturbation response metric of each video segment A local confidence tuple and a global index set are constructed using the residual stability index RSI, and nonlinear normalization is performed based on the statistical mean of the training set. ; in, This represents the original global value of the j-th indicator, where j=1 is CEF, j=2 is D, and j=3 is RSI. Let represent the statistical mean of the training set for the j-th indicator; This represents the scaling factor; This represents the j-th index value after normalization; Define the final credibility score S: ; Weights are automatically generated based on indicator stability. : ; Where S represents the final credibility fusion score; Indicates the normalization index The fusion weights for the final score S; This represents the variance of the j-th normalized index at the fragment level; By defining an expected misjudgment risk function that integrates the false positive rate and the false negative rate, the optimal decision boundary threshold that minimizes risk loss is statistically learned on a dynamic sliding window, and the video authenticity classification decision is made based on this threshold.
6. The video generated content credibility detection method according to claim 5, characterized in that, The feedback execution process for performing secondary perturbation analysis and fusion scoring on samples with ambiguous boundaries is as follows: when the absolute difference between the final reliable score and the adaptive judgment threshold is less than the set value, the feedback process is triggered, the perturbation response and multi-scale consistency analysis are re-executed to generate a secondary score, and the first score and the secondary score are weighted and fused to generate an enhanced result and update historical indicators to optimize threshold learning.
7. The video generation content credibility detection method according to claim 1, characterized in that, The process of fine-tuning the model parameters by combining the binary cross-entropy loss and the KL divergence distribution alignment loss is as follows: Take the video V to be processed, obtain its final reliable score S, if it satisfies If the video is considered a potentially high-value ambiguous sample, then the following fusion-based pseudo-label generation function is used: ; in, This indicates increased tolerance for false label detection; This represents the automatically generated pseudo-label value. If it can be confirmed as a fake, it is 1; if it is real, it is 0; if it is uncertain, it is "uncertain" and will not be used in training. Construct a disturbance response memory cache unit: Using similarity clustering algorithm to analyze disturbance response and consistency Cluster analysis is performed to identify new forgery style groups with significant structural differences, and a local representation center vector C is constructed for each forgery style. k ; Where M is the perturbation response memory cache unit, containing the residual tensor, perturbation response, consistency kernel feature, and pseudo-label of all automatically labeled samples; C k Let be the cluster center vector of the k-th class of fake style in the perturbation response space; Select samples with high reliability of pseudo-labels ( ), enter fine-tuning set D fine And perform lightweight fine-tuning using the following loss function: ; in, This represents the overall loss function for fine-tuning the training. Represents the weighting coefficients of the KL divergence loss term; This indicates that the current model is in the input fragment F i The probability distribution of the predicted output; This represents the binary cross-entropy loss; This indicates that the current detection model detects video segment F. i Output forgery credibility score; This represents the KL divergence.
8. A video-generated content credibility detection system, used to execute the video-generated content credibility detection method according to any one of claims 1 to 7, characterized in that, include: Memory; processor; A computer program stored in the memory and capable of running on the processor; The processor executes a computer program to implement the video generation content credibility detection method according to any one of claims 1 to 7.
Citation Information
Patent Citations
A Method and System for Identifying Abnormal Group Behavior Based on Multi-Scale Temporal Information Fusion
CN112016500B
Generative adversarial video super-resolution reconstruction and reconstructed image authenticity identification method
CN112070665A
Video stream dynamic fragment encryption and block chain evidence storage method
CN120416543A