Generated video detection method and device based on physically-driven space-time modeling, equipment and medium
By using a physics-driven spatiotemporal modeling method and the MMD deep kernel training model to extract video features, the problems of unstable performance and insufficient generalization of generated video detection are solved, and efficient video authenticity determination is achieved.
Patent Information
- Application Number
- CN202510791559.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-13
- Publication Date
- 2025-09-12
AI Technical Summary
Existing generated video detection methods suffer from unstable detection performance and insufficient cross-domain generalization, making it difficult to effectively distinguish real videos from synthetic videos.
A method based on physics-driven spatiotemporal modeling is adopted. By extracting the normalized spatiotemporal gradient features of generated videos and real videos, the MMD deep kernel training model is used to increase the inter-class distance and reduce the intra-class distance, and a difference measurement model is established to determine whether the video is generated by AI.
The accuracy and cross-domain generalization ability of generated video detection are improved, and the authenticity and security of videos can be reliably determined.
Smart Images

Figure CN120635783A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of digital video content identification, and in particular to a generated video detection method based on physics-driven spatiotemporal modeling. Background Art
[0002] In recent years, generative techniques based on diffusion models have made significant progress in synthesizing highly realistic videos. Diffusion frameworks such as Sora can generate temporal content that is visually indistinguishable from natural video, demonstrating great potential in fields such as virtual reality and film and television production. However, while high-quality synthetic video technology offers convenience, it also creates opportunities for malicious uses such as deepfakes and synthetic media fraud, posing a significant threat to information security and public opinion. Therefore, reliable AI-generated video detection technology is urgently needed.
[0003] Existing methods for detecting generated videos are mainly divided into physical constraint-based detection methods and model-based detection methods. Physical constraint-based detection methods use the physical laws of natural videos, such as motion continuity and brightness consistency, to detect anomalies in synthetic videos. However, since high-quality synthetic videos often contain only very subtle non-physical features, these methods lack accuracy. Model-based detection methods rely on large-scale supervised learning, using local light flow artifacts, appearance consistency statistics, or training specialized classifiers to distinguish real videos from synthetic videos. However, these methods rely heavily on the generative models covered by the training set, making it difficult to generalize to videos generated by new generative models. Summary of the Invention
[0004] To address the above problems, the present invention proposes a generative video detection method based on physics-driven spatiotemporal modeling, which mainly solves the problems of unstable detection performance and insufficient cross-domain generalization of existing generative video detection methods.
[0005] In order to solve the above technical problems, the technical solutions of the present invention are as follows:
[0006] A generative video detection method based on physics-driven spatiotemporal modeling includes the following steps:
[0007] Step 1: According to the pre-defined physical-driven spatiotemporal modeling formula, the normalized spatiotemporal gradient features of the generated video and the real video are extracted, which are defined as the generated video features and the real video features respectively.
[0008] Step 2: Establish a measurement model based on normalized spatiotemporal gradient feature differences, and train the depth kernel of the measurement model with the generated video features and the real video features. The optimization goal of the training process includes increasing the inter-class distance between the generated video features and the real video features, and reducing the intra-class distance of the real video features, to obtain the trained MMD depth kernel;
[0009] Step 3: For the detection task of the video to be tested, the difference between the video to be tested and the real video is calculated using the trained MMD depth kernel, which is defined as a difference value; and whether the video to be tested is generated by AI is determined based on the difference value.
[0010] In some embodiments, the physical driven spatiotemporal modeling formula is:
[0011] (1);
[0012] Where, Indicates time Next video frame The joint probability density of Directly estimated by the score network of the pre-trained diffusion model, Displacement by adjacent frames approximate( ), Ensure the stability of the denominator value.
[0013] In some implementations, the metric model employs a deep kernel:
[0014] (2);
[0015] Where, Indicates that the video is continuous Normalized spatiotemporal gradient feature sequence over frames; is a deep neural network mapping; and Both are bandwidth parameters and Gaussian kernel; is the balance coefficient; kernel parameter set To be learned.
[0016] In some embodiments, the optimization goal is:
[0017] (3);
[0018] Among them are
[0019] (4);
[0020] (5);
[0021] (6);
[0022] Where, and represent the real videos and generated videos in the training set respectively.
[0023] In some embodiments, the difference value is:
[0024] (7);
[0025] Where, and Represent the real video and the video to be detected in the reference set respectively; Indicates the Normalized spatiotemporal gradient features of reference videos; Indicates the video to be detected Normalized spatiotemporal gradient characteristics; is a kernel function, for example , mapping the video features to a reproducing kernel Hilbert space .
[0026] In some implementations, the process of generating a video detection is:
[0027] (8);
[0028] Where, is a decision threshold (usually set to 1); for the video to be detected The maximum mean difference between the normalized spatiotemporal gradient features of the real video in the reference set and the normalized spatiotemporal gradient features of the real video in the reference set is estimated. This function calculates the difference value by comparing With threshold The detection result is determined by the size relationship between : if it is greater than the threshold, the detection result is the generated video; otherwise, the detection result is the real video.
[0029] The beneficial effects of the present invention are: by extracting normalized spatiotemporal gradient features based on physics-driven spatiotemporal modeling, and training the deep kernel with the optimization goal of increasing the inter-class distance between generated video and real video features and reducing the intra-class distance of real video features, the problems of unstable detection performance and insufficient cross-domain generalization of existing methods are overcome, and it can accurately determine whether the video is generated by AI, providing solid protection for the credibility and security of video content. BRIEF DESCRIPTION OF THE DRAWINGS
[0030] Figure 1This is a flow chart of a method for generating video detection based on physical-driven spatiotemporal modeling disclosed in an embodiment of the present invention;
[0031] Figure 2 Experimental data on applying NSG-VD disclosed in the embodiments of the present invention to AI-generated video detection; DETAILED DESCRIPTION
[0032] To make the objectives, technical solutions, and advantages of the present invention more clear and distinct, the present invention is further described in detail below with reference to the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are merely for the purpose of explaining the present invention and are not intended to limit the present invention. It should also be noted that, for ease of description, the accompanying drawings only illustrate portions relevant to the present invention, not all of the present invention.
[0033] This embodiment proposes a method for generating video detection based on physical driven spatiotemporal modeling, such as Figure 1 As shown, the following steps are included:
[0034] Step 1: According to the pre-defined physical-driven spatiotemporal modeling formula, the score network based on the pre-trained diffusion model approximately extracts the normalized spatiotemporal gradient features of the generated video and the real video, which are defined as the generated video features and the real video features, respectively.
[0035] In this embodiment, the above physical driven spatiotemporal modeling formula is:
[0036] (1);
[0037] Where, Indicates time Next video frame The joint probability density of and time partial derivatives There are difficulties in implementation. It is directly estimated by the score network of the pre-trained diffusion model, while Displacement by adjacent frames Approximate calculation, that is In order to avoid the instability of the denominator of formula (1), we introduce .
[0038] Step 2: Establish a measurement model based on normalized spatiotemporal gradient feature differences, and train the depth kernel of the measurement model with the generated video features and the real video features. The optimization goal of the training process includes increasing the inter-class distance between the generated video features and the real video features, and reducing the intra-class distance of the real video features, to obtain the trained MMD depth kernel;
[0039] Maximum Mean Discrepancy (MMD) is a metric used to measure the difference between two probability distributions, P and Q. It is particularly used in machine learning and statistics, particularly in parameter-free and kernel methods. In this example, MMD is used as a metric model to measure the difference between two distributions.
[0040] In this embodiment, the measurement model uses a deep kernel:
[0041] (2);
[0042] Where, Indicates that the video is continuous Normalized spatiotemporal gradient feature sequence over frames; is a deep neural network mapping; and Both are bandwidth parameters and Gaussian kernel; is the balance coefficient; kernel parameter set To be learned.
[0043] In this embodiment, the optimization goal of the training process is:
[0044] (3);
[0045] Among them are
[0046] (4);
[0047] (5);
[0048] (6);
[0049] Where, and Represent the real videos and generated videos in the training set respectively; Indicates that it is based on a parameter set The MMD kernel, express The coefficient of difference between the two distributions is It represents the test power of the optimization objective and characterizes the certainty of the distribution difference. A higher test power indicates a greater certainty of the distribution difference.
[0050] The architecture design and training process of the MMD deep kernel are as follows:
[0051] Defining the MMD depth kernel For a neural network equipped with a feature extractor , the normalized spatiotemporal gradient features extracted in step 1 are used as the features of the input video, and the network It consists of a hidden layer transformer and a multi-layer perceptron. During the training process, the normalized spatiotemporal gradient features are first extracted from the real video and the generated video based on step 1 as the features of the input video, and then the deep kernel network is trained by maximizing the target in the equation Used to generate video detection scenes.
[0052] Step 3: For the detection task of the video to be tested, the difference between the video to be tested and the real video is calculated using the trained MMD depth kernel, which is defined as a difference value; and whether the video to be tested is generated by AI is determined based on the difference value.
[0053] In this embodiment, the difference value is:
[0054] (7);
[0055] Where, and Represent the real video and the video to be detected in the reference set respectively; Indicates the Normalized spatiotemporal gradient features of reference videos; Indicates the video to be detected Normalized spatiotemporal gradient characteristics; is a kernel function, for example , mapping the video features to a reproducing kernel Hilbert space .
[0056] In this embodiment, the detection process of the video to be tested is as follows:
[0057] (8);
[0058] Where, is a decision threshold (usually set to 1); for the video to be detected The maximum mean difference between the normalized spatiotemporal gradient features of the test video and the normalized spatiotemporal gradient features of the real video in the reference set is estimated. The larger the difference value, the greater the difference between the test video and the reference sample (real video). The function calculates the difference value by comparing With threshold The detection result is determined by the size relationship between : if it is greater than the threshold, the detection result is the generated video; otherwise, the detection result is the real video.
[0059] The following experimental data proves the effectiveness of the method proposed by the present invention. Figure 2 .
[0060] Figure 2 The experimental comparison chart of the proposed method in generated video detection applied to multiple generated video datasets shows that the proposed method (NSG-VD) outperforms traditional model-based methods (NPR, TALL, STIL, DeMamba) in detecting videos generated by various video generation models.
[0061] The above embodiments are intended only to illustrate the technical concepts and features of the present invention. Their purpose is to enable those skilled in the art to understand the contents of the present invention and implement them accordingly. They are not intended to limit the scope of protection of the present invention. Any equivalent changes or modifications made based on the essence of the present invention are intended to be covered by the scope of protection of the present invention.
Claims
1. A generative video detection method based on physics-driven spatiotemporal modeling, characterized in that: The following steps are involved: Step 1: According to the pre-defined physical-driven spatiotemporal modeling formula, the normalized spatiotemporal gradient features of the generated video and the real video are extracted, which are defined as the generated video features and the real video features respectively; Step 2: Establish a measurement model based on normalized spatiotemporal gradient feature differences, and train the depth kernel of the measurement model with the generated video features and the real video features. The optimization goal of the training process includes increasing the inter-class distance between the generated video features and the real video features, and reducing the intra-class distance of the real video features, to obtain the trained MMD depth kernel; Step 3: For the detection task of the video to be tested, the difference between the video to be tested and the real video is calculated using the trained MMD depth kernel, which is defined as a difference value; and whether the video to be tested is generated by AI is determined based on the difference value.
2. The method according to claim 1, characterized in that The physical driven spatiotemporal modeling formula in step 1 is: (1); Where, Indicates time Next video frame The joint probability density of Directly estimated by the score network of the pre-trained diffusion model, Displacement by adjacent frames approximate( ), Ensure the stability of the denominator value.
3. The method according to claim 1, characterized in that The depth kernel in step 2 is defined as: (2); Where, Indicates that the video is continuous Normalized spatiotemporal gradient feature sequence over frames; is a deep neural network mapping; and Both are bandwidth parameters and Gaussian kernel; is the balance coefficient; kernel parameter set To be learned.
4. The method according to claim 1, wherein The optimization objective in step 2 is to maximize the following expression: (3); Among them are (4); (5); (6); Where, and represent the real videos and generated videos in the training set respectively.
5. The method according to claim 1, characterized in that The difference value in step 3 is calculated as follows: (7); Where, and Represent the real video and the video to be detected in the reference set respectively; Indicates the Normalized spatiotemporal gradient features of reference videos; Indicates the video to be detected Normalized spatiotemporal gradient characteristics; is a kernel function, for example , mapping the video features to a reproducing kernel Hilbert space .
6. The method according to claim 1, characterized in that In step 3, the detection result is output through the judgment function: (8); Where, is a decision threshold (usually set to 1); for the video to be detected The maximum mean difference between the normalized spatiotemporal gradient features of the real video in the reference set and the normalized spatiotemporal gradient features of the real video in the reference set is estimated. This function calculates the difference value by comparing With threshold The detection result is determined by the size relationship: if it is greater than the threshold, the detection result is the generated video; Otherwise, the detection result is a real video.