Video authenticity detection method and device, electronic equipment, medium and program product
Patent Information
- Application Number
- CN202610800483.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-04
- Publication Date
- 2026-09-29
AI Technical Summary
[0007]本申请提供一种视频真实性检测方法、装置、电子设备、介质及程序产品,以解决相关技术中泛化能力差、抗干扰性弱、适配性不足和鲁棒性弱的问题,提升了视频检测的可靠性、实用性和针对性,具有良好的泛化能力
;
Smart Images

Figure CN122842007A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence video detection technology, and in particular to a video authenticity detection method, device, electronic device, medium and program product. Background Technology
[0002] With the rapid development of Large Video Generative Models (LVGMs), models such as OpenAI's Sora and Pika AI Video Generation Platform have fundamentally changed video synthesis technology. Unlike earlier generation techniques that required specialized skills and specific datasets, current LVGMs can directly generate realistic, temporally coherent, and semantically consistent videos using simple text or image cues, significantly improving visual fidelity and narrative complexity while lowering the technical barrier to media generation. However, this also brings potential risks. Realistic yet deceptive videos can be widely misused, causing harm in the spread of misinformation, fraud, and various digital deception scenarios. Therefore, detection technology for LVGM-generated content has become a critical requirement for ensuring digital security.
[0003] Early generative content detection methods mainly relied on model-specific artifacts. For image detection, they focused on statistical inconsistencies in spectral anomalies or sensor noise, while for video detection, they focused on temporal artifacts. However, these artifact-based methods cannot be adapted to modern LVGMs—models that achieve high semantic fidelity and temporal coherence through global and long-range attention mechanisms, eliminating the low-level periodic artifacts that traditional detection methods rely on.
[0004] In recent years, the detection paradigm has shifted to leverage the inherent characteristics of the generation process, mainly falling into two categories: semantic inversion and generative inversion. Semantic inversion methods are based on the assumption of an unusually strong alignment between generated content and reverse-engineered text prompts, using cosine similarity metrics embedded in CLIP (Contrastive Language-Image Pre-training). Generative inversion methods, on the other hand, utilize the fact that generated images are easier to reconstruct than real images, measuring error through an inversion-reconstruction loop. Currently, these image detection principles have been adapted to the video domain, but core limitations still exist.
[0005] In related technologies, video-generated content detection methods are still limited by their reliance on two-dimensional appearance features, resulting in weak generalization ability across generators, and two-dimensional artifacts are easily masked by real-world interference such as compression and scaling. Meanwhile, modern text-to-video models possess advanced creative reasoning capabilities, generating content that goes beyond literal cues, weakening the core assumptions of semantic alignment methods and reducing their long-term applicability.
[0006] In summary, the relevant technologies suffer from poor generalization ability, weak anti-interference ability, and insufficient adaptability. How to overcome the limitations of dependence on two-dimensional appearance features, achieve efficient and robust detection of LVGMs-generated content, and improve the ability to identify the authenticity of digital content has become the focus and challenge of current research and urgently needs to be solved. Summary of the Invention
[0007] This application provides a video authenticity detection method, apparatus, electronic device, medium, and program product to solve the problems of poor generalization ability, weak anti-interference ability, insufficient adaptability, and weak robustness in related technologies, thereby improving the reliability, practicality, and pertinence of video detection and having good generalization ability.
[0008] To achieve the above objectives, the first aspect of this application proposes a video authenticity detection method, comprising the following steps: The video to be detected is acquired, and the video to be detected is processed to obtain a dynamic three-dimensional point cloud sequence; Multi-scale geometric feature extraction is performed on the dynamic three-dimensional point cloud sequence to obtain a high-dimensional feature sequence with spatiotemporal structure; Global spatiotemporal relationship modeling is performed on the high-dimensional feature sequence to obtain modeled features, and global aggregation is performed on the modeled features to obtain global aggregation results. The authenticity detection result of the video to be detected is obtained based on the global aggregation results.
[0009] According to one embodiment of this application, the step of processing the video to be detected to obtain a dynamic three-dimensional point cloud sequence includes: Depth prediction is performed on each frame of the video to be detected to obtain an initial depth map for each frame; based on the initial depth map for each frame, an optimization parameter set is constructed, and with dense photometric loss as the objective function, the initial depth map for each frame is iteratively optimized by gradient descent through a sliding window to obtain an optimized depth map for each frame. A dynamic 3D point cloud sequence is obtained based on the optimized depth map corresponding to each frame of the image.
[0010] According to one embodiment of this application, obtaining a dynamic 3D point cloud sequence based on the optimized depth map corresponding to each frame image includes: The optimized depth map corresponding to each frame of the image is back-projected into three-dimensional points in the camera coordinate system to obtain the three-dimensional point cloud corresponding to each frame of the image in the camera coordinate system. Based on a preset transformation strategy, the initial 3D point cloud corresponding to each frame of the image is transformed to a unified world global coordinate system, so as to obtain the 3D point cloud corresponding to each frame of the image in the world coordinate system. The dynamic 3D point cloud sequence is generated based on the 3D point cloud corresponding to each frame of the image in the world coordinate system.
[0011] According to one embodiment of this application, the step of backprojecting the optimized depth map corresponding to each frame of image into three-dimensional points in the camera coordinate system to obtain the three-dimensional point cloud corresponding to each frame of image in the camera coordinate system includes: Based on a preset back-projection formula, the optimized depth map corresponding to each frame of the image is back-projected into 3D points in the camera coordinate system to obtain the 3D point cloud corresponding to each frame of the image in the camera coordinate system. The preset back-projection formula is as follows: ; in, For the first t 3D point coordinates in the frame camera coordinate system To optimize the pixel depth values of the depth map, For the first t Frame-optimized camera intrinsic parameter matrix u The horizontal pixel coordinates v These are the pixel coordinates in the vertical direction.
[0012] According to one embodiment of this application, the preset transformation strategy is: ; in, Let be the coordinates of a 3D point in the world coordinate system in frame t. For the first t Frame-optimized camera-to-world transformation matrix W Using the world coordinate system, C Let be the camera coordinate system.
[0013] According to one embodiment of this application, the step of performing global spatiotemporal relationship modeling on the high-dimensional feature sequence to obtain the modeled features includes: Based on a preset self-attention calculation formula, a global spatiotemporal relationship model is performed on the high-dimensional feature sequence to obtain the modeled features. The preset self-attention calculation formula is as follows: ; in, Q For query vector, K For key vectors, V For value vectors, for Q and K Dimensions This is the scaling factor.
[0014] The video authenticity detection method proposed in this application involves processing the video to be detected to obtain a dynamic 3D point cloud sequence, extracting multi-scale geometric features to obtain a high-dimensional feature sequence, performing global spatiotemporal relationship modeling to obtain modeled features, and then globally aggregating them. The authenticity detection result of the video to be detected is then obtained based on the global aggregation result. This solves the problems of poor generalization ability, weak anti-interference ability, insufficient adaptability, and weak robustness in related technologies, improving the reliability, practicality, and targeting of video detection, and demonstrating good generalization ability.
[0015] To achieve the above objectives, a second aspect of this application provides a video authenticity detection device, comprising: The acquisition module acquires the video to be detected and processes the video to be detected to obtain a dynamic three-dimensional point cloud sequence. The extraction module performs multi-scale geometric feature extraction on the dynamic three-dimensional point cloud sequence to obtain a high-dimensional feature sequence with spatiotemporal structure; The detection module performs global spatiotemporal relationship modeling on the high-dimensional feature sequence to obtain modeled features, performs global aggregation on the modeled features to obtain a global aggregation result, and obtains the authenticity detection result of the video to be detected based on the global aggregation result.
[0016] According to one embodiment of this application, the acquisition module is specifically used for: Depth prediction is performed on each frame of the video to be detected to obtain an initial depth map for each frame; based on the initial depth map for each frame, an optimization parameter set is constructed, and with dense photometric loss as the objective function, the initial depth map for each frame is iteratively optimized by gradient descent through a sliding window to obtain an optimized depth map for each frame. A dynamic 3D point cloud sequence is obtained based on the optimized depth map corresponding to each frame of the image.
[0017] According to one embodiment of this application, the acquisition module is specifically used for: The optimized depth map corresponding to each frame of the image is back-projected into three-dimensional points in the camera coordinate system to obtain the three-dimensional point cloud corresponding to each frame of the image in the camera coordinate system. Based on a preset transformation strategy, the initial 3D point cloud corresponding to each frame of the image is transformed to a unified world global coordinate system, so as to obtain the 3D point cloud corresponding to each frame of the image in the world coordinate system. The dynamic 3D point cloud sequence is generated based on the 3D point cloud corresponding to each frame of the image in the world coordinate system.
[0018] According to one embodiment of this application, the acquisition module is specifically used for: Based on a preset back-projection formula, the optimized depth map corresponding to each frame of the image is back-projected into 3D points in the camera coordinate system to obtain the 3D point cloud corresponding to each frame of the image in the camera coordinate system. The preset back-projection formula is as follows: ; in, For the first t 3D point coordinates in the frame camera coordinate system To optimize the pixel depth values of the depth map, For the first t Frame-optimized camera intrinsic parameter matrix u The horizontal pixel coordinates v These are the pixel coordinates in the vertical direction.
[0019] According to one embodiment of this application, the preset transformation strategy is: ; in, Let be the coordinates of a 3D point in the world coordinate system in frame t. For the first t Frame-optimized camera-to-world transformation matrix W Using the world coordinate system, C Let be the camera coordinate system.
[0020] According to one embodiment of this application, the detection module is specifically used for: Based on a preset self-attention calculation formula, a global spatiotemporal relationship model is performed on the high-dimensional feature sequence to obtain the modeled features. The preset self-attention calculation formula is as follows: ; in, Q For query vector, K For key vectors, V For value vectors, for Q and K Dimensions This is the scaling factor.
[0021] The video authenticity detection device proposed in this application processes the video to be detected to obtain a dynamic 3D point cloud sequence, extracts multi-scale geometric features to obtain a high-dimensional feature sequence, performs global spatiotemporal relationship modeling to obtain modeled features, performs global aggregation, and obtains the authenticity detection result of the video to be detected based on the global aggregation result. This solves the problems of poor generalization ability, weak anti-interference ability, insufficient adaptability, and weak robustness in related technologies, improves the reliability, practicality, and targeting of video detection, and has good generalization ability.
[0022] To achieve the above objectives, a third aspect of this application provides an electronic device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the video authenticity detection method as described in the above embodiments.
[0023] To achieve the above objectives, a fourth aspect of this application provides a computer-readable storage medium having a computer program stored thereon, which is executed by a processor to implement the video authenticity detection method as described in the above embodiments.
[0024] To achieve the above objectives, a fifth aspect of this application provides a computer program product, which, when executed by a processor, implements the video authenticity detection method as described in the above embodiments.
[0025] Additional aspects and advantages of this application will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of this application. Attached Figure Description
[0026] The above and / or additional aspects and advantages of this application will become apparent and readily understood from the following description of the embodiments taken in conjunction with the accompanying drawings, wherein: Figure 1 This is a flowchart of a video authenticity detection method provided according to an embodiment of this application; Figure 2 This is a flowchart of an iterative bundle adjustment method according to an embodiment of this application; Figure 3 This is a schematic diagram of a hybrid network architecture for three-dimensional spatiotemporal feature extraction and classification according to an embodiment of this application; Figure 4 This is a flowchart of a video authenticity detection method according to an embodiment of this application; Figure 5 This is a block diagram of a video authenticity detection device provided according to an embodiment of this application; Figure 6 This is a schematic diagram of the structure of an electronic device provided according to an embodiment of this application. Detailed Implementation
[0027] The embodiments of this application are described in detail below. Examples of these embodiments are shown in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain this application, and should not be construed as limiting this application.
[0028] The following description, with reference to the accompanying drawings, outlines a video authenticity detection method, apparatus, electronic device, medium, and program product based on embodiments of this application. Before introducing the video authenticity detection method of the embodiments of this application, let's briefly introduce the AI-generated video detection technology solution in related technologies.
[0029] Currently, AI-generated video detection technologies largely rely on 2D (Two-Dimensional) appearance artifact analysis, such as pixel domain statistical anomalies, temporal discontinuities, and rendering traces specific to the generation model. These methods have significant technical drawbacks: First, they exhibit strong model specificity and poor cross-generator generalization ability; 2D artifact features learned by a particular generation model are almost ineffective against novel, unseen video generation models. Second, they lack robustness and are easily affected by disturbances in the real-world scene. Compression, additive noise, and resolution adjustments during video transmission can easily mask shallow 2D appearance artifacts, leading to a significant drop in detection accuracy. Third, they fail to address the core defects of generated videos. While existing large video generation models (LVGMs) can synthesize visually realistic 2D videos, they do not explicitly model the 3D geometric structure and physical causal relationships of the real world, resulting in systematic 3D (Three-Dimensional) spatiotemporal geometric inconsistencies in the generated videos. Therefore, the relevant technologies, which only capture the generation defects at the two-dimensional level, have core problems such as poor generalization ability across generation models and susceptibility to actual interference from video compression, additive noise, watermarks, etc. The essential defect of generated videos lies in the fact that their generation process does not explicitly model the three-dimensional structure and physical causal relationship of the real world, resulting in inconsistencies in their three-dimensional structure and spatiotemporal dynamic changes. Moreover, such inconsistencies will not disappear due to two-dimensional disturbances.
[0030] Based on the above problems, and because real videos are two-dimensional projections of three-dimensional physical scenes onto a camera imaging model, the entire process is subject to strict geometric and physical constraints, and the corresponding three-dimensional structure has spatiotemporal consistency; while large video generation models synthesize videos through a high-dimensional latent space diffusion process, lacking a modeling process of the real three-dimensional world, and cannot maintain the spatiotemporal consistency of the three-dimensional structure, often exhibiting three-dimensional defects such as depth instability, abnormal occlusion relationships, and inconsistent rigidity. This application's embodiments utilize an end-to-end technical process to reconstruct a 3D structure from a 2D video. Then, through the analysis and discrimination of 3D spatiotemporal features, high-precision detection of the generated video is achieved. This provides a novel technical path to address the technical pain points of traditional 2D detection methods: starting from the physical formation principle of real videos, it addresses the core pain point of 3D geometric modeling defects in large video generation models. The detection logic is upgraded from "2D appearance artifact matching" to "3D spatiotemporal geometric consistency verification," proposing a 3D spatiotemporal detector (3D-STD, 3D Spacetime Transformer Decoder). This detector extracts 3D spatiotemporal anomaly features through dynamic 3D scene reconstruction of uncalibrated videos, achieving high-precision detection of AI-generated videos. Simultaneously, it solves the technical problems of poor cross-generator generalization and weak anti-perturbation ability in traditional methods. Furthermore, the detection process is executed end-to-end, requiring no additional information such as camera calibration or scene priors, thus adapting to the unconstrained video detection needs in real-world scenarios.
[0031] Specifically, Figure 1 This is a flowchart of a video authenticity detection method according to an embodiment of this application.
[0032] like Figure 1 As shown, the video authenticity detection method includes the following steps: In step S101, the video to be detected is acquired and processed to obtain a dynamic three-dimensional point cloud sequence.
[0033] Among them, dynamic 3D point cloud sequence refers to an ordered set of 3D point clouds that continuously change over time in a unified world coordinate system after 3D reconstruction of an uncalibrated original 2D video.
[0034] Specifically, in this embodiment, the video to be detected is acquired, and the video is processed frame by frame and the camera and scene geometric parameters are jointly optimized to restore the three-dimensional geometric structure corresponding to the video. The optimized geometric parameters are then converted into a three-dimensional point cloud sequence in a unified world coordinate system to achieve a standardized representation of the three-dimensional structure.
[0035] Furthermore, in some embodiments, processing the video to be detected to obtain a dynamic 3D point cloud sequence includes: performing depth prediction on each frame of the video to be detected to obtain an initial depth map corresponding to each frame; constructing an optimization parameter set based on the initial depth map corresponding to each frame, and using dense photometric loss as the objective function, iteratively optimizing the initial depth map corresponding to each frame through gradient descent with a sliding window to obtain an optimized depth map corresponding to each frame; and obtaining a dynamic 3D point cloud sequence based on the optimized depth map corresponding to each frame.
[0036] The initial depth map refers to the intermediate relative depth data output after depth prediction is completed for each frame of the video image. Dense luminance loss refers to a robust objective function constructed with full-pixel-level luminance consistency as its core.
[0037] Specifically, this embodiment processes the input video sequence frame by frame, using a monocular depth estimation model to perform dense depth prediction on each frame, obtaining an initial depth map for each frame, laying the foundation for subsequent accurate reconstruction. Furthermore, to eliminate the scale ambiguity problem of monocular depth estimation, the 3D geometric structure information corresponding to the video is established. Based on the Bundle Adjustment method, camera intrinsic parameters, camera pose, and scene depth are jointly optimized. Through pixel-level photometric consistency constraints, all geometric parameters are iteratively corrected to obtain a 3D geometric reconstruction result with consistent metric (real physical scale attribute). Even further, to achieve standardized representation of the 3D structure between video frames and facilitate subsequent spatiotemporal consistency analysis, the optimized depth information and camera parameters are subjected to coordinate transformation. First, it is back-projected into a 3D point cloud in the camera coordinate system, and then converted into a 3D point cloud sequence in a unified world coordinate system, ensuring the spatiotemporal correlation of the 3D structure between frames.
[0038] Furthermore, monocular depth estimation is a crucial step in recovering 3D depth information from a 2D image. This application embodiment selects the MiDaS monocular depth estimation model (a type of monocular depth estimation model) based on a visual Transformer. This model has cross-scene depth prediction capabilities and can perform high-precision dense depth prediction on video frames of different scenes and resolutions. This application embodiment uses the input video sequence... Perform a frame-by-frame traversal, where N is the total number of video frames. For the first The frame's RGB (Red-Green-Blue) image, , These represent the image height and width, respectively, with 3 representing the number of RGB channels. Each frame is input into a monocular depth estimation model for dense depth prediction. Through feature extraction and depth regression, the initial depth map corresponding to that frame is obtained, i.e., a depth mapping relationship is established using the MiDaS model. ,in This is the initial depth map for this frame. As a relative depth map, the pixel values represent inverse depth (parallax) and are normalized within a specified range. They have scale and offset ambiguity and cannot directly represent the real 3D geometry, but can provide a reliable geometric prior for subsequent joint optimization. The optimization process will be based on this prior to obtain the real depth information.
[0039] Furthermore, in this embodiment, based on the initial depth map of a single frame, a ternary optimization parameter set is constructed, which includes camera intrinsic parameters K, camera pose P, and scene dense depth D. With dense photometric loss as the objective function, the parameter set is jointly optimized through a sliding window + gradient descent iteration method to solve the scale ambiguity problem of monocular depth estimation and obtain metric-consistent 3D geometric reconstruction parameters. This process is a pixel-level dense optimization and is highly sensitive to subtle 3D geometric inconsistencies in the generated video.
[0040] Specifically, such as Figure 2 As shown, Figure 2 This is a flowchart of an iterative beam adjustment method according to an embodiment of this application. The iterative beam adjustment method includes the following steps: applying a length of [missing information] to the input video frame sequence. sliding window Perform local frame selection, starting with the first frame of the window. Construct a local 3D scene based on the reference frame; based on the currently estimated camera parameters. K , pose P With depth map D The reference frame pixels are back-projected into 3D space through 3D scene mapping, and then a virtual observation image is generated through 2D reprojection. The Charbonnier loss function is further used to iteratively optimize camera parameters, pose and depth estimation by using the photometric error between the reprojected image and the original frame in the window as a constraint. Global parameter optimization of the video sequence is achieved through progressive iteration of the sliding window over the entire video sequence.
[0041] This application embodiment performs camera intrinsic parameter modeling. Camera intrinsic parameters characterize the inherent optical properties of the camera. To adapt to possible abnormal changes in intrinsic parameters in the generated video, the intrinsic parameters are considered as optimizable variables for each frame. For the t-th frame of the video, a 4-parameter vector is used. Indicates the camera's focal length ( , ) and principal point coordinates ( , ), and construct the intrinsic parameter matrix. : ; in, and The camera is in , Focal length of direction, and These are the coordinates of the principal point on the camera's imaging plane.
[0042] Furthermore, in this embodiment, the camera intrinsic parameters are regarded as optimizable variables for each frame to adapt to the problem of abnormal changes in intrinsic parameters that may exist in AI-generated videos. The initial value of the intrinsic parameter matrix is set to a reasonable default value according to the image resolution, and then iteratively corrected through the optimization process.
[0043] Furthermore, for the t-th frame of the video, this embodiment uses a 6-dimensional vector from the Lie algebra se(3). Let represent the camera pose, and its transformation relationship with the rigid body transformation matrix is as follows: ; in, The translation component represents the change in the camera's position in three-dimensional space. Represents the rotation component, characterizing the camera's orientation change in the form of an axial angle; The hat operator maps a 6-dimensional Lie algebra vector to the corresponding 4×4 homogeneous transformation matrix. The rigid body transformation matrix is used to transform the camera coordinate system to the world coordinate system. To eliminate global coordinate system blur, the pose of the first frame of the video is fixed as the identity matrix, and the poses of subsequent frames are optimized based on this frame.
[0044] Furthermore, the embodiments of this application address the video... Frame, with initial depth map Using the initial values, construct a frame-level dense depth map. By treating the inverse depth of each pixel as an independent optimization variable, pixel-level fine-grained depth optimization is achieved. To balance 3D reconstruction accuracy and computational efficiency, a sliding window strategy is employed to perform block optimization on video frames, with the window size adjusted by setting a beam. Regarding the video number Frame, defines the current optimization window That is, only for the first Frame to the The local frame sequence of the frame is jointly optimized; the depth, intrinsic parameters, and pose of all frames within the window are integrated into a unified parameter set. The pose of the first frame of the window is fixed as the anchor point to ensure the consistency of 3D geometric reconstruction within the window.
[0045] Furthermore, in this embodiment, the Charbonnier robust loss is used as a basis to calculate the loss between any two frames within the window. The pixel-level photometric error is primarily determined by the geometric transformation process of "backprojection-coordinate transformation-reprojection." A photometric loss function is constructed by comparing the pixel differences between the real frame and the reprojected predicted frame. Specifically, for each frame... any pixel Through the current internal reference and depth 2D pixels are back-projected onto the camera 3D points in coordinate system : ; in, For pixels In frame Depth values in a depth map For frames The inverse of the camera intrinsic matrix, which projects two-dimensional pixel coordinates into three-dimensional spatial points along the camera's line of sight.
[0046] Furthermore, to achieve coordinate system transformation of 3D points between different frames, the camera... Converting 3D points in a coordinate system to a camera The transformation relationship for a 3D point in a coordinate system is as follows: ; in, For camera To the camera The relative transformation matrix is derived from the world coordinate system transformation matrix of the two frames.
[0047] Furthermore, 3D points Through the camera Internal reference Reprojection to frame From the image plane, obtain the homogeneous pixel coordinates. 2D pixel coordinates are obtained through perspective division. : ; This application embodiment aligns the sub-pixel coordinates. Perspective normalization is performed to obtain two-dimensional pixel coordinates: ; Furthermore, to achieve joint optimization of camera intrinsics, pose, and scene depth, a photometric consistency loss function based on Charbonnier robust loss is constructed. This function uses the pixel difference between the real frame and the reprojected predicted frame as the optimization objective, calculating the frame... Pixels With frames Reprojected pixels Charbonnier loss ( (To be a very small constant, avoiding values within the square root of 0), the total luminous loss is obtained by summing the luminous errors of all valid pixels within the window. : ; in, To optimize the set of frame pairs within the sliding window, For frames Projecting from the middle to the frame The effective set of pixels within the image boundary The loss function is Charbonnier, which is robust and can effectively resist interference such as changes in illumination and local occlusion.
[0048] Furthermore, in this embodiment, a gradient descent optimizer (such as Adam (Adaptive Moment Estimation)) is used to iteratively optimize the photometric consistency loss function, continuously correcting the camera intrinsic parameters, camera pose, and scene depth parameters until the loss converges or a preset number of iterations is reached, thus obtaining the optimized camera intrinsic parameters. Camera pose and scene depth map The updated formula is: ; in, The learning rate controls the step size for parameter updates; the number of iterations is set. Repeat the process of "photometric loss calculation - gradient solution - parameter update" until the loss converges or the upper limit of the number of iterations is reached, to obtain the optimized camera intrinsic parameters for each frame within the window. Pose matrix Dense depth map .
[0049] Furthermore, the video frame sequence is input into the sliding window frame by frame in chronological order, and the operation of "sliding window construction - dense photometric loss calculation - gradient descent iterative optimization" is repeated to complete the 3D geometric parameter optimization of all video frames, and output the optimized camera intrinsic parameters, pose, and depth parameters of the entire sequence.
[0050] Furthermore, in some embodiments, obtaining a dynamic 3D point cloud sequence based on the optimized depth map corresponding to each frame image includes: backprojecting the optimized depth map corresponding to each frame image into 3D points in the camera coordinate system to obtain a 3D point cloud corresponding to each frame image in the camera coordinate system; transforming the initial 3D point cloud corresponding to each frame image to a unified world global coordinate system based on a preset transformation strategy to obtain a 3D point cloud corresponding to each frame image in the world coordinate system; and generating a dynamic 3D point cloud sequence based on the 3D point cloud corresponding to each frame image in the world coordinate system.
[0051] Optionally, in some embodiments, backprojecting the optimized depth map corresponding to each frame of an image into 3D points in the camera coordinate system to obtain a 3D point cloud corresponding to each frame of an image in the camera coordinate system includes: based on a preset backprojection formula, backprojecting the optimized depth map corresponding to each frame of an image into 3D points in the camera coordinate system to obtain a 3D point cloud corresponding to each frame of an image in the camera coordinate system, wherein the preset backprojection formula is: ; in, For the first t 3D point coordinates in the frame camera coordinate system To optimize the pixel depth values of the depth map, For the first t Frame-optimized camera intrinsic parameter matrix u The horizontal pixel coordinates v These are the pixel coordinates in the vertical direction.
[0052] Optionally, in some embodiments, the preset transformation strategy is: ; in, For the first t 3D point coordinates in the frame world coordinate system For the first t Frame-optimized camera-to-world transformation matrix W Using the world coordinate system, C Let be the camera coordinate system.
[0053] In this context, a 3D point cloud refers to a set of discrete points representing the 3D geometric structure of a scene, obtained by reconstructing a single 2D image. The preset transformation strategy can be a user-defined strategy, a strategy obtained through a finite number of experiments, or a strategy obtained through a finite number of computer simulations.
[0054] Specifically, in this embodiment, based on the optimized 3D geometric parameters obtained by iterative beam adjustment, the 2D image of each frame is back-projected into a 3D point cloud in a unified world coordinate system to generate a dynamic 3D point cloud sequence.
[0055] Furthermore, in this embodiment of the application, a three-dimensional point cloud back projection is performed in the camera coordinate system. For the video... t Frames, based on optimized depth maps and internal reference For each pixel with a valid depth value Back projection yields the camera 3D points in coordinate system The three-dimensional points corresponding to all valid pixels constitute a three-dimensional point cloud in the camera's t-coordinate system. .
[0056] Furthermore, in order to achieve a unified representation of point clouds across all frames, an optimized pose transformation matrix is used. Mapping the 3D point cloud from the camera coordinate system to the world coordinate system yields... All the transformed 3D points constitute a point cloud in the world coordinate system. This application embodiment integrates the 3D point clouds in the world coordinate system of each frame to finally obtain a 3D point cloud sequence in a unified world coordinate system: This sequence preserves the three-dimensional geometric structure of the video and the spatiotemporal dynamic relationship between frames, providing standardized input for subsequent feature extraction. The 3D spatiotemporal geometric anomalies of the generated video will manifest as features such as structural inconsistency, depth shift, and occlusion anomalies in this point cloud sequence.
[0057] In step S102, multi-scale geometric feature extraction is performed on the dynamic three-dimensional point cloud sequence to obtain a high-dimensional feature sequence with spatiotemporal structure.
[0058] Among them, the high-dimensional feature sequence refers to the temporal feature vector sequence containing the three-dimensional spatiotemporal structure information of the scene, which is formed by aggregating and splicing the multi-scale geometric features of each frame of the dynamic three-dimensional point cloud sequence in time order.
[0059] Specifically, in order to accurately capture the geometric structural features of 3D point clouds, this application embodiment uses the PointNet++ network (hierarchical point cloud feature extraction network) to process the 3D point cloud sequence in a unified coordinate system. Through a hierarchical feature extraction process, multi-scale geometric features of the point cloud are extracted, which not only preserves the fine local geometric details, but also captures the global structural features, and outputs a high-dimensional feature sequence with spatiotemporal structure.
[0060] Furthermore, since 3D point clouds are unordered and unstructured data, traditional convolutional neural networks cannot process them directly. This application uses PointNet++ as the backbone network to perform hierarchical multi-scale feature extraction on dynamic 3D point cloud sequences. The core achieves point cloud downsampling and feature upscaling through a Set Abstraction (SA) layer, while preserving local details and global structural features of the 3D geometry. The final SA layer does not perform global pooling, thus preserving the spatiotemporal sequence structure of the features. Specifically, this application divides the 3D point cloud sequence in a unified world coordinate system into batch inputs according to time windows. ,in For batch size, The number of points in a single frame of the point cloud. The feature dimension of a point includes 3D coordinates and color features. This application uses the PointNet++ network as the backbone for 3D feature extraction. This network is specifically designed for unordered 3D point clouds and achieves multi-scale geometric feature extraction through a hierarchical set abstraction (SA) layer. The core process includes farthest point sampling, local neighborhood grouping, and local feature learning. Farthest Point Sampling (FPS) refers to sampling the centroid set from the input point cloud according to the principle of the farthest spatial distance. This reduces the number of point clouds and computational cost while ensuring the spatial coverage of the centroid set, guaranteeing the comprehensiveness of subsequent feature extraction. Local neighborhood grouping can be sphere query grouping. Using each centroid as the center of a sphere and setting a fixed radius, a sphere query is performed to divide the neighboring points around the centroid into a local point cloud group, achieving the division of local 3D geometric structures. Local feature learning refers to extracting the relative coordinates and feature information of points within each local point cloud group using a mini-PointNet, and then obtaining the aggregated features of the centroid through max pooling, achieving the invariance of the unordered point cloud arrangement while simultaneously improving the feature dimension.
[0061] Furthermore, the first two SA layers reduce the number of centroids and continuously increase the feature dimension through downsampling, while the third SA layer does not perform global pooling, outputting a high-dimensional spatiotemporal feature sequence containing 64 centroids. ,in As the final feature dimension, this feature sequence serves as a spatiotemporally salient descriptor for the 3D point cloud, preserving the 3D geometric structure and inter-frame spatiotemporal relationships of the point cloud, and is used as input for subsequent encoders. Specifically, through the concatenation of multiple ensemble abstraction layers, multi-scale geometric feature extraction from local to global perspectives is achieved. Pre-sequence layers capture fine local geometric details, while post-sequence layers capture global 3D structural features, ultimately outputting a high-dimensional feature sequence that preserves spatiotemporal structure. .
[0062] In step S103, global spatiotemporal relationship modeling is performed on the high-dimensional feature sequence to obtain the modeled features, and global aggregation is performed on the modeled features to obtain the global aggregation result. The authenticity detection result of the video to be detected is obtained based on the global aggregation result.
[0063] Optionally, in some embodiments, global spatiotemporal relationship modeling is performed on the high-dimensional feature sequence to obtain the modeled features, including: global spatiotemporal relationship modeling is performed on the high-dimensional feature sequence based on a preset self-attention calculation formula to obtain the modeled features, wherein the preset self-attention calculation formula is: ; in, , , These are the query vector, key vector, and value vector, respectively, obtained from the input feature sequence through a linear transformation. for and Dimensions The scaling factor is used to avoid the vanishing gradient problem of softmax (Softmax Function, normalized exponential function) caused by excessively large vector inner products; the softmax function is used to calculate the attention weights between features to capture long-range spatiotemporal dependencies between frames.
[0064] Global aggregation refers to the feature integration operation that performs mean pooling along the temporal dimension on a high-dimensional feature sequence, fusing the temporal features into a single global feature vector.
[0065] Specifically, in order to analyze the spatiotemporal consistency of 3D features and achieve detection and discrimination, this embodiment of the application utilizes a Transformer network to model the global spatiotemporal relationship of high-dimensional feature sequences, captures long-range spatiotemporal dependencies between frames through a self-attention mechanism, then globally aggregates the modeled features, inputs them into a multilayer perceptron to complete binary classification and discrimination, and finally outputs the video authenticity detection result. Further, this embodiment of the application uses the high-dimensional feature sequences output by the PointNet++ network... The input token is directly used as the input token of the Transformer network. The 3D spatial location information and inter-frame spatiotemporal information of the point cloud are fully integrated into the features, eliminating the need for additional position encoding. In this embodiment, the input token undergoes a linear transformation to obtain the query vector. Q Key vector K, value vector VThis paper calculates attention weights between tokens using a scaled dot product self-attention mechanism, capturing multi-dimensional long-range dependencies between features. Through parallel computation of multi-head self-attention, it fuses multi-dimensional spatiotemporal correlation features to achieve comprehensive analysis of 3D structure and dynamic changes. Furthermore, in this embodiment, the output of multi-head self-attention is fed into a layer normalization and feed-forward network (FFN) to form a single Transformer encoder layer. By stacking multiple encoder layers, deep spatiotemporal context reasoning is performed on the feature sequence, fully modeling the 3D geometric change relationships between frames to obtain a context-aware discriminative feature sequence that fuses long-range spatiotemporal correlations. This process can accurately capture 3D geometric instabilities across frames in the generated video, such as abrupt changes in object depth, abnormal occlusion relationships, and inconsistent rigidity, which are long-range spatiotemporal anomalies and are a core component for AI-generated video detection.
[0066] Furthermore, in this embodiment, the context-aware discriminative feature sequence output by the Transformer encoder is globally aggregated, transforming high-dimensional features into a single global feature vector. Then, an MLP (Multi-Layer Perceptron) classifier is used to achieve binary classification of real videos and AI-generated videos. Specifically, mean pooling is performed along the sequence dimension on the modeled feature sequence output by the Transformer model, aggregating the sequence features into a single global feature vector that integrates three-dimensional geometry and spatiotemporal correlation information within the entire time window. The global feature vector is input to an MLP classifier composed of fully connected layers and activation functions. The classifier adopts a structure of "fully connected layer + ReLU (Rectified Linear Unit) activation layer + dropout layer + single-output fully connected layer," where the last layer is a single-output node, outputting a binary classification logit (Logistic Unit) value. The logit value output by the MLP classifier is mapped to a sigmoid function. The system sets a classification threshold for each interval. When the output value is greater than the threshold, the input video is determined to be an AI-generated video (positive class); when the output value is less than the threshold, the input video is determined to be a real video (negative class). The detection results of all time windows of the video are fused to output the final video detection result.
[0067] like Figure 3 As shown, Figure 3This is a schematic diagram of a hybrid network architecture for 3D spatiotemporal feature extraction and classification according to an embodiment of this application. The dynamic 3D point cloud sequence is divided into input samples according to time windows. First, the disordered 3D point cloud is subjected to hierarchical multi-scale feature extraction through the PointNet++ backbone network to obtain a high-dimensional feature sequence that preserves the spatiotemporal structure. Then, the feature sequence is input into the Transformer encoder, and the long-range spatiotemporal context dependency is modeled through a multi-head self-attention mechanism to obtain discriminative features that integrate 3D geometry and spatiotemporal information. Finally, the feature sequence is globally aggregated and input into the MLP classifier to achieve binary classification of real video and AI-generated video, and the final detection result is output. This hybrid architecture can accurately capture subtle and systematic 3D spatiotemporal geometric anomalies in generated videos.
[0068] Therefore, the embodiments of this application break through the two-dimensional limitations of traditional generated video detection, and perform detection and analysis from the essential perspective of three-dimensional geometric consistency, rather than two-dimensional surface artifacts, which greatly improves the essentiality and reliability of the detection method and fundamentally solves the technical pain points of traditional methods; it realizes uncalibrated end-to-end three-dimensional geometric reconstruction without the need for any prior information such as camera intrinsics and pose, and can directly process unconstrained raw videos, making it suitable for real and complex application scenarios, thus improving the practicality and scenario adaptability of the method; through the combination of PointNet++ and Transformer, it realizes local three-dimensional geometric features and global... The joint modeling of spatiotemporal correlation features not only accurately captures the multi-scale geometric details of point clouds but also effectively analyzes the three-dimensional spatiotemporal consistency between frames, making detection and discrimination more targeted. It has strong robustness to common interferences in practical applications such as video compression, additive Gaussian noise, and text watermarking, and the three-dimensional geometric consistency features will not disappear due to perturbations at the two-dimensional pixel level. It has good generalization ability to unknown new video generation models. Most generated videos lack the ability to model three dimensions. The detection signals captured by the embodiments of this application have cross-model universality, can adapt to the development of new generation technologies, and avoid the dependence of traditional methods on the specificity of generation models.
[0069] To facilitate those skilled in the art to further understand the video authenticity detection method proposed in the embodiments of this application, the following is combined with... Figure 4 Further explanation is needed.
[0070] like Figure 4 As shown, Figure 4This is a flowchart of a video authenticity detection method according to an embodiment of this application. The video authenticity detection method includes the following steps: First, through a dynamic 3D scene mapping module, a 2D video frame sequence is converted into a dynamic 3D point cloud sequence in a global coordinate system, completing the feature dimensionality upgrade from "2D appearance" to "3D geometry"; then, through a spatiotemporal feature extraction and classification module, PointNet++ is used to capture the local and global geometric features of the 3D point cloud, and a Transformer encoder is used to model the long-range spatiotemporal dependencies between frames; finally, a classifier is used to achieve accurate detection of AI-generated videos. It can achieve high-precision, high-generalization, and high-robust detection of AI-generated videos under unconstrained conditions without camera calibration or scene priors. It is applicable to T2V (Text to Video) / I2V (Image to Video) generated video detection of mainstream large video generation models such as OpenSora (Open-Sora Framework, an open-source Sora-like video generation framework), Sora, and Pika. The video authenticity detection method proposed in this application can be implemented based on deep learning frameworks such as PyTorch (Torch deep learning framework), and can be deployed on computing power platforms such as servers. The detection process supports end-to-end processing, and can detect input uncalibrated videos and output real / generated judgment results. It is suitable for various AI-generated video detection application scenarios such as digital content authentication, false information tracing, and media content review.
[0071] It should be noted that the specific implementation examples described in the embodiments of this application are merely illustrative of the methods and steps of this application. Those skilled in the art can make corresponding modifications, additions, or variations to the described specific implementation steps (i.e., adopt similar alternative methods), but without departing from the principles and essence of this application or exceeding the scope defined by the appended claims. The scope of this application is limited only by the appended claims.
[0072] The video authenticity detection method proposed in this application involves processing the video to be detected to obtain a dynamic 3D point cloud sequence, extracting multi-scale geometric features to obtain a high-dimensional feature sequence, performing global spatiotemporal relationship modeling to obtain modeled features, and then globally aggregating them. The authenticity detection result of the video to be detected is then obtained based on the global aggregation result. This solves the problems of poor generalization ability, weak anti-interference ability, insufficient adaptability, and weak robustness in related technologies, improving the reliability, practicality, and targeting of video detection, and demonstrating good generalization ability.
[0073] Next, the video authenticity detection device proposed according to the embodiments of this application is described with reference to the accompanying drawings.
[0074] Figure 5This is a block diagram of a video authenticity detection device according to an embodiment of this application.
[0075] like Figure 5 As shown, the video authenticity detection device 10 includes: an acquisition module 100, an extraction module 200, and a detection module 300, wherein, The acquisition module 100 acquires the video to be detected and processes it to obtain a dynamic 3D point cloud sequence. The extraction module 200 performs multi-scale geometric feature extraction on the dynamic 3D point cloud sequence to obtain a high-dimensional feature sequence with spatiotemporal structure. The detection module 300 performs global spatiotemporal relationship modeling on the high-dimensional feature sequence to obtain the modeled features, performs global aggregation on the modeled features to obtain the global aggregation result, and obtains the authenticity detection result of the video to be detected based on the global aggregation result.
[0076] According to one embodiment of this application, the acquisition module 100 is specifically used for: Depth prediction is performed on each frame of the video to be detected to obtain an initial depth map for each frame. Based on the initial depth map for each frame, an optimization parameter set is constructed, and the initial depth map for each frame is iteratively optimized by gradient descent through a sliding window with dense photon loss as the objective function to obtain the optimized depth map for each frame. A dynamic 3D point cloud sequence is obtained based on the optimized depth map corresponding to each frame of the image.
[0077] According to one embodiment of this application, the acquisition module 100 is specifically used for: The optimized depth map corresponding to each frame of the image is back-projected into three-dimensional points in the camera coordinate system to obtain the three-dimensional point cloud corresponding to each frame of the image in the camera coordinate system. Based on a preset transformation strategy, the initial 3D point cloud corresponding to each frame of the image is transformed to a unified world global coordinate system, so as to obtain the 3D point cloud corresponding to each frame of the image in the world coordinate system. A dynamic 3D point cloud sequence is generated based on the 3D point cloud corresponding to each frame of the image in the world coordinate system.
[0078] According to one embodiment of this application, the acquisition module 100 is specifically used for: Based on a preset backprojection formula, the optimized depth map corresponding to each frame of the image is backprojected into 3D points in the camera coordinate system to obtain the 3D point cloud corresponding to each frame of the image in the camera coordinate system. The preset backprojection formula is as follows: ; in, For the first t 3D point coordinates in the frame camera coordinate system To optimize the pixel depth values of the depth map, For the first t Frame-optimized camera intrinsic parameter matrix u The horizontal pixel coordinates v These are the pixel coordinates in the vertical direction.
[0079] According to one embodiment of this application, the preset transformation strategy is as follows: ; in, Let be the coordinates of a 3D point in the world coordinate system in frame t. For the first t Frame-optimized camera-to-world transformation matrix W Using the world coordinate system, C Let be the camera coordinate system.
[0080] According to one embodiment of this application, the detection module 300 is specifically used for: Based on a pre-defined self-attention calculation formula, a global spatiotemporal relationship model is performed on the high-dimensional feature sequence to obtain the modeled features. The pre-defined self-attention calculation formula is as follows: ; in, Q For query vector, K For key vectors, V For value vectors, for Q and K Dimensions This is the scaling factor.
[0081] It should be noted that the foregoing explanation of the video authenticity detection method embodiment also applies to the video authenticity detection device of this embodiment, and will not be repeated here.
[0082] The video authenticity detection device proposed in this application processes the video to be detected to obtain a dynamic 3D point cloud sequence, extracts multi-scale geometric features to obtain a high-dimensional feature sequence, performs global spatiotemporal relationship modeling to obtain modeled features, performs global aggregation, and obtains the authenticity detection result of the video to be detected based on the global aggregation result. This solves the problems of poor generalization ability, weak anti-interference ability, insufficient adaptability, and weak robustness in related technologies, improves the reliability, practicality, and targeting of video detection, and has good generalization ability.
[0083] Figure 6 This is a schematic diagram of an electronic device provided in an embodiment of the present invention. The electronic device may include: The memory 601, the processor 602, and the computer program stored on the memory 601 and capable of running on the processor 602.
[0084] When the processor 602 executes the program, it implements the video authenticity detection method provided in the above embodiments.
[0085] Furthermore, electronic devices also include: Communication interface 603 is used for communication between memory 601 and processor 602.
[0086] The memory 601 is used to store computer programs that can run on the processor 602.
[0087] The memory 601 may include high-speed RAM (Random Access Memory) memory, and may also include non-volatile memory, such as at least one disk storage.
[0088] If the memory 601, processor 602, and communication interface 603 are implemented independently, then the communication interface 603, memory 601, and processor 602 can be interconnected via a bus to complete communication between them. The bus can be an ISA (Industry Standard Architecture) bus, a PCI (Peripheral Component Interconnect) bus, or an EISA (Extended Industry Standard Architecture) bus, etc. The bus can be divided into address bus, data bus, control bus, etc. For ease of representation, Figure 6 The bus is represented by a single thick line, but this does not mean that there is only one bus or one type of bus.
[0089] Optionally, in a specific implementation, if the memory 601, processor 602, and communication interface 603 are integrated on a single chip, then the memory 601, processor 602, and communication interface 603 can communicate with each other through an internal interface.
[0090] The processor 602 may be a CPU (Central Processing Unit), an ASIC (Application Specific Integrated Circuit), or one or more integrated circuits configured to implement embodiments of the present invention.
[0091] This invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the video authenticity detection method described above.
[0092] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the steps in any of the above-described video authenticity detection method embodiments.
[0093] Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of technical features indicated. Thus, a feature defined as "first" or "second" may explicitly or implicitly include at least one of that feature. In the description of this application, "multiple" means at least two, such as two, three, etc., unless otherwise explicitly specified.
[0094] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., refer to specific features, structures, materials, or characteristics described in connection with that embodiment or example, which are included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.
[0095] Although embodiments of this application have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting this application. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of this application.
Claims
1. A method for detecting the authenticity of a video, characterized in that, include: The video to be detected is acquired, and the video to be detected is processed to obtain a dynamic three-dimensional point cloud sequence; Multi-scale geometric feature extraction is performed on the dynamic three-dimensional point cloud sequence to obtain a high-dimensional feature sequence with spatiotemporal structure; Global spatiotemporal relationship modeling is performed on the high-dimensional feature sequence to obtain modeled features, and global aggregation is performed on the modeled features to obtain global aggregation results. The authenticity detection result of the video to be detected is obtained based on the global aggregation results.
2. The method according to claim 1, characterized in that, The process of processing the video to be detected to obtain a dynamic 3D point cloud sequence includes: Depth prediction is performed on each frame of the video to be detected to obtain an initial depth map for each frame; based on the initial depth map for each frame, an optimization parameter set is constructed, and with dense photometric loss as the objective function, the initial depth map for each frame is iteratively optimized by gradient descent through a sliding window to obtain an optimized depth map for each frame. A dynamic 3D point cloud sequence is obtained based on the optimized depth map corresponding to each frame of the image.
3. The method according to claim 2, characterized in that, The step of obtaining a dynamic 3D point cloud sequence based on the optimized depth map corresponding to each frame image includes: The optimized depth map corresponding to each frame of the image is back-projected into three-dimensional points in the camera coordinate system to obtain the three-dimensional point cloud corresponding to each frame of the image in the camera coordinate system. Based on a preset transformation strategy, the initial 3D point cloud corresponding to each frame of the image is transformed to a unified world global coordinate system, so as to obtain the 3D point cloud corresponding to each frame of the image in the world coordinate system. The dynamic 3D point cloud sequence is generated based on the 3D point cloud corresponding to each frame of the image in the world coordinate system.
4. The method according to claim 3, characterized in that, The step of backprojecting the optimized depth map corresponding to each frame of image into 3D points in the camera coordinate system to obtain the 3D point cloud corresponding to each frame of image in the camera coordinate system includes: Based on a preset back-projection formula, the optimized depth map corresponding to each frame of the image is back-projected into 3D points in the camera coordinate system to obtain the 3D point cloud corresponding to each frame of the image in the camera coordinate system. The preset back-projection formula is as follows: ; in, For the first t 3D point coordinates in the frame camera coordinate system To optimize the pixel depth values of the depth map, For the first t Frame-optimized camera intrinsic parameter matrix u The horizontal pixel coordinates are... v These are the pixel coordinates in the vertical direction.
5. The method according to claim 3, characterized in that, The preset transformation strategy is as follows: ; in, Let be the coordinates of a 3D point in the world coordinate system in frame t. For the first t Frame-optimized camera-to-world transformation matrix W Using the world coordinate system, C Let be the camera coordinate system.
6. The method according to claim 1, characterized in that, The process of performing global spatiotemporal relationship modeling on the high-dimensional feature sequence to obtain the modeled features includes: Based on a preset self-attention calculation formula, a global spatiotemporal relationship model is performed on the high-dimensional feature sequence to obtain the modeled features. The preset self-attention calculation formula is as follows: ; in, Q For query vector, K For key vectors, V For value vectors, for Q and K Dimensions This is the scaling factor.
7. A video authenticity detection device, characterized in that, include: The acquisition module acquires the video to be detected and processes the video to be detected to obtain a dynamic three-dimensional point cloud sequence. The extraction module performs multi-scale geometric feature extraction on the dynamic three-dimensional point cloud sequence to obtain a high-dimensional feature sequence with spatiotemporal structure; The detection module performs global spatiotemporal relationship modeling on the high-dimensional feature sequence to obtain modeled features, performs global aggregation on the modeled features to obtain a global aggregation result, and obtains the authenticity detection result of the video to be detected based on the global aggregation result.
8. An electronic device, characterized in that, include: A memory, a processor, and a computer program stored in the memory and executable on the processor, the processor executing the program to implement the video authenticity detection method as described in any one of claims 1-6.
9. A computer-readable storage medium having a computer program stored thereon, characterized in that, The program is executed by the processor to implement the video authenticity detection method as described in any one of claims 1-6.
10. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the video authenticity detection method as described in any one of claims 1-6.