Video detection model training method and apparatus, video detection method and apparatus, and device
By perturbing key points of the target object in a video frame sequence, video frame sequence samples are constructed and the model is trained. This solves the problem of low detection accuracy of existing video detection models and achieves efficient recognition and generalization capabilities for deepfake videos.
Patent Information
- Application Number
- PCT/CN2025/091660
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-07-10
- Filing Date
- 2025-04-28
- Publication Date
- 2026-01-15
AI Technical Summary
Existing video detection models trained in the technology do not have high accuracy in detecting deepfake videos and lack generalization ability.
By perturbing the key points of the target object in the video frame sequence, video frame sequence samples are constructed. Then, computer vision technology and machine learning methods are used to train a video detection model to simulate the random jitter and drift of key points in deepfake videos, thereby improving the detection capability of the model.
The trained video detection model can accurately determine whether a video is a deepfake, improving detection accuracy and generalization ability, and can identify anomalies in unseen deepfake videos.
Smart Images

Figure CN2025091660_15012026_PF_FP_ABST
Abstract
Description
Training methods for video detection models, video detection methods, devices and equipment
[0001] This application claims priority to Chinese Patent Application No. 202410923946.9, filed on July 10, 2024, entitled “Training Method for Video Detection Model, Video Detection Method, Apparatus and Device”, the entire contents of which are incorporated herein by reference. Technical Field
[0002] This application relates to the field of artificial intelligence technology, and in particular to a training method for a video detection model, a video detection method, an apparatus, and a device. Background Technology
[0003] With the development of artificial intelligence technology, more and more videos are being forged using Deepfake (specifically referring to deepfake technology), making deepfake video detection increasingly important. Deepfake video detection refers to the process of detecting videos forged using Deepfake technology.
[0004] In related technologies, training samples are constructed using techniques such as discarding, repeating, and fusing, and a video detection model for detecting the authenticity of videos is trained based on these training samples. For example, for a given video frame sequence, related technologies construct training samples by employing processes such as randomly deleting video frames, copying and adding video frames, and fusing video frames.
[0005] However, the similarity between the training samples obtained using related technologies and deepfake videos is low, resulting in low detection accuracy of the trained video detection model. Summary of the Invention
[0006] This application provides a method for training a video detection model, a video detection method, an apparatus, and a device. The technical solutions provided by this application may include the following.
[0007] According to one aspect of the embodiments of this application, a method for training a video detection model is provided, the method comprising:
[0008] Obtain a video frame sequence, the video frame sequence including multiple video frames related to the target object, ordered chronologically;
[0009] For m video frames in the video frame sequence, the key points of the target object in the m video frames are perturbed respectively to obtain a video frame sequence sample. The features related to the key points of the target object show anomalies in the changes of multiple video frames included in the video frame sequence sample, where m is a positive integer.
[0010] A video detection model is trained based on the video frame sequence samples to obtain a trained video detection model, which is used to detect whether there are any anomalies in the video.
[0011] According to one aspect of the embodiments of this application, a video detection method is provided, the method comprising:
[0012] Acquire at least one video frame sequence of a video, the video frame sequence comprising multiple video frames related to the target object in chronological order;
[0013] For any video frame sequence in the at least one video frame sequence, a first feature map of the video frame sequence is obtained by a video detection model. The first feature map is used to indicate the changes of features related to the key points of the target object in the video frame sequence in the time dimension.
[0014] For any video frame sequence in the at least one video frame sequence, a classification result of the video frame sequence is obtained through a video detection model. The classification result of the video frame sequence is used to indicate whether there are abnormal changes in the features related to the key points of the target object in the video frame sequence. The video detection model is trained with video frame sequence samples. For m video frames in the video frame sequence samples, the key points of the target object are perturbed. The features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence samples, where m is a positive integer.
[0015] Based on the classification results corresponding to the at least one video frame sequence, the detection result of the video is obtained, and the detection result is used to indicate whether there is an anomaly in the video.
[0016] According to one aspect of the embodiments of this application, a training apparatus for a video detection model is provided, the apparatus comprising:
[0017] A frame sequence acquisition module is used to acquire a video frame sequence, which includes multiple video frames related to the target object, ordered chronologically.
[0018] The key point perturbation module is used to perturb the key points of the target object in each of the m video frames in the video frame sequence to obtain a video frame sequence sample. The features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence sample, where m is a positive integer.
[0019] The detection model training module is used to train a video detection model based on the video frame sequence samples to obtain a trained video detection model, which is used to detect whether there are any anomalies in the video.
[0020] According to one aspect of the embodiments of this application, a video detection apparatus is provided, the apparatus comprising:
[0021] A frame sequence acquisition module is used to acquire at least one video frame sequence of a video, wherein the video frame sequence includes multiple video frames related to the target object and ordered chronologically.
[0022] The classification result acquisition module is used to acquire the classification result of any video frame sequence in the at least one video frame sequence through a video detection model. The classification result of the video frame sequence is used to indicate whether there are abnormal changes in the features related to the key points of the target object in the multiple video frames included in the video frame sequence. The video detection model is trained on video frame sequence samples. For m video frames in the video frame sequence samples, the key points of the target object are perturbed, and the features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence samples, where m is a positive integer.
[0023] The inspection result acquisition module is used to acquire the detection result of the video based on the classification results corresponding to the at least one video frame sequence, and the detection result is used to indicate whether there is an anomaly in the video.
[0024] According to one aspect of the embodiments of this application, a computer device is provided, the computer device including a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement the training method of the video detection model described above, or to implement the video detection method described above.
[0025] According to one aspect of the embodiments of this application, a computer-readable storage medium is provided, wherein a computer program is stored in the storage medium, the computer program being loaded and executed by a processor to implement the training method of the video detection model described above, or to implement the video detection method described above.
[0026] According to one aspect of the embodiments of this application, a computer program product is provided, comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium and executes the computer program, causing the computer device to perform the training method for the video detection model described above, or to perform the video detection method described above.
[0027] The technical solutions provided in this application embodiment may include the following beneficial effects.
[0028] For a video frame sequence related to a target object, video frame sequence samples are constructed by perturbing the key points of the target object. This simulates the "random jitter and drift of key points" phenomenon commonly found in deepfake videos (videos constructed using deepfake technology). By training a video detection model using these video frame sequence samples, the trained video detection model can detect "random jitter and drift of key points." This allows the trained video detection model to accurately determine whether the video corresponding to the video frame sequence is a deepfake video based on whether the video frame sequence exhibits "random jitter and drift of key points," thereby improving the detection accuracy of the video detection model.
[0029] In addition, since the trained video detection model has the ability to detect "random jitter and drift of key points", it can also detect the anomalies of "key points" in unseen deepfake videos to determine the authenticity of the video, thereby effectively improving the generalization of the video detection model. Attached Figure Description
[0030] Figure 1 is a schematic diagram of a video detection system provided in an embodiment of this application;
[0031] Figure 2 is a schematic diagram of a video detection model provided in one embodiment of this application;
[0032] Figure 3 is a schematic diagram of a video detection model provided in another embodiment of this application;
[0033] Figure 4 is a schematic diagram of an adapter provided in one embodiment of this application;
[0034] Figure 5 is a flowchart of a training method for a video detection model provided in one embodiment of this application;
[0035] Figure 6 is a schematic diagram of the FFD (Facial Feature Drift) phenomenon provided in an embodiment of this application;
[0036] Figure 7 is a flowchart of a method for obtaining video frame sequence samples according to an embodiment of this application;
[0037] Figure 8 is a schematic diagram of a method for obtaining video frame sequence samples according to an embodiment of this application;
[0038] Figure 9 is a flowchart of a video detection method provided in an embodiment of this application;
[0039] Figures 10 to 14 are schematic diagrams of an information authentication interface provided in one embodiment of this application;
[0040] Figure 15 is a schematic diagram of a first feature map and a second feature map provided in an embodiment of this application;
[0041] Figure 16 is a block diagram of a training apparatus for a video detection model provided in one embodiment of this application;
[0042] Figure 17 is a block diagram of a training apparatus for a video detection model provided in another embodiment of this application;
[0043] Figure 18 is a block diagram of a video detection device provided in an embodiment of this application;
[0044] Figure 19 is a schematic diagram of the structure of a computer device provided in one embodiment of this application. Detailed Implementation
[0045] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0046] Artificial intelligence (AI) is the theory, methods, technology, and application systems that use digital computers or machines controlled by digital computers to simulate, extend, and expand human intelligence, perceive the environment, acquire knowledge, and use that knowledge to achieve optimal results. In other words, AI is a comprehensive technology within computer science that attempts to understand the essence of intelligence and produce a new kind of intelligent machine that can react in a way similar to human intelligence. AI studies the design principles and implementation methods of various intelligent machines, enabling them to possess the functions of perception, reasoning, and decision-making.
[0047] Artificial intelligence (AI) is a comprehensive discipline encompassing a wide range of fields, including both hardware and software technologies. Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, pre-trained model technology, operating / interactive systems, and mechatronics. Among these, pre-trained models, also known as large-scale models or foundational models, can be widely applied to downstream tasks across various AI fields after fine-tuning. AI software technologies primarily include computer vision, speech processing, natural language processing, and machine learning / deep learning.
[0048] Computer vision (CV) is a science that studies how to enable machines to "see." More specifically, it refers to machine vision, which uses cameras and computers to replace human eyes in tasks such as target recognition, tracking, and measurement, and further performs image processing to create images more suitable for human observation or transmission to instruments for detection. As a scientific discipline, computer vision researches related theories and technologies, attempting to build artificial intelligence systems capable of extracting information from images or multidimensional data. Large model technology has brought significant changes to the development of computer vision. Pre-trained models in the field of vision, such as Swin-Transformer, ViT (Vision Transformer), V-MOE (Vision Mixture-of-Experts), and MAE (Masked Autoencoders), can be fine-tuned and quickly and widely applied to specific downstream tasks. Computer vision technologies typically include image processing, image recognition, image semantic understanding, image retrieval, OCR (Optical Character Recognition), video processing, video semantic understanding, video content / behavior recognition, 3D object reconstruction, 3D (Three-Dimensional) technology, virtual reality, augmented reality, simultaneous localization and mapping (SLAM), and other technologies.
[0049] Key technologies in speech technology include Automatic Speech Recognition (ASR), Text-to-Speech (TTS), and voiceprint recognition. Enabling computers to hear, see, speak, and feel is the future direction of human-computer interaction, with speech emerging as one of the most promising methods. Large-scale modeling has revolutionized speech technology; pre-trained models such as WavLM and UniSpeech, which utilize the Transformer architecture, possess strong generalization and versatility, enabling them to excel in various speech processing tasks.
[0050] Natural Language Processing (NLP) is an important field within computer science and artificial intelligence. It studies the theories and methods for enabling effective communication between humans and computers using natural language. NLP deals with natural language, the language people use in daily life, and is closely related to linguistics, while also involving computer science and mathematics. Pre-trained models, a crucial technique for model training in artificial intelligence, evolved from large language models in NLP. After fine-tuning, large language models can be widely applied to downstream tasks. NLP techniques typically include text processing, semantic understanding, machine translation, question answering, and knowledge graphs.
[0051] Machine learning (ML) is a multidisciplinary field involving probability theory, statistics, approximation theory, convex analysis, and algorithm complexity theory. It specifically studies how computers can simulate or implement human learning behavior to acquire new knowledge or skills and reorganize existing knowledge structures to continuously improve their performance. Machine learning is the core of artificial intelligence and the fundamental way to endow computers with intelligence; its applications span all areas of artificial intelligence. Machine learning and deep learning typically include techniques such as artificial neural networks, belief networks, reinforcement learning, transfer learning, inductive learning, and instruction-based learning. Pre-trained models are the latest development in deep learning, integrating all of these techniques.
[0052] Autonomous driving technology refers to vehicles driving themselves without driver intervention. It typically includes technologies such as high-precision mapping, environmental perception, computer vision, behavioral decision-making, path planning, and motion control. Autonomous driving encompasses various development paths, including single-vehicle intelligence, vehicle-to-infrastructure (V2I) communication, and networked cloud control. Autonomous driving technology has broad application prospects, currently focusing on logistics, public transportation, taxis, and intelligent transportation systems, and is expected to see further development in the future.
[0053] With the research and advancement of artificial intelligence (AI) technology, AI is being studied and applied in various fields, such as smart homes, smart wearable devices, virtual assistants, smart speakers, smart marketing, autonomous driving, drones, digital twins, virtual humans, robots, AI-generated content (AIGC), conversational interaction, smart healthcare, smart customer service, and game AI. It is believed that with the development of technology, AI will be applied in more fields and play an increasingly important role.
[0054] The technical solutions provided in this application relate to computer vision and machine learning technologies in artificial intelligence to achieve deepfake video detection. For example, this application can utilize computer vision technology to construct video frame sequence samples, and then use computer vision and machine learning technologies to train a video detection model based on the video frame sequence samples to obtain a video detection model that can be used to detect deepfake videos.
[0055] Deepfake video detection refers to the process of detecting deepfake videos (such as videos with fake faces). Forgery techniques used in video detection include at least one of the following AIGC (Artificial Intelligence Generated Content) techniques: face replacement, face driving, face fusion, etc. Deepfake video detection can also be called video editing detection. A deepfake video is a video created using deepfake technology. Video forgery detection can refer to the process of detecting forged videos.
[0056] Deepfake technology: specifically refers to deepfake technology (widely known as AI face swapping), which represents fake images generated based on deep learning methods (such as images forged based on faces). This technology can create fake images that do not exist in reality. Deepfake technology is generally divided into face-swapping and face-reenactment.
[0057] An adapter is a module that can be inserted into a pre-trained neural network (such as the pre-trained model mentioned above) to fine-tune the network to adapt it to new tasks or datasets. An adapter typically contains one or two layers of a neural network, such as fully connected layers or convolutional layers, which perform transformations on the input data during the network's forward propagation. By training the inserted adapter, we can adapt the pre-trained neural network to new tasks without changing the weights of the original neural network. This method is widely used in transfer learning and multi-task learning.
[0058] Blending: This is a data augmentation method based on image fusion. In the field of deepfake detection, it specifically refers to the process of creating deepfake images by fusing the facial regions of two images (such as Poisson fusion). The resulting deepfake images can then be used together with real images to train video detection models.
[0059] Video-level blending refers to the process of performing blending operations at the video level (dimension) to generate a sequence of deepfake images. Unlike image-level blending, video-level blending has received less research in the field of deepfake detection. A deepfake image sequence is a sequence composed of deepfake images.
[0060] The technical solutions provided in this application are applicable to any scenario requiring deepfake video detection, such as identity verification (e.g., verification via recorded video), access control (e.g., facial recognition access control), fake video detection, video detection model training, and construction of training data for training video detection models. The technical solutions provided in this application can improve the detection accuracy and generalization ability of video detection models.
[0061] For example, in an identity verification scenario, after obtaining the user-uploaded input video, the client uses a trained video detection model to perform deepfake detection on the input video to determine whether it is fake or edited. If the input video is genuine, identity verification is performed based on it; if the input video is fake (e.g., fake / edited video), verification fails or the user is prompted to re-upload the video. The trained video detection model has the ability to detect "random jitter and drift of key points," and it is trained using video frame sequence samples constructed based on the technical solution provided in this application. Features related to the user's key points exhibit abnormal changes across multiple video frames included in the video frame sequence sample; that is, the user's key points exhibit random jitter and drift across multiple video frames included in the video frame sequence sample.
[0062] The video detection model provided in the embodiments of this application will be described below.
[0063] Please refer to Figure 1, which shows a schematic diagram of a video detection system provided in one embodiment of this application. The video detection system may include a model training device 10 and a model usage device 20.
[0064] The model training device 10 can be an electronic device such as a mobile phone, desktop computer, tablet computer, laptop computer, PC (Personal Computer), vehicle terminal, server, intelligent robot, smart TV, multimedia playback device, or other electronic devices with strong computing power. This application embodiment does not limit this. The server can be a single server, a server cluster consisting of multiple servers, or a cloud computing service center.
[0065] The model training device 10 is used to train the video detection model 30. The video detection model 30 is an AI model used to detect anomalies in videos (such as detecting whether a video is a deepfake). An AI model is a mathematical model based on artificial intelligence technology that can automatically process and analyze input data and output corresponding results. AI models include models with a large number of parameters trained through deep learning algorithms and artificial neural networks, such as the aforementioned large models, large language models, pre-trained models, and pre-trained neural networks.
[0066] The video detection model 30 takes the video frame sequence corresponding to the video as input and outputs the classification result of the video frame sequence. This classification result is used to indicate whether there are anomalies in the video frame sequence. Therefore, based on the classification result of the video frame sequence, it can determine whether there are anomalies in the video, thereby realizing deepfake video detection. Optionally, the model training device 10 uses machine learning to train the video detection model 30 using video frame sequence samples to obtain a trained video detection model 30 with better performance.
[0067] The trained video detection model 30 described above can be deployed on the model-using device 20 to provide video detection services (such as deepfake video detection). The model-using device 20 can be an electronic device such as a mobile phone, desktop computer, tablet computer, laptop computer, personal computer, vehicle terminal, server, intelligent robot, smart TV, multimedia playback device, or other electronic devices with strong computing power. This application embodiment does not limit this.
[0068] Optionally, the model uses a client installed and running a target application on device 20. This target application may include at least one of the following: video detection application, authentication application, access control application, model training application, social entertainment application, game application, shopping application, or payment application. Optionally, the aforementioned target application supports the deployment and use of the trained video detection model 30.
[0069] For example, referring to Figure 1, the model training device 10 can obtain a video frame sequence sample with "random jitter and drift" by perturbing the key points of the target object in at least one video frame in the video frame sequence, and use the video frame sequence sample to train the video detection model 30, so that the trained video detection model 30 has the ability to detect "random jitter and drift". The trained video detection model 30 can be deployed in the model using device 20 to provide deepfake video detection services.
[0070] In some embodiments, please refer to FIG2, which shows a schematic diagram of a video detection model provided in one embodiment of this application. The video detection model 30 may include: a temporal feature extraction network 301 and a spatial feature extraction network 302.
[0071] The temporal feature extraction network 301 is used to extract features from video frame sequence samples to obtain a first feature map of the target object in the temporal dimension. The first feature map can be used to indicate the changes in the features of the target object in the video frame sequence samples in the temporal dimension. The temporal feature extraction network 301 takes word embeddings of the video frame sequence samples as input and the first feature map as output.
[0072] In one example, the temporal feature extraction network 301 is constructed based on a first pre-trained neural network 301a and a first adapter 301b. The first pre-trained neural network 301a is used to extract features from video frame sequence samples to obtain feature data, and the first adapter 301b is used to adapt the first pre-trained neural network 301a to the temporal task, that is, to transform the features extracted by the first pre-trained neural network 301a to obtain a first feature map related to the time dimension.
[0073] Optionally, the first pre-trained neural network 301a is a pre-trained neural network, which can be built based on any neural network that can be used to process video, such as Transformer, Swin-Transformer, ViT, V-MOE, MAE, CNN (Convolutional Neural Networks), RNN (Recurrent Neural Networks), ResNet (Residual Networks), CLIP (Contrastive Language-Image Pre-training, a text-image pre-training model based on contrastive learning), etc. The first adapter 301b can be inserted into the first pre-trained neural network 301a, for example, the first adapter 301b can serve as the output layer of the first pre-trained neural network 301a to output the first feature map.
[0074] The spatial feature extraction network 302 is used to extract features from the output of the temporal feature extraction network 301 to obtain a second feature map of the target object in the spatial dimension. The second feature map can be used to indicate the spatial dimension changes of the features corresponding to the target object in the video frame sequence samples. The spatial feature extraction network 302 takes the first feature map as input and the second feature map as output.
[0075] In one example, the spatial feature extraction network 302 is constructed based on a second pre-trained neural network 302a and a second adapter 302b. The second pre-trained neural network 302a is used to extract features from the first feature map to obtain feature data, and the second adapter 302b is used to adapt the second pre-trained neural network 302a to the spatial task, that is, to transform the feature data extracted by the second pre-trained neural network 302a to obtain a second feature map related to the spatial dimension.
[0076] Optionally, the second pre-trained neural network 302a is also a pre-trained neural network, which can be built based on any neural network that can be used to process video, such as Transformer, Swin-Transformer, ViT, V-MOE, MAE, CNN, RNN, ResNet, CLIP, etc. The second adapter 302b can be inserted into the second pre-trained neural network 302a, for example, the second adapter 302b can serve as the output layer of the second pre-trained neural network 302a to output the second feature map.
[0077] In one example, the network structure of the first pre-trained neural network 301a is the same as that of the second pre-trained neural network 302a, and the network structure of the first adapter 301b is the same as that of the second adapter 302b, but the tasks of the first adapter 301b and the second adapter 302b are different. This application embodiment does not limit this.
[0078] For example, referring to Figure 3, taking a video detection model 30 built based on ViT as an example, the video detection model 30 includes 14 blocks. Each block includes a temporal feature extraction network, a spatial feature extraction network, a normalization layer (norm), and a multilayer perceptron (MLP). The temporal feature extraction network includes a first pre-trained network 301a inserted into the first adapter 301b (i.e., the T-Adapter), and the spatial feature extraction network includes a second pre-trained network 302a inserted into the second adapter 302b (i.e., the S-Adapter). The first pre-trained network 301a and the second pre-trained network 302a without adapters can both be built based on the encoder in ViT, such as both including a normalization layer and a multi-head attention mechanism layer. Optionally, the pre-trained weights on CLIP (i.e., weights obtained through pre-training) are used to initialize the first pre-trained network 301a and the second pre-trained network 302a to transfer the pre-trained weights of ViT to the video processing task.
[0079] Referring to Figure 4, the first adapter 301b and the second adapter 302b may each include two fully connected layers (FC): fully connected layer 1 and fully connected layer 2. The first adapter 301b and the second adapter 302b may also each include one rectifier layer (such as GELU, Gaussian error linear unit).
[0080] Optionally, the video detection model 30 also includes an I3D Head (a network structure for extracting video features, such as overlaying features in the time dimension to obtain a spatiotemporal feature map) and a classifier. The I3D Head takes the outputs corresponding to 14 blocks as input and the spatiotemporal feature map of the video frame sequence samples as output. The classifier is used to classify the video frame sequence samples based on the spatiotemporal feature map of the video frame sequence samples to obtain the classification result of the video frame sequence samples. The classifier takes the output of the I3D Head (i.e., the spatiotemporal feature map) as input and the classification result as output.
[0081] It should be noted that the video detection model 30 shown in Figures 2 to 4 is only exemplary and illustrative. The specific structure of the video detection model 30 is not limited in the embodiments of this application, and it can be set and adjusted according to actual usage requirements.
[0082] The following are embodiments of the method of this application, which illustrate the training method of the video detection model 30. For details not disclosed in the embodiments of the method of this application, please refer to the above embodiments.
[0083] Please refer to Figure 5, which shows a flowchart of a training method for a video detection model provided in an embodiment of this application. The execution entity of each step of the method can be the model training device 10 shown in Figure 1. The method may include the following steps (501-503).
[0084] Step 501: Obtain a video frame sequence, which includes multiple video frames related to the target object, ordered chronologically.
[0085] A video frame sequence can refer to a sequence of multiple video frames arranged chronologically. In this embodiment, the video frame sequence is related to a target object; that is, each video frame in the sequence is related to the target object. For example, each video frame in the sequence includes the target object; or, each video frame includes a part of the target object, such as a face; or, at least two video frames in the sequence include the target object. This embodiment does not limit the number of video frames in the sequence; it can be set and adjusted according to actual usage requirements. For example, the number of video frames in the sequence can be 7, 8, 9, 10, etc.
[0086] Optionally, during the training of the video detection model, the aforementioned target object refers to the object to be perturbed at key points, such as any object included in the video frame sequence; during deepfake video detection, the target object can refer to the object of interest in the video frame sequence, such as the object to be authenticated, the object suspected of being forged (e.g., face swapping, fabrication, AI generation), the object suspected of being edited, etc. This application embodiment does not limit the type of target object, and the type of target object includes at least one of the following: human, animal.
[0087] In one example, the aforementioned video frame sequence can be constructed based on real video. For instance, multiple video frames related to the target object can be randomly extracted from the real video to construct the video frame sequence; alternatively, multiple consecutive video frames related to the target object can be extracted from the real video to construct the video frame sequence; or, a first number of video frames can be sampled from the real video, and then a second number of consecutive video frames can be randomly selected from the first number of video frames to construct the video frame sequence. This application does not limit the method for constructing the video frame sequence.
[0088] The first quantity is greater than the second quantity, and both quantities can be set and adjusted according to actual usage needs. A real video refers to a video that has not been forged, while a forged video refers to a video that has been forged (e.g., processed using deepfake technology). Optionally, a real video refers to a video captured on the face of a target object. For example, in an identity verification scenario, this real video is a video captured on the face of a target object using a camera.
[0089] For example, during the training of the video detection model, multiple real videos can be obtained from a database. For any given real video, 100 video frames are sampled from each real video, and 8 consecutive video frames are randomly selected from these 100 video frames to form a video frame sequence. The training of the video detection model is an iterative process, such as adjusting the parameters of the video detection model with a batch of samples each time, where each batch of samples includes at least one sample. This application embodiment uses a certain iterative process of the video detection model as an example for illustration.
[0090] Step 502: For m video frames in the video frame sequence, the key points of the target object in each of the m video frames are perturbed to obtain a video frame sequence sample. The features related to the key points of the target object show anomalies in the changes of multiple video frames contained in the video frame sequence sample, where m is a positive integer.
[0091] Video frame sequence samples refer to the samples used to train the aforementioned video detection model, which can be obtained by adjusting the aforementioned video frame sequence. For example, the aforementioned m video frames can refer to all video frames in the video frame sequence, that is, by perturbing the key points of the target object in each video frame of the video frame sequence to obtain video frame sequence samples; the aforementioned m video frames can also refer to a portion of the video frame sequence, such as m being greater than or equal to a first threshold, that is, by perturbing the key points of the target object in a portion of the video frame sequence to obtain video frame sequence samples. The first threshold can be set and adjusted based on empirical values, such as the first threshold being half the number of video frames in the video frame sequence. When the aforementioned m video frames are a portion of the video frame sequence, the video frame sequence can be randomly sampled to obtain m distinct video frames.
[0092] In this application's embodiments, "perturbation" can refer to a process of changing position, such as perturbing keypoints, which can refer to changing the position of keypoints in a video frame. Keypoints refer to points used to locate and identify parts of a target object. This application's embodiments do not limit the type of keypoints. For example, when the video detection model is used to detect the facial region of a target object, the aforementioned keypoints can be facial keypoints, such as keypoints corresponding to eyebrows, eyes, nose, and mouth; when the video detection model is used to detect the human posture of a target object, the aforementioned keypoints can be limb keypoints, such as keypoints corresponding to the head, legs, hands, and torso; when the video detection model is used to detect the palm of a target object, the aforementioned keypoints can be hand keypoints, such as keypoints corresponding to fingertips, palm, and palm prints; when the video detection model is used to detect the iris of a target object, the aforementioned keypoints can be iris keypoints, such as keypoints used to locate and identify the iris.
[0093] Features related to keypoints can refer to features extracted from keypoints. Optionally, for facial keypoints, features related to the keypoints of the target object refer to facial features; for limb keypoints, features related to the keypoints of the target object refer to pose features; for hand keypoints, features related to the keypoints of the target object refer to palm features; and for iris keypoints, features related to the keypoints of the target object refer to iris features.
[0094] In this embodiment, an anomaly can refer to a temporal inconsistency in features, used to indicate whether a video frame sequence sample has been forged. For example, the changes in the facial features of the target object in the multiple video frames contained in the video frame sequence were originally consistent in time, that is, the changes were natural and continuous. However, after the facial key points were perturbed, the changes in the facial key points of the target object in the multiple video frames contained in the video frame sequence sample became inconsistent in time, with unnatural jitter and drift. This caused subtle inconsistencies in the position and shape of facial organs, and thus the changes in the facial features of the target object in the multiple video frames contained in the video frame sequence sample were also inconsistent in time, that is, there was an anomaly. In this embodiment, this anomaly is called FFD (Facial Feature Drift).
[0095] Most deepfake algorithms generate results frame-by-frame. For example, deepfake algorithms typically use frame-by-frame face swapping to forge video frames and obtain fake video frames. During the deepfake face swapping process, the detection results (facial keypoints) of a face detector are often used for alignment, thereby enabling affine transformation and other subsequent face swapping processes. However, since most face detectors currently identify and detect faces frame-by-frame, the temporal consistency of facial keypoints across multiple fake video frames cannot be guaranteed. This leads to subtle inconsistencies in the positions of facial keypoints between different fake video frames, resulting in the FFD (Facial Feature Defect) phenomenon. Furthermore, since the inner face after a deepfake face swap is usually generated by a deep neural network (such as a generative adversarial network), it is inherently difficult to ensure that the positions of facial keypoints in different fake video frames are perfectly aligned.
[0096] For example, referring to Figure 6, for a fake video, even if two consecutive video frames appear relatively still, subtle unnatural drifts and jitters may be observed in facial features (e.g., eyes, nose, etc.). For example, for the real video frame and the fake video frame in dashed box 601, when the first real video remains unchanged and the second real video is faked into a fake video frame, there is obvious drift in the target object's eyeballs.
[0097] In one example, the keypoints in each video frame can be randomly perturbed first, and then the above video frame sequence samples can be generated through image-level self-blending to simulate unnatural jitter and drift of the keypoints, thereby constructing a video frame sequence sample with high similarity to the forged video. For example, the above video frame sequence sample can be a sample constructed by simulating the FFD phenomenon, such as by randomly perturbing the facial keypoints in each video frame to cause the facial keypoints to shift in the temporal dimension, and then generating the above video frame sequence sample through image-level self-blending. Since this data augmentation method is constructed specifically for the characteristics of deepfake video generation, it can effectively simulate the temporal inconsistencies of deepfake videos, thereby effectively improving the detection accuracy and generalization of video detection models.
[0098] For example, as shown in FIG7, step 502 may further include the following sub-steps:
[0099] Step 502a: For any video frame among the m video frames, perturb each pixel in the video frame to obtain a perturbed video frame. There is a difference between the pixels in the perturbed video frame and the pixels in the video frame.
[0100] Optionally, randomly perturbing each pixel (including keypoints) in a video frame can cause each pixel to jitter randomly, thus simulating the random jitter and drift of keypoints. Since perturbation can change the position of pixels, at the same position, there are differences between pixels in the perturbed video frame and pixels in the corresponding video frame, such as different pixel values.
[0101] In one example, the process of obtaining the perturbed video frame can be as follows:
[0102] 1. Obtain the affine transformation matrix. The affine transformation matrix uses perturbation parameters as elements, which are used to change the position of the pixels.
[0103] An affine transformation matrix is a matrix used to perform affine transformations. Affine transformations can be used to linearly transform the coordinates of a pixel to change its position. For example, if changing the position of a pixel involves at least one of the following: rotation, translation, or scaling, then the affine transformation matrix can be used to implement at least one of these: rotation, translation, or scaling. Optionally, changing the position of a pixel may also include at least one of the following: horizontal flipping or cropping.
[0104] For example, an affine transformation matrix can individually implement any one of rotation, translation, or scaling; it can also simultaneously implement any two of rotation, translation, or scaling; and it can even simultaneously implement rotation, translation, and scaling. This application does not limit this. Rotation can be the process of rotating a pixel around a point in a video frame; translation can refer to the process of moving a pixel based on its original position in the video frame; and scaling can refer to the process of decreasing or increasing the distance between pixels.
[0105] Optionally, the perturbation parameters follow at least one of the following distributions: uniform distribution, Gaussian distribution, and Poisson distribution. For example, the perturbation parameters can be sampled according to a uniform distribution for the current video frame, and then an affine transformation matrix, denoted as A, can be constructed using the perturbation parameters as elements.
[0106] 2. Perform an affine transformation on the video frame based on the affine transformation matrix to obtain the perturbed video frame.
[0107] Optionally, for each pixel in a video frame, multiplying the pixel's coordinates by the affine transformation matrix yields the changed position of that pixel. After perturbing each pixel in the video frame, a perturbed video frame is generated. For example, a perturbed video frame can be represented as follows: I warp =AI;
[0108] Among them, I warp Let I be the perturbed video frame (represented by the set of pixels after perturbing), and let A be the affine transformation matrix.
[0109] For key points in a video frame, the corresponding perturbation process can be represented as follows: L * R =AL R ;
[0110] Among them, L * R For key point L R After affine transformation, the keypoints (e.g., represented by their positions in video frames) are indicated by R, which indicates the type of keypoint, such as facial keypoints.
[0111] This application embodiment simulates random jitter and drift of key points by randomly perturbing key points in m video frames, which can improve the similarity between video frame sequence samples and deepfake videos, thereby improving the ability of video detection models to identify deepfake videos.
[0112] Step 502b: Fuse the perturbation video frame and the video frame to obtain a fake video frame. There are differences between the key points of the target object in the fake video frame and the key points of the target object in the video frame. There are no differences between the non-key points of the target object in the fake video frame and the non-key points of the target object in the video frame.
[0113] Non-critical points refer to pixels in a video frame other than keypoints. Optionally, the position of a keypoint of the target object in a forged video frame differs from the position of a keypoint of the target object in the corresponding video frame, i.e., keypoints drift. Conversely, the position of a non-critical point of the target object in a forged video frame does not differ from the position of a non-critical point of the target object in the corresponding video frame, i.e., non-critical points do not drift.
[0114] Since this application embodiment only focuses on the jitter of key points and does not consider the jitter of non-key points, a mask can be constructed to keep non-key points unchanged while changing key points, thereby improving the similarity between video frame sequence samples and deepfake videos. For example, the process of generating fake video frames may include the following:
[0115] 1. Obtain multiple keypoint groups of the target object in the video frame, with each keypoint group including multiple keypoints.
[0116] A keypoint group refers to a set of keypoints. Optionally, the keypoints in each keypoint group are related. For example, each keypoint group corresponds to the area occupied by a part of the target object, that is, each keypoint group includes the keypoints of the part corresponding to that keypoint group, and the keypoints in the keypoint group can jointly reflect information such as the shape, size, and position of that part. For example, for facial keypoints, the target object may correspond to 6 keypoint groups, the two eyebrows to two keypoint groups, the two eyes to two keypoint groups, the nose to one keypoint group, and the mouth to one keypoint group.
[0117] Optionally, key points (such as the coordinates of key points) or multiple key point groups of the target object are detected from the video frame by a key point detection model. The key point detection model can be built based on any of the following neural networks: DCNN (Deep CNN), TCNN (Tweaked CNN), or DAN (Deep Alignment Networks).
[0118] 2. For any keypoint group among multiple keypoint groups, construct a mask image corresponding to the video frame based on the keypoint group. The mask image is used to indicate the area of the keypoint group in the video frame.
[0119] In other words, a mask image is constructed for each keypoint group. A mask image can be an image composed of mask values, where each pixel in the mask image is represented by a mask value of 0 to 1 (i.e., a binary image). A mask image with all mask values of 0 appears white, and a mask image with all mask values of 1 appears black. The size of the mask image is the same as the size of the video frame.
[0120] The shape of the region containing the keypoint group in the video frame corresponds to the shape of the part corresponding to the keypoint group. The region containing the keypoint group in the video frame may also include the region containing the part corresponding to the keypoint group in the video frame; this embodiment does not limit this. For example, the mask value of each pixel within the region corresponding to the keypoint group is between 0 and 1, and it is white; the mask value of each pixel outside the region corresponding to the keypoint group is 1, and it is black.
[0121] In one example, the process of constructing a mask image may include the following:
[0122] (1) For any pixel in a video frame, obtain the first distance between the pixel and the area where the key point group is located.
[0123] The region where a keypoint group is located refers to the area of the keypoint group within a video frame. In this embodiment, the first distance is used to indicate the distance between a pixel and the region where the keypoint group is located.
[0124] In one example, the distance between a pixel and each keypoint is first obtained, then the minimum distance is determined from multiple distances, and this minimum distance is determined as the first distance.
[0125] In one example, the geometric center of the keypoint group is calculated based on the position of each keypoint in the keypoint group, and the distance between the pixel and the geometric center is determined as the first distance.
[0126] For example, the first distance can be represented as follows: dist(x, y, L) r );
[0127] Here, `dist()` is a function used to calculate the distance between two points, where `x` and `y` are the x and y coordinates of the pixel in the video frame, and `L`... r Used to represent the geometric center (e.g., the coordinates of the geometric center) corresponding to the key point group r.
[0128] (2) Determine the mask value of the pixel based on the first distance. The mask value of the pixel is negatively correlated with the first distance.
[0129] Optionally, the range of the mask value in the embodiments of this application can be [0, 1].
[0130] In one example, a maximum distance is set for each keypoint group. This is achieved by obtaining the maximum distance between the edge of the corresponding part of the keypoint group and the geometric center of the keypoint group, and determining the maximum distance as any number greater than or equal to this maximum value. For example, for pixels belonging to the corresponding part of a keypoint group, the corresponding mask value ranges from [0, 1], and the mask value of the pixel is negatively correlated with the aforementioned first distance. For pixels not belonging to the corresponding part of a keypoint group, the corresponding mask value is 0.
[0131] For example, the process of determining the mask value may include the following:
[0132] Divide the first distance by the maximum distance to obtain the first quotient. If the first quotient is greater than or equal to 1, set the mask value of the pixel to 1. If the first quotient is less than 1 but greater than or equal to 0, set the mask value to the first quotient.
[0133] For each keypoint group, the mask value of each pixel in the corresponding mask image can be represented as follows:
[0134] Among them, M r (x, y) represents the mask value of any pixel in the mask image, where x and y are the horizontal and vertical coordinates of that pixel in the video frame. `clip()` is a function used to restrict the value to a specified range. `dist()` is a function used to calculate the distance between two points. `fdist` r It refers to the maximum distance corresponding to keypoint group r, where R is a set of multiple keypoint groups.
[0135] The clip() function, clip(value, min, max), restricts value to the range [min, max]. If value is less than min, clip() returns min; if value is greater than max, clip() returns max; if value is within the range [min, max], clip() returns value.
[0136] The above formula allows the pixel's mask value to smoothly transition from 1 (fully included in the blend, i.e., requiring fusion) to 0 (excluded from the blend, i.e., not requiring fusion) as the distance to the keypoint increases, until the maximum distance is reached. By fusing the perturbed video frame and the video frame based on the mask image, the closer pixels in the forged video frame are to the keypoint, the more obvious the perturbation becomes. This accurately simulates the jitter and drift of the keypoint, thereby improving the similarity between the video frame sequence samples and the deepfake video.
[0137] (3) Based on the mask values of each pixel in the video frame, a mask image is constructed.
[0138] Optionally, the mask image can be constructed by replacing the pixel values of each pixel in the video frame with the corresponding mask values.
[0139] For each group of key points, a mask image is constructed.
[0140] 3. Under the constraints of the mask images corresponding to multiple key point groups, the perturbed video frames and video frames are fused to obtain the forged video frames.
[0141] The mask values in the mask image can be used to indicate the proportion of each pixel in the perturbed video frame contributed to the forged video frame. Optionally, for each mask image, under the constraints of the mask image, the perturbed video frame and the video frame are fused to obtain the mixed video frame corresponding to the mask image. Then, the mixed video frames of each mask image are fused to obtain the forged video frame.
[0142] For example, the process may include the following:
[0143] (1) For any mask image in the mask images corresponding to multiple key point groups, adjust the perturbation video frame based on the mask image to obtain the first mixed video frame, and adjust the video frame based on the inverse mask image of the mask image to obtain the second mixed video frame. The sum of the mask values of the pixels at the same position in the mask image and the inverse mask image is 1.
[0144] The first mixed video frame is used to indicate the amount of forged video frames used on the perturbed video frames. The second mixed video frame is used to indicate the amount of forged video frames used on the video frames.
[0145] Optionally, for any pixel in the perturbed video frame, the pixel is multiplied by its corresponding mask value in the mask image to obtain the first mixed video frame. For example, the first mixed video frame can be represented as follows: (M r ·I warp ); where M r This is the mask image corresponding to the key point group r.
[0146] Optionally, for any pixel in a video frame, the pixel is multiplied by the mask value corresponding to that pixel in the inverse mask image to obtain a second mixed video frame. For example, the second mixed video frame can be represented as follows: (1-M) r )·I; where M r For the mask image corresponding to key point group r, (1-M) r ) is the inverse mask image corresponding to the key point group r.
[0147] (2) Based on the first mixed video frame and the second mixed video frame, the mixed video frame of the mask image is obtained.
[0148] Optionally, for the first and second mixed video frames, the pixel values of pixels at the same position are added together to obtain the mixed video frame of the mask image. For example, the mixed video frame can be represented as follows:
[0149] I' r =(M r ·I warp )+((1-M r )·I), r∈R;
[0150] Among them, M r The I value used on each pixel in the mixed video frame is determined. warp The amount of I is used to achieve the mixing effect at the location corresponding to the key point group r.
[0151] (3) Fuse the mixed video frames corresponding to multiple mask images to obtain fake video frames.
[0152] Optionally, a weighted fusion is performed on multiple mixed video frames to obtain a forged video frame. This ensures that the features corresponding to the target object in each mixed video frame are evenly fused, thereby improving the quality of the forged video frame. For example, a forged video frame can be represented as follows:
[0153] Each keypoint group r is determined by its blending weight α. r Contribute to I′, thereby evenly integrating all features of the target object into the final forged video frame, with mixing weight α. r The settings and adjustments can be made according to actual usage needs, and this application embodiment does not limit this.
[0154] Step 502c: Replace m video frames with the corresponding fake video frames to obtain a video frame sequence sample.
[0155] For any one of the m video frames, replace it with the corresponding fake video frame. After all m video frames have been replaced, a video frame sequence sample is obtained. Key points of the target object exhibit jitter and drift in the video frame sequence sample. During the replacement process, the video frames in the video frame sequence sample are kept in chronological order.
[0156] Step 503: Train a video detection model based on video frame sequence samples to obtain a trained video detection model. The trained video detection model is used to detect whether there are any anomalies in the video.
[0157] The trained video detection model can be used to provide fake video detection services (such as deepfake video detection services). In the embodiments of this application, if there are anomalies in the video frame sequence samples, it can be determined that the video corresponding to the video frame sequence samples is abnormal, that is, the video can be determined to be a fake video, such as a fake video, an edited video, a deepfake video, etc.
[0158] In one example, a video detection model can be used to obtain the classification results of video frame sequence samples, and then the video detection model can be trained based on the classification results. This process may include the following:
[0159] 1. Obtain the first feature map of the video frame sequence sample through the video detection model. The first feature map is used to indicate the change of the corresponding features of the target object in the video frame sequence sample in the time dimension.
[0160] Optionally, the video frame sequence samples can be converted into word embedding block sequences, and then the temporal feature extraction network in the video detection model can be used to extract features from the word embedding block sequences to obtain the first feature map. For example, referring to Figure 2, for a word embedding block sequence e, its size is (T, N+1, D), where T is the number of video frames in the video frame sequence samples, N+1 is the number of word embedding blocks corresponding to each video frame, and D is the number of channels. T and N+1 can be set and adjusted according to actual usage requirements.
[0161] Optionally, the word embedding block sequence e may also include position embedding and temporal embedding, where position embedding is used to indicate the position of each word embedding block and temporal embedding is used to indicate the time of each word embedding block.
[0162] By extracting features from the word embedding block sequence e using the temporal feature extraction network 301, the first feature map e can be obtained. t First feature map e t The dimensions are also (T, N+1, D), and the first feature map e t The feature map slices (N+1, D) in the video frame sequence are stacked in the time dimension to indicate how the features of the target object in the video frame sequence sample change in the time dimension, such as whether the feature changes are continuous and natural.
[0163] It should be noted that the working process of each block in the video detection model is the same. This application embodiment uses the working process of a certain block in the video detection model (such as the first block) as an example for illustration.
[0164] 2. Based on the first feature map, the video detection model obtains the second feature map. The second feature map is used to indicate the spatial dimension changes of the features corresponding to the target object in the video frame sequence samples.
[0165] Optionally, the second feature map can be obtained by extracting features from the first feature map using a spatial feature extraction network in the video detection model. Since the spatial feature extraction network operates in a spatial dimension, the first feature map can be transformed to a spatial dimension before processing it. For example, this process may include the following:
[0166] (1) Perform a dimension transformation on the first feature map to obtain the transformed first feature map.
[0167] Optionally, the feature map slices in the transformed first feature map are stacked in spatial dimensions. For example, referring to Figure 2, for the first feature map e t By performing a dimensionality transformation, we can obtain the transformed first feature map e′. t The first feature map after transformation, e′ t The size is (N+1, T, D), and the first feature map after transformation is e′. t The feature map slices (T, D) are stacked in the time dimension. For example, the first feature map can be mapped to the transformed first feature map, or the first feature map can be transformed to the transformed first feature map through a convolutional layer. This application embodiment does not limit this.
[0168] (2) The second feature map is obtained by using the video detection model based on the transformed first feature map.
[0169] Optionally, the second feature map can be obtained by extracting features from the transformed first feature map using a spatial feature extraction network in the video detection model. For example, referring to Figure 2, the transformed first feature map e′ is processed by the spatial feature extraction network 302. t By performing feature extraction, the second feature map e can be obtained. s Second feature map e s The size is (N+1, T, D) to indicate the spatial variation of the features of the target object in multiple video frames contained in the video frame sequence sample, such as whether the feature changes are continuous and natural.
[0170] To ensure that the video detection model can capture both space and time, embodiments of this application explicitly change the channel dimension of the input, so that the temporal feature extraction network and the spatial feature extraction network can be applied to the spatial and temporal dimensions of their inputs, respectively.
[0171] 3. Based on the second feature map, the video detection model obtains the classification results of the video frame sequence samples. The classification results of the video frame sequence samples are used to indicate whether there are any abnormalities in the changes of features related to the key points of the target object in the multiple video frames contained in the video frame sequence samples.
[0172] Optionally, the output of the last block in the video detection model is used as the final second feature map. Then, the I3D Head in the video detection model is used to extract features from the second feature map to obtain the spatiotemporal feature map. Finally, the classifier in the video detection model is used to classify the spatiotemporal feature map to obtain the classification result of the video frame sequence sample.
[0173] The classification results can be used to indicate whether a video frame sequence sample is forged, or whether there is keypoint jitter or drift. For example, for facial keypoints, the classification results of the video frame sequence sample are used to indicate whether the facial features of the target object exhibit FFD (Facial Feature Dispersion). Optionally, if the classification results indicate anomalies, the video frame sequence sample can be determined to be forged.
[0174] 4. Train the video detection model based on the classification results to obtain the trained video detection model.
[0175] Optionally, based on the classification result, a loss function value for the video detection model is determined. The loss function value is used to indicate the detection accuracy of the video detection model. For example, if the classification result is a binary classification result: 1 and 0, where 1 indicates the presence of an anomaly and 0 indicates the absence of an anomaly, and the label of the video frame sequence sample is 1 (i.e., an anomaly exists), then a loss function value can be constructed based on the difference between the classification result and the label. For example, a cross-entropy loss function or a mean squared error loss function can be used.
[0176] After obtaining the loss function value, the video detection model can be trained based on the loss function value to obtain the trained video detection model. Since the first and second pre-trained neural networks in the video detection model have been pre-trained, training can be performed only on the first and second adapters to adapt the video detection model to the deepfake video detection task. This reduces the training difficulty and complexity of the video detection model, thereby improving the training efficiency of the video detection model.
[0177] For example, the parameters of the first pre-trained neural network and the second pre-trained neural network can be fixed, and the parameters of the first adapter and the second adapter can be adjusted according to the loss function value to obtain the trained video detection model. The training of the video detection model is an iterative process, and the termination condition of the iteration includes at least one of the following: minimizing the loss function value, the number of iterations being greater than or equal to a threshold, the loss function value being less than or equal to a threshold, etc., which are not limited in this embodiment.
[0178] This application embodiment adjusts the parameters of the first adapter and the second adapter, enabling the video detection model to learn the ability to detect anomalies (i.e., unnatural jitter and drift of key points) in video frame sequence samples simultaneously in both spatial and temporal dimensions. This helps to improve the accuracy of anomaly detection.
[0179] Optionally, both the first adapter and the second adapter can be lightweight adapters (i.e., with fewer network parameters), which helps to reduce the structural complexity of the video detection model, making it easier to build and train the video detection model.
[0180] In summary, the technical solution provided in this application, for video frame sequences related to a target object, constructs video frame sequence samples by perturbing the key points of the target object. This simulates the "random jitter and drift of key points" phenomenon commonly found in deepfake videos (videos constructed using deepfake technology). By training a video detection model using these video frame sequence samples, the trained video detection model can detect "random jitter and drift of key points." Consequently, the trained video detection model can accurately determine whether the video corresponding to the video frame sequence is a deepfake video based on whether the video frame sequence exhibits the "random jitter and drift of key points," thereby improving the detection accuracy of the video detection model.
[0181] In addition, since the trained video detection model has the ability to detect "random jitter and drift of key points", it can also detect the anomalies of "key points" in unseen deepfake videos to determine the authenticity of the video, thereby effectively improving the generalization of the video detection model.
[0182] In some embodiments, when the aforementioned key points are facial key points of the target object, the features of the target object are facial features. The technical solution provided in this application embodiment can construct video frame sequence samples exhibiting FFD (Free Front-Ended Defects). Then, by training a video detection model using these video frame sequence samples, a video detection model capable of accurately detecting FFD can be obtained. For example, referring to Figure 8, the construction of video frame sequence samples may include the following:
[0183] 1. Obtain the video frame sequence, which does not exhibit FFD (Free Frame Deformation).
[0184] Optionally, a video frame sequence refers to a sequence constructed based on the face of a target object, such as multiple video frames related to the face of the target object, ordered chronologically. The video frame sequence is sampled from a real video.
[0185] 2. For any video frame 801 in the video frame sequence, perform key point detection on video frame 801 to obtain the facial key points 802 of the target object.
[0186] Optionally, facial key points 802 are divided into multiple key point groups, such as two eyebrows corresponding to one key point group, two eyes corresponding to one key point group, the nose corresponding to one key point group, and the mouth corresponding to one key point group.
[0187] 3. Based on facial key points 802, construct the mask image 803 of video frame 801.
[0188] Mask image 803 includes four mask images: an eyebrow mask image, an eye mask image, a nose mask image, and a mouth mask image, each corresponding to a keypoint group. The eyebrow mask image indicates the area occupied by the two eyebrows in video frame 801, the eye mask image indicates the area occupied by the two eyes in video frame 801, the nose mask image indicates the area occupied by the nose in video frame 801, and the mouth mask image indicates the area occupied by the mouth in video frame 801.
[0189] Optionally, for any one of the four keypoint groups, a mask image corresponding to the video frame is constructed based on the keypoint group. The method for constructing the mask image is the same as in the above embodiment and will not be repeated here. A mask image is constructed for each keypoint group.
[0190] 4. Construct the affine transformation matrix by randomly sampling the perturbation parameters.
[0191] The affine transformation matrix uses perturbation parameters as elements, which are used to change the position of the pixels. These perturbation parameters follow a uniform distribution. The methods for changing the position of the pixels include at least one of the following: rotation, translation, and scaling.
[0192] 5. Perform an affine transformation on video frame 801 according to the affine transformation matrix to obtain perturbed video frame 804.
[0193] 6. Under the constraint of the mask image 803, the perturbation video frame 804 and the video frame 801 are fused to obtain the forged video frame 805.
[0194] For any one of the mask images—eyebrow mask image, eye mask image, nose mask image, and mouth mask image—adjust the perturbation video frame 804 based on the mask image to obtain a first mixed video frame, and adjust the video frame 801 based on the inverse mask image corresponding to the mask image to obtain a second mixed video frame; based on the first mixed video frame and the second mixed video frame, obtain the mixed video frame corresponding to the mask image; fuse the mixed video frames corresponding to the four mask images respectively to obtain the forged video frame 805.
[0195] 7. Replace each video frame in the video frame sequence with the corresponding fake video frame 805 in turn to obtain a video frame sequence sample. This video frame sequence sample exhibits FFD phenomenon.
[0196] The embodiments of this application do not limit the order in which the mask image 803 and the perturbation video frame 804 are constructed.
[0197] 8. By training a video detection model using the video frame sequence samples, a video detection model that can accurately detect FFD phenomena can be obtained.
[0198] In summary, the technical solution provided in this application, for video frame sequences related to a target object, constructs video frame sequence samples by perturbing the facial key points of the target object, simulating the "random jitter and drift of key points" phenomenon commonly found in deepfake videos (videos constructed using deepfake technology). Then, by training a video detection model using these video frame sequence samples, the trained video detection model can be equipped with the ability to detect the "FFD phenomenon," thereby improving the detection accuracy of the video detection model.
[0199] In addition, since the trained video detection model has the ability to detect the "FFD phenomenon", it can also detect the abnormality of "facial key points" in unseen deepfake videos to determine the authenticity of the video, thereby effectively improving the generalization of the video detection model.
[0200] Please refer to Figure 9, which shows a flowchart of a video detection method provided in an embodiment of this application. The execution subject of each step of the method can be the model-using device 20 shown in Figure 1. The method can include the following steps (901-903).
[0201] Step 901: Obtain at least one video frame sequence of the video, the video frame sequence including multiple video frames related to the target object in chronological order.
[0202] The aforementioned video can refer to any video to be detected. For example, in an identity verification scenario, the aforementioned video can refer to a video uploaded by the user that records images of the face, palm, iris, etc.
[0203] During the video detection process described above, the number of video frame sequences extracted can be set according to actual usage requirements. For example, multiple video frame sequences can be randomly extracted from the video to prevent missed detections and avoid misidentifying forged parts in the video. For instance, five video frame sequences can be extracted sequentially from the video at set intervals, with each video frame sequence consisting of eight consecutive video frames.
[0204] For any content not described in the embodiments of this application, please refer to the above embodiments, and it will not be repeated here.
[0205] Step 902: For any video frame sequence in at least one video frame sequence, obtain the classification result of the video frame sequence through the video detection model. The classification result of the video frame sequence is used to indicate whether there are abnormal changes in the features related to the key points of the target object in the multiple video frames contained in the video frame sequence. The video detection model is trained on the video frame sequence samples. For m video frames in the video frame sequence samples, the key points of the target object are perturbed, and the features related to the key points of the target object show abnormal changes in the multiple video frames contained in the video frame sequence samples, where m is a positive integer.
[0206] The video detection model in this embodiment refers to the trained video detection model. The process of obtaining video frame sequence samples and training the video detection model using video frame sequence samples is the same as described in the above embodiments. For content not described in this embodiment, please refer to the above embodiments, which will not be repeated here.
[0207] In one example, the process of obtaining the classification results of a video frame sequence may include the following:
[0208] 1. Obtain the first feature map of the video frame sequence through the video detection model. The first feature map is used to indicate the changes of features related to the key points of the target object in the video frame sequence in the time dimension.
[0209] Optionally, for any video frame sequence, the video frame sequence is converted into a word embedding block sequence, and then the word embedding block sequence is used to extract features through the temporal feature extraction network in the trained video detection model to obtain the first feature map.
[0210] Optionally, when the key points are facial key points of the target object, the above features are facial features of the target object.
[0211] 2. Based on the first feature map, the video detection model obtains the second feature map. The second feature map is used to indicate the changes in the spatial dimension of the features related to the key points of the target object in the video frame sequence.
[0212] Optionally, the second feature map can be obtained by extracting features from the first feature map using the spatial feature extraction network in the trained video detection model. Since the spatial feature extraction network operates in the spatial dimension, the first feature map can be transformed into a spatial dimension before processing it.
[0213] For example, the process may include the following: performing a dimensional transformation on the first feature map to obtain a transformed first feature map, wherein the feature map slices in the transformed first feature map are stacked in spatial dimensions; and obtaining a second feature map based on the transformed first feature map using a video detection model, such as by extracting features from the transformed first feature map using a spatial feature extraction network in the video detection model.
[0214] 3. Based on the second feature map, the video detection model is used to obtain the classification results of the video frame sequence.
[0215] Optionally, the output of the last block in the video detection model is used as the final second feature map. Then, the I3D Head in the video detection model is used to extract features from the second feature map to obtain a spatiotemporal feature map. Finally, the classifier in the video detection model is used to classify the spatiotemporal feature map to obtain the classification result of the video frame sequence.
[0216] The classification results can be used to indicate whether a video frame sequence is fake. For example, for facial keypoints, the classification results of a video frame sequence can be used to indicate whether the facial features of the target object exhibit FFD (Facial Feature Deformation).
[0217] Step 903: Obtain the detection result of the video based on the classification results corresponding to at least one video frame sequence. The detection result is used to indicate whether there is an anomaly in the video.
[0218] Optionally, if at least one video frame sequence contains an abnormal video frame sequence, it can be determined that the video is abnormal, that is, the video is fake.
[0219] For example, for facial key points, if at least one video frame sequence contains a video frame sequence exhibiting FFD (Facial Failure Distortion), then the video can be determined to be a deepfake video.
[0220] In some embodiments, referring to Figures 10 to 14, taking an identity verification scenario as an example, in response to a user's identity verification operation, the client displays an information authentication interface 1001. The information authentication interface 1001 displays multiple input options, such as name input, ID number input, and ID card photo input. After the user completes information input, the information authentication interface 1001 updates and displays prompt information to remind the user to record video under certain conditions. In response to the user triggering the start of video recording, the information authentication interface 1001 displays a recording area, which can display the video recorded by the user in real time. After the user completes the recording of the facial video as required, the client calls a trained video detection model to detect the facial video to determine whether it is a real video. If the facial video is a real video, and both information authentication and facial detection pass, the information authentication interface 1001 displays an authentication success notification to inform the user that authentication has been successful.
[0221] In some embodiments, for a sequence of repeated video frames constructed by copying a single image, and a sequence of forged video frames, the first feature map and the second feature map extracted by the trained video detection model are shown in Figure 15. For each video frame in the repeated video frame sequence 1501, the corresponding first feature map is the same (i.e. there is no change in the time dimension), and the corresponding second feature map is also the same (i.e. there is no change in the spatial dimension).
[0222] For each video frame in the forged video frame sequence 1502, its corresponding first feature map can reflect reasonable time-related motion (such as mouth movement), and its corresponding second feature map can also reflect reasonable spatial-related motion. This proves that the video detection model in this application embodiment can learn spatial and temporal information simultaneously.
[0223] In some embodiments, experiments were conducted on several commonly used datasets to compare generalization performance with related techniques. These datasets include: FaceForensics++ (FF++), Deepfake Detection Challenge (DFDC), a preview version of DFDC (DFDCP), DeepfakeDetection (DFD), and CelebDF (CDF).
[0224] The embodiments of this application are trained on the FF+ dataset and tested on other datasets for generalization. As shown in Table 1, the technical solution provided by the embodiments of this application can effectively improve the generalization of the video detection model.
[0225] Table 1
[0226] As shown in Table 2, a comparison of the network parameters of video detection models in related technologies reveals that the video detection model in this application embodiment has the fewest parameters and performs best on the CDF-v2 and DFDC datasets.
[0227] Table 2
[0228] In summary, the technical solution provided in this application, since the video detection model can detect whether there are anomalies in the frequency frame sequence samples in both spatial and temporal dimensions, is beneficial to improving the accuracy of anomaly detection.
[0229] Furthermore, the trained video detection model's ability to detect random jitter and drift of key points improves its accuracy. Additionally, for unseen deepfake videos, the trained model can also detect anomalies in key points to determine the video's authenticity, thus effectively enhancing its generalization ability.
[0230] The following are embodiments of the apparatus described in this application, which can be used to execute the embodiments of the method described in this application. For details not disclosed in the apparatus embodiments of this application, please refer to the embodiments of the method described in this application.
[0231] Referring to Figure 16, a block diagram of a training apparatus for a video detection model according to an embodiment of this application is shown. This apparatus has the functionality to implement the method example described above. The apparatus can be a computer device as described above, or it can be installed within a computer device. As shown in Figure 16, the apparatus 1600 includes: a frame sequence acquisition module 1601, a key point perturbation module 1602, and a detection model training module 1603.
[0232] The frame sequence acquisition module 1601 is used to acquire a video frame sequence, which includes multiple video frames related to the target object and ordered chronologically.
[0233] The key point perturbation module 1602 is used to perturb the key points of the target object in each of the m video frames in the video frame sequence to obtain a video frame sequence sample. The features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence sample, where m is a positive integer.
[0234] The detection model training module 1603 is used to train a video detection model based on the video frame sequence samples to obtain a trained video detection model, which is used to detect whether there are any anomalies in the video.
[0235] In some embodiments, as shown in FIG17, the key point perturbation module 1602 includes: a pixel perturbation submodule 1602a, a video frame fusion module 1602b, and a sample acquisition submodule 1602c.
[0236] The pixel perturbation submodule 1602a is used to perturb each pixel in any one of the m video frames to obtain a perturbed video frame, wherein the pixels in the perturbed video frame are different from the pixels in the original video frame.
[0237] The video frame fusion module 1602b is used to fuse the disturbed video frame and the video frame to obtain a fake video frame. The key points of the target object in the fake video frame are different from the key points of the target object in the video frame. The non-key points of the target object in the fake video frame are not different from the non-key points of the target object in the video frame.
[0238] The sample acquisition submodule 1602c is used to replace the m video frames with the fake video frames corresponding to the m video frames respectively, so as to obtain the video frame sequence sample.
[0239] In some embodiments, the video frame fusion module 1602b is configured to:
[0240] Obtain multiple key point groups of the target object in the video frame, each key point group including multiple key points;
[0241] For any one of the multiple keypoint groups, a mask image of the video frame is constructed based on the keypoint group, and the mask image is used to indicate the area of the keypoint group in the video frame;
[0242] Under the constraints of the mask images corresponding to the multiple key point groups, the perturbed video frame and the video frame are fused to obtain the forged video frame.
[0243] In some embodiments, the video frame fusion module 1602b is further configured to:
[0244] For any pixel in the video frame, obtain the first distance between the pixel and the region where the keypoint group is located;
[0245] Based on the first distance, the mask value of the pixel is determined, and the mask value of the pixel is negatively correlated with the first distance;
[0246] The mask image is constructed based on the mask values of each pixel in the video frame.
[0247] In some embodiments, the video frame fusion module 1602b is further configured to:
[0248] For any mask image in the mask images corresponding to the multiple key point groups, the perturbed video frame is adjusted based on the mask image to obtain a first mixed video frame, and the video frame is adjusted based on the inverse mask image of the mask image to obtain a second mixed video frame. In the mask image and the inverse mask image, the sum of the mask values of pixels at the same position is 1.
[0249] Based on the first mixed video frame and the second mixed video frame, the mixed video frame of the mask image is obtained;
[0250] The fake video frame is obtained by fusing multiple masked images and their corresponding mixed video frames.
[0251] In some embodiments, the video frame fusion module 1602b is further configured to:
[0252] For any pixel in the perturbed video frame, multiply the pixel value by the mask value corresponding to the pixel in the mask image to obtain the first mixed video frame;
[0253] For any pixel in the video frame, the pixel is multiplied by the mask value corresponding to the pixel in the inverse mask image to obtain the second mixed video frame.
[0254] In some embodiments, the pixel perturbation submodule 1602a is further configured to:
[0255] Obtain an affine transformation matrix, wherein the affine transformation matrix has perturbation parameters as elements, and the perturbation parameters are used to change the position of the pixel, wherein the method of changing the position of the pixel includes at least one of the following: rotation, translation, scaling;
[0256] The perturbed video frame is obtained by performing an affine transformation on the video frame according to the affine transformation matrix.
[0257] In some embodiments, as shown in FIG17, the detection model training module 1603 further includes: a feature map acquisition submodule 1603a, a classification result acquisition submodule 1603b, and a detection model training submodule 1603c.
[0258] The feature map acquisition submodule 1603a is used to acquire a first feature map of the video frame sequence sample through the video detection model. The first feature map is used to indicate the change of the feature corresponding to the target object in the video frame sequence sample in the time dimension.
[0259] The feature map acquisition module 1603a is further configured to obtain a second feature map based on the first feature map using the video detection model. The second feature map is used to indicate the spatial dimension changes of the features corresponding to the target object in the video frame sequence samples.
[0260] The classification result acquisition submodule 1603b is used to obtain the classification result of the video frame sequence sample based on the second feature map through the video detection model. The classification result of the video frame sequence sample is used to indicate whether there are any abnormalities in the changes of features related to the key points of the target object in the multiple video frames contained in the video frame sequence sample.
[0261] The detection model training submodule 1603c is used to train the video detection model based on the classification results to obtain the trained video detection model.
[0262] In some embodiments, the feature map acquisition module 1603a is further configured to:
[0263] Perform a dimensionality transformation on the first feature map to obtain the transformed first feature map;
[0264] The video detection model obtains the second feature map based on the transformed first feature map; wherein the feature map slices in the first feature map are stacked in the time dimension, and the feature map slices in the transformed first feature map are stacked in the spatial dimension.
[0265] In some embodiments, the video detection model includes a temporal feature extraction network and a spatial feature extraction network. The temporal feature extraction network is constructed based on a first pre-trained neural network and a first adapter, and the spatial feature extraction network is constructed based on a second pre-trained neural network and a second adapter. The first adapter is used to obtain the first feature map by combining with the first pre-trained neural network, and the second adapter is used to obtain the second feature map by combining with the second pre-trained neural network. The detection model training submodule 1603c is further configured to:
[0266] Based on the classification results, determine the loss function value of the video detection model;
[0267] By fixing the parameters of the first pre-trained neural network and the second pre-trained neural network, and adjusting the parameters of the first adapter and the second adapter according to the loss function value, the trained video detection model is obtained.
[0268] In some embodiments, when the key point is a facial key point of the target object, the feature is a facial feature of the target object.
[0269] In summary, the technical solution provided in this application, for video frame sequences related to a target object, constructs video frame sequence samples by perturbing the key points of the target object. This simulates the "random jitter and drift of key points" phenomenon commonly found in deepfake videos (videos constructed using deepfake technology). By training a video detection model using these video frame sequence samples, the trained video detection model can detect "random jitter and drift of key points." Consequently, the trained video detection model can accurately determine whether the video corresponding to the video frame sequence is a deepfake video based on whether the video frame sequence exhibits the "random jitter and drift of key points," thereby improving the detection accuracy of the video detection model.
[0270] In addition, since the trained video detection model has the ability to detect "random jitter and drift of key points", it can also detect the anomalies of "key points" in unseen deepfake videos to determine the authenticity of the video, thereby effectively improving the generalization of the video detection model.
[0271] Referring to Figure 18, a block diagram of a video detection apparatus provided in one embodiment of this application is shown. This apparatus has the functionality to implement the method example described above. The apparatus can be a computer device as described above, or it can be installed within a computer device. As shown in Figure 18, the apparatus 1800 includes: a frame sequence acquisition module 1801, a classification result acquisition module 1802, and an inspection result acquisition module 1803.
[0272] The frame sequence acquisition module 1801 is used to acquire at least one video frame sequence of a video, the video frame sequence including multiple video frames related to the target object and ordered in chronological order.
[0273] The classification result acquisition module 1802 is used to acquire the classification result of any video frame sequence in the at least one video frame sequence through a video detection model. The classification result of the video frame sequence is used to indicate whether there are abnormal changes in the features related to the key points of the target object in the multiple video frames included in the video frame sequence. The video detection model is trained on video frame sequence samples. For m video frames in the video frame sequence samples, the key points of the target object are perturbed, and the features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence samples, where m is a positive integer.
[0274] The inspection result acquisition module 1803 is used to acquire the detection result of the video based on the classification results corresponding to the at least one video frame sequence, and the detection result is used to indicate whether there is an anomaly in the video.
[0275] In some embodiments, the classification result acquisition module 1802 is further configured to:
[0276] The video detection model is used to obtain a first feature map of the video frame sequence. The first feature map is used to indicate the changes of features related to the key points of the target object in the video frame sequence in the time dimension.
[0277] Based on the first feature map, the video detection model obtains a second feature map, which is used to indicate the changes of features related to the key points of the target object in the video frame sequence in the spatial dimension.
[0278] The video detection model obtains the classification result of the video frame sequence based on the second feature map.
[0279] In some embodiments, the classification result acquisition module 1802 is further configured to:
[0280] Perform a dimensionality transformation on the first feature map to obtain the transformed first feature map;
[0281] The video detection model obtains the second feature map based on the transformed first feature map; wherein the feature map slices in the first feature map are stacked in the time dimension, and the feature map slices in the transformed first feature map are stacked in the spatial dimension.
[0282] In some embodiments, the video detection model includes a temporal feature extraction network and a spatial feature extraction network. The temporal feature extraction network is constructed based on a first pre-trained neural network and a first adapter, and the spatial feature extraction network is constructed based on a second pre-trained neural network and a second adapter. The first adapter is used to obtain the first feature map by combining with the first pre-trained neural network, and the second adapter is used to obtain the second feature map by combining with the second pre-trained neural network.
[0283] In summary, the technical solution provided in this application, since the video detection model can detect whether there are anomalies in the frequency frame sequence samples in both spatial and temporal dimensions, is beneficial to improving the accuracy of anomaly detection.
[0284] Furthermore, the trained video detection model's ability to detect random jitter and drift of key points improves its accuracy. Additionally, for unseen deepfake videos, the trained model can also detect anomalies in key points to determine the video's authenticity, thus effectively enhancing its generalization ability.
[0285] It should be noted that the apparatus provided in the above embodiments is only illustrated by the division of the above functional modules when implementing its functions. In actual applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the device can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and the specific implementation process can be found in the method embodiments, which will not be repeated here.
[0286] Please refer to Figure 19, which shows a schematic diagram of a computer device provided in one embodiment of this application. This computer device can be any electronic device with data computing, processing, and storage functions, and can be implemented as a model training device 10 or a model usage device 20 in the implementation environment shown in Figure 1. Specifically, it may include the following:
[0287] The computer device 1900 includes a central processing unit (such as a CPU, GPU, or FPGA) 1901, a system memory 1904 including RAM (Random-Access Memory) 1902 and ROM (Read-Only Memory) 1903, and a system bus 1905 connecting the system memory 1904 and the central processing unit 1901. The computer device 1900 also includes a basic input / output system 1906 to facilitate information transfer between various devices within the server, and a mass storage device 1907 for storing the operating system 1913, application programs 1914, and other program modules 1915.
[0288] The basic input / output system 1906 includes a display 1908 for displaying information and an input device 1909 for user input, such as a mouse or keyboard. Both the display 1908 and the input device 1909 are connected to the central processing unit 1901 via an input / output controller 1910 connected to the system bus 1905. The basic input / output system 1906 may also include the input / output controller 1910 for receiving and processing input from multiple other devices such as a keyboard, mouse, or electronic stylus. Similarly, the input / output controller 1910 also provides output to a display screen, printer, or other types of output devices.
[0289] The mass storage device 1907 is connected to the central processing unit 1901 via a mass storage controller (not shown) connected to the system bus 1905. The mass storage device 1907 and its associated computer-readable media provide non-volatile storage for the computer device 1900. That is, the mass storage device 1907 may include computer-readable media (not shown) such as a hard disk or a CD-ROM (Compact Disc Read-Only Memory) drive.
[0290] Without loss of generality, the computer-readable medium may include computer storage media and communication media. Computer storage media include volatile and non-volatile, removable and non-removable media implemented using any method or technology for storing information such as computer-readable instructions, data structures, program modules, or other data. Computer storage media include RAM, ROM, EPROM (Erasable Programmable Read-Only Memory), EEPROM (Electrically Erasable Programmable Read-Only Memory), flash memory or other solid-state storage technologies, CD-ROM, DVD (Digital Video Disc) or other optical storage, magnetic tape cassettes, magnetic tape, disk storage, or other magnetic storage devices. Of course, those skilled in the art will recognize that the computer storage medium is not limited to the above-mentioned types. The system memory 1904 and mass storage device 1907 described above can be collectively referred to as memory.
[0291] According to an embodiment of this application, the computer device 1900 can also be connected to a remote computer on a network, such as the Internet. That is, the computer device 1900 can be connected to the network 1912 via the network interface unit 1911 connected to the system bus 1905, or the network interface unit 1911 can be used to connect to other types of networks or remote computer systems (not shown).
[0292] The memory also includes a computer program stored in the memory and configured to be executed by one or more processors to implement the training method of the video detection model or the video detection method described above.
[0293] In some embodiments, a computer-readable storage medium is also provided, wherein a computer program is stored therein, which, when executed by a processor, implements a training method for a video detection model or the aforementioned video detection method.
[0294] Optionally, the computer-readable storage medium may include: ROM (Read-Only Memory), RAM (Random-Access Memory), SSD (Solid State Drives), or optical disc, etc. The random access memory may include ReRAM (Resistance Random Access Memory) and DRAM (Dynamic Random Access Memory).
[0295] In some embodiments, a computer program product is also provided, the computer program product comprising a computer program stored in a computer-readable storage medium. A processor of a computer device reads the computer program from the computer-readable storage medium, and the processor executes the computer program, causing the computer device to perform a training method for a video detection model or the aforementioned video detection method.
[0296] It should be noted that, in this application embodiment, before and during the collection of user-related data, a prompt interface, pop-up window, or voice prompt message can be displayed. This prompt interface, pop-up window, or voice prompt message is used to inform the user that their relevant data is being collected. This ensures that the application only begins executing the steps related to collecting user-related data after receiving confirmation from the user regarding the prompt interface or pop-up window; otherwise (i.e., without receiving confirmation from the user), the steps to collect user-related data end, meaning no user-related data is collected. In other words, all user data collected in this application is processed strictly in accordance with the requirements of relevant national laws and regulations. The informed consent or separate consent of the personal information subject is obtained only with the user's consent and authorization. Subsequent data use and processing are conducted within the scope of laws, regulations, and the authorization of the personal information subject. Furthermore, the collection, use, and processing of relevant user data must comply with the relevant laws, regulations, and standards of the relevant countries and regions. For example, the videos, video frame sequences, and features involved in this application are all obtained with full authorization.
[0297] It should be understood that "multiple" as used herein refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. Furthermore, the step numbers described herein are merely illustrative of one possible execution order. In some other embodiments, the steps may not be executed in numerical order, such as two steps with different numbers being executed simultaneously, or two steps with different numbers being executed in the reverse order of the illustration. This application does not limit this.
[0298] The above description is merely an exemplary embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A method for training a video detection model, the method being executed by a computer device, the method comprising: Obtain a video frame sequence, the video frame sequence including multiple video frames related to the target object, ordered chronologically; For m video frames in the video frame sequence, the key points of the target object in the m video frames are perturbed respectively to obtain a video frame sequence sample. The features related to the key points of the target object show anomalies in the changes of multiple video frames included in the video frame sequence sample, where m is a positive integer. A video detection model is trained based on the video frame sequence samples to obtain a trained video detection model, which is used to detect whether there are any anomalies in the video.
2. The method according to claim 1, wherein, The process of perturbing the key points of the target object in the m video frames to obtain video frame sequence samples includes: For any video frame among the m video frames, perturb each pixel in the video frame to obtain a perturbed video frame, wherein there is a difference between the pixels in the perturbed video frame and the pixels in the video frame; By fusing the disturbed video frame and the video frame, a fake video frame is obtained. The key points of the target object in the fake video frame differ from the key points of the target object in the video frame. The non-key points of the target object in the fake video frame do not differ from the non-key points of the target object in the video frame. The m video frames are replaced with the fake video frames corresponding to the m video frames respectively to obtain the video frame sequence sample.
3. The method according to claim 2, wherein, The process of fusing the disturbed video frame and the video frame to obtain the forged video frame includes: Obtain multiple key point groups of the target object in the video frame, each key point group including multiple key points; For any one of the multiple keypoint groups, a mask image of the video frame is constructed based on the keypoint group, and the mask image is used to indicate the area of the keypoint group in the video frame; Under the constraints of the mask images corresponding to the multiple key point groups, the perturbed video frame and the video frame are fused to obtain the forged video frame.
4. The method according to claim 3, wherein, The step of constructing the mask image corresponding to the video frame based on the key point group includes: For any pixel in the video frame, obtain the first distance between the pixel and the region where the keypoint group is located; Based on the first distance, the mask value of the pixel is determined, and the mask value of the pixel is negatively correlated with the first distance; The mask image is constructed based on the mask values of each pixel in the video frame.
5. The method according to claim 3 or 4, wherein, The process of fusing the perturbed video frame and the video frame to obtain the forged video frame, under the constraints of the mask images corresponding to the multiple key point groups, includes: For any mask image in the mask images corresponding to the multiple key point groups, the perturbed video frame is adjusted based on the mask image to obtain a first mixed video frame, and the video frame is adjusted based on the inverse mask image of the mask image to obtain a second mixed video frame. In the mask image and the inverse mask image, the sum of the mask values of pixels at the same position is 1. Based on the first mixed video frame and the second mixed video frame, the mixed video frame of the mask image is obtained; The fake video frame is obtained by fusing multiple masked images and their corresponding mixed video frames.
6. The method according to claim 5, wherein, The step of adjusting the perturbed video frame based on the mask image to obtain a first mixed video frame, and adjusting the video frame based on the inverse mask image corresponding to the mask image to obtain a second mixed video frame, includes: For any pixel in the perturbed video frame, multiply the pixel value by the mask value corresponding to the pixel in the mask image to obtain the first mixed video frame; For any pixel in the video frame, the pixel is multiplied by the mask value corresponding to the pixel in the inverse mask image to obtain the second mixed video frame.
7. The method according to any one of claims 2 to 6, wherein, The process of perturbing each pixel in the video frame to obtain a perturbed video frame includes: Obtain an affine transformation matrix, wherein the affine transformation matrix has perturbation parameters as elements, and the perturbation parameters are used to change the position of the pixel, wherein the method of changing the position of the pixel includes at least one of the following: rotation, translation, scaling; The perturbed video frame is obtained by performing an affine transformation on the video frame according to the affine transformation matrix.
8. The method according to any one of claims 1 to 7, wherein, The process of training a video detection model based on the video frame sequence samples to obtain the trained video detection model includes: The video detection model is used to obtain a first feature map of the video frame sequence sample, and the first feature map is used to indicate the change of the feature corresponding to the target object in the video frame sequence sample in the time dimension. The video detection model obtains a second feature map based on the first feature map. The second feature map is used to indicate the spatial dimension changes of the features corresponding to the target object in the video frame sequence samples. The video detection model obtains the classification result of the video frame sequence sample based on the second feature map. The classification result of the video frame sequence sample is used to indicate whether there are abnormalities in the changes of the features related to the key points of the target object in the multiple video frames contained in the video frame sequence sample. The video detection model is trained based on the classification results to obtain the trained video detection model.
9. The method according to claim 8, wherein, The step of obtaining a second feature map based on the first feature map using the video detection model includes: Perform a dimensionality transformation on the first feature map to obtain the transformed first feature map; The second feature map is obtained by the video detection model based on the transformed first feature map; In this process, the feature map slices in the first feature map are stacked in the time dimension, and the feature map slices in the transformed first feature map are stacked in the spatial dimension.
10. The method according to claim 8 or 9, wherein, The video detection model includes a temporal feature extraction network and a spatial feature extraction network. The temporal feature extraction network is constructed based on a first pre-trained neural network and a first adapter. The spatial feature extraction network is constructed based on a second pre-trained neural network and a second adapter. The first adapter is used to obtain the first feature map by combining with the first pre-trained neural network, and the second adapter is used to obtain the second feature map by combining with the second pre-trained neural network. The step of training the video detection model based on the classification result to obtain the trained video detection model includes: Based on the classification results, determine the loss function value of the video detection model; By fixing the parameters of the first pre-trained neural network and the second pre-trained neural network, and adjusting the parameters of the first adapter and the second adapter according to the loss function value, the trained video detection model is obtained.
11. The method according to any one of claims 1 to 10, wherein, When the key point is a facial key point of the target object, the feature is the facial feature of the target object.
12. A video detection method, the method being executed by a computer device, the method comprising: Acquire at least one video frame sequence of a video, the video frame sequence comprising multiple video frames related to the target object in chronological order; For any video frame sequence in the at least one video frame sequence, a classification result of the video frame sequence is obtained through a video detection model. The classification result of the video frame sequence is used to indicate whether there are abnormal changes in the features related to the key points of the target object in the multiple video frames included in the video frame sequence. The video detection model is trained with video frame sequence samples. For m video frames in the video frame sequence samples, the key points of the target object are perturbed, and the features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence samples, where m is a positive integer. Based on the classification results corresponding to the at least one video frame sequence, the detection result of the video is obtained, and the detection result is used to indicate whether there is an anomaly in the video.
13. The method according to claim 12, wherein, The step of obtaining the classification result of the video frame sequence through the video detection model includes: The video detection model is used to obtain a first feature map of the video frame sequence. The first feature map is used to indicate the changes of features related to the key points of the target object in the video frame sequence in the time dimension. Based on the first feature map, the video detection model obtains a second feature map, which is used to indicate the changes of features related to the key points of the target object in the video frame sequence in the spatial dimension. The video detection model obtains the classification result of the video frame sequence based on the second feature map.
14. The method according to claim 13, wherein, The step of obtaining a second feature map based on the first feature map using the video detection model includes: Perform a dimensionality transformation on the first feature map to obtain the transformed first feature map; The second feature map is obtained by the video detection model based on the transformed first feature map; In this process, the feature map slices in the first feature map are stacked in the time dimension, and the feature map slices in the transformed first feature map are stacked in the spatial dimension.
15. The method according to any one of claims 12 to 14, wherein, The video detection model includes a temporal feature extraction network and a spatial feature extraction network. The temporal feature extraction network is constructed based on a first pre-trained neural network and a first adapter. The spatial feature extraction network is constructed based on a second pre-trained neural network and a second adapter. The first adapter is used to obtain the first feature map by combining with the first pre-trained neural network, and the second adapter is used to obtain the second feature map by combining with the second pre-trained neural network.
16. A training apparatus for a video detection model, the apparatus comprising: A frame sequence acquisition module is used to acquire a video frame sequence, which includes multiple video frames related to the target object, ordered chronologically. The key point perturbation module is used to perturb the key points of the target object in each of the m video frames in the video frame sequence to obtain a video frame sequence sample. The features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence sample, where m is a positive integer. The detection model training module is used to train a video detection model based on the video frame sequence samples to obtain a trained video detection model, which is used to detect whether there are any anomalies in the video.
17. A video detection device, the device comprising: A frame sequence acquisition module is used to acquire at least one video frame sequence of a video, wherein the video frame sequence includes multiple video frames related to the target object and ordered chronologically. The classification result acquisition module is used to acquire the classification result of any video frame sequence in the at least one video frame sequence through a video detection model. The classification result of the video frame sequence is used to indicate whether there are abnormal changes in the features related to the key points of the target object in the multiple video frames included in the video frame sequence. The video detection model is trained on video frame sequence samples. For m video frames in the video frame sequence samples, the key points of the target object are perturbed, and the features related to the key points of the target object show abnormal changes in the multiple video frames included in the video frame sequence samples, where m is a positive integer. The inspection result acquisition module is used to acquire the detection result of the video based on the classification results corresponding to the at least one video frame sequence, and the detection result is used to indicate whether there is an anomaly in the video.
18. A computer device comprising a processor and a memory, the memory storing a computer program, the computer program being loaded and executed by the processor to implement a training method for a video detection model as claimed in any one of claims 1 to 11, or a video detection method as claimed in any one of claims 12 to 15.
19. A computer-readable storage medium storing a computer program, the computer program being loaded and executed by a processor to implement a training method for a video detection model as claimed in any one of claims 1 to 11, or a video detection method as claimed in any one of claims 12 to 15.
20. A computer program product comprising a computer program stored in a computer-readable storage medium, wherein a processor reads from and executes the computer program to implement a training method for a video detection model as claimed in any one of claims 1 to 11, or a video detection method as claimed in any one of claims 12 to 15.
Citation Information
Patent Citations
Video content identification method and device, storage medium and electronic equipment
CN111241985A
Head decoration processing method and device based on artificial intelligence
CN111563868A
False face video identification method and system and readable storage medium
CN111967427A
Video anomaly detection model training method, video anomaly detection method and device
CN113435432A
Video generation method, video model training method and electronic equipment
CN117633296A
Cited By
Depth forgery detection method based on active high-frequency texture stripping and difference graph reasoning
CN122199525A