Forgery detection method, model training method, equipment, storage medium and product
By extracting multi-frame face images from video, performing clustering analysis and obtaining context information, as a priori constraint of the forged detection model, the problem of difficulty in detecting deep forged content in the prior art is solved, and higher detection accuracy and robustness are achieved.
Patent Information
- Application Number
- CN202510374642.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-26
- Publication Date
- 2025-06-27
AI Technical Summary
The prior art is difficult to effectively detect deep forged content, especially in videos, and traditional single-frame image forged detection methods have become difficult to meet the needs.
By extracting multiple frames of face images from the video to be detected, clustering analysis is performed, and context information of the face image, including clustering tags, face features, relative position information and timing features, it is used as a priori constraint of the forged detection model, forged detection is performed.
It improves the accuracy and robustness of forgery detection, reduces misjudgment, and can more accurately identify forgery behavior, breaking through the limitations of single-frame detection.
Smart Images

Figure CN120219936A_ABST
Abstract
Description
Technical Field
[0001] One or more embodiments of this specification relate to the field of image processing technology, and in particular, to a forgery detection method, a training method for a forgery detection model, an electronic device, a computer-readable storage medium, and a computer program product. Background Art
[0002] With the rapid development of artificial intelligence technology, the application of deep learning models in the field of image generation has made breakthrough progress, driving the innovation of digital content creation. However, deepfake content based on technologies such as generative adversarial networks and diffusion models presents the characteristics of hyper-realism, scene generalization, and industrial production, and the forged faces generated by it have broken through the defense boundaries of traditional identification means. Such technologies have been misused in scenarios such as the spread of false news and identity theft fraud, seriously threatening public information security, personal privacy rights and interests, and the social trust system.
[0003] Traditional forgery detection methods mostly rely on single-frame images for analysis. However, due to the high authenticity and concealment of deepfake content, traditional single-frame image forgery detection methods are no longer sufficient to meet the requirements. Summary of the Invention
[0004] In view of this, one or more embodiments of this specification provide a forgery detection method, a training method for a forgery detection model, an electronic device, a computer-readable storage medium, and a computer program product.
[0005] To achieve the above object, one or more embodiments of this specification provide the following technical solutions:
[0006] According to a first aspect of one or more embodiments of this specification, a forgery detection method is proposed, including:
[0007] Determine at least two frames of images to be detected from the video to be detected, and perform face extraction on each of the at least two frames of images to be detected to obtain multiple face images;
[0008] Cluster the multiple face images, and determine the cluster labels to which each face image belongs, where all face images belonging to the same cluster label are divided into the same group;
[0009] Obtain the context information of at least one group of face images, where the context information includes at least one of the following: the cluster label, the facial features of the face image, the relative position information of the face image in the image to be detected to which it belongs, and the temporal sequence features of the image to be detected to which the face image belongs;
[0010] Use the context information of the at least one set of face images as a prior constraint for the trained forgery detection model, and control the forgery detection model to perform forgery detection on the at least one set of face images to obtain a forgery detection result.
[0011] According to a second aspect of the embodiments of the present specification, a method for training a forgery detection model is provided, including:
[0012] Obtain a plurality of video samples labeled with supervision labels;
[0013] For each video sample, determine at least two frames of images to be detected from the video sample, perform face extraction on the at least two frames of images to be detected to obtain multiple face images, and perform clustering on the multiple face images to determine the clustering labels to which each of the face images belongs. Divide all the face images belonging to the same clustering label into the same group, and further obtain the context information of at least one set of face images. The context information includes at least one of the following: the clustering label, the facial features of the face image, the relative position information of the face image in the image to be detected to which it belongs, and the temporal feature of the image to be detected to which the face image belongs;
[0014] Input the at least one set of face images corresponding to the video sample and their context information into the forgery detection model to be trained, so that the forgery detection model refers to the context information of the at least one set of face images, performs forgery detection on the at least one set of face images to obtain a prediction result, and uses minimizing the error between the prediction result and the supervision label as an optimization goal to train the forgery detection model; wherein, the trained forgery detection model is applied to the forgery detection method described in the first aspect.
[0015] According to a third aspect of the embodiments of the present specification, an electronic device is provided, including:
[0016] A processor;
[0017] A memory for storing instructions executable by the processor;
[0018] Wherein, when the processor executes the executable instructions, it is used to implement the method described in the first aspect or the second aspect.
[0019] According to a fourth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, on which a computer program is stored, and when the program is executed by a processor, it implements the steps of the method described in the first aspect or the second aspect.
[0020] According to a fifth aspect of the embodiments of the present specification, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the steps of the method described in the first aspect or the second aspect.
[0021] The technical solutions provided by the embodiments of this specification may include the following beneficial effects:
[0022] In the embodiments of this specification, by combining face extraction, clustering analysis, and context information in the video, the forgery detection method can effectively improve the accuracy and robustness of detection. On the one hand, by extracting multiple frames of face images from the video and performing clustering, it is ensured that the facial images of the same person are correctly classified, which helps to reduce misjudgments. On the other hand, based on the use of at least one context information such as clustering labels, facial features, relative position information, and temporal features, prior information reference is provided for the forgery detection model, thereby breaking through the limitation of the single-frame detection's dependence on local texture and more accurately identifying forgery behaviors.
[0023] It should be understood that the above general description and the following detailed description are only exemplary and explanatory, and cannot limit this specification. Brief Description of the Drawings
[0024] Figure 1 is a flowchart of a method for training a forgery detection model provided by an exemplary embodiment.
[0025] Figure 2 is a schematic diagram of the training of a forgery detection model provided by an exemplary embodiment.
[0026] Figure 3 is a schematic structural diagram of a forgery detection model provided by an exemplary embodiment.
[0027] Figure 4 is a flowchart of a forgery detection method provided by an exemplary embodiment.
[0028] Figure 5 is a schematic structural diagram of an electronic device provided by an exemplary embodiment. Detailed Description of the Embodiments
[0029] Here, the exemplary embodiments will be described in detail, and the examples are shown in the drawings. When the following description refers to the drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with one or more embodiments of this specification. On the contrary, they are only examples of devices and methods consistent with some aspects of one or more embodiments of this specification as detailed in the appended claims.
[0030] It should be noted that: in other embodiments, the steps of the corresponding method are not necessarily executed in the order shown and described in this specification. In some other embodiments, the steps included in the method may be more or less than those described in this specification. In addition, a single step described in this specification may be decomposed into multiple steps for description in other embodiments; and multiple steps described in this specification may also be combined into a single step for description in other embodiments.
[0031] The user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this specification are all information and data that have been authorized by the user or fully authorized by all parties. And the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or refuse.
[0032] Based on the problems in the related art, the embodiments of this specification provide a method for training a forgery detection model and a forgery detection method based on the forgery detection model.
[0033] Please refer to Figure 1 and Figure 2 , the embodiments of this specification provide a method for training a forgery detection model, which can be executed by an electronic device, and the electronic device includes but not limited to a physical server, a server cluster, a cloud server, a smart phone / mobile phone, a tablet computer, a personal digital assistant (PDA), a laptop computer, and a desktop computer, etc. The training method includes:
[0034] In S101, obtain a plurality of video samples labeled with supervision labels.
[0035] In this step, by obtaining video samples with supervision labels, real data is provided for subsequent training to help the model learn how to distinguish forged and real video content.
[0036] Exemplarily, the supervision label includes at least one of the following: a first label for describing whether there is a forged portrait in the video sample, and a second label for describing whether there is a forgery situation for each person in the video sample.
[0037] The first label helps the forgery detection model to be trained to judge whether a video is forged as a whole, and is applicable to the judgment of the authenticity of the whole video. If there are forged characters in the video (for example, the face is replaced, the image is tampered with, etc.), the model should be able to judge that the video is forged and output a forgery label; otherwise, it outputs a non-forgery label. The first label enables the forgery detection model to be trained to judge the authenticity of the video at a high level, focusing on whether there is forged false information in the video, not limited to forgeries in a single frame, but covering the entire video.
[0038] The second label focuses on the forgery situation of each character in the video and helps the model judge which specific characters are forged from a fine-grained perspective. It allows the forgery detection model to be trained to more accurately identify the forged characters in the video and provide more specific forgery information. For example, one of the faces in the video may be forged through face-swapping technology, while the other characters are real. Through the second label, it can help the forgery detection model to be trained to distinguish the authenticity of each character. It is applicable to scenarios where there are multiple characters in the video, improving the meticulousness of forgery detection.
[0039] In S102, for each video sample, at least two frames of images to be detected are determined from the video sample, face extraction is performed on the at least two frames of images to be detected respectively to obtain multiple face images, and clustering is performed on the multiple face images to determine the cluster labels to which each face image belongs. All face images belonging to the same cluster label are divided into the same group, and then the context information of at least one group of face images is obtained. The context information includes at least one of the following: cluster label, facial features of the face image, relative position information of the face image in the image to be detected to which it belongs, and temporal features of the image to be detected to which the face image belongs.
[0040] The inventors found that in actual forgery scenarios, there are often one or more characters in the video to be detected, where the appearance duration of different characters is random, the face size is random, the relative position of the face in the video frame is random, and the angle of the facial orientation is random. And for forged videos made using deep forgery technology, restricted by the current forgery technology and forgery process, there are often larger forgery traces when the forged face is in a large-angle downward, upward, or shaking head situation, and it is easier to distinguish forgery compared to facial data with a frontal orientation. Based on the characteristics of the actual forgery scenario, before training the forgery detection model to be trained, for each video sample, by performing step S102, extracting multiple face images from each video sample, performing clustering, and generating the context information of each face image, it can effectively enhance the judgment basis of the model when facing forgery scenarios, thereby improving the accurate detection ability of the model for forged videos.
[0041] The following gives an exemplary description of determining at least two frames of images to be detected from each video sample:
[0042] In a possible implementation, the electronic device may randomly extract at least two frames of images to be detected from the video sample. Randomly selecting at least two frames of images to be detected helps the model learn forgery behaviors in more different scenarios and conditions. Since the video content and background are dynamically changing, random extraction can better cover different time periods and changing conditions, making the training data more representative.
[0043] In another possible implementation, the electronic device samples at least two frames of images to be detected from the video sample according to a preset sampling rate. Using a fixed sampling rate can more precisely control the quantity of the input data of the model, ensuring the balance and training efficiency of the training data set. The specific value of the sampling rate can be specifically set according to the actual application scenario, and this embodiment does not make any limitation thereto.
[0044] In yet another possible implementation, the electronic device may sample multiple candidate images from the video sample according to a preset sampling rate, then perform quality detection on each candidate image in at least one dimension to obtain the quality score of each candidate image in at least one dimension, and finally determine the candidate images whose quality scores in at least one dimension meet the preset quality conditions as the images to be detected. Among them, the preset quality conditions can be specifically set according to the actual application scenario, and this embodiment does not make any limitation thereto. Through quality detection, it can be ensured that the images used for training and forgery detection have high quality, avoiding the negative impact of low-quality images (such as blurred or images containing a large amount of background noise) on training.
[0045] Exemplarily, a trained quality detection model corresponds to each dimension. The electronic device can input each candidate image into the quality detection model corresponding to each dimension for quality detection, so as to obtain the quality score of this candidate image in this dimension. For example, there are N candidate images and K quality detection models corresponding to K dimensions. Inputting each candidate image into the K quality detection models corresponding to K dimensions for quality detection respectively, K quality scores corresponding to this candidate image can be obtained, so that at least two high-quality images to be detected can be screened out from the N candidate images based on the K quality scores corresponding to each candidate image.
[0046] It can be understood that the quality detection models corresponding to each dimension can directly apply the trained models in related technologies, or, training samples labeled with labels can also be collected, and then the quality detection models to be trained can be trained by using a supervised learning method. This embodiment does not make any limitation thereto.
[0047] Exemplarily, the above-mentioned dimensions include at least one of the following: detail loss, visual information fidelity, human motion artifact loss, texture consistency, and illumination consistency, etc., but not limited thereto. Among them, (1) Detail loss is used to measure the loss degree of high-frequency information (such as edges, textures, and tiny structures) in the candidate image. For example, in video compression, super-resolution reconstruction, or forgery generation tasks, the details of the original video frame may be lost due to algorithm processing. Detail loss usually manifests as image blurring, unclear edges, or texture distortion. (2) Visual information fidelity is used to measure the authenticity of the candidate image in human visual perception. Visual fidelity emphasizes the subjective perception of the image by the human visual system, rather than simply pixel-level differences. It not only focuses on details but also covers global features such as color, illumination, and contrast. (3) Human motion artifact loss is used to measure the artifacts (such as motion blur, ghosting, and edge breaks) generated in the candidate image due to inaccurate motion estimation or compensation. These artifacts are usually related to human motion or camera motion. Motion blur is the blurring caused by the misalignment of moving objects between frames, ghosting is the repeated appearance of moving objects between frames, and edge breaks are the discontinuity of the edges of moving objects. (4) Texture consistency is used to measure whether the facial and background textures in the candidate image are natural and coherent. (5) Illumination consistency is used to measure whether the illumination and shadows in the candidate image are reasonable and consistent.
[0048] Exemplarily, after confirming at least two images to be detected in the above manner, face decoupling can be performed on the at least two images to be detected, that is, face extraction is performed on each of the at least two frames of images to be detected to obtain multiple face images. In a possible implementation manner, for each frame of the image to be detected, a face detection algorithm (such as a multi-task cascaded convolutional network, Haar feature cascaded classifier, etc.) can be applied to detect and locate all the faces in the image to be detected. Then, for the detected face regions in each frame, face images are extracted by cropping. Finally, each frame of the image to be detected will be extracted into independent face images. Even if there are multiple faces in one frame, these faces will be processed as independent images.
[0049] Furthermore, multiple face images can be clustered to determine the cluster labels to which each face image belongs, and all face images belonging to the same cluster label are grouped into the same group. In a possible implementation, the electronic device can use a trained feature extraction model (such as a face recognition model based on a convolutional neural network) to extract the face features of each face image. The feature extraction model can generate the face features of each face image by analyzing the key facial features in the face image (such as the relative positions of eyes, nose, and mouth, facial contours, etc.); then, based on the similarity between the face features of any two face images, multiple face images are clustered to determine the cluster labels to which each face image belongs. Clustering algorithms such as K-means, DBSCAN (Density-Based Spatial Clustering of Applications with Noise), and hierarchical clustering can be used to implement the above clustering process, but not limited to this. The cluster labels help identify whether there are multiple frames of the same person in the video. Especially when there are multiple similar people in the video frames, these faces can be classified as the same person through clustering. The face images belonging to the same group usually represent the images of the same person. It can be understood that the feature extraction model can directly apply the trained models in related technologies, or it can also collect the training samples labeled with labels and then use the supervised learning method to train the feature extraction model to be trained. This embodiment does not make any restrictions on this.
[0050] Through face image extraction, feature extraction, and similarity-based clustering, the above process can identify the same person in different frames of the video and assign a unified cluster label to each same person. The cluster labels not only help identify and group the same person but also provide important information for subsequent forgery detection, helping the model identify whether there are forged face images in the video.
[0051] For example, assume there are 5 frames of images to be detected, and 10 face images are extracted. After being processed by the feature extraction and clustering algorithms, the following cluster labels may be obtained:
[0052] Face image Cluster label Face 1 1 Face 2 2 Face 3 1 Face 4 3 Face 5 2 Face 6 1 Face 7 3 Face 8 2 Face 9 1 Face 10 3
[0053] Among them, the cluster label "1" indicates the same person A (for example, person A in the video). The cluster label "2" indicates the same person B. The cluster label "3" indicates the same person C.
[0054] Exemplarily, one or more than one group of face images can be selected for forgery detection according to actual needs. Then the electronic device can obtain context information of at least one group of face images, and the context information includes at least one of the following: clustering labels, facial features of each face image, relative position information of each face image in the image to be detected to which it belongs, and time sequence features of the image to be detected to which each face image belongs.
[0055] (1) In practical applications, the forgery detection method can be applied to various scenarios, such as news reports, movie clips, social media videos, etc. These scenarios may contain a variety of situations from a single person to multiple people. The clustering labels obtained by the clustering method can flexibly adapt to these different scenarios. Regardless of whether the video contains a single person, multiple protagonists, or a group of people, the model can detect forgery for each person based on the situation indicated by the clustering label. The processing is not limited by the number of people. No matter how many faces appear in the video, the model can detect the face of each person in a targeted manner based on the clustering results, thereby improving the accuracy and efficiency of detection.
[0056] (2) The facial features of face images help the model learn how to recognize the difference between natural facial movements and forged facial movements. For example, when a person lowers his head or shakes his head at a large angle, the forged image usually shows unnatural facial distortion, while the real image maintains a more natural change. The model can effectively distinguish the forged behavior through facial features. For example, in a forged video, the face of the "protagonist" character usually occupies a larger proportion, while the faces of secondary characters or background characters account for a relatively small proportion. Forgery technology often performs more forgery operations on the "protagonist" character. By analyzing the proportion of each character's face in the video, the forgery detection model can more effectively identify forgery traces.
[0057] Exemplarily, the facial features of the face image include at least one of the following: facial posture features, face proportion, and mouth opening and closing degree. The electronic device can use facial key point detection technology to extract facial key points of each face image, and the facial key points include but are not limited to eyes, nose, mouth, chin, etc.
[0058] Regarding the facial pose features of a face image, an electronic device can determine them based on the difference between the facial key points in the face image and the facial key points in a reference pose. For example, by calculating the distance between the facial key points in the face image and the facial key points in the reference pose, the facial orientation in the face image (such as front, side, head down, head up, etc.) can be determined. It can be understood that the embodiments of this specification do not impose any restrictions on the representation form of facial pose features, and can be specifically set according to the actual application scenario. For example, 0, 1, 2, 3, .., m facial orientation encodings can be given according to different facial orientations. Assuming that the facial orientation encoding corresponding to the facial key points in the reference pose is 0, after determining the difference between the facial key points in the face image and the facial key points in the reference pose, based on the pre-set mapping relationship between the difference and the facial orientation encoding, the facial orientation encoding corresponding to the face image is determined.
[0059] Regarding the degree of mouth opening and closing in a face image, an electronic device can determine the distance between the upper and lower lips based on the key points of the upper lip and the lower lip of the face in the face image, and then determine the degree of mouth opening and closing based on the ratio of the distance between the upper and lower lips of the face in the face image to a specified side of the face in the face image (such as the facial height). It can be understood that the embodiments of this specification do not impose any restrictions on the representation form of the degree of mouth opening and closing, and can be specifically set according to the actual application scenario. For example, the opening state encodings corresponding to different intervals of the ratio can be set, which are different opening state encodings of 0, 1, 2, 3, …, n. Furthermore, after the electronic device calculates the ratio used to represent the degree of mouth opening and closing of the face image, according to the interval where the ratio is located and the opening state encodings corresponding to the above different intervals, the opening state encoding corresponding to the face image is determined. The degree of mouth opening and closing is an important feature for judging forged images. Especially in deepfake technology, forged images often have significant problems in the details of the mouth. For example, the mouth in a forged image may appear unnatural, with inconsistent opening and closing, or the degree of opening and closing of the mouth is too large or too small. By analyzing the degree of mouth opening and closing, the model can identify these abnormalities and improve the accuracy of forgery detection.
[0060] The face proportion is used to describe the relative size of a face in an image to be detected. If the face size of the face image is the same as the display size of the face image in the image to be detected to which the face image belongs, then the face proportion is determined based on the ratio between the size of the face image and the size of the image to be detected to which the face image belongs. It can be understood that the embodiments of this specification do not impose any restrictions on the representation form of the face proportion, and can be specifically set according to the actual application scenario. For example, different face proportion encodings corresponding to different intervals of the face proportion can be set, which are different face proportion encodings of 0, 1, 2, 3, …, n respectively. Then, after calculating the face proportion of the face image, the electronic device can determine the face proportion encoding corresponding to the face image according to the interval where the face proportion is located and the face proportion encodings corresponding to the above different intervals. Alternatively, the face proportion information can also be directly used.
[0061] (3) In a forged video, especially in scenarios involving face swapping or deepfake technology, usually the "main character" will appear in the central position of the video, while secondary characters or background characters may appear in the edge area or the corner of the video. The relative position information obtained in this specification can help the model identify which character is the "main character" in the video. Since forgery techniques often perform more forgery operations on the "main character", identifying the positions of these main characters can help the model focus on these key characters and perform more refined forgery detection. The analysis of relative position helps to determine which characters' images are more important, especially in multi-character scenarios. The main character usually appears in the center of the video and appears frequently in multiple frames, and this position feature helps the model better capture the features related to forgery.
[0062] It can be understood that the embodiments of this specification do not impose any restrictions on the representation form of the relative position information, and can be specifically set according to the actual application scenario. For example, different relative position encodings corresponding to different image regions of the image to be detected can be set, which are different relative position encodings of 0, 1, 2, 3, …, n respectively. Then, after calculating the image region where the face image is located in the image to be detected to which the face image belongs, the electronic device can use the relative position encoding corresponding to the image region as the relative position information of the face image.
[0063] (4) Temporal features can help the model identify whether the actions of the characters in the video are coherent. Especially when the face makes large movements (such as bowing the head, raising the head, shaking the head), forgery techniques usually result in inconsistent or unnatural facial movements of the characters. The introduction of temporal features allows the model to consider the facial changes between different time points, so as to identify whether there is a temporal anomaly caused by forgery techniques.
[0064] In S103, at least one set of face images corresponding to the video sample and their context information are input into the forgery detection model to be trained. The forgery detection model refers to the context information of at least one set of face images to perform forgery detection on at least one set of face images to obtain a prediction result, and takes minimizing the error between the prediction result and the supervision label as the optimization objective to train the forgery detection model.
[0065] Among them, the context information of each set of face images corresponding to the video sample has multiple functions in the forgery detection model: (1) The context information enhances the model's ability to judge forged images by providing a richer description of each frame in the video. (2) The context information (such as temporal information and relative position information) can help the model understand the natural dynamics and spatial distribution of the people in the video. (3) When training the model, by inputting the context information, the model can refer to more features for learning and can more accurately identify forgery behaviors. Finally, using the context information to effectively supervise the forgery detection model can greatly improve the performance of forgery detection, especially when facing complex video forgery techniques.
[0066] In this step, at least one set of face images corresponding to the video sample and their context information can be used as the input for training the forgery detection model. By comparing the error between the prediction result and the supervision label, the loss function is used to minimize this error to optimize the parameters of the forgery detection model. The forgery detection model performs backpropagation based on the error between the prediction result and the supervision label. The smaller the error, the stronger the prediction ability of the forgery detection model. Through continuous iteration, the forgery detection model gradually learns how to better use the context information for forgery detection in each training. Finally, the forgery detection model can identify subtle forgery traces in the video, such as forgery behaviors that are not easily detectable in a single-frame image.
[0067] Exemplarily, if the model input includes at least one set of face images corresponding to the video sample and their context information, the corresponding supervision label can be the first label corresponding to the video sample, or the second label corresponding to each group of images, or the first label and the second label can also be used in combination. This embodiment does not make any restrictions on this.
[0068] Exemplarily, please refer to Figure 3 , the forgery detection model includes a feature extraction network (such as a convolutional neural network, a recurrent neural network, or a long short-term memory network, etc.) and a forgery detection network based on the Transformer architecture; the feature extraction network is used to extract features from at least one set of face images to obtain forgery detection features; the forgery detection network based on the Transformer architecture is used to perform forgery detection based on the fusion result of the context information of at least one set of face images and the forgery detection features to obtain a prediction result.
[0069] It is understandable that this embodiment does not impose any restrictions on the fusion method between context information and forgery detection features, and can be specifically set according to the actual application scenario. For example, the context information and forgery detection features can be directly concatenated, or the forgery detection features and context information can be adaptively fused by weighting, or the forgery detection features and context information can be directly added, and so on.
[0070] In some embodiments, after the forgery detection model is trained, the forgery detection model can be deployed to an electronic device that executes the forgery detection method. The electronic device includes but is not limited to physical servers, server clusters, cloud servers, smart phones / mobile phones, tablet computers, personal digital assistants (PDAs), laptop computers, and desktop computers, etc. Please refer to Figure 4 , which shows a schematic flowchart of a forgery detection method. The method includes:
[0071] In S401, at least two frames of images to be detected are determined from the video to be detected, and face extraction is respectively performed on the at least two frames of images to be detected to obtain multiple face images.
[0072] In this step, first, at least two frames of images to be detected are extracted from the video to be detected. There are various ways to select video frames.
[0073] In a possible implementation manner, the electronic device can randomly extract at least two frames of images to be detected from the video to be detected.
[0074] In another possible implementation manner, the electronic device samples at least two frames of images to be detected from the video to be detected according to a preset sampling rate. The specific value of the sampling rate can be specifically set according to the actual application scenario, and this embodiment does not impose any restrictions on it.
[0075] In yet another possible implementation manner, the electronic device can sample multiple candidate images from the video to be detected according to a preset sampling rate, then perform quality detection on each candidate image in at least one dimension to obtain the quality scores of each candidate image in at least one dimension, and finally determine the candidate images whose quality scores in at least one dimension meet the preset quality conditions as the images to be detected. Among them, the preset quality conditions can be specifically set according to the actual application scenario, and this embodiment does not impose any restrictions on it. Through quality detection, it can be ensured that the images used for forgery detection have high quality and avoid low-quality images (such as blurred or images containing a large amount of background noise) from affecting the detection results.
[0076] Exemplarily, each dimension corresponds to a trained quality detection model, and the electronic device can input each candidate image into the quality detection model corresponding to each dimension for quality detection, thereby obtaining the quality score of the candidate image in the dimension. For example, if there are N candidate images and quality detection models corresponding to K dimensions, each candidate image is input into the quality detection model corresponding to the K dimensions for quality detection, and K quality scores corresponding to the candidate image can be obtained, so that at least two high-quality images to be detected can be screened out from the N candidate images based on the K quality scores corresponding to each candidate image.
[0077] Exemplarily, the above-mentioned dimensions include at least one of the following: detail loss, visual information fidelity, character motion artifact loss, texture consistency, and illumination consistency, etc., but are not limited thereto. In other words, the quality detection models corresponding to the K dimensions include a detail loss detection model, a visual information fidelity detection model, a character motion artifact loss detection model, a texture consistency detection model, and an illumination consistency detection model, etc.
[0078] Among them, detail loss is used to measure the degree of loss of high-frequency information (such as edges, textures, and tiny structures) in the candidate image, especially the tiny features in the image, such as texture, facial details (such as the subtle parts of the eyes and mouth), etc. Details in fake videos or images are often lost due to image compression, face-changing technology, or image synthesis, causing the image to appear blurry or unnatural. If the texture of a face in a video frame is partially blurred, the detail loss detection model will give a lower quality score.
[0079] Visual information fidelity is used to measure the degree of authenticity of candidate images in human visual perception. Visual fidelity emphasizes the subjective perception of the human visual system on images rather than simple pixel-level differences. It not only focuses on details, but also covers global features such as color, lighting, and contrast. The visual information fidelity detection model evaluates the similarity between an image and a real image, focusing on features such as image color, contrast, and sharpness. Images with high fidelity should have realistic visual effects and appear to blend naturally with the scene. If the lighting in a frame is not coordinated with the background, or the color of the character does not match the background, the visual information fidelity detection model will give a lower quality score.
[0080] The artifact loss of human motion is used to measure the artifacts (such as smear, ghosting, edge breakage) generated in the candidate image due to inaccurate motion estimation or compensation. These artifacts are usually related to human motion or camera motion. Smear is the blur caused by the misalignment of moving objects between frames. Ghosting is the repeated appearance of moving objects between frames. Edge breakage is the discontinuity of the edges of moving objects. If in the candidate image, the facial expression of a certain character does not match the head movement, or the person suddenly deforms, this unnatural movement will be recognized by the artifact loss detection model of human motion and a lower quality score will be given.
[0081] Texture consistency is used to measure whether the facial and background textures in the candidate image are natural and coherent. If the texture of the human face and the background in the video do not blend well, or the texture contrast between the face and other image regions is significantly different, the texture consistency model will give a lower quality score.
[0082] Illumination consistency is used to measure whether the illumination and shadows in the candidate image are reasonable and consistent. If in the video, the illumination direction of the person is inconsistent with that of the background, or the facial shadows do not seem to match the actual illumination, the illumination consistency model will give a lower quality score.
[0083] Through multi-dimensional quality detection, the electronic device can more accurately evaluate the quality of each frame of candidate image, ensuring high reliability when the selected image is used for subsequent forgery detection.
[0084] Exemplarily, after confirming at least two images to be detected in the above manner, face decoupling can be performed on at least two images to be detected, that is, face extraction is performed on at least two frames of images to be detected respectively to obtain multiple face images. In a possible implementation, for each frame of the image to be detected, a face detection algorithm (such as YOLO, MTCNN, Haar Cascade, etc. based on deep learning) can be applied to detect and locate all the faces in the image to be detected. Then, for the detected face regions in each frame, face images are extracted by cropping. Finally, each frame of the image to be detected will be extracted into independent face images. Even if there are multiple faces in one frame, these faces will be processed as independent images.
[0085] In S402, clustering is performed on multiple face images to determine the clustering labels to which each face image belongs. Among them, all the face images belonging to the same clustering label are divided into the same group.
[0086] In this step, for the face images extracted in step S401, the electronic device groups the faces using a clustering algorithm. The purpose of clustering is to classify similar face images into the same group, which is usually based on the features of the face (such as facial contour, expression, eye position, etc.). Each group of face images is assigned a clustering label, and all face images belonging to the same clustering label are regarded as facial images of the same person. Through clustering, it can help the forgery detection model clearly distinguish different people, avoid confusing different people in the video, and thus improve the accuracy of forgery detection.
[0087] In a possible implementation, the electronic device can use a trained feature extraction model (such as a face recognition model based on a convolutional neural network) to extract the face features of each face image. The feature extraction model can generate the face features of each face image by analyzing the key facial features (such as the relative positions of eyes, nose, mouth, facial contour, etc.) in the face image; then, based on the similarity between the face features of any two face images, cluster multiple face images to determine the clustering label to which each face image belongs.
[0088] In S403, obtain the context information of at least one group of face images. The context information includes at least one of the following: clustering label, facial features of the face image, relative position information of the face image in the image to be detected to which it belongs, and temporal features of the image to be detected to which the face image belongs.
[0089] Exemplarily, one or more groups of face images can be selected for forgery detection according to actual needs.
[0090] In the first possible implementation, the group of face images with the largest number of face images within the group can be selected for forgery detection, which helps to select the most typical images for detection and increase the coverage of forgery detection. When there are multiple different people in the video and the frequency of appearance of people in each group is different, selecting the group with the largest number can ensure more comprehensive forgery detection.
[0091] In a second possible implementation, selection can be made based on the user's selection operation. For example, during video playback or display, the user manually selects one or more faces in the video frames, which can be achieved through the user interface. The user may mark the person to be detected by clicking with the mouse, touch screen operation, or other interaction methods. Through previous clustering analysis, the electronic device identifies the cluster group where the face image selected by the user is located. Each cluster group contains a set of face images that are similar in image features and usually represent the same person. Once the cluster group where the selected face is located is determined, the electronic device can perform forgery detection based on all the face images in this group. These face images may come from different video frames, but since they belong to the same person, they should be considered together during forgery detection. All the face images in this group will be passed as input to the forgery detection model.
[0092] In a third possible implementation, among all the cluster groups, select the group with the greatest difference from other groups. Assume that the face images in this group are more affected during forgery processing. By analyzing the differences between cluster groups through clustering, potential forgery features can be identified. This strategy is applicable to the situation where forgery features in forged videos may cause obvious differences between images and other images. Selecting this group with a larger difference for detection can help discover the features of forged videos.
[0093] In a fourth possible implementation, forgery techniques may leave forgery traces when facial expressions change, especially when a person makes large emotional expressions (such as smiling, frowning, etc.). Therefore, the cluster groups containing images with various facial expressions can be selected for detection to capture the forgery traces that may be exposed by the expression changes. If the forged video contains frequent facial expression changes (such as happiness, anger, sorrow, joy, etc.), preferentially selecting these images can help detect the forgery traces caused by the expression changes.
[0094] In a fifth possible implementation, forgery detection can also be performed based on the face images of all cluster groups.
[0095] After determining at least one group of face images to be detected, the context information of the at least one group of face images to be detected can be obtained. The context information includes at least one of the following: cluster label, facial features of each face image, relative position information of each face image in the image to be detected to which it belongs, and temporal features of the image to be detected to which each face image belongs.
[0096] (1) The cluster label can help the forgery detection model identify whether different frames in the video belong to the same person from a global perspective, which is crucial for forgery detection because forgery behavior may cause inconsistent forgery traces between multiple frames.
[0097] (2) Facial features can help the forgery detection model focus on facial details, especially to identify the naturalness of facial movements (such as lowering the head, raising the head, shaking the head), and can effectively identify abnormalities in forged faces.
[0098] Exemplarily, the facial features of a face image include at least one of the following: face pose feature, face proportion, and mouth opening / closing degree. The electronic device can use face key point detection technology to extract the face key points of each face image. The face key points include but are not limited to eyes, nose, mouth, chin, etc.
[0099] Regarding the face pose feature of a face image, the electronic device can determine it based on the difference between the face key points in the face image and the face key points in the reference pose. For example, by calculating the distance between the face key points in the face image and the face key points in the reference pose, the face orientation in the face image (such as front, side, lowering the head, raising the head, etc.) can be determined.
[0100] Regarding the mouth opening / closing degree in a face image, the electronic device can determine the distance between the upper and lower lips based on the upper lip key point and the lower lip key point of the face in the face image, and then determine the mouth opening / closing degree based on the ratio between the distance between the upper and lower lips of the face in the face image and the specified side (such as the facial height) of the face in the face image.
[0101] The face proportion is used to describe the relative size of the face in the image to be detected. If the size of the face image is the same as the display size of the face image in the image to be detected to which the face image belongs, then the face proportion is determined based on the ratio between the size of the face image and the size of the image to be detected to which the face image belongs.
[0102] (3) The relative position information enables the forgery detection model to determine which person's image is more important, especially in a multi-person scenario. The main character usually appears in the center of the video and appears frequently in multiple frames. This position feature helps the model better capture the features related to forgery.
[0103] (4) The temporal features enhance the forgery detection model's understanding of the dynamic information of the video. Especially when facing forgery situations such as large-angle head lowering and head raising, it can identify unnatural facial changes in forged videos.
[0104] Rich context information provides additional "clues" for forgery detection, especially in complex video scenarios (such as multi-person interaction, complex background, etc.), which helps improve the adaptability of the model in different forgery scenarios.
[0105] In S404, the context information of at least one set of face images is used as a prior constraint for the trained forgery detection model to control the forgery detection model to perform forgery detection on at least one set of face images, and a forgery detection result is obtained.
[0106] In this step, in addition to inputting at least one set of face images into the trained forgery detection model, the context information of this set of images obtained from step S403 (such as at least one of clustering labels, facial features, relative position information, and temporal features) is also input into the trained forgery detection model as prior information. In this way, the forgery detection model not only relies on traditional image features but also can refer to context information to perform forgery detection. The forgery detection model will use this context information to guide the forgery detection process and finally output the result of forgery detection (such as determining whether a certain face image is forged).
[0107] In this embodiment, the context information provides rich prior knowledge. The forgery detection model can combine the features of the image itself with the context information to improve the accuracy of forgery detection. The context information helps the forgery detection model better understand the spatio-temporal relationship in the video, especially in complex forgery behaviors, and can improve the robustness of forgery detection.
[0108] In one implementation, the electronic device can input at least one set of face images and their context information into the forgery detection model, so that the forgery detection model extracts forgery detection features from at least one set of face images, and performs forgery detection based on the fusion result of the context information of at least one set of face images and the forgery detection features to obtain a forgery detection result. In this embodiment, by inputting at least one set of face images and their context information into the forgery detection model, the context information provides a richer context, such as facial position, temporal features, etc., which can help the model understand the relationship between the face in the image and other image contents. Fusing the context information with the extracted face features helps the model better identify the differences between forged faces and real faces, especially in complex forgery scenarios, and improves the accuracy of forgery detection.
[0109] Exemplarily, the forgery detection model includes a feature extraction network and a forgery detection network based on the Transformer architecture. The forgery detection features are obtained by the feature extraction network extracting features from at least one set of face images; the forgery detection result is obtained by the forgery detection network based on the Transformer architecture performing forgery detection based on the fusion result.
[0110] A forgery detection network based on the Transformer architecture, the core of which lies in making full use of the self-attention mechanism of the Transformer. This mechanism can effectively capture long-range dependencies in the data and has strong adaptability to the complex data patterns and feature associations involved in forgery detection. Through the self-attention mechanism, it is possible to focus more on key information and reduce the interference of irrelevant information, thereby improving the accuracy of detection.
[0111] Exemplarily, the forgery detection result includes at least one of the following: a first result for describing whether there is a forged portrait in the video to be detected, and a second result for describing whether the people indicated by each group of face images are forged. Among them, the first result is considered from the entire video and is an overall judgment (for example, "forged" or "genuine"), indicating whether the entire video contains forged human faces. By providing an overall judgment on whether there are forged faces in the entire video, the first result helps users quickly understand the overall quality and credibility of the video. The second result provides the forgery detection result for each clustering group (for example, marked as "forged" or "genuine") and can provide more detailed forgery detection information; for example, if there are multiple people in the video, it can accurately determine which person is forged and which is genuine. When the user is concerned about a specific person, the second result can provide a detailed detection report for each group, meeting the user's needs in personalized detection.
[0112] The various technical features in the above embodiments can be combined arbitrarily as long as there is no conflict or contradiction between the features. However, due to space limitations, they are not described one by one. Therefore, any combination of the various technical features in the above embodiments also falls within the scope disclosed in this specification.
[0113] In some embodiments, the embodiments of this specification also provide an electronic device, including: a processor; a memory for storing instructions executable by the processor; wherein, the processor realizes the method described in any one of the above by running the executable instructions.
[0114] Figure 5 It is a schematic structural diagram of a device provided by an exemplary embodiment. Please refer to Figure 5, at the hardware level, the device includes a processor 502, an internal bus 504, a network interface 506, a memory 508, and a non-volatile memory 510. Of course, it may also include other hardware required for other functions. One or more embodiments of this specification can be implemented in software. For example, the processor 502 reads the corresponding computer program from the non-volatile memory 510 into the memory 508 and then runs it. Of course, in addition to the software implementation, one or more embodiments of this specification do not exclude other implementation methods, such as logic devices or a combination of software and hardware, etc. That is to say, the execution subject of the following processing flow is not limited to each logic unit, and can also be hardware or a logic device.
[0115] In some embodiments, the forgery detection device can be applied to a device as Figure 5 shown to implement the technical solutions of this specification. Among them, the forgery detection device may include:
[0116] A face image acquisition module, configured to determine at least two frames of images to be detected from the video to be detected, and perform face extraction on each of the at least two frames of images to be detected to obtain multiple face images.
[0117] A clustering module, configured to cluster the multiple face images, determine the clustering labels to which each face image belongs, where all face images belonging to the same clustering label are divided into the same group.
[0118] A context acquisition module, configured to acquire context information of at least one group of face images, where the context information includes at least one of the following: the clustering label, the facial features of the face image, the relative position information of the face image in the image to be detected to which it belongs, and the temporal feature of the image to be detected to which the face image belongs.
[0119] A forgery detection module, configured to use the context information of the at least one group of face images as a prior constraint of a trained forgery detection model, and control the forgery detection model to perform forgery detection on the at least one group of face images to obtain a forgery detection result.
[0120] Exemplarily, the forgery detection module is specifically configured to input the at least one group of face images and their context information into the forgery detection model, so that the forgery detection model extracts forgery detection features from the at least one group of face images, and perform forgery detection based on the fusion result of the context information of the at least one group of face images and the forgery detection features to obtain a forgery detection result.
[0121] Exemplarily, the forgery detection result includes: a first result for describing whether there is a forged portrait in the video to be detected, and / or, a second result for describing whether the person indicated by each group of face images is forged.
[0122] Exemplarily, the forgery detection model includes a feature extraction network and a forgery detection network based on the Transformer architecture. The forgery detection features are obtained by the feature extraction network extracting features from the at least one set of face images. The forgery detection result is obtained by the forgery detection network based on the Transformer architecture performing forgery detection based on the fusion result.
[0123] Exemplarily, the facial features of the face image include at least one of the following: face pose feature, face proportion, and mouth opening and closing degree. Among them, the face pose feature of the face image is determined based on the difference between the face key points in the face image and the face key points of the reference pose; the size of the face image is the same as the display size of the face image in the to-be-detected image to which the face image belongs, and the face proportion is determined based on the ratio between the size of the face image and the size of the to-be-detected image to which the face image belongs; the mouth opening and closing degree of the face image is determined based on the ratio between the distance between the upper and lower lips of the face in the face image and the specified side of the face in the face image.
[0124] Exemplarily, the clustering module is specifically configured to use the trained feature extraction model to extract the face features of each face image; based on the similarity between the face features of any two face images, cluster the multiple face images to determine the clustering label to which each face image belongs.
[0125] Exemplarily, the face image acquisition module is specifically configured to sample multiple candidate images from the to-be-detected video according to a preset sampling rate; perform quality detection on each candidate image in at least one dimension to obtain the quality score of each candidate image in the at least one dimension; determine the candidate images whose quality scores in the at least one dimension meet the preset quality conditions as the to-be-detected images.
[0126] Exemplarily, a trained quality detection model corresponds to each dimension, and the quality score of each candidate image in each dimension is obtained by inputting the candidate image into the quality detection model corresponding to that dimension for quality detection.
[0127] Exemplarily, the at least one dimension is determined from the following dimensions: detail loss, visual information fidelity, and artifact loss of human motion; among them, the detail loss is used to measure the degree of loss of high-frequency information in the candidate image; the visual information fidelity is used to measure the real degree of the candidate image in human eye visual perception; the artifact loss of human motion is used to measure the artifacts generated in the candidate image due to inaccurate motion estimation.
[0128] In some embodiments, the training device of the forgery detection model can be applied to a device as shown in Figure 5 to implement the technical solutions of this specification. Among them, the training device of the forgery detection model may include:
[0129] A sample acquisition module, configured to acquire a plurality of video samples labeled with supervision labels.
[0130] A sample preprocessing module, for each video sample, determining at least two frames of images to be detected from the video sample, performing face extraction on each of the at least two frames of images to obtain a plurality of face images, and clustering the plurality of face images to determine the clustering labels to which each of the face images belongs, dividing all the face images belonging to the same clustering label into the same group, and further obtaining the context information of at least one group of face images, where the context information includes at least one of the following: the clustering label, the facial features of the face image, the relative position information of the face image in the image to be detected to which it belongs, and the temporal feature of the image to be detected to which the face image belongs.
[0131] A model training module, configured to input at least one group of face images corresponding to the video sample and their context information into a forgery detection model to be trained, so that the forgery detection model refers to the context information of the at least one group of face images, performs forgery detection on the at least one group of face images to obtain a prediction result, and takes minimizing the error between the prediction result and the supervision label as an optimization objective to train the forgery detection model.
[0132] Exemplarily, the forgery detection model includes a feature extraction network and a forgery detection network based on the Transformer architecture; the feature extraction network is configured to perform feature extraction on the at least one group of face images to obtain forgery detection features; the forgery detection network based on the Transformer architecture is configured to perform forgery detection based on the fusion result of the context information of the at least one group of face images and the forgery detection features to obtain the prediction result.
[0133] The implementation processes of the functions and roles of each module in the above device are specifically described in the implementation processes of the corresponding steps in the above method, and will not be elaborated here.
[0134] Based on the same concept as the above method, this specification also provides an electronic device, including: a processor; a memory for storing executable instructions that can be executed by the processor; wherein, the processor realizes the steps of the method as described in any of the above embodiments by running the executable instructions.
[0135] Based on the same concept as the above method, this specification also provides a computer-readable storage medium, on which computer instructions are stored, and when the instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0136] Computer-readable media include both permanent and non-permanent, removable and non-removable media, and information storage can be implemented by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassette tapes, disk storage, quantum memory, graphene-based storage media or other magnetic storage devices, or any other non-transmission media that can be used to store information accessible by a computing device. As defined herein, computer-readable media do not include transitory media such as modulated data signals and carrier waves.
[0137] Based on the same concept as the above method, this specification also provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the method described in any of the above embodiments are implemented.
[0138] The above are only the preferred embodiments of one or more embodiments of this specification, and are not intended to limit one or more embodiments of this specification. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of one or more embodiments of this specification shall be included within the scope of protection of one or more embodiments of this specification.
Claims
1. A counterfeit detection method, comprising: Determine at least two frames of images to be detected from the video to be detected, and perform face extraction on the at least two frames of images to be detected respectively to obtain multiple face images; Clustering the plurality of face images to determine a cluster label to which each of the face images belongs, wherein all face images belonging to the same cluster label are classified into the same group; Acquire context information of at least one group of face images, the context information comprising at least one of the following: the clustering label, the facial features of the face images, the relative position information of the face images in the to-be-detected images to which they belong, and the temporal features of the to-be-detected images to which the face images belong; The context information of the at least one group of face images is used as a priori constraints of a trained forgery detection model, and the forgery detection model is controlled to perform forgery detection on the at least one group of face images to obtain a forgery detection result.
2. The method according to claim 1, wherein the context information of the at least one group of face images is used as a priori constraints of a trained forgery detection model, and the forgery detection model is controlled to perform forgery detection on the at least one group of face images to obtain a forgery detection result for the video to be detected, comprising: The at least one group of facial images and their context information are input into the forgery detection model, so that the forgery detection model performs feature extraction on the at least one group of facial images to obtain forgery detection features, and performs forgery detection based on the fusion result of the context information of the at least one group of facial images and the forgery detection features to obtain a forgery detection result.
3. The method according to claim 1 or 2, wherein the forgery detection result comprises: A first result for describing whether there is a forged portrait in the video to be detected, and / or a second result for describing whether the person indicated by each group of facial images is forged.
4. The method according to claim 2, wherein the forgery detection model comprises a feature extraction network and a forgery detection network based on a Transformer architecture; The forgery detection feature is obtained by extracting features from the at least one group of face images by the feature extraction network; The forgery detection result is obtained by performing forgery detection based on the fusion result by the forgery detection network based on the Transformer architecture.
5. The method according to claim 1, wherein the facial features of the facial image include at least one of the following: facial posture features, facial proportion, and mouth opening and closing degree; in, The facial posture features of the facial image are determined based on the difference between the facial key points in the facial image and the facial key points of the reference posture; The size of the face image is the same as the display size of the face image in the image to be detected to which it belongs, and the face proportion is determined based on the ratio between the size of the face image and the size of the image to be detected to which the face image belongs; The degree of mouth opening and closing of the facial image is determined based on the ratio between the distance between the upper and lower lips of the face in the facial image and the designated side of the face in the facial image.
6. The method according to claim 1, wherein clustering the plurality of face images to determine a cluster label to which each of the face images belongs comprises: Extracting facial features of each of the facial images using the trained feature extraction model; Based on the similarity between the facial features of any two of the facial images, the multiple facial images are clustered to determine the cluster label to which each of the facial images belongs.
7. The method according to claim 1, wherein determining at least two frames of images to be detected from the video to be detected comprises: According to a preset sampling rate, sampling is performed from the video to be detected to obtain multiple frames of candidate images; Performing quality detection on at least one dimension of each of the candidate images to obtain a quality score of each of the candidate images in the at least one dimension; A candidate image whose quality score in the at least one dimension meets a preset quality condition is determined as the image to be detected.
8. According to the method of claim 7, each dimension corresponds to a trained quality detection model, and the quality score of each candidate image in each dimension is obtained by inputting the candidate image into the quality detection model corresponding to the dimension for quality detection; And / or, the at least one dimension is determined from the following dimensions: detail loss, visual information fidelity and character motion artifact loss; wherein, The detail loss is used to measure the degree of loss of high-frequency information in the candidate image; The visual information fidelity is used to measure the authenticity of the candidate image in human visual perception; The character motion artifact loss is used to measure artifacts generated in the candidate image due to inaccurate motion estimation.
9. A method for training a forgery detection model, comprising: Obtain multiple video samples annotated with supervision labels; For each video sample, at least two frames of images to be detected are determined from the video sample, faces are extracted from the at least two frames of images to be detected respectively to obtain multiple face images, and the multiple face images are clustered to determine the clustering label to which each face image belongs, and all face images belonging to the same clustering label are divided into the same group, thereby obtaining context information of at least one group of face images, wherein the context information includes at least one of the following: the clustering label, the facial features of the face image, the relative position information of the face image in the image to be detected to which it belongs, and the temporal features of the image to be detected to which the face image belongs; Inputting at least one group of face images corresponding to the video sample and context information thereof into a forgery detection model to be trained, so that the forgery detection model refers to the context information of the at least one group of face images, performs forgery detection on the at least one group of face images to obtain a prediction result, and trains the forgery detection model with minimizing the error between the prediction result and the supervision label as an optimization goal; The trained forgery detection model is applied to the forgery detection method according to any one of claims 1 to 8.
10. The method according to claim 9, wherein the forgery detection model comprises a feature extraction network and a forgery detection network based on a Transformer architecture; The feature extraction network is used to extract features from the at least one group of face images to obtain forgery detection features; The forgery detection network based on the Transformer architecture is used to perform forgery detection based on the context information of the at least one group of face images and the fusion result of the forgery detection feature to obtain the prediction result.
11. An electronic device, comprising: processor; A memory for storing processor-executable instructions; wherein the processor implements the steps of the method according to any one of claims 1 to 10 by executing the executable instructions.
12. A computer-readable storage medium having computer instructions stored thereon, which, when executed by a processor, implement the steps of the method according to any one of claims 1 to 10.
13. A computer program product, comprising a computer program / instruction, which, when executed by a processor, implements the steps of the method according to any one of claims 1 to 10.