Deep fake detection method and system based on graph convolution and multi-scale hint fusion
By using graph convolution and multi-scale prompt fusion technology in depth forgery detection, the shortcomings of existing methods in global-local information modeling and information fusion are solved, and higher detection accuracy and generalization capabilities are achieved.
Patent Information
- Application Number
- CN202411427606.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-14
- Publication Date
- 2025-05-09
- Estimated Expiration
- 2044-10-14
Smart Images

Figure CN118941936B_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of deep fake detection, and in particular to a deep fake detection method and system based on graph convolution and multi-scale prompt fusion. Background Art
[0002] The term deepfake refers to the creation of highly realistic fake images and videos through deep learning algorithms, especially generative adversarial networks (GANs), which are difficult for ordinary users to identify. Deepfake technology is widely used in industries such as entertainment and film production, but more often it brings serious social problems such as identity theft and privacy infringement; therefore, it is crucial to develop effective deepfake detection technology that aims to distinguish between real images or videos and fake content to protect personal privacy and public safety.
[0003] At present, although detection methods based on deep convolutional neural networks (CNNs) can identify certain specific types of forged content, they are limited by the local receptive field and find it difficult to capture global dependencies, causing the model to focus only on certain local areas, which limits the model's generalization ability. In addition, in the RGB color space, some subtle traces of forgery are difficult to detect, causing existing models to easily ignore these key areas.
[0004] In response to the first point, some studies have turned to using Vision Transformer (ViT) to enhance the global perception of the model; however, these ViT-based methods divide images into patches of fixed size, which may cause continuous facial features such as eyes to be scattered into different patches, thus affecting the coherence of global information; in addition, each patch is small in size, contains less information, and lacks clear semantic associations between patches; these problems lead to suboptimal global information modeling and unnecessary similarity calculations. At the same time, due to the insufficient modeling of local features by ViT, using only ViT may ignore some key local areas. Therefore, existing methods lack flexibility and robustness in global-local information modeling.
[0005] Regarding the second point, in order to make up for the defects in the RGB space, some studies introduced frequency domain information as an auxiliary modality. In the frequency space, some forgery traces that are difficult to detect in the RGB space will become very obvious; however, these methods usually only fuse multimodal data through simple splicing or addition operations, failing to fully utilize their respective advantages, resulting in the loss of important details; in addition, these methods only use the last feature map of the model for classification, however, different forgery traces have different sizes, and the forgery traces in the shallow feature map of the model may be gradually lost as the model propagates, causing the model to ignore key information from the shallow layer when making the final prediction; therefore, the fusion of RGB and frequency information in the prior art is not fine-grained enough.
[0006] Therefore, despite the introduction of global and frequency information, existing deep fake detection methods still have problems of low accuracy and weak generalization ability. Summary of the invention
[0007] In order to overcome the shortcomings of the above-mentioned prior art, the present invention provides a deep fake detection method and system based on graph convolution and multi-scale prompt fusion, which overcomes the limitations brought by fixed shape patch division, enhances the adaptability of the model to different forgery patterns, and realizes more fine-grained fusion of RGB and frequency information, thereby improving the accuracy and generalization ability of deep fake detection.
[0008] To achieve the above objectives, one or more embodiments of the present invention provide the following technical solutions:
[0009] A first aspect of the present invention provides a deep fake detection method based on graph convolution and multi-scale hint fusion.
[0010] Deep fake detection method based on graph convolution and multi-scale hint fusion, including:
[0011] Obtain the face image to be detected;
[0012] Input the face image into the trained detection model, perform binary classification on whether it is fake or not, and obtain the deep fake detection result;
[0013] Among them, the detection model extracts graph features and frequency features in face images, uses an adaptive graph convolution module to perform graph convolution on the graph features, aggregates neighbor node information to form group tokens, and uses a multi-scale prompt fusion module to generate prompt tokens from graph features and frequency features; finally, the spliced tokens are input into ViT for classification.
[0014] Furthermore, the detection model includes a dual-branch feature extraction network, an adaptive graph convolution module, a multi-scale cue fusion module and a classification module;
[0015] The dual-branch feature extraction network, based on CNN, extracts graph features and frequency features in face images;
[0016] The adaptive graph convolution module forms group tokens based on the adaptively constructed graph structure and graph convolution;
[0017] The multi-scale prompt fusion module generates prompt tokens based on the prompt learner;
[0018] The classification module performs classification based on ViT.
[0019] Furthermore, the dual-branch feature extraction network includes two branches, which are used to extract graph features and frequency features respectively;
[0020] The image features are extracted from the face image using CNN;
[0021] The frequency features are extracted from the frequency domain representation of the face image using CNN.
[0022] Furthermore, the adaptive graph convolution module includes adaptive graph construction and graph aggregation;
[0023] The adaptive graph construction adaptively constructs the graph structure through relative position encoding and feature similarity of graph features;
[0024] The graph aggregation aggregates related node information through graph convolution to obtain group tokens containing rich semantic information.
[0025] Furthermore, the graph structure is composed of a node matrix and an adjacency matrix, and the specific construction steps are:
[0026] Treat each pixel in the graph feature as a node and construct a node matrix;
[0027] Using the local aggregator, the query, key and value matrices of the node matrix are calculated, and then the similarity scores between the nodes are calculated. Based on the similarity scores, the adjacency matrix is constructed.
[0028] Furthermore, the multi-scale cue fusion module sets a cue learner for each feature to locate key information of the input feature and convert it into a cue token;
[0029] The prompt learner uses spatial attention and channel attention to extract key information and obtains the corresponding prompt token through linear mapping.
[0030] Furthermore, the classification module inputs the concatenated tokens into ViT for global feature extraction and feature fusion, and performs classification based on the fused features.
[0031] A second aspect of the present invention provides a deep fake detection system based on graph convolution and multi-scale hint fusion.
[0032] Deep fake detection system based on graph convolution and multi-scale hint fusion, including:
[0033] The image acquisition module is configured to: acquire a face image to be detected;
[0034] The forgery detection module is configured to: input the face image into the trained detection model, perform binary classification on whether it is forged, and obtain a deep forgery detection result;
[0035] Among them, the detection model extracts graph features and frequency features in face images, uses an adaptive graph convolution module to perform graph convolution on the graph features, aggregates neighbor node information to form group tokens, and uses a multi-scale prompt fusion module to generate prompt tokens from graph features and frequency features; finally, the spliced tokens are input into ViT for classification.
[0036] The third aspect of the present invention provides a computer-readable storage medium having a program stored thereon, which, when executed by a processor, implements the steps in the deep fake detection method based on graph convolution and multi-scale prompt fusion as described in the first aspect of the present invention.
[0037] The fourth aspect of the present invention provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor, wherein when the processor executes the program, the steps in the deep fake detection method based on graph convolution and multi-scale prompt fusion as described in the first aspect of the present invention are implemented.
[0038] One or more of the above technical solutions have the following beneficial effects:
[0039] (1) Global-local information modeling
[0040] Since ViT is weak in processing local features, the present invention designs a hybrid architecture that combines the advantages of CNN and ViT.
[0041] First, CNN is used to extract local features of the image, and then each pixel in the feature map generated by CNN is regarded as a token and input into ViT to extract global information. In this way, an effective combination of local and global features is achieved. In order to more fully extract global and frequency features, the present invention proposes two modules: adaptive graph convolutional network (AGCN) and multi-scale prompt fusion (MSPF).
[0042] (2) Global Information Modeling
[0043] For global information extraction, this paper proposes an adaptive graph convolutional network (AGCN) module.
[0044] Through relative position relationships and feature similarities, AGCN can aggregate related node information through graph convolution based on the adaptively constructed graph structure to form group tokens with richer semantic information, effectively constructing a more robust global information representation.
[0045] Through AGCN, each token can adaptively aggregate information from its neighboring nodes to contain richer semantic information; after all tokens have aggregated information, the associations between them are clearer, enhancing the global modeling capability of the model; this method not only overcomes the limitations brought by fixed-shape patch division, but also enhances the model's adaptability to different forgery patterns.
[0046] (3) Frequency feature fusion
[0047] For frequency information extraction, the present invention proposes a multi-scale prompt fusion (MSPF) module.
[0048] MSPF extracts frequency information of different scales by prompting the learner. This information is input into ViT in the form of prompt tokens as supplements and group tokens. It interacts with the group tokens through the self-attention mechanism to achieve the fusion of RGB-frequency information of different scales. At the same time, the present invention adds additional classification tokens to interact with the above two tokens, comprehensively considers multi-modal and multi-scale information for final classification; through MSPF, RGB and frequency modal information are fused through the self-attention mechanism, and relevant information is flexibly selected for fusion; in addition, the classification token can interact with tokens of different scales and modalities, and make final classification based on comprehensive information from multiple aspects, making the multi-modal information fusion process more fine-grained and flexible.
[0049] Advantages of additional aspects of the present invention will be given in part in the following description, and in part will become obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0050] The accompanying drawings in the specification, which constitute a part of the present invention, are used to provide a further understanding of the present invention. The exemplary embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute improper limitations on the present invention.
[0051] Figure 1 This is a structural diagram of the detection model of the first embodiment. DETAILED DESCRIPTION
[0052] It should be noted that the following detailed descriptions are illustrative and are intended to provide further explanation of the present application. Unless otherwise specified, all technical and scientific terms used in the present invention have the same meanings as those commonly understood by those skilled in the art to which the present application belongs.
[0053] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present application. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "include" are used in this specification, it indicates the presence of features, steps, operations, devices, components and / or combinations thereof.
[0054] Terminology explanation:
[0055] 1. Deepfake: Deepfake technology is an artificial intelligence technology that uses a machine learning model called "generative adversarial network" (GAN) to merge and superimpose images or videos onto the source images or videos, and uses neural network technology to perform large-sample learning to splice a person's voice, facial expressions and body movements into false content. The most common form of deepfake is AI face-changing technology, in addition to voice simulation, face synthesis, video generation, etc. Its emergence makes it possible to tamper with or generate highly realistic and difficult-to-distinguish audio and video content, and observers ultimately cannot distinguish the authenticity with the naked eye.
[0056] 2. Frequency information of the image: An indicator of the intensity of the grayscale value changes in the image, which is the gradient of the grayscale in the plane space.
[0057] Embodiment 1
[0058] In one embodiment of the present disclosure, a deep fake detection method based on graph convolution and multi-scale hint fusion is provided, comprising the following steps:
[0059] Step S1: Obtain a face image to be detected;
[0060] Step S2: Input the face image into the trained detection model, perform binary classification on whether it is forged or not, and obtain the deep fake detection result;
[0061] Among them, the detection model extracts graph features and frequency features in face images, uses an adaptive graph convolution module to perform graph convolution on the graph features, aggregates neighbor node information to form group tokens, and uses a multi-scale prompt fusion module to generate prompt tokens from graph features and frequency features; finally, the spliced tokens are input into ViT for classification.
[0062] Furthermore, the detection model includes a dual-branch feature extraction network, an adaptive graph convolution module, a multi-scale cue fusion module and a classification module;
[0063] The dual-branch feature extraction network, based on CNN, extracts graph features and frequency features in face images;
[0064] The adaptive graph convolution module forms group tokens based on the adaptively constructed graph structure and graph convolution;
[0065] The multi-scale prompt fusion module generates prompt tokens based on the prompt learner;
[0066] The classification module performs classification based on ViT.
[0067] Furthermore, the dual-branch feature extraction network includes two branches, which are used to extract graph features and frequency features respectively;
[0068] The image features are extracted from the face image using CNN;
[0069] The frequency features are extracted from the frequency domain representation of the face image using CNN.
[0070] Furthermore, the adaptive graph convolution module includes adaptive graph construction and graph aggregation;
[0071] The adaptive graph construction adaptively constructs the graph structure through relative position encoding and feature similarity of graph features;
[0072] The graph aggregation aggregates related node information through graph convolution to obtain group tokens containing rich semantic information.
[0073] Furthermore, the graph structure is composed of a node matrix and an adjacency matrix, and the specific construction steps are:
[0074] Treat each pixel in the graph feature as a node and construct a node matrix;
[0075] Using the local aggregator, the query, key and value matrices of the node matrix are calculated, and then the similarity scores between the nodes are calculated. Based on the similarity scores, the adjacency matrix is constructed.
[0076] Furthermore, the multi-scale cue fusion module sets a cue learner for each scale feature map output by the CNN, so as to locate the key information of the input features and convert them into cue tokens;
[0077] The prompt learner uses spatial attention and channel attention to extract key information and obtains the corresponding prompt token through linear mapping.
[0078] Furthermore, the classification module inputs the concatenated tokens into ViT for global feature extraction and feature fusion, and performs classification based on the fused features.
[0079] The following is a detailed description of the implementation process of the deep fake detection method based on graph convolution and multi-scale prompt fusion in this embodiment.
[0080] This embodiment provides a more specific detection model DINA. The input of DINA is an RGB face image. , and then through the calculation process of the detection model DINA, the prediction result of the image is finally output ,in represents the model predicts that the image is real, Represents an image that is fake; Figure 1 As shown in the figure, DINA consists of four parts: a dual-branch feature extraction network, an adaptive graph convolution module (AGCN), a multi-scale prompt fusion (MSPF) module, and a classification module. The following is a detailed description of each model:
[0081] 1. Dual-branch feature extraction network
[0082] like Figure 1 As shown, for the input RGB image , first, use the SRM filter to obtain its frequency domain representation , dual-branch feature extraction consists of two CNN feature extractors and Composed of, they do not share parameters, respectively and As input, its output is (The specific number is determined by the CNN feature extractor) feature maps and ,in, represent The first feature map of the output, and the rest are similar. As The last feature map of contains rich local information, but lacks global information modeling. The input is fed into the adaptive graph convolution module to enhance the global information. In addition, other feature maps come from different scales or different modalities, which contain unique forgery information, so these feature maps are fed into the multi-scale hint fusion module to extract these key information.
[0083] The function of this part is to use CNN to extract features of different scales. In this embodiment, both branches use efficientnet-b4 as the feature extractor, which outputs 5 feature maps. For the acquisition of frequency domain features, the SRM filter is used to obtain .
[0084] Note that the main function of this part is to extract features for subsequent modules, so it is also feasible to use other CNN models or frequency feature conversion methods.
[0085] 2. Adaptive Graph Convolution Module (AGCN)
[0086] like Figure 1 As shown, the extracted feature map Input into AGCN to aggregate neighbor node information to form a group token .
[0087] The role of this module is to dynamically aggregate the features of its related neighbor nodes for each token, thereby forming an expression containing rich semantic information; this module consists of two parts: adaptive graph construction and graph aggregation. Each pixel in the graph is regarded as a node of the graph. Adaptive graph construction can adaptively form the adjacency matrix of the graph according to the relative position relationship and feature similarity between nodes; while graph aggregation aggregates information through the message propagation mechanism based on the adjacency matrix of the previous step.
[0088] 1. Adaptive graph construction
[0089] Given input features , are the number of channels, height and width of the input features respectively, and each pixel is regarded as a node of the graph. The size of the node matrix is converted to ,in After that, an adjacency matrix is constructed based on the relative position relationship and feature similarity between nodes. , specifically:
[0090] First, use a local aggregator to Mapping to query matrix , key matrix , value matrix :
[0091]
[0092] in, , as well as are the respective local aggregators, which consist of depthwise separable convolutional layers and linear projection layers.
[0093] Afterwards, based on , calculate the similarity matrix between nodes :
[0094]
[0095] in, represent The element in the i-th row and j-th column is the similarity between the two corresponding positions in the feature map. is the similarity matrix between all nodes. and Respectively represent and The elements of the i-th and j-th rows of Represents a matrix transpose operation; is a learnable relative position encoding that represents the spatial association between positions i and j.
[0096] The above formula calculates the similarity between different nodes from the perspective of feature similarity and relative position relationship. According to the obtained similarity matrix Constructing the set of neighbor nodes of the graph :
[0097]
[0098] in, are the neighbors of node i, for The i-th row of is the index corresponding to the largest element of topk, that is, for each node, select k nodes that are most similar to it as its neighbors.
[0099] Will Convert it into the form of an adjacency matrix and get the adjacency matrix , according to the node matrix and the adjacency matrix , we can construct a graph .
[0100] 2. Graph Aggregation
[0101] Based on the graph , a 3-layer graph convolutional network is used to aggregate relevant information for each node, and the t+1th layer is calculated by the following formula:
[0102]
[0103] in, , is the identity matrix, yes The degree matrix of is the weight matrix of this layer.
[0104] Through graph convolutional networks, each node can broadcast its own information. On the other hand, each node can also receive information from its neighboring nodes.
[0105] Through the above two parts, the final group token expression can be obtained .
[0106] 3. Multi-Scale Prompt Fusion (MSPF) Module
[0107] like Figure 1As shown, all remaining feature maps are input into MSPF to form prompt tokens and , corresponding to the two modes of RGB and frequency, so it is also called RGB prompt token and frequency hint tokens .
[0108] The function of this module is to learn the corresponding prompt tokens according to the input feature map. MSPF mainly consists of multiple prompt learner networks, which have the same structure but do not share parameters; each prompt learner network can locate the key information of the input feature map and convert it into a prompt token; the input of this module is Feature Map as well as , each feature map corresponds to a prompt learner.
[0109] Feature map For example, first, use spatial attention and channel attention to extract its key information:
[0110]
[0111]
[0112] in, is the sigmoid function, and For spatial and channel attention, LANet and SENet are used to locate key information and are not the main content of this module. It is also feasible to use other attention mechanisms.
[0113] Afterwards, the corresponding prompt token is obtained through linear mapping:
[0114]
[0115] in, represents the global average pooling operation, It consists of two linear layers.
[0116] Similarly, we can get the prompt tokens corresponding to other feature maps. The tokens corresponding to the two modalities RGB and frequency are concatenated to get two prompt tokens. and .
[0117] 4. Classification Module
[0118] like Figure 1 As shown, the group token , prompt token and And additional classification tokens (Learnable parameters) are spliced together and sent to a 2-layer ViT network for global feature extraction and feature fusion, that is, feature interaction, to obtain the classification token after interaction ; The classification tokens after interaction The prediction results of the model are obtained by sending them to the classification head, and the binary cross entropy loss (BCE) is used to guide model learning.
[0119] Specifically, the three types of tokens obtained above and a learnable classification token CLS are concatenated to obtain ,Will It is sent to a 2-layer ViT network for feature interaction and fusion; in ViT, through the self-attention mechanism, and and Interact with other systems to gather more information from different scales and frequencies. Interaction, fusion of RGB, frequency and multi-scale information, and obtain the classification token after interaction ; Finally, CLS is sent to the classification head to obtain the prediction results of the model.
[0120] In order to verify the effect of the method in this embodiment, a cross-dataset evaluation strategy was adopted. The detection model was trained on the FF++(c23) dataset, and then tested on the Celeb-DF V2 and DFDC datasets. The area under the receiver operating characteristic curve (AUC) was used as the evaluation indicator. This indicator is used for binary classification and represents the probability that the model ranks positive samples (fake faces) before negative samples (real faces). The comparison results are shown in Table 1.
[0121] Table 1 Comparison results of the method in this embodiment with other algorithms
[0122]
[0123] As can be seen from the table, compared with the baseline model efficientnet-b4, F3-Net and LRL introduced frequency information, HFI-Net and F2Trans further introduced global information, and their effects were significantly improved, indicating that global information and frequency information are of great help to improve model performance; DINA uses AGCN and MSPF modules to solve the problems of these methods being insufficiently flexible in global information modeling and insufficiently granular in the process of frequency feature fusion. The results show that DINA has greater advantages than similar models HFI-Net and F2Trans, proving the effectiveness of the two modules.
[0124] Embodiment 2
[0125] In one embodiment of the present disclosure, a deep fake detection system based on graph convolution and multi-scale prompt fusion is provided, including:
[0126] The image acquisition module is configured to: acquire a face image to be detected;
[0127] The forgery detection module is configured to: input the face image into the trained detection model, perform binary classification on whether it is forged, and obtain a deep forgery detection result;
[0128] Among them, the detection model extracts graph features and frequency features in face images, uses an adaptive graph convolution module to perform graph convolution on the graph features, aggregates neighbor node information to form group tokens, and uses a multi-scale prompt fusion module to generate prompt tokens from graph features and frequency features; finally, the spliced tokens are input into ViT for classification.
[0129] Embodiment 3
[0130] The purpose of this embodiment is to provide a computer-readable storage medium.
[0131] A computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the steps in the deep fake detection method based on graph convolution and multi-scale prompt fusion as described in the first embodiment of the present disclosure.
[0132] Embodiment 4
[0133] The purpose of this embodiment is to provide an electronic device.
[0134] An electronic device includes a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, the steps of the deep fake detection method based on graph convolution and multi-scale prompt fusion as described in the first embodiment of the present disclosure are implemented.
[0135] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. For those skilled in the art, the present invention may have various modifications and variations. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of the present invention shall be included in the protection scope of the present invention.
Claims
1. A deep fake detection method based on graph convolution and multi-scale hint fusion, characterized in that: include: Obtain the face image to be detected; Input the face image into the trained detection model, perform binary classification on whether it is fake or not, and obtain the deep fake detection result; The detection model extracts graph features and frequency features from the face image, performs graph convolution on the graph features using an adaptive graph convolution module, aggregates neighbor node information to form group tokens, and generates prompt tokens from the graph features and frequency features using a multi-scale prompt fusion module; finally, the tokens concatenated from the group tokens, prompt tokens, and learnable classification tokens are input into ViT for classification; The detection model includes a dual-branch feature extraction network, an adaptive graph convolution module, a multi-scale prompt fusion module and a classification module; The dual-branch feature extraction network, based on CNN, extracts graph features and frequency features in face images; The adaptive graph convolution module forms group tokens based on the adaptively constructed graph structure and graph convolution; The multi-scale prompt fusion module generates prompt tokens based on the prompt learner; The classification module performs classification based on ViT; The adaptive graph convolution module includes adaptive graph construction and graph aggregation; The adaptive graph construction adaptively constructs the graph structure through relative position encoding and feature similarity of graph features; The graph aggregation aggregates neighbor node information through graph convolution to obtain group tokens containing rich semantic information; The multi-scale prompt fusion module sets a prompt learner for each feature to locate the key information of the input graph features and frequency features and convert them into prompt tokens; The prompt learner uses spatial attention and channel attention to extract key information and obtains corresponding prompt tokens through linear mapping; The classification module inputs the spliced tokens into ViT for global feature extraction and feature fusion, and performs classification based on the fused features.
2. The deep fake detection method based on graph convolution and multi-scale hint fusion as claimed in claim 1, characterized in that: The dual-branch feature extraction network includes two branches, which are used to extract graph features and frequency features respectively; The image features are extracted from the face image using CNN; The frequency features are extracted from the frequency domain representation of the face image using CNN.
3. The deep fake detection method based on graph convolution and multi-scale hint fusion as claimed in claim 1, characterized in that: The graph structure is composed of a node matrix and an adjacency matrix. The specific construction steps are as follows: Treat each pixel in the graph feature as a node and construct a node matrix; Using the local aggregator, the query, key and value matrices of the node matrix are calculated, and then the similarity scores between the nodes are calculated. Based on the similarity scores, the adjacency matrix is constructed.
4. A deep fake detection system based on graph convolution and multi-scale prompt fusion, adopting the deep fake detection method based on graph convolution and multi-scale prompt fusion as described in any one of claims 1-3, characterized in that: include: The image acquisition module is configured to: acquire a face image to be detected; The forgery detection module is configured to: input the face image into the trained detection model, perform binary classification on whether it is forged, and obtain a deep forgery detection result; The detection model extracts graph features and frequency features in the face image, performs graph convolution on the graph features using an adaptive graph convolution module, aggregates neighbor node information to form group tokens, and generates prompt tokens from the graph features and frequency features using a multi-scale prompt fusion module; Finally, the concatenated tokens of the group token, the prompt token, and the learnable classification token are input into ViT for classification.
5. An electronic device, comprising: a memory for non-transitory storage of computer readable instructions; a processor for executing the computer-readable instructions; Wherein, when the computer-readable instructions are executed by the processor, the method according to any one of claims 1 to 3 is executed.
6. A storage medium, characterized in that: The computer-readable instructions are non-transitorily stored, wherein when the computer-readable instructions are executed by a computer, the method according to any one of claims 1 to 3 is performed.
Citation Information
Patent Citations
Deep forgery detection method and corresponding device
CN115311525A
Human body posture estimation method and system for learning rich visual features based on multi-scale Transform
CN117115855A