Face forgery detection method and system based on space-time hyperbolic hypergraph

The spatio-temporal hypergraph network addresses the challenge of capturing complex facial video structures by modeling intra-frame and inter-frame relationships, improving detection accuracy and generalization in facial forgery detection.

CN120318889AActive Publication Date: 2025-07-15NANJING UNIV OF POSTS & TELECOMM

Patent Information

Application Number
CN202510750870.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-06
Publication Date
2025-07-15
Estimated Expiration
2045-06-06

AI Technical Summary

Technical Problem

When facing deep forgery technology, existing face forgery detection technology is difficult to effectively explore the complex space-time structure and semantic information of face videos, resulting in insufficient detection accuracy and generalization capabilities, which makes it difficult to meet actual needs.

Method used

The face forgery detection method based on space-time hyperbolic hypergraph is adopted, and the face area is divided by the face feature partition extraction module, the space-time hypergraph is constructed, and the hyperbolic hypergraph learning module is used to extract and classify feature learning capabilities.

Benefits of technology

It realizes more accurate spatial and temporal feature extraction, improves the accuracy and generalization ability of face forgery detection, and can better identify authentic face videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120318889A_ABST
    Figure CN120318889A_ABST
Patent Text Reader

Abstract

The invention discloses a face forgery detection method and system based on a space-time hyperbolic hypergraph in the technical field of image processing, and aims to solve the problem that the internal feature relation of a face structure is difficult to capture in the prior art. The method comprises the following steps: inputting a face video of which the authenticity is to be detected into a trained face counterfeiting detection model; performing feature extraction on the face video of which the authenticity is to be detected through a face feature partition extraction module to obtain face local features; according to the face local features, constructing a space-time hypergraph through a space-time node hypergraph construction module; performing semantic classification feature extraction on the space-time hypergraph through a hyperbolic hypergraph learning module to obtain classification features; and inputting the classification features into a classifier to obtain a true or false detection result. According to the invention, high-accuracy and high-generalization face counterfeiting detection can be realized, the method can be widely applied to the fields of security and protection, finance and the like, and reliable guarantee is provided for related face recognition application.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of image processing, and particularly to a face forgery detection method and system based on a spatio-temporal hyperbolic hypergraph. Background Art

[0002] Traditional face video forgery generally can only achieve simple facial replacement effects, and the results are not only rough but also easily recognizable. Recently, with the rapid development of deep learning in the field of computer vision, face forgery technology has also made certain breakthroughs. Existing advanced face deep forgery technologies, such as generative adversarial networks, diffusion models, etc., can not only achieve high-precision face swapping, but also perform accurate face attribute editing and realistic face synthesis, making a large number of fake face images and videos reach the level of being indistinguishable from the real ones. Although these deep face forgery technologies can enrich the form of content creation and open up a space for content creativity, they also pose potential threats to society and the general public.

[0003] Therefore, face forgery detection technology has emerged, aiming to accurately judge the authenticity of face images or videos. Early detection methods had certain effects on simple forgery samples, but they would be stretched when facing the increasingly updated deep forgery technologies, and detection methods with low accuracy often led to serious security problems. After the rise of deep learning, face forgery detection technology has been improved to a certain extent. However, due to the still lack of sufficient exploration of the spatio-temporal complex structure and semantic information of face videos, there are still shortcomings in detection accuracy and generalization ability, and thus it is difficult to meet the actual needs.

[0004] Although some existing face forgery detection models have begun to attempt to consider the spatio-temporal complex structure and semantic information of face videos, most of them perform feature learning and processing in the Euclidean space. When representing data with complex hierarchical structures and highly sparse characteristics, the Euclidean space often faces great challenges, and thus it is difficult to capture the internal feature connections of the face structure. In practical applications, face videos often not only contain pixel-level detailed information, but also have obvious hierarchical structures, such as the spatial hierarchical relationships between different facial organs and the temporal and hierarchical relationships between different frames in the video. Therefore, the problem of how to effectively mine face structured information and make full use of its hierarchical relationships for feature extraction and discrimination remains unresolved. Summary of the Invention

[0005] The purpose of the present invention is to overcome the deficiencies in the prior art and provide a face forgery detection method and system based on a spatio-temporal hyperbolic hypergraph, aiming to model the intra-frame structure, inter-frame temporality, and cross-frame similarity of face videos in the form of a spatio-temporal hypergraph, and introduce a hyperbolic hypergraph network to focus on the learning of feature hierarchical relationships, so as to achieve highly accurate and strongly generalized face forgery detection.

[0006] To solve the above technical problems, the present invention is implemented by the following technical solutions:

[0007] In a first aspect, the present invention provides a face forgery detection method based on a spatio-temporal hyperbolic hypergraph, including:

[0008] Input the face video to be detected for authenticity into a trained face forgery detection model: extract features from the face video to be detected for authenticity through a face feature partition extraction module to obtain local face features; construct a spatio-temporal hypergraph according to the local face features through a spatio-temporal node hypergraph construction module; extract semantic classification features from the spatio-temporal hypergraph through a hyperbolic hypergraph learning module to obtain classification features; input the classification features into a classifier to obtain a detection result of true or false.

[0009] Optionally, the face feature partition extraction module includes a face region extraction module and a local feature extraction module;

[0010] The face region extraction module includes a face key point detector and a facial region divider; the face key point detector locates, according to the face video to be detected for authenticity , the key points of each frame of image ; the facial region divider divides the face into facial regions according to the key points ; where , , , , represents the number of frames of the face video to be detected for authenticity, represents the frame number, represents the th frame of image, represents the th frame of image, the th key point, represents the th frame of image, the th facial region;

[0011] The local feature extraction module extracts features from each facial region using a pre-trained convolutional neural network to obtain the feature vector of the facial region , and outputs the local face feature , where , , represents the set of feature vectors of the th frame of image, the facial regions, .

[0012] Optionally, the spatio-temporal node hypergraph construction module uses local face features as nodes and constructs hyperedges using a rule-based hyperedge construction method and a similarity-based hyperedge construction method to obtain a spatio-temporal hypergraph;

[0013] The expression of the spatio-temporal hypergraph is:

[0014] ;

[0015] where represents the node set, , represents the th node, represents the number of frames of the face video whose authenticity is to be detected, represents the total number of facial regions, represents the hyperedge set, , represents the time-dimension hyperedge, represents the region-dimension hyperedge, represents the similarity hyperedge.

[0016] Optionally, the rule-based hyperedge construction method establishes connection relationships between nodes from the time dimension and the region dimension to obtain time-dimension hyperedges and region-dimension hyperedges;

[0017] The time-dimension hyperedge is constructed for nodes in different frames with the same facial region, and the expression of the time-dimension hyperedge is:

[0018] ;

[0019] where represents the th frame of the image, the th facial region's feature vector corresponding node, represents the th frame of the image, the th facial region's feature vector corresponding node;

[0020] The region-dimension hyperedge is constructed for different facial regions of nodes in the same frame, and the expression of the region-dimension hyperedge is:

[0021] ;

[0022] where represents the th frame of the image, the th facial region's feature vector The corresponding node, represents the th feature vector of the th facial region in the th frame image;

[0023] The similarity-based hyperedge construction method selects the nodes most similar to this node through a node similarity function to obtain a similarity hyperedge;

[0024] The expression of the similarity hyperedge is:

[0025] ;

[0026] where represents the th node, represents the set of most similar nodes, represents the th similar node of , represents represents th similar node of , represents represents th similar node of .

[0027] Optionally, the hyperbolic hypergraph learning module includes a hyperbolic exponential mapping layer, a hyperbolic learning module, and a hyperbolic logarithmic mapping layer;

[0028] Semantic classification feature extraction from the spatio-temporal hypergraph by the hyperbolic hypergraph learning module includes:

[0029] Using the hyperbolic exponential mapping layer to map the feature vector of the spatio-temporal hypergraph from the Euclidean space to the hyperbolic space to obtain the mapped feature;

[0030] According to the mapped feature, using the hyperbolic learning module for semantic classification feature extraction to obtain hyperbolic space classification features;

[0031] Using the hyperbolic logarithmic mapping layer to map the hyperbolic space classification features back to the Euclidean space to obtain classification features.

[0032] Optionally, the expression of the hyperbolic exponential mapping layer is as follows:

[0033] ;

[0034] where represents the hyperbolic curvature, represents the origin, denotes the exponential mapping function with as the origin and as the hyperbolic curvature, denotes the eigenvector of the spatio-temporal hypergraph in Euclidean space;

[0035] The expression of the hyperbolic logarithmic mapping layer is as follows:

[0036] ;

[0037] where, denotes the hyperbolic space classification feature, denotes the logarithmic mapping function with as the origin and as the hyperbolic curvature.

[0038] Optionally, the hyperbolic learning module at least includes a set of serially connected hyperbolic feature transformation layers, hyperbolic neighborhood aggregation layers, and non-linear activation layers;

[0039] Using the hyperbolic learning module for semantic classification feature extraction includes:

[0040] Using the hyperbolic feature transformation layer to perform a linear transformation on the mapped features to obtain linearly transformed features;

[0041] According to the linearly transformed features and the linearly transformed features of the neighborhood nodes corresponding to the linearly transformed features, using the hyperbolic neighborhood aggregation layer for weighted aggregation to obtain aggregated features;

[0042] Using the non-linear activation layer to perform a non-linear transformation on the aggregated features to obtain hyperbolic space classification features.

[0043] Optionally, the expression of the hyperbolic feature transformation layer is as follows:

[0044] ;

[0045] where, denotes the eigenvector of the th node in the th group of hyperbolic feature transformation layers, denotes the eigenvector of the th node in the th group of hyperbolic feature transformation layers, denotes the weight matrix of the th group of hyperbolic feature transformation layers, denotes the bias term of the th group of hyperbolic feature transformation layers, denotes the Möbius scalar multiplication;

[0046] The expression of the hyperbolic neighborhood aggregation layer is as follows:

[0047] ;

[0048] where, represents the feature of the th node, represents 's aggregated feature, represents the hyperbolic curvature, represents the th node and the th node's hyperedge, represents the weight of the hyperedge , represents the set of features of the neighborhood nodes of the th node, represents the as the origin and as the logarithmic mapping function of the hyperbolic curvature, represents the as the origin and as the exponential mapping function of the hyperbolic curvature;

[0049] The expression of the non - linear activation layer is as follows:

[0050] ;

[0051] where, represents the input of the non - linear activation layer, represents the output of the non - linear activation layer, represents the th group of hyperbolic curvature of the non - linear activation layer, represents the th group of hyperbolic curvature of the non - linear activation layer, represents the as the origin and as the exponential mapping function of the hyperbolic curvature, represents the as the origin and as the logarithmic mapping function of the hyperbolic curvature, represents the activation function.

[0052] Optionally, the classifier includes 1 fully - connected layer. The input size of the fully - connected layer is the size of the classification feature, and the output size of the fully - connected layer is 2, representing the probabilities of being detected as true and false respectively;

[0053] The training process of the face forgery detection model includes:

[0054] Construct a training dataset, where the training dataset includes real face video data and face video data generated using forgery techniques. All video data is equipped with corresponding labels for distinguishing real data from forged data;

[0055] According to the training dataset, calculate the cross-entropy loss of the face forgery detection model;

[0056] Based on the gradient descent method, iteratively update and train the face forgery detection model to minimize the cross-entropy loss, and obtain a trained face forgery detection model;

[0057] The calculation formula of the cross-entropy loss is:

[0058] ;

[0059] where, represents the label, represents the detection result, represents the detection result and the label the loss between them.

[0060] In a second aspect, the present invention provides a face forgery detection system based on a spatio-temporal hyperbolic hypergraph, including:

[0061] A data acquisition module for: acquiring a face video to be detected for authenticity;

[0062] A face feature partition extraction module for: extracting features from the face video to be detected for authenticity to obtain local face features; obtaining a detection result;

[0063] A spatio-temporal node hypergraph construction module for: constructing a spatio-temporal hypergraph according to the local face features;

[0064] A hyperbolic hypergraph learning module for: extracting semantic classification features from the spatio-temporal hypergraph to obtain classification features;

[0065] A detection module for: inputting the classification features into a classifier to obtain a detection result.

[0066] Compared with the prior art, the beneficial effects achieved by the present invention:

[0067] 1. The face forgery detection method based on spatio-temporal hyperbolic hypergraph provided by the present invention uses a face feature partition extraction module to divide the facial region by locating key points and deeply extract features, fully mining the exclusive features of each facial region; through a spatio-temporal node hypergraph construction module, using local feature data as nodes and combining rules and similarity to construct hyperedges, comprehensively capturing complex high-order relationships in the spatio-temporal dimension and integrating information; through a hyperbolic hypergraph learning module, mapping the spatio-temporal hypergraph to the hyperbolic space for hyperbolic hypergraph convolution aggregation, and using the unique geometric properties of the hyperbolic space to focus on the learning of feature hierarchical relationships and improve the ability of feature learning;

[0068] 2. The face forgery detection system based on spatio-temporal hyperbolic hypergraph provided by the present invention realizes face forgery detection based on spatio-temporal hyperbolic hypergraph by setting a data acquisition module, a face feature partition extraction module, a spatio-temporal node hypergraph construction module, a hyperbolic hypergraph learning module and a detection module, can more accurately extract the spatio-temporal consistency features of face forgery, achieve higher accuracy and stronger generalization ability, and has practical significance and good application prospects. BRIEF DESCRIPTION OF THE DRAWINGS

[0069] Figure 1 is a flowchart of the face forgery detection method based on spatio-temporal hyperbolic hypergraph provided by an embodiment of the present invention;

[0070] Figure 2 is a schematic structural diagram of a face forgery detection model provided by an embodiment of the present invention;

[0071] Figure 3 is a schematic structural diagram of the face feature partition extraction module provided by an embodiment of the present invention;

[0072] Figure 4 is a schematic structural diagram of the spatio-temporal node hypergraph construction module provided by an embodiment of the present invention;

[0073] Figure 5 is a schematic structural diagram of the hyperbolic hypergraph learning module provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0074] The technical solutions of the present invention will be described in detail below with reference to the drawings and specific embodiments. It should be understood that the specific features in the embodiments of the present invention are detailed descriptions of the technical solutions of the present invention, rather than limitations on the technical solutions of the present invention. Without conflict, the technical features in the embodiments of the present invention and the embodiments can be combined with each other.

[0075] It should be noted that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Additionally, the character " / " in this text generally indicates that the associated objects before and after are in an "or" relationship.

[0076] Embodiment 1:

[0077] An embodiment of the present invention discloses a face forgery detection method based on a spatio-temporal hyperbolic hypergraph, referring to Figure 1 and Figure 2 as shown, specifically including the following steps:

[0078] S1, input the face video to be detected for authenticity into the trained face forgery detection model:

[0079] S2, extract features from the face video to be detected for authenticity through the face feature partition extraction module to obtain local face features;

[0080] S3, construct a spatio-temporal hypergraph according to the local face features through the spatio-temporal node hypergraph construction module;

[0081] S4, extract semantic classification features from the spatio-temporal hypergraph through the hyperbolic hypergraph learning module to obtain classification features;

[0082] S5, input the classification features into a classifier to obtain a detection result of true or false.

[0083] Specifically, in step S1, the training process of the face forgery detection model includes:

[0084] First, construct a training data set. The training data set includes real face video data and face video data generated using forgery techniques. All video data are equipped with corresponding labels to distinguish real data from forged data; in this embodiment, data is obtained from publicly available face forgery detection data sets such as FaceForensics++ (FF++), DeepFake DetectionChallenge (DFDC), Celeb-DF (CDF), etc.; detect faces using the Dlib tool, perform alignment processing on the detected faces using affine transformation, set a fixed margin (e.g., 1.3) for the face area and then perform a cropping operation, and finally scale the cropped face images to a fixed pixel size (in this example, set to 256*256 pixels);

[0085] Second, input the real face video data and the face video data generated using existing forgery techniques into the pre-constructed face forgery detection model to obtain the model prediction results , where the label of the real face video data is 0, and the label of the face video data generated by existing forgery techniques is 1;

[0086] 3. Calculate the cross-entropy loss of the face forgery detection model according to the training data set; the calculation formula of the cross-entropy loss is:

[0087] ;

[0088] where represents the label, represents the detection result, represents the detection result and the loss between the label ;

[0089] 4. Based on the gradient descent method, perform iterative update training on the face forgery detection model to minimize the cross-entropy loss and obtain a trained face forgery detection model.

[0090] In step S2, as shown in Figure 3 , the face feature partition extraction module includes a face region extraction module and a local feature extraction module; the face feature partition extraction module divides the input face video to be detected for authenticity into regions through the face region extraction module and the local feature extraction module, and generates face local features.

[0091] The face region extraction module includes a face key point detector and a facial region divider; the face key point detector locates the key points of each frame of image according to the face video to be detected for authenticity ; the facial region divider divides the face into facial regions according to the key points ; where represents the number of facial regions, ; where , , , represents the number of frames of the face video to be detected for authenticity, represents the frame number, represents the th frame image, represents the th frame image, represents the th key point in the th frame image, represents the

[0092] th facial region in the th frame image; the local feature extraction module uses a pre-trained convolutional neural network according to each facial region Perform feature extraction to obtain the facial region of the feature vector , and output the local face feature , where , , represents the th frame image feature vector set of the .

[0093] The local feature extraction module performs feature extraction on multiple facial regions to generate local face features; the local feature extraction module uses a pre-trained convolutional neural network for each facial region to perform feature extraction to obtain the facial region of the feature vector , and output the local face feature , where represents the th frame image feature vector set of the

[0094] In step S3, as shown in Figure 4 , the spatio-temporal node hypergraph construction module uses the local face feature as a node and constructs hyperedges using a rule-based hyperedge construction method and a similarity-based hyperedge construction method to obtain a spatio-temporal hypergraph; the rule-based hyperedge construction method establishes a clear connection relationship from two dimensions of time and region; the similarity-based hyperedge construction method finds the most similar nodes for the nodes and constructs hyperedges

[0095] The expression of the spatio-temporal hypergraph is

[0096] ;

[0097] where represents the node set , represents the th node represents the number of frames of the face video to be detected for authenticity represents the total number of facial regions represents the hyperedge set , represents the hyperedge in the time dimension represents the hyperedge in the region dimension represents the similarity hyperedge

[0098] The time - dimension hyper - edge is constructed for nodes in different frames with the same facial region, and the expression of the time - dimension hyper - edge is:

[0099] ;

[0100] Among them, represents the feature vector of the th facial region in the th frame image corresponding node, represents the feature vector of the th facial region in the th frame image corresponding node.

[0101] The region - dimension hyper - edge is constructed for different facial regions of nodes in the same frame, and the expression of the region - dimension hyper - edge is:

[0102] ;

[0103] Among them, represents the feature vector of the th facial region in the th frame image corresponding node, represents the feature vector of the th facial region in the th frame image corresponding node.

[0104] The similarity - based hyper - edge construction method selects the nodes most similar to the node through the node similarity function to obtain the similarity hyper - edge; the similarity can be calculated by ; the expression of the similarity hyper - edge is:

[0105] ;

[0106] Among them, represents the th node, represents the set of the most similar nodes to represents the th similar node of ; represents the th similar node of ; represents the th similar node of ; represents the node The eigenvector of represents the node The eigenvector of represents the norm operation.

[0107] In step S4, referring to Figure 5 As shown, the hyperbolic hypergraph learning module includes a hyperbolic exponential mapping layer, a hyperbolic learning module, and a hyperbolic logarithmic mapping layer;

[0108] Semantic classification feature extraction of the spatio-temporal hypergraph is performed through the hyperbolic hypergraph learning module, including:

[0109] Using the hyperbolic exponential mapping layer, the eigenvector of the spatio-temporal hypergraph is mapped from the Euclidean space to the hyperbolic space to obtain the mapped feature;

[0110] According to the mapped feature, semantic classification feature extraction is performed using the hyperbolic learning module to obtain hyperbolic space classification features;

[0111] Using the hyperbolic logarithmic mapping layer, the hyperbolic space classification features are mapped back to the Euclidean space to obtain classification features.

[0112] The hyperbolic learning module includes L groups of cascaded hyperbolic feature transformation layers, hyperbolic neighborhood aggregation layers, and non-linear activation layers; the hyperbolic feature transformation layer is used to perform linear transformation on the features of the hypergraph nodes; the hyperbolic neighborhood aggregation layer is used to aggregate the information within the node neighborhood to obtain surrounding environment feature information; the non-linear activation layer is used to introduce non-linearity and dynamically adjust the curvature to enhance the model expression and adaptation ability.

[0113] Semantic classification feature extraction using the hyperbolic learning module includes:

[0114] Using the hyperbolic feature transformation layer to perform linear transformation on the mapped feature to obtain a linearly transformed feature;

[0115] According to the linearly transformed feature and the linearly transformed features of the neighborhood nodes corresponding to the linearly transformed feature, weighted aggregation is performed using the hyperbolic neighborhood aggregation layer to obtain an aggregated feature;

[0116] Using the non-linear activation layer to perform non-linear transformation on the aggregated feature to obtain hyperbolic space classification features.

[0117] The expression of the hyperbolic exponential mapping layer is as follows:

[0118] ;

[0119] Where represents the hyperbolic curvature, represents the origin, represents taking with the origin at as the exponential mapping function of the hyperbolic curvature, represents the eigenvector of the spacetime hypergraph in Euclidean space.

[0120] The expression of the hyperbolic logarithmic mapping layer is as follows:

[0121] ;

[0122] where represents the hyperbolic space classification feature, represents with the origin at as the logarithmic mapping function of the hyperbolic curvature.

[0123] The expression of the hyperbolic feature transformation layer is as follows:

[0124] ;

[0125] where represents the eigenvector of the th node in the th group of hyperbolic feature transformation layers, represents the eigenvector of the th node in the th group of hyperbolic feature transformation layers, represents the weight matrix of the th group of hyperbolic feature transformation layers, represents the bias term of the th group of hyperbolic feature transformation layers, represents the Möbius scalar multiplication.

[0126] The expression of the hyperbolic neighborhood aggregation layer is as follows:

[0127] ;

[0128] where represents the feature of the th node, represents after-aggregation feature of represents the hyperbolic curvature, represents the th node and the th node between the hyperedge, represents the hyperedge weight of represents the th node's neighborhood nodes' feature set, represents with the origin at The logarithmic mapping function of hyperbolic curvature denotes with as the origin and with as the exponential mapping function of hyperbolic curvature

[0129] The expression of the non - linear activation layer is as follows:

[0130] ;

[0131] wherein represents the input of the non - linear activation layer, represents the output of the non - linear activation layer, represents the hyperbolic curvature of the nth group of non - linear activation layers, represents the hyperbolic curvature of the mth group of non - linear activation layers, denotes with as the origin and with as the exponential mapping function of hyperbolic curvature, denotes with as the origin and with as the logarithmic mapping function of hyperbolic curvature, represents the activation function

[0132] In step S5, the classifier includes 1 fully - connected layer. The input size of the fully - connected layer is the size of the classification feature, and the output size of the fully - connected layer is 2, representing the probabilities of being detected as true and false respectively.

[0133] Embodiment 2:

[0134] Based on the same inventive concept as Embodiment 1, the present invention embodiment discloses a face forgery detection system based on spatio - temporal hyperbolic hypergraphs, including:

[0135] A data acquisition module, configured to: acquire a face video to be detected for authenticity;

[0136] A face feature partition extraction module, configured to: extract features from the face video to be detected for authenticity to obtain local face features; obtain a detection result;

[0137] A spatio - temporal node hypergraph construction module, configured to: construct a spatio - temporal hypergraph according to the local face features;

[0138] A hyperbolic hypergraph learning module, configured to: extract semantic classification features from the spatio - temporal hypergraph to obtain classification features;

[0139] A detection module, configured to: input the classification features into a classifier to obtain a detection result.

[0140] For the specific function implementation of each of the above modules, refer to the relevant content in the method of Embodiment 1, which will not be elaborated here.

[0141] Those skilled in the art should understand that the embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of an all-hardware embodiment, an all-software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present invention can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memories, CD-ROMs, optical memories, etc.) that contain computer-usable program code.

[0142] The present invention is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the present invention. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a system for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0143] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, so that the instructions stored in the computer-readable memory generate a manufactured article including an instruction system that implements the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0144] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in Figure 1 one or more of the flows Figure 1 or blocks or combinations of blocks.

[0145] The embodiments of the present invention have been described above in conjunction with the accompanying drawings. However, the present invention is not limited to the above specific embodiments. The above specific embodiments are merely illustrative rather than restrictive. Under the inspiration of the present invention, those of ordinary skill in the art can also make many forms without departing from the spirit of the present invention and the scope protected by the claims. All of these fall within the protection scope of the present invention.

Claims

1. A face forgery detection method based on a spatio-temporal hyperbolic hypergraph, characterized in that Including: Input the face video to be detected for authenticity into a trained face forgery detection model: extract features from the face video to be detected for authenticity through a face feature partition extraction module to obtain local face features; construct a spatio-temporal hypergraph based on the local face features through a spatio-temporal node hypergraph construction module; extract semantic classification features from the spatio-temporal hypergraph through a hyperbolic hypergraph learning module to obtain classification features; Input the classification features into a classifier to obtain a detection result of true or false.

2. The face forgery detection method based on spatio-temporal hyperbolic hypergraph according to claim 1, wherein The face feature partition extraction module includes a face region extraction module and a local feature extraction module; The face region extraction module includes a face key point detector and a facial region divider; the face key point detector locates the key points of each frame of the face video to be detected for authenticity ; the facial region divider divides the face into facial regions according to the key points ; where represents the number of frames of the face video to be detected for authenticity, represents the frame number, represents the -th frame image, represents the -th key point in the -th frame image, and represents the -th facial region in the -th frame image; The local feature extraction module, based on each facial region , uses a pre-trained convolutional neural network to perform feature extraction, obtaining the feature vector of the facial region , and outputs the local face feature , where , , represents the feature vector set of the th facial region in the th frame image, .

3. The face forgery detection method based on spatio-temporal hyperbolic hypergraph according to claim 1, characterized in that, The spatio-temporal node hypergraph construction module uses the local face features as nodes and constructs hyperedges using a rule-based hyperedge construction method and a similarity-based hyperedge construction method to obtain a spatio-temporal hypergraph; The expression of the spatio-temporal hypergraph is: ; Among them, represents the node set, , represents the th node, represents the number of frames of the face video whose authenticity is to be detected, represents the total number of facial regions, represents the hyperedge set, , represents the time - dimension hyperedge, represents the region - dimension hyperedge, represents the similarity hyperedge.

4. The face forgery detection method based on spatio-temporal hyperbolic hypergraph according to claim 3, characterized in that, The rule-based hyperedge construction method establishes connection relationships between nodes from the time dimension and the region dimension to obtain time dimension hyperedges and region dimension hyperedges; The time dimension hyperedges are constructed for nodes in different frames with the same facial region, and the expression of the time dimension hyperedges is: ; Among them, represents the feature vector of the nth facial region in the corresponding node, and represents the feature vector of the nth facial region in the corresponding node; The region dimension hyperedges are constructed for different facial regions of nodes in the same frame, and the expression of the region dimension hyperedges is: ; Among them, represents the th frame image of the feature vector of the corresponding node, represents the th frame image of the feature vector of the corresponding node; The similarity-based hyperedge construction method selects the nodes most similar to the node through a node similarity function to obtain a similarity hyperedge; ​ The expression of the similarity hyperedges is: ; Among them, represents the th node, represents the set of most similar nodes, represents the th similar node of represents represents th similar node of represents represents th similar node.

5. The face forgery detection method based on spatio-temporal hyperbolic hypergraph according to claim 1, characterized in that, The hyperbolic hypergraph learning module includes a hyperbolic exponential mapping layer, a hyperbolic learning module, and a hyperbolic logarithm mapping layer; Extracting semantic classification features from the spatio-temporal hypergraph through the hyperbolic hypergraph learning module includes: Using the hyperbolic exponential mapping layer to map the feature vectors of the spatio-temporal hypergraph from the Euclidean space to the hyperbolic space to obtain the mapped features; According to the mapped features, using the hyperbolic learning module to perform semantic classification feature extraction to obtain hyperbolic space classification features; Using the hyperbolic logarithm mapping layer to map the hyperbolic space classification features back to the Euclidean space to obtain classification features.

6. The face forgery detection method based on a spatio-temporal hyperbolic hypergraph according to claim 5, wherein, The expression of the hyperbolic exponential mapping layer is as follows: ; Among them, represents the hyperbolic curvature, represents the origin, represents taking as the origin and as the exponential mapping function of the hyperbolic curvature, represents the eigenvector of the spatio-temporal hypergraph in the Euclidean space; The expression of the hyperbolic logarithm mapping layer is as follows: ; Among them, represents the hyperbolic space classification feature, represents as the origin and is the logarithmic mapping function with as the hyperbolic curvature.

7. The face forgery detection method based on the spatio-temporal hyperbolic hypergraph according to claim 5, wherein The hyperbolic learning module includes at least one set of serially connected hyperbolic feature transformation layers, hyperbolic neighborhood aggregation layers, and non-linear activation layers; Using the hyperbolic learning module to perform semantic classification feature extraction includes: Using the hyperbolic feature transformation layer to perform a linear transformation on the mapped features to obtain linearly transformed features; According to the linearly transformed features and the linearly transformed features of the neighborhood nodes corresponding to the linearly transformed features, using the hyperbolic neighborhood aggregation layer to perform weighted aggregation to obtain aggregated features; Using the non-linear activation layer to perform a non-linear transformation on the aggregated features to obtain hyperbolic space classification features.

8. The face forgery detection method based on spatio-temporal hyperbolic hypergraph according to claim 7, wherein The expression of the hyperbolic feature transformation layer is as follows: ; Among them, represents the eigenvector of the th node in the th group of hyperbolic feature transformation layers, represents the eigenvector of the th node in the th group of hyperbolic feature transformation layers, represents the weight matrix of the th group of hyperbolic feature transformation layers, represents the bias term of the th group of hyperbolic feature transformation layers, represents the Möbius scalar multiplication; The expression of the hyperbolic neighborhood aggregation layer is as follows: ; Among them, represents the feature of the th node, represents the aggregated feature of, represents the hyperbolic curvature, represents the th node and the th node's hyperedge, represents the hyperedge 's weight, represents the feature set of the neighborhood nodes of the th node, represents the logarithmic mapping function with as the origin and as the hyperbolic curvature, represents the exponential mapping function with as the origin and as the hyperbolic curvature; The expression of the non-linear activation layer is as follows: ; Among them, represents the input of the non-linear activation layer, represents the output of the non-linear activation layer, represents the hyperbolic curvature of the nth group of non-linear activation layers, represents the hyperbolic curvature of the mth group of non-linear activation layers, represents the exponential mapping function with as the origin and as the hyperbolic curvature, represents the logarithmic mapping function with as the origin and as the hyperbolic curvature, represents the activation function.

9. The face forgery detection method based on spatio-temporal hyperbolic hypergraph according to claim 1, characterized in that, The classifier contains 1 fully connected layer, the input size of the fully connected layer is the size of the classification features, and the output size of the fully connected layer is 2, respectively representing the probabilities of being detected as true and false; The training process of the face forgery detection model includes: Construct a training dataset, where the training dataset includes real face video data and face video data generated using forgery techniques, and all video data are equipped with corresponding labels for distinguishing real data from forged data; Calculate the cross-entropy loss of the face forgery detection model according to the training dataset; Iteratively update and train the face forgery detection model based on the gradient descent method to minimize the cross-entropy loss and obtain a trained face forgery detection model; The calculation formula of the cross-entropy loss is: ; Among them, represents the label, represents the detection result, represents the detection result and the label the loss between them.

10. A face forgery detection system based on a spatio-temporal hyperbolic hypergraph, characterized in that, including: A data acquisition module for: acquiring a face video to be detected for authenticity; A face feature partition extraction module for: extracting features from the face video to be detected for authenticity to obtain local face features; obtaining a detection result; A spatio-temporal node hypergraph construction module for: constructing a spatio-temporal hypergraph according to the local face features; A hyperbolic hypergraph learning module for: extracting semantic classification features from the spatio-temporal hypergraph to obtain classification features; A detection module for: inputting the classification features into a classifier to obtain a detection result.

Citation Information

Patent Citations

  • Visual object retrieval method and device based on hyperbolic hypergraph convolution

    CN116665006A

  • Deep forgery detection method based on face geometrical relationship reasoning

    CN116758604A

  • ResNet and LSTM (Long Short Term Memory)-based multi-dimensional spatio-temporal feature fusion deep counterfeit video detection method and device

    CN119339220A

  • Facial key point-based counterfeit speaking face detection method and system

    CN119964215A

  • Line laser vision sensing device

    KR1020260136242A

Cited By

  • Image forgery hierarchical attribution method based on hyperbolic space and VMama

    CN122336524A

  • A Hierarchical Attribution Method for Image Forgery Based on Hyperbolic Space and VMamba

    CN122336524B