Method for detecting a fake face based on an agent-guided hybrid expert network and related device
The method for detecting fake faces by using a proxy-guided hybrid expert network solves the problem of model bias towards certain fake face types, improves the robustness and generalization ability of fake face detection, and can effectively identify diverse fake face patterns.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- INST OF AUTOMATION CHINESE ACAD OF SCI
- Filing Date
- 2026-01-28
- Publication Date
- 2026-04-28
AI Technical Summary
In existing technologies, fake face detection models tend to be biased towards certain types of fake faces, making it difficult to fully cover diverse fake face patterns and lacking robustness.
A fake face detection method based on agent-guided hybrid expert network is adopted. By acquiring the face image to be detected, the similarity between the expert-extracted features and the learnable agent vectors under multiple preset modes is calculated using a hybrid expert feature extractor. The model is trained by combining batch regularization loss, cross-entropy loss, agent optimization loss and feature reconstruction loss to improve generalization ability.
It achieves stable detection performance when facing various types of forgery, can comprehensively cover diverse forgery patterns, and improves the robustness and generalization ability of the model.
Smart Images

Figure CN121600581B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of computer vision technology, and more specifically, to a method and related equipment for detecting fake faces based on agent-guided hybrid expert networks. Background Technology
[0002] In recent years, with the rapid development of face generation and editing technologies (such as image generation technologies based on generative adversarial networks and style transfer), virtual makeup, face reshaping, and post-production face replacement in film and television have been widely applied. However, if these technologies are misused, they may cause serious social problems. For example, identity forgery, the spread of false information, and deepfake videos may lead to public opinion misguidance and fraud risks. Therefore, how to effectively and reliably detect forged face content has become a key issue in ensuring information security and social trust.
[0003] In related technologies, various enhancement strategies exist, such as latent space-based data augmentation methods to expand feature distribution; feature decoupling methods to separate the semantics of real and fake features; and spatiotemporal consistency constraint methods to model changes in forgery traces between video frames. However, different deepfake algorithms exhibit significant differences in the features they forge at the spatial, temporal, and texture levels. Feature representation learning-based methods in related technologies still suffer from insufficient robustness when facing novel or complex forgery patterns. Furthermore, some prototype-based detection methods can model intra-class invariance by constructing prototype vectors for real and fake classes. However, these methods often rely on manually designed forged feature samples, resulting in a limited number of samples and making the model prone to bias towards certain forgery types, failing to comprehensively cover diverse forgery patterns. Summary of the Invention
[0004] This disclosure provides a method and related equipment for detecting fake faces based on agent-guided hybrid expert networks, in order to at least solve the problem in the above-mentioned related technologies that the models tend to be biased towards certain types of fake faces and are difficult to fully cover diverse fake face patterns.
[0005] According to a first aspect of the present disclosure, a method for detecting fake faces is provided, applied to a fake face detection model including a hybrid expert feature extractor. The method includes: acquiring a face image to be detected; inputting the face image to be detected into the hybrid expert feature extractor to obtain expert-extracted features; calculating the similarity between the expert-extracted features and k learnable proxy vectors in each of a plurality of preset modes; for each preset mode, calculating the average value of the similarity between the expert-extracted features and the k learnable proxy vectors in that preset mode; determining the maximum average value among the average values of the plurality of similarities corresponding one-to-one with the plurality of preset modes; and determining the classification label of the preset mode to which the maximum average value belongs as the authenticity detection result of the face image, wherein the classification label is real or fake.
[0006] Optionally, the hybrid expert feature extractor comprises multiple expert sub-networks; the fake face detection model is trained as follows: a current batch of samples is acquired, wherein the current batch of samples contains multiple face image samples, each face image sample corresponding to a pattern label, the pattern label indicating whether the corresponding face image sample is a fake face sample and the corresponding fake pattern, the multiple preset patterns being the patterns indicated by the pattern labels corresponding to the multiple face image samples; each face image sample is input into the hybrid expert feature extractor to obtain the features extracted by the training experts and the weights assigned to each expert sub-network in the multiple expert sub-networks. The weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples. Calculate the batch-level regularization term loss. Based on the features extracted by the training expert and the corresponding pattern labels for each face image sample from the multiple face image samples, the cross-entropy loss is calculated. Based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable proxy vectors in each of the multiple preset modes, the proxy optimization loss is calculated. and feature reconstruction loss Based on the batch-level regularization term loss The cross-entropy loss The aforementioned agent optimization loss and the feature reconstruction loss The parameters of the fake face detection model are adjusted for training.
[0007] Optionally, the agent optimization loss is calculated based on the expert-extracted features corresponding to each face image sample in the plurality of face image samples and the k learnable agent vectors in each of the plurality of preset modes. The method includes: for each face image sample among the multiple face image samples, calculating a first similarity between the expert-extracted features corresponding to that face image sample and k learnable proxy vectors in a preset pattern indicated by the pattern label corresponding to that face image sample; calculating a second similarity between the expert-extracted features corresponding to that face image sample and the k learnable proxy vectors in other preset patterns, wherein the classification labels of the other preset patterns are different from the classification labels of the pattern label corresponding to that face image sample; and calculating the proxy optimization loss based on the first similarity and the second similarity corresponding to each face image sample among the multiple face image samples. .
[0008] Optionally, the feature reconstruction loss is calculated based on the features extracted by the trained expert corresponding to each face image sample in the plurality of face image samples and the k learnable proxy vectors in each of the plurality of preset modes. This includes: for each face image sample among the multiple face image samples, extracting features based on the trained expert corresponding to that face image sample, and calculating the masked features. Based on the features of the masked And the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the face image sample, to calculate the reconstructed features. Based on the features extracted and reconstructed by trained experts corresponding to each of the multiple face image samples, features are obtained from the multi-face image samples. Calculate the feature reconstruction loss .
[0009] Optionally, it further includes: for the fake training expert extraction features whose corresponding pattern labels are fake labels among the multiple training expert extraction features corresponding to one-to-one with the multiple face image samples, using the spherical linear interpolation method to expand the fake training expert extraction features to obtain Slerp features, wherein the Slerp features correspond to Slerp pattern labels; and calculating the proxy optimization loss based on the training expert extraction features corresponding to each face image sample in the multiple face image samples and the k learnable proxy vectors in each of the multiple preset modes. and feature reconstruction loss This includes: calculating the proxy optimization loss based on the expert-extracted features corresponding to each face image sample in the plurality of face image samples, the Slerp features, k learnable proxy vectors in each preset mode in the plurality of preset modes, and k learnable proxy vectors in the mode indicated by the Slerp mode label. and the feature reconstruction loss .
[0010] Optionally, the weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples... Calculate the batch-level regularization term loss. ,include:
[0011] The batch-level regularization term loss is calculated using the following formula. :
[0012]
[0013] in, The current batch of samples contains multiple face image samples. This refers to the m-th face image sample among the multiple face image samples. for The weights assigned to each of the plurality of expert subnetworks. This represents the standard deviation calculation. This indicates the operation of averaging.
[0014] According to a second aspect of the present disclosure, a forged face detection apparatus is provided, comprising: an image acquisition module configured to acquire a face image to be detected; a feature extraction module configured to input the face image to be detected into a hybrid expert feature extractor included in a forged face detection model to obtain expert-extracted features; a similarity calculation module configured to calculate the similarity between the expert-extracted features and k learnable proxy vectors in each of a plurality of preset modes; an average value calculation module configured to calculate, for each preset mode, the average value of the similarity between the expert-extracted features and the k learnable proxy vectors in that preset mode; a maximum average value determination module configured to determine the maximum average value among the average values of the plurality of similarities corresponding one-to-one with the plurality of preset modes; and a authenticity detection result determination module configured to determine the classification label of the preset mode corresponding to the maximum average value as the authenticity detection result of the face image, wherein the classification label is real or forged.
[0015] Optionally, the hybrid expert feature extractor comprises multiple expert sub-networks; the fake face detection model is trained as follows: a current batch of samples is acquired, wherein the current batch of samples contains multiple face image samples, each face image sample corresponding to a pattern label, the pattern label indicating whether the corresponding face image sample is a fake face sample and the corresponding fake pattern, the multiple preset patterns being the patterns indicated by the pattern labels corresponding to the multiple face image samples; each face image sample is input into the hybrid expert feature extractor to obtain the features extracted by the training experts and the weights assigned to each expert sub-network in the multiple expert sub-networks. The weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples. Calculate the batch-level regularization term loss. Based on the features extracted by the training expert and the corresponding pattern labels for each face image sample from the multiple face image samples, the cross-entropy loss is calculated. Based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable proxy vectors in each of the multiple preset modes, the proxy optimization loss is calculated. and feature reconstruction loss Based on the batch-level regularization term loss The cross-entropy loss The aforementioned agent optimization loss and the feature reconstruction loss The parameters of the fake face detection model are adjusted for training.
[0016] Optionally, the agent optimization loss is calculated based on the expert-extracted features corresponding to each face image sample in the plurality of face image samples and the k learnable agent vectors in each of the plurality of preset modes. The method includes: for each face image sample among the multiple face image samples, calculating a first similarity between the expert-extracted features corresponding to that face image sample and k learnable proxy vectors in a preset pattern indicated by the pattern label corresponding to that face image sample; calculating a second similarity between the expert-extracted features corresponding to that face image sample and the k learnable proxy vectors in other preset patterns, wherein the classification labels of the other preset patterns are different from the classification labels of the pattern label corresponding to that face image sample; and calculating the proxy optimization loss based on the first similarity and the second similarity corresponding to each face image sample among the multiple face image samples. .
[0017] Optionally, the feature reconstruction loss is calculated based on the features extracted by the trained expert corresponding to each face image sample in the plurality of face image samples and the k learnable proxy vectors in each of the plurality of preset modes. This includes: for each face image sample among the multiple face image samples, extracting features based on the trained expert corresponding to that face image sample, and calculating the masked features. Based on the features of the masked And the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the face image sample, to calculate the reconstructed features. Based on the features extracted and reconstructed by trained experts corresponding to each of the multiple face image samples, features are obtained from the multi-face image samples. Calculate the feature reconstruction loss .
[0018] Optionally, it further includes: for the fake training expert extraction features whose corresponding pattern labels are fake labels among the multiple training expert extraction features corresponding to one-to-one with the multiple face image samples, using the spherical linear interpolation method to expand the fake training expert extraction features to obtain Slerp features, wherein the Slerp features correspond to Slerp pattern labels; and calculating the proxy optimization loss based on the training expert extraction features corresponding to each face image sample in the multiple face image samples and the k learnable proxy vectors in each of the multiple preset modes. and feature reconstruction loss This includes: calculating the proxy optimization loss based on the expert-extracted features corresponding to each face image sample in the plurality of face image samples, the Slerp features, k learnable proxy vectors in each preset mode in the plurality of preset modes, and k learnable proxy vectors in the mode indicated by the Slerp mode label. and the feature reconstruction loss .
[0019] Optionally, the weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples... Calculate the batch-level regularization term loss. ,include:
[0020] The batch-level regularization term loss is calculated using the following formula. :
[0021]
[0022] in, The current batch of samples contains multiple face image samples. This refers to the m-th face image sample among the multiple face image samples. for The weights assigned to each of the plurality of expert subnetworks. This represents the standard deviation calculation. This indicates the operation of averaging.
[0023] According to a third aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing processor-executable instructions; wherein the processor is configured to execute the instructions to implement a fake face detection method according to the present disclosure.
[0024] According to a fourth aspect of the present disclosure, a computer-readable storage medium is provided that, when instructions in the computer-readable storage medium are executed by a processor of an electronic device, enables the electronic device to perform a fake face detection method according to the present disclosure.
[0025] According to a fifth aspect of the present disclosure, a computer program product is provided, including a computer program that, when executed by a processor, implements a method for detecting fake faces according to the present disclosure.
[0026] The technical solutions provided by the embodiments of this disclosure have at least the following beneficial effects:
[0027] In this disclosure, by introducing a surrogate learning mechanism, namely by maintaining a learnable surrogate vector, the characteristics of real patterns and various forgery patterns can be effectively represented. This enables the establishment of a learnable surrogate space between real features and forgery features, thereby improving the model's generalization ability. In other words, it ensures that the model can maintain stable detection performance when facing various forgery types or when no forgery types are encountered, thus ensuring comprehensive coverage of diverse forgery patterns and achieving more robust real-fake discrimination.
[0028] It should be understood that the above general description and the following detailed description are exemplary and explanatory only, and are not intended to limit this disclosure. Attached Figure Description
[0029] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure, and are not intended to unduly limit this disclosure.
[0030] Figure 1 This is a flowchart illustrating a method for detecting fake faces according to exemplary embodiments of the present disclosure;
[0031] Figure 2This is a schematic diagram illustrating the extraction of face data from raw video using a Multi-task Cascaded Convolutional Networks (MTCNN) extractor according to an exemplary embodiment of the present disclosure.
[0032] Figure 3 This is a schematic diagram illustrating the structure of a hybrid expert feature extractor according to an exemplary embodiment of the present disclosure;
[0033] Figure 4 This is a schematic diagram illustrating an agent vector learning mechanism according to an exemplary embodiment of the present disclosure;
[0034] Figure 5 This is a block diagram illustrating a fake face detection apparatus according to an exemplary embodiment of the present disclosure;
[0035] Figure 6 This is a block diagram illustrating an electronic device according to exemplary embodiments of the present disclosure. Detailed Implementation
[0036] To enable those skilled in the art to better understand the technical solutions of this disclosure, the technical solutions in the embodiments of this disclosure will be clearly and completely described below with reference to the accompanying drawings.
[0037] It should be noted that the terms "first," "second," etc., used in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. The embodiments described in the following examples do not represent all embodiments consistent with this disclosure. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this disclosure as detailed in the appended claims.
[0038] It should be noted that the phrase "at least one of several items" in this disclosure refers to three parallel cases: "any one of the several items", "a combination of any number of the several items", and "all of the several items". For example, "including at least one of A and B" includes the following three parallel cases: (1) including A; (2) including B; (3) including A and B. Another example is "performing at least one of step one and step two", which means the following three parallel cases: (1) performing step one; (2) performing step two; (3) performing both step one and step two.
[0039] Figure 1This is a flowchart illustrating a fake face detection method according to an exemplary embodiment of the present disclosure, applied to a fake face detection model including a hybrid expert feature extractor.
[0040] Reference Figure 1 In step 101, the face image to be detected can be obtained. For example, but not limited to, MTCNN can be used to obtain the face data to be detected from the image. Figure 2 This is a schematic diagram illustrating the extraction of face data from raw video using an MTCNN extractor according to an exemplary embodiment of the present disclosure.
[0041] In step 102, the face image to be detected can be input into the hybrid expert feature extractor to obtain expert-extracted features.
[0042] In step 103, the similarity between the expert-extracted features and the k learnable proxy vectors in each of the multiple preset patterns can be calculated.
[0043] For example, multiple preset modes can include real modes and fake modes. Real modes are modes that have not been faked; that is, any mode featuring a face image that has not been faked or altered can be considered a real mode. Fake modes refer to face images that have been faked or altered. Furthermore, depending on the method of fakery, fake modes can be categorized into different types, such as, but not limited to: face-swapping modes (e.g., replacing one person's face image with another person's face image), identity-swapping modes (e.g., replacing one person's identity with another person's identity), color-changing modes (e.g., skin whitening), etc.
[0044] In step 104, for each preset mode, the average similarity between the expert-extracted features and the k learnable proxy vectors in that preset mode can be calculated.
[0045] For example, suppose multiple preset patterns form a set C, then for each preset pattern... The expert-extracted features and the preset pattern can be calculated using the following formula. The average similarity among the k learnable agent vectors:
[0046]
[0047] in, This indicates that experts extract features. With preset mode The average similarity among the k learnable agent vectors. Preset mode The next One proxy vector, This is the cosine similarity function.
[0048] In step 105, the maximum average value among the average values of multiple similarities corresponding to multiple preset patterns can be determined.
[0049] In step 106, the category label of the preset mode corresponding to the maximum average value can be determined as the authenticity detection result of the face image, wherein the category label can be real or fake.
[0050] It should be noted that in this disclosure, there is only one real mode, and the category label corresponding to the real mode is "real". The forgery mode can include multiple forgery types, but the category label corresponding to all forgery modes is "forgery". For example, as mentioned above, forgery modes can include, but are not limited to: face-swapping mode, identity-swapping mode, color-swapping mode, etc., but the category label corresponding to these different types of forgery modes is the same, namely "forgery".
[0051] According to exemplary embodiments of this disclosure, the hybrid expert feature extractor described above can be based on a frozen VisionTransformer (ViT), and a hybrid expert layer can be embedded in each Transformer Block. This allows for dynamic selection of the most suitable expert module through a gating mechanism to extract diverse features adapted to different forgery patterns. Furthermore, the hybrid expert feature extractor constructed based on the frozen ViT backbone can also include multiple expert subnetworks. Each expert subnetwork can employ the same lightweight bottleneck convolutional structure. For example, each expert subnetwork may include... Convolution is used for feature transformation and Convolution is used to capture clues to spatial forgery.
[0052] The aforementioned fake face detection model can be trained in the following way:
[0053] First, the current batch of samples can be obtained. This batch can contain multiple face image samples, each with a corresponding pattern label. This pattern label indicates whether the corresponding face image sample is a forged face sample and the corresponding forgery pattern. The aforementioned preset patterns can be the patterns indicated by the pattern labels corresponding to the multiple face image samples.
[0054] Then, each face image sample can be input into the hybrid expert feature extractor to obtain the features extracted by the trained experts and the weights assigned to each expert subnetwork in the multiple expert subnetworks. Next, weights can be assigned to each expert subnetwork based on each face image sample from multiple face image samples. Calculate the batch-level regularization term loss. .
[0055] Next, we can extract features and corresponding pattern labels from each face image sample in multiple face image samples using trained experts, and then calculate the cross-entropy loss. Furthermore, based on the features extracted by the trained expert corresponding to each face image sample from multiple face image samples, and the k learnable agent vectors in each of the multiple preset modes, the agent optimization loss can be calculated. and feature reconstruction loss .
[0056] Finally, the calculated batch regularization term loss can be used as a basis. Cross-entropy loss Losses from agent optimization and feature reconstruction loss The parameters of the fake face detection model are adjusted for training. For example, the total loss function can be calculated using the following formula. Therefore, it is possible to base this on the total loss function. Adjust the parameters of the fake face detection model for training:
[0057]
[0058] In this way, by jointly optimizing multiple loss functions during the training phase, the robustness and generalization ability of the model can be enhanced.
[0059] For example, the features extracted by the trained experts and the weights assigned to each of the multiple expert subnetworks can be calculated using a hybrid expert feature extractor in the following manner. :
[0060] First, features can be extracted from each face image sample. Then, features can be analyzed. The gated input vector is obtained by performing average pooling and reshaping operations. Next, the pre-activation score for each expert subnetwork can be calculated:
[0061]
[0062] in, For the gated weight matrix, Here is the noise weight matrix. This represents the standard normally distributed noise term. Indicates assignment to the first The pre-activation score of an expert subnetwork.
[0063] Then, a Top-k noise gating mechanism can be used to dynamically select the expert subnetwork, thereby retaining only the top-k noise subnetwork. The expert subnetwork with the highest pre-activation score is selected, and the remaining expert subnetworks can be set to negative infinity. ):
[0064]
[0065] in, Indicates the selection of input vectors Center front The value of the largest element is set, and all other elements are set to negative infinity. ).
[0066] Next, the gating score can be calculated and normalized:
[0067]
[0068] Among them, the gating score It can be the selection weight of the corresponding expert subnetwork in the hybrid expert network for the corresponding feature.
[0069] Then, the features extracted by the training expert can be calculated using the following formula. :
[0070]
[0071] in, Indicates allocation to the first The weights of each expert subnetwork Indicates the first The output of an expert subnetwork Indicates input features The network was restructured to meet the needs of each expert subnetwork.
[0072] Figure 3 This is a schematic diagram illustrating the structure of a hybrid expert feature extractor according to an exemplary embodiment of the present disclosure. (Refer to...) Figure 3 GateG() represents the aforementioned gate weight matrix. and noise weight matrix .
[0073] It should be noted that, in order to avoid the gated network being concentrated in only a few expert subnetworks, a batch-level regularization term loss can be introduced. To constrain the distribution balance of gating scores.
[0074] According to an exemplary embodiment of this disclosure, the batch regularization term loss can be calculated using the following formula. :
[0075]
[0076] in, This refers to multiple face image samples included in the current batch of samples. Let m be the m-th face image sample from a set of multiple face image samples. for The weights assigned to each expert subnetwork in the multiple expert subnetworks. This represents the standard deviation calculation. This represents the mean operation, and... and Used to measure the balance of gating assignments.
[0077] Alternatively, the output of the expert subnetwork can be used as a global feature representation. It can be L2 normalized (L2 Norm) to obtain Then, the normalized features It can be input into a linear classifier and can be lost through cross-entropy loss. Optimize.
[0078] In this disclosure, feature expansion strategies may also be employed, such as, but not limited to, spherical linear interpolation (SLEP) methods, to expand the original feature space to generate synthetic features that simulate unseen fake samples, thereby enhancing the generalization ability of the model.
[0079] According to an exemplary embodiment of this disclosure, for fake training expert extraction features whose corresponding pattern labels are fake labels among multiple training expert extraction features corresponding to multiple face image samples one-to-one, the fake training expert extraction features can be extended using the spherical linear interpolation method to obtain Slerp features, wherein each Slerp feature can correspond to a Slerp pattern label. For example, each Slerp feature can correspond to its own Slerp pattern label, and the Slerp pattern labels corresponding to different Slerp features can be different from each other.
[0080] Then, based on the features extracted by the trained expert, the Slerp features, the k learnable proxy vectors in each preset mode from multiple face image samples, and the k learnable proxy vectors in the mode indicated by the Slerp mode label, the proxy optimization loss can be calculated. and feature reconstruction loss .
[0081] For example, for each forged feature It can randomly sample another fake feature And can extract interpolation coefficients. Then, the interpolation features can be calculated using the slerp method:
[0082]
[0083]
[0084] in and All of these can be L2-normalized features. This interpolation method can effectively generate synthetic features located between two fake samples, thereby filling the gaps in the feature space, i.e., it can achieve implicit expansion to cover potential unseen fake types.
[0085] Next, the generated synthetic features, i.e., slerp features, can be combined with the aforementioned... By merging the sets, an expanded training set can be obtained. .
[0086] It should be noted that, in this disclosure, for each pattern indicated by the pattern labels corresponding to the real sample features and fake sample features (including fake features generated based on slerp) in the extended training feature set, a set of learnable surrogate vectors can be established. Then, optimization can be performed based on the similarity relationship between the features and the surrogate vectors to ensure that the surrogate vectors maintain a balanced and unbiased feature representation for different fake patterns.
[0087] According to an exemplary embodiment of this disclosure, for each face image sample among multiple face image samples, a first similarity can be calculated between the expert-extracted features corresponding to that face image sample and k learnable proxy vectors under a preset pattern indicated by the pattern label corresponding to that face image sample. Furthermore, a second similarity can also be calculated between the expert-extracted features corresponding to that face image sample and the k learnable proxy vectors under other preset patterns. The classification labels of the other preset patterns are different from the classification labels of the pattern label corresponding to that face image sample. Then, a proxy optimization loss can be calculated based on the first and second similarities corresponding to each face image sample among the multiple face image samples. .
[0088] For example, suppose we expand the training feature set. The total includes There are three modes: the real mode, the fake modes included in the original training set, and the newly added Slerp-based fake modes. For each mode... All can be maintained A learnable agent vector These surrogate vectors can be used to represent the feature space of a pattern.
[0089] For expanding the training feature set The first in Features A positive similarity set can be defined. That is, features can be calculated. With features The corresponding pattern label indicates the k learnable agent vectors in the preset pattern. Cosine similarity between :
[0090]
[0091] in, It can represent pattern-level ground truth values, such as: real pattern, fake pattern 1, fake pattern 2, fake pattern 3, ..., slerp-based fake pattern, etc.
[0092] You can also define a negative similarity set. That is, features can be calculated. With each agent vector Cosine similarity between them:
[0093]
[0094] in, Represents the label-level ground truth, i.e. It can represent category labels (e.g., real / fake), and, It can represent and Different categories of agent vector sets. For example, if but It can be a set of proxy vectors for all forgery patterns (e.g., forgery pattern 1, forgery pattern 2, forgery pattern 3, ..., slerp-based forgery patterns); or, if but It can be a set of proxy vectors for all real-world patterns.
[0095] The proxy optimization loss function can be calculated using the following formula. :
[0096]
[0097] in, Indicates the expansion of the training feature set The total number of features in This represents the number of elements in the positive similarity set. This represents the number of elements in the negative similarity set. This loss function can maintain a relatively strong coarse-grained boundary between real and fake data, while encouraging diversity and distribution of proxy vectors in the feature space, thereby enabling the capture of various heterogeneous forgery clues.
[0098] According to an exemplary embodiment of this disclosure, for each face image sample among multiple face image samples, features can be extracted based on the trained expert corresponding to that face image sample, and the masked features can be calculated. Then, it is possible to base the masked features on... And the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the face image sample, to calculate the reconstructed features. Next, features can be extracted and reconstructed based on the training of experts corresponding to each face image sample from multiple face image samples. Calculate the feature reconstruction loss .
[0099] For example, this can be done by expanding the training feature set. Each feature in Perform a random masking operation, and the mask ratio can be set to... This allows us to obtain the masked features. Then, the mask feature can be... Input to encoder This generates preliminary feature representations, and the output features can be L2 normalized.
[0100] Next, the normalized features can be compared with the corresponding class's proxy vector set. The attention score is calculated by performing an inner product operation on the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the normalized feature. Then, the attention score can be calculated by normalizing along the proxy vector dimension. The function obtains the attention weight vector :
[0101]
[0102] in, Indicates the first The attention distribution vector corresponding to each sample mask features The representation obtained by the encoder, For the corresponding category proxy vector (feature) The transpose of the k learnable proxy vectors in the preset mode indicated by the corresponding mode label.
[0103] Next, attention weights can be used. Proxy vectors for positive samples (feature The k learnable proxy vectors in the preset mode indicated by the corresponding mode label are weighted and fused, and can be input into the decoder. In this process, feature reconstruction is performed to obtain reconstructed features. :
[0104]
[0105] Among them, symbols This indicates an element-wise weighted operation.
[0106] Then, the feature reconstruction loss can be obtained by calculating the mean squared error between each original feature and the corresponding reconstructed feature. :
[0107]
[0108] in, To expand the training feature set The total number of samples in the sample, Original features These are the features obtained after proxy reconstruction.
[0109] In this way, by masking some features and redistributing attention weights based on their similarity to the proxy vector, the proxy vector can be guided to reconstruct the extended features through a linear mapping, thereby enhancing its ability to capture subtle forgery clues. Figure 4 This is a schematic diagram illustrating an agent vector learning mechanism according to an exemplary embodiment of the present disclosure.
[0110] Figure 5 This is a block diagram illustrating a fake face detection device 500 according to an exemplary embodiment of the present disclosure.
[0111] Reference Figure 5 The fake face detection device 500 may include an image acquisition module 501, a feature extraction module 502, a similarity calculation module 503, an average value calculation module 504, a maximum average value determination module 505, and a authenticity detection result determination module 506.
[0112] The image acquisition module 501 can acquire images of faces to be detected.
[0113] The feature extraction module 502 can input the face image to be detected into the hybrid expert feature extractor to obtain expert-extracted features.
[0114] The similarity calculation module 503 can calculate the similarity between the expert-extracted features and the k learnable proxy vectors in each of the multiple preset modes.
[0115] For each preset mode, the average value calculation module 504 can calculate the average similarity between the expert-extracted features and the k learnable agent vectors in that preset mode.
[0116] The maximum average value determination module 505 can determine the maximum average value among the average values of multiple similarities corresponding to multiple preset modes.
[0117] The authenticity detection result determination module 506 can determine the authenticity detection result of the face image by the category label of the preset mode corresponding to the maximum average value. The category label can be real or fake.
[0118] According to exemplary embodiments of this disclosure, the above-described fake face detection model can be trained in the following manner:
[0119] First, the current batch of samples can be obtained. This batch can contain multiple face image samples, each with a corresponding pattern label. This pattern label indicates whether the corresponding face image sample is a forged face sample and the corresponding forgery pattern. The aforementioned preset patterns can be the patterns indicated by the pattern labels corresponding to the multiple face image samples.
[0120] Then, each face image sample can be input into the hybrid expert feature extractor to obtain the features extracted by the trained experts and the weights assigned to each expert subnetwork in the multiple expert subnetworks. Next, weights can be assigned to each expert subnetwork based on each face image sample from multiple face image samples. Calculate the batch-level regularization term loss. .
[0121] Next, we can extract features and corresponding pattern labels from each face image sample in multiple face image samples using trained experts, and then calculate the cross-entropy loss. Furthermore, based on the features extracted by the trained expert corresponding to each face image sample from multiple face image samples, and the k learnable agent vectors in each of the multiple preset modes, the agent optimization loss can be calculated. and feature reconstruction loss .
[0122] Finally, the calculated batch regularization term loss can be used as a basis. Cross-entropy loss Losses from agent optimization and feature reconstruction loss The parameters of the fake face detection model were adjusted for training.
[0123] According to an exemplary embodiment of this disclosure, the batch regularization term loss can be calculated using the following formula. :
[0124]
[0125] in, This refers to multiple face image samples included in the current batch of samples. Let m be the m-th face image sample from a set of multiple face image samples. for The weights assigned to each expert subnetwork in the multiple expert subnetworks. This represents the standard deviation calculation. This represents the mean operation, and... and Used to measure the balance of gating assignments.
[0126] According to an exemplary embodiment of this disclosure, for fake training expert extraction features whose corresponding pattern labels are fake labels among multiple training expert extraction features corresponding to multiple face image samples one-to-one, the fake training expert extraction features can be extended using the spherical linear interpolation method to obtain Slerp features, wherein each Slerp feature can correspond to a Slerp pattern label. For example, each Slerp feature can correspond to its own Slerp pattern label, and the Slerp pattern labels corresponding to different Slerp features can be different from each other.
[0127] Then, based on the features extracted by the trained expert, the Slerp features, the k learnable proxy vectors in each preset mode from multiple face image samples, and the k learnable proxy vectors in the mode indicated by the Slerp mode label, the proxy optimization loss can be calculated. and feature reconstruction loss .
[0128] According to an exemplary embodiment of this disclosure, for each face image sample among multiple face image samples, a first similarity can be calculated between the expert-extracted features corresponding to that face image sample and k learnable proxy vectors under a preset pattern indicated by the pattern label corresponding to that face image sample. Furthermore, a second similarity can also be calculated between the expert-extracted features corresponding to that face image sample and the k learnable proxy vectors under other preset patterns. The classification labels of the other preset patterns are different from the classification labels of the pattern label corresponding to that face image sample. Then, a proxy optimization loss can be calculated based on the first and second similarities corresponding to each face image sample among the multiple face image samples. .
[0129] According to an exemplary embodiment of this disclosure, for each face image sample among multiple face image samples, features can be extracted based on the trained expert corresponding to that face image sample, and the masked features can be calculated. Then, it is possible to base the masked features on... And the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the face image sample, to calculate the reconstructed features. Next, features can be extracted and reconstructed based on the training of experts corresponding to each face image sample from multiple face image samples. Calculate the feature reconstruction loss .
[0130] Figure 6 This is a block diagram illustrating an electronic device 600 according to an exemplary embodiment of the present disclosure.
[0131] Reference Figure 6 The electronic device 600 includes at least one memory 601 and at least one processor 602. The at least one memory 601 stores instructions that, when executed by the at least one processor 602, perform a fake face detection method according to an exemplary embodiment of the present disclosure.
[0132] As an example, electronic device 600 may be a PC, tablet, personal digital assistant, smartphone, or other device capable of executing the aforementioned instructions. Here, electronic device 600 is not necessarily a single electronic device, but may be a collection of any devices or circuits capable of executing the aforementioned instructions (or instruction sets) individually or in combination. Electronic device 600 may also be part of an integrated control system or system manager, or may be configured to interconnect with a portable electronic device locally or remotely (e.g., via wireless transmission) through an interface.
[0133] In electronic device 600, processor 602 may include a central processing unit (CPU), a graphics processing unit (GPU), a programmable logic device, a dedicated processor system, a microcontroller, or a microprocessor. By way of example and not limitation, processor may also include analog processors, digital processors, microprocessors, multi-core processors, processor arrays, network processors, etc.
[0134] The processor 602 can execute instructions or code stored in the memory 601, which can also store data. Instructions and data can also be sent and received via a network through a network interface device, which can employ any known transmission protocol.
[0135] The memory 601 may be integrated with the processor 602, for example, by placing RAM or flash memory within an integrated circuit microprocessor. Alternatively, the memory 601 may include a separate device, such as an external disk drive, a storage array, or other storage device that can be used by any database system. The memory 601 and the processor 602 may be operatively coupled, or may communicate with each other, for example, via I / O ports, network connections, etc., enabling the processor 602 to read files stored in the memory.
[0136] In addition, the electronic device 600 may also include a video display (such as a liquid crystal display) and a user interaction interface (such as a keyboard, mouse, touch input device, etc.). All components of the electronic device 600 can be interconnected via a bus and / or network.
[0137] According to exemplary embodiments of this disclosure, a computer-readable storage medium may also be provided, which, when executed by a processor of an electronic device, enables the electronic device to perform the aforementioned method for detecting fake faces. Examples of computer-readable storage media include: read-only memory (ROM), random access programmable read-only memory (PROM), electrically erasable programmable read-only memory (EEPROM), random access memory (RAM), dynamic random access memory (DRAM), static random access memory (SRAM), flash memory, non-volatile memory, CD-ROM, CD-R, CD+R, CD-RW, CD+RW, DVD-ROM, DVD-R, DVD+R, DVD-RW, DVD+RW, DVD-RAM, BD-ROM, BD-R, BD-R LTH, BD-RE, Blu-ray or optical disc storage, hard disk drive (HDD), solid-state drive (SSD), card storage (such as multimedia cards, secure digital (SD) cards, or ultra-fast digital (XD) cards), magnetic tape, floppy disk, magneto-optical data storage device, optical data storage device, hard disk, solid-state drive, and any other device configured to store a computer program and any associated data, data files, and data structures in a non-transitory manner and to provide the computer program and any associated data, data files, and data structures to a processor or computer so that the processor or computer can execute the computer program. The computer program in the aforementioned computer-readable storage medium can run in an environment deployed in computer devices such as clients, hosts, agent devices, servers, etc. Furthermore, in one example, the computer program and any associated data, data files, and data structures are distributed across a networked computer system, such that the computer program and any associated data, data files, and data structures are stored, accessed, and executed in a distributed manner through one or more processors or computers.
[0138] According to exemplary embodiments of the present disclosure, a computer program product may also be provided, including a computer program that, when executed by a processor, implements the fake face detection method according to the present disclosure.
[0139] According to the agent-guided hybrid expert network-based fake face detection method and related equipment disclosed herein, by introducing an agent learning mechanism, that is, by maintaining a learnable agent vector, the characteristics of real patterns and various fake patterns can be effectively represented. In this way, a learnable agent space can be established between real features and fake features, thereby improving the generalization ability of the model. That is, it can ensure that the model can maintain stable detection performance when facing various fake types or no fake types, and can also ensure comprehensive coverage of diverse fake patterns, achieving more robust real and fake discrimination.
[0140] According to exemplary embodiments of this disclosure, the hybrid expert feature extractor described above can be based on a frozen VisionTransformer (ViT) and can embed hybrid expert layers in each Transformer Block. In this way, the most suitable expert module can be dynamically selected through a gating mechanism to extract diverse features adapted to different forgery patterns.
[0141] According to exemplary embodiments of this disclosure, during the training phase, the robustness and generalization ability of the model can be enhanced by jointly optimizing multiple loss functions.
[0142] According to exemplary embodiments of this disclosure, batch-level regularization term loss is introduced. This can prevent the gating network from being concentrated in only a few expert subnetworks, thereby constraining the distribution balance of gating scores.
[0143] According to exemplary embodiments of this disclosure, by employing a feature expansion strategy, the original feature space can be expanded to generate synthetic features simulating unseen forged samples, thereby enhancing the model's generalization ability. For example, by employing spherical linear interpolation (SLEP) to expand the original feature space, synthetic features located between two forged samples can be effectively generated, thereby filling the gaps in the feature space, i.e., implicit expansion can be achieved to cover potential unseen forged types.
[0144] According to an exemplary embodiment of this disclosure, the proxy optimization loss function It can maintain a relatively strong coarse-grained boundary between reality and forgery, while encouraging the diversity and distribution of proxy vectors in the feature space, thereby enabling the capture of various heterogeneous forgery clues.
[0145] According to an exemplary embodiment of this disclosure, by masking some features and reallocating attention weights based on their similarity to the proxy vector, the proxy vector can be guided to reconstruct the extended features through a linear mapping, thereby enhancing its ability to capture subtle forgery clues.
[0146] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the following claims.
[0147] It should be understood that this disclosure is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of this disclosure is limited only by the appended claims.
Claims
1. A method for detecting fake faces, applied to a fake face detection model including a hybrid expert feature extractor, characterized in that, The method includes: Obtain the face image to be detected; The face image to be detected is input into the hybrid expert feature extractor to obtain expert-extracted features; Calculate the similarity between the expert-extracted features and k learnable agent vectors in each of the multiple preset modes; For each preset mode, calculate the average similarity between the expert-extracted features and k learnable agent vectors in that preset mode; Determine the maximum average value among the average values of multiple similarities corresponding to the multiple preset patterns; The category label of the preset mode corresponding to the maximum average value is determined as the authenticity detection result of the face image, wherein the category label is real or fake; The hybrid expert feature extractor comprises multiple expert sub-networks; The fake face detection model was trained in the following way: Obtain the current batch of samples, wherein the current batch of samples contains multiple face image samples, each face image sample has a corresponding pattern label, the pattern label is used to indicate whether the corresponding face image sample is a fake face sample and the corresponding fake pattern, and the multiple preset patterns are the patterns indicated by the pattern labels corresponding to the multiple face image samples. Each face image sample is input into the hybrid expert feature extractor to obtain the features extracted by the trained experts and the weights assigned to each of the multiple expert sub-networks. ; The weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples. Calculate the batch-level regularization term loss. ; Based on the features extracted by the training expert and the corresponding pattern labels for each face image sample from the multiple face image samples, the cross-entropy loss is calculated. ; Based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable agent vectors in each of the multiple preset modes, the agent optimization loss is calculated. and feature reconstruction loss ; Based on the batch regularization term loss The cross-entropy loss The aforementioned agent optimization loss and the feature reconstruction loss The parameters of the fake face detection model are adjusted for training.
2. The method as described in claim 1, characterized in that, The agent optimization loss is calculated based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable agent vectors in each of the multiple preset modes. ,include: For each face image sample among the multiple face image samples, calculate the first similarity between the training expert extracted features corresponding to the face image sample and the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the face image sample. Calculate the second similarity between the expert-extracted features corresponding to the face image sample and k learnable proxy vectors in other preset modes, wherein the classification labels of the other preset modes are different from the classification labels of the mode labels corresponding to the face image sample. Based on the first similarity and the second similarity corresponding to each face image sample in the multiple face image samples, the proxy optimization loss is calculated. .
3. The method as described in claim 1, characterized in that, The feature reconstruction loss is calculated based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable proxy vectors in each of the multiple preset modes. ,include: For each face image sample among the multiple face image samples, the masked features are calculated based on the features extracted by the trained expert corresponding to that face image sample. ; Based on the masked features And the k learnable proxy vectors in the preset pattern indicated by the pattern label corresponding to the face image sample, to calculate the reconstructed features. ; Based on the features extracted and reconstructed by trained experts for each face image sample from the multiple face image samples, features are obtained. Calculate the feature reconstruction loss .
4. The method as described in claim 1, characterized in that, Also includes: For the fake training expert extraction features whose pattern label is a fake label among the multiple training expert extraction features corresponding to the multiple face image samples, the spherical linear interpolation method is used to expand the fake training expert extraction features to obtain slerp features, wherein the slerp features correspond to slerp pattern labels. The agent optimization loss is calculated based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable agent vectors in each of the multiple preset modes. and feature reconstruction loss ,include: Based on the expert-extracted features corresponding to each face image sample in the multiple face image samples, the Slerp features, the k learnable proxy vectors in each preset mode of the multiple preset modes, and the k learnable proxy vectors in the mode indicated by the Slerp mode label, the proxy optimization loss is calculated. and the feature reconstruction loss .
5. The method as described in claim 1, characterized in that, The weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples. Calculate the batch-level regularization term loss. ,include: The batch-level regularization term loss is calculated using the following formula. : in, The current batch of samples contains multiple face image samples. This refers to the m-th face image sample among the multiple face image samples. for The weights assigned to each of the plurality of expert subnetworks. This represents the standard deviation calculation. This indicates the operation of averaging.
6. A device for detecting fake faces, characterized in that, include: The image acquisition module is configured to acquire images of the faces to be detected. The feature extraction module is configured to input the face image to be detected into the hybrid expert feature extractor contained in the fake face detection model to obtain expert-extracted features; The similarity calculation module is configured to calculate the similarity between the expert-extracted features and k learnable proxy vectors in each of the multiple preset modes; The average value calculation module is configured to calculate the average similarity between the expert-extracted features and k learnable agent vectors in the preset mode for each preset mode; The maximum average value determination module is configured to determine the maximum average value among the average values of multiple similarities corresponding to the multiple preset modes; The authenticity detection result determination module is configured to determine the category label of the preset mode corresponding to the maximum average value as the authenticity detection result of the face image, wherein the category label is real or fake; The hybrid expert feature extractor comprises multiple expert sub-networks; The fake face detection model was trained in the following way: Obtain the current batch of samples, wherein the current batch of samples contains multiple face image samples, each face image sample has a corresponding pattern label, the pattern label is used to indicate whether the corresponding face image sample is a fake face sample and the corresponding fake pattern, and the multiple preset patterns are the patterns indicated by the pattern labels corresponding to the multiple face image samples. Each face image sample is input into the hybrid expert feature extractor to obtain the features extracted by the trained experts and the weights assigned to each of the multiple expert sub-networks. ; The weights assigned to each expert subnetwork based on each face image sample from the multiple face image samples. Calculate the batch-level regularization term loss. ; Based on the features extracted by the training expert and the corresponding pattern labels for each face image sample from the multiple face image samples, the cross-entropy loss is calculated. ; Based on the expert-extracted features corresponding to each face image sample in the multiple face image samples and the k learnable agent vectors in each of the multiple preset modes, the agent optimization loss is calculated. and feature reconstruction loss ; Based on the batch regularization term loss The cross-entropy loss The aforementioned agent optimization loss and the feature reconstruction loss The parameters of the fake face detection model are adjusted for training.
7. An electronic device, characterized in that, include: processor; Memory used to store the processor's executable instructions; The processor is configured to execute the instructions to implement the fake face detection method as described in any one of claims 1 to 5.
8. A computer-readable storage medium, characterized in that, When the instructions in the computer-readable storage medium are executed by the processor of the electronic device, the electronic device is able to perform the fake face detection method as described in any one of claims 1 to 5.
9. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by the processor, it implements the fake face detection method as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Face recognition method and system for dense crowd
CN118918628A
Face forgery detection method and system based on reconstruction learning and hybrid expert mode
CN119625812A