A domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP

By adopting the cross-modal fusion guide CLIP method in face anti-counterfeiting technology, the content description is introduced into prompts and visual features, which solves the problem of performance degradation in cross-data set experiments, and improves the system's security and domain generalization performance.

CN119339447BActive Publication Date: 2025-05-16INST OF AUTOMATION CHINESE ACAD OF SCI
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411423929.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-10-12
Publication Date
2025-05-16
Estimated Expiration
2044-10-12

AI Technical Summary

Technical Problem

The existing face anti-counterfeiting technology has deteriorated performance in cross-dataset experiments, and the domain generalization-based method leads to distortion of semantic feature structures, which has poor results.

Method used

The cross-modal fusion guided CLIP (CMFG-CLIP) method is used to introduce the content description of the sample into the prompts, and visual features are guided through fusion to enhance the model's understanding of content attributes and improve the generalization of visual features.

Benefits of technology

It improves the security of the face recognition system, effectively fights malicious physical media attacks, improves domain generalization performance, and avoids distortion of semantic feature structures.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119339447B_ABST
    Figure CN119339447B_ABST
Patent Text Reader

Abstract

The present invention discloses a domain generalized face anti-counterfeiting method guided by CLIP through cross-modal fusion. The method comprises inputting an image sample into a preprocessing unit for preprocessing, including face detection and data enhancement of the detected face image; generating text features including category description and content description for each text sample based on a fine-grained prompt template unit, generating visual features for the preprocessed image sample based on a domain-aware adapter unit, fusing the text features with the visual features to form a cross-modal joint feature representation, and training a face anti-counterfeiting prediction network using the joint feature representation; inputting a real face image processed by the preprocessing unit into a trained face anti-counterfeiting prediction network for inference prediction. The present invention utilizes text features to dynamically adjust visual features, rather than directly editing visual features, thereby enhancing the model's understanding and association of content and category semantics, and enhancing the domain generalization capability of face anti-counterfeiting technology.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the field of computer vision-human face liveness detection, and in particular relates to a domain generalized human face anti-counterfeiting method guided by cross-modal fusion CLIP. Background Art

[0002] The essence of face liveness detection is a binary classification task, that is, judging whether a given face image is from a real person or a forged face. Existing face anti-counterfeiting technologies (FAS) have achieved significant performance in experiments with data sets where the training and test data are from the same domain, but their performance has seriously declined in cross-dataset experiments due to the large distribution differences between different domains. Face anti-counterfeiting based on domain generalization aims to alleviate the impact of distribution differences by accessing multiple source domains. In the past, two strategies were most commonly used. One is to rely on domain labels to learn a domain-invariant feature space through adversarial training, and this feature space is also generalized to unseen domains. The other is based on examples, generating generalized features from features unrelated to liveness through decoupled representation learning or metric learning. However, face anti-counterfeiting based on these two strategies inevitably leads to the distortion of semantic feature structures, and the effect of domain generalization is not good.

[0003] In recent years, CLIP (Contrastive Language-Image Pretraining) based methods use category text as the weight of the classifier to adjust the visual features. However, relying solely on category-level cue engineering cannot effectively distinguish different types of attacks. Summary of the invention

[0004] To solve the above technical problems, the present invention proposes a domain generalized face anti-counterfeiting method called Cross-ModalFusion Guided CLIP (CMFG-CLIP), which introduces the content description of the sample into the prompt and guides the visual features in a fusion manner, aiming to improve the generalization of visual features by enhancing the model's understanding of content attributes.

[0005] To achieve the above purpose, the technical solution adopted by the present invention is as follows:

[0006] A domain generalized face anti-counterfeiting method guided by cross-modal fusion CLIP, the method comprising:

[0007] Inputting the image sample into the preprocessing unit for preprocessing, including face detection and data enhancement of the detected face image;

[0008] Generate text features including category description and content description for each text sample based on the fine-grained hint template unit, generate visual features for the preprocessed image sample based on the image encoder, fuse the text features with the visual features to form a cross-modal joint feature representation, and use the joint feature representation to train the face anti-counterfeiting prediction network;

[0009] The real face image processed by the preprocessing unit is input into the trained face anti-counterfeiting prediction network for inference prediction.

[0010] The beneficial effects of the present invention are:

[0011] The present invention solves the problem of using text descriptions without category semantics in the FAS task in large visual-text models. At the same time, it builds a bridge between the pre-trained CLIP and the FAS task based on domain generalization (DG), thereby improving the security of the face recognition system and effectively countering malicious physical media attacks. BRIEF DESCRIPTION OF THE DRAWINGS

[0012] Figure 1 A flowchart of the steps performed by the pre-processing unit of the present invention;

[0013] Figure 2 A network framework diagram of a domain generalized face anti-counterfeiting method guided by CLIP through cross-modal fusion according to the present invention;

[0014] Figure 3 A flowchart of the steps to be executed for the training unit;

[0015] Figure 4 Schematic diagram of the steps to perform for the prediction unit. DETAILED DESCRIPTION

[0016] The method of the present invention will be further described below in conjunction with the accompanying drawings.

[0017] The present invention provides a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP, the method comprises preprocessing, fine-grained prompt template (Fine-Grained Prompt Template, FGPT), domain-aware adapter (Domain-Aware Adapter, DA-Adapter), training and prediction to perform corresponding steps:

[0018] Preprocessing, which performs face detection on the input image, crops the face image into a fixed size, and performs data augmentation with random size cropping and horizontal flipping;

[0019] Fine-grained prompt template (FGPT), located in the text branch, consists of two language concepts: content and category. It encourages the classifier to understand visual features from text semantics and directly adjust visual features through multimodal fusion to achieve generalization. FGPT uses a fixed template to generate prompts, introduces the content description of each sample into the text branch for the first time, and improves the generalization ability of visual features by enhancing the model's understanding and association of content and category semantics.

[0020] The Domain-Aware Adapter (DA-Adapter), which is inserted into the image encoder and updated together with the downstream task, serves as a parameter-efficient transfer learning method to adapt the pre-trained CLIP to the FAS task. It contains an attention-based pooling layer that adaptively focuses on domain-invariant semantics to adapt to the frozen visual features.

[0021] The training process mainly trains the face anti-counterfeiting prediction network to enable it to have the function of predicting the authenticity of faces;

[0022] The prediction process inputs the preprocessed face image into the trained network to obtain the predicted value of the face authenticity.

[0023] Specifically, the present invention provides a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP, the method comprising:

[0024] Inputting the image sample into the preprocessing unit for preprocessing, including face detection and data enhancement of the detected face image;

[0025] Generate text features including category description and content description for each text sample based on the fine-grained hint template unit, generate visual features for the preprocessed image sample based on the image encoder, fuse the text features with the visual features to form a cross-modal joint feature representation, and use the joint feature representation to train the face anti-counterfeiting prediction network;

[0026] The real face image processed by the preprocessing unit is input into the trained face anti-counterfeiting prediction network for inference prediction.

[0027] like Figure 1 As shown, the preprocessing includes the following steps:

[0028] Detect whether the image sample contains a face. If not, abandon the image sample, otherwise proceed to the next step.

[0029] After adjusting the image samples where faces are detected to a fixed size, random size cropping and horizontal flipping are performed as data augmentation operations to output the preprocessed images.

[0030] like Figure 2 and Figure 3As shown, the fine-grained prompt template unit is implemented in the following steps:

[0031] Given a batch of B text samples , generate a category description in a fixed template format for each text sample , generate a content description for each text sample through MiniGPT-4 ;

[0032] Describe the category and content description , which are converted into category word embedding vectors through the word segmenter and embedding layer respectively and content word embedding vector ,in , Represents the dimension of the category word embedding vector, , Represents the dimension of the content word embedding vector;

[0033] ,

[0034] in, represents the word embedding vector after being processed by the word segmenter and the embedding layer, Represents the embedding layer function, which is used to convert the segmented text into a vector representation. Represents a tokenizer function, which is used to split text into basic units. Indicates category description or content description ,in The value is , Corresponding category, Corresponding content;

[0035] Embedding the category words into vectors and content word embedding vector Through the text encoder , get the category text features and content text features ,in , The dimension representing the categorical text features, Dimensions representing textual features of content.

[0036] The image encoder includes a parallel Transformer module and a domain-aware adapter. The domain-aware adapter includes three convolutional layers and an attention pooling layer, expressed as:

[0037] ,

[0038] ,

[0039] in Represents the visual features of the image sample before passing through the Transformer module, Representation layer normalization, For domain-aware adapters, Represents the adaptive visual features output by the domain-aware adapter; MHSA represents the multi-head self-attention mechanism, which is used to process information in parallel on multiple subspaces; MLP represents a multi-layer perceptron, which consists of multiple fully connected layers. Represents the image encoder The final visual features obtained are Represents the dimension of visual features.

[0040] The adaptive visual features are obtained by the following steps:

[0041] After flattening the image sample input to the image encoder into a 1D token sequence, the first patch token is obtained ,in represents the number of patch tokens, d represents the embedding dimension of the image sample, passing through the first convolutional layer of the domain-aware adapter. First Patch Token The embedding dimension changes from Reduce to a new feature dimension , get the second patch token ;

[0042] Restore Second Patch Token 2D structure ,in Represents the side length of one dimension in a 2D structure, The second patch token representing the restored 2D structure The size of , the second patch token Through the second convolutional layer Get patch embed

[0043] Based on the attention pooling layer Embed the patch Summarizes the content representation of the entire image , the content indicates and patch embedding Connect along the embedding dimension and pass through the third convolutional layer , and change the feature dimension from Adjust back , and finally generate adaptive visual features :

[0044] ,

[0045] ,

[0046] ,

[0047] ,

[0048] In the formula, Represents a concatenation operation along the embedding dimension.

[0049] The step of fusing text features with visual features to form a cross-modal joint feature representation includes:

[0050] The final visual features and categorical text features , content text features Linearly map to the joint embedding space with embedding dimension d, and transform the final visual features along the embedding dimension d and categorical text features and content text features Perform fusion operation to obtain fusion features :

[0051] ,

[0052] In the formula, represents the feature concatenation operation along the batch dimension;

[0053] The fusion features Input into two types of linear classifiers to predict the probability of being alive or fake:

[0054] ,

[0055] in, represents the classification loss, represents the cross entropy loss, represents a fully connected layer followed by softmax, Labels indicating live or fake faces;

[0056] The final visual features and categorical text features Concatenate according to the embedding dimension d to obtain positive image-text feature pairs ; Use the hard negative sampling strategy from ALBEF to create negative image feature pairs , negative text feature pairs ,in represents negative samples, Indicates the batch size:

[0057] ,

[0058] All positive image-text feature pairs and negative image feature pairs , negative text feature pairs Connect them to get joint features based on batch and embedding dimensions ;

[0059] ,

[0060] The face anti-counterfeiting prediction network is optimized by predicting the matching and mismatching probabilities of the joint feature P:

[0061] ,

[0062] in, is the image-text matching loss function, represents a fully connected layer followed by softmax, Represents matched and unmatched image-text pairs.

[0063] The attention-based pooling layer Embed the patch Summarizes the content representation of the entire image ,include:

[0064] Computing patch embeddings The mean and standard deviation of and ;

[0065] Obtain normalized content features through normalization , and through learnable query parameters Mapping to Query ;

[0066] Patch embedding at each spatial location Respectively through the learnable key parameters and learnable value parameters , mapped to the key Sum ;

[0067] Calculation query With all keys The dot product of , used to calculate the attention score, expressed as:

[0068] ,

[0069] ,

[0070] ,

[0071] ,

[0072] In the formula, Represents the scaling factor, which is used to prevent the gradient from disappearing or exploding due to excessive dot product. represents the new feature dimension, represents the number of heads in the multi-head attention mechanism, The function is used to convert the attention score into a probability distribution. Represents another learnable parameter matrix that is used to sum the weighted value vector Mapping back to the original feature space, represents the final domain-invariant content representation and serves as the output of the attention-based pooling layer, i.e., the content representation of the entire image: Represents the value vector after weighted aggregation through the attention mechanism.

[0073] The training of the face anti-counterfeiting prediction network comprises the following steps:

[0074] The fusion features are trained using classification loss, and the predicted joint features are trained using matching loss. The classification loss is used to distinguish between real faces and fake faces, and the matching loss is used to optimize the matching degree of image-text pairs.

[0075] The total training loss is the sum of the classification loss and the matching loss. It is used to determine whether the total training loss converges. If the convergence condition is met, the training is terminated; otherwise, the next round of training is continued.

[0076] The Adam optimizer is used to update the face anti-counterfeiting prediction network parameters according to the calculated gradient to obtain a trained face anti-counterfeiting prediction network.

[0077] like Figure 4 As shown, the prediction process is used to perform the following steps:

[0078] Inputting the real face image to be predicted into the preprocessing unit for operation;

[0079] The text content and preprocessed real face images are input into the trained face anti-counterfeiting prediction network (CMFG-CLIP) to extract visual features and text features, and perform feature fusion to form a cross-modal joint feature representation, which is input into the classifier to perform classification prediction of real faces and forged faces.

[0080] The specific embodiments described above further illustrate the objectives, technical solutions and beneficial effects of the present invention in detail. It should be understood that the above description is only a specific embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A domain generalized face anti-counterfeiting method guided by cross-modal fusion CLIP, characterized in that: The method comprises: Inputting the image sample into the preprocessing unit for preprocessing, including face detection and data enhancement of the detected face image; Based on the fine-grained hint template unit, a text feature including a category description and a content description is generated for each text sample, and based on the image encoder, a visual feature is generated for the preprocessed image sample, and the text feature is fused with the visual feature to form a cross-modal joint feature representation, and the joint feature representation is used to train the face anti-counterfeiting prediction network; the fine-grained hint template unit is implemented according to the following steps: Given a batch of B text samples , generate a category description in a fixed template format for each text sample , generate a content description for each text sample through MiniGPT-4 ; Describe the category and content description , which are converted into category word embedding vectors through the word segmenter and embedding layer respectively and content word embedding vector ,in , represents the dimension of the category word embedding vector, , Represents the dimension of the content word embedding vector; , in, represents the word embedding vector after being processed by the word segmenter and the embedding layer, Represents the embedding layer function, which is used to convert the segmented text into a vector representation. Represents a tokenizer function, which is used to split text into basic units. Indicates category description or content description ,in The value is , Corresponding category, Corresponding content; Embedding the category words into vectors and content word embedding vector Through the text encoder , get the category text features and content text features ,in , The dimension representing the categorical text features, Dimensions representing the textual characteristics of the content; The real face image processed by the preprocessing unit is input into the trained face anti-counterfeiting prediction network for inference prediction.

2. According to claim 1, a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP is characterized in that: The pre-processing comprises the following steps: Detect whether the image sample contains a face. If not, abandon the image sample, otherwise proceed to the next step. After adjusting the image samples where faces are detected to a fixed size, random size cropping and horizontal flipping are performed as data augmentation operations to output the preprocessed images.

3. According to claim 1, a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP is characterized in that: The image encoder includes a parallel Transformer module and a domain-aware adapter, which includes three convolutional layers and an attention pooling layer, expressed as: , , in Represents the visual features of the image sample before passing through the Transformer module, Representation layer normalization, For domain-aware adapters, Representing the adaptive visual features output by the domain-aware adapter; MHSA stands for multi-head self-attention mechanism, which is used to process information in parallel on multiple subspaces. MLP stands for multi-layer perceptron, which consists of multiple fully connected layers. Represents the image encoder The final visual features obtained are .

4. According to claim 3, a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP is characterized in that: The adaptive visual features are obtained by the following steps: After flattening the image sample input to the image encoder into a 1D token sequence, the first patch token is obtained ,in represents the number of patch tokens, d represents the embedding dimension of the image sample, passing through the first convolutional layer of the domain-aware adapter. First Patch Token The embedding dimension changes from Reduce to a new feature dimension , get the second patch token ; Restore Second Patch Token 2D structure ,in Represents the side length of one dimension in a 2D structure, The second patch token representing the restored 2D structure The size of , the second patch token Through the second convolutional layer Get patch embed Based on the attention pooling layer Embed the patch Summarizes the content representation of the entire image , the content indicates and patch embedding Connect along the embedding dimension and pass through the third convolutional layer , and change the feature dimension from Adjust back , and finally generate adaptive visual features : , , , , In the formula, Represents a concatenation operation along the embedding dimension.

5. According to claim 1, a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP is characterized in that: The step of fusing text features with visual features to form a cross-modal joint feature representation includes: The final visual features and categorical text features , content text features Linearly map to a joint embedding space with embedding dimension d The final visual features are transformed along the embedding dimension d and categorical text features and content text features Perform fusion operation to obtain fusion features : , In the formula, represents the feature concatenation operation along the batch dimension; The fusion features Input into two types of linear classifiers to predict the probability of being alive or fake: , in, represents the classification loss, represents the cross entropy loss, represents a fully connected layer followed by softmax, Labels indicating live or fake faces; The final visual features and categorical text features Concatenate according to the embedding dimension d to obtain positive image-text feature pairs ; Use the hard negative sampling strategy from ALBEF to create negative image feature pairs , negative text feature pairs ,in represents negative samples, Indicates the batch size: , All positive image-text feature pairs and negative image feature pairs , negative text feature pairs Connect them to get joint features based on batch and embedding dimensions ; , The face anti-counterfeiting prediction network is optimized by predicting the matching and mismatching probabilities of the joint feature P: , in, is the image-text matching loss function, represents a fully connected layer followed by softmax, Represents matched and unmatched image-text pairs.

6. According to claim 4, a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP is characterized in that: The attention-based pooling layer Embed the patch Summarizes the content representation of the entire image ,include: Computing patch embeddings The mean and standard deviation of and ; Obtain normalized content features through normalization , and through learnable query parameters Mapping to Query ; Patch embedding at each spatial location Respectively through the learnable key parameters and learnable value parameters , mapped to the key Sum ; Calculation query With all keys The dot product of , used to calculate the attention score, expressed as: , , , , In the formula, Represents the scaling factor, which is used to prevent the gradient from disappearing or exploding due to excessive dot product. represents the new feature dimension, represents the number of heads in the multi-head attention mechanism, The function is used to convert the attention score into a probability distribution. Represents another learnable parameter matrix that is used to sum the weighted value vector Mapping back to the original feature space, represents the final domain-invariant content representation and serves as the output of the attention-based pooling layer, i.e., the content representation of the entire image: Represents the value vector after weighted aggregation through the attention mechanism.

7. According to claim 1, a domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP is characterized in that: The training of the face anti-counterfeiting prediction network comprises the following steps: The fusion features are trained using classification loss, and the predicted joint features are trained using matching loss. The classification loss is used to distinguish between real faces and fake faces, and the matching loss is used to optimize the matching degree of image-text pairs. The total training loss is the sum of the classification loss and the matching loss. It is used to determine whether the total training loss converges. If the convergence condition is met, the training is terminated; otherwise, the next round of training is continued. The Adam optimizer is used to update the face anti-counterfeiting prediction network parameters according to the calculated gradient to obtain a trained face anti-counterfeiting prediction network.

8. The domain generalization face anti-counterfeiting method guided by cross-modal fusion CLIP according to claim 1, characterized in that: The prediction process is used to perform the following steps: Inputting the real face image to be predicted into the preprocessing unit for operation; The text content and preprocessed real face images are input into the trained face anti-counterfeiting prediction network to extract visual features and text features, and perform feature fusion to form a cross-modal joint feature representation, and perform classification prediction of real faces and forged faces.

Citation Information

Patent Citations

  • Human face in-vivo detection model training method, human face in-vivo detection method and human face in-vivo detection device

    CN115761839A