Model training method, face authentic identification method, device, equipment, medium and product

By generating strong facial feature maps and combining them with text description templates, a training set of image-text pairs is constructed to train a multimodal anti-spoofing model, which solves the problem of poor face anti-spoofing performance and achieves effective recognition and identification of highly realistic fake faces.

CN121746840APending Publication Date: 2026-03-27CHINA MOBILE FINANCIAL TECHNOLOGY CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-10
Publication Date
2026-03-27

AI Technical Summary

Technical Problem

In existing technologies, the face authentication performance is poor, especially since small deep learning models are prone to overfitting to specific forgery methods, resulting in weak generalization and cross-domain capabilities, making it difficult to effectively identify highly realistic forged face images.

Method used

By utilizing a feature extraction network to generate strong facial feature maps and combining them with preset text description templates, an image-text pair training set is constructed to train a multimodal anti-spoofing model, including an image encoder, a text encoder, and an image-text decoder. This enables the automated generation of image-text descriptions for forged images and increases the scale of the learning data.

Benefits of technology

It improves the effectiveness of face authentication, enabling the identification of forgery clues that are not visible to the human eye. It constructs a large-scale image and text training dataset, enhances the learning ability of the multimodal authentication model, and improves the accuracy of authentication.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121746840A_ABST
    Figure CN121746840A_ABST
Patent Text Reader

Abstract

The invention provides a model training method, a face authentic identification method and device, equipment, a medium and a product, relates to the technical field of artificial intelligence, and aims to solve the problem that the face authentic identification effect is poor. The method comprises the following steps: performing feature recognition processing on a face image sample by using a feature extraction network to obtain a face strong feature mapping graph; on the basis of the face strong feature mapping graph and a preset text description template, an image-text pair training set is generated, the image-text pair training set comprises multiple pieces of training data, and the training data comprises the face strong feature mapping graph and corresponding text description information; a target training set is utilized to train a multi-mode authentic identification large model, a trained multi-mode authentic identification large model is obtained, the target training set comprises the image-text pair training set, and the multi-mode authentic identification large model comprises an image encoder, a text encoder and an image-text decoder. According to the invention, the face authentic identification effect can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the field of artificial intelligence technology, specifically relating to a model training method, a face authentication method, a device, equipment, medium, and product. Background Technology

[0002] The use of artificial intelligence (AI) technology to forge customer facial images, videos, and audio information to impersonate others and commit telecommunications fraud and illegal financial arbitrage poses significant security risks to mobile internet, public safety, and cybersecurity regulation, and is a current security challenge for facial recognition. Among related technologies, AI-generated content (AIGC) facial forgery detection technology primarily employs small deep learning models. However, these small deep learning models have relatively small network structures and parameter sizes, resulting in limited learning capacity and a small dataset. Consequently, they are prone to overfitting to specific forgery methods, leading to poor facial authentication performance. Summary of the Invention

[0003] This application provides a model training method, a face authentication method, an apparatus, a device, a medium, and a product to address the problem of poor face authentication performance.

[0004] To solve the above-mentioned technical problems, this application is implemented as follows: In a first aspect, embodiments of this application provide a model training method, including: A feature extraction network is used to perform feature recognition processing on face image samples to obtain a strong feature map of the face. Based on the strong facial feature map and the preset text description template, a training set of image-text pairs is generated. The training set of image-text pairs includes multiple training data, which includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. The multimodal anti-spoofing model is trained using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

[0005] Secondly, embodiments of this application provide a face authentication method, including: The face image to be identified is input into the trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result; The trained multimodal anti-spoofing model is the multimodal anti-spoofing model obtained by the model training method described in the first aspect.

[0006] Thirdly, embodiments of this application provide a model training apparatus, comprising: The processing module is used to perform feature recognition processing on face image samples using a feature extraction network to obtain a strong feature map of the face; The generation module is used to generate a training set of image-text pairs based on the strong facial feature map and a preset text description template. The training set of image-text pairs includes multiple training data, and the training data includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. The training module is used to train the multimodal anti-spoofing model using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

[0007] Fourthly, embodiments of this application provide a face authentication device, comprising: The inference module is used to input the face image to be identified into the trained multimodal anti-spoofing model for inference and obtain the anti-spoofing result; The trained multimodal anti-spoofing model is the multimodal anti-spoofing model obtained by the model training method described in the first aspect.

[0008] Fifthly, embodiments of this application provide an electronic device, including a transceiver and a processor, wherein the processor is used for: A feature extraction network is used to perform feature recognition processing on face image samples to obtain a strong feature map of the face. Based on the strong facial feature map and the preset text description template, a training set of image-text pairs is generated. The training set of image-text pairs includes multiple training data, which includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. The multimodal anti-spoofing model is trained using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

[0009] Sixthly, embodiments of this application provide an electronic device, including a transceiver and a processor, wherein the processor is used for: The face image to be identified is input into the trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result; The trained multimodal anti-spoofing model is the multimodal anti-spoofing model obtained by the model training method described in the first aspect.

[0010] In a seventh aspect, embodiments of this application provide an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor, wherein when the program is executed by the processor, it implements the steps of the model training method as described in the first aspect above; or, when the program is executed by the processor, it implements the steps of the face authentication method as described in the second aspect above.

[0011] Eighthly, embodiments of this application provide a computer-readable storage medium storing a computer program, wherein when executed by a processor, the computer program implements the steps of the model training method as described in the first aspect above; or, when executed by a processor, the computer program implements the steps of the face authentication method as described in the second aspect above.

[0012] Ninthly, embodiments of this application provide a computer program product, including computer instructions, which, when executed by a processor, implement the steps of the model training method as described in the first aspect above; or, when executed by a processor, the computer instructions implement the steps of the face authentication method as described in the second aspect above.

[0013] In this embodiment, a feature extraction network is used to perform feature recognition processing on face image samples to obtain a strong face feature map. Based on the strong face feature map and a preset text description template, an image-text pair training set is generated. The image-text pair training set includes multiple training data, which includes the strong face feature map and corresponding text description information. The text description information corresponding to different training data is different. A multimodal anti-spoofing model is trained using a target training set to obtain a trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features. In this way, by using a feature extraction network to perform feature recognition processing on face image samples, a strong face feature map is obtained. Based on the strong face feature map and a preset text description template, a training set of image-text pairs is generated, which enables the automatic generation of image-text descriptions of forged images. Furthermore, forged clue images that are not visible to the human eye can also generate text descriptions of forged face regions. Constructing a large-scale image-text training dataset can increase the scale of the learning data for the multimodal forgery detection model, thereby improving the effect of face forgery detection. Attached Figure Description

[0014] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments of this application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0015] Figure 1 This is a flowchart of a model training method provided in an embodiment of this application; Figure 2 This is a schematic diagram of a training multimodal anti-spoofing model provided in an embodiment of this application; Figure 3 This is a schematic diagram of generating a strong facial feature mapping map provided in an embodiment of this application; Figure 4 This is a schematic diagram illustrating a method for generating a training set of image-text pairs according to an embodiment of this application; Figure 5 This is an application flowchart of a face authentication method provided in an embodiment of this application; Figure 6 This is a flowchart of a face authentication method provided in an embodiment of this application; Figure 7 This is a schematic diagram of a multimodal counterfeit detection model inference provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application; Figure 9 This is a schematic diagram of the structure of a face authentication device provided in an embodiment of this application; Figure 10 This is one of the structural schematic diagrams of an electronic device provided in the embodiments of this application; Figure 11 This is a second schematic diagram of the structure of an electronic device provided in the embodiments of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0017] For ease of understanding, the following describes some aspects of the embodiments of this application: Criminals steal personal facial images or videos and then use digital facial generation to create fake faces, impersonating customers and verifying their identities through facial recognition. Such individual fraud or arbitrage activities pose a significant risk to the lives and property of the public. In particular, the popularity of large-scale model technology has further lowered the threshold for various AIGC facial forgery detection technologies, leading to frequent cases of successful cracking. There is an urgent need to accelerate the development of AIGC facial forgery recognition technology using new technologies and methods to solve the increasingly serious problem of reliable facial authentication.

[0018] In related technologies, image-level forgery detection methods mainly target single frames of images or videos to be inspected. Based on different detection principles, this technology can be divided into spatial domain-based and frequency domain-based detection methods. These methods utilize various variants of deep learning convolutional neural networks to extract artifact features from the spatial or frequency domains of the image, such as texture features, statistical features, frequency features, and multi-scale features. These features are then combined with a classification network to achieve AIGC face forgery detection. However, these methods have limited generalization capabilities, and as fake faces become increasingly realistic, their robustness is insufficient, easily missing deeply faked images or videos. Furthermore, related detection technologies not only explore forgery detection features in images and videos but also combine multimodal information such as frequency and audio to identify forgery clues in images, but these methods still do not effectively solve the problem.

[0019] Currently, the performance of AIGC face spoofing and face-swapping recognition technology is not ideal. Especially with the prevalence of large-scale model technology, the technical threshold for digital face spoofing (AIGC) is decreasing, leading to a proliferation of spoofing techniques and methods. Regardless of the base spoofing model or pre- and post-processing, any alteration to the generated deepfake images results in significant missed detections. This is primarily because related AIGC face spoofing detection technologies mainly rely on small deep learning models with relatively small network structures and parameter sizes, resulting in limited learning capacity and a small learning dataset. Consequently, they are prone to overfitting to specific spoofing methods. Correspondingly, there are numerous small recognition models or methods for different types or characteristics of AI face-swapping images. In summary, the main drawbacks of these small recognition model technologies are poor generalization and weak cross-domain capabilities; when faced with spoofing methods not encountered during training, the model's detection accuracy drops rapidly.

[0020] Therefore, in recent years, AIGC face forgery detection technology based on multimodal large models has gradually emerged. This technology mainly leverages the general image semantic understanding capabilities of multimodal large models, and enhances the recognition ability of AIGC face forgery detection through incremental training on vertical domain AIGC face forgery data. In implementation, this type of method typically involves manually describing AIGC face forgery clues using various descriptive methods as text descriptions, along with video data as input to the multimodal large model. The incrementally trained large model is then used to detect AIGC face forgery videos. The advantages of this method are its large model parameter scale and the general semantic advantage of multimodality, which improves the model's accuracy to some extent. However, this method also has limitations. For forgery clues that are discernible to the human eye, textual semantic descriptions of image forgery clues can be achieved through human annotation or semi-automatic techniques using large models. However, current forgery technology is developing rapidly, and the realism of images and videos is very high, making it difficult for the human eye to distinguish them. Therefore, manual or semi-automatic annotation cannot adequately describe forgery clues, and the aforementioned models are unsuitable for such situations, necessitating new large model technologies to address this need.

[0021] In this application, a model training method, a face authentication method, a device, an equipment, a medium, and a product are proposed to solve the problem of poor face authentication performance.

[0022] See Figure 1 , Figure 1 This is a flowchart of a model training method provided in an embodiment of this application, such as... Figure 1 As shown, the method includes the following steps: Step 101: Use a feature extraction network to perform feature recognition processing on the face image samples to obtain a strong feature map of the face.

[0023] In this step, the feature extraction network can be a deep learning network structure, such as a convolutional neural network (CNN, like ResNet and VGG) or a visual Transformer. Its function is to receive raw image pixel data as input, and through multiple layers of nonlinear transformations, abstract and output feature representations that can characterize the high-level semantics and discriminative patterns of the image. It can be understood that the feature extraction network is a pre-trained model with preliminary face authentication capabilities.

[0024] The aforementioned face image samples can be data used for model training, including real face images and fake face images generated by AIGC technology. The fake face images can include samples with high realism and forgery traces that are difficult for the human eye to directly identify.

[0025] The aforementioned feature recognition process can involve the feature extraction network performing forward propagation calculations on face image samples to extract and retain the spatial feature representations of the intermediate layers.

[0026] The aforementioned strong facial feature map can be a heatmap generated based on class activation mapping technology, corresponding to the spatial size of facial image samples. Each pixel value in the strong facial feature map can characterize the contribution or correlation strength of the corresponding image region to the model's decision of forgery or authenticity. High-scoring regions in the strong facial feature map are areas strongly correlated with forgery clues, visualizing the spatial distribution of subtle forgery traces that are difficult for the human eye to detect.

[0027] Step 102: Based on the strong facial feature map and the preset text description template, generate a training set of image-text pairs. The training set of image-text pairs includes multiple training data. The training data includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different.

[0028] In this step, the aforementioned preset text description template can be a series of predefined structured string formats, containing replaceable variable placeholders, such as {location} and {anomaly type}. This text description template can be used to convert visual positioning information into natural language descriptions, for example: "A trace of {anomaly type} was detected at {location}".

[0029] The training set for the above image-text pairs can be a dataset where each data point is a paired sample, consisting of an image and a text describing the content of that image. This is the standard data format for training multimodal large models.

[0030] The aforementioned text description information may be a specific text statement generated by filling the text description template with information from strongly correlated regions in the strong facial feature mapping map, and its content clearly indicates the semantic parts and / or trace nature of the forgery.

[0031] It is understandable that the different text description information corresponding to the different training data mentioned above is used to define a one-to-many data construction strategy. That is, for the same original face image sample, multiple different text descriptions can be generated in one or more of the following ways: (a) using different description templates; (b) focusing on different high-scoring regions in the strong feature map; (c) using different description granularities or perspectives for the same region.

[0032] Step 103: Train the multimodal anti-spoofing model using the target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

[0033] In this step, the target training set mentioned above can be the complete dataset ultimately used to train the multimodal anti-counterfeiting model, including the aforementioned automatically generated image-text pair training set.

[0034] The aforementioned image encoder can be an architecture such as a visual Transformer, used to encode face image samples into a high-dimensional visual feature vector, which condenses the global and local visual information of the image.

[0035] The aforementioned text encoder can be a pre-trained language model such as BERT, used to encode the input text description into a semantic feature vector with the same dimension as the visual feature vector, which condenses the semantic information of the text.

[0036] The aforementioned image-text decoder can be a module based on the Transformer decoder architecture. Its core is a cross-modal attention mechanism, which is used to receive feature vectors of images and text, and to enable deep interaction and alignment of the two features through attention calculation, ultimately outputting a fused multimodal joint feature representation.

[0037] For example, during training, the multimodal anti-spoofing model reads in an image-text pair (such as a fake image with the descriptive text "artificial synthetic traces exist on the side of the nose bridge"). The image encoder is used to extract image features, the text encoder is used to extract text features, and the cross-modal attention mechanism in the image-text decoder learns to establish the association between specific visual patterns of the side of the nose bridge in the image and concepts such as the side of the nose bridge and artificial synthetic traces in the text. By training on a massive number of such image-text pairs, the model learns to associate subtle visual anomalies with accurate semantic descriptions.

[0038] In this embodiment, a strong feature map of the face is obtained by using a feature extraction network to perform feature recognition processing on face image samples. Based on the strong feature map of the face and a preset text description template, a training set of image-text pairs is generated, which enables the automatic generation of image-text descriptions of forged images. Furthermore, forged clue images that are not visible to the human eye can also generate text descriptions of forged face regions. Constructing a large-scale image-text training dataset can increase the scale of the learning data for the multimodal forgery detection model, thereby improving the effect of face forgery detection.

[0039] For example, Figure 2 This is a schematic diagram of a training multimodal anti-spoofing model provided in an embodiment of this application, as shown below. Figure 2 As shown, the specific steps are as follows: Step 1: Collect two types of data: a domain-specific large dataset based on covert forgery descriptions and a dataset of forgery descriptions visible to the human eye, and clean and standardize the data.

[0040] Step 2: Generate multi-dimensional, formatted forgery description text for a single forged image, providing a multi-perspective description of the same forgery feature.

[0041] Step 3: The image encoder selects a model that supports fine-grained feature extraction, such as ViT, inputs the preprocessed fake image, extracts the visual features of the image, and outputs an image feature vector of dimension D, as shown in formula (1). The text encoder uses a pre-trained language model such as BERT to segment and tokenize the fake descriptive text, extracts global text semantic features, and outputs a text feature vector of the same dimension as the image encoder (D-dimensional), as shown in formula (2).

[0042]

[0043] Step 4: Use a text-image decoder, such as the decoder of a large language model, to parse the information in the text description. Cross-modal attention (CA) and feed-forward network (FFN) process image features and text features, enabling bidirectional interaction between image features and text features, and achieving text-image feature fusion, as shown in formula (3).

[0044]

[0045] Step 5: Through multiple rounds of iterative training, text labels are automatically generated iteratively. Through this process, the model parameters are continuously updated to enhance the accuracy of text description and model discrimination.

[0046] Optionally, the step of using a feature extraction network to perform feature recognition processing on face image samples to obtain a strong face feature map includes: Feature extraction networks are used to extract features from face image samples to obtain feature map clusters of the face image samples; The feature map clusters are processed by global average pooling and fully connected layers to obtain the weight vector of the forgery category; Based on the weight vector, the feature maps of each channel in the feature map cluster are weighted and summed to obtain a forgery clue distribution map; The forged clue distribution map is upsampled to the size of the face image sample to obtain the target distribution map; Based on the correlation metric score in the target distribution map, the target distribution map is mapped to the region with the highest correlation metric score in the face image sample to obtain a strong face feature mapping map.

[0047] Specifically, the aforementioned feature map cluster can be the output of the last convolutional layer in the feature extraction network. It is a three-dimensional tensor with dimensions [H, W, C]. Here, H and W are the height and width of the feature map, respectively, representing the spatial dimension; C is the number of channels, and each channel can be regarded as an independent "feature detector" used to respond to a specific local pattern or texture in the face image sample, such as edges, textures of specific frequencies, eye structures, etc.

[0048] The global average pooling described above can be a spatial pooling operation, which calculates the average value of all pixel values ​​for each channel of the feature map cluster, thereby compressing the two-dimensional spatial information of each channel into a scalar value.

[0049] The above fully connected layer processing can be performed based on the classification layer at the end of the feature extraction network.

[0050] The weight vector for the aforementioned forgery category can be the column of weight parameters in the fully connected layer that corresponds to the category label "forgery".

[0051] The aforementioned forgery clue distribution map can be a two-dimensional single-channel image obtained through the weighted summation calculation described above. The value of each pixel (x, y) on the forgery clue distribution map represents the comprehensive response of all feature channels in the corresponding downsampled region of the original image, weighted according to their importance. It can be understood that the higher the value of the forgery clue distribution map, the greater the probability that the spatial location is a forgery clue.

[0052] The aforementioned upsampling can be used to enlarge low-resolution images to high resolution, ensuring that the spatial dimensions of the forgery clue distribution map are consistent with the original face image sample, facilitating pixel-level alignment and mapping. The target distribution map described above can be a forgery clue distribution map with the same size as the original face image sample after upsampling.

[0053] The aforementioned correlation metric score can be the value of each pixel in the target distribution map, used to characterize the correlation strength between the pixel location and the forgery category.

[0054] The above mapping can be a process of segmenting high-score regions from the target distribution map according to preset rules, such as setting a score threshold, and highlighting these regions on the original face image samples.

[0055] The aforementioned strong facial feature mapping map can be a visualization result that overlays the most relevant regions of the target distribution map onto the original facial image sample in the form of a heat map, and can be used to clearly indicate the precise location of forgery traces.

[0056] For example, Figure 3 This is a schematic diagram of generating a strong facial feature map according to an embodiment of this application, such as... Figure 3 As shown, the steps for generating a strong facial feature map are as follows: Step 1: Input a face image and use a feature extraction network, such as the commonly used ResNet or VGG, to obtain the output feature map of the last convolutional layer—a cluster of feature maps with several channels, usually with dimensions [H, W, C], where H and W are the height and width of the feature map, and C is the number of channels.

[0057] Step 2: The feature map clusters of several channels are transformed into weight vectors with the shape [C, num_classes] after passing through the feature densification and class mapping weight network, such as global average pooling and classification layer (fully connected layer), where W[:, c] represents the weight vector of class c.

[0058] Step 3: Weighted summation, for each category c, for each channel feature map at different spatial locations Multiply by the corresponding weight Summing yields the distribution map of forged clues. The formula is shown below: ; Step 4: By upsampling the forged clue distribution map to the size of the input image and combining it with the relevance metric score of the forged features in the distribution map, the map is mapped to the region in the original face image that is most relevant to a specific category, thus obtaining a strong face feature map.

[0059] In this embodiment, the original feature extraction network can be designed as described above to achieve the desired result. Figure 1 The successful construction of a distribution map of forgery clues of equal size, combined with the correlation metric score of forgery features in the distribution map, maps and identifies strongly correlated areas of face forgery clues in the original image. This transforms hidden or traceless forgery clues in the original face image sample from invisible to the human eye into explicit regional clues, enabling the textual description of those invisible, highly realistic hidden forged face images. This solves the problem of not being able to achieve textual description of forgery clues in large-scale hidden or traceless forged face images.

[0060] Optionally, the preset text description template includes multiple sub-templates, and the step of generating a text-image pair training set based on the strong facial feature map and the preset text description template includes: The location information of the forged region is determined based on the strong facial feature mapping map. Multiple text descriptions are generated based on the multiple sub-templates and the location information, and the multiple text descriptions correspond one-to-one with the multiple sub-templates; The multiple text descriptions are combined one by one with the strong facial feature map to obtain a text-image pair training set, which includes at least one pair of the strong facial feature map and the text description.

[0061] Specifically, the aforementioned forged region can be a connected region formed by pixels whose relevance score exceeds a preset threshold in the face strong feature map. The aforementioned location information can be a digital description of the forged region, such as semantic location information and geometric location information. The semantic location information can be mapped to standardized facial part names, such as "left cheek," "nose wing," and "right eyebrow arch," by combining it with a face parsing model or a face keypoint model. The geometric location information can be the coordinates of the region's circumscribed rectangle, center point coordinates, or region mask.

[0062] The aforementioned sub-templates can be a set of predefined formatted strings with different sentence structures but similar semantics. Each sub-template can be a complete sentence skeleton containing variable placeholders (such as {location}) for embedding specific location information. For example, sub-template 1: "There are obvious traces of synthesis in {location}."; sub-template 2: "The image exhibits unnatural properties in the {location} region."; sub-template 3: "Tampering was detected in {location}."

[0063] The aforementioned text description information can be a complete natural language sentence formed by filling the location information into the corresponding placeholder of a sub-template.

[0064] The one-to-one correspondence mentioned above means that each sub-template will be combined with the location information independently once to generate a unique text description. Thus, N sub-templates can generate N different text descriptions for the same set of location information.

[0065] For example, Figure 4 This is a schematic diagram of a method for generating a training set of image-text pairs provided in an embodiment of this application, such as... Figure 4 As shown, this paper maps the strongly correlated regions of forged clues to the semantic descriptions of face regions, proposes an automated format for constructing descriptions of forged regions and forged clues, and realizes an automated text description generation technology for forged text. This allows for the construction of a large-scale training set of image-text pairs describing hidden forged clues. The specific steps are as follows: Step 1: Through the strong feature mapping map of the face, the high-scoring area is the face location area corresponding to category c, thus obtaining the localization information of this area.

[0066] Step 2: Predefine n formatted templates for describing forged clues, such as "{XX} is a significantly tampered area", "{XX} has obvious signs of tampering", "The area involved in the significant tampering includes {XX}", etc.

[0067] Step 3: Based on the location information of high-scoring regions of category c and the correlation between them and the target category, and combined with n predefined description templates, n text descriptions of the images are automatically generated, thereby realizing the automated generation of large-scale image-text pair training sets.

[0068] In this implementation, by mapping the strongly correlated regions of forgery clues in the distribution map to the original image and combining them with a formatted semantic description template, that is, taking the high-scoring regions in the distribution map of forgery clues as the core, combining their spatial location and correlation with forgery clues, and combining them with a predefined description template, an automated text and image description generation technology for forgery images is realized. Thus, text descriptions of forgery clue images that are not visible to the human eye can also be generated for face forgery regions.

[0069] Optionally, the step of training the multimodal anti-spoofing model using the image and text training set to obtain the trained multimodal anti-spoofing model includes: The multimodal anti-spoofing model is trained based on the objective loss function and the image-text pair training set to obtain the trained multimodal anti-spoofing model. The objective loss function is used to increase the Euclidean distance between fake samples and real samples in the image-text pair training set.

[0070] Specifically, the aforementioned objective loss function can be an objective function used to guide the optimization of parameters in a large multimodal anti-spoofing model during training.

[0071] It is understandable that the above training can be a process of iteratively adjusting the learnable parameters of the image encoder, text encoder, and image-text decoder in the multimodal anti-spoofing large model through backpropagation algorithm and gradient descent optimizer, so that the overall loss of the model on the image-text pair training set is continuously reduced.

[0072] For example, the target loss function L is shown in formula (4) below, which increases the distance between fake and real samples, for example, by measuring it using Euclidean distance. When the sample , When samples belong to the same class, minimize their distance, as shown in formula (5); when samples , When they belong to different categories, maximize the distance between them, as shown in formula (6), where the minimum distance is... The choice should be based on the characteristics of the dataset; you can start by trying numbers between 0.2 and 0.5.

[0073] (4) (5) (6) In this implementation, the multimodal anti-counterfeiting model is optimized by increasing the Euclidean distance between fake and real samples in the training set. This forces the model to learn more discriminative feature representations, effectively separating real and fake samples in the feature space. This significantly reduces the classification error rate and improves the accuracy of anti-counterfeiting.

[0074] Optionally, the target training set further includes a preset dataset, which includes multiple image samples and text descriptions corresponding to each of the multiple image samples, and the text descriptions corresponding to each of the multiple image samples are obtained by manual annotation.

[0075] Specifically, the aforementioned preset dataset can be a pre-built training dataset.

[0076] Understandably, the aforementioned pre-set dataset can systematically cover a variety of known and typical forgery methods and techniques, such as facial manipulation, face replacement, and face reproduction, supplemented by precise human descriptions.

[0077] In this implementation, the target training set also includes a preset dataset, which allows for the introduction of more accurately labeled manually labeled data. This helps to more reasonably define the core distribution of real samples and the distribution of known fake samples in the feature space, so that the decision boundary that the model can learn is more reasonable and generalized, thereby avoiding overfitting to certain specific biases that may exist in automated data.

[0078] For example, Figure 5 This is an application flowchart of a face authentication method provided in an embodiment of this application, such as... Figure 5 As shown, this paper first proposes to construct a distribution map of forgery clues, and then uses scores to quantify and identify the strongly related regions of forgery clues in the distribution map. Next, it maps the strongly related regions of forgery clues to the semantic descriptions of face regions, and proposes an automated format for constructing descriptions of forgery regions and forgery clues, thereby realizing an automated text description generation technology for forgery. This allows for the construction of a large-scale training set of image-text pairs describing hidden forgery clues.

[0079] Then, based on the domain-specific large dataset of covert forgery descriptions, along with the dataset of forgery descriptions visible to the human eye, the image-text pairs are used as input to the multimodal large model. The images are used as input to the image encoder, and the text descriptions are used as input to the text encoder. Both are decoded by the large language model. The image-text pairs adopt a one-to-many construction mode, that is, the same image may correspond to multiple automatically generated forgery description texts. After training with large-scale image-text pairs, a deep forgery detection large model is obtained.

[0080] In the inference phase, based on the deep anti-spoofing model proposed in this application, images are used as input to obtain textual semantic descriptions of the images. This proposal proposes a semantic parsing model to ultimately output a conclusion on whether the image is forged, along with a description of the forged region. Furthermore, a new loss function is proposed in the model to further increase the distance between forged and real samples. Finally, multi-round iterative training is proposed, with iterative automatic generation of text labels. Through this process, the accuracy of the text description and the model is continuously increased.

[0081] Specifically, it includes the following steps: Step 1: Constructing a distribution map of forged clues and identifying strongly correlated regions: First, a feature extraction network is designed to generate feature map clusters with several channels from the original image. For these feature map clusters, a feature densification and category mapping weight network is further designed. Through network weight calculation, a forgery clue distribution map is obtained from the feature map clusters. Furthermore, by upsampling the forgery clue distribution map to the size of the original input image and combining it with the correlation metric score of the forgery features in the distribution map, it is mapped to the original face image region to identify the strong correlation region of face forgery clues in the original image for the next step.

[0082] Step 2: Automatically generate formatted semantic descriptions for areas strongly related to forged clues: This application proposes a technique for automatically generating formatted semantic descriptions of regions strongly correlated with forgery clues. This technique enables the generation of textual descriptions of face forgery regions even in images where forgery clues cannot be observed, overcoming the limitation of related technologies that can only describe observable forgery regions with text. A formatted template for describing forgery regions and forgery clues is designed, using high-scoring regions in the forgery clue distribution map as the core. Combining their spatial location and correlation with the target category, various structured texts are automatically generated based on the predefined description template.

[0083] Step 3: Training a large-scale AIGC face deep anti-spoofing model based on a large-scale automated forgery description image set: Based on the aforementioned automated method, for AIGC face forgery images with visually imperceptible forgery clues that cannot be described in text by relevant technologies, a large-scale forgery description image-text set is automatically identified and generated. This is combined with a face image forgery description image-text set containing visually observable forgery clues to construct a one-to-many image-text pair of "image-multi-dimensional forgery description". For the input of this large-scale one-to-many image-text dataset, a large-scale face forgery detection model network is designed. This involves extracting visual features from the image through an image encoder, extracting global semantic features based on a text encoder, parsing information from the description using the text encoder, and achieving image-text feature fusion through a cross-modal attention mechanism. Through iterative generation of image-text pairs and iterative model training using large-scale data, the model gains the ability to distinguish between genuine and fake images and output interpretable conclusions.

[0084] Step 4: Design of a counterfeit detection reasoning model based on a generative semantic parsing model: Based on the aforementioned deep anti-counterfeiting pre-trained large model, visual language multimodal coding features containing forgery traces are extracted. A semantic parsing module is designed to convert the aforementioned multimodal coding features into structured text descriptions. That is, combined with the design of the inference classification layer network, the structured inference results and descriptions containing the true or false conclusions and details of the forged area are realized in the output, thus achieving interpretable inference from visual features to text anti-counterfeiting conclusions.

[0085] See Figure 6 , Figure 6 This is a flowchart of a face authentication method provided in an embodiment of this application, such as... Figure 6 As shown, the method includes the following steps: Step 601: Input the face image to be identified into the trained multimodal anti-spoofing model for inference and obtain the anti-spoofing result; The trained multimodal anti-spoofing model is a multimodal anti-spoofing model obtained based on the above model training method.

[0086] Specifically, the aforementioned face image to be identified can be a digital face image or video frame whose authenticity needs to be determined and whose source is unknown or yet to be confirmed.

[0087] In this embodiment, the multimodal anti-spoofing model obtained based on the above model training method performs face anti-spoofing on the face image to be identified, thereby improving the accuracy of anti-spoofing.

[0088] Optionally, the step of inputting the face image to be identified into the trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result includes: The face image to be identified is input into the trained multimodal anti-spoofing model for feature extraction, resulting in multimodal encoded features. The multimodal encoded features are converted into structured text descriptions to obtain semantic feature vectors; The authentication result is generated based on the multimodal coding features and the semantic feature vector. The authentication result includes at least one of the following: the authenticity conclusion corresponding to the face image to be identified and the text description of the forged region corresponding to the face image to be identified.

[0089] Specifically, the aforementioned multimodal coding features can be feature vectors or feature sequences formed by the multimodal anti-spoofing model after processing the input image.

[0090] The above conclusions on authenticity can be binary or multi-class classifications of the image's authenticity, such as "real" or "fake".

[0091] The text description of the aforementioned forged area can be a specific textual explanation of the problematic area in the image when the conclusion is that it is forged, such as "there is an unnatural shadow transition below the right eyelid".

[0092] In this implementation, by using multimodal coding features and semantic feature vectors together as the basis for generating the authentication results, the multimodal authentication model can integrate deep visual evidence and high-level semantic understanding, thereby improving the reliability of the authentication results.

[0093] For example, Figure 7 This is a schematic diagram of a multimodal counterfeit detection model inference provided in an embodiment of this application, such as... Figure 7 As shown, it includes the following steps: Step 1: After preprocessing the image to be inferred, input it into the semantic parsing module trained on the deep anti-counterfeiting model. The aforementioned multimodal coding features are converted into structured text descriptions, and the structured semantic feature vector is output.

[0094] Step 2: Combining image features with semantic parsing results, the inference layer of the semantic parsing model, such as a classification network, outputs whether the decision result is fake, and generates a text description of the fake region (if the decision result is fake).

[0095] See Figure 8 , Figure 8This is a schematic diagram of the structure of a model training device provided in an embodiment of this application, as shown below. Figure 8 As shown, the model training device 800 includes: Processing module 801 is used to perform feature recognition processing on face image samples using a feature extraction network to obtain a strong feature map of the face; The generation module 802 is used to generate a text-image pair training set based on the strong facial feature map and a preset text description template. The text-image pair training set includes multiple training data, and the training data includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. Training module 803 is used to train the multimodal anti-spoofing model using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

[0096] Optionally, the processing module 801 includes: The first extraction unit is used to extract features from the face image sample using a feature extraction network to obtain the feature map cluster of the face image sample; The processing unit is used to perform global average pooling and fully connected layer processing on the feature map cluster to obtain the weight vector of the forgery category; The summation unit is used to perform weighted summation of the feature maps of each channel in the feature map cluster based on the weight vector to obtain a forgery clue distribution map. A sampling unit is used to upsample the forged clue distribution map to the size of the face image sample to obtain a target distribution map; The mapping unit is used to map the target distribution map to the region with the highest correlation metric score in the face image sample based on the correlation metric score in the target distribution map, so as to obtain a strong feature mapping map of the face.

[0097] Optionally, the preset text description template includes multiple sub-templates, and the generation module 802 includes: The determining unit is used to determine the location information of the forged region based on the strong facial feature mapping map; The first generation unit is configured to generate multiple text descriptions based on the multiple sub-templates and the location information, wherein the multiple text descriptions correspond one-to-one with the multiple sub-templates; The combination unit is used to combine the plurality of text descriptions with the strong facial feature map one by one to obtain a text-image pair training set, wherein the text-image pair training set includes at least one pair of the strong facial feature map and the text description.

[0098] Optionally, the training module 803 includes: The training unit is used to train the multimodal anti-spoofing model based on the target loss function and the image-text pair training set, so as to obtain the trained multimodal anti-spoofing model. The target loss function is used to increase the Euclidean distance between fake samples and real samples in the image-text pair training set.

[0099] Optionally, the target training set further includes a preset dataset, which includes multiple image samples and text descriptions corresponding to each of the multiple image samples, and the text descriptions corresponding to each of the multiple image samples are obtained by manual annotation.

[0100] It should be noted that the model training apparatus provided in this application embodiment is an apparatus capable of executing the above-described model training method. Therefore, all implementation methods in the above-described model training method embodiments are applicable to this apparatus and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.

[0101] See Figure 9 , Figure 9 This is a schematic diagram of the structure of a face authentication device provided in an embodiment of this application, as shown below. Figure 9 As shown, the face authentication device 900 includes: The inference module 901 is used to input the face image to be identified into the trained multimodal anti-spoofing model for inference and obtain the anti-spoofing result. The trained multimodal anti-spoofing model is a multimodal anti-spoofing model obtained based on the model training method described in any one of claims 1-5.

[0102] Optionally, the inference module 901 includes: The second extraction unit is used to input the face image to be identified into the trained multimodal anti-spoofing model for feature extraction, and obtain multimodal encoded features. The conversion unit is used to convert the multimodal encoded features into a structured text description to obtain a semantic feature vector; The second generation unit is used to generate a counterfeit detection result based on the multimodal coding features and the semantic feature vector. The counterfeit detection result includes at least one of the authenticity conclusion corresponding to the face image to be identified and the text description of the counterfeit region corresponding to the face image to be identified.

[0103] It should be noted that the face authentication device provided in this application embodiment is a device capable of executing the above-described face authentication method. Therefore, all implementation methods in the above-described face authentication method embodiments are applicable to this device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.

[0104] For details, see Figure 10 As shown in the figure, this application embodiment also provides an electronic device, including a bus 1001, a transceiver 1002, an antenna 1003, a bus interface 1004, a processor 1005, and a memory 1006.

[0105] Processor 1005, used for: A feature extraction network is used to perform feature recognition processing on face image samples to obtain a strong feature map of the face. Based on the strong facial feature map and the preset text description template, a training set of image-text pairs is generated. The training set of image-text pairs includes multiple training data, which includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. The multimodal anti-spoofing model is trained using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

[0106] exist Figure 10 In this document, a bus architecture (represented by bus 1001) is used. Bus 1001 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 1005 and memory represented by memory 1006. Bus 1001 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 1004 provides an interface between bus 1001 and transceiver 1002. Transceiver 1002 may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 1005 is transmitted over a wireless medium via antenna 1003, which further receives data and transmits it to processor 1005.

[0107] Processor 1005 is responsible for managing bus 1001 and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 1006 can be used to store data used by processor 1005 during operation.

[0108] Optionally, the processor 1005 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).

[0109] Optionally, the processor 1005 is specifically used for: Feature extraction networks are used to extract features from face image samples to obtain feature map clusters of the face image samples; The feature map clusters are processed by global average pooling and fully connected layers to obtain the weight vector of the forgery category; Based on the weight vector, the feature maps of each channel in the feature map cluster are weighted and summed to obtain a forgery clue distribution map; The forged clue distribution map is upsampled to the size of the face image sample to obtain the target distribution map; Based on the correlation metric score in the target distribution map, the target distribution map is mapped to the region with the highest correlation metric score in the face image sample to obtain a strong face feature mapping map.

[0110] Optionally, the preset text description template includes multiple sub-templates, and the processor 1005 is specifically used for: The location information of the forged region is determined based on the strong facial feature mapping map. Multiple text descriptions are generated based on the multiple sub-templates and the location information, and the multiple text descriptions correspond one-to-one with the multiple sub-templates; The multiple text descriptions are combined one by one with the strong facial feature map to obtain a text-image pair training set, which includes at least one pair of the strong facial feature map and the text description.

[0111] Optionally, the processor 1005 is specifically used for: The multimodal anti-spoofing model is trained based on the objective loss function and the image-text pair training set to obtain the trained multimodal anti-spoofing model. The objective loss function is used to increase the Euclidean distance between fake samples and real samples in the image-text pair training set.

[0112] Optionally, the target training set further includes a preset dataset, which includes multiple image samples and text descriptions corresponding to each of the multiple image samples, and the text descriptions corresponding to each of the multiple image samples are obtained by manual annotation.

[0113] It should be noted that the electronic device provided in this application embodiment is a device capable of executing the above-described model training method. Therefore, all implementation methods in the above-described model training method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.

[0114] For details, see Figure 11 As shown in the figure, this application embodiment also provides an electronic device, including a bus 1101, a transceiver 1102, an antenna 1103, a bus interface 1104, a processor 1105, and a memory 1106.

[0115] Processor 1105, used for: The face image to be identified is input into the trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result; The trained multimodal anti-spoofing model is a multimodal anti-spoofing model obtained based on the above-mentioned model training method.

[0116] exist Figure 11 In this document, a bus architecture (represented by bus 1101) is used. Bus 1101 may include any number of interconnected buses and bridges, linking various circuits including one or more processors represented by processor 1105 and memory represented by memory 1106. Bus 1101 may also link various other circuits such as peripheral devices, voltage regulators, and power management circuits, which are well known in the art and therefore will not be described further herein. Bus interface 1104 provides an interface between bus 1101 and transceiver 1102. Transceiver 1102 may be a single element or multiple elements, such as multiple receivers and transmitters, providing a unit for communicating with various other devices over a transmission medium. Data processed by processor 1105 is transmitted over a wireless medium via antenna 1103, which further receives data and transmits it to processor 1105.

[0117] Processor 1105 is responsible for managing bus 1101 and general processing, and can also provide various functions, including timing, peripheral interface, voltage regulation, power management, and other control functions. Memory 1106 can be used to store data used by processor 1105 during operation.

[0118] Optionally, the processor 1105 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or a complex programmable logic device (CPLD).

[0119] Optionally, the processor 1105 is specifically used for: The face image to be identified is input into the trained multimodal anti-spoofing model for feature extraction, resulting in multimodal encoded features. The multimodal encoded features are converted into structured text descriptions to obtain semantic feature vectors; The authentication result is generated based on the multimodal coding features and the semantic feature vector. The authentication result includes at least one of the following: the authenticity conclusion corresponding to the face image to be identified and the text description of the forged region corresponding to the face image to be identified.

[0120] It should be noted that the electronic device provided in this application embodiment is a device capable of executing the above-described face spoofing detection method. Therefore, all implementation methods in the above-described face spoofing detection method embodiments are applicable to this electronic device and can achieve the same or similar beneficial effects. To avoid repetition, this embodiment will not elaborate further.

[0121] This invention also provides an electronic device, including: a processor, a memory, and a program stored in the memory and executable on the processor. When the program is executed by the processor, it implements the various processes of the above-described model training method or face authentication method embodiments and achieves the same technical effect. To avoid repetition, it will not be described again here.

[0122] This application also provides a computer-readable storage medium storing a computer program. When executed by a processor, this computer program implements the various processes of the above-described model training method or face authentication method embodiments, achieving the same technical effects. To avoid repetition, further details are omitted here. The computer-readable storage medium may include read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk.

[0123] This application also provides a computer program product, including computer instructions. When executed by a processor, the computer instructions implement the various processes of the above-described model training method or face authentication method, and achieve the same technical effect. To avoid repetition, they will not be described again here.

[0124] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Unless otherwise specified, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element.

[0125] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods of the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk) and includes several instructions to cause a terminal (which may be a mobile phone, computer, server, air conditioner, or network device, etc.) to execute the methods described in the various embodiments of this application.

[0126] The embodiments of this application have been described above with reference to the accompanying drawings. However, this application is not limited to the specific embodiments described above. The specific embodiments described above are merely illustrative and not restrictive. Those skilled in the art can make many other forms under the guidance of this application without departing from the spirit and scope of the claims, and all of these forms are within the protection scope of this application.

Claims

1. A model training method, characterized in that, include: A feature extraction network is used to perform feature recognition processing on face image samples to obtain a strong feature map of the face. Based on the strong facial feature map and the preset text description template, a training set of image-text pairs is generated. The training set of image-text pairs includes multiple training data, which includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. A multimodal anti-spoofing model is trained using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

2. The method according to claim 1, characterized in that, The process of using a feature extraction network to perform feature recognition processing on face image samples to obtain a strong face feature map includes: Feature extraction networks are used to extract features from face image samples to obtain feature map clusters of the face image samples; The feature map clusters are processed by global average pooling and fully connected layers to obtain the weight vector of the forgery category; Based on the weight vector, the feature maps of each channel in the feature map cluster are weighted and summed to obtain a forgery clue distribution map; The forged clue distribution map is upsampled to the size of the face image sample to obtain the target distribution map; Based on the correlation metric score in the target distribution map, the target distribution map is mapped to the region with the highest correlation metric score in the face image sample to obtain a strong face feature mapping map.

3. The method according to claim 1, characterized in that, The preset text description template includes multiple sub-templates. The generation of a text-image pair training set based on the strong facial feature map and the preset text description template includes: The location information of the forged region is determined based on the strong facial feature mapping map. Multiple text descriptions are generated based on the multiple sub-templates and the location information, and the multiple text descriptions correspond one-to-one with the multiple sub-templates; The multiple text descriptions are combined one by one with the strong facial feature map to obtain a text-image pair training set, which includes at least one pair of the strong facial feature map and the text description.

4. The method according to claim 1, characterized in that, The step of training the multimodal anti-spoofing model using the aforementioned image and text training set to obtain the trained multimodal anti-spoofing model includes: The multimodal anti-spoofing model is trained based on the objective loss function and the image-text pair training set to obtain the trained multimodal anti-spoofing model. The objective loss function is used to increase the Euclidean distance between fake samples and real samples in the image-text pair training set.

5. The method according to claim 1, characterized in that, The target training set also includes a preset dataset, which includes multiple image samples and text descriptions corresponding to each of the multiple image samples. The text descriptions corresponding to each of the multiple image samples are obtained by manual annotation.

6. A method for face authentication, characterized in that, include: The face image to be identified is input into the trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result; The trained multimodal anti-spoofing model is a multimodal anti-spoofing model obtained based on the model training method described in any one of claims 1-5.

7. The method according to claim 6, characterized in that, The process of inputting the face image to be identified into a trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result includes: The face image to be identified is input into the trained multimodal anti-spoofing model for feature extraction, resulting in multimodal encoded features. The multimodal encoded features are converted into structured text descriptions to obtain semantic feature vectors; The authentication result is generated based on the multimodal coding features and the semantic feature vector. The authentication result includes at least one of the following: the authenticity conclusion corresponding to the face image to be identified and the text description of the forged region corresponding to the face image to be identified.

8. A model training device, characterized in that, include: The processing module is used to perform feature recognition processing on face image samples using a feature extraction network to obtain a strong feature map of the face; The generation module is used to generate a training set of image-text pairs based on the strong facial feature map and a preset text description template. The training set of image-text pairs includes multiple training data, and the training data includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. The training module is used to train the multimodal anti-spoofing model using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

9. A face authentication device, characterized in that, include: The inference module is used to input the face image to be identified into the trained multimodal anti-spoofing model for inference and obtain the anti-spoofing result; The trained multimodal anti-spoofing model is a multimodal anti-spoofing model obtained based on the model training method described in any one of claims 1-5.

10. An electronic device, characterized in that, Includes a transceiver and a processor, the processor being used for: A feature extraction network is used to perform feature recognition processing on face image samples to obtain a strong feature map of the face. Based on the strong facial feature map and the preset text description template, a training set of image-text pairs is generated. The training set of image-text pairs includes multiple training data, which includes the strong facial feature map and the corresponding text description information. The text description information corresponding to different training data is different. A multimodal anti-spoofing model is trained using a target training set to obtain the trained multimodal anti-spoofing model. The target training set includes the image-text pair training set. The multimodal anti-spoofing model includes an image encoder, a text encoder, and an image-text decoder. The outputs of the image encoder and the text encoder are connected to the input of the image-text decoder. The image encoder is used to extract visual features of the image, the text encoder is used to extract semantic features of the text description, and the image-text decoder is used to fuse the visual features and the semantic features.

11. An electronic device, characterized in that, Includes a transceiver and a processor, the processor being used for: The face image to be identified is input into the trained multimodal anti-spoofing model for inference to obtain the anti-spoofing result; The trained multimodal anti-spoofing model is a multimodal anti-spoofing model obtained based on the model training method described in any one of claims 1-5.

12. An electronic device, characterized in that, include: A processor, a memory, and a program stored in the memory and executable on the processor, wherein the program or instructions, when executed by the processor, implement the steps of the model training method as described in any one of claims 1-5; or, the program, when executed by the processor, implements the steps of the face authentication method as described in claim 6 or 7.

13. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that, when executed by a processor, implements the steps of the model training method as described in any one of claims 1 to 5; or, when executed by a processor, the computer program implements the steps of the face authentication method as described in claim 6 or 7.

14. A computer program product, characterized in that, The method includes computer instructions that, when executed by a processor, implement the steps of the model training method as described in any one of claims 1 to 5; or, when executed by a processor, the computer instructions implement the steps of the face authentication method as described in claim 6 or 7.